Emergency vehicle priority and queuing optimization cooperative control method

By dynamically generating signal timing strategies using the LSTM-Transformer-PPO model, the problem of coordinating rapid passage of emergency vehicles with optimized queuing of social vehicles during peak traffic hours is solved, thereby improving the operational efficiency and emergency response capabilities of the urban road network.

CN121214701AActive Publication Date: 2025-12-26CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Application Number
CN202511759932.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2025-12-26
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

During peak traffic hours, emergency vehicles are prone to congestion in the city's core areas and main roads. Existing signal control strategies are unable to achieve rapid passage for emergency vehicles and optimized queuing for other vehicles, resulting in a disruption of the fragile balance of traffic flow, increasing the difficulty of emergency vehicle passage and negative traffic effects.

Method used

A dual-branch neural network model integrating LSTM and Transformer architectures is adopted, combined with a PPO agent, to predict the future lane queuing evolution trend and emergency vehicle trajectory status through multi-source traffic state data, dynamically generate signal timing strategies, and achieve coordinated control of emergency vehicle priority passage and social vehicle queuing optimization.

Benefits of technology

It significantly improves the operational efficiency of intersections under complex traffic loads, enhances the response efficiency of emergency vehicles, alleviates delays and queue accumulation of social vehicles, and achieves balanced traffic flow in multiple directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214701A_ABST
    Figure CN121214701A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent traffic, discloses an emergency vehicle priority and queuing optimization cooperative control method, and solves the problem of cooperation of emergency priority and social vehicle dispersion at a peak-saturated intersection. The method comprises the following steps: constructing a road network model containing a center intersection and four peripheral intersections, and collecting multi-source traffic data; a data set is established after preprocessing, and an LSTM-Transform double-branch model is trained to predict queuing and emergency tracks; constructing an intelligent agent based on PPO, and combining segmented reward training; and deploying an output signal decision. According to the invention, emergency low delay is guaranteed, social vehicle queuing is optimized, peak intersection efficiency is improved, and the method is suitable for high-saturation scenes; according to the method, the social vehicle congestion is effectively relieved while the rescue vehicles are guaranteed to pass preferentially, and the overall operation efficiency and the emergency response capability of the intersection in the complex traffic environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent transportation, in particular to an emergency vehicle priority and queuing optimization collaborative control method. BACKGROUND

[0002] During the peak traffic period, the core area and main roads of the city generally bear high-intensity traffic load. At this time, emergency vehicles performing emergency tasks are extremely prone to be trapped in congestion nodes and bottleneck sections, resulting in a significant increase in response time, and thus exacerbating the risk of life and property loss. Therefore, it is of urgent and practical necessity to implement signal priority right-of-way for emergency vehicles at saturated flow intersections to ensure their safe and efficient passage. However, under saturated flow conditions, the absolute priority passage behavior of emergency vehicles (such as forced green light extension or red light truncation based on phase preemption strategy) will significantly interfere with the existing signal timing scheme of the intersection, breaking the fragile balance of traffic flow. Especially when the queue lengths of each approach are close to the saturation critical point, such interference can easily induce a sharp increase in queues in a particular direction, leading to further deterioration of congestion. Therefore, how to collaboratively achieve the dual goals of ensuring the rapid and safe passage of emergency vehicles and simultaneously implementing timing optimization and queue relief for social vehicles has become the core challenge of peak period emergency priority control. The particularity of the intersection during the peak period lies in that the queues of each direction are generally under high pressure, and the traffic flow arrival rate fluctuates greatly. This not only increases the difficulty of passage for emergency vehicles, but also amplifies the negative effects of the priority strategy on traffic. Compared to traditional signal control research, which focuses on smooth transition of timing schemes in steady-state scenarios, the essence of the problem in the emergency priority scenario is how to complete the spatio-temporal collaborative two-stage operation within a very small time window. 1. Right-of-way priority control: Through signal priority control mechanism, absolute right-of-way is given to emergency vehicles to ensure their priority passage at the intersection.

[0003] 2. Queuing optimization control: After passing through the intersection, the emergency vehicle needs to implement dynamic optimization of signal timing in real time, actively regulate the queues of the disturbed phases through phase sequence reconstruction and adaptive adjustment of green light duration, and optimize the traffic efficiency of the intersection during the peak period.

[0004] Therefore, the signal control strategy in the emergency priority scenario should be able to minimize the delay of emergency vehicles and efficiently dissipate and spatially balance the queues of social vehicles in a saturated flow environment through multi-objective optimization mechanism. This requires the strategy not only to respond to the needs of emergency vehicles, but also to accurately predict and quickly intervene in the disturbance caused by the priority behavior to social traffic flow, ensuring rescue efficiency while maximizing the resilience level of the road network.

[0005] Currently, the research on the right-of-way priority control of emergency vehicles at home and abroad mainly forms three method systems: traditional transition strategy (including direct conversion, green light duration adjustment, stage conversion and smooth transition), optimization control model (signal control method based on multi-objective function) and intelligent adaptive algorithm (control strategy with dynamic optimization capability). Although the traditional strategy is practical in the conventional traffic scene, it faces the risk of dynamic response delay and queue out of control under the saturated flow condition in the peak period. Although the optimization control model can construct a multi-objective function, it is difficult to realize the coordinated control of complex intersections due to the algorithm efficiency bottleneck caused by the diversity of targets in the saturated state. Although the intelligent adaptive algorithm is committed to the coordinated optimization of emergency priority and social vehicle dredging, it still has defects such as high computational complexity and insufficient real-time performance. The existing research generally lacks the following three aspects of comprehensive consideration: insufficient adaptability in saturated scenarios, lack of multi-agent collaborative optimization, and limited proactive control capability. SUMMARY

[0006] The purpose of the embodiment of the application is to provide an emergency vehicle priority and queue optimization collaborative control method, which realizes the multi-objective collaborative optimization of minimizing the delay of social vehicles, equalizing the queue and preventing overflow while ensuring the low delay of emergency vehicles. The response efficiency of emergency vehicles and the import lane queue efficiency in the peak period are improved.

[0007] To solve the above technical problems, the technical scheme adopted by the application is an emergency vehicle priority and queue optimization collaborative control method, which is performed according to the following steps: S1, a road intersection road network topology structure model is constructed and multi-source traffic state data is collected; S2, the collected data is preprocessed and multi-source heterogeneous feature extraction is performed, and a standardized feature vector is constructed; S3, a time series data set is constructed based on the preprocessed data; S4, a double-branch traffic prediction model integrating LSTM and Transformer architecture is established: S5, the double-branch traffic prediction model is trained; a multi-task mean square error loss function including queue evolution prediction loss and emergency vehicle running state prediction loss is used to optimize the parameters of the double-branch traffic prediction model until convergence; S6, a PPO agent is constructed, and the state space of the PPO agent integrates real-time traffic state, historical observation data and the prediction results of the double-branch traffic prediction model; S7, a segmented reward function is designed to guide the PPO agent to perform multi-objective collaborative optimization, and the dynamic weight coefficient is used to balance the emergency vehicle priority and social vehicle optimization targets; S8, training the PPO agent; the PPO agent interacts with the traffic simulation environment to sample trajectory data, determines the state-action excess return through the advantage function estimation, and updates the parameters of the policy network and the value network in combination with the clipping optimization target; S9, deploying the trained double-branch traffic prediction model, inputting real-time collected traffic state data, obtaining future lane-by-lane queuing evolution trend and emergency vehicle trajectory prediction results; the PPO agent outputs signal phase selection and green light duration decision based on real-time traffic state data and the prediction results, controls the traffic signal to execute the corresponding timing scheme, and the signal phase selection only switches between the four green light main phases, realizing the priority of emergency vehicles and the optimization of social vehicle queuing.

[0008] Further, the center intersection adopts a four-stage eight-phase timing scheme, which includes two straight left phases in the east-west direction and two straight and left turn phases in the south-north direction; the lengths of the four entrance lanes of the center intersection are 800m, 450m, 386m and 400m respectively, and each entrance lane selects one straight lane and one left turn lane as the research and control object; the traffic simulation environment uses the SUMO simulation platform, and the simulation duration is 1 hour. The multi-source traffic state data output by the simulation is saved as a CSV file.

[0009] Further, the S2 is specifically to construct the queue / flow ratio and the synthetic feature of net growth queue through feature engineering, remove redundant information through correlation analysis, mutual information evaluation, VIF multicollinearity detection and feature importance screening, and then normalize the features to obtain a standardized feature vector; The standardized feature vector is : (1) Wherein, is a very small constant, is the mean of each dimension feature in the training set, is the standard deviation of each dimension feature in the training set, is the feature vector before standardization.

[0010] Further, in S3, the specific process of constructing the time series data set is: Let the original time series data be: (2) Wherein, represents the input all-lane feature vector at time t, represents the output queue length at time t, represents the time series length. Establish the input sequence for each time step and predict target sequence : (3) (4) The final dataset for: (5) in, Indicates the length of the input time series. Indicates the length of the output time series. Furthermore, the dataset is divided into a training set and a test set in a 7:3 ratio.

[0011] Furthermore, in S4, the dual-branch traffic prediction model includes an LSTM encoder, a Transformer encoder, and a dual-branch prediction head. The LSTM encoder captures the local temporal dynamics of traffic flow, the Transformer encoder captures the global spatiotemporal dependence, and the dual-branch prediction head outputs the prediction results of the future time period lane queuing evolution trend and emergency vehicle trajectory status, respectively. The LSTM encoder has 112 hidden layers and a total of 3 layers. The Transformer encoder has 3 layers, 4 attention heads, a feedforward layer size of 128, a dropout rate of 0.1, an output window of 5, and a state prediction output dimension of 2 for each emergency vehicle.

[0012] Furthermore, the multi-task mean squared error loss function in S5 Specifically: (6) in, These are uncertainties between tasks, used to represent the difficulty of each task. This is a regularization term for the uncertainty variable, used to constrain the model's estimation of uncertainty. Predicting loss for queue evolution: (7) in, and They represent the first The predicted value output at the predicted time, the first prediction time. Output the true value at each predicted time point; Predicted losses based on emergency vehicle operational status: Assuming each emergency vehicle It contains 4 predictor variables, totaling For vehicles, the emergency prediction dimension is: , then: (8) where, and denote the predicted value output of the th emergency vehicle, the true value output of the th emergency vehicle, respectively, denotes the emergency prediction dimension.

[0013] Further, the final state of the PPO agent in S6 is denoted as: (9) where, is the normalized original input state at time , containing the queue length, vehicle waiting time, lane flow, signal state one-hot encoding, intersection throughput, simulation time, and lane, position, speed of the emergency vehicle of the main road; is an additional flag to assist the strategy in determining when to make decisions based on prediction information; is a time sequence correlation encoding vector that abstracts the future steps; is a vector concatenation operation.

[0014] Further, the segmented reward function in S7 includes the combination of emergency vehicle reward , social vehicle reward , and phase switching penalty : (10) where, is the adaptive weight coefficient between the emergency priority and the social vehicle queue optimization two types of returns at time , and , specifically: (11) where denotes the distance of the emergency vehicle to the intersection, exp is the natural exponential function, is the maximum delay constant, is a parameter to control the steepness of the curve.

[0015] Further, the emergency vehicle reward is specifically: (12) where, and denote the distance of the emergency vehicle to the intersection at time The signal for the lane is green or yellow; a value of 1 indicates the corresponding light color, otherwise a value of 0. , They are respectively and Component weights; Indicates the speed increment. for The weight, The increment of the potential function, for The weights; The social vehicle reward Specifically: (13) (14) (15) in, This represents the increase in outflow volume for this step; These are the component weights for maximum queueing potential difference, normalized delay, queueing balance potential difference, and normalized number of departing vehicles, respectively. To shape the overall strength of society; The maximum observable outflow rate; Average lane waiting time; The maximum delay constant; The difference of the maximum queuing potential function; The difference of the potential function for queuing equilibrium; Core rewards for social vehicles; Shaping rewards for the potential function of social vehicles; The phase switching penalty Specifically: (16) in, This indicates the duration of the green light. The effective distance threshold indicating emergency priority Indicates the tolerance threshold. Indicates the phase switching penalty weight. This indicates the upper limit of tolerance for the effective distance threshold for emergency priority.

[0016] Furthermore, the process of determining the dominance function in S8 is specifically as follows: (17) In the state Take action below The immediate reward received afterward; is a discount factor, used to attenuate the influence of future values; is the estimate of the value network for the next state ; is the estimate of the value network for the current state ; The total optimization objective of the PPO agent is : (18) where, represents an increased policy entropy term to encourage exploration, and represent the value loss coefficient and the entropy coefficient, respectively; is the policy clipping objective; is the value function loss.

[0017] Compared with the prior art, the beneficial effects of the present application include the following aspects: the present application constructs a multi-dimensional traffic state perception system covering the inlets of the center intersection and its four directly associated peripheral intersections, integrates key features such as queue length, traffic volume, signal timing state and emergency vehicle operation information, and realizes fine modeling and real-time analysis of the traffic flow evolution process at the lane level. On this basis, the PPO (Proximal Policy Optimization) multi-objective reinforcement learning algorithm is used to dynamically generate differentiated signal timing strategies according to the predicted traffic state parameters, realize dynamic priority allocation for emergency vehicles and adaptive optimization of phase timing for social vehicles, and significantly improve the operating efficiency and response capability of the intersection under complex traffic load.

[0018] The present application designs a double-branch neural network model integrating LSTM and Transformer architecture. The model receives standardized multi-feature time series data at the input layer, uses the LSTM module to capture the basic time sequence dynamic characteristics of traffic flow, simultaneously introduces the self-attention mechanism of Transformer to strengthen the extraction ability of key spatiotemporal features, and cooperates the advantages of the two through a customized gating fusion mechanism. The model adopts a double-branch structure to output the prediction results of lane queue evolution trend and emergency vehicle state parameters respectively, meeting the differentiated needs of heterogeneous traffic subjects in prediction accuracy and timeliness.

[0019] The present application constructs a segmented reward function mechanism, decouples the signal control process into two cooperative optimization stages of emergency vehicle priority passing (TSC-Priority) and social vehicle queuing and dredging (TSC-Recovery), and realizes multi-objective joint optimization by means of the time sequence decision-making ability of the PPO algorithm. The design effectively avoids the 'dimension disaster' problem caused by the high dimension of the state space in the traditional method. The present application can significantly improve the passing efficiency of rescue vehicles, relieve the delay and queue accumulation of social vehicles, realize the balanced dredging of multi-direction traffic flow and the improvement of overall traffic capacity in the high saturation flow scenario of urban road network, especially in the peak period and emergency response situation, and has wide application prospect in intelligent traffic control of complex urban road network. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, below will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is the overall flow chart of the method of the present application; Figure 2 is the SUMO simulation road network topology map; Figure 3 is the key feature correlation heat map; Figure 4 is the center controlled intersection lane-level queue evolution prediction result map; (a) is the E00_0 lane queue prediction value and true value comparison (MAE=0.43, RMSE=0.73); (b) is the E00_1 lane queue prediction value and true value comparison (MAE=1.01, RMSE=1.69); (c) is the E04_0 lane queue prediction value and true value comparison (MAE=0.33, RMSE=0.62); (d) is the E04_1 lane queue prediction value and true value comparison (MAE=0.97, RMSE=1.74); (e) is the E05_0 lane queue prediction value and true value comparison (MAE=0.21, RMSE=0.42); (f) is the E05_1 lane queue prediction value and true value comparison (MAE=0.47, RMSE=0.83); (g) is the E06_0 lane queue prediction value and true value comparison (MAE=0.29, RMSE=0.55); (h) is the E06_1 lane queue prediction value and true value comparison (MAE=0.84, RMSE=1.55); Figure 5 is the center controlled intersection lane-level queue evolution prediction convergence curve diagram; Figure 6 is an emergency vehicle trajectory state prediction convergence curve; Figure 7 is an emergency vehicle trajectory state parameter prediction error scatter plot; wherein (a) is an emergency vehicle position prediction scatter plot (ty2.0, original meters); (b) is an emergency vehicle speed prediction scatter plot (ty2.0, original meters / second); (c) is an emergency vehicle position prediction scatter plot (ty1.0, original meters); (d) is an emergency vehicle speed prediction scatter plot (ty1.0, original meters / second); Figure 8 is a central controlled intersection lane level queue equilibrium heat map; Figure 9 is a PPO agent entropy value change curve. DETAILED DESCRIPTION

[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0023] The embodiment provides an emergency vehicle priority and queue optimization collaborative control method, specifically an intersection emergency vehicle priority and queue optimization collaborative control method based on fusion LSTM-Transformer-PPO. The embodiment faces the "priority-dredging" integrated signal control framework of the saturated intersection in the peak period. Through the fusion LSTM-Transformer algorithm model architecture, multi-feature input is completed to realize queue evolution prediction and emergency vehicle running state prediction. Based on the PPO algorithm, a signal control strategy is dynamically generated. The diversity of the optimization target and the subject is fully considered. Based on the traffic and queue characteristics of the city peak period, the road right priority of the emergency vehicle and the queue optimization process of the social vehicle are separated and connected. Through the method of deep reinforcement learning, queue evolution prediction is implemented for each import lane. At the same time, the position and speed of the emergency vehicle and other trajectory information are predicted. According to the prediction result, the signal light is controlled in stages. After the smooth passage of the emergency vehicle, the queue optimization of the social vehicle is quickly connected.

[0024] In some specific embodiments, the specific process of the emergency vehicle priority and queue optimization collaborative control method is as shown in Figure 1 S1, a road intersection road network topology structure model is constructed and multi-source traffic state data is collected. The function of the research scene and data collection module specifically includes road network information acquisition, flow collection and preprocessing, simulation running, and export of traffic state information. ​

[0025] S101, build a road intersection road network topology. In this embodiment, the research scenario is the intersection area under the saturated traffic of urban road network; including determining the number, type, lane number, speed interval, signal timing scheme of the intersection, and the traffic flow of each intersection, etc., and the road network topology is as shown in Figure 2 , Figure 3 .

[0026] The road intersection road network topology model is the associated area of five intersections (including the center intersection and the four peripheral intersections directly connected thereto), the control intersection is a four-stage eight-phase timing scheme, including two straight left phases in the east-west direction, two straight and left turn phases in the south-north direction. The lengths of the four entrance lanes of the center intersection are 800m, 450m, 386m and 400m respectively. Each direction is a city trunk road, and there are multiple straight and left turn lanes. There are non-motor vehicle lanes, right turn exclusive lanes and pedestrian crosswalks, which are typical urban trunk road signal intersections, and there is a large amount of traffic during peak hours, and the speed limit of the trunk road is 50-60km / h. According to the previous data investigation, a typical straight and left turn lane is selected for each entrance lane as the main research and control object during the specific implementation of this embodiment. A number of traffic paths are set, respectively from one peripheral intersection to another peripheral intersection, and pass through the center intersection, and the traffic demand of each path represents the saturated traffic characteristics. The road network and traffic flow parameter files established in SUMO are saved for subsequent reading and use of LSTM-Transformer traffic prediction model and PPO control optimization model. The road network topology is as shown in Figure 2 .

[0027] S102, obtain traffic state data. Use the road network and traffic data to perform 1 hour of simulation test, obtain real-time traffic flow data, and export the traffic state information (including queue length, traffic volume, phase state, phase remaining time, vehicle headway (average, median, 85% quantile), vehicle headway at stop line, emergency vehicle state information (speed, position, phase information, vehicle headway) of each intersection entrance lane required for model training from SUMO, save as a CSV file for subsequent data preprocessing.

[0028] S103, data preprocessing.

[0029] In this embodiment, the preprocessing of data is to obtain the original data (mainly for the required feature engineering for queue prediction) of the queue length, traffic flow, signal light related information, and the corresponding headway and headway time of each approach of the five intersections, and other feature information required for queue prediction. Through multi-dimensional feature analysis, the system evaluates the feature value and feature coherence, reduces redundant information, and improves prediction efficiency. Finally, the correlation index between different features is output.

[0030] In some specific embodiments, the data processing process includes the original traffic state data related to queue prediction obtained from SUMO in the input S102 step, covering the queue length, traffic flow, and signal light related information of the five intersections, as well as the headway and headway time of each approach. The pre-processing side then conducts feature engineering driven by domain knowledge, constructs synthetic features such as queue / flow ratio and net growth queue to enhance prediction sensitivity and interpretability. On this basis, through correlation analysis, mutual information evaluation, and VIF multicollinearity detection, combined with random forest importance, SHAP value interpretation, and stability selection, the system evaluates feature value and correlation, reduces data redundancy, and improves modeling efficiency; according to the principles of information theory and statistical significance, key prediction factors are identified lane by lane; for example, on the target prediction lane E04_0, the focus is on its own lane queue state (E04_0_queue) and upstream traffic flow (E04_0_up_E08_1_queue). The output side forms an index matrix and scoring results (including mutual information score, VIF report, importance and stability ranking, and SHAP contribution overview) about the correlation between different features, and determines the key feature set and explanation framework for each lane, providing data-driven basis for intelligent transportation system optimization, and saving as a new CSV file for subsequent LSTM-Transformer traffic prediction model reading. See Figure 4 (a)-(h), the model performs well in queue length prediction, with ideal prediction effect. The prediction error in most scenarios does not exceed one vehicle, and the error in the remaining cases is basically controlled within the range of 1 to 2 vehicles, reflecting the high precision and strong practicality of the model.

[0031] S2, system initialization and configuration module of the dual-branch traffic prediction model combining LSTM and Transformer architecture. First, the prediction model in this embodiment is composed of two algorithm parts (LSTM and Transformer) and adopts a single encoder + dual-branch decoding time series prediction framework, which is composed of an LSTM encoder (local time series modeling), a Transformer encoder (global dependence modeling) and a dual-branch prediction head (queue prediction and emergency vehicle trajectory state prediction). The specific architecture principle and prediction design of the prediction model will be described in detail in the following S5 step. In this step, the data loading will be completed by reading the previously created feature engineering CSV file, road network configuration, vehicle flow definition and model saving file and parameter configuration.

[0032] S201, read configuration file and parameters. Load traffic flow data and road network topology information (including the file after data preprocessing of the CSV file saved by the configuration file S102, and various road network, vehicle flow and other files and parameters generated when the road network topology structure is established S101), define the mapping relationship between intersections and the mapping relationship between each entrance and the signal lamp. Set the output window size of the LSTM-Transformer traffic prediction model to 5, the input window size to 30, the batch size to 32 and the training round each time (50 rounds for the first training); S202, emergency vehicle configuration. Define the type information and list data of emergency vehicles, and design the mapping and driving route of emergency vehicles on the lane.

[0033] S203, save the configuration of the LSTM-Transformer traffic prediction model. Save the specific configuration information of the model as a JSON file, including the size of the input window and the output window, the output dimension of the emergency vehicle data, the center lane and the emergency vehicle list and other information.

[0034] S3, feature engineering processing module. This module is to preprocess the input intersection and emergency vehicle state information, and normalize and standardize the multi-feature data; since the feature engineering is multi-source heterogeneous data, this module will perform multi-source heterogeneous feature extraction and state representation. In this embodiment, the prediction problem of the traffic intersection is regarded as a multi-source heterogeneous time series modeling task: on the one hand, the main queue data (the queue length, flow, phase remaining time of each entrance lane, etc.) and the queue and signal state of the upstream adjacent intersection are introduced, and on the other hand, the full-process trajectory information (lane ID, signal state, position, speed, remaining phase, etc.) of the emergency vehicle (EV) is integrated. These data sources contain not only continuous numerical features of different scales and distributions at the same time step, but also high-dimensional category and state features obtained through one-hot conversion, which are then spliced into fixed-length time series input through a sliding window. After the above multi-modal input is normalized and standardized, it is sent to a dual-branch LSTM-Transformer framework to capture the spatial correlation and time dependence of time series information, so as to realize the joint prediction of the behaviors of the regular queue and the emergency vehicle in a complex traffic scene.

[0035] S301, first, initialize the feature container, and define the global feature matrix as an empty list.

[0036] S302, construct a feature vector; traverse all lanes and time steps, and for any lane , time step , construct a set of input feature vectors :

[0037] Among them: : the queue length of the current lane; : one-hot encoding of the signal state of the current lane (red, green, yellow); : the signal phase remaining time of the current lane; : the road section flow of the current lane; : the queue length of the corresponding lane of the upstream associated intersection; : one-hot encoding of the signal state of the corresponding lane of the upstream associated intersection; : the signal phase remaining time of the corresponding lane of the upstream associated intersection; In addition, in the above feature vector is the emergency vehicle trajectory feature, and for each emergency vehicle , the following is extracted:

[0038] wherein: : the current lane of the emergency vehicle; : the position of the emergency vehicle in the current lane; : the instantaneous speed; : the signal light status of the current lane.

[0039] S303, standardization processing; set the total time step as , the feature vector input at each time is , and the data matrix composed of all sample matrices is:

[0040]

[0041] wherein, is a lane set, represents the total number of lanes.

[0042] All input features need to be normalized before entering the LSTM-Transformer traffic prediction model to avoid the adverse effects of numerical scale differences on model training, especially the scale-sensitive parameter optimization in the LSTM and attention module. To accelerate convergence and improve training stability, the present embodiment applies the following standardization formula to all numerical features: let the mean and standard deviation of each dimension feature in the training set be , represents a d-dimensional real vector space, then the normalized feature vector is:

[0043] wherein, is a very small constant to prevent division by zero, is the time for the th lane dimensional input feature vector.

[0044] Therefore, for each dimension , the specific mean and standard deviation are calculated as:

[0045] N represents the Nth time step; represents the th observation (or time step) in the The original numerical value in the feature dimension.

[0046] The normalized sample is obtained as:

[0047] Here is the numerical stability term, which is usually set to .

[0048] represents the d-th component of the input feature vector of the l-th lane at time t, and in the embodiment, the corresponding normalized lane queue length in the feature arrangement is .

[0049] S4, a data set construction module is established. This module defines and initializes the data set of the lane queue, completes the basic parameter setting, loads the queue and storage feature data, calculates each function, converts data, and forms the final model training input multi-feature data.

[0050] S401, after completing the multi-source feature extraction and standardization processing, it is necessary to organize the continuous time series into supervised learning samples acceptable to the LSTM-Transformer traffic prediction model, which specifically includes an input sequence and a corresponding prediction target. Let the original time series data be:

[0051] : represents all lane feature vectors at time t : represents the output queue length (prediction target) at the corresponding time To construct a time series prediction task, a sliding window method is used to generate training samples. For each time step, the input sequence and prediction target sequence of the sample are defined as:

[0052]

[0053] The final obtained data set is:

[0054] When constructing data, the start point of the time window is traversed , and each group is added to the training set as a sample. Wherein represents the time series length, represents the input time series length, represents the output time series length, The "maximum value of start index minus 1" when representing the sliding window pattern sample, equivalently, it embodies the remaining time steps that the window can still slide. Therefore, Indicates the number of samples that can be formed.

[0055] S402, tensor conversion. For batch training of input deep learning model, all input features and output targets need to be converted to floating point type tensor representation:

[0056]

[0057] wherein, Indicates the matrix space of the input time window sample, Indicates the vector space of the output (prediction) time window, and the floating point number is 32.

[0058] The final formed training sample pair is:

[0059] wherein, Indicates the number of data samples processed at a time (batch size), Indicates the Cartesian product space of the small batch training sample.

[0060] In addition, in order to adapt to the batch training process of the model, all samples will be further organized into mini-batches of fixed size, and loaded using the DataLoader mechanism. The division of training and test sets is usually 7:3 Dataloader split to ensure the generalization ability evaluation of the model on unseen data.

[0061] S5, a dual-branch traffic prediction model architecture integrating LSTM and Transformer architecture is established; the data set is used for training to predict the future period lane queue evolution trend and emergency vehicle trajectory state. The LSTM_Transformer model is defined, the model is initialized, the encoding is set, and the dual output branch is constructed.

[0062] S501, initialization function. Set the input dimension of the LSTM_Transformer prediction model to d, the LSTM encoder hidden layer to 112, the total number of layers to 3, the Transformer encoder layer to 3, the attention head number to 4, the feedforward layer size to 128, the dropout rate to 0.1, the output window to 5, and the state prediction output dimension of each emergency vehicle to 2.

[0063] S502, establish an LSTM model encoding module; The LSTM is used to model local temporal dependencies, and the LSTM gating unit at time step t includes input gate, forget gate, output gate, candidate state, update state and hidden state output, respectively, and the specific calculation content is as follows: Input gate:

[0064] Forget gate:

[0065] Output gate:

[0066] Candidate state:

[0067] Update state:

[0068] Output hidden state:

[0069] Wherein: : input feature vector at time : hidden state (output) at time : previous time hidden state (previous time output); : previous time cell state; : : : input gate, forget gate, output gate; : candidate cell state; : cell state at time : weight matrix and bias of input gate; : weight matrix and bias of candidate state; : weight matrix and bias of output gate; : weight matrix and bias of forget gate; : sigmoid activation function; ​​​​​ tanh: hyperbolic tangent (tanh) activation function; : element-wise multiplication.

[0070] Output sequence As the next stage input. According to the parameter setting of the last step initialization function, set the encoder parameters (input dimension, hidden layer and dropout rate, etc.). In addition, set its position encoding module, the relevant setting parameters are consistent with the encoder.

[0071] S503, establish the Transformer encoder module. The Transformer encoder acts on global dependence modeling, that is, for capturing the global space-time dependence relationship of traffic flow, the input design (pasted matrix) is:

[0072] Wherein, is a matrix set composed of all real elements with a shape of rows, columns. Then, the LSTM output (hidden state) of the time step and the position encoding of the time step are added element by element to obtain the representation with position information , which is used in the following Transformer self-attention module:

[0073] The self-attention mechanism of the Transformer model is the key to global time sequence rule learning, and the dimension of feature engineering determines the number of attention:

[0074] The attention setting of the embodiment is:

[0075] is the transpose of the matrix Then the multi-head attention expression is:

[0076] The expression of its residual connection and feedforward network is designed as:

[0077] Wherein: : the output or input matrix of the last layer, and , the number of rows is the length of the time series, and the dimension (The original multi-lane and emergency lane features are linearly embedded or LSTM encoded in this embodiment and then mapped to this dimension); : Projection matrix, projecting to Query, Key, Value space; : Scaling factor; : The attention output of the th head; : The weight of linear projection after multi-head concatenation; : Respectively represent the first layer weight and bias Similarly); : Activation function.

[0078] The final model will output as the final prediction output of the Transformer model (take the last time step). In addition, the encoder settings and the initialized model remain the same.

[0079] S504, double-branch prediction structure. The ZT output by the Transformer module is connected to two fully connected layers: respectively, the queue length prediction output:

[0080] Emergency vehicle state prediction output:

[0081] Where: respectively represent the output weight and bias of the queue prediction ).

[0082] S505, loss function and training target: the double-branch prediction structure of the LSTM-Transformer traffic prediction model completes the multi-layer prediction task by reading the multi-feature engineering, so the multi-task mean squared error loss function is used for joint training: Task 1: Queue evolution prediction loss, using mean squared error (MSE) to measure the difference between the predicted queue length and the true value:

[0083] Where, and respectively represent the predicted value output at the th prediction moment and the true value output at the th prediction moment.

[0084] Task 2: Emergency vehicle running state prediction loss, assuming each emergency vehicle contains 4 prediction variables (see Equation 2), totaling vehicles, then the emergency prediction dimension is The emergency vehicle prediction loss is defined using the MSE form:

[0085] where and represent the predicted value output of the th emergency vehicle and the true value output of the th emergency vehicle, respectively, represents the emergency prediction dimension.

[0086] Finally, the total loss function is obtained by weighting the multi-task:

[0087] where is the weight parameter representing the importance preference for the queuing length and emergency prediction task, is a set of positive real numbers. To enhance the stability and adaptability of the model, introduce the inter-task uncertainty variable representing the difficulty of each task, is the uncertainty regularization term used to constrain the estimate of uncertainty; the optimization objective becomes:

[0088] This loss form is derived from multi-task Gaussian uncertainty modeling, which helps to automatically balance the weights.

[0089] S6, evaluation index setting. To comprehensively measure the performance of the model in the traffic prediction task, the following commonly used indicators are used in the present embodiment: training loss (Huber Loss):

[0090] Mean Absolute Error (MAE):

[0091] Mean Absolute Percentage Error (MAPE):

[0092] S7, LSTM-Transformer traffic prediction model training. The training main process of the LSTM-Transformer traffic prediction model is divided into two stages of training and testing: in the training stage, the model calculates the prediction output by forward propagation, evaluates the prediction error of the queue length and emergency vehicle state by using the loss function, updates the parameters by back propagation and gradient clipping, and saves the model and output training loss regularly; in the testing stage, the model fixes the parameters for inference, collects and returns the comparison data of the queue length prediction, emergency vehicle state prediction and true value, which is used for performance evaluation. The whole process adopts modular design, supports independent training and verification of multiple lanes, and finally outputs can be used to analyze the performance of the model in the traffic prediction task.

[0093] S8, building PPO agent. The embodiment is based on the proximal policy optimization PPO (Proximal Policy Optimization) algorithm to establish an emergency priority and queue optimization system, which realizes the double target optimization of emergency vehicles and social vehicles by controlling the phase selection and phase length of traffic signal lights. The system is simulated in the SUMO environment, the state is composed of multi-source features, the action is the signal phase selection and duration, and the reward function combines the EV priority potential function and the SV traffic efficiency. The embodiment realizes a PPO multi-objective optimization algorithm for traffic signal control, which combines the SUMO simulation environment, the traffic prediction model (LSTM-Transformer) and the multi-objective reward function to execute the priority strategy and optimize the signal light control.

[0094] S801, multi-source heterogeneous feature extraction and state representation. In the embodiment, the state space input of the PPO agent is composed of three time series information streams: first, the multi-source static features, which fuse the real-time queue, flow and signal state of the main road and the upstream lane, as well as the position, speed and other trajectory information of each emergency vehicle; second, the historical data, which extracts the original sequence and statistics (maximum, minimum, mean, variance) of the queue, delay and signal sequence through short-time (5 steps) and long-time (30 steps) sliding windows to capture the evolution trend of traffic flow; third, multi-source heterogeneous prediction information, which is first output by the Transformer to generate the queue and emergency vehicle trajectory at the future step, and then encoded into a fixed vector by a small LSTM inside the PPO and combined with the effectiveness flag, so that the strategy can review the past, be based on the present, and also be able to understand the traffic situation that will happen soon, so as to generate a more comprehensive, robust and safe and efficient signal control strategy at each decision-making moment.

[0095] S8011, main road and upstream vehicle features. For each main road lane At time t, a six-dimensional vector is extracted (35) in: : Normalized time Lane queue length is used to measure the maximum queue length and balance of a lane; Vehicle waiting time, used to calculate features such as average delay; : for time Lane flow; : One-hot encoding for red, yellow, and green light states; : Control the throughput of the intersection per unit time; Simulation time indicates periodicity. Dimensions are The real vector space. Additionally, all different superscripts below... Both represent vector spaces of different dimensions (determined by the superscript).

[0096] S8012, Characteristics of Emergency Vehicles. Regarding the... Emergency vehicles, if they are If it occurs at a certain time, then: (36) in: The lane where the emergency vehicle is currently located; The location of the lane; Instantaneous velocity; The one-hot code of the lane where the emergency vehicle is located; Dimensions are The real vector space.

[0097] S8013, Global State and Normalization. Concatenate all the above static features to obtain the original global state: (37) in, Each of them and Representing time respectively For each lane The constructed single-lane feature vector and the single-vehicle input feature vector of each emergency vehicle, The total number of lanes. This represents the total feature dimension for all predicted lanes and emergency vehicles. For dimension The real vector space.

[0098] For each dimension According to the training set mean Standard deviation Standardize: (38) and Representing time respectively No. The normalized and unnormalized values ​​of each input dimension are given, and the normalized vector is denoted as . , It is a numerically stable term.

[0099] S8014. Historical observation data, which includes short-window historical sequences and long-window sliding statistics. To reflect traffic evolution trends, both short-window and long-window histories are maintained in the environment. Where the short window length is... In this embodiment, =5, as detailed below:

[0100] in, These represent the queue, delay, and corresponding one-hot encoded concatenated information of the signal at the past five steps. Additionally, That is, at time The set of queue lengths for all lanes. and Similarly, represent time respectively. The average delay set for all lanes and the one-hot set of traffic light coding information.

[0101] Let the length of the long window be This implementation method takes =30 Calculate the sliding statistic:

[0102] in These represent the team leader and the delay time, respectively. In Record the maximum, minimum, average, and standard deviation of the normalized queue length and vehicle delay for each lane in the last 30 steps. Vectorize the features of the short and long windows, flatten them, and record them as follows: .

[0103] S8015, Multi-source heterogeneous prediction information. The LSTM-Transformer traffic prediction model is called to obtain future... (Predicted sequence with a value of 5 steps):

[0104] in, Represents a set of numbers whose elements are real numbers and have lines and A set of columns of matrices, Representing the future Predictive information on pedestrian lane queuing and emergency vehicle trajectories (location, speed), This is the total dimension for predicting all lane queues and the trajectory information such as the position and speed of emergency vehicles, and To avoid the PPO algorithm failing to truly understand the temporal information contained in the prediction results due to direct state concatenation, this embodiment designs a small LSTM network within the PPO for temporal encoding. The LSTM encoder uses parameters... For time input sequence The result of mapping 3D embedding vector:

[0105] in," This indicates a function notation that takes parameters as conditions, and is not a vector concatenation.

[0106] At the same time, valid flag information is introduced:

[0107] S8016, Final State Representation. The normalized original state, prediction validity, and LSTM encoding are concatenated to form the complete state for input to the PPO policy network:

[0108] in, For vector concatenation operations, the above... , , The three elements are concatenated according to their feature dimensions to form a complete state vector, which constitutes the single-step input of the PPO agent for action decision-making. Additional signs will be added based on the characteristics of the main road and emergency situations. The assisting strategy determines when to make decisions based on predictive information, encoding vectors. Extracting the future The temporal correlation of steps avoids the semantic sparsity caused by directly splicing high-dimensional prediction sequences. The total dimension of the policy network input vector is represented by the time step. Normalized original input state , 1 validity flag and prediction embedding Three parts are spliced together: S802, reinforcement learning environment and action conversion.

[0109] S8021, action definition. This embodiment contains two types of action selection in control, i.e. phase selection and time length control, which makes real-time action selection and control of the controlled signal light main phase and green light time length according to the real-time changes of the traffic flow and queue of each lane. The specific phase selection and time length control are as follows:

[0110]

[0111] Among them only switch between the four "green light main phases", the number corresponds to the green light state allowed to run in the SUMO timing file, indicates the phase number of the intersection geometry and traffic flow direction, which only releases the four main passing phases. and respectively represent the preset minimum and maximum green light time length, indicates the green light duration, which is in integer seconds, indicating how long this green light lasts.

[0112] S8022, environment conversion. The execution of the action space in the environment is divided into two steps, the action of the agent is first cut to the specified green light phase through traci, and then the simulation is advanced seconds (read the state every second and add the time); then cut into the corresponding yellow light phase, last for a fixed transition time ; finally, according to the accumulated phase time and the final phase, it is judged whether the whole simulation is ended.

[0113] S803, segmented reward function design, the reward function includes an emergency vehicle priority reward item and a social vehicle queuing and dredging reward item, which is used to guide the PPO agent to carry out multi-objective collaborative optimization; The overall reward is composed of emergency vehicle reward , social vehicle reward and phase switching penalty linear combination:

[0114] Among them, is the time is used for adaptive weight coefficient between emergency priority and social vehicle queuing optimization two types of returns, and , the emergency vehicle weight is given priority, which is dynamically calculated and adjusted by the potential function.

[0115] S8031, Emergency vehicle reward. Reward the emergency vehicle that meets the priority release distance :

[0116] where, and are the components of the reward function, and is 1 if the emergency vehicle is in the green or yellow lane at time , otherwise 0, and are the component weights of and , respectively. is the speed increment, encouraging the emergency vehicle to accelerate, and is the weight. is the potential function increment, is the weight, which is specified as follows: ,

[0117] where is the distance of the emergency vehicle to the intersection, is a normalization constant, and finally is normalized online with mean and variance and truncated to .

[0118] S8032, Social vehicle reward. Considering the maximum queue, delay, equilibrium, and throughput of all social vehicles, a potential function mechanism is adopted to decompose the complex optimization objective of peak hours into two parts: absolute core optimization and potential plasticity. The equilibrium of potential difference is achieved through annealing: Reward 1: Social vehicle core reward. The core reward of social vehicles adopts an absolute reward expression form. Since the optimization control environment of PPO requires an immediate reward at each time step, rather than a large reward after the entire cycle, we use (the increment of the number of vehicles that flow out at this step) to measure the outflow, rather than a cumulative total. The delay is also normalized using (the average waiting time of all vehicles at this step), rather than the average at the end of the cycle. In this way, PPO can receive immediate feedback on whether it has done well or poorly at each step, ensuring that the policy gradient update under Markov Decision Process (MDP) is effective. The specific expression is as follows:

[0119] Reward two: social vehicle potential function shaping reward: the maximum queue, equilibrium potential function shaping of social vehicles guarantees the invariance of their strategies, guarantees that PPO focuses on the queue peak and imbalance, and does not change the optimal strategy set. First, define the potential function of the maximum queue and the equilibrium index as the maximum queue potential and the imbalance potential , which is expressed as follows:

[0120]

[0121] The classical potential function shaping reward at time is given as:

[0122] wherein is the potential function discount coefficient, is the potential function calculated by the state .

[0123] Based on the classical potential function difference theory, the above defined maximum queue potential and the imbalance potential are expressed as the following maximum queue potential function difference and the queue equilibrium potential function difference :

[0124]

[0125] Finally, the two parts of the potential function difference are integrated by weighting, and is used to represent the second component of the social vehicle target reward, that is, the shaping item based on the potential function difference "maximum queue length and equilibrium", which measures whether the maximum queue and queue imbalance at time is improved compared with , and is used to guide the strategy to reduce the bottleneck and imbalance while meeting the emergency priority, which is as follows:

[0126] Therefore, the total reward of the social vehicle is the weighted sum of the core reward and the potential function shaping reward:

[0127] wherein: : lane average waiting delay, which measures the delay cost; : maximum delay constant; : the increment of outflow; : the maximum observable outflow; : the normalization constant, usually the maximum queue length of a lane; : the standard deviation of queue length, measuring the equilibration; : the maximum queue length; : the discount factor of shaping (usually consistent with the PPO main value), ensuring the potential value decay smooth; : the component weight of the maximum queue potential difference, the normalized delay, the queue equilibration potential difference and the normalized number of departing vehicles (throughput), respectively; : the overall strength of social shaping.

[0128] S8033, dynamically adjust the potential function shaping. This embodiment introduces a periodic annealing mechanism in the emergency priority strategy to dynamically adjust the shaping weight . The emergency arrives immediately to the maximum value , the reward guidance of the "queue equilibration" and "preventing overflow" of social vehicles offsets the risk of destroying the queue balance caused by the emergency switching; after the emergency leaves, linear annealing is performed in the next 5 main phase periods, which will smoothly recover from the maximum value to the baseline level, so as to ensure the rapid passage of emergency vehicles while smoothly transitioning back to regular traffic optimization.

[0129]

[0130] wherein, represents the number of cycles passed after the emergency vehicle leaves the control intersection, represents the annealing cycle, a total of cycles are transitioned, represents the baseline weight returned after annealing is completed. Ensure emergency vehicle priority moment, maximize queue equilibration, after the emergency vehicle leaves, gradually weaken the importance of "equilibration", avoid long-term excessive pursuit of equilibration and loss of traffic efficiency.

[0131] S8034, phase switching penalty. If the phase changes and the duration is too short, then:

[0132] wherein, represents the duration of this green light, represents the effective distance threshold of emergency priority, represents the tolerance threshold, if the switching duration is too short in the control process, This indicates the upper limit of tolerance for the effective distance threshold for emergency priority, then execution... punish, This represents the phase switching penalty weight.

[0133] S8035, EV priority weight Dynamically adjusted. Sigmoid weights are designed based on the location of emergency vehicles and their distance relative to the tail end of the queue in the corresponding lane at the controlled intersection.

[0134] in The distance from the emergency vehicle to the intersection is represented by exp, which is the natural exponential function with parameters. To control the steepness of the curve, dynamic balancing EV is prioritized over SV optimization.

[0135] S804, Algorithm Training. S is calculated using PPO. t Generational Strategy Network and value network The process involves three steps: data sampling, estimating advantage, and optimizing objectives.

[0136] S8041, Data Sampling. Under the current strategy... The system interacts with the environment to obtain a trajectory:

[0137] :time The state; Actions sampled based on the old strategy; :action Instant rewards obtained in the environment; The new state reached after an action; The length of an episode (or the truncation length).

[0138] S8042, Dominance Function Estimation. Let... Indicates the state Next, execute The "excess returns" obtained:

[0139] In state Take action below The immediate reward received afterward; Discount factor: Used to reduce the impact of future value; Value network for the next state The estimate; Value network for the current state The estimate; One-step TD goal refers to a single-step bootstrapping reward. S8043, PPO Optimization Objective. The core of PPO is to limit the update magnitude. The following formulas respectively represent the optimization objective pruning and the MSE fitting of the TD objective. :

[0140]

[0141] in, For the first The reward objective of each step is used to train the value function. for The strategy of tailoring the target, For strategy parameters, For the trimming operator, For parameters The value function loss, i.e., the loss of parameters value network For return target Calculate the mean squared error.

[0142] Finally, the overall optimization loss objective is expressed as:

[0143] in, This represents the clipping factor, and the clipping range is []. ],ensure When the deviation is too large, the target is no longer increased, thereby controlling the magnitude of the strategy update. This indicates the addition of a policy entropy term to encourage exploration. and These represent the value loss coefficient and the entropy coefficient, respectively.

[0144] S9. Training Process: During the training process of PPO, the agent first adopts the current policy. Interact with the environment to continuously collect a segment of trajectory data. And perform necessary normalization or prediction enhancement on the state; then utilize the value network Calculate the one-step TD target and advantage estimation ; then multiple rounds of joint optimization of policy and value networks are performed on the batch of data - the policy network performs gradient ascent on the clipping target, and the value network minimizes mean squared error, with an additional policy entropy regularization to encourage exploration; after each round of update, the new policy parameters are synchronized to , the sampling buffer is emptied, and the next round of sampling-new cycle is entered, until the policy converges or the training budget is reached.

[0145] This embodiment optimizes the control strategy of traffic signal lights by combining the SUMO simulation environment, LSTM-Transformer traffic prediction model, and PPO reinforcement learning agent. The goal of the model is to optimize the multi-objective function by designing an appropriate reward strategy. It improves traffic flow and the efficiency of emergency vehicle passage.

[0146] This embodiment addresses the technical problem of urban traffic signal control and emergency vehicle priority passage, and develops an intelligent traffic signal control system based on deep reinforcement learning in the Python language environment. The system integrates PPO and LSTM-Transformer traffic prediction model to analyze the dynamic characteristics of traffic flow and optimize the signal light control strategy in real time, achieving efficient management and control of traffic flow. The system architecture includes three core modules: multi-source heterogeneous data-driven traffic flow prediction module, PPO agent training module, and real-time signal optimization control module.

[0147] This embodiment uses deep reinforcement learning technology to realize intelligent control of intersection signal lights. It accurately predicts the queue length and emergency vehicle operating status of each lane in the future period through the LSTM-Transformer traffic prediction model, and uses the prediction results as state input to PPO for decision-making. It innovatively designs a reward mechanism that integrates traffic status, emergency vehicles, and signal optimization, enabling the agent to autonomously optimize the signal control strategy, improving the passage efficiency of emergency vehicles in saturated traffic environment, reducing the impact of emergency vehicle priority passage on social vehicles, and improving the overall traffic conditions. Compared with traditional priority passage and signal optimization schemes, this invention takes into account emergency response efficiency and intersection passage efficiency, realizes fully data-driven dynamic regulation, and has stronger environmental adaptability and intelligence level. As shown in Figure 5 , during the training process of the invention, the training loss and validation loss corresponding to the lane-level queue evolution prediction task of the model both show a rapid downward trend with the increase of iteration number, and finally tend to be stable, indicating that the model has converged well and can provide reliable prior information for subsequent queue behavior analysis. As shown in Figure 6 ​As shown, the prediction loss of the position, speed and lane state of the emergency vehicle also presents a continuous decrease and finally keeps stable, proving that the front trajectory prediction module adopted by the present application has usability and robustness, and can effectively support subsequent decision control. Figure 7 As shown in (a)-(d) (Ty1.0 and ty2.0 in the figure represent the types of emergency vehicles, corresponding to ambulances and fire engines respectively), the prediction error scatter distribution of the two types of emergency vehicle trajectory state parameters in the embodiment is closely around the zero point, and the mean value is close to zero, and the dispersion degree is small, which shows that there is no significant systematic deviation in the prediction result, and the variance is effectively controlled, and the prediction accuracy meets the design requirements. Figure 8 As shown, through the queuing balance heat map, it can be observed that in the peak congestion scene, the distribution of vehicles across lanes is more uniform, verifying that the equalization shaping method proposed by the present application can indeed effectively improve the queuing distribution between lanes and improve the balance of the overall traffic flow. Figure 9 As shown, in the reinforcement learning training process, the entropy value of the PPO strategy of the embodiment gradually decreases from a high level, and finally remains small amplitude oscillation in a small range, and the change process shows that the agent has fully explored in the early learning stage, and then gradually converges to a stable and efficient strategy, indicating that the training process is reasonable and effective.

[0148] Each embodiment in the specification is described in a related manner, and the same and similar parts between each embodiment can be referred to each other, and each embodiment mainly explains the difference from other embodiments. Especially, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can refer to the part of the method embodiment.

[0149] The above only describes the preferred embodiments of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An emergency vehicle priority and queuing optimization collaborative control method, characterized in that, According to the following steps: S1, construct a road intersection road network topology model and collect multi-source traffic state data; S2, pre-process the collected data and extract multi-source heterogeneous features to construct a standardized feature vector; S3, construct a time series data set based on the pre-processed data; S4, establish a dual-branch traffic prediction model that integrates LSTM and Transformer architecture: S5, train the dual-branch traffic prediction model; use a multi-task mean square error loss function that includes queue evolution prediction loss and emergency vehicle operating state prediction loss to optimize the parameters of the integrated dual-branch traffic prediction model until convergence; S6, construct a PPO agent, the state space of which integrates real-time traffic state, historical observation data, and the prediction results of the dual-branch traffic prediction model; S7, design a segmented reward function to guide the PPO agent to perform multi-objective collaborative optimization, and balance the emergency vehicle priority and social vehicle optimization objectives through dynamic weight coefficients; S8, train the PPO agent; the PPO agent interacts with the traffic simulation environment to sample trajectory data, determines the state-action excess return through the advantage function estimate, and updates the policy network and value network parameters in combination with the clipping optimization objective; S9, deploy the trained dual-branch traffic prediction model, input real-time collected traffic state data, and obtain future lane-by-lane queuing evolution trends and emergency vehicle trajectory prediction results; the PPO agent outputs signal phase selection and green light duration decisions based on real-time traffic state data and the prediction results, controls the traffic signal lights to execute the corresponding timing scheme, and the signal phase selection only switches between the four green light main phases, achieving emergency vehicle priority and social vehicle queuing optimization.

2. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, The intersection road network topology model in S1 is the associated area of five intersections, the center intersection uses a four-stage eight-phase timing scheme, which includes two straight left phases in the east-west direction and two straight and left turn phases in the north-south direction; the lengths of the four approaches of the center intersection are 800m, 450m, 386m, and 400m respectively, and one straight lane and one left turn lane are selected for each approach as the research and control object; the traffic simulation environment uses the SUMO simulation platform, and the simulation duration is 1 hour, the multi-source traffic state data output by the simulation is saved as a CSV file.

3. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, S2 specifically constructs composite features such as queue / flow ratio and net growth queue through feature engineering, removes redundant information through correlation analysis, mutual information evaluation, VIF multicollinearity detection, and feature importance screening, and then normalizes the features to obtain a standardized feature vector; where the normalized feature vector is: (1) wherein, is a very small constant, is the mean of each dimension feature in the training set, is the standard deviation of each dimension feature in the training set, is the feature vector before normalization.

4. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, In S3, the specific process of constructing the time series data set is as follows: Let the original time series data be: (2) wherein, represents an input all-lane feature vector at time instant, represents an output queue length at time instant, represents a time series length; establishing an input sequence for each time step and a predicted target sequence : (3) (4) The final dataset Is: (5) wherein, denotes the input time series length, denotes the output time series length, And the data set is divided into training set and test set in the ratio of 7:

3.

5. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, In the S4, the double-branch traffic prediction model comprises an LSTM encoder, a Transformer encoder and a double-branch prediction head, the LSTM encoder captures local time sequence dynamic characteristics of traffic flow, the Transformer encoder captures global space-time dependence, and the double-branch prediction head respectively outputs prediction results of lane-by-lane queuing evolution trend and trajectory state of emergency vehicles in a future period. The LSTM encoder hidden layer is 112, and the total number of layers is 3. The Transformer encoder has 3 layers, 4 attention heads, a feedforward layer size of 128, a dropout rate of 0.1, an output window of 5, and a state prediction output dimension of 2 for each emergency vehicle.

6. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, Multi-task mean squared error loss function in S5 Specifically: (6) where, is the inter-task uncertainty variable, which represents the difficulty of each task, is the regularizer of uncertainty variable, which constrains the model’s estimation of uncertainty, is the queue evolution prediction loss: (7) wherein, and respectively represent the predicted value output of the first prediction time, and the true value output of the first prediction time. Emergency vehicle operational state prediction loss: Assume each emergency vehicle contains 4 prediction variables, totaling vehicles, then the emergency prediction dimension is Then: (8) wherein, and respectively represent the predicted value output of the th emergency vehicle, the true value output of the th emergency vehicle, represents the emergency prediction dimension.

7. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, Final state of the PPO agent in S6 is represented as: (9) wherein, is the normalized raw input state, containing the queue length of the main road lane, vehicle waiting time, lane flow, signal state one-hot encoding, intersection throughput, simulation time, and the lane, position, and speed of emergency vehicles; is the normalized raw input state, containing the queue length of the main road lane, vehicle waiting time, lane flow, signal state one-hot encoding, intersection throughput, simulation time, and the lane, position, and speed of emergency vehicles; is an additional flag to assist the strategy in determining when to make decisions based on prediction information; is the refined future is the time sequence correlation encoding vector of the future is a vector splicing operation. 8.The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, S7 piecewise reward function including emergency vehicle reward , social vehicle reward combined with phase switch penalty ​ (10) Wherein, is the time An adaptive weight coefficient is used between the two types of rewards of emergency priority and social vehicle queuing optimization, and Specifically, (11) wherein represents the distance of the emergency vehicle to the intersection, exp is the natural exponential function, is a maximum delay constant, is a parameter that controls the steepness of the curve.

9. The emergency vehicle priority and platooning optimization cooperative control method of claim 8, wherein, The emergency vehicle reward Specifically: (12) wherein, and denotes 1 if the emergency vehicle is in the same lane as the lane signal is green or yellow, and 0 otherwise, and , are the component weights of and respectively; denotes the speed increment, is the weight of , is the potential function increment, is the weight of . The social vehicle reward Specifically: (13) (14) (15) wherein, is the increment of the outflow traffic volume of the current step; are the component weights of the maximum queue potential difference, the normalized delay, the queue equilibrium potential difference, and the normalized number of departing vehicles, respectively; is the overall strength of the social shaping; is the maximum observable outflow traffic volume; is the average waiting delay of the lane; is the maximum delay constant; is the maximum queue potential function difference; is the queue equilibrium potential function difference; is the core reward of the social vehicle; is the potential function shaping reward of the social vehicle; The phase switch penalty Specifically: (16) wherein, represents the current green light duration, represents the emergency priority effective distance threshold, represents the tolerance threshold, represents the phase switch penalty weight, represents the tolerance upper limit of the emergency priority effective distance threshold.

10. The emergency vehicle priority and queuing optimization cooperative control method of claim 1, wherein, In the S8, the advantage function estimation determination process is specifically as follows: (17) to take action in state the immediate reward obtained after; is a discount factor that attenuates the influence of future values; is the estimate of the value network for the next state ; is the estimate of the value network for the current state ;​ The total optimization goal of the PPO agent is : (18) wherein, represents an increased policy entropy term for encouraging exploration, and respectively represent a value loss coefficient and an entropy coefficient; is a policy clipping target; is a value function loss.

Citation Information

Patent Citations

  • Intersection priority control method based on emergency lane

    CN112216131A

  • Method for improving traffic passing efficiency by utilizing intelligent network connection vehicle

    CN112700642A

  • Emergency signal priority and social vehicle dynamic path induction collaborative optimization method

    CN113724510A

  • Emergency vehicle priority passing control method based on parameterized reinforcement learning

    CN119832755A

  • Short-time rail transit flow prediction method and device based on decoupling process

    CN120125052A

Cited By

  • Signal lamp duration estimation method based on depth image prior guided by physical information

    CN121600496A

  • Signal light duration estimation method based on physical information guided deep image prior

    CN121600496B