Start-stop channel decision-making method and system based on reinforcement learning
By constructing a traffic flow interruption probability model and Markov decision-making process based on reinforcement learning, optimizing the start and stop decision-making of highway emergency lanes, solving problems such as insufficient scientific decision-making and lagging response in the existing technology, and achieving more efficient and reasonable use of emergency lanes.
Patent Information
- Application Number
- CN202510194389.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
AI Technical Summary
The existing highway emergency lane start and stop decisions rely on manual monitoring, which has problems such as insufficient scientific nature, lagging response, inability to accurately predict and waste of resources.
The start-stop channel decision-making method based on reinforcement learning is adopted, and by constructing a traffic flow interruption probability model and Markov decision-making process, the Q-Learning algorithm is used to optimize the start-stop decision-making logic, so as to achieve accurate prediction of future traffic conditions and reasonable use strategies for emergency lanes.
It improves the scientificity and efficiency of emergency lanes, reduces resource waste, ensures traffic capacity at critical moments, and reduces the risk of traffic congestion and the possibility of accidents.
Smart Images

Figure CN119992836A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of highway traffic control, and in particular relates to a start-stop channel decision method and system based on reinforcement learning. Background Art
[0002] As an important part of the modern transportation system, highways play a key role in promoting economic exchanges between regions and improving transportation efficiency. However, with the rapid growth of the number of vehicles, highway congestion is becoming increasingly serious, especially at key sections such as ramp entrances and exits and bridge entrances, where vehicles gather significantly, resulting in limited traffic capacity. To meet this challenge, many countries and regions have introduced the concept of emergency lanes as life-saving channels, emergency evacuation channels, and temporary lanes to relieve traffic pressure in special circumstances.
[0003] Emergency lanes are usually strictly restricted to vehicles that are only used for emergency missions, such as ambulances, fire trucks, and police cars, to ensure that these vehicles can quickly reach the scene of the accident or the location where help is needed. In addition, under certain conditions, such as when a major traffic accident or a large-scale event causes a surge in traffic, the management department will consider temporarily opening the emergency lane to ordinary vehicles to speed up evacuation and reduce pressure on the main road.
[0004] Currently, most highway management departments rely on manual monitoring and experience to decide whether to enable the emergency lane. This method mainly uses video surveillance systems in multiple sections to observe, and operators make decisions based on real-time road conditions and personal experience. Although this method can meet the needs to a certain extent, they also expose some significant problems:
[0005] ① Lack of scientificity: Since decisions are mainly based on personal experience and intuition, they are prone to controversy, and the decision-making criteria may be inconsistent among different operators, which may lead to unnecessary confusion.
[0006] ② Response lag: When relying on manual monitoring, it takes a long time to identify changes in traffic conditions and respond, which may miss the best time and affect the effect of activating the emergency lane.
[0007] ③ Unable to make accurate predictions: Without a mathematical model to discover the specific conditions under which congestion is about to occur, it is impossible to provide early warning, so the emergency lane is often not activated until congestion has already occurred, weakening its preventive and proactive nature.
[0008] ④ Waste of resources: If the emergency lane is frequently or improperly activated, it will not only reduce its effectiveness as a life-saving channel, but may also consume excessive human and material resources, and will also have a negative impact on normal traffic order. Summary of the invention
[0009] The present invention aims to address the deficiencies in the start-stop decision-making of emergency lanes on existing highways, and proposes a start-stop channel decision-making method and system based on reinforcement learning. An efficient traffic flow interruption probability model is constructed using a reinforcement learning algorithm, which can accurately predict the traffic conditions in the future and plan the use strategy of the emergency lane in advance. By optimizing the start-stop decision logic, unnecessary resource waste is reduced, and at the same time, the rationality and efficiency of the use of the emergency lane are improved, ensuring that the lifeline is unobstructed at critical times.
[0010] In order to achieve the above object, the present invention adopts the following technical solutions:
[0011] The present invention provides a start-stop channel decision method based on reinforcement learning, comprising:
[0012] Constructing a traffic flow interruption probability model, describing the relationship between the congestion probability and traffic flow parameters, and determining the warning vehicle speed and warning flow; the traffic flow parameters include flow, density and speed;
[0013] The start-stop channel decision problem is converted into a Markov decision process, and the state space, action space and state transition probability of the Markov decision process are determined;
[0014] The Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision.
[0015] Preferably, the construction of a traffic flow interruption probability model, describing the relationship between the congestion probability and the traffic flow parameters, and determining the warning vehicle speed and warning flow rate includes:
[0016] The Brilon model is used to construct the traffic flow interruption probability model, which is expressed as:
[0017] ;
[0018] in: Indicates flow rate, Indicates capacity, which refers to the maximum number of vehicles that can pass through a road or a section of a road in a unit of time. Indicates the observed flow Under these conditions, the capacity Less than or equal to The probability of is in the time interval The traffic volume observed at means The number of time intervals, The traffic flow reaches The number of time intervals, is the congestion interval A collection of;
[0019] Read the flow, speed and timestamp data of each observation point, and iterate and calculate the data at each time point , and then obtain the congestion probability and traffic flow distribution;
[0020] Draw speed and congestion probability distribution maps and flow and congestion probability distribution maps to determine warning vehicle speed and warning flow.
[0021] Preferably, before constructing the traffic flow interruption probability model, the method further includes:
[0022] Acquiring traffic flow data; the traffic flow data includes flow, density and speed in the traffic flow;
[0023] Using Origin software, the road congestion level RCL is used as the label ColorBar to color the corresponding data in the flow-speed distribution diagram in the traffic flow basic diagram. Combined with the visualization results, threshold analysis is performed to determine the congestion speed threshold.
[0024] Preferably, the step of converting the start-stop channel decision problem into a Markov decision process and determining the state space, action space and state transition probability of the Markov decision process includes:
[0025] Define the state space of the Markov decision process It is expressed as: ,in, represents the number of vehicles passing through the video observation point at time t, represents the traffic density of the video observation point at time t, represents the average vehicle speed at the video observation point at time t;
[0026] Defining the action space of a Markov decision process It is expressed as: ,in, Indicates that the emergency lane is not enabled. Indicates that the emergency lane is activated;
[0027] The state transition probability is determined by adopting a random state transition strategy, that is, after the algorithm takes an action, the environment gives an immediate reward and then transfers to the next state through a random state generator.
[0028] Preferably, the Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision, including:
[0029] The Q-Learning algorithm is used to solve the Markov decision process.
[0030] In the Q-Learning algorithm, the reward function is designed as:
[0031] ;
[0032] in, ,
[0033] ,
[0034] ,
[0035] in, represents the comprehensive reward function, Rewards for congestion relief, Enable penalties for non-bottleneck periods, To punish frequent switching, represents the congestion coefficient at time t, represents the congestion coefficient at time t+10, is the scaling factor, and are penalty coefficients, is the switch penalty function,
[0036] It is expressed as:
[0037] ,
[0038] in is a decay factor used to calculate the weight of each state change in the historical record, Display history Middle The state of the secondary emergency lane indicates that the emergency lane is Is it enabled? 1 means enabled, 0 means disabled. is an indicator function used to determine the Second and Whether the status of the secondary emergency lane has changed, if ,but ,otherwise, ;
[0039] In the Q-Learning algorithm, the Q table is constructed as follows: the Q table is a two-dimensional array, in which the rows represent the state space of the Markov decision process and the columns represent the action space of the Markov decision process;
[0040] In the Q-Learning algorithm, actions are selected using an ε-greedy strategy.
[0041] Preferably, the Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision, including:
[0042] The observed state space data is compressed by a data compression algorithm to reduce the state space to 10 3 Magnitude, as the S* input to the Q-learning algorithm;
[0043] The Q-learning algorithm observes the current state of the environment S*, selects actions based on the ε-greedy algorithm and the Q table, and outputs the actions to the environment;
[0044] The environment takes the current S and action as input, outputs the congestion level Cl, and based on the congestion level Cl based on the reward function Calculate rewards;
[0045] Update Q value;
[0046] Repeat the above steps until the maximum action step is reached;
[0047] Repeat multiple rounds until the maximum number of rounds is reached.
[0048] Preferably, the updating of the Q value comprises:
[0049] Based on the Bellman equation, the Q value in the Q table is updated as follows:
[0050] ;
[0051] in, is the current Q value when taking action A in state S, is the learning rate, is the discount factor, is the new state reached after executing action A, It is a new state The maximum Q value of all possible actions.
[0052] Preferably, in the solving process,
[0053] The learning rate The initial setting is 0.1, which then decreases gradually over time;
[0054] The discount factor Set between 0 and 1;
[0055] The exploration rate ε has an initial value of 1 and gradually decreases over time.
[0056] The present invention also provides a start-stop channel decision system based on reinforcement learning, which is used to implement the above-mentioned start-stop channel decision method based on reinforcement learning. The system includes:
[0057] Traffic flow interruption probability model building module, used to build a traffic flow interruption probability model, characterize the relationship between congestion probability and traffic flow parameters, and determine the warning speed and warning flow;
[0058] A Markov decision process conversion module, used to convert the start-stop channel decision problem into a Markov decision process, and determine the state space, action space and state transition probability of the Markov decision process;
[0059] Reinforcement learning algorithm model building module, used to build a reinforcement learning algorithm model, including designing reward functions, building Q tables, and determining Q value update strategies and parameter settings;
[0060] The calculation and output module is used to solve the reinforcement learning algorithm model to obtain the start-stop channel decision.
[0061] Preferably, the system further comprises:
[0062] Data preprocessing module, used to
[0063] Acquiring traffic flow data; the traffic flow data includes flow, density and speed in the traffic flow;
[0064] Using Origin software, the road congestion level RCL is used as the label ColorBar to color the corresponding data in the flow-speed distribution diagram in the traffic flow basic diagram. Combined with the visualization results, threshold analysis is performed to determine the congestion speed threshold.
[0065] The beneficial effects of the present invention are:
[0066] The present invention provides a start-stop channel decision method and system based on reinforcement learning, which has significant advantages and effects in structural features and technical implementation compared with the prior art. First, the present invention provides a solid theoretical basis for the start-stop of the emergency lane by introducing a traffic flow interruption probability model (such as the Brilon model). This data-driven method can quantify the auxiliary decision-making process and ensure that decisions in different situations are consistent and explainable. At the same time, a standardized decision-making process is established to overcome the subjectivity and inconsistency problems caused by traditional reliance on manual experience and intuition, and improve the quality and reliability of decision-making.
[0067] Secondly, the present invention uses Markov decision process (MDP) modeling and Q-Learning algorithm to achieve real-time monitoring and rapid response to traffic conditions, and can make the best decision in the shortest time, effectively preventing or alleviating the occurrence of traffic congestion. The intelligent agent automatically selects the best action according to the current environmental state, without human intervention, reducing human delays and ensuring the immediacy and accuracy of decision execution. In addition, by constructing and training a reinforcement learning model, the present invention can accurately predict changes in traffic flow in the future, plan the use strategy of emergency lanes in advance, avoid passive response afterwards, and issue congestion warnings in time based on the set warning speed and warning flow, helping management departments take preventive measures and reduce congestion risks.
[0068] Furthermore, the present invention fully considers the safety of enabling the emergency lane when designing the reward function. By setting up a reasonable reward and penalty mechanism, it encourages the use of the emergency lane under safe conditions, while avoiding the safety hazards caused by frequent switching. Continuous monitoring of traffic flow parameters and learning optimization combined with historical data make the management of the emergency lane more stable and orderly, reducing the possibility of traffic accidents caused by drivers suddenly changing their driving paths. This not only strengthens safety protection, but also reduces the risk of accidents.
[0069] Finally, the present invention avoids overly frequent operations by precisely controlling the activation time of the emergency lane, reduces the consumption of human and material resources, and also improves the effectiveness of the emergency lane as a life-saving channel. Through intelligent regulation of traffic flow, the present invention not only improves road capacity, but also promotes the efficient operation of the entire transportation system, reduces vehicle waiting time and energy consumption, and brings economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 The flow and speed distribution diagram of each observation point during the preprocessing of traffic flow data in the embodiment of the present invention;
[0071] Figure 2 The density and speed distribution diagram of each observation point during the preprocessing of traffic flow data in the embodiment of the present invention;
[0072] Figure 3 It is a time distribution diagram of congestion probability calculated by the traffic flow interruption probability model constructed based on the Brilon model in an embodiment of the present invention;
[0073] Figure 4 A distribution diagram of speed and congestion probability calculated in an embodiment of the present invention;
[0074] Figure 5 A flow rate and congestion probability distribution diagram calculated in an embodiment of the present invention;
[0075] Figure 6 A schematic diagram of the interaction process between an intelligent agent and a traffic flow interruption probability model provided in an embodiment of the present invention;
[0076] Figure 7 A schematic diagram of a cumulative reward learning curve calculated in an embodiment of the present invention;
[0077] Figure 8 A schematic diagram of a process flow of a start-stop channel decision method based on reinforcement learning provided by an embodiment of the present invention;
[0078] Fig. 9 A reinforcement learning-based start-stop channel decision system architecture is provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0080] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0081] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0082] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0083] It should be emphasized here that the step marks mentioned below are not intended to limit the order of the steps, but it should be understood that the steps can be executed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be executed simultaneously.
[0084] The embodiment of the present invention provides a start-stop channel decision method based on reinforcement learning, which is as follows:
[0085] S1. Construct a traffic flow interruption probability model to characterize and analyze traffic congestion.
[0086] The mathematical description of congestion probability involves three basic parameters in traffic flow theory: flow, density, and speed. In traffic flow theory, there is a complex nonlinear relationship between these three parameters. Specifically, an increase in vehicle density usually leads to a decrease in vehicle speed. When the density exceeds a certain critical value, traffic congestion will occur. The temporary activation of emergency lanes needs to be based on accurate congestion predictions to avoid unnecessary confusion and safety hazards. In existing research, traffic flow probability models mainly include models based on cellular automata, models based on fluid dynamics, models based on queuing theory, etc. The embodiment of the present invention first constructs a traffic flow interruption probability model to characterize and analyze the traffic congestion state. The specific process includes:
[0087] S11. Data preprocessing
[0088] First, the existing traffic flow data is preprocessed, and the potential related properties in the data are evaluated and determined with the help of the basic traffic flow map and the road congestion level RCL of each road section. Using Origin software, the road congestion level RCL is used as the label ColorBar to color the corresponding data in the "flow and speed distribution map" in the basic traffic flow map to intuitively present the statistical relationship between the basic traffic flow information and the road congestion level RCL, and then combined with the visualization results to perform threshold analysis and determine the congestion speed threshold. For details, see Figure 1 and Figure 2 .
[0089] from Figure 1 It can be seen that when the vehicle speed is above 60km / h, the traffic is basically in a free flow state. When the speed gradually decreases to below 60km / h, the congestion level gradually increases. As the speed decreases and the density increases, the congestion state shows a significant increasing trend. When the average vehicle speed drops below 40km / h, it shows a significant congestion state. Figure 2 It can also be seen from the flow speed distribution diagram of each observation point that when the vehicle speed drops below 40 km / h, the road congestion level tends to the maximum value and increases significantly. Based on the above analysis, the embodiment of the present invention determines the congestion speed threshold as 40 km / h.
[0090] It should be noted that the congestion speed threshold, as an environmental parameter, indirectly affects the strategy update through state space partitioning and reward function design. In the Q-Learning algorithm, as the basic parameter of the environmental feedback mechanism, its value selection will significantly affect the convergence direction and the final strategy characteristics.
[0091] S12. Constructing a traffic flow interruption probability model
[0092] To construct a traffic flow interruption probability model, we first need to select a suitable model to describe it. In terms of model selection, there are mainly the following options, including the Brilon model, logistic regression, random forest and neural network. The specific characteristics of each model are shown in Table 1 below.
[0093] Table 1 Analysis of the main features of various models
[0094]
[0095] In terms of model selection, the embodiment of the present invention makes targeted selections based on the characteristics. The main analysis of the requirements for building this model includes three aspects: first, the model has a theoretical basis for certain traffic flow characteristics to improve the interpretability of the model; second, the model can make full use of the mined data on a limited data set; third, it can provide real-time, reliable, and high-precision predictions in terms of prediction accuracy. The embodiment of the present invention preferably uses the Brilon model to construct a traffic flow interruption probability model, as follows:
[0096] (1)
[0097] in: Indicates the flow rate, which has been obtained in S11; Capacity refers to the maximum number of vehicles that can pass through a road or a section of a road in a unit of time. This is usually called the road's traffic capacity and is an indicator of the upper limit of the traffic flow that a road can accommodate; Indicates the observed flow Under these conditions, the capacity Less than or equal to probability; is in the interval Traffic flow observed at means The number of time intervals; The traffic flow reaches When a congestion is considered separately, Set to 1; is the congestion interval A collection of .
[0098] It should be noted that both time interval and congestion interval represent time attributes, but the congestion interval is only a specific manifestation of the time interval.
[0099] S13. Numerical calculation and model analysis
[0100] Read the flow, speed, timestamp and other data of each observation point, and iterate and calculate the data of each time point , where i=5, indicating sampling from time points t to t+5. Through iterative calculation, the congestion probability and traffic flow distribution are obtained, such as Figure 3 shown.
[0101] according to Figure 3 It can be seen that, in general, the probability of congestion is high between time step 0 and time step 140, and the probability of congestion has been significantly reduced after time step 150. Comparing the video information, it can be seen that this time period is in the peak interval of driving, and the traffic flow is large. The correlation can be significantly observed from the time distribution diagram of traffic flow, speed and traffic density. Specifically observing the trend of congestion probability, it can be found that as time goes by, the probability of traffic congestion at the peak stage near time step 69 has a large degree of decline and fluctuation. Therefore, the congestion probability P shows a significant downward trend between time steps 63 and 69, but after the emergency lane exceeds its traffic capacity carrying range, the congestion probability rises significantly. This phenomenon shows that even if the emergency lane is opened on this section of road, the change trend of the congestion probability P can intuitively reflect that its road carrying capacity can no longer support the smooth traffic of this section.
[0102] Furthermore, the embodiment of the present invention attempts to explore the relationship between the congestion probability and traffic flow parameters such as speed and flow that can be directly obtained by the device, so as to further reveal the correlation between the congestion probability and traffic flow parameters. Based on this, a distribution diagram of speed and congestion probability and a distribution diagram of flow and congestion probability are drawn, see Figure 4 and Figure 5 .
[0103] from Figure 4 It can be found that there are several important features between speed and congestion probability that can assist decision-making. (1) When the average speed is less than about 60 km / h, the congestion probability increases significantly, basically in the range of more than 80%. However, as the speed decreases, the change in the congestion probability is basically stable in the range of 80% to 100%, with a small increase. It can be seen that the warning speed of this section is 60 km / h. When the speed drops below 60 km / h, it is an important signal with a high probability of congestion. (2) When the speed is at a higher level (greater than 80 km / h), the congestion probability is more distributed in the range of more than 80% and approaching 0. It can be seen that when the speed is high, the congestion probability shows a polarization trend, and there is still a certain probability of congestion at this time.
[0104] from Figure 5It can be observed that there are also several features between traffic flow and congestion probability that facilitate decision-making. In general, the probability of congestion is positively correlated with traffic flow. When the traffic flow is greater than 1,600 vehicles / h to about 2,000 vehicles / h, the overall probability of congestion shows a rapid upward trend. When the traffic flow is greater than 2,000 vehicles, the probability of congestion rapidly increases to around 80%, and then the rate of increase slows down. Based on this, it can be considered that the traffic flow of 1,600 to 2,000 vehicles / h is the warning traffic flow for this section, and when the traffic flow is greater than 1,600 vehicles / h, the probability of congestion will increase significantly.
[0105] The above-mentioned warning vehicle speed and warning traffic volume can provide a theoretical basis and efficient decision-making assistance for the reasonable activation of emergency lanes on highways.
[0106] The results and analysis show that the embodiment of the present invention uses the Brilon model to establish a traffic flow interruption probability model, which can provide a theoretical basis for decision makers to temporarily activate emergency channels.
[0107] S2. Markov decision process
[0108] Since the above decision problem has a significant Markov characteristic, that is, the state at the next moment depends only on the state and decision at the current moment, and has nothing to do with the state and decision in the past, the embodiment of the present invention transforms the problem into a typical Markov decision process (MDP), and thus solves it using the reinforcement learning method.
[0109] Markov Decision Process (MDP) is used to simulate the decision-making process of decision makers in an uncertain environment. The MDP model can help decision makers evaluate the long-term effects of different decision strategies and select the optimal strategy to maximize or minimize a certain performance indicator, such as reward or cost. The specific interaction process is as follows Figure 6 As shown in the figure, the decision-maker agent obtains the current state and reward from the environment, interacts with the environment (traffic flow interruption probability model) by performing actions (opening and closing the emergency lane), generates the current moment reward and the next state based on the designed reward function, and returns them to the agent, thus completing a time step interaction process.
[0110] The following describes and constructs the model based on the basic framework of Markov decision process. A complete Markov decision process consists of six tuples: Composition, including status ,action , state transition probability , reward function ,Strategy and the discount factor . The reward function Built by algorithms, strategies is the output of the algorithm, the discount factor Defined by the algorithm, let's first define three basic parameters: state ,action And the state transition probability .
[0111] S21. Define State Space
[0112] The state space is the set of all possible states in the decision process. For a single video observation point, according to the data collected by S11, the state It can be defined as a vector containing flow, density and speed .in, represents the number of vehicles passing through the video observation point at time t, represents the traffic density of the video observation point at time t, Represents the average vehicle speed at the video observation point at time t.
[0113] S22. Define Action Space
[0114] The action space contains a set of two actions, namely .in, Indicates that the emergency lane is not enabled. Indicates that the emergency lane is activated.
[0115] S23, state transition probability
[0116] The state transfer in the embodiment of the present invention adopts a random state transfer strategy, that is, after the algorithm takes an action, the environment gives an immediate reward and then transfers to the next state through a random state generator.
[0117] S3, Q-Learning algorithm construction
[0118] S31. Algorithm selection
[0119] In the embodiment of the present invention, the Q-Learning algorithm is selected as the solution algorithm based on the following considerations:
[0120] (1) The Q-Learning algorithm is a typical model-free reinforcement learning algorithm. The algorithm can learn strategies through interaction without prior knowledge of the transition probability of the environment. It can learn better strategies in actual traffic problems where it is difficult to obtain accurate state transition probabilities.
[0121] (2) As a reinforcement learning algorithm, the Q-Learning algorithm has interpretability that other deep reinforcement learning algorithms do not have. The algorithm clearly lists the Q values of actions in each possible state in the form of a Q table, which can intuitively and transparently display the decision logic, which is conducive to the decision maker who is "in the loop" to understand the supporting basis of the algorithm decision;
[0122] (3) The Q-Learning algorithm supports offline learning, allowing the agent to update the Q value based on historical data without real-time environmental feedback. It is suitable for the offline learning reality of the video historical data learning strategy in this problem.
[0123] The following is a detailed introduction to the implementation of the Q-Learning algorithm in this paper.
[0124] S32. Reward Function Design
[0125] The design of the reward function is the core of the algorithm's convergence. When designing the reward function, we should ensure that the reward function can accurately reflect the effect of enabling the emergency lane and guide the decision-making system to learn effective strategies.
[0126] (1) Congestion relief rewards,
[0127] (2)
[0128] in, Rewards for congestion relief, represents the congestion coefficient at time t, Indicates the congestion coefficient at time t+10. The opening of the emergency lane should be able to alleviate the future congestion risk to a certain extent. If the opening of the emergency lane can reduce the congestion coefficient 10 minutes later, a positive reward will be given. Parameters It is a scaling factor used to adjust the intensity of the reward. It should be noted that the congestion coefficient can reflect the degree of traffic congestion at a specific time, and its acquisition methods include but are not limited to the following:
[0129] Traffic flow sensors: Sensors installed on the road can monitor traffic flow and speed in real time, thereby calculating the congestion factor.
[0130] Camera monitoring: By analyzing video data captured by traffic cameras, road congestion can be estimated.
[0131] Mobile device data: Using GPS data from smartphones or other mobile devices, a large amount of information about vehicle locations and movement speeds can be collected to estimate congestion factors.
[0132] Traffic models and algorithms: By building a traffic flow model and combining historical data with real-time data, algorithms are used to predict and calculate the congestion coefficient.
[0133] (2) Penalty is enabled during non-bottleneck periods.
[0134] (3)
[0135] in, represents the penalty for opening the emergency lane during the non-bottleneck period. During the non-bottleneck period, the number of times the emergency lane is opened should be reduced as much as possible to reduce management risks. Therefore, at the current time i If it is less than 2, a penalty will be imposed to avoid opening the emergency channel at will. is a penalty coefficient used to impose a penalty when the emergency lane is opened during non-bottleneck periods (i.e. when traffic is not congested). The larger the coefficient, the heavier the penalty, thus encouraging drivers not to use the emergency lane at will when traffic is not congested.
[0136] The value should be determined based on actual conditions and policy requirements. Usually, this value needs to be set through experiments, simulations, or expert consultation to ensure that it can effectively prevent the abuse of emergency lanes during non-bottleneck periods while not causing excessive negative impact on the use of emergency lanes. The following factors can be considered when setting the value:
[0137] Traffic rules and regulations: Determine appropriate penalties based on local traffic laws.
[0138] Frequency of emergency lane usage: If the emergency lane is frequently abused, a higher penalty factor may be required.
[0139] Traffic volume and congestion: In areas with high traffic volume or frequent congestion, stricter penalties may be warranted.
[0140] Social and economic benefits: Consider the impact of punitive measures on society and the economy to ensure their rationality and feasibility.
[0141] By setting this value reasonably, the use of the emergency lane can be effectively managed to ensure that it plays a role when it is really needed, while reducing unnecessary traffic interference.
[0142] (3) Penalty for frequent switching,
[0143] The frequency of opening emergency passages should be reduced, and excessive opening and closing of emergency passages should be punished.
[0144] (4)
[0145] in, To punish frequent switching, Indicates the history of emergency lane status changes, is the penalty coefficient, which is used in the frequent switching penalty, indicating the degree of penalty for each frequent switching of the emergency lane. Specifically, The larger the value, the heavier the penalty for frequently switching the emergency lane, thus encouraging less frequent switching of the emergency lane.
[0146] is the switch penalty function, and its expression is:
[0147] (5)
[0148] in, Is a decay factor used to calculate the weight of each state change in the historical record, indicating the change from the current moment To The time interval between state changes, specifically, The larger the value is, the closer the state change is to the previous one, and the greater the impact on the frequent switching penalty is.
[0149] Display history Middle The state of the secondary emergency lane is a binary variable indicating whether the emergency lane is Whether it is enabled, 1 means enabled, 0 means disabled.
[0150] is an indicator function used to determine the Second and Whether the status of the secondary emergency lane has changed. If so, ,but ,otherwise, , this function is used to count the number of emergency lane status changes in the historical records.
[0151] (4) Comprehensive reward function,
[0152] (6).
[0153] Integrate the above three reward functions to form the final comprehensive reward function that integrates multiple considerations . This reward function can reflect the decision-making effect under different traffic conditions from multiple perspectives. While encouraging the agent to enable the emergency channel during traffic congestion, it avoids the abuse of the emergency channel during non-bottleneck periods. At the same time, considering that the emergency channel is an important unit of public transportation, it avoids the misleading and hidden safety hazards caused by repeated opening and closing to the driver, so the reward function has strong interpretability.
[0154] S33. Q-table construction
[0155] In the Q-learning algorithm, the Q table is the core data structure used to store the expected cumulative reward of each state-action pair, that is, the Q value. The update of the Q value is the most critical step in the learning process, which enables the agent to gradually learn the optimal strategy for taking a specific action in a specific state.
[0156] The Q table is set as a two-dimensional array, where the rows represent the state space and the columns represent the action space. In this problem, the state is represented by The possible situations in the traversal are obtained, and the action space has 2 values A sample Q table is shown in Table 2.
[0157] Table 2 Example of Q table
[0158]
[0159] S4, Solving based on Q-Learning algorithm
[0160] In the Q-learning algorithm, the Q-value update strategy is the core of the algorithm, which determines how the agent improves its value assessment of each state-action pair based on experience. The main steps include: initializing the Q table, selecting actions, executing actions and observing results, and calculating Q-value updates. The main core mechanism is the Q-value update based on the Bellman formula and the exploration and utilization based on the ε-greedy strategy. The specific implementation process is as follows:
[0161] S41. Data Compression
[0162] Since Q-learning requires the construction of a Q table, the table contains all the state-action space. The input state data The state space is about 10 6 Therefore, if the original data is used, data explosion will occur easily. Therefore, data compression is needed first. Data compression mainly uses the segmentation method to convert each dimensional data into an integer discrete value between [0,10], thereby converting the state space from 10 6 The magnitude is compressed to 10 3 Magnitude. It includes the following steps:
[0163] S411, determine the number of segments. Divide the data into 10 segments according to the maximum and minimum values of each dimension data.
[0164] S412, calculating the interval width. After determining the number of segments, the width of each interval is calculated.
[0165] S413, defining interval boundaries: Define the boundaries of each interval according to the interval width.
[0166] S414. Write a discretization function that can accept data of various dimensions and convert them into discrete values in the range of [0,10].
[0167] S415: Update the state space. Discretization functions are applied to data of each dimension to complete data compression.
[0168] S42, determine the Q value update strategy
[0169] The Q-value update strategy is based on the Bellman equation, which is used to continuously approximate the long-term value Q-value of taking a specific action in each state. The update formula is as follows:
[0170] (7)
[0171] in, is the current Q value when taking action A in state S, The learning rate is set to 0.99. The discount factor is set to 0.9, is the new state reached after executing action A. In this problem, the state is randomly sampled. is the maximum Q-value of all possible actions in the new state.
[0172] S43. Exploration and Exploitation
[0173] The Q-learning algorithm needs to find a balance between exploration (trying new or unknown actions) and exploitation (taking the best known action), which is achieved through the ε-greedy strategy.
[0174] Exploration: Randomly select an action with a certain probability (ε < 1), which helps discover better actions but may degrade performance in the short term.
[0175] Utilization: Select the action with the highest current Q value with probability (1-ε) to use known information to improve the current decision performance.
[0176] S44, parameter setting
[0177] 1. Learning Rate This parameter determines the rate at which new information covers previous information. The learning rate is initially set to 0.1 and then gradually decreases over time.
[0178] 2. Discount Factor This parameter determines the importance of future rewards to the current moment and is usually set between 0 and 1. In the embodiment of the present invention, it is set to 0.9.
[0179] 3. Exploration rate ε. It is used to define how the agent selects corresponding actions based on the Q table. The initial value of this parameter is 1, and it gradually decreases over time to gradually reduce the number of random explorations so that the agent can use the learned experience to explore more efficiently.
[0180] S45, Algorithm Flow
[0181] Based on the above settings, the overall solution process of the algorithm is divided into the following steps:
[0182] S451, optimize the traffic congestion environment model. The state space of the model is compressed by a data compression algorithm to reduce its space complexity to 10 3 Magnitude, as the S* input to the Q-learning algorithm;
[0183] S452, algorithm training. The Q-learning algorithm observes the current environment state S*, selects actions based on the ε-greedy algorithm and the Q table, and outputs the actions to the environment;
[0184] S453, the environment takes the current S and action as input and outputs the congestion level Cl. A reward is generated according to the congestion level Cl;
[0185] S454, updating the Q value. Specifically, updating the Q value in the Q table according to the Bellman equation;
[0186] S455, repeat the above steps until the maximum action step max_steps is reached;
[0187] S456, repeat multiple rounds until the maximum number of rounds max_episode is reached.
[0188] It should be noted that the congestion level is calculated based on the environment state S and the action, which usually involves a model or algorithm that can predict or evaluate the congestion situation based on the current traffic conditions and the actions taken. For example, models in traffic flow theory, such as fluid dynamics models or models based on cellular automata, can be used to simulate changes in traffic flow and calculate the congestion level accordingly. The calculated congestion level is returned to the environment as feedback information to evaluate the results of the action and provide a basis for the next decision.
[0189] The pseudo code of Q-learning algorithm is shown in Table 3.
[0190] Table 3 Q-Learning algorithm flow
[0191]
[0192] S46. Experimental results analysis
[0193] The embodiment of the present invention builds a development framework based on Windows 10 professional and Anaconda 3 environment, and completes programming implementation, training and verification based on Python 3.9. The experimental process records the cumulative reward function curve of each training round and draws it. The experiment is carried out for 300 rounds, and three groups of hyperparameters are tested. The cumulative reward function is as follows Figure 7 shown. Figure 7 In the example, after the algorithm reaches 100 training rounds, the cumulative rewards of each training round gradually tend to be stable, which indicates that the algorithm has gradually converged. In order to compare the performance of the algorithm under different hyperparameters, the embodiment of the present invention trains the algorithm under three hyperparameter combinations. Figure 7 As you can see, when the learning rate parameter =0.1, =0.9, the training effect is the best, and the algorithm convergence speed and the maximum value of the accumulated reward are better than other parameters to a certain extent. From the verification analysis, it can also be found that under high density and low speed conditions, the trained agent tends to enable the emergency lane to alleviate congestion; while under good traffic conditions, it chooses not to enable it to avoid unnecessary waste of resources.
[0194] In summary, the present invention provides a start-stop channel decision method based on reinforcement learning, which aims to optimize the activation and closing strategy of the emergency lane of the highway through an intelligent algorithm. Figure 8 , including the following steps:
[0195] A1, collect and pre-process traffic flow data and determine the congestion level RCL, and perform visual analysis to determine the congestion speed threshold; the traffic flow data includes flow, speed and density;
[0196] Preprocessing refers to compressing traffic flow data through data compression algorithms to reduce its spatial complexity to 10 3 Magnitude;
[0197] A2, select the Brilon algorithm to build a traffic flow interruption probability model, calculate the congestion probability, provide a theoretical basis for start-stop decision-making, convert the start-stop channel decision problem into a Markov decision process, and define the state space and action space;
[0198] A3, initialize the Q-learning algorithm, set the initial Q table and reward function, define the learning rate, discount factor and exploration rate ε;
[0199] A3, start the training cycle, the agent observes the current environment state S*, and selects action A according to the ε-greedy strategy, and receives the new state S' and reward R after execution;
[0200] A4, use the Bellman equation to update the Q value and improve the value assessment under specific conditions; analyze the cumulative reward curve and adjust the hyperparameters to optimize the training effect;
[0201] A6. End training when the maximum number of rounds max_episode or other stopping conditions are reached.
[0202] In order to ensure that those skilled in the art can reproduce the present invention based on the described content, a preferred embodiment is described in detail below.
[0203] (1) System environment construction
[0204] First, the development framework was built using the Windows 10 Professional operating system and the Anaconda 3 environment, and Python 3.9 was used to complete the programming implementation. This environment provides the necessary software foundation for training and validating the model.
[0205] (2) Data preprocessing
[0206] Data preprocessing is one of the important steps of the present invention, which involves cleaning, converting and analyzing traffic flow data. By reading the flow rate, speed, timestamp and other information of each observation point, the data sampling at each time point is obtained through iterative calculation. For these raw data, the segmentation method is used to discretize them into integer values between [0,10], thereby effectively compressing the state space from about 10 6 The magnitude is reduced to 10 3 It solves the data explosion problem that may occur in the Q-learning algorithm.
[0207] (3) Constructing a traffic flow interruption probability model
[0208] Based on the Brilon model, the present invention constructs a traffic flow interruption probability model to evaluate the probability that the road capacity is less than or equal to the actual flow under different flow conditions. This model helps determine the warning speed (60km / h) and warning flow (1600 vehicles / h) as the threshold conditions for activating the emergency lane.
[0209] (4) Reinforcement learning model design
[0210] (4.1) State space definition
[0211] The state space is represented by a vector containing flow, density, and speed. For example, the state at time t can be expressed as ,in represents the number of vehicles at time t, represents the traffic density, Indicates average vehicle speed.
[0212] (4.2) Action space definition
[0213] The action space includes two actions: a1 (enable the emergency channel) and a2 (disable the emergency channel).
[0214] (4.3) Reward function design
[0215] The reward function is designed to reflect the effect of enabling the emergency lane and guide the decision system to learn effective strategies. It comprehensively considers three aspects: congestion relief reward, non-bottleneck opening penalty, and frequent opening and closing penalty, and finally forms a multi-integrated comprehensive reward function.
[0216] (4.4) Q table initialization and update
[0217] The Q table is a two-dimensional array with rows representing state space and columns representing action space. Initially, all Q values are set to zero. As the training process progresses, the Q values are continuously updated according to the Bellman equation to approximate the long-term value of taking specific actions in each state.
[0218] (4.5) Exploration and Exploitation Strategy
[0219] The ε-greedy strategy is used to balance the relationship between exploring unknown actions and utilizing existing knowledge. At the beginning, ε is set to 1, allowing more random exploration; over time, ε gradually decreases, allowing the agent to rely more on learned experience to make decisions.
[0220] (5) Parameter setting and result analysis
[0221] This example conducted 300 rounds of experiments and tested three different hyperparameter combinations. The results showed that when the learning rate was 0.1 and the discount factor was 0.9, the cumulative reward was the largest and the convergence speed was fast. In addition, the experiment also recorded the cumulative reward function curve of each round and drew relevant charts, such as the cumulative reward learning curve Figure 7 shown.
[0222] (6) Model deployment and application
[0223] After sufficient training, the model can be deployed in the actual environment to monitor traffic conditions in real time and automatically make the best decision. Once it detects that traffic flow exceeds the set threshold or severe congestion is expected in the future, it will issue a warning in time and recommend the use of emergency lanes, thereby effectively preventing or alleviating potential traffic jams.
[0224] In summary, this embodiment demonstrates how to use reinforcement learning algorithms to achieve intelligent management of emergency lanes on highways, which improves road capacity and safety while reducing unnecessary waste of resources.
[0225] Based on the same inventive concept, the embodiment of the present invention also provides a start-stop channel decision system based on reinforcement learning, see Fig. 9 , which includes:
[0226] The data preprocessing module is used to obtain traffic flow data; the traffic flow data includes the flow, density and speed in the traffic flow. The Origin software is used to use the road congestion level RCL as the label ColorBar to color the corresponding data in the flow-speed distribution diagram of the traffic flow basic diagram, and the threshold analysis is performed in combination with the visualization results to determine the congestion speed threshold;
[0227] Traffic flow interruption probability model building module, used to build a traffic flow interruption probability model, characterize the relationship between congestion probability and traffic flow parameters, and determine the warning speed and warning flow;
[0228] A Markov decision process conversion module, used to convert the start-stop channel decision problem into a Markov decision process, and determine the state space, action space and state transition probability of the Markov decision process;
[0229] The reinforcement learning algorithm model building module is used to build the reinforcement learning algorithm model, including designing the reward function, building the Q table, and determining the Q value update strategy and parameter setting;
[0230] The calculation and output module is used to solve the reinforcement learning algorithm model to obtain the start-stop channel decision.
[0231] It is worth pointing out that the system embodiment corresponds to the above-mentioned method embodiment, and the implementation methods of the above-mentioned method embodiment are all applicable to the system embodiment and can achieve the same or similar technical effects, so they will not be repeated here.
[0232] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0233] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0234] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0235] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A start-stop channel decision method based on reinforcement learning, characterized in that: include: Construct a traffic flow interruption probability model to describe the relationship between congestion probability and traffic flow parameters, and determine the warning speed and warning flow; The traffic flow parameters include flow, density and speed; The start-stop channel decision problem is converted into a Markov decision process, and the state space, action space and state transition probability of the Markov decision process are determined; The Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision.
2. The start-stop channel decision method based on reinforcement learning according to claim 1, characterized in that: The traffic flow interruption probability model is constructed to describe the relationship between the congestion probability and the traffic flow parameters, and to determine the warning vehicle speed and warning flow, including: The Brilon model is used to construct the traffic flow interruption probability model, which is expressed as: ; in: Indicates flow rate, Indicates capacity, which refers to the maximum number of vehicles that can pass through a road or a certain section of the road in a unit of time. Indicates the observed flow Under these conditions, the capacity Less than or equal to The probability of is in the time interval The traffic volume observed at means The number of time intervals, The traffic flow reaches The number of time intervals, is the congestion interval A collection of; Read the flow, speed and timestamp data of each observation point, and iterate and calculate the data at each time point , and then obtain the congestion probability and traffic flow distribution; Draw speed and congestion probability distribution maps and flow and congestion probability distribution maps to determine warning vehicle speed and warning flow.
3. The start-stop channel decision method based on reinforcement learning according to claim 2, characterized in that: Before constructing the traffic flow interruption probability model, the method further includes: Acquiring traffic flow data; the traffic flow data includes flow, density and speed in the traffic flow; Using Origin software, the road congestion level RCL is used as the label ColorBar to color the corresponding data in the flow-speed distribution diagram in the traffic flow basic diagram. Combined with the visualization results, threshold analysis is performed to determine the congestion speed threshold.
4. The start-stop channel decision method based on reinforcement learning according to claim 1, characterized in that: The step of converting the start-stop channel decision problem into a Markov decision process and determining the state space, action space and state transition probability of the Markov decision process includes: Define the state space of the Markov decision process It is expressed as: ,in, represents the number of vehicles passing through the video observation point at time t, represents the traffic density of the video observation point at time t, represents the average vehicle speed at the video observation point at time t; Defining the action space of a Markov decision process It is expressed as: ,in, Indicates that the emergency lane is not enabled. Indicates that the emergency lane is activated; The state transition probability is determined by adopting a random state transition strategy, that is, after the algorithm takes an action, the environment gives an immediate reward and then transfers to the next state through a random state generator.
5. The start-stop channel decision method based on reinforcement learning according to claim 4 is characterized in that: The Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision, including: The Q-Learning algorithm is used to solve the Markov decision process. In the Q-Learning algorithm, the reward function is designed as: ; in, , , , in, represents the comprehensive reward function, Rewards for congestion relief, Enable penalties for non-bottleneck periods, To punish frequent switching, represents the congestion coefficient at time t, represents the congestion coefficient at time t+10, is the scaling factor, and are penalty coefficients, is the switch penalty function, It is expressed as: , in is a decay factor used to calculate the weight of each state change in the historical record, Display history Middle The state of the secondary emergency lane indicates that the emergency lane is Is it enabled? 1 means enabled, 0 means disabled. is an indicator function used to determine the Second and Whether the status of the secondary emergency lane has changed, if ,but ,otherwise, ; In the Q-Learning algorithm, the Q table is constructed as follows: the Q table is a two-dimensional array, in which the rows represent the state space of the Markov decision process and the columns represent the action space of the Markov decision process; In the Q-Learning algorithm, actions are selected using an ε-greedy strategy.
6. The method for starting and stopping channels based on reinforcement learning according to claim 5, characterized in that: The Markov decision process is solved with the warning vehicle speed and the warning flow rate as constraints to obtain the start-stop channel decision, including: The observed state space data is compressed by a data compression algorithm to reduce the state space to 10 3 Magnitude, as the S* input to the Q-learning algorithm; The Q-learning algorithm observes the current state of the environment S*, selects actions based on the ε-greedy algorithm and the Q table, and outputs the actions to the environment; The environment takes the current S and action as input, outputs the congestion level Cl, and based on the congestion level Cl based on the reward function Calculate rewards; Update Q value; Repeat the above steps until the maximum action step is reached; Repeat multiple rounds until the maximum number of rounds is reached.
7. The method for starting and stopping channels based on reinforcement learning according to claim 6, characterized in that: The updating Q value comprises: Based on the Bellman equation, the Q value in the Q table is updated as follows: ; in, is the current Q value when taking action A in state S, is the learning rate, is the discount factor, is the new state reached after executing action A, It is a new state The maximum Q value of all possible actions.
8. The method for starting and stopping channels based on reinforcement learning according to claim 7, characterized in that: In the solution process, The learning rate The initial setting is 0.1, which then decreases gradually over time; The discount factor Set between 0 and 1; The exploration rate ε has an initial value of 1 and gradually decreases over time.
9. A start-stop channel decision system based on reinforcement learning, characterized in that: For implementing the start-stop channel decision method based on reinforcement learning according to any one of claims 1 to 8, the system comprises: Traffic flow interruption probability model building module, used to build a traffic flow interruption probability model, characterize the relationship between congestion probability and traffic flow parameters, and determine the warning speed and warning flow; A Markov decision process conversion module, used to convert the start-stop channel decision problem into a Markov decision process, and determine the state space, action space and state transition probability of the Markov decision process; Reinforcement learning algorithm model building module, used to build a reinforcement learning algorithm model, including designing reward functions, building Q tables, and determining Q value update strategies and parameter settings; The calculation and output module is used to solve the reinforcement learning algorithm model to obtain the start-stop channel decision.
10. The start-stop channel decision system based on reinforcement learning according to claim 9, characterized in that: The system further comprises: Data preprocessing module, used to Acquiring traffic flow data; the traffic flow data includes flow, density and speed in the traffic flow; Using Origin software, the road congestion level RCL is used as the label ColorBar to color the corresponding data in the flow-speed distribution diagram in the traffic flow basic diagram. Combined with the visualization results, threshold analysis is performed to determine the congestion speed threshold.
Citation Information
Cited By
Method for deciding whether emergency lanes are opened or not on expressway under large-flow road condition
CN120708411A
Highway hard shoulder management and control method and system oriented to traffic toughness
CN121545349A