Three-dimensional road network traffic guidance and emergency rescue cooperative management method for underground road

By establishing a visual energy field driving behavior decision-making model and multi-agent simulation, combined with a coupled optimization model, traffic flow dispersion and emergency resource optimization are achieved, solving the problem of secondary accident prevention in complex underground road traffic accidents and improving the operational resilience of the underground road network.

CN117373243BActive Publication Date: 2026-04-28CHINA MERCHANTS CHONGQING COMM RES & DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MERCHANTS CHONGQING COMM RES & DESIGN INST
Filing Date
2023-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Complex underground roads are prone to traffic accidents under high traffic volume scenarios, especially fires caused by electric vehicles. Existing fire-fighting facilities are unable to effectively extinguish these fires, leading to escalation of traffic accidents and affecting the operational performance of underground roads.

Method used

By establishing a driving behavior decision-making model based on visual energy fields, and combining multi-agent simulation and coupled optimization models, traffic flow can be dispersed and emergency resources can be optimized to prevent secondary accidents and improve the resilience of the transportation system.

Benefits of technology

Effectively disperse traffic flow, prevent secondary accidents, maintain the normal operation of the underground road network, and enhance the resilience of the transportation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117373243B_ABST
    Figure CN117373243B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional road network traffic guidance and emergency rescue cooperative management method of underground roads, and comprises the following steps: S1, a driving behavior decision model considering surrounding vehicles and underground road environment is established based on visual energy field; S2, a three-dimensional road network traffic state evolution is predicted based on the driving behavior decision model through a traffic accident scene driven multi-agent simulation; S3, a coupling optimization model of road network signal control path guidance and resource allocation dynamic scheduling is established considering the time and space characteristics of secondary accident risk for the purpose of dispersing traffic flow and preventing secondary accidents; and S4, the three-dimensional road network traffic of the region where the underground roads are located under the traffic event is managed and controlled based on the coupling optimization model. The application can realize the cooperation of traffic flow dispersion and secondary accident prevention, and ensure the normal operation of the three-dimensional road network traffic of the region where the underground roads are located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of road traffic, specifically to a method for coordinated management of traffic guidance and emergency rescue in underground road networks. Background Technology

[0002] Underground roads, as an important means of transportation development and expansion in the high-density core areas of megacities, are gradually evolving into a systematic and networked system in the form of complex tunnels with multiple access points and ring-radial configurations. In terms of traffic operation functions, complex underground roads and elevated roads at ground level play similar roles. However, the unique road environment of underground roads leads to differences in their traffic characteristics.

[0003] First, to avoid underground spaces already occupied by buildings and subways, the alignment of complex underground road curves and merging / diversion zones is often restricted. Furthermore, changes in lighting and the sidewall effect within tunnels create visual strain and a sense of oppression for drivers, leading to frequent minor traffic accidents and other emergencies. Second, after an emergency occurs, due to changes in lighting and the sidewall effect, drivers exhibit different following and lane-changing behaviors in underground roads compared to above-ground areas, making secondary accidents highly likely in high-traffic scenarios. With the increasing popularity of electric vehicles, if an electric vehicle is involved in a traffic accident and its battery is damaged, it can quickly catch fire. Existing fire-fighting facilities within tunnels are often insufficient to effectively extinguish battery fires, and the semi-enclosed nature of tunnels can potentially cause large-scale fires, escalating the original traffic accident. Therefore, traffic accidents on complex underground roads, especially in high-traffic scenarios, can easily escalate into multi-vehicle accidents or fires, leading to a prolonged and severe degradation of the operational performance of underground roads.

[0004] Therefore, in response to traffic accidents on complex underground roads, a collaborative management method for traffic guidance and emergency rescue in the three-dimensional road network of underground roads is needed, which can both disperse traffic flow and prevent secondary accidents, and maintain the traffic operation status of the three-dimensional road network in the area where the underground roads are located. Summary of the Invention

[0005] In view of this, the purpose of this invention is to overcome the defects in the prior art and provide a method for coordinated management of traffic guidance and emergency rescue in the three-dimensional road network of underground roads, which can achieve coordinated traffic flow dispersion and prevention of secondary accidents, and ensure the normal operation of the three-dimensional road network in the area where the underground roads are located.

[0006] The method for coordinated management of traffic guidance and emergency rescue in a three-dimensional road network for underground roads according to the present invention includes the following steps:

[0007] S1. Based on the visual energy field, establish a driving behavior decision model that takes into account the surrounding vehicles and underground road environment;

[0008] S2. Through multi-agent simulation driven by traffic accident scenarios, based on driving behavior decision-making models, predict the evolution of traffic state in three-dimensional road networks;

[0009] S3. To address the issues of dispersed traffic flow and the prevention of secondary accidents, and considering the spatiotemporal characteristics of secondary accident risks, a coupled optimization model for road network signal control path guidance and dynamic scheduling of resource allocation is established.

[0010] S4. Based on the aforementioned coupled optimization model, traffic control is implemented in the three-dimensional road network area where underground roads are located during traffic incidents.

[0011] Furthermore, step S1 specifically includes:

[0012] By combining the driver's visual brightness and the tunnel sidewalls, a ground-state field strength model of the road environment in different sections is established;

[0013] Based on the dynamic field strength of traffic flow and the ground state field strength of road environment formed by trajectory data, the driver's visual energy field is quantified to form a decision model for tunnel following and lane changing behavior.

[0014] The vehicle trajectory in natural and simulated driving is discretized in time and space, and the following and lane-changing behavior decision variables are used to form driving process state-action sequential decision data.

[0015] A driving behavior decision model is constructed using a combined approach of forward and inverse reinforcement learning.

[0016] Initialize a random policy, sample in the road environment and agent simulation model, and merge the sampled trajectory with natural and simulated driving decision data to achieve the inverse reinforcement learning process;

[0017] A visual energy field function for an agent is generated using a deep neural network. The agent's policy is then updated based on the obtained visual energy field function. This process is iterated continuously. By comparing and evaluating simulated sampling data with natural and simulated driving behavior decision data, an iteration termination rule is formed.

[0018] Furthermore, step S2 specifically includes:

[0019] To address the spatiotemporal variations in traffic demand on complex underground roads, a multi-agent simulation platform is used to construct different traffic scenarios. Based on a driving behavior decision-making model, operational strategies for traffic accidents, road network traffic guidance, signal control, and emergency resource scheduling are implemented to simulate the evolution of road network traffic flow. Among these, signal control includes signal control at at-grade intersections and signal control for underground road ramps.

[0020] Furthermore, a coupled optimization model for road network signal control path guidance and dynamic scheduling of resource allocation is established, specifically including:

[0021] Based on the evolution of traffic conditions in complex underground road networks under traffic accidents, and combining the signal control parameters of regional road network intersections and underground road merging zones with the intersection turning ratio under path guidance, a multi-objective programming model is established to address the total delay of road network segments and the travel time in merging zones.

[0022] Based on the evolution of traffic conditions in complex underground road networks under traffic accidents, the average value of the road network conditions is obtained. With constraints such as roadside parking capacity for road rescue, fire-fighting resources, and medical resources, and with the goal of resource allocation consumption, a static planning model for the spatial allocation of rescue vehicles in the first stage is established.

[0023] Based on the stochastic demand generated by the dynamic evolution of road network status and the spatiotemporal distribution of secondary accident probability, and given the spatial configuration of road rescue vehicles, fire fighting and medical resources, a stochastic dynamic programming model for the second stage of resource scheduling is established with the time to reach the accident location as the objective.

[0024] The static planning model for the spatial configuration of rescue vehicles in the first stage is combined with the stochastic dynamic planning model for the scheduling of various resources in the second stage to obtain a two-stage planning model.

[0025] Different weight coefficients are assigned to the objective functions of the multi-objective programming model and the two-stage programming model, and the weights are summed to obtain the configured objective function. Based on the different types of traffic accidents, coupling constraints are set according to the constraints in the multi-objective programming model and the two-stage programming model. The configured objective function is used as the co-objective function of the coupled optimization model, and the coupling constraints are used as the constraints of the coupled optimization model to form the coupled optimization model.

[0026] Furthermore, the objective function of the multi-objective programming model includes a first objective function and a second objective function;

[0027] The first objective function is:

[0028]

[0029] in, Indicates total delays in the ground road network; It is the set of signal control nodes of the ground road network, and the random demand loaded on the road network during time period t is Λ. t ;k represents the decision-making stage number, k0 is the initial stage, and the traffic state and control strategy of stage k are represented by x respectively. k and u k ;x k+1 x under demand Λt k and u k Decision, i.e., x k+1 ~P(xk u k ,|Λ t The transformation matrix P is derived from the cellular transport model (CTM) of traffic flow; d k Let γ be the total network delay in stage k, where K represents the control stage of the plan. k This is the discount factor. Represents the mathematical expectation;

[0030] The second objective function is:

[0031]

[0032] Where TTS represents the total travel time of the complex underground road; k represents the cell of the complex underground road, l k ρ k (t) and q k (t) represents the length of cell k, the traffic density in time period t, and the queue length, respectively; n is the total number of cells; Δt is the simulation step size; and T is the total time period.

[0033] Furthermore, the objective function of the two-stage planning model includes a third objective function;

[0034] The third objective function is:

[0035]

[0036] Where Z represents the sum of emergency resource allocation consumption and minimum emergency rescue time; V S Let x represent the set of candidate locations for emergency resource points. i and d i Γ1 and Γ2 represent the quantity of emergency resources stored and the procurement consumption at node i in the road network, respectively; Q(x, Γ1, Γ2) represents the minimum emergency rescue time, and Γ1 and Γ2 are parameters that control road interruption and traffic demand uncertainty in the road network, respectively.

[0037] Furthermore, the collaborative objective function C is determined according to the following formula:

[0038]

[0039] Where ω1, ω2 and ω3 are weighting coefficients.

[0040] Furthermore, step S4 specifically includes:

[0041] Based on traffic flow simulation, a traffic environment of a three-dimensional road network in a complex underground road area is constructed to simulate the traffic flow evolution after a traffic accident. Intelligent agents in the environment are established using traffic lights and rescue vehicles to realize the interaction between intelligent agents and the environment.

[0042] Based on the Traffic Conflict Index (TTC), secondary accidents are randomly generated to simulate emergency resource allocation and random dynamic scheduling. A deep Q-network is established with the total delay of nodes and road segments, the probability of traffic accidents based on TTC, and the time for rescue vehicles to arrive at the accident location as Q values. The quantification of the probability of accidents can realize the reliability design of road network signal control and path guidance.

[0043] Within the inner loop of a multi-agent deep Q-network, policy optimization is performed based on the value function. A flexible actor-critic framework is adopted, leveraging the ease of parameter sharing in deep neural networks to collaboratively train the actor and critic modules.

[0044] Based on experience replay, critics estimate the value function by evaluating the quality of the adopted strategies; actors update strategy parameters and take actions based on the information provided by critics, achieving highly robust road network operation and control decision outputs; during the retrospective evaluation of the value function, a coupled optimization model is used for rapid estimation, and the road network operation and control strategies are retrospectively updated.

[0045] The beneficial effects of this invention are as follows: This invention discloses a method for coordinated management of traffic guidance and emergency rescue in an underground road network. By optimizing regional road network traffic signal control and route guidance to disperse traffic flow, and by optimizing the allocation and dynamic scheduling of emergency resources to prevent secondary accidents, the two are implemented in synergy, thereby maintaining traffic operation in the complex underground road network under traffic incidents and improving the operational resilience of the three-dimensional transportation system. Attached Figure Description

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0047] Figure 1 This is a schematic diagram of the collaborative management method of the present invention;

[0048] Figure 2 (a) is a schematic diagram of the underground roadway environment according to the present invention;

[0049] Figure 2 (b) is a schematic diagram of the coupling of the strategy planning path of the present invention at shared road segments and nodes;

[0050] Figure 3 (a) is a schematic diagram showing the positional relationship between the tunnel sidewall and the vehicle according to the present invention;

[0051] Figure 3 (b) is a schematic diagram of the dynamic energy field of traffic flow in the following scenario of the present invention;

[0052] Figure 3 (c) is a schematic diagram of the dynamic energy field of traffic flow in the lane change scenario of the present invention;

[0053] Figure 4 This is a schematic diagram of the driving behavior decision-making model framework of the present invention, which combines forward and reverse reinforcement learning.

[0054] Figure 5 This is a schematic diagram illustrating the construction of a multi-agent simulation scenario and the inference of traffic conditions according to the present invention;

[0055] Figure 6 This is a schematic diagram illustrating the parameter analysis of the coupling optimization model of the present invention;

[0056] Figure 7 This is a schematic diagram of the data and model hybrid-driven three-dimensional road network operation and control of the present invention. Detailed Implementation

[0057] The present invention will be further described below with reference to the accompanying drawings, as shown in the figures:

[0058] The method for coordinated management of traffic guidance and emergency rescue in a three-dimensional road network for underground roads according to the present invention includes the following steps:

[0059] S1. Based on the visual energy field, establish a driving behavior decision model that takes into account the surrounding vehicles and underground road environment;

[0060] S2. Through multi-agent simulation driven by traffic accident scenarios, based on driving behavior decision-making models, predict the evolution of traffic state in three-dimensional road networks;

[0061] S3. To address the issues of dispersed traffic flow and the prevention of secondary accidents, and considering the spatiotemporal characteristics of secondary accident risks, a coupled optimization model for road network signal control path guidance and dynamic scheduling of resource allocation is established.

[0062] S4. Based on the aforementioned coupled optimization model, traffic control is implemented in the three-dimensional road network area where underground roads are located during traffic incidents.

[0063] like Figure 2 As shown in (a), at the traffic flow level, the following and lane-changing behaviors of drivers under the influence of road environment factors such as lighting changes, curves, and sidewalls in underground roads differ significantly from those on surface roads, thus affecting the modeling of traffic flow evolution and control strategies; such as Figure 2 As shown in (b), the present invention takes into account the risks of secondary accidents, electric vehicle fires, etc., as well as the coupling of traffic flow paths under signal control path guidance and emergency rescue vehicle paths under dynamic scheduling at shared road sections and nodes, thereby achieving the synergy of traffic flow dispersion and secondary accident prevention.

[0064] In this embodiment, in step S1, as follows: Figure 3As shown, the visual energy field is divided into the environmental ground state and the traffic flow dynamics energy field. The environmental ground state mainly represents the energy state of the road infrastructure, whose energy distribution is independent of the traffic flow state and does not change over time. The traffic flow dynamics mainly represent the energy state of surrounding vehicles, whose energy distribution is affected by factors such as the relative position of vehicles, their relative motion state, and traffic density; this energy level changes over time.

[0065] By combining driver visual brightness and tunnel sidewall data, a ground-state field strength model of the road environment is established for different sections, such as curves and merging / diverging zones. Based on the dynamic field strength of traffic flow and the ground-state field strength of the road environment formed from trajectory data, the driver's visual energy field is quantified, and based on this, a decision-making model for tunnel following and lane-changing behaviors is formed.

[0066] We utilize inverse reinforcement learning to learn and train the visual energy field function for self-generated complex underground road driving behavior, and construct a behavioral decision-making model framework:

[0067] First, the vehicle trajectory in natural and simulated driving is discretized in time and space, and behavioral decision variables such as following and changing lanes are used to form state-action sequential decision data for the driving process.

[0068] Then, a driving behavior decision-making model framework is constructed using a collaborative approach of forward and inverse reinforcement learning, such as... Figure 4 As shown.

[0069] By initializing a random policy, sampling is performed in the road environment and the agent simulation model. The sampled trajectories and natural and simulated driving decision data are then merged and used together to achieve the inverse reinforcement learning process.

[0070] A deep neural network is used to generate a visual energy field function for the intelligent agent. Based on this function, the agent's strategy is updated, and the process is iterated continuously. An iteration termination rule is established by comparing simulated sampling data with natural and simulated driving behavior decision data. The intelligent agent can be a physical entity, such as a robot or an autonomous vehicle, or a virtual entity, such as a computer program or a virtual assistant. Existing technologies are used for the intelligent agent, and its specific type and complexity can be selected based on actual working conditions.

[0071] In this embodiment, in step S2, considering the spatiotemporal variations in traffic demand on complex underground roads, different traffic scenarios (such as...) are constructed using a multi-agent simulation platform. Figure 5 As shown in the figure, based on the driving behavior decision model, the system simulates the evolution of road network traffic flow by considering operational strategies for traffic accidents, road network traffic guidance, signal control, and emergency resource scheduling. Among these, signal control includes signal control at at-grade intersections and signal control for underground road ramps.

[0072] In this embodiment, in step S3, as follows: Figure 6As shown, based on the traffic state evolution of complex underground road networks under traffic accidents, and combining the signal control parameters of regional road network intersections and underground road merging zones with the intersection turning ratio under path guidance, a multi-objective programming model is established to address the total delay of road segments and the travel time of merging zones. The resulting optimal total delay of road segments and the travel time of merging zones provide the foundation for constructing the value function in the backtracking evaluation of deep reinforcement learning control strategies.

[0073] Based on the evolution of traffic conditions in complex underground road networks under traffic accidents, the average value of the road network conditions is obtained. With constraints such as roadside parking capacity for road rescue, fire and medical resources, and the sum of resource allocation consumption and expected time to reach the traffic accident location as the objective function, a static programming model for the spatial allocation of rescue vehicles in the first stage is established.

[0074] Based on the stochastic demand generated by the dynamic evolution of road network status and the spatiotemporal distribution of secondary accident probabilities, and given the spatial configuration of road rescue vehicles, fire fighting and medical resources, a stochastic dynamic programming model for the second stage of resource scheduling (stochastic dynamic programming model for multiple resource scheduling) is established with the time to reach the accident location as the objective function. The obtained optimal rescue time provides the basis for constructing the value function of the deep reinforcement learning control strategy.

[0075] The static planning model for the spatial configuration of rescue vehicles in the first stage is combined with the stochastic dynamic planning model for the scheduling of various resources in the second stage to obtain a two-stage planning model.

[0076] Considering the priority weights of the two strategies, construct a collaborative optimization objective function: assign different weight coefficients to the objective function of the multi-objective programming model and the objective function of the two-stage programming model, and sum the weights to obtain the configured objective function;

[0077] Based on the different types of traffic accidents, coupling constraints are set according to the constraints in the multi-objective programming model and the two-stage programming model. The configured objective function is used as the collaborative objective function of the coupled optimization model, and the coupling constraints are used as the constraints of the coupled optimization model to form a coupled optimization model. Among them, for minor traffic accidents involving electric vehicles, the priority of roads, medical care, and fire rescue needs to be considered; for other minor traffic accidents, the priority of regional traffic flow dispersion strategies needs to be considered.

[0078] The objective function of the multi-objective programming model includes a first objective function and a second objective function;

[0079] The first objective function is:

[0080]

[0081] in, Indicates total delays in the ground road network; It is the set of signal control nodes of the ground road network, and the random demand loaded on the road network during time period t is Λ. t ;k represents the decision-making stage number, k0 is the initial stage, and the traffic state and control strategy of stage k are represented by x respectively. k and u k ;x k+1 From demand Λ t x below k and u k Decision, i.e., x k+1 ~P(x k u k ,||Λ t The transformation matrix P is derived from the cellular transport model (CTM) of traffic flow; d k Let γ be the total network delay in stage k, where K represents the control stage of the plan. k This is the discount factor. Represents the mathematical expectation;

[0082] The underlying model for the above optimization problem is CTM (Cell Transfer Model):

[0083]

[0084]

[0085]

[0086] Among them, f ij (t), β ij (t) and These represent the traffic flow, turning ratio, and green light ratio of upstream road segment i and downstream road segment j connected by node m during time period t. λ i (t) represents the total flow through node m on road segment i. l represents the cell of road segment i, ρ l (t) and f l (t) represents the traffic density and outflow of cell l in time period t, respectively. Δt and Δx l Q represents the simulation step size and the length of cell l, respectively. l The passage capacity of cell l. Congestion density of cell (l+1). and Let represent the free flow velocity of cell l and the backward propagation velocity of congestion, respectively.

[0087] Calculation of delays:

[0088]

[0089]

[0090]

[0091]

[0092] d l (t), d i (t), d t,m and d t These represent cell delay, segment delay, node delay, and total network delay for time period t, respectively. m Let I represent the set of upstream road segments i, and let I represent the set of all road segments.

[0093] The second objective function is:

[0094]

[0095] Where TTS represents the total travel time of the complex underground road; k represents the cell of the complex underground road, l k ρ k (t) and q k (t) represents the length of cell k, the traffic density in time period t, and the queue length, respectively; n is the total number of cells; Δt is the simulation step size; and T is the total time period.

[0096] The underlying model for the above optimization problem is CTM (Cell Transfer Model):

[0097]

[0098] q k (t+1)=q k (t)+Δt·(w k (t)-r k (t))

[0099] Where, φ k (t), r k (t) and w k (t) represent the outflow, ramp inflow rate, and external inflow demand of cell k in time period t, respectively. β k This indicates the outgoing ratio of the outgoing ramp cell k.

[0100] The objective function of the two-stage planning model includes a third objective function;

[0101] The third objective function is:

[0102]

[0103] Where Z represents the sum of emergency resource allocation consumption and minimum emergency rescue time; VS Let x represent the set of candidate locations for emergency resource points. i and d i Γ1 and Γ2 represent the quantity of emergency resources stored and the procurement consumption at node i in the road network, respectively; Q(x, Γ1, Γ2) represents the minimum emergency rescue time, and Γ1 and Γ2 are parameters that control road interruption and traffic demand uncertainty in the road network, respectively.

[0104] Phase 1: Emergency Resource Allocation, Constraints:

[0105] Roadside parking capacity constraints for roadside assistance: ∑ i∈V f i y i ≤G

[0106] Constraints on medical and fire-fighting resources:

[0107] Among them, f i The fixed cost for constructing emergency resource storage node i. i The variable is 0-1. If i is selected as the emergency resource storage node, y i =1; conversely, y i =0. C i Let G be the storage capacity of emergency resource storage node i. Let G be the total investment in emergency resource allocation. Let V or N represent the set of road network nodes.

[0108] Phase Two: Minimizing Emergency Rescue Time

[0109]

[0110] Wherein, the first term represents the time of rescue for an existing demand, and the second term represents the time of rescue for a random demand. E or A represents the set of road network segments, f ij and c ij V represents the traffic flow and unit consumption for transporting emergency resources on road segment (i, j), respectively. d z represents the set of nodes generated by random demand. i and s i Let represent the demand and compensation consumption at node i generated by the random demand, respectively.

[0111] The constraints for the second-stage optimization problem are as follows:

[0112] Secondary accidents generate random dynamic demands:

[0113]

[0114]

[0115] Here, E1 represents the set of road sections with a high risk of secondary accidents.ij For a variable of 0-1, if a secondary accident occurs on road segment (i, j), r ij =1; conversely, r ij =0.

[0116] Description of the dynamic characteristics of macroscopic traffic flow in the road network:

[0117]

[0118]

[0119] in, and Let represent the traffic flow, available stock, demand, and random demand generation rate of the k-th type of emergency resource transportation traffic flow on road segment (j, i) or node i during time period t, respectively. D is the duration of the time period. The number of rescue trailers prepared in advance at node i, This represents the remaining proportion of emergency resource transportation traffic flow for the kth type under event type s.

[0120] A collaborative objective function can be constructed by summing the weights or by normalizing each objective function separately. The collaborative objective function C is determined according to the following formula:

[0121]

[0122] Where ω1, ω2, and ω3 are weighting coefficients. Appropriate weight values ​​can be set for the three objective functions according to the actual working conditions.

[0123] In this embodiment, in step S4, as follows: Figure 7 As shown, a traffic environment of a three-dimensional road network in a complex underground road area is constructed based on traffic flow simulation. The traffic flow evolution state after a traffic accident is simulated. Intelligent agents in the environment are established using traffic lights, rescue vehicles, etc., to realize the interaction between intelligent agents and the environment.

[0124] Based on the Traffic Conflict Index (TTC), secondary accidents are randomly generated to simulate emergency resource allocation and random dynamic scheduling. A deep Q-network is established with the total delay of nodes and road segments, the probability of traffic accidents based on TTC, and the time for rescue vehicles to arrive at the accident location as Q values. The quantification of the probability of accidents enables the reliability design of road network signal control (including underground road ramp control) and path guidance.

[0125] Within the inner loop of a multi-agent deep Q-network, policy optimization is performed based on the value function. A flexible actor-critic framework is adopted, leveraging the ease of parameter sharing in deep neural networks to collaboratively train the actor and critic modules.

[0126] Based on experience replay, critics estimate the value function by evaluating the quality of the adopted strategies; actors update strategy parameters and take actions based on the information provided by critics, thereby achieving highly robust road network operation and control decision outputs; during the backtesting evaluation of the value function, a coupled optimization model is used for rapid estimation, thereby backtesting the road network operation and control strategies and achieving a hybrid data and model-driven approach.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for collaborative management of traffic guidance and emergency rescue in a three-dimensional road network of underground roads, characterized in that: Includes the following steps: S1. Based on the visual energy field, establish a driving behavior decision-making model that considers surrounding vehicles and the underground road environment; step S1 specifically includes: By combining the driver's visual brightness and the tunnel sidewalls, a ground-state field strength model of the road environment in different sections is established; Based on the dynamic field strength of traffic flow and the ground state field strength of road environment formed by trajectory data, the driver's visual energy field is quantified to form a decision model for tunnel following and lane changing behavior. The vehicle trajectory in natural and simulated driving is discretized in time and space, and the following and lane-changing behaviors are used as decision variables to form state-action sequential decision data for the driving process. A driving behavior decision model is constructed using a combined approach of forward and inverse reinforcement learning. Initialize a random policy, sample in the road environment and agent simulation model, and merge the sampled trajectory with natural and simulated driving decision data to achieve the inverse reinforcement learning process; The visual energy field function of the agent is generated by using a deep neural network. The agent's policy is updated based on the obtained visual energy field function. The process is iterated continuously. The simulation sampling is compared and evaluated with natural and simulated driving behavior decision data to form an iteration termination rule. S2. Through multi-agent simulation driven by traffic accident scenarios, based on a driving behavior decision model, predict the evolution of traffic state in a three-dimensional road network; step S2 specifically includes: To address the spatiotemporal variations in traffic demand on complex underground roads, a multi-agent simulation platform is used to construct different traffic scenarios. Based on a driving behavior decision-making model, operational strategies are developed for traffic accidents, road network traffic guidance, signal control, and emergency resource scheduling to simulate the evolution of road network traffic flow. Among these, signal control includes signal control at at-grade intersections and signal control for underground road ramps. S3. To address the issues of dispersed traffic flow and the prevention of secondary accidents, and considering the spatiotemporal characteristics of secondary accident risks, a coupled optimization model for road network signal control path guidance and dynamic scheduling of resource allocation is established. Establish a coupled optimization model for road network signal control path guidance and dynamic scheduling of resource allocation, specifically including: Based on the evolution of traffic conditions in complex underground road networks under traffic accidents, and combining the signal control parameters of regional road network intersections and underground road merging zones with the intersection turning ratio under path guidance, a multi-objective programming model is established to address the total delay of road network segments and the travel time in merging zones. Based on the evolution of traffic conditions in complex underground road networks under traffic accidents, the average value of the road network conditions is obtained. With constraints such as roadside parking capacity for road rescue, fire-fighting resources, and medical resources, and with the goal of resource allocation consumption, a static planning model for the spatial configuration of rescue vehicles in the first stage is established. Based on the stochastic demand generated by the dynamic evolution of road network status and the spatiotemporal distribution of secondary accident probability, and given the spatial configuration of road rescue vehicles, fire fighting and medical resources, a stochastic dynamic programming model for the second stage of resource scheduling is established with the time to reach the accident location as the objective. The static planning model for the spatial configuration of rescue vehicles in the first stage is combined with the stochastic dynamic planning model for the scheduling of various resources in the second stage to obtain a two-stage planning model. Different weight coefficients are assigned to the objective functions of the multi-objective programming model and the two-stage programming model, and the weights are summed to obtain the configured objective function. Based on the different types of traffic accidents, coupling constraints are set according to the constraints in the multi-objective programming model and the two-stage programming model. The configured objective function is used as the co-objective function of the coupled optimization model, and the coupling constraints are used as the constraints of the coupled optimization model to form the coupled optimization model. The objective function of the multi-objective programming model includes a first objective function and a second objective function; The first objective function is: ; in, Indicates total delays in the ground road network; It is a set of ground road network signal control nodes, during the time period The random demand loaded onto the road network is ; Indicates the decision-making stage number. For the initial stage, stage The traffic conditions and control strategies are respectively represented as: and ; From demand Below and Decision, that is Transformation matrix It is derived from the cellular transport model (CTM) of traffic flow; For the stage Overall road network delays Indicates the control phase of the plan. This is the discount factor. Represents the mathematical expectation; The second objective function is: ; in, The total travel time for the complex underground roads; A cell representing a complex underground road system. , and They represent cells respectively Length, in time period Traffic density and queue length; The total number of cells; This is for simulating step size; Total time period; The objective function of the two-stage planning model includes a third objective function; The third objective function is: ; in, The sum of emergency resource allocation consumption and minimum emergency rescue time; This represents the set of candidate locations for emergency resource points. and These respectively represent the nodes in the road network. The quantity of emergency resources stored and their procurement consumption; Indicates the minimum emergency rescue time. and These are parameters for controlling road network disruptions and traffic demand uncertainty; S4. Based on the aforementioned coupled optimization model, traffic in the three-dimensional road network area where underground roads are located is controlled under traffic incidents.

2. The method for coordinated management of traffic guidance and emergency rescue in an underground road network according to claim 1, characterized in that: The collaborative objective function is determined according to the following formula. : ; in, , as well as These are the weighting coefficients.

3. The method for coordinated management of traffic guidance and emergency rescue in an underground road network according to claim 1, characterized in that: Step S4 specifically includes: Based on traffic flow simulation, a traffic environment of a three-dimensional road network in a complex underground road area is constructed to simulate the traffic flow evolution after a traffic accident. Intelligent agents in the environment are established using traffic lights and rescue vehicles to realize the interaction between intelligent agents and the environment. Based on the Traffic Conflict Index (TTC), secondary accidents are randomly generated to simulate emergency resource allocation and random dynamic scheduling. A deep Q-network is established with the total delay of nodes and road segments, the probability of traffic accidents based on TTC, and the time for rescue vehicles to arrive at the accident location as Q values. The quantification of the probability of accidents enables the reliable design of road network signal control and path guidance. Within the inner loop of a multi-agent deep Q-network, policy optimization is performed based on the value function. A flexible actor-critic framework is adopted, leveraging the ease of parameter sharing in deep neural networks to collaboratively train the actor and critic modules. Based on experience replay, critics estimate the value function by evaluating the quality of the adopted strategies; actors update strategy parameters and take actions based on the information provided by critics, achieving highly robust road network operation and control decision outputs; during the retrospective evaluation of the value function, a coupled optimization model is used for rapid estimation, and the road network operation and control strategies are retrospectively updated.

Citation Information

Patent Citations

  • Highway tunnel operation management and control system

    CN112885100A

  • Tunnel digital twinning scene construction method and computer equipment

    CN113538863A