A Collaborative Control Method for Merging Zones Based on RainbowDQN and Driving Risk Field
By using a collaborative control method combining Rainbow DQN and traffic risk field, the signal control in the merging zone of expressways is dynamically optimized, solving the problem of mismatch between logic and real-time risk in existing technologies, and improving traffic efficiency and safety.
Patent Information
- Application Number
- CN202511027393.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing signal control methods for merging zones on expressways suffer from a mismatch between logic and real-time risks, leading to a decline in the quality of mixed traffic flow and an increase in merging risks.
A collaborative control method based on Rainbow DQN and driving risk field is adopted. By acquiring traffic operation data, a state space, action space and reward function are constructed to dynamically control the main line speed limit and ramp signals, and to collaboratively control the driving trajectory of CAV vehicles. Traffic environment perception information is acquired by using millimeter-wave radar, pressure sensors and other devices to construct driving risk field, refine the state and action space of the agent, and design a reward function to optimize signal control.
It improves traffic efficiency and safety in the merging zone of expressways, reduces merging risks, and ensures the overall traffic quality of mixed traffic flows.
Smart Images

Figure CN120748204B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of traffic safety technology, specifically a merging zone collaborative control method based on Rainbow DQN and driving risk field. Background Technology
[0002] Merging zones on expressways are prone to traffic conflicts and congestion due to the complex merging behavior and significant speed differences of vehicles. This seriously threatens urban traffic efficiency and road safety, and traffic accidents can cause substantial economic losses. Furthermore, the mixed traffic flow of connected vehicles (CAVs) and conventional vehicles (HVs) further complicates vehicle interactions in merging zones, making control methods that target only CAVs or traffic equipment less effective. Therefore, optimizing signal control methods in merging zones on expressways is crucial for improving traffic safety, capacity, and reducing economic losses.
[0003] Existing signal control methods for merging zones on expressways reduce the overall traffic quality of mixed traffic flows and increase merging risks due to a mismatch between control logic and real-time risks. Summary of the Invention
[0004] The purpose of this invention is to address the problem that existing expressway merging zone signal control methods reduce the overall traffic quality of mixed traffic flow and increase merging risks due to the mismatch between control logic and real-time risks. This invention provides a merging zone collaborative control method based on Rainbow DQN and driving risk field.
[0005] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0006] A collaborative control method for merging zones based on Rainbow DQN and driving risk field includes the following steps:
[0007] Step 1: Obtain traffic operation data for the main line;
[0008] Step 2: Obtain the vehicle queue length n of the ramp. fleet ;
[0009] Step 3: Combine traffic operation data and vehicle queue length n fleet The features are encoded separately and then input into the Rainbow DQN model to obtain the control action output by the model. The merging zone is then coordinated and controlled based on the control action.
[0010] The Rainbow DQN model is trained using a defined state space S, action space A, and reward function r. The state space S is represented as follows:
[0011] S = (s1, s2, s3)
[0012] Among them, s1 is the state space of the mainline traffic flow, s2 is the state space of the ramp, and s3 is the state space of the mainline vehicles.
[0013] The action space A is represented as follows:
[0014] A = (A1, A2, A3)
[0015] Among them, A1 is the action space of mainline variable speed limit control, A2 is the action space of ramp signal control, and A3 is the action space of mainline CAV vehicle control.
[0016] The reward function r is expressed as:
[0017] r = w s r s +w f r f +w acd r acd +w risk r risk +w balence r balence +w coop r coop
[0018] Where, r s Main speed reward factor, r f r is the reward factor for ramp queue length. acd r is the accident penalty factor. risk As a driving risk penalty factor, r balence Main turn balance control reward factor, r coop For CAV vehicle collaborative reward factor, w s w f w acd w risk w balence w coop These are the influence coefficients corresponding to each factor.
[0019] Furthermore, the traffic operation data includes the average traffic flow speed v. avg The speed v of the vehicle along the lane i The acceleration a of the vehicle along the lane direction i The vehicle's x-coordinate i The vehicle's ordinate y i And the vehicle's mass m i .
[0020] Furthermore, the state space s1 of the linear traffic flow is represented as:
[0021] s1=(v avg ,v max ,n total ,n coop )
[0022] The state space s2 of the ramp is represented as follows:
[0023] s2=(n fleet ,t green )
[0024] The state space s3 of the mainline vehicle is represented as follows:
[0025] s3=(s v1 ,s v2 ,s v3 ,......s vi ,......,s vn )
[0026] s vi =(v i ,a i ,x i ,y i ,R i )
[0027] Among them, v avg v represents the average operating speed across different lanes, n is the maximum number of lanes, and v max The speed limit for the main lane, n total The total number of vehicles on the main line, n coop The number of vehicles responding to coordination commands on the main line, n fleet t represents the length of the queue at the ramp. green The green light duration for the ramp signal, s vi Let R be the set of operating states of a single vehicle, where i is the vehicle number and R is the set of operating states of a single vehicle. i This is the driving risk index.
[0028] Furthermore, the action space A1 of the mainline variable speed limit control is represented as follows:
[0029] A1=(V min V1, V2, ... V j ......,V max )
[0030] The action space A2 for ramp signal control is represented as:
[0031] A2=(t min ,t1,t2,......t k ......,t max )
[0032] The action space A3 of the mainline CAV vehicle intelligent agent is represented as follows:
[0033] A3 = (change, keep, speedup)
[0034] V j =V min +10(j-1)
[0035] t k =t min +5(k-1)
[0036] Where j = 0, 1, 2, ..., k = 0, 1, 2, ..., V min Minimum driving speed on the main line, V max The maximum speed limit for the main line design, t min The minimum green light duration, t max The maximum green light duration is indicated by "change" to move to the inner lane, "keep" to maintain the current driving state, and "speedup" to accelerate to the current maximum speed limit on the main road.
[0037] Furthermore, the mainline speed reward factor r s Represented as:
[0038]
[0039] The ramp queue length reward factor r f Represented as:
[0040]
[0041] The accident penalty factor r acd Represented as:
[0042]
[0043] The driving risk penalty factor r risk Represented as:
[0044]
[0045] The main turn balance control reward factor r balence Represented as:
[0046]
[0047] The CAV vehicle collaborative reward factor r coop Represented as:
[0048]
[0049] Among them, vmax The current maximum speed limit for the main line, n max α is the maximum queue length of the ramp. road r is the speed bonus coefficient for the lane. f n is the reward factor for ramp queue length. avg ω represents the average queue length at the ramp. v The main line control influences the weighting coefficient, ω f φ is the weighting coefficient for the influence of ramp control. coop For collaborative weights, τ is the average delay for CAV vehicles to execute cooperative commands. coop This represents the delay tolerance coefficient.
[0050] Furthermore, the driving risk index R i Represented as:
[0051]
[0052] Among them, E a E represents the average field strength of driving risk. i Let i be the total risk of vehicle i while it is in motion.
[0053] Furthermore, the total risk E of vehicle i while it is in motion i Represented as:
[0054] E i =E snow +E locomotion +ω heavy R HV
[0055] Among them, E snow For strong driving potential energy, E locomotion R is the potential field strength of the vehicle. HV For areas with high vehicle risk, ω heavy The value is the heavy vehicle impact factor. When the target vehicle is a heavy vehicle, its value is 1, and when the target vehicle is not a heavy vehicle, its value is 0.
[0056] Furthermore, the field strength E of the driving potential energy field snow Represented as:
[0057]
[0058] Where, φ N φ represents the visibility of the road where the vehicle is located. P φ is the slope of the road where the vehicle is located. R This is the radius of the curve on the road where the vehicle is located. For standard road visibility, For standard road slope, Let ψ1, ψ2, and ψ3 be the standard road curve radius, and τ be the undetermined constant coefficients. static α is the critical threshold for safe distance. snow γ represents the influence coefficient of ice and snow in cold regions. speed v is the velocity influence coefficient. i Let a be the speed of vehicle i traveling along the road direction. i Let δ be the acceleration of vehicle i. a δ is the acceleration influence coefficient. snow The thickness of thin ice or snow. Let T1 be the potential energy field intensity factor generated by the lane divider, T2 be the potential energy field intensity factor generated by the boundary of the central median, and T3 be the potential energy field intensity factor generated by the road boundary line. Let q be an undetermined coefficient. 1-i Let q be the perpendicular distance from the center of mass of vehicle i to the nearest white dashed line in the lane where the vehicle is located. 2-i Let b be the vertical distance from the center of mass of vehicle i to the boundary of the central divider. i Let μ be the distance between the centroid of vehicle i and the road boundary line on the side where the vehicle is located, and σ1 and σ2 be the coefficients of variation. F The proportion of the potential energy field intensity generated by the road boundary line, t snow Let f(δ) be the road surface temperature. snow ,t snow ) as δ snow and t snow is the independent variable.
[0059] Furthermore, the vehicle kinematic potential field strength E locomotion Represented as:
[0060]
[0061] M i =m i (1.566×10 -14 ·v 6.687 +0.3345)
[0062] Among them, M i Let θ be the equivalent mass of vehicle i, (x0, y0) be the spatial coordinates of the target vehicle's center of mass, and θ be the equivalent mass of vehicle i. car Let ρ be the angle between a point around the vehicle and the vehicle's center of mass. driver For the driver coefficient, A p For driver characteristics, the value ranges from 0 to 1, λ loc β loc ξ loc The coefficients are undetermined, |l′| and |l| are intermediate variables, (x, y) are the vehicle's position coordinates, and α is the coefficient of the unknown coefficient. v These are undetermined parameters related to speed.
[0063] Furthermore, the heavy vehicle risk field R HV Represented as:
[0064]
[0065] Among them, M HV For heavy vehicle weight, M std For standard axle load mass, Q is the instantaneous acceleration of the loaded vehicle, Qco is the real-time CO emission rate, and Q crit For the CO safety threshold, α HV1 α HV2 α HV3 These are the weighting coefficients for each factor.
[0066] The beneficial effects of this invention are:
[0067] This application, by considering the mixed traffic flow including CAVs on the mainline of the merging zone, mainline traffic flow information, and ramp queuing information, refines the state space and action space of the cooperative control agent. It proposes a reward function that considers traffic efficiency, traffic safety, and cooperative interaction, thereby dynamically controlling the mainline speed limit, dynamically coordinating queuing vehicles on ramps, and guiding the driving trajectories of CAV vehicles on the mainline, effectively improving traffic efficiency and safety. It ensures the overall traffic quality of the mixed traffic flow and reduces merging risks. Attached Figure Description
[0068] Figure 1 This is a schematic diagram of the overall scheme of this application;
[0069] Figure 2 This is a schematic diagram of the hardware layout and control results in the embodiments of this application. Detailed Implementation
[0070] It should be noted that, where there is no conflict, the various embodiments disclosed in this application can be combined with each other.
[0071] Specific Implementation Method 1: The merging zone collaborative control method based on RainbowDQN and driving risk field described in this implementation method includes the following steps:
[0072] Step 1: Obtain traffic operation data for the main line;
[0073] Step 2: Obtain the vehicle queue length n of the ramp. fleet ;
[0074] Step 3: Combine traffic operation data and vehicle queue length n fleet The features are encoded separately and then input into the Rainbow DQN model to obtain the control action output by the model. The merging zone is then coordinated and controlled based on the control action.
[0075] The Rainbow DQN model is trained using a defined state space S, action space A, and reward function r. The state space S is represented as follows:
[0076] S = (s1, s2, s3)
[0077] Among them, s1 is the state space of the mainline traffic flow, s2 is the state space of the ramp, and s3 is the state space of the mainline vehicles.
[0078] The action space A is represented as follows:
[0079] A = (A1, A2, A3)
[0080] Among them, A1 is the action space of mainline variable speed limit control, A2 is the action space of ramp signal control, and A3 is the action space of mainline CAV vehicle control.
[0081] The reward function r is expressed as:
[0082] r = w s r s +w f r f +w acd r acd +w risk r risk +w balence r balence +w coop r coop
[0083] Where, r s Main speed reward factor, r f r is the reward factor for ramp queue length. acd r is the accident penalty factor. risk As a driving risk penalty factor, r balence Main turn balance control reward factor, r coop For CAV vehicle collaborative reward factor, w s w f w acd w risk w balence w coop These are the influence coefficients corresponding to each factor.
[0084] Global traffic environment perception information is obtained through roadside equipment and vehicle-mounted sensors:
[0085] Millimeter-wave radar, pressure sensors, and laser remote-sensing road condition sensors are deployed along the main road. When a vehicle enters the detection range, these devices monitor and store vehicle information in real time, collecting one month's worth of traffic data for the merging area of the expressway. Millimeter-wave radar captures vehicle motion information, providing real-time feedback on the vehicle's position in space and its relative distance to surrounding objects. Based on the radar's deployment location and the actual positions of lane and road boundaries, this data is transmitted to the traffic signal control center to calculate the vehicle's positional relationship with these boundaries. Pressure sensors detect the actual weight of vehicles. Laser remote-sensing road condition sensors utilize laser remote sensing technology and multispectral measurement principles. By emitting infrared light and analyzing the reflected light, they accurately measure the thickness of water, ice, and snow on the road surface, and can independently measure water and ice content, as well as road surface temperature. Loop sensors count n vehicles on the main road. total And calculate the average lane speed v avg The vehicle is equipped with speed sensors, acceleration sensors, and a coordinate positioning system, which can transmit real-time operating status data (v... i ,a i ,x i ,y i The data is transmitted to the roadside computing unit for unified processing, where v i a represents the speed of vehicle i along the lane direction. i Let x be the acceleration of vehicle i along the lane direction. i ,y i These are the x and y coordinates of vehicle i, respectively. Meanwhile, coil detectors positioned at fixed locations are used in the ramp area to obtain the ramp vehicle queue length n. fleet Then all the information is uploaded to the Roadside Unit (RSU), which transmits the data to the traffic signal control center for analysis and processing.
[0086] Constructing a driving risk field based on vehicle operation data:
[0087] Based on the vehicle operation data collected in step one, and combined with the influence of the cold-region ice and snow environment and factors such as lane dividing lines, road boundary lines, and heavy vehicle driving, a driving risk field is constructed to quantify driving risks in a fine-grained manner.
[0088] (1) Potential energy field of vehicle movement
[0089] The potential energy field strength of a vehicle can be expressed as:
[0090] E snow =(μ L L i +μ F F i )·Pr ·P s (1)
[0091] In the formula: E snow Let L be the driving potential energy field strength of vehicle i. i Let μ be the potential energy field intensity generated by the lane dividing line on vehicle i. L Let μ be the proportion of the potential energy field intensity generated by the lane dividing line, Fi be the potential energy field intensity generated by vehicle i at the road boundary line, and μ be the proportion of the potential energy field intensity generated by the lane dividing line. F P represents the proportion of the potential energy field intensity generated by the road boundary line. r P is an environmental static influencing factor. s These are environmental dynamic influencing factors.
[0092] Road conditions can be assessed by road visibility φ N Road slope φ P and road curve radius φ R These parameters can be obtained directly by consulting relevant literature for evaluation. The definition of environmental static impact factors is as follows:
[0093]
[0094] In the formula, τ static ψ1, ψ2, ψ3 are undetermined constant coefficients, and τ static ψ1, ψ2, ψ3>0, φ N Let φ be the road visibility (m) of vehicle i on the road it is traveling on. P The road gradient (%) of the road where vehicle i is traveling. R Let be the radius of the road curve on the road where vehicle i is traveling, in meters. Standard road visibility is measured in meters (m) under clear weather and sufficient sunlight. The standard road slope is the slope value of a straight road, expressed as a percentage. The standard road curve radius is taken as the road curve radius value of a straight road, in meters.
[0095] In cold regions, expressways are easily covered by snow in winter, and can even form thin ice surfaces, posing a significant driving risk to those traveling for work or leisure. Therefore, this invention considers the impact of icy and snowy environments on vehicle movement in cold regions and proposes dynamic environmental influencing factors:
[0096]
[0097] In the formula, P s As a dynamic environmental influencing factor, α snow The influence coefficient of ice and snow in cold regions, μ f γ is the road surface friction coefficient. speedv is the velocity influence coefficient. i Let δ be the speed of vehicle i traveling along the road direction. a Let a be the acceleration influence coefficient. i Let be the acceleration of vehicle i.
[0098] This application proposes using cold-region environmental characteristics (ice and snow thickness and temperature) to characterize the road surface friction coefficient in cold regions, thereby expressing the driving risks in cold regions:
[0099]
[0100] In the formula, δ snow The thickness of thin ice or snow, in mm, t snow Road surface temperature, °C These are coefficients to be determined.
[0101] The potential energy field generated by lane dividers restricts vehicles to the center of their lanes to reduce lateral drift, but allows lane changes. The potential energy field generated by the median strip, while restricting vehicles to the center of their lanes, prevents them from crossing lanes. A Gaussian-like model is used to describe the intensity of the potential energy fields generated by the lane dividers and median strip boundaries:
[0102]
[0103] In the formula, T1 is the potential energy field intensity factor generated by the lane dividing line, T2 is the potential energy field intensity factor generated by the boundary of the central divider, and q 1-i Let q be the perpendicular distance (m) from the centroid of vehicle i to the nearest white dashed line in the lane where the vehicle is located. 2-i Let m be the vertical distance from the center of mass of vehicle i to the boundary of the central divider. σ1 and σ2 are coefficients representing the rate at which the potential energy field strength increases or decreases as vehicle i approaches or moves away from the lane divider.
[0104] The closer a vehicle is to the road boundary line, the greater the potential energy field strength, and the higher the risk. The formula for calculating the potential energy field strength at the road boundary line is:
[0105]
[0106] In the formula, T3 is the potential energy field intensity factor generated by the road boundary line, and b i Let be the distance, in meters, between the centroid of vehicle i and the road boundary line on the side where the vehicle is located. Considering icy and snowy conditions, road boundary lines, lane dividers, and other road conditions, the driving potential energy field strength can be expressed as:
[0107]
[0108] (2) Vehicle track
[0109] When vehicles travel on the road, they all have some influence on other adjacent vehicles. The strength of the potential field formed by these vehicles depends on the vehicle's own properties and its motion state parameters. Regarding vehicle properties, an equivalent mass is used to describe them, which is related to the actual vehicle weight and speed, as shown in the following formula.
[0110] M i =m i (1.566×10 -14 ·v i 6.687 +0.3345) (8)
[0111] In the formula, M i Let m be the equivalent mass of vehicle i. i For the actual mass of the vehicle, v i Let be the current speed of vehicle i. This shows that the equivalent mass increases with speed, meaning that the equivalent mass of the same vehicle at high speed is much greater than that at low speed.
[0112] Regarding motion parameters, the intensity and distribution of the vehicle's kinematic potential field differ for different velocities and accelerations. The intensity of the vehicle's kinematic potential field at a point in space is mainly related to the target vehicle's velocity *v*, acceleration *a*, and the distance *l* from that point to the vehicle. Generally, the following formula is used to calculate the distance, but this will result in an equal distance *l* regardless of the angle from which the vehicle is approached:
[0113]
[0114] Therefore, a pseudo-distance is introduced, namely the following formula, to correct the distance in actual space:
[0115]
[0116] In the formula, (x0, y0) represents the spatial coordinates of the target vehicle's center of mass, and τ safe The critical threshold for safe distance, v i α represents the vehicle's current speed. v This represents undetermined parameters related to speed.
[0117] During vehicle movement, attractive and repulsive forces are generated between vehicles within a certain distance. For example, when following another vehicle, if the distance between vehicles is small, the following vehicle will appropriately reduce its speed, which is equivalent to the preceding vehicle exerting a repulsive force on it. When the distance between vehicles is large, the following vehicle will increase its speed to reach the desired speed, which is equivalent to the preceding vehicle exerting an attractive force on it. The potential field strength model of vehicle motion is constructed as follows:
[0118]
[0119] In the formula, θ car Let a be the angle between a point around the vehicle and the vehicle's center of mass. i Let ρ be the acceleration of vehicle i currently along the road direction. driver For the driver coefficient, A p This represents driver characteristics, with values ranging from 0 to 1. Values closer to 1 indicate a more aggressive driving style. loc β loc ξ loc These are coefficients to be determined.
[0120] (3) Risk field of heavy vehicles
[0121] This application uses a heavy vehicle risk field to quantify the driving risk of heavy vehicles. The heavy vehicle risk field quantifies the contribution of a single heavy vehicle to the driving risk field at a specific spatiotemporal point. Its mass term adopts a fourth-order relationship (based on ESAL theory), which better reflects the actual laws of road damage. It reflects the mechanical damage (based on the fourth-order law) and braking inertia risk of heavy vehicles to the road surface. The greater the mass, the higher the road fatigue damage and collision kinetic energy. The acceleration term uses an exponential function to characterize the nonlinear risk surge during rapid acceleration and deceleration. Rapid acceleration leads to turbulence in following traffic flow, and rapid deceleration triggers a chain reaction of braking from following vehicles. The greater the absolute value of acceleration, the nonlinear increase in risk contribution. The exhaust gas term uses a hyperbolic tangent function to limit the concentration saturation effect (marginal risk decreases beyond the threshold). Exhaust gas accumulation leads to decreased visibility and driver hypoxia; the risk increases linearly with the proportion of concentration exceeding the threshold. In summary, its expression is:
[0122]
[0123] In the formula, M HV M represents the mass of the loaded vehicle (in tons). std For standard axle load mass, refer to the "Specifications for Design of Asphalt Pavement of Highway" (JTG D50-2017). Instantaneous acceleration of the loaded vehicle (unit: m / s²) 2 ), measured by the vehicle-mounted acceleration sensor, λ HV Q is the factor affecting the acceleration of a loaded vehicle. CO The real-time CO emission rate (unit: g / s) can be measured using a vehicle exhaust emission analyzer. Q crit The CO safety threshold is set according to the "Detailed Guidelines for Ventilation Design of Highway Tunnels". α HV1 α HV2 α HV3 These are the weighting coefficients for each factor, which can be dynamically adjusted based on real-time environmental data.
[0124] Taking all the above risk factors into account, the total risk of the vehicle while in motion is as follows:
[0125] E synthetic =E snow +E locomotion +ω heavy R HV (13)
[0126] In the formula, ω heavy The value is the heavy vehicle impact factor. When the target vehicle is a heavy vehicle, its value is 1, and when the target vehicle is not a heavy vehicle, its value is 0.
[0127] Driving risk intensity is an absolute indicator with a wide numerical range, making it difficult to directly evaluate driving risk. Therefore, it is processed and defined as a relative indicator called the "Driving Risk Index" to assess the magnitude of risk during driving. A higher Driving Risk Index indicates a greater likelihood of risk to the vehicle at that moment. The expression for the Driving Risk Index is:
[0128]
[0129] In the formula, R i E represents the driving risk index of vehicle i. a E represents the average field strength of driving risk. i The total risk of vehicle i while it is in motion
[0130] Step 3: Construct the reward function, action space, and state space based on the relevant data of the main line and ramps.
[0131] Based on traffic operation data collected or calculated from the mainline and ramps, a state space S, action space A, and reward function r are constructed to meet the training requirements of the reinforcement learning model. The cooperative control agent defined in this application includes three regulation dimensions: mainline variable speed limit control, ramp signal control, and mainline CAV vehicle control.
[0132] State space S
[0133] Based on the defined intelligent agent, the state space of the mainline traffic flow is s1 = (v avg ,v max ,n total ,n coop ), where v avg The average operating speed across different lanes is represented by v, where N represents the maximum number of lanes. max This represents the maximum speed limit for the current mainline lane, where n total This represents the total number of vehicles on the main line, where n coop This represents the number of vehicles responding to coordination commands on the mainline. The state space of the ramp is s2 = (n fleet ,t green ), where n fleet t represents the length of the queue at the ramp.green This represents the green light duration of the ramp signal. The state space for mainline vehicles is s3 = (s... v1 ,s v2 ,s v3 ,......s vi ,......,s vn ) represents the set of vehicle states on the main line, while s vi =(v i ,a i ,x i ,y i ,R i Let represent the set of operating states of a single vehicle, where i is the vehicle number. Then the total state space is S = (s1, s2, s3).
[0134] Action Space A
[0135] Based on the defined intelligent agent, the action space of the mainline variable speed limit control is A1 = (V min V1, V2, ... V j ......,V max ), where V j =V min +10(j-1), unit km / h, j=0,1,2......,V min Indicates the minimum design speed for the main line, V max This represents the maximum design speed limit for the main line, and the operating space for ramp signal control is A2 = (t min ,t1,t2,......t k ......,t max ), where t k =t min +5(k-1), unit s, k = 0, 1, 2, ..., t min t represents the minimum green light duration. max Let A represent the maximum green light duration. The action space of the mainline CAV vehicle agent is A3 = (change, keep, speedup), where change means changing lanes to the inner lane, keep means maintaining the current state, and speedup means accelerating to the current maximum speed limit on the mainline. Therefore, the total action space is A = (A1, A2, A3).
[0136] reward function r
[0137] The reward function for the same expressway merging zone is jointly determined by the mainline traffic flow status, ramp queue length and signal timing, and mainline CAV driving status, and mainly includes three types of rewards: traffic efficiency, traffic safety, and cooperative interaction.
[0138] Traffic efficiency rewards
[0139] If we expect the average speed of the mainline traffic flow to remain at a high level, then:
[0140]
[0141] In the formula, r s This represents the main speed reward factor, α. road This indicates the speed bonus coefficient for the lane.
[0142] If we want the queue length of the ramp to be as small as possible, then we have:
[0143]
[0144] Where, r f Represented as the ramp queue length reward factor, n avg The average queue length at the ramp is [number of vehicles].
[0145] Traffic safety awards
[0146] If a traffic accident occurs after the action is taken, a penalty will be imposed:
[0147]
[0148] Where, r acd This is a penalty factor for accidents.
[0149] At the same time, the driving risks of the vehicle should be considered:
[0150]
[0151] Where, r risk This is a penalty factor for driving risks.
[0152] Collaborative Interaction Rewards
[0153] When implementing signal control, it is also necessary to consider balancing the mainline traffic efficiency with the ramp queuing pressure, which leads to:
[0154]
[0155] Where, r balence Main turn balance control reward factor, ω v The main line control influences the weighting coefficient, ω f v is the weighting coefficient for the influence of ramp control. max This is the current maximum speed limit for the main storyline.
[0156] CAV vehicles are also encouraged to respond to signal coordination commands to improve merging efficiency:
[0157]
[0158] Where, r coop For CAV vehicle collaborative reward factor, φ coop For collaborative weights, The average delay, s, for CAV vehicles to execute coordinated commands is automatically calculated by the control center based on the coordinate changes transmitted from the vehicle to the RSU. τ coop This is the delay tolerance coefficient, used to control the decay rate of the delay penalty.
[0159] Therefore, the final reward function is:
[0160] r = w s r s +w f r f +w acd r acd +w risk r risk +w balence r balence +w coop r coop (twenty one)
[0161] Where r is the comprehensive reward function, w s w f w acd w risk w balence w coop These are the influence coefficients corresponding to each factor.
[0162] Step 4: Build and train the RainbowDQN deep reinforcement learning model.
[0163] RainbowDQN is a deep reinforcement learning algorithm that uses a dual Q-network as its core, consisting of a main Q-network Q = (s, a; θ) and a target Q-network Q. target The value function of an action is estimated using the formula Q = (s, a; θ'), where s represents the "state," describing the current state of the agent's environment, and a represents the "action," the behavioral choice the agent can make given the current state. The neural network architectures in both networks are identical: the main Q-network Q = (s, a; θ) is used to select actions, while the target Q-network Q'... target=(s,a;θ') is used to estimate the target Q-value and update the main Q-network parameters. In the network structure, a Dueling DQN is introduced, dividing the neural network into two branches: a value function and a dominance function. This transforms the output Q-value into the sum of state value and action value, improving the accuracy of Q-value prediction. Subsequently, a Distributive DQN is used, employing a value distribution instead of the value function in traditional DQNs to represent all potential outcomes of the state-action pair (a,s). The output is transformed from Q-values into a distribution of Q-values, and the optimal action is selected. A noisy network adds Gaussian noise to the network parameters of these two modules to improve the model's exploration performance. Finally, the main Q-network parameter update uses KL divergence as the loss function to measure the difference between the predicted distribution and the target distribution. A multi-step learning method is used in the calculation of the Q-value distribution. Training combines a multi-step sampling priority replay method, using the absolute value of the temporal difference (TD-error) as a priority indicator to prioritize the sampling of training samples. The target Q-network is updated at fixed step intervals using a weighted average of the parameters of the main Q-network and the current parameters of the target Q-network. The neural network in the RainbowDQN algorithm mainly consists of a main Q-network Q = (s, a; θ) (where θ is the corresponding neural network parameter) and a target network Q... target = (s, a; θ') (θ' is the corresponding neural network parameter), and the neural network parameter is updated by prioritizing experience playback.
[0164] Specifically, during the training of the RainbowDQN model for signal control in the merging zone of expressways, the agent, at each simulation step (total number of steps T), collects the current global traffic environment state s at time t. t Make the next action choice a t The state space of each object in the merging region is therefore changed to s t+1 This refers to state transition, during which the agent receives a reward r from the environment. t As feedback for the executed action, the interaction process (s) is then processed. t a t r t s t+1 d) These samples are stored as experience in the experience replay pool and used as training samples for subsequent parameter updates of the neural network. However, when |D| > M, the lowest priority sample is deleted. Then, a priority experience replay strategy is used to draw a batch of B samples (s) from the experience pool. t a t r t s t+1 ,d), where s t Indicates state, a t Indicates based on state The action to be performed, r t Indicates the execution of action a t Rewards for environmental feedback, s t+1 Indicates the execution of action a t Then, based on the state transition probability, the new state is denoted by d, indicating whether it is the final state. For each sample drawn, the following calculation is performed:
[0165] First, the optimal action a is computed using the main Q-network in the Double DQN network. * The formula is as follows:
[0166] a * =argmax a Q(s t+N ,a;θ) (22)
[0167] In the formula, Q(s) t+N ,a;θ) represents in s t+N At a given time, the main Q network takes action. The expected value is calculated using a Distributional DQN network, as shown in the following formula:
[0168]
[0169] In the formula, x represents the value of the x-th sampling point, and p(z) x |s t+N-1 ,a t+N-1 ;θ) represents the main Q-network in s t+N-1 In the state at time , take action a t+N-1 The probability of the x-th sampling point.
[0170] Subsequently, a multi-step learning strategy is used to compute the Q of the target Q-network. target Value distribution, the formula is as follows:
[0171]
[0172] in For an agent in state s t+N Next, the target Q-network uses the optimal action a provided by the main Q-network. * The calculated value is γ, which is the discount factor. For an agent in state s t The true value of the target.
[0173] Next, the cross-entropy loss function is used to calculate the error between the true value distribution and the estimated value distribution. To use the cross-entropy loss function, the value distribution calculated by the target Q-network and the value distribution of the main Q-network must be aligned to the same range, i.e., [Z]. min Zmax Between ] . After adjustment The distance to the two adjacent sample points on the left and right is used as the weight. The probability is assigned to the left and right adjacent sample points, thus obtaining the true target probability based on the original sample points. The calculation formula using the cross-entropy loss function is as follows:
[0174]
[0175] in, p represents the true probability of the x-th sample point calculated by the target network. x (s t ,a t θ) represents the estimated probability of the x-th sample point calculated using the main Q-network. Since the RainbowDQN algorithm uses both Dueling DQN and Distributional DQN in its neural network output, p x (s t ,a t The formula for calculating θ is as follows:
[0176]
[0177] Finally, stochastic gradient descent is used to update the parameter θ.
[0178]
[0179] To mitigate the overestimation phenomenon of the main Q-network, the target Q-network serves as an auxiliary network to assist in parameter training. The target Q-network parameters are updated using both delayed and soft update strategies. Delayed updates, also known as generational updates, are performed at fixed time intervals. Soft updates, on the other hand, update the target Q-network parameters as a weighted average of the main Q-network parameters and the current target Q-network parameters, as shown in the following formula:
[0180] θ′=τθ+(1-τ)θ′ (28)
[0181] In the formula, τ is the smoothing coefficient, representing the degree of influence of the main Q network parameters on the target network.
[0182] The above detailed process is repeated N times to complete parameter updates and model training.
[0183] Step 5: Use the trained deep reinforcement learning model for signal control.
[0184] The trained RainbowDQN model is deployed to the traffic signal control center and communicates in real time with the roadside unit and signal controller via the API interface. The detection equipment continuously collects traffic-related data information of the main line and ramps, which is converted into standardized feature vectors by the preprocessing module and then input into the model. The model outputs an optimal control action every time the ramp signal cycle is completed. The signal controller receives the action and executes it immediately, and synchronously transmits cooperative instructions to the CAV vehicles through the RSU.
[0185] Example 1
[0186] according to Figure 2 The diagram shows the equipment deployment. Traffic data was collected from weekday mornings from 8:00 AM to 10:00 AM for one month within a certain expressway merging area. The actual scenario was modeled using SUMO simulation software. The main road is a two-way six-lane road with a maximum speed limit of 120 km / h and a minimum speed limit of 60 km / h. It also has a single-lane ramp connection. The maximum green light duration is 30 seconds, and the minimum green light duration is 5 seconds.
[0187] The constructed action space is: A1 = (60, 70, 80, 90, 100, 110, 120), unit km / h.
[0188] A2 = (5, 10, 15, 20, 25, 30), unit: seconds; A3 = (0, 1, 2), where 0 represents "change" (change to the inner lane), 1 represents "keep" (keep the current driving state), and 2 represents "speedup" (accelerate to the current main lane's maximum speed limit). The total action space is A = (A1, A2, A3).
[0189] The main hyperparameters of RainbowDQN reinforcement learning are set as follows: learning rate 0.001, discount factor 0.95, batch size 16, initial exploration probability 0.99, final exploration probability 0.01, minimum value -100, maximum value 100, number of intervals 51, replay experience capacity 100000, multi-step learning step size 2, target network update frequency 50, and smoothing factor 0.001. Using specific state-space information as input and action selection as output, the simulation step size per cycle is set to 3600. The model is trained, and finally, the network parameters of the training cycle with the highest average reward value are selected as the final network parameters.
[0190] Ultimately, the control effect of the merging zone at a certain moment when the next green light signal on the ramp is obtained as follows: Figure 2As shown in the diagram, when the ramp green light illuminates, the Rainbow DQN network, trained on the traffic flow information collected in the previous cycle (i.e., the state space S, where the driving risk distribution of each vehicle is illustrated by the ellipse in the diagram), outputs a variable speed limit of 90 km / h and a ramp green light duration of 15 seconds. Vehicles using traffic information monitoring (CAVs) within the monitoring area receive information from the roadside traffic safety units (RSUs). For example, the CAV in the outermost lane receives the message "change lanes," while the CAV in the middle lane receives the message "maintain current state." The CAVs can then adjust accordingly. This achieves coordinated control of traffic lights, mainline variable speed limit signs, and upstream CAVs on the mainline.
[0191] This application proposes a control method for merging zones on expressways that considers the driving risk field. By taking into account the dynamic and static influence factors of the external environment, a driving potential energy field is established, and a heavy vehicle risk field is established considering the driving risk of heavy vehicles. Finally, a driving risk index is proposed, which can more accurately quantify the driving risk of merging zones on expressways.
[0192] This application designs a collaborative control method for merging zones on expressways based on Rainbow DQN deep reinforcement learning. By considering the mixed traffic flow including CAVs on the mainline of the merging zone, mainline traffic flow information, and ramp queuing information, the state space and action space of the collaborative control agent are refined. A reward function that considers traffic efficiency, traffic safety, and collaborative interaction is proposed, thereby dynamically controlling the mainline speed limit, dynamically coordinating queuing vehicles on ramps, and guiding the driving trajectory of CAV vehicles on the mainline, effectively improving traffic efficiency and traffic safety.
[0193] It should be noted that the specific embodiments are merely explanations and illustrations of the technical solution of the present invention and should not be used to limit the scope of protection. Any modifications made in accordance with the claims and specification of the present invention that are only partial should still fall within the protection scope of the present invention.
Claims
1. A collaborative control method for merging zones based on Rainbow DQN and driving risk field, characterized in that... Includes the following steps: Step 1: Obtain traffic operation data for the main line; Step 2: Obtain the vehicle queue length of the ramp. ; Step 3: Combine traffic operation data and vehicle queue length The features are encoded separately and then input into the Rainbow DQN model to obtain the control action output by the model. The merging zone is then coordinated and controlled based on the control action. The Rainbow DQN model uses a defined state space. Action space and reward function The state space is obtained through training. Represented as: in, The state space of the main line traffic flow. For the state space of the ramp, The state space of the main line vehicle; The action space Represented as: in, The motion space for the main line variable speed limit control. For the action space of ramp signal control, The motion space for the main line CAV vehicle control; The reward function Represented as: in, The main speed reward factor, The reward factor for ramp queue length. As a penalty factor for accidents, As a driving risk penalty factor, The main turn balance control reward factor For CAV vehicle collaborative reward factor, , , , , , These are the influence coefficients corresponding to each factor; The state space of the linear traffic flow Represented as: The state space of the ramp Represented as: The state space of the mainline vehicle Represented as: in, The average operating speed for different lanes, The maximum number of lanes, The speed limit for the main lanes, The total number of vehicles on the main line. The number of vehicles responding to coordination commands on the main line. This refers to the length of the queue at the ramp. The green light duration for ramp signals, This is a set of operating states for a single vehicle. Assign vehicle number, This refers to the driving risk index; The driving risk index Represented as: in, This represents the average strength of driving risk across all scenarios. For vehicles Total risk during driving; The total risk of vehicle i while it is in motion Represented as: in, To ensure strong driving potential energy in every field. Let be the potential field strength of the vehicle. For areas with high-risk heavy vehicles, The value of is the heavy vehicle impact factor. When the target vehicle is a heavy vehicle, its value is 1, and when the target vehicle is a non-heavy vehicle, its value is 0. The heavy vehicle risk field Represented as: in, For the weight of heavy vehicles, For standard axle load mass, For the instantaneous acceleration of the heavy vehicle, for Real-time emission rate for Safety threshold , , These are the weighting coefficients for each factor.
2. The merging zone collaborative control method based on Rainbow DQN and driving risk field according to claim 1, characterized in that... The traffic operation data includes average traffic flow speed. Vehicle speed along the lane direction Vehicle acceleration along the lane direction The horizontal coordinate of the vehicle The vehicle's longitudinal coordinate and the quality of the vehicle .
3. The merging zone collaborative control method based on Rainbow DQN and driving risk field according to claim 2, characterized in that... The operating space of the main line variable speed limit control Represented as: Action space of ramp signal control Represented as: The action space of the main CAV vehicle intelligent agent Represented as: in, , , The minimum driving speed on the main line. The maximum speed limit is designed for the main line. The minimum green light duration, The maximum green light duration, To change lanes to the inside lane, To maintain the current driving state, To accelerate to the current maximum speed limit on the main road.
4. The merging zone collaborative control method based on Rainbow DQN and driving risk field according to claim 3, characterized in that... The main line speed reward factor Represented as: The ramp queue length reward factor Represented as: The accident penalty factor Represented as: The driving risk penalty factor Represented as: The main turn balance control reward factor Represented as: The CAV vehicle collaborative reward factor Represented as: in, This is the current maximum speed limit for the main storyline. This represents the maximum queue length of the ramp. The speed bonus coefficient for the lane. The reward factor for ramp queue length. The average queue length at the ramp. The main line controls the weighting coefficients. The weighting coefficient for the impact of ramp control. For collaborative weights, The average latency for CAV vehicles to execute cooperative commands. This represents the delay tolerance coefficient.
5. The merging zone collaborative control method based on Rainbow DQN and driving risk field according to claim 4, characterized in that... The driving potential energy field strength Represented as: in, Visibility of the road where the vehicle is located. This refers to the slope of the road where the vehicle is located. This is the radius of the curve on the road where the vehicle is located. For standard road visibility, For standard road slope, The standard road curve radius, , , For undetermined constant coefficients, The critical threshold for safe distance, The influence coefficient of ice and snow in cold regions. For speed influence coefficient, For vehicles Speed of travel along the road For vehicles vehicle acceleration, For acceleration influence coefficient, The thickness of thin ice or snow. , For undetermined coefficients, The potential energy field intensity factor generated by the lane dividing line. The potential energy field intensity factor generated at the boundary of the central dividing zone. The potential energy field intensity factor generated by the road boundary line. For vehicles The center of gravity is the perpendicular distance from the lane where the vehicle is located to the nearest white dashed line. For vehicles The vertical distance from the centroid to the boundary of the central divider. For vehicles The distance between the center of gravity and the road boundary line on the side where the vehicle is located. , The coefficient of variation, The proportion of the potential energy field intensity generated by the road boundary line. For road surface temperature, For and is the independent variable.
6. The merging zone collaborative control method based on Rainbow DQN and driving risk field according to claim 5, characterized in that... The vehicle kinematic potential field strength Represented as: in, For vehicles Equivalent quality, The spatial coordinates of the target vehicle's center of mass. Let be the angle between a point around the vehicle and the vehicle's center of mass. For driver coefficient, For driver characteristics, the value ranges from 0 to 1. , , For undetermined coefficients, and As an intermediate variable, These are the vehicle's position coordinates. These are undetermined parameters related to speed.
Citation Information
Patent Citations
Confluence area conflict identification early warning method and system based on driving risk field
CN119028174A
Bus dynamic scheduling method and system based on vehicle management
CN120673615A