Proportional variable speed limit management and control method for mixed traffic flow

By constructing a probabilistic variable speed limit control method based on Markov decision model and dual-delay depth deterministic strategy gradient algorithm in a hybrid traffic flow environment, the implementation effect and stability of the speed limit strategy in a hybrid traffic flow is solved, and the coordinated speed limit between autonomous driving vehicles and artificial driving vehicles is achieved, and the stability and safety of traffic flow are improved.

CN120299248AActive Publication Date: 2025-07-11HARBIN INST OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510522595.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-11
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The implementation effect of existing variable speed limiting methods in hybrid traffic flow environments is greatly affected by driver compliance, and large-scale speed limiting instructions are prone to traffic instability, and the speed limiting strategy in the hybrid mode between autonomous driving vehicles and artificial driving vehicles is insufficient.

Method used

A probabilistic variable speed limit control method based on Markov decision model and dual-delay depth deterministic strategy gradient algorithm is adopted to obtain vehicle and environmental information through roadside terminals, a probabilistic variable speed limit strategy is constructed, speed limit instructions are sent for autonomous vehicles, and autonomous vehicles are used to guide artificial driving vehicles to follow the speed limit strategy and optimize traffic flow.

Benefits of technology

Improve the speed limit execution effect, reduce traffic instability caused by driver disobeying speed limits, and improve the stability and road safety of hybrid traffic flows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299248A_ABST
    Figure CN120299248A_ABST
Patent Text Reader

Abstract

The invention discloses a proportion-variable speed-limiting management and control method for mixed traffic flow, and belongs to the technical field of intelligent traffic management and control. In order to solve the problem of improving the speed limit execution effect, the method comprises the following steps of: constructing a state, an action and a reward function in a probability variable speed limit management and control strategy based on a Markov decision model; constructing a probability variable speed limit management and control strategy solving model based on a double-delay depth deterministic strategy gradient algorithm, and solving an optimal probability variable speed limit management and control strategy of the current traffic state; constructing a structural equation model, calculating the automatic driving separation rate of all the automatic driving vehicles in the control road section under the optimal probability variable speed limit control strategy, and enabling the automatic driving vehicles with the overhigh automatic driving separation rate not to be used for receiving a speed limit control instruction; and according to the speed limit proportion and the current traffic state, constructing a speed limit management and control guiding capability index, performing management and control benefit evaluation on each speed limit execution vehicle combination, and selecting the speed limit execution vehicle combination with the optimal management and control benefit as an execution object of a final speed limit strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent traffic control, and particularly relates to a proportional variable speed limit control method for mixed traffic flow. Background Art

[0002] In a high-traffic environment, the non-periodic congestion that occurs on highways not only seriously reduces the road traffic efficiency but may also trigger a chain reaction, increasing the probability of traffic accidents.

[0003] As a traffic management strategy that dynamically adjusts the road speed limit value based on the traffic environment, the variable speed limit method aims to reduce the speed difference of vehicles near traffic congestion areas, balance the traffic flow distribution, thereby improving the traffic capacity of sections and alleviating congestion. However, there are still some limitations in this control mode: (1) Existing variable speed limit control strategies rely on roadside signs for release. When the driver compliance is low, it is difficult to ensure the speed limit effect. (2) The speed limit value usually adopts relatively obvious integer values instead of the optimal dynamic speed limit value based on real-time traffic conditions, which reduces the accuracy of the speed limit strategy to a certain extent. (3) Temporarily issuing unified speed limit information to all vehicles on the entire section may cause behaviors such as driver nervousness and sudden braking, which instead exacerbates the instability of the traffic flow.

[0004] With the rapid development of autonomous driving technology, a mixed traffic flow environment in which autonomous driving vehicles and human-driven vehicles coexist has gradually formed and will continue to exist for a long time in the future. Autonomous driving vehicles have the ability to interact with the highway variable speed limit control platform in real time, and can be used to receive speed limit instructions and guide human-driven vehicles to follow the variable speed limit strategy to optimize the traffic flow. However, currently, the fully autonomous driving technology is not yet mature, and autonomous driving vehicles are still in the stage of coexistence of manual control and automatic control. When an autonomous driving vehicle receives a speed limit instruction, some drivers may choose to disengage from the autonomous driving mode and switch to manual driving due to distrust of the autonomous driving system. This behavior may cause additional speed fluctuations, thereby affecting the stability of the traffic flow.

[0005] Therefore, how to fully consider the driving characteristics of autonomous driving vehicles in a mixed traffic flow environment, design a more intelligent, precise, and dynamically adaptable variable speed limit strategy, and improve traffic flow stability and road safety has become an important research direction for the construction of intelligent connected highways and one of the key issues for the optimization of the future traffic management system. Summary of the Invention

[0006] The problem to be solved by the present invention is to improve the speed limit execution effect and reduce the traffic instability caused by drivers' non-compliance with speed limits, and a proportional variable speed limit control method for mixed traffic flow is proposed.

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] A proportional variable speed limit control method for mixed traffic flow, comprising the following steps:

[0009] S1. Obtain vehicle information and environmental information of the current section through a roadside terminal, obtain the travel data of autonomous vehicles through information interaction, and retrieve the historical travel data of autonomous vehicles in the section for a certain period of time to construct a section historical traffic data set;

[0010] S2. Based on the Markov decision model, construct the state, action, and reward functions in the probabilistic variable speed limit control strategy;

[0011] S3. Construct a probabilistic variable speed limit control strategy solution model based on the twin-delayed deep deterministic policy gradient algorithm, train the probabilistic variable speed limit control strategy solution model with the section historical traffic data obtained in step S1, obtain a trained probabilistic variable speed limit control strategy solution model, and solve the optimal probabilistic variable speed limit control strategy for the current traffic state;

[0012] S4. Construct a structural equation model, calculate the autonomous driving disengagement rate of all autonomous vehicles in the controlled section under the optimal probabilistic variable speed limit control strategy, and autonomous vehicles with an excessive autonomous driving disengagement rate will not be used to receive speed limit control instructions;

[0013] S5. Construct a speed limit control guidance ability index based on the speed limit ratio and the current traffic state, evaluate the control benefits of each speed limit execution vehicle combination, and select the speed limit execution vehicle combination with the optimal control benefit as the execution object of the final speed limit strategy.

[0014] Further, the specific implementation method of step S1 includes the following steps:

[0015] S1.1. The roadside terminal obtains the vehicle information of the current section, including the driving parameters and position information of all vehicles in the section. The driving parameters specifically include the speed v of vehicle i i and acceleration a i , and the position information refers to the position coordinates (x i , y i ) of the centroid of vehicle i;

[0016] S1.2. The roadside terminal obtains the environmental information of the current section, including traffic environmental information and natural environmental information. The traffic environmental information includes the current road occupancy rate q of the section and the average driving speed of vehicles in the sub-section Among them, the current section is divided into a controlled section and a congested section according to the difference in the average driving speed of vehicles in different sub-sections. When the average driving speed of vehicles in the sub-section If so, the sub - road section is determined as a congested road section r con , α is the congestion judgment threshold, and v * is the free - flow speed; for the congested road section r con the sub - road sections within 1000 meters upstream will be used as controlled road sections r vsl for speed - limit control; the natural environment information includes weather, light, and road humidity;

[0017] S1.3. The roadside terminal obtains vehicle type information through information interaction, including manned vehicles or autonomous vehicles, and collects the personal attributes of the drivers of each autonomous vehicle and the vehicle travel attributes. Among them, the personal attributes of the driver include the driver's age group, gender, and the current autonomous driving supervision status, and the vehicle travel attributes include the total travel duration and the autonomous driving travel duration.

[0018] Furthermore, the specific implementation method of step S2 includes the following steps:

[0019] S2.1. The state consists of 4 elements: the road occupancy rate of the congested road section r con the average driving speed of vehicles on the congested road section r the road occupancy rate of the controlled road section r con the average driving speed of vehicles on the controlled road section r state vsl the road occupancy rate of the controlled road section r the average driving speed of vehicles on the controlled road section r vsl the average driving speed of vehicles on the controlled road section r state

[0020] S2.2. Set the action A to consist of two continuous variables: the speed - limit value and the speed - limit ratio. The speed - limit value v ∈ [v min , v max , where v min is the minimum road - section speed - limit value, and v max is the maximum road - section speed - limit value; the speed - limit ratio P ∈ (0, 1), and the speed - limit ratio is the proportion of autonomous vehicles receiving speed - limit instructions in the current control cycle. The action A = (v, p) constitutes a two - dimensional continuous action;

[0021] S2.3. Set the rewards to include four indicators: vehicle delay, vehicle CO2 emissions, fuel consumption, and driving risk;

[0022] S2.3.1. The expression of the vehicle delay index is:

[0023]

[0024] where D i is the delay time of vehicle i, L is the total road - section length, and respectively represent the average driving speed and free flow speed of vehicle i;

[0025] S2.3.2. The expression for the vehicle CO2 emission index is:

[0026] E i (t) = max (E0, e1 + e2v i (t) + e3(v i (t)) 2 + e4a i (t) + e5(a i (t)) 2 + e6v i (t)a i (t))

[0027]

[0028] where, E i (t) is the CO2 emission rate of vehicle i at time t, v i (t) and a i (t) are the speed and acceleration of vehicle i at time t respectively, E i is the CO2 emission of vehicle i during the observation time T, and the remaining parameters are specifically E0 = 5.11×10 -1 , e1 = 5.53×10 -1 , e2 = 1.61×10 -1 , e3 = -2.89×10 -2 , e4 = 2.66×10 -1 , e5 = 5.11×10 -1 , e6 = 1.83×10 -1 ;

[0029] S2.3.3. The expression for the fuel consumption index is:

[0030] VSP i (t) = v i (t)(1.1a i (t) + 0.132) + 0.000302(v i (t)) 3 ,

[0031]

[0032] where, VSP i (t) is the specific power of vehicle i at time t, NFR i (t) is the fuel consumption rate of vehicle i at time t, NFR i is the fuel consumption of vehicle i during the observation time T;

[0033] S2.3.4. The expression of the driving risk index is as follows:

[0034]

[0035]

[0036] Among them, TTC i (t) represents the time required for a collision to occur when the speed difference of vehicle i remains unchanged, x i-1 (t) and x i (t) represent the positions of vehicle i - 1 and vehicle i at time t respectively, l is the vehicle body length, v i-1 (t) and v i (t) represent the speeds of vehicle i - 1 and vehicle i at time t respectively, TTC i is the total cumulative time of collision occurrence of vehicle i within the observation time T;

[0037] S2.3.5. Construct the weighted reward R as follows:

[0038]

[0039] Among them, ω1, ω2, ω3, ω4 are the weight coefficients of the vehicle delay index, vehicle CO2 emission index, fuel consumption index, and driving risk index respectively, and N is the total number of vehicles.

[0040] Furthermore, the specific implementation method of step S3 includes the following steps:

[0041] S3.1. Construct the network components of the probability variable speed limit control strategy solution model based on the double - delay deep deterministic policy gradient algorithm, including the policy network μ(S t |θ μ ), the evaluation network Q(S t ,A t |θ Q ), the target network of the policy network and the target network of the evaluation network are respectively Among them, the evaluation network and the target network of the evaluation network are composed of two sub - networks, and constitute the evaluation network Q(S t ,A t |θ Q ), and constitute the target network of the evaluation network Initialize the parameters of all networks;

[0042] S3.2. Collect experience data and store it in the experience replay pool: For the state S at time step t t, the model selects an action A according to the current policy network μ(S t |θ μ ), where t is a noise term to ensure exploration; According to the interaction between the action A

[0043] and the environment, the next time-step state S t and the reward R t+1 are obtained. The (S t , A t , R t , S t , S t+1 )) is stored in the experience replay pool;

[0044] S3.3. Sample a batch of data from the experience replay pool: Each time of update, a batch of data is randomly drawn from the experience replay pool;

[0045] S3.4. Use the batch of data sampled in step S3.3 to update the parameters of the evaluation network First, calculate the target Q value The calculation formula is as follows:

[0046]

[0047] where γ is the discount factor, is the action generated by the target network,

[0048] Then, adopt the mean squared error loss function L(θ Q ) to update the parameters of the evaluation network. The expression is:

[0049]

[0050] where m is the size of the batch sampled in S3.3;

[0051] S3.5. Use the updated evaluation network to optimize the policy network: Calculate the gradient of the policy network through the evaluation network. The goal is to maximize the Q value of the evaluation network and update the parameters θ μ of the policy network. The calculation formula is as follows:

[0052]

[0053] where is the gradient of the policy network, is the gradient of the evaluation network for the action, is the gradient of the policy network for the network parameters;

[0054] S3.6. Perform soft update of the target network regularly: The soft update method of the target network parameters is as follows:

[0055]

[0056] where τ is the soft update factor, is the target network parameter of the policy network, is the target network parameter of the evaluation network, θ μ is the policy network parameter, θ Q is the evaluation network parameter;

[0057] S3.7. Iteratively execute steps S3.2 - S3.6 until the predetermined number of training times is reached or the model converges, and obtain the trained probability variable speed limit control strategy solution model;

[0058] S3.8. Use the trained probability variable speed limit control strategy solution model obtained in step S3.7 to solve the optimal control strategy (v vsl , p vsl ) under the current traffic state.

[0059] Furthermore, the specific implementation method of step S4 includes the following steps:

[0060] S4.1. Construct a structural equation calculation model, and the expression is:

[0061] η = Bη + Γξ + ζ

[0062] Z = Λ z ξ + δ

[0063] W = Λ w η + ε

[0064] where η represents the matrix of endogenous latent variables, ξ represents the matrix of exogenous latent variables, and ζ is the matrix of disturbance terms in the structural equation; B represents the influence coefficient matrix of exogenous latent variables on endogenous latent variables, Γ describes the interaction relationship between endogenous latent variables, Z and W are the observed manifest variable matrices corresponding to ξ and η respectively, Λ z and Λ w are the measurement error terms of Z and W respectively;

[0065] S4.2. Construct a structural equation path diagram, and set the internal interaction relationship between exogenous latent variables and endogenous latent variables as: The autonomous driving disengagement rate η1 is affected by the natural environment ξ1, traffic environment ξ2, driving parameters ξ3, driver's personal attributes ξ4, trip attributes ξ5, and control strategy ξ6;

[0066] S4.3. According to the structural equation formula and the structural equation path diagram, the equation group is calculated simultaneously, and the explicit variable coding values ​​obtained by processing the data in the historical traffic data set of the road section are substituted into the simultaneous equation group to fit the coefficient matrix B, Γ, ζ, Λ in the structural equation. z ,Λ w , δ, ε;

[0067] S4.4. The controlled road section r obtained in step S3 vsl The variable information corresponding to each autonomous driving vehicle is brought into the structural equation to obtain the autonomous driving disengagement rate η1 of each autonomous driving vehicle;

[0068] S4.5. When the autonomous driving disengagement rate η1 of any autonomous driving vehicle is higher than the specified speed limit control threshold, it will not be used to receive speed limit control instructions. Except for the autonomous driving vehicles selected to receive speed limit instructions, the remaining autonomous driving vehicles will be equivalent to manually driven vehicles in this round of control and accept guidance from the speed limit enforcement vehicle.

[0069] Furthermore, the specific implementation method of step S5 includes the following steps:

[0070] S5.1. Construct the speed limit control and guidance capability index to obtain the speed limit control and guidance capability Φ of the autonomous driving vehicle i to the manually driven vehicle j ij The specific calculation method is as follows:

[0071] Φ ij =f(d ij ,Δv ij ,ρ ij ,κ i )

[0072] Among them, d ij is the distance along the road between the autonomous driving vehicle i and the manually driven vehicle j, Δv ij is the speed difference between i and j along the road direction, ρ ij is the traffic flow density of the sub-segments near i and j, κ i is the stability factor of the speed limit instruction currently executed by i, f(·) is the function of the comprehensive control effect, which is a weighted linear combination or nonlinear function;

[0073] S5.2. Calculate the comprehensive control benefit under the speed limit control guidance capability index, set in the process of speed limit guidance, if the comprehensive guidance value of the guided manually driven vehicle is lower than the preset threshold value φ * , then the vehicle is considered to have not been effectively guided. The calculation method is as follows:

[0074] S5.2.1. Calculate the effective guidance rate: Let the number of manually driven vehicles in all controlled sections be N h, let the comprehensive guidance value of the j-th human-driven vehicle be Then the effective guidance rate The expression of is:

[0075]

[0076] Among them, I(·) is an indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise;

[0077] S5.2.2. Calculate the comprehensive guidance value: Based on the fact that each human-driven vehicle may be guided by multiple autonomous vehicles at the same time, let the set of autonomous vehicles be Vehicle For the guidance ability Φ of the human-driven vehicle j ij The weight of the action is χ ij , then the comprehensive guidance value of vehicle j The calculation formula is as follows:

[0078]

[0079] Among them, χ ij ∈[0,1], which is set according to factors such as the distance attenuation function, relative speed, or sensing range;

[0080] S5.2.3. Final evaluation value of the scheme: To comprehensively measure the advantages and disadvantages of different implementation schemes, the comprehensive management and control benefit index G vsl is introduced and calculated by weighted calculation of the effective guidance rate and the average comprehensive guidance value:

[0081]

[0082] Among them, λ1 and λ2 are weight parameters, and the emphasis target is set according to the system design requirements.

[0083] Advantages and effects of the present invention:

[0084] A proportional variable speed limit control method for hybrid traffic flow according to the present invention aims at the problems of insufficient adaptability of the existing variable speed limit method in the hybrid traffic environment, large influence of the execution effect on driver compliance, and easy traffic fluctuations caused by large-scale speed limit instructions. A refined proportional speed limit scheme is constructed based on the reinforcement learning method. By sending speed limit instructions to some autonomous vehicles, the autonomous vehicles are used to guide the human-driven vehicles to follow the speed limit strategy, improve the speed limit execution effect, and reduce traffic instability caused by driver non-compliance with the speed limit. Brief description of the drawings

[0085] Figure 1 is a flowchart of a proportional variable speed limit control method for hybrid traffic flow according to the present invention;

[0086] Figure 2 This is the structural equation path diagram of the present invention. Specific Embodiments

[0087] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. Usually, the components of the specific embodiments of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0088] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed present invention, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0089] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and are described in detail in conjunction with the attached Figure 1 and the attached Figure 2 as follows:

[0090] Example 1:

[0091] A proportional variable speed limit control method for mixed traffic flow includes the following steps:

[0092] S1. Obtain the vehicle information and environmental information of the current road section through a roadside terminal, obtain the travel data of autonomous vehicles through information interaction, and retrieve the historical travel data of autonomous vehicles on the road section for a certain period of time to construct a historical traffic data set for the road section;

[0093] Further, the specific implementation method of step S1 includes the following steps:

[0094] S1.1. The roadside terminal obtains the vehicle information of the current road section, including the driving parameters and position information of all vehicles in the road section. The driving parameters specifically include the speed v of vehicle i i , acceleration a i , and the position information refers to the position coordinates (x i , y i ) where the centroid of vehicle i is located;

[0095] S1.2. The roadside terminal obtains the environmental information of the current road section, including traffic environment information and natural environment information. The traffic environment information includes the current road occupancy rate q of the road section and the average driving speed of vehicles within the sub-road section. Among them, the current road section is divided into a controlled section and a congested section according to the difference in the average driving speed of vehicles in different sub-road sections. When the average driving speed of vehicles within the sub-road section is less than or equal to αv, then this sub-road section is determined as a congested section r, con where α is the congestion judgment threshold and v * is the free flow speed; for the congested section r con the sub-road sections within 1000 meters upstream will be used as the controlled section r vsl for speed limit control; the natural environment information includes weather, light, and road humidity.

[0096] S1.3. The roadside terminal obtains vehicle type information through information interaction, including human-driven vehicles or autonomous vehicles, and collects the personal attributes of the drivers of each autonomous vehicle and the vehicle trip attributes. Among them, the personal attributes of the drivers include the driver's age group, gender, and the current autonomous driving supervision status, and the vehicle trip attributes include the total trip duration and the autonomous driving trip duration.

[0097] S2. Based on the Markov decision model, construct the state, action, and reward functions in the variable probability speed limit control strategy;

[0098] Furthermore, the specific implementation method of step S2 includes the following steps:

[0099] S2.1. Set the state to be composed of 4 elements: the road occupancy rate of the congested section r con , the average driving speed of vehicles in the congested section r , the road occupancy rate of the controlled section r con , and the average driving speed of vehicles in the controlled section r , the state vsl , the road occupancy rate of the controlled section r , and the average driving speed of vehicles in the controlled section r vsl ; State

[0100] S2.2. Set the action A to be composed of two continuous variables: the speed limit value and the speed limit ratio. The speed limit value v ∈ [v min , v max , where v min is the minimum road section speed limit value and v max is the maximum road section speed limit value; the speed limit ratio P ∈ (0, 1), and the speed limit ratio is the proportion of autonomous vehicles receiving speed limit instructions within the current control period. The action A = (v, p) constitutes a two-dimensional continuous action;

[0101] S2.3. Set the rewards to include four indicators: vehicle delay, vehicle CO2 emissions, fuel consumption, and driving risk;

[0102] These indicators reflect the overall performance of the traffic system. When training the model, single-reward training can be adopted according to requirements, or weighted training can be carried out according to weights;

[0103] S2.3.1. The expression for the vehicle delay indicator is:

[0104]

[0105] where D i is the delay time of vehicle i, L is the total road section length, and represent the average driving speed and free flow speed of vehicle i, respectively;

[0106] S2.3.2. The expression for the vehicle CO2 emissions indicator is:

[0107] E i (t) = max (E0, e1 + e2v i (t) + e3(v i (t)) 2 + e4a i (t) + e5(a i (t)) 2 + e6v i (t)a i (t))

[0108]

[0109] where E i (t) is the CO2 emission rate of vehicle i at time t, v i (t) and a i (t) are the speed and acceleration of vehicle i at time t, respectively, E i is the CO2 emissions of vehicle i during the observation time T, and the remaining parameters are specifically E0 = 5.11×10 -1 , e1 = 5.53×10 -1 , e2 = 1.61×10 -1 , e3 = -2.89×10 -2 , e4 = 2.66×10 -1 , e5 = 5.11×10 -1 , e6 = 1.83×10 -1 ;

[0110] S2.3.3. The expression for the fuel consumption indicator is:

[0111] VSP i VSP(t) = v i (t)(1.1a i (t) + 0.132) + 0.000302(v i (t)) 3 ,

[0112]

[0113] where, VSP i (t) is the specific power of vehicle i at time t, and NFR i (t) is the fuel consumption rate of vehicle i at time t, and NFR i is the fuel consumption of vehicle i during the observation time T;

[0114] S2.3.4. The expression of the driving risk index is:

[0115]

[0116] where, TTC i (t) represents the time required for a collision to occur when the speed difference of vehicle i remains unchanged, and x i-1 (t) and x i (t) represent the positions of vehicle i - 1 and vehicle i at time t respectively, l is the vehicle body length, and v i-1 (t) and v i (t) represent the speeds of vehicle i - 1 and vehicle i at time t respectively, and TTC i is the total cumulative time of collisions of vehicle i during the observation time T;

[0117] S2.3.5. Construct the weighted reward R as follows:

[0118]

[0119] where, ω1, ω2, ω3, ω4 are the weight coefficients of the vehicle delay index, vehicle CO2 emission index, fuel consumption index, and driving risk index respectively, and N is the total number of vehicles.

[0120] S3. Construct a probability variable speed limit control strategy solution model based on the double - delay deep deterministic policy gradient algorithm, and use the historical traffic data of the road section obtained in step S1 to train the probability variable speed limit control strategy solution model to obtain a trained probability variable speed limit control strategy solution model, and solve the optimal probability variable speed limit control strategy for the current traffic state;

[0121] Furthermore, the specific implementation method of step S3 includes the following steps:

[0122] S3.1. Construct the network components of the probability variable speed limit control strategy solution model based on the Double Delayed Deep Deterministic Policy Gradient algorithm, including the policy network μ(S t |θ μ ), the evaluation network Q(S t , A t |θ Q ), and the target networks of the policy network and the evaluation network are respectively Among them, the evaluation network and the target network of the evaluation network are composed of two sub-networks, and constitute the evaluation network Q(S t , A t |θ Q ), and constitute the target network of the evaluation network Initialize the parameters of all networks;

[0123] S3.2. Collect experience data and store it in the experience replay pool: For the state S at time step t t , the model selects an action A t |θ μ ) according to the current policy network μ(S t ), where is the noise term to ensure exploration;

[0124] According to the action A t and the interaction with the environment, obtain the next time step state S t+1 and the reward R t , and store (S t , A t , R t , S t+1 )) in the experience replay pool;

[0125] S3.3. Sample a batch of data from the experience replay pool: Each time of update, randomly draw a batch of data from the experience replay pool;

[0126] S3.4. Update the parameters of the evaluation network using the batch data sampled in step S3.3 , first calculate the target Q value The calculation formula is as follows:

[0127]

[0128] Among them, γ is the discount factor, is the action generated by the target network,

[0129] Then adopt the mean square error loss function L(θQ ) Update the evaluation network parameters, with the expression:

[0130]

[0131] where m is the size of the sampling batch in S3.3.

[0132] S3.5. Optimize the policy network using the updated evaluation network: Calculate the gradient of the policy network through the evaluation network, with the goal of maximizing the Q-value of the evaluation network, and update the parameters θ of the policy network μ , and the calculation formula is as follows:

[0133]

[0134] where, is the gradient of the policy network, is the gradient of the evaluation network with respect to the action, is the gradient of the policy network with respect to the network parameters;

[0135] Furthermore, minimize this loss so that the policy network can select actions with higher Q-values, thereby optimizing the policy.

[0136] S3.6. Periodically perform a soft update of the target network: The soft update method of the target network parameters is as follows:

[0137]

[0138] where τ is the soft update factor, is the target network parameter of the policy network, is the target network parameter of the evaluation network, θ μ is the policy network parameter, θ Q is the evaluation network parameter;

[0139] S3.7. Iteratively execute steps S3.2 - S3.6 until reaching the predetermined number of training times or the model converges, to obtain a trained probability variable speed limit control strategy solution model;

[0140] S3.8. Use the trained probability variable speed limit control strategy solution model obtained in step S3.7 to solve the optimal control strategy (v vsl , p vsl ) under the current traffic state.

[0141] S4. Construct a structural equation model, calculate the autonomous driving disengagement rate of all autonomous vehicles in the controlled section under the optimal probability variable speed limit control strategy, and autonomous vehicles with too high an autonomous driving disengagement rate will not be used to receive speed limit control instructions;

[0142] Furthermore, the specific implementation method of step S4 includes the following steps:

[0143] S4.1. Construct a structural equation calculation model, the expression is:

[0144] η=Bη+Γξ+ζ

[0145] Z=Λ z ξ+δ

[0146] W=Λ w η+ε

[0147] Among them, η represents the endogenous latent variable matrix, ξ represents the exogenous latent variable matrix, ζ is the interference term matrix in the structural equation; B represents the influence coefficient matrix of exogenous latent variables on endogenous latent variables, Γ describes the interaction relationship between endogenous latent variables, Z and W are the observed explicit variable matrices corresponding to ξ and η respectively, Λ z With Λ w are the measurement error terms of Z and W respectively;

[0148] The specific variable division is shown in Table 1:

[0149] Table 1:

[0150]

[0151]

[0152] When processing the historical travel data of the autonomous driving vehicle, the explicit variables involved in each record are encoded. The explicit variables are divided into two categories: ordinal and categorical. Ordinal variables are normalized and assigned values ​​according to their numerical range, and categorical variables are encoded in binary form. The detailed explicit variable encoding method is shown in Table 2:

[0153] Table 2:

[0154]

[0155]

[0156] S4.2. Construct a structural equation path diagram and set the internal relationship between the exogenous latent variables and the endogenous latent variables as follows: the autonomous driving disengagement rate η1 is affected by the natural environment ξ1, the traffic environment ξ2, the driving parameters ξ3, the driver's personal attributes ξ4, the trip attributes ξ5, and the control strategy ξ6;

[0157] S4.3. According to the structural equation formula and the structural equation path diagram, the equation group is calculated simultaneously, and the explicit variable coding values ​​obtained by processing the data in the historical traffic data set of the road section are substituted into the simultaneous equation group to fit the coefficient matrix B, Γ, ζ, Λ in the structural equation. z ,Λw , δ, ε;

[0158] S4.4. Substitute the variable information corresponding to each autonomous vehicle in the controlled section r obtained in step S3 into the structural equation to obtain the autonomous driving disengagement rate η1 of each autonomous vehicle; vsl within the controlled section r obtained in step S3 into the structural equation to obtain the autonomous driving disengagement rate η1 of each autonomous vehicle;

[0159] S4.5. When the autonomous driving disengagement rate η1 of any autonomous vehicle is higher than the specified speed limit control threshold, it will not be used to receive the speed limit control instruction. Except for the autonomous vehicles selected to receive the speed limit instruction, the remaining autonomous vehicles are equivalently regarded as manually driven vehicles in this round of control and receive the guiding effect from the speed limit execution vehicle.

[0160] S5. Construct a speed limit control guiding ability index based on the speed limit ratio and the current traffic state, evaluate the control benefits of each combination of speed limit execution vehicles, and select the combination of speed limit execution vehicles with the best control benefits as the execution object of the final speed limit strategy.

[0161] Further, the specific implementation method of step S5 includes the following steps:

[0162] S5.1. Construct a speed limit control guiding ability index to obtain the speed limit control guiding ability Φ of autonomous vehicle i for manually driven vehicle j ij The specific calculation method is as follows:

[0163] Φ ij = f(d ij , Δv ij , ρ ij , κ i )

[0164] where d ij is the distance along the road driving direction between autonomous vehicle i and manually driven vehicle j, Δv ij is the speed difference along the road driving direction between i and j, ρ ij is the traffic flow density of the sub - section near i and j, κ i is the stability factor of i currently executing the speed limit instruction, and f(·) is a function of the comprehensive control effect, which is a weighted linear combination or a non - linear function;

[0165] In the form of linear weighting, it can be written as:

[0166]

[0167] where β1, β2, β3, β4 are normalized weight coefficients, which can be set or trained according to historical data or strategy design.

[0168] To measure the guiding role of each autonomous vehicle in the implementation of the speed limit policy, an index of speed limit control and guidance ability is proposed. This ability is affected by factors such as its driving state, position, surrounding traffic flow density, and stability of the vehicle distance. The core of the guiding ability is represented by the control field, which represents the control radiation range of an autonomous vehicle in space.

[0169] S5.2. Calculate the comprehensive control benefit under the speed limit control and guidance ability index. During the speed limit guidance process, if the comprehensive guidance value of the manually driven vehicle under guidance is lower than the preset threshold φ * , it is considered that the vehicle has not received effective guidance, and the calculation method is as follows:

[0170] S5.2.1. Calculate the effective guidance rate: Denote the number of manually driven vehicles in all controlled sections as N h , and let the comprehensive guidance value of the j-th manually driven vehicle be Then the effective guidance rate The expression of is:

[0171]

[0172] where I(·) is the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise;

[0173] S5.2.2. Calculate the comprehensive guidance value: Based on the fact that each manually driven vehicle may be simultaneously guided by multiple autonomous vehicles, let the set of autonomous vehicles be vehicle For the guiding ability Φ ij of the manually driven vehicle j, the weight is χ ij , then the comprehensive guidance value of vehicle j is calculated as follows:

[0174]

[0175] where χ ij ∈[0,1], and is set according to factors such as the distance attenuation function, relative speed, or perception range;

[0176] S5.2.3. Final evaluation value of the scheme: To comprehensively measure the advantages and disadvantages of different implementation schemes, an index of comprehensive control benefit G vsl is introduced, which is calculated by weighted calculation of the effective guidance rate and the average comprehensive guidance value:

[0177]

[0178] where λ1 and λ2 are weight parameters, and the emphasis target is set according to the system design requirements.

[0179] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0180] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A variable proportion speed limit control method for mixed traffic flows, characterized in that, It includes the following steps: S1. Obtain the vehicle information and environmental information of the current road section through the roadside terminal, obtain the travel data of the autonomous vehicle through information interaction, and retrieve the historical travel data of the autonomous vehicle on the road section for a certain period of time to construct the historical traffic dataset of the road section; S2. Based on the Markov decision model, construct the state, action, and reward functions in the probability variable speed limit control strategy; S3. Construct a probability variable speed limit control strategy solution model based on the twin-delayed deep deterministic policy gradient algorithm, and use the historical traffic data of the road section obtained in step S1 to train the probability variable speed limit control strategy solution model to obtain a trained probability variable speed limit control strategy solution model, and solve the optimal probability variable speed limit control strategy for the current traffic state; S4. Construct a structural equation model, calculate the autonomous driving disengagement rate of all autonomous vehicles in the controlled road section under the optimal probability variable speed limit control strategy, and autonomous vehicles with an excessive autonomous driving disengagement rate will not be used to receive speed limit control instructions; S5. Construct a speed limit control guidance ability index based on the speed limit ratio and the current traffic state, evaluate the control benefits of each speed limit execution vehicle combination, and select the speed limit execution vehicle combination with the best control benefit as the execution object of the final speed limit strategy.

2. The variable proportion speed limit control method for mixed traffic flow according to claim 1, characterized in that The specific implementation method of step S1 includes the following steps: S1.

1. The roadside terminal obtains the vehicle information of the current road section, including the driving parameters and location information of all vehicles within the road section. The driving parameters specifically include the speed v of vehicle i i , acceleration a i , and the location information refers to the location coordinates (x i , y i ) of the centroid of vehicle i; S1.

2. The roadside terminal obtains the environmental information of the current road section, including traffic environmental information and natural environmental information. The traffic environmental information includes the current road occupancy rate q of the road section and the average driving speed of vehicles within the sub-road section. Among them, the current road section is divided into a controlled road section and a congested road section according to the difference in the average driving speed of vehicles in different sub-road sections. When the average driving speed of vehicles within the sub-road section is less than or equal to αvf, then this sub-road section is determined as a congested road section r com , where α is the congestion judgment threshold and vf * is the free flow speed; for the congested road section r con , the sub-road sections within 1000 meters upstream will be used as the controlled road section r vsl for speed limit control; the natural environmental information includes weather, light, and road humidity. S1.

3. The roadside terminal obtains vehicle type information through information interaction, including human-driven vehicles or autonomous vehicles, and collects the personal attributes of the drivers of each autonomous vehicle and the vehicle travel attributes. Among them, the personal attributes of the driver include the driver's age group, gender, and the current autonomous driving supervision status, and the vehicle travel attributes include the total travel duration and the autonomous driving travel duration.

3. A variable proportion speed limit control method for mixed traffic flow according to claim 1 or 2, characterized in that The specific implementation method of step S2 includes the following steps: S2.

1. The setting status consists of 4 elements: the road occupancy rate of congested section r con of congested section r con average vehicle driving speed regulated section r vsl road occupancy rate of regulated section r vsl average vehicle driving speed of status S2.

2. The setting action A consists of two continuous variables, namely the speed limit value and the speed limit ratio. The speed limit value v ∈ [v min , v max , where v min is the minimum speed limit value of the section, and v max is the maximum speed limit value of the section; The speed limit ratio P ∈ (0, 1), and the speed limit ratio is the proportion of autonomous vehicles receiving speed limit instructions in the current control cycle. The action A = (v, p) constitutes a two-dimensional continuous action; S2.

3. Set the rewards to include four indicators: vehicle delay, vehicle CO2 emissions, fuel consumption, and driving risk; S2.3.

1. The expression of the vehicle delay index is: Among them, D i is the delay time of vehicle i, L is the total road section length, and represent the average driving speed and free flow speed of vehicle i respectively; S2.3.

2. The expression of the vehicle CO2 emissions index is: E i (t) = max(E0, e1 + e2v i (t) + e3(v i (t)) 2 + e4a i (t) + e5(a i (t)) 2 + e6v i (t)a i (t)) Among them, E i (t) is the CO2 emission rate of vehicle i at time t, v i (t) and a i (t) are the speed and acceleration of vehicle i at time t respectively, E i is the CO2 emission of vehicle i within the observation time T, and the remaining parameters are specifically E0 = 5.11×10 -1 , e1 = 5.53×10 -1 , e2 = 1.61×10 -1 , e3 = -2.89×10 -2 , e4 = 2.66×10 -1 , e5 = 5.11×10 -1 , e6 = 1.83×10 -1 ; S2.3.

3. The expression of the fuel consumption index is: VSP i v(t) = v i v(t)(1.1a i v(t) + 0.132) + 0.000302(v i v(t)) 3 , Among them, VSP i (t) is the specific power of vehicle i at time t, and NFR i (t) is the fuel consumption rate of vehicle i at time t, and NFR i is the fuel consumption of vehicle i within the observation time T; S2.3.

4. The expression of the driving risk index is: Among them, TTC i (t) represents the time required for a collision to occur when the speed difference of vehicle i remains unchanged, x i-1 (t) and x i (t) represent the positions of vehicle i - 1 and vehicle i at time t respectively, l is the vehicle body length, v i-1 (t) and v i (t) represent the speeds of vehicle i - 1 and vehicle i at time t respectively, TTC i is the total cumulative time of collision of vehicle i within the observation time T; S2.3.

5. Construct the weighted reward R as follows: Among them, ω1, ω2, ω3, and ω4 are the weight coefficients of the vehicle delay index, vehicle CO2 emissions index, fuel consumption index, and driving risk index respectively, and N is the total number of vehicles.

4. A method for variable proportion speed limit control for mixed traffic flow according to claim 3, characterized in that, The specific implementation method of step S3 includes the following steps: S3.

1. Construct the network components of the probability-variable speed limit control strategy solution model based on the double-delay deep deterministic policy gradient algorithm, including the policy network μ(S t |θ μ ), the evaluation network Q(S t , A t |θQ ) . The target network of the policy network and the target network of the evaluation network are respectively Among them, the evaluation network and the target network of the evaluation network are composed of two sub-networks and constitute the evaluation network Q(S t , A t |θ Q ), and constitute the target network of the evaluation network Initialize the parameters of all networks; S3.

2. Collect empirical data and store it in the empirical replay pool: For the state S at time step t t , the model selects an action A according to the current policy network μ(S t |θμ), where t is a noise term to ensure exploration; ​ According to action A t The interaction with the environment yields the next time-step state S t+1 and reward R t , (S t , A t , R t , S t+1 )) is stored in the experience replay pool; S3.

3. Sample a batch of data from the experience replay pool: Each time an update is made, randomly extract a batch of data from the experience replay pool; S3.

4. Update the parameters of the evaluation network using the batch data sampled in step S3.3 First, calculate the target Q value The calculation formula is as follows: where γ is the discount factor, is the action generated by the target network, Then, the mean squared error loss function \(L(\theta Q )\) is used to update the evaluation network parameters, and the expression is as follows: Among them, m is the size of the sampled batch in S3.3; S3.

5. Optimize the policy network using the updated evaluation network: Calculate the gradient of the policy network through the evaluation network, with the goal of maximizing the Q-value of the evaluation network and updating the parameters θ of the policy network μ , and the calculation formula is as follows: Among them, is the gradient of the policy network, is the gradient of the evaluation network for the action, is the gradient of the policy network for the network parameters; S3.

6. Periodically perform a soft update of the target network: The soft update method of the target network parameters is as follows: where τ is the soft update factor, are the target network parameters of the policy network, are the target network parameters of the evaluation network, and θ μ are the policy network parameters, and θ Q are the evaluation network parameters; S3.

7. Iteratively execute steps S3.2 - S3.6 until the predetermined number of training times is reached or the model converges, and obtain the trained probability variable speed limit control strategy solution model; S3.

8. Use the trained variable probability speed limit control strategy solution model obtained in step S3.7 to solve the optimal control strategy (v vsl , p vsl ) under the current traffic state.

5. A variable-ratio speed limit control method for mixed traffic flow according to claim 4, characterized in that, The specific implementation method of step S4 includes the following steps: S4.

1. Construct a structural equation calculation model, and the expression is: η = Bη + Γξ + ζ Z = Λ z ξ + δ W = Λ w η + ε Among them, η represents the matrix of endogenous latent variables, ξ represents the matrix of exogenous latent variables, and ζ is the matrix of disturbance terms in the structural equation; B represents the influence coefficient matrix of exogenous latent variables on endogenous latent variables, Γ describes the interaction relationship between endogenous latent variables, Z and W are the observed manifest variable matrices corresponding to ξ and η respectively, Λ z and Λ w are the measurement error terms of Z and W respectively; S4.

2. Construct a structural equation path diagram, and set the internal action relationship between exogenous latent variables and endogenous latent variables as: the autonomous driving disengagement rate η1 is affected by the natural environment v1, traffic environment ξ2, driving parameters ξ3, driver's personal attributes ξ4, trip attributes ξ5, and control strategy ξ6; S4.

3. According to the structural equation formula and the structural equation path diagram, the equation group is calculated simultaneously, and the explicit variable coding values ​​obtained by processing the data in the historical traffic data set of the road section are substituted into the simultaneous equation group to fit the coefficient matrix B, Γ, ζ, Λ in the structural equation. z ,Λ w , δ, ε; S4.

4. Substitute the variable information corresponding to each autonomous vehicle within the controlled section r vsl obtained in step S3 into the structural equation to obtain the autonomous driving disengagement rate η1 of each autonomous vehicle; S4.

5. When the autonomous driving disengagement rate η1 of any autonomous vehicle is higher than the specified speed limit control threshold, it will not be used to receive speed limit control instructions. Except for the autonomous vehicles selected to receive speed limit instructions, the remaining autonomous vehicles are equivalently regarded as manually driven vehicles in this round of control and receive the guiding effect from speed limit execution vehicles.

6. The variable proportion speed limit control method for mixed traffic flow according to claim 5, wherein, The specific implementation method of step S5 includes the following steps: S5.

1. Construct the speed limit control and guidance ability index to obtain the speed limit control and guidance ability Φ of the autonomous vehicle i for the human-driven vehicle j ij The specific calculation method is as follows: Φ ij = f(d ij , Δv ij , ρ ij , κ i ) Among them, d ij is the distance along the road driving direction between the autonomous vehicle i and the human-driven vehicle j, and Δv ij is the speed difference along the road driving direction between i and j, and ρ ij is the traffic flow density of the sub-road section near i and j, and κ i is the stability factor of the speed limit instruction currently executed by i, and f(·) is a function of the comprehensive control effect, which is a weighted linear combination or a non-linear function; S5.

2. Calculate the comprehensive control benefit under the speed limit control and guidance ability index. During the speed limit guidance process, if the comprehensive guidance value of the manually driven vehicle under guidance is lower than the preset threshold φ * , it is considered that the vehicle has not received effective guidance, and the calculation method is as follows: S5.2.

1. Calculate the effective guidance rate: Denote the number of human-driven vehicles in all controlled sections as N h , and let the comprehensive guidance value of the j-th human-driven vehicle be Then the effective guidance rate has the following expression: where I(·) is the indicator function, which takes the value of 1 when the condition is satisfied, otherwise 0; S5.2.

2. Calculate the comprehensive guidance value: Based on the fact that each manually driven vehicle may be guided by multiple autonomous vehicles simultaneously, let the set of autonomous vehicles be vehicle For the guidance ability Φ of the manually driven vehicle j ij The weight of the action is χ ij , then the comprehensive guidance value of vehicle j The calculation formula is as follows: where χ ij ∈ [0, 1], which is set according to factors such as distance attenuation function, relative speed, or sensing range; S5.2.

3. Final evaluation value of the solution: To comprehensively measure the advantages and disadvantages of different implementation solutions, the comprehensive management and control benefit index G is introduced. vsl , which is calculated by weighted calculation of the effective guidance rate and the average comprehensive guidance value: where λ1 and λ2 are weight parameters, and the emphasis target is set according to the system design requirements.

Citation Information

Patent Citations

  • Dynamic cooperative management and control method for variable speed limit of automatic driving special lane and universal lane in confluence area on expressway

    CN113096416A

  • Trajectory planning method and system facing mixed environment

    CN116946188A

  • Highway differential variable speed limit control method and device based on deep reinforcement learning and storage medium

    CN117496721A

  • CAV speed guidance system and method based on deep reinforcement learning

    CN117612396A

  • Variable speed limit strategy test method considering driver compliance in mixed traffic flow

    CN117994984A