Traffic confluence control method and system
By collecting and analyzing traffic flow data, determining the traffic flow status of the target control, and adjusting vehicle behavior using optimization models and multi-agent reinforcement learning models, the problems of delays and unstable behaviors in ramp confluence control are solved, and more efficient and safe traffic flow management is achieved.
Patent Information
- Application Number
- CN202411879556.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing ramp merging control methods cannot effectively solve the ramp merging problem caused by dynamic changes in traffic flow, and in mixed traffic flows, human driving behavior may lead to the failure of target control traffic signals.
Traffic flow data is collected through roadside detectors and vehicle-mounted terminals, and the target control traffic flow state is determined in combination with the theory of the basic traffic flow chart. The optimization model is used to determine intelligent connected vehicles that need to participate in collaboration, and the control instructions of cooperative vehicles and speed instructions of ramp vehicles are generated in real time. The driving speed of ramp vehicles is adjusted through the multi-agent reinforcement learning model.
Effectively reduce delays and energy consumption in ramp merging, avoid frequent unstable behaviors such as braking and acceleration, improve traffic flow efficiency and safety, and is suitable for hybrid traffic scenarios where manual driving and autonomous driving vehicles coexist.
Smart Images

Figure CN119942780A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent transportation systems and relates to a traffic merging control method and system. Background Art
[0002] With the rapid development of urbanization and industrialization, the ramp merging areas of highways, urban expressways and other roads have become high-incidence areas of congestion. In these areas, the traffic flow of the main line and the ramp needs to merge in a limited space, and unreasonable merging will lead to unstable traffic flow, further causing a series of problems such as congestion and traffic accidents. Traditional traffic management modes, such as simple traffic signal control or variable speed limit control, can no longer meet the growing and complex needs of highway traffic. Therefore, it has become an urgent task to realize more sophisticated and intelligent control technology for ramp merging areas.
[0003] Ramp merging control technology is an important breakthrough in the informationization and intelligentization of traffic in recent years. This technology mainly detects and controls the traffic flow of ramps and main roads through real-time data collection and advanced data analysis algorithms. It can comprehensively analyze the multi-dimensional information such as the speed, traffic volume, and vehicle spacing of each lane in the merging area, and determine the appropriate control strategies such as lane change and speed adjustment through the model. Therefore, in-depth research on ramp merging control technology can not only promote the development of highway traffic management to a higher level of informationization and intelligence, but also have a profound impact on alleviating traffic congestion and improving road transportation efficiency and safety.
[0004] Existing ramp merging control methods often fail to solve the ramp merging problem with dynamically changing traffic flow from the macro and micro traffic flow levels. The introduction of target-controlled traffic signals provides a new solution to this problem, but in mixed traffic flows, the interference caused by human driving behavior may make target-controlled traffic signals ineffective, and further research is needed on the collaboration between vehicles and the interaction between different types of traffic participants. Summary of the invention
[0005] The purpose of the present invention is to provide a traffic merging control method and system, which can effectively reduce the delay and energy consumption in ramp merging and avoid unstable behaviors such as frequent braking and acceleration.
[0006] In order to solve the above technical problems, the present invention is implemented by adopting the following technical solutions.
[0007] In a first aspect, the present invention provides a traffic merging control method, comprising the following steps:
[0008] Collecting traffic flow data of the main line and ramps through roadside detectors and vehicle-mounted terminals; determining a target control traffic flow state based on the traffic flow data;
[0009] According to the number of vehicles, vehicle types and vehicle position distribution in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, the intelligent networked vehicles that need to participate in the coordination are determined through a pre-built optimization model, and the determined intelligent networked vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal platooning solution;
[0010] According to the target control traffic flow state and the optimal formation plan, the control time and control instructions of the cooperative vehicles are generated in real time; the control instructions are issued to the cooperative vehicles through the communication system of the road test processor, instructing the cooperative vehicles to slow down within the control time, thereby forming a sufficient gap for the ramp vehicles to merge into the main line;
[0011] The ramp vehicle speed is planned through a multi-agent reinforcement learning model to obtain a speed instruction; after receiving the speed instruction, the ramp vehicle adjusts its speed to merge into the main line.
[0012] In combination with the first aspect, further, the specific steps of determining the target controlled traffic flow state according to the traffic flow data are:
[0013] The traffic flow data is combined with the basic traffic flow graph theory to determine the target control traffic flow state; the traffic flow data includes vehicle type, position, speed, headway and headway; the traffic flow parameters of the current traffic flow state of the main line include flow, density, speed, headway and headway, respectively, q A , k A 、v A 、h A and A ; The traffic flow parameters of the target controlled traffic flow state include flow, density, speed and headway, which are respectively represented by q c , k c 、v C and h C ; Among them, the current traffic flow state of the main line is called state A, and the target controlled traffic flow state is called state C;
[0014] According to the number of vehicles n in the main line fleet p 、The headway h of state A A Get the cycle length C c As shown below:
[0015] C c =h A ·n p ;
[0016] Among them, the number of vehicles in the main line fleet is n p 、State A headway h AIt is directly obtained through the road side detector and is a known quantity;
[0017] The duration of the gap, G, is:
[0018] G=n p h(h A -h C );
[0019] Among them, h C is the headway time of state C, G is the duration of the gap;
[0020] The number of ramp vehicles that can merge into the gap n G for:
[0021]
[0022] Among them, n G The number of ramp vehicles that can merge into the gap, g m is the minimum headway time required for a ramp vehicle to merge into the main line; then the cycle length C of a ramp vehicle merging into the main line c The number of vehicles passing through the ramp q ramp As shown below:
[0023]
[0024] Thus, the ideal target traffic control state headway h is obtained ideal , as shown below:
[0025] h ideal =h A -q ramp ·h A ·g m ;
[0026] Assuming that the total number of vehicles on the main line to be planned is m, in order to form an effective gap between vehicles on the main line to be planned, the following formula should be satisfied:
[0027]
[0028] Among them, [·] means rounding up, np min represents the minimum fleet size required to form sufficient gaps; according to the above formula, np min The value range of the vehicle headway h in state C is determined by the following formula: C The minimum value limit of ;
[0029]
[0030] Among them, h Cmin Indicates the headway h of state C CThe limit headway time;
[0031] Control the headway h of traffic conditions according to the ideal target ideal and the limit headway h Cmin , get the headway h of state C C , as shown below:
[0032] h C =max(h ideal ,h Cmin );
[0033] According to the obtained state C headway h C And the basic relationship expression of traffic flow, calculate the speed v of state C C , density k C And the flow rate q C , the basic relationship expression of traffic flow is as follows:
[0034]
[0035] q C =k C ·v C ;
[0036] Among them, T represents the safe headway, l is the vehicle length, v0 is the free flow speed, and s0 is the minimum safe stopping distance, all of which are known quantities.
[0037] In combination with the first aspect, further, a specific method for constructing the optimization model is as follows:
[0038] Taking minimizing the delay D of all vehicles passing through the merging point as the optimization goal, the established optimization model is as follows:
[0039]
[0040] Among them, ω main and ω ramp They represent the weights of vehicle delays on the main line and ramps, m refers to the number of vehicles on the main line to be planned, and n refers to the number of ramp vehicles that need to be coordinated. Refers to the delay of the i-th mainline vehicle, i = 1, 2, 3, ..., m, Refers to the delay of the jth ramp vehicle, j = 1, 2, 3, ..., n;
[0041] The objective function of the optimization model is:
[0042]
[0043] Among them, x AC It represents the distance required for the convoy to form a gap; t ACIndicates the time required to change from the current traffic flow state of the main line to the target controlled traffic state; h r It means that when there are no vehicles on the main lane, the ramp vehicles move at the average ramp speed v r Drive at a constant speed and maintain the headway; h C Indicates the headway time of the target controlled traffic state; h A The headway time indicating the current traffic flow status of the main line;
[0044] The optimization model needs to meet the following constraints: only networked autonomous driving vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point is greater than or equal to the distance x required for the convoy to form a gap. AC , which is 0 <x AC ≤d.
[0045] In combination with the first aspect, further, a specific method for obtaining the optimal formation solution is as follows:
[0046] The optimal formation scheme is determined by using genetic algorithm, which includes the following steps:
[0047] According to the type of vehicles to be planned and the distribution of vehicle positions, possible formation schemes are generated according to the finite exhaustive method to form the initial population P of the genetic algorithm = {P1, P2, ..., P z}, each individual P z Indicates a specific formation plan;
[0048] For each individual P z Substitute into the optimization model for calculation and obtain the fitness value f(P z );
[0049]
[0050] From the initial population P = {P1, P2, ..., P z} Select b individuals with high fitness as parents; randomly pair two parent individuals with a crossover probability p cross Determine whether to perform a crossover operation; if so, use a multi-point crossover method to randomly select multiple crossover points and alternately exchange the gene fragments between these points;
[0051] For each offspring individual, with mutation probability p mutation Determine whether to perform a mutation operation; if so, use a random mutation method to randomly select a gene position and change its value to a randomly generated valid value;
[0052] Recalculate the fitness of the new population generated by the crossover and mutation operations; determine whether the maximum number of iterations has been reached or the population fitness change is less than the set threshold. If so, terminate the iteration and select the individual P with the highest fitness.fmax As the optimal formation solution; if not, return to each individual P z Substitute into the optimization model for calculation and obtain the fitness value f(P z ) step to restart the iterative update of the population.
[0053] In combination with the first aspect, further, the specific steps of generating the control time and control instructions of the cooperative vehicles in real time are:
[0054] Using the basic relationship expression of traffic flow and the determined target to control the speed v of traffic flow state C , speed v C This is the deceleration control instruction. According to the principle of traffic wave propagation, when the cooperative vehicle receives the deceleration to v C After the instruction, the time required for the state A to change to the state C is t AC ;
[0055]
[0056] Among them, v AC Indicates the wave speed of transition from state A to state C;
[0057]
[0058] The distance x required for cooperative vehicles to form a gap AC ;
[0059]
[0060] The distance d between the cooperative vehicle and the merging point is detected by the detector on the main line side, and the control time t of sending the deceleration control command to the cooperative vehicle is control As shown in the following formula;
[0061]
[0062] When the control time t control Upon arrival, a deceleration control instruction is sent to the cooperating vehicles to ensure that the gap formed meets the merging needs of ramp vehicles while not disrupting the overall stability of the main line traffic flow.
[0063] In combination with the first aspect, further, a specific method for planning the driving speed of ramp vehicles is as follows:
[0064] According to the relative position and speed difference between the ramp vehicle and the preceding vehicle and the gap information of the main line, the optimal movement speed of the ramp vehicle is determined by using the pre-built multi-agent reinforcement learning model, that is, the speed instruction is obtained;
[0065] The speed instruction is issued to the intelligent connected vehicles on the ramp by using the road test processor, and the ramp vehicles travel at the optimal movement speed.
[0066] In a second aspect, the present invention provides a traffic merging control system, comprising:
[0067] The target controlled traffic flow state determination module is configured to collect traffic flow data of the main line and the ramp through the road side detector and the vehicle terminal; determine the target controlled traffic flow state according to the traffic flow data;
[0068] The optimal platooning solution module is configured to determine the intelligent networked vehicles that need to participate in the coordination through a pre-built optimization model according to the number of vehicles, vehicle types and vehicle position distribution in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, and the determined intelligent networked vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal platooning solution;
[0069] The instruction generation module is configured to generate control time and control instructions for the cooperative vehicles in real time according to the target control traffic flow state and the optimal formation plan; the control instructions are issued to the cooperative vehicles through the communication system of the road test processor, instructing the cooperative vehicles to slow down within the control time, thereby forming a sufficient gap for the ramp vehicles to merge into the main line;
[0070] The reinforcement learning module is configured to plan the driving speed of ramp vehicles through a multi-agent reinforcement learning model to obtain speed instructions; after receiving the speed instructions, the ramp vehicles adjust their speeds so as to merge into the main line.
[0071] In a third aspect, the present invention proposes a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned traffic merging control method are implemented.
[0072] In a fourth aspect, the present invention provides a computer device, comprising:
[0073] Memory for storing computer programs;
[0074] A processor is used to execute the computer program to implement the steps of the above-mentioned traffic merging control method.
[0075] In a fifth aspect, a computer program product includes a computer program, which implements the steps of the above-mentioned traffic merging control method when executed by a processor.
[0076] Compared with the prior art, the present invention has the following beneficial effects:
[0077] (1) The present invention uses the basic traffic flow diagram to guide cooperative vehicles to adjust their speed and position in real time, reserve appropriate merging gaps for ramp vehicles, reduce the risk of collision and scratching during the merging process, and improve driving safety in mixed traffic environments.
[0078] (2) Through active control and intelligent collaboration, the present invention enables vehicles to better cope with complex and changing traffic conditions, reduces the driver's decision-making burden, and improves the stability of autonomous driving vehicles in mixed traffic environments.
[0079] (3) The present invention uses a multi-agent reinforcement learning method to dynamically adjust the speed of ramp vehicles in real time to ensure that ramp vehicles merge into the main line smoothly and safely, thereby improving the overall traffic flow efficiency. It can effectively reduce delays and energy consumption in ramp merging, avoid frequent unstable behaviors such as braking and acceleration in traditional ramp merging, and achieve a greener and more environmentally friendly mode of transportation.
[0080] (4) The present invention has wide adaptability and can be applied to mixed traffic scenarios where manually driven vehicles and autonomous vehicles coexist. Through multi-agent reinforcement learning strategies, it can adapt to different traffic densities, traffic flows, and vehicle types, showing strong flexibility and practicality.
[0081] (5) The present invention is used for coordinated ramp merging control and has important application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 Schematic diagram of the flow of the traffic merging control method according to Embodiment 1 of the present invention;
[0083] Figure 2 This is a schematic diagram of information collection in Example 1 of the present invention;
[0084] Figure 3 This is a schematic diagram of target controlled traffic flow state in Example 1 of the present invention, wherein Figure 3 a is a schematic diagram of the main line gap and the spatiotemporal information of the convoy after the vehicle formation. Figure 3 b is a schematic diagram of converting into a spatiotemporal sequence to help ramp vehicles perform speed planning;
[0085] Figure 4 Schematic diagram of the reinforcement learning training process of Example 1 of the present invention. DETAILED DESCRIPTION
[0086] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. The embodiments of the present invention and the technical features in the embodiments may be combined with each other unless there is a conflict.
[0087] The term "and / or" is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " generally indicates that the related objects are in an "or" relationship.
[0088] Example 1
[0089] like Figure 1 As shown, this embodiment introduces a traffic merging control method, including the following steps:
[0090] Step S1: The traffic flow data of the main line and ramps are collected through roadside detectors and vehicle-mounted terminals. The traffic flow data is combined with the basic traffic flow graph theory to determine the target control traffic flow state; the traffic flow data includes vehicle type, position, speed, headway and headway distance, etc. Among them, the current traffic flow state of the main line is called state A, and its traffic flow parameters include flow rate, density, speed, headway and headway distance are q A , k A 、v A 、h A and A The target controlled traffic flow state is called state C, and its traffic flow parameters include flow, density, speed, and headway, which are represented by q c , k c 、v C and h C .
[0091] Step S2: According to the number of vehicles, vehicle types and vehicle location distribution in the main line to be planned, combined with the traffic flow parameters of the target controlled traffic flow state C, the intelligent connected vehicles that need to participate in the coordination are determined through a pre-built optimization model. The determined intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal formation plan.
[0092] Step S3: According to the target control traffic flow state C and the optimal formation plan, the control time and control instructions of the cooperative vehicles are generated in real time. The control instructions are issued to the cooperative vehicles through the communication system of the road test processor, instructing them to slow down within the control time, so as to form enough gaps for ramp vehicles to merge into the main line.
[0093] Step S4: The multi-agent reinforcement learning model pre-built in the road test processor plans the driving speed of the ramp vehicle according to the actual position and speed of each ramp vehicle, that is, obtains the speed instruction; after receiving the speed instruction from the road test processor, the ramp vehicle adjusts the speed to ensure that the ramp vehicle safely merges into the main line within the predetermined time window (i.e., the duration of the gap G).
[0094] The multi-agent reinforcement learning model of the present invention dynamically adjusts the driving routes of each ramp vehicle according to their actual position and speed to ensure the safety and smoothness of the merging process. The process is optimized through a real-time feedback mechanism to ensure that ramp vehicles merge safely within a given time window (i.e., the duration of the gap G).
[0095] In a specific implementation of this embodiment, in step S1, the method for determining the target controlled traffic flow state includes the following steps:
[0096] Step S11: According to the number of vehicles n in the main line fleet p 、State A headway h A Get the cycle length C c As shown below:
[0097] C c =h A ·n p
[0098] Among them, the number of vehicles in the main line fleet is n p 、State A headway h A It can be directly obtained through the road side detector and is a known quantity.
[0099] The duration G of the gap is:
[0100] G=n p ·(h A -h C )
[0101] Among them, h C is the headway time of the target controlled traffic flow state C, and G is the duration of the gap.
[0102] The number of ramp vehicles that can merge into the gap n G for:
[0103]
[0104] Among them, n G The number of ramp vehicles that can merge into the gap, g m is the minimum headway time required for a ramp vehicle to merge into the main line. Then the cycle length C of a ramp vehicle merging into the main line is c The amount of traffic on the ramp that can pass through q ramp As shown below:
[0105]
[0106] Therefore, the headway h of the ideal target control traffic state can be obtained: ideal , as shown below:
[0107] h ideal =h A -q ramp ·h A ·g m
[0108] It should be noted that the cycle length C c It can be compared to a traffic light cycle. For ramp vehicles, the duration of the gap G represents a green light, which allows them to pass; the time beyond the duration of the gap G represents a red light, which allows mainline vehicles to pass, but not ramp vehicles.
[0109] Here n G to q ramp The conversion is: Assuming G = 10s, C c =25s, n G This is the number of cars that can enter within 10 seconds, which is actually the number of cars that can enter within 25 seconds (because cars cannot pass during the 15 seconds of the red light phase). ramp , is to put "25s through n G car" is converted into "passing q within 1 hour ramp car".
[0110] Step S12: Assuming that the total number of vehicles in the main line to be planned is m, in order to form an effective gap between the vehicles in the main line to be planned, the following formula should be satisfied:
[0111]
[0112] Among them, [·] means rounding up, np min It represents the minimum fleet size required to form sufficient gaps. According to the above formula, np min The value range of , then the headway h of state C can be determined by the following formula C The minimum value limit h Cmin .
[0113]
[0114] Among them, h Cmin The headway h of state C is referred to as C The limiting headway time.
[0115] Step S13: Control the headway h of the traffic state C according to the ideal target obtained in step S11 ideal and the limit headway h obtained in step S12 Cmin , get the headway h of the target controlled traffic state C C , as shown below:
[0116] h C=max(h ideal ,h Cmin )
[0117] Step S14: h obtained from step S13 C , using the traffic flow basic diagram derived from the IDM model to describe the basic relationship of traffic flow, calculate the speed v of the target control traffic flow state C C , density k C , flow rate q C , the basic relationship expression of traffic flow is as follows:
[0118]
[0119] Wherein, T represents the safe headway, l is the vehicle length, v0 is the free flow speed, and s0 is the minimum safe stopping distance, which are all known quantities in the present invention.
[0120] In a specific implementation of this embodiment, in step S2, solving the optimal formation solution by optimizing the model includes the following steps:
[0121] Step S21: Identify vehicles to be planned in the vehicle detection area of the main line, and convert the detection results into binary strings using unsigned binary integers, where the networked autonomous driving vehicle CAV that can serve as a cooperative vehicle is represented by "1" and the manually driven vehicle is represented by "0", thereby obtaining a mathematical expression of the type of vehicles to be planned and the distribution of vehicle positions; the vehicle detection area is the total coverage area of several detectors set on the side of the main line;
[0122] Step S22: Taking minimizing the delay D of all vehicles passing through the merging point as the optimization goal, the established optimization model is as follows:
[0123]
[0124] Among them, ω main and ω ramp They represent the weights of vehicle delays on the main line and ramps, m refers to the number of vehicles on the main line to be planned, and n refers to the number of ramp vehicles that need to be coordinated. Refers to the delay of the i-th mainline vehicle, i = 1, 2, 3, ..., m, Refers to the delay of the jth ramp vehicle, j = 1, 2, 3, ..., n;
[0125] The delay of the i-th mainline vehicle is defined as the excess time spent by the mainline vehicle platoon, calculated as follows:
[0126]
[0127] in, is the arrival time of the i-th mainline vehicle at the merging point in the platoon, It refers to the arrival time of the i-th mainline vehicle arriving at the merging point without platooning;
[0128] Assume that the first vehicle to be planned in this planning is a cooperative vehicle, that is, i = 1, and the distance from the merging point is d; after receiving the signal to decelerate to v C After the instruction, the time required to change from state A to state C is t AC The expression is:
[0129]
[0130] Among them, v AC represents the wave speed from state A to state C, v A Indicates the vehicle speed in state A; s A Indicates the headway distance in state A:
[0131]
[0132] Therefore, the distance x required for the cooperating vehicles to form a gap AC It can be expressed as:
[0133]
[0134] When the cooperative vehicles move at a constant speed v A Driving dx AC The actual time it takes for the last vehicle in the main line to reach the merging point when it starts to slow down after a distance of Expressed as:
[0135]
[0136] In order to be universal, the actual time for the i-th vehicle in the main line to arrive at the merging point in the platoon is expressed as:
[0137]
[0138] If no platoon control operation is performed, the mainline vehicles will always remain in state A, moving at a constant speed v A Drive and maintain a headway distance of h A ; In this case, the arrival time of the i-th mainline vehicle at the merging point without platooning is It is expressed as:
[0139]
[0140] The delay of the jth ramp vehicle is defined as the arrival time of the jth ramp vehicle at the merging point in the platooning situation The arrival time of the jth ramp vehicle arriving at the merging point without platooning The difference is expressed as:
[0141]
[0142] For the first vehicle in the cooperative vehicle fleet, i.e., i=1, its arrival time at the merging point is the end time of the gap, i.e., the arrival time of the last ramp vehicle j=n at the merging point; therefore, we have:
[0143]
[0144] When there are no vehicles on the main lane, the ramp vehicles move at the average ramp speed v r Drive at a constant speed and keep the headway h r ; Referring to existing research, the arrival time of the jth ramp vehicle at the merging point is defined as:
[0145]
[0146] According to the motion plan, the ramp vehicle arrives at the merging point in state C, with a headway time of h. C , therefore, the actual arrival time of the ramp vehicle at the merging point is:
[0147]
[0148] Combined with the above derivation, the optimization model objective function can be written as:
[0149]
[0150] In addition, the optimization model must also meet the following constraints: only networked autonomous vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point must be greater than or equal to the distance x required for the convoy to form a gap AC , which is 0 <x AC ≤d.
[0151] Step S23: Based on the mathematical expression of the types of vehicles to be planned and the distribution of vehicle positions obtained in step S21, a genetic algorithm is used to determine the optimal formation plan.
[0152] In a specific implementation of this embodiment, in step S23, the optimal formation scheme is determined using a genetic algorithm, including the following process:
[0153] Step S231: Based on the mathematical expressions of the types of vehicles to be planned and the distribution of vehicle positions obtained in step S21, possible formation schemes are generated according to the finite exhaustive method to form an initial population P of the genetic algorithm = {P1, P2, ..., P z}, each individual P z Represents a specific formation plan.
[0154] Step S232: For each individual P z Substitute the optimization model established in step S22 for calculation to obtain the fitness value f(P z )
[0155]
[0156] Step S233: From the initial population P = {P1, P2, ..., P z} Select b individuals with high fitness as parents. Randomly pair two parent individuals with a crossover probability p cross Determine whether to perform a crossover operation. If so, use a multi-point crossover method to randomly select multiple crossover points and alternately exchange the gene fragments between these points.
[0157] Step S234: For each offspring individual, the mutation probability p mutation Determine whether to perform a mutation operation. If so, use the random mutation method to randomly select a gene position and change its value to a randomly generated valid value.
[0158] Step S235: Recalculate the fitness of the new population generated by the crossover operation and the mutation operation. Determine whether the maximum number of iterations has been reached or the population fitness change is less than the set threshold. If so, terminate the iteration and select the individual P with the highest fitness. fmax As the optimal formation solution; if not, return to step S232 and restart the iterative update of the population.
[0159] In a specific implementation of this embodiment, in step S3, the road test processor is a key component of the present invention, which receives vehicle dynamic data and traffic environment information in real time, runs the control method and system proposed by the present invention, and issues control instructions to the networked autonomous driving vehicle.
[0160] In a specific implementation of this embodiment, in step S3, the control time and control instructions of the cooperative vehicles are generated in real time, which specifically includes the following processes:
[0161] Step S31: Using the basic traffic flow relationship expression described in step S13 and the target determined in step S14 to control the speed v of the traffic flow state C C , speed v C This is the deceleration control instruction. According to the principle of traffic wave propagation, when the cooperative vehicle receives the deceleration to v C After the instruction, the time required for the state A to change to the state C is t AC .
[0162]
[0163] Among them, v ACIndicates the wave speed of transition from state A to state C.
[0164]
[0165] The distance x required for cooperative vehicles to form a gap AC .
[0166]
[0167] The distance d between the cooperative vehicle and the merging point is detected by the detector on the main line side, and the control time t of sending the deceleration control command to the cooperative vehicle is control As shown in the following formula.
[0168]
[0169] Step S32: When the control time t control Upon arrival, a deceleration control instruction is sent to the cooperating vehicles to ensure that the gap formed meets the merging needs of ramp vehicles while not disrupting the overall stability of the main line traffic flow.
[0170] In a specific implementation of this embodiment, in step S4, the method for planning the driving speed of ramp vehicles specifically includes the following steps:
[0171] Step S41: According to the relative position and speed difference between the ramp vehicle and the preceding vehicle and the gap information of the main line, the optimal movement speed of the ramp vehicle is determined by using the pre-built multi-agent reinforcement learning model, that is, the speed instruction is obtained.
[0172] Step S42: using the road test processor in step S3 to issue the speed instruction to the intelligent connected vehicle on the ramp, the ramp vehicle travels at the optimal movement speed.
[0173] It should be noted that the multi-agent reinforcement learning model in the present invention is an existing multi-agent reinforcement learning model. The present invention defines input, output and reward functions based on the existing multi-agent reinforcement learning model. The specific process of defining input, output and reward functions is shown in the specific steps of the following step S41.
[0174] like Figure 4 As shown, step S41 specifically includes the following steps:
[0175] Step S411, define the simulation environment of the traffic merging scene and initialize the agent. For the intelligent connected vehicle e on the ramp, use the communication function of the intelligent connected vehicle to obtain the status information of the intelligent connected vehicle
[0176]
[0177] Among them, xe represents the distance from the intelligent connected vehicle e to the convergence point, v e represents the driving speed of the intelligent connected vehicle e, x lea d er represents the distance from the preceding vehicle of the intelligent connected vehicle e to the merging point, v lea d er Indicates the speed of the vehicle ahead, G info Indicates the start and end time of the gap. The first four quantities can be obtained through the sensors configured in the intelligent connected vehicle. info It can be obtained through the interaction between the intelligent networked vehicle and the road test processor described in step S3.
[0178] Define the action space of the intelligent connected vehicle e:
[0179]
[0180] Among them, a e represents the acceleration of vehicle e, Represents the action space of the intelligent connected vehicle e.
[0181] Define the reward function of the intelligent connected vehicle e:
[0182] r e =w comfort r comfort +w gap r gap +w safety r safety +w effiency r effiency
[0183] Among them, w comfort ,w gap ,w safety ,w effiency They correspond to the comfort evaluation values r comfort , gap evaluation value r gap , safety assessment value r safety , efficiency evaluation value r effiency The weight, r e Represents the reward value. The five evaluation values are calculated by the following method:
[0184] (1) Comfort evaluation value r comfort It is defined as the absolute negative value of the jerk of the intelligent connected vehicle, and the calculation formula is:
[0185]
[0186] Among them, Δa e It represents the acceleration difference of the intelligent connected vehicle e at two adjacent moments, and Δt represents the time interval.
[0187] (2) If the intelligent connected vehicle merges into the gap, then r gap =10, otherwise r gap =0.
[0188] (3) If the intelligent connected vehicle e collides, then r safety = -10, otherwise r safety =0.
[0189] (4) Efficiency evaluation value r effiency It is defined as the negative value of the absolute difference between the vehicle speed and the expected speed, and the calculation formula is:
[0190] r efficiency =-|v e -v desired |
[0191] Among them, v e is the speed of the intelligent connected vehicle e, v desired is the expected speed of the vehicle on the ramp.
[0192] Step S412: At each time step, the agent obtains current state information from the environment And select the action space based on the current state information Action Space Including adjusting vehicle speed.
[0193] Step S413: The agent executes the selected action space The environment changes, and the reward value r is fed back e .
[0194] In step S414, the agent updates its strategy based on the obtained rewards and new status, in order to obtain higher cumulative rewards in the future.
[0195] Step S415, determine whether the strategy for updating the new state converges or reaches the expected performance, if so, end the training; otherwise, return to step S412.
[0196] In a specific implementation of this embodiment, as Figure 2 As shown, the present invention uses several detectors placed on the road side to detect the traffic flow data of the main line and the ramp. Specifically, the ramp detector A1 is used to monitor the traffic flow of the entrance ramp, and two detectors are installed on the main road side, namely, detector B1 and detector B2. Detector B1 is used to measure the main line traffic flow and record the corresponding macro state B, and detector B2 is used to identify the type and location of the vehicle to be planned.
[0197] In a specific implementation of this embodiment, as Figure 3As shown in FIG. 3 , step S3 further clarifies the control time of the cooperative vehicle and the start and end time of the gap formed by the cooperative vehicle according to the target control traffic flow state and the optimal formation plan, and sends a deceleration command to the cooperative vehicle when the control time arrives, so that the traffic flow of the main line forms a separated fleet and a gap that can be used by ramp vehicles. When the cooperative vehicle decelerates, a fleet is formed in the gap formation area on the main line. This process is iteratively applied to control the deceleration of other cooperative vehicles, and gaps are cyclically generated for the entrance ramp vehicles to use. Therefore, as shown in FIG. (3a), the spatiotemporal information on the main line is divided into two categories: occupied spatiotemporal areas (i.e., areas occupied by fleets) and unoccupied spatiotemporal areas (i.e., areas occupied by gaps). This information can be analogized to the phase of a traffic light, where the occupied spatiotemporal area represents the red phase and the unoccupied spatiotemporal area represents the green phase. Next, the control center captures these two types of information and encodes them into a spatiotemporal sequence, which is then transmitted to all vehicles on the entrance ramp, as shown in FIG. (3b).
[0198] Example 2
[0199] Based on the same inventive concept as that of Embodiment 1, this embodiment introduces a traffic merging control system, including:
[0200] The target controlled traffic flow state determination module is configured to collect traffic flow data of the main line and the ramp through the road side detector and the vehicle terminal; determine the target controlled traffic flow state according to the traffic flow data;
[0201] The optimal platooning solution module is configured to determine the intelligent networked vehicles that need to participate in the coordination through a pre-built optimization model according to the number of vehicles, vehicle types and vehicle position distribution in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, and the determined intelligent networked vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal platooning solution;
[0202] The instruction generation module is configured to generate control time and control instructions for the cooperative vehicles in real time according to the target control traffic flow state and the optimal formation plan; the control instructions are issued to the cooperative vehicles through the communication system of the road test processor to instruct the cooperative vehicles to slow down within the control time, thereby forming a sufficient gap for the ramp vehicles to merge into the main line;
[0203] The reinforcement learning module is configured to plan the driving speed of ramp vehicles through a multi-agent reinforcement learning model to obtain speed instructions; after receiving the speed instructions, the ramp vehicles adjust their speeds so as to merge into the main line.
[0204] Example 3
[0205] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned traffic merging control method are implemented.
[0206] Example 4
[0207] Based on the same inventive concept as other embodiments, this embodiment introduces a computer device, including: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the above-mentioned traffic merging control method.
[0208] Example 5
[0209] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including a computer program, which implements the steps of the above-mentioned traffic merging control method when executed by a processor.
[0210] The present invention uses the target control traffic flow state to guide cooperative vehicles to actively reserve appropriate merging gaps, and uses a multi-agent reinforcement learning strategy to dynamically adjust the driving behavior of ramp vehicles in real time to ensure that ramp vehicles merge smoothly and safely into the main line. This method effectively reduces delays and energy consumption in ramp merging, avoids frequent braking, acceleration and other unstable behaviors, and improves the efficiency, safety and environmental protection of the transportation system.
[0211] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0212] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0213] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0214] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0215] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make many forms under the guidance of the present invention, which all fall within the protection of the present invention.
Claims
1. A traffic merging control method, characterized in that: The following steps are involved: Collecting traffic flow data of the main line and ramps through roadside detectors and vehicle-mounted terminals; determining a target control traffic flow state based on the traffic flow data; According to the number of vehicles, vehicle types and vehicle position distribution in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, the intelligent networked vehicles that need to participate in the coordination are determined through a pre-built optimization model, and the determined intelligent networked vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal platooning solution; According to the target control traffic flow state and the optimal formation plan, the control time and control instructions of the cooperative vehicles are generated in real time; the control instructions are issued to the cooperative vehicles through the communication system of the road test processor, instructing the cooperative vehicles to slow down within the control time, thereby forming a sufficient gap for the ramp vehicles to merge into the main line; The ramp vehicle speed is planned through a multi-agent reinforcement learning model to obtain speed instructions; After receiving the speed instruction, the ramp vehicle adjusts its speed so as to merge into the main line.
2. The traffic merging control method according to claim 1, characterized in that: The specific steps of determining the target controlled traffic flow state according to the traffic flow data are: The traffic flow data is combined with the basic traffic flow graph theory to determine the target control traffic flow state; the traffic flow data includes vehicle type, position, speed, headway and headway; the traffic flow parameters of the current traffic flow state of the main line include flow, density, speed, headway and headway, respectively, q A , k A 、v A 、h A and A ; The traffic flow parameters of the target controlled traffic flow state include flow, density, speed and headway, which are respectively represented by q c , k c 、v C and h C ; Among them, the current traffic flow state of the main line is called state A, and the target controlled traffic flow state is called state C; According to the number of vehicles n in the main line fleet p 、The headway h of state A A Get the cycle length C c As shown below: C c =h A ·n p ; Among them, the number of vehicles in the main line fleet is n p 、State A headway h A It is directly obtained through the road side detector and is a known quantity; The duration of the gap, G, is: G=n p ·(h A -h C ); Among them, h C is the headway time of state C, G is the duration of the gap; The number of ramp vehicles that can merge into the gap n G for: Among them, n C The number of ramp vehicles that can merge into the gap, g m is the minimum headway time required for a ramp vehicle to merge into the main line; then the cycle length C of a ramp vehicle merging into the main line c The number of vehicles passing through the ramp q ramp As shown below: Thus, the ideal target traffic control state headway h is obtained ideal , as shown below: h ideal =h A -q ramp ·h A ·g m ; Assuming that the total number of vehicles on the main line to be planned is m, in order to form an effective gap between vehicles on the main line to be planned, the following formula should be satisfied: Among them, [·] means rounding up, np min represents the minimum fleet size required to form sufficient gaps; according to the above formula, np min The value range of the vehicle headway h in state C is determined by the following formula: C The minimum value limit of ; Among them, h Cmin Indicates the headway h of state C C The limit headway time; Control the headway h of traffic conditions according to the ideal target ideal and the limit headway h Cmin , get the headway h of state C C , as shown below: h C =max(h ideal ,h Cmin ); According to the obtained state C headway h C And the basic relationship expression of traffic flow, calculate the speed v of state C C , density k C And the flow rate q C , the basic relationship expression of traffic flow is as follows: q C =k C ·v C ; Among them, T represents the safe headway, l is the vehicle length, v0 is the free flow speed, and s0 is the minimum safe stopping distance, all of which are known quantities.
3. The traffic merging control method according to claim 1, characterized in that: The specific method of constructing the optimization model is as follows: Taking minimizing the delay D of all vehicles passing through the merging point as the optimization goal, the established optimization model is as follows: Among them, ω main and ω ramp They represent the weights of vehicle delays on the main line and ramps, m refers to the number of vehicles on the main line to be planned, and n refers to the number of ramp vehicles that need to be coordinated. Refers to the delay of the i-th mainline vehicle, i = 1, 2, 3, ..., m, Refers to the delay of the jth ramp vehicle, j = 1, 2, 3, ..., n; The objective function of the optimization model is: Among them, x AC It represents the distance required for the convoy to form a gap; t AC Indicates the time required to change from the current traffic flow state of the main line to the target controlled traffic state; h r It means that when there are no vehicles on the main lane, the ramp vehicles move at the average ramp speed v r Drive at a constant speed and maintain the headway; h C Indicates the headway time of the target controlled traffic state; h A The headway time indicating the current traffic flow status of the main line; The optimization model needs to meet the following constraints: only networked autonomous driving vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point is greater than or equal to the distance x required for the convoy to form a gap. AC , which is 0 <x AC ≤d.
4. The traffic merging control method according to claim 3, characterized in that: The specific method to obtain the optimal formation plan is: The optimal formation scheme is determined by using genetic algorithm, which includes the following steps: According to the type of vehicles to be planned and the distribution of vehicle positions, possible formation schemes are generated according to the finite exhaustive method to form the initial population P of the genetic algorithm = {P1, P2, ..., P z }, each individual P z Indicates a specific formation plan; For each individual P z Substitute into the optimization model for calculation and obtain the fitness value f(P z ); From the initial population P = {P1, P2, ..., P z } Select b individuals with high fitness as parents; randomly pair two parent individuals with a crossover probability p cross Determine whether to perform a crossover operation; if so, use a multi-point crossover method to randomly select multiple crossover points and alternately exchange the gene fragments between these points; For each offspring individual, with mutation probability p mutation Determine whether to perform a mutation operation; if so, use a random mutation method to randomly select a gene position and change its value to a randomly generated valid value; Recalculate the fitness of the new population generated by the crossover and mutation operations; determine whether the maximum number of iterations has been reached or the population fitness change is less than the set threshold. If so, terminate the iteration and select the individual P with the highest fitness. fmax As the optimal formation solution; if not, return to each individual P z Substitute into the optimization model for calculation and obtain the fitness value f(P z ) step to restart the iterative update of the population.
5. The traffic merging control method according to claim 2, characterized in that: The specific steps for generating the control time and control instructions of cooperative vehicles in real time are as follows: Using the basic relationship expression of traffic flow and the determined target to control the speed v of traffic flow state C , speed v C This is the deceleration control instruction. According to the principle of traffic wave propagation, when the cooperative vehicle receives the deceleration to v C After the instruction, the time required for the state A to change to the state C is t AC ; Among them, v AC Indicates the wave speed of transition from state A to state C; The distance x required for cooperative vehicles to form a gap AC ; The distance d between the cooperative vehicle and the merging point is detected by the detector on the main line side, and the control time t of sending the deceleration control command to the cooperative vehicle is control As shown in the following formula; When the control time t control Upon arrival, a deceleration control instruction is sent to the cooperative vehicle.
6. The traffic merging control method according to claim 2, characterized in that: The specific method for planning the vehicle speed on the ramp is: According to the relative position and speed difference between the ramp vehicle and the preceding vehicle and the gap information of the main line, the optimal movement speed of the ramp vehicle is determined by using the pre-built multi-agent reinforcement learning model, that is, the speed instruction is obtained; The speed instruction is issued to the intelligent connected vehicles on the ramp by using the road test processor, and the ramp vehicles travel at the optimal movement speed.
7. A traffic merging control system, characterized in that: include: A target controlled traffic flow state determination module is configured to collect traffic flow data of the main line and ramps through roadside detectors and vehicle-mounted terminals; Determining a target controlled traffic flow state according to the traffic flow data; The optimal platooning solution module is configured to determine the intelligent networked vehicles that need to participate in the coordination through a pre-built optimization model according to the number of vehicles, vehicle types and vehicle position distribution in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, and the determined intelligent networked vehicles that need to participate in the coordination are called cooperative vehicles, so as to obtain the optimal platooning solution; The instruction generation module is configured to generate control time and control instructions for the cooperative vehicles in real time according to the target control traffic flow state and the optimal formation plan; the control instructions are issued to the cooperative vehicles through the communication system of the road test processor, instructing the cooperative vehicles to slow down within the control time, thereby forming a sufficient gap for the ramp vehicles to merge into the main line; The reinforcement learning module is configured to plan the driving speed of ramp vehicles through a multi-agent reinforcement learning model to obtain speed instructions; after receiving the speed instructions, the ramp vehicles adjust their speeds so as to merge into the main line.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the traffic merging control method described in any one of claims 1 to 6 are implemented.
9. A computer device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program to implement the steps of the traffic merging control method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the traffic merging control method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Vehicle variable speed limit control optimization method, system, medium and equipment
CN115862322A
Dynamic cooperative confluence control method and system suitable for mixed traffic flow
CN117409564A
Motorcade length control method and system for network automatic driving
CN117437804A
Integrated control method for confluence bottleneck area based on dynamic division of lane-level cells
CN117671955A
Dynamic platoon formation method under mixed autonomous vehicles flow
US20220351625A1
Cited By
Cooperative control method for vehicles in diverging and converging areas
CN120877527A
A merging area vehicle cooperative control method
CN120877527B
Expressway ramp confluence control method and device based on formation cooperation
CN122135577A