A traffic merging control method and system
By collecting traffic flow data and utilizing optimization models and multi-agent reinforcement learning models, the speed and position of intelligent connected vehicles are adjusted in real time, solving the problem of unstable traffic flow during ramp merging and achieving smooth merging of vehicles on ramps and efficient management of the traffic system.
Patent Information
- Application Number
- CN202411879556.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing ramp merging control methods cannot effectively cope with dynamic changes in traffic flow, leading to unstable traffic flow, frequent braking and acceleration, and other unstable behaviors, which cannot meet the needs of intelligent and efficient management of highway traffic.
Traffic flow data is collected by roadside detectors and vehicle terminals. Using optimization models and multi-agent reinforcement learning models, control and speed commands for cooperating vehicles are generated in real time to guide intelligent connected vehicles to coordinate and adjust their speed and position to form appropriate merging gaps and ensure that vehicles on ramps smoothly merge into the main line.
It reduces delays and energy consumption during ramp merging, improves traffic flow safety and efficiency, adapts to different traffic densities and volumes, reduces the decision-making burden on drivers, and achieves a greener and more environmentally friendly mode of transportation.
Smart Images

Figure CN119942780B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent transportation systems technology, and relates to a traffic merging control method and system. Background Technology
[0002] With rapid urbanization and industrialization, merging ramps on highways and urban expressways have become high-risk areas for congestion. In these areas, traffic flows from the main road and ramps need to merge within a limited space; improper merging can lead to traffic instability, further causing congestion, traffic accidents, and a host of other problems. Traditional traffic management models, such as simple traffic signal control or variable speed limit control, are no longer sufficient to meet the growing and increasingly complex demands of highway traffic. Therefore, developing more precise and intelligent control technologies for merging ramps has become an urgent task.
[0003] Ramp merging control technology represents a significant breakthrough in traffic informatization and intelligentization in recent years. This technology primarily utilizes real-time data acquisition and advanced data analysis algorithms to detect and control traffic flow on ramps and main roads. It comprehensively analyzes multi-dimensional information such as vehicle speed, traffic volume, and vehicle spacing within the merging area, and determines appropriate lane change and speed adjustment control strategies through modeling. Therefore, in-depth research into ramp merging control technology will not only promote the development of highway traffic management towards a higher level of informatization and intelligentization but will also have a profound impact on alleviating traffic congestion and improving road transport efficiency and safety.
[0004] Existing ramp merging control methods often fail to address the dynamic merging problem at both the macro and micro traffic flow levels. The introduction of target-controlled traffic signals offers a novel solution, but in mixed traffic flows, interference from human driving behavior can render these signals ineffective, necessitating further research into vehicle cooperation and interactions among different types of traffic participants. Summary of the Invention
[0005] The purpose of this invention is to provide a traffic merging control method and system that can effectively reduce delays and energy consumption during ramp merging and avoid frequent braking, acceleration and other unstable behaviors.
[0006] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.
[0007] In a first aspect, the present invention proposes a traffic merging control method, comprising the following steps:
[0008] Traffic flow data for the mainline and ramps is collected using roadside detectors and on-board terminals; the target traffic flow state is determined based on the traffic flow data.
[0009] Based on the number of vehicles, vehicle types, and vehicle location distribution in the main line to be planned, and combined with the traffic flow parameters of the target controlled traffic flow state, the intelligent connected vehicles that need to participate in the coordination are determined through a pre-built optimization model. The determined intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, thereby obtaining the optimal formation scheme.
[0010] Based on the target traffic flow status and optimal platooning scheme, control time and control instructions for cooperating vehicles are generated in real time. The control instructions are issued to the cooperating vehicles through the communication system of the road test processor, guiding the cooperating vehicles to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line.
[0011] The speed of vehicles on the ramp is planned by a multi-agent reinforcement learning model, and a speed command is generated; after receiving the speed command, the vehicles on the ramp adjust their speed in order to merge into the main line.
[0012] In conjunction with the first aspect, the specific steps for determining the target control traffic flow state based on the traffic flow data are as follows:
[0013] Using traffic flow data combined with basic traffic flow graph theory, the target control traffic flow state is determined. The traffic flow data includes vehicle type, location, speed, headway, and spacing. The traffic flow parameters for the current mainline traffic flow state include flow rate, density, speed, headway, and spacing, each represented by q. A k A v A h A and s A The traffic flow parameters for controlling traffic flow state include flow rate, density, speed, and headway, which are respectively represented as q. c k c v C and h C The current traffic flow state of the main line is referred to as state A, and the target controlled traffic flow state is referred to as state C.
[0014] Based on the number of vehicles n in the mainline convoy p The headway h of vehicle in state A A Obtain the period length C c As shown in the following formula:
[0015] C c =h A ·n p ;
[0016] Among them, the number of vehicles in the main line convoy is n p State A: Headway h AThe data is obtained directly from the roadside detector and is a known quantity.
[0017] The duration G of the gap is:
[0018] G = n p h(h A -h C );
[0019] Among them, h C G represents the headway in state C, and G represents the duration of the gap.
[0020] The number of ramp vehicles that can merge into within the gap, n G for:
[0021]
[0022] Where, n G The number of vehicles that can merge into the ramp within the gap, g m Let C be the minimum headway required for a vehicle to merge onto the main line from a ramp; then the cycle length C of a vehicle merging onto a ramp is... c Traffic flow q of the ramp passing through ramp As shown in the following formula:
[0023]
[0024] Therefore, the headway h of the ideal target traffic control state is obtained. ideal As shown below:
[0025] h ideal =h A -q ramp ·h A ·g m ;
[0026] Assuming the total number of vehicles on the planned main line is m, in order to create effective gaps between vehicles on the planned main line, the following formula should be satisfied:
[0027]
[0028] Where [·] represents rounding up, np min This represents the minimum convoy size required to form sufficient gaps; np can be obtained from the above formula. min The range of values for h is then determined by the following formula for the headway h in state C. C The minimum value limit;
[0029]
[0030] Among them, h Cmin The headway h represents the distance between the front and rear of the vehicle in state C. CThe clearance between the front and rear of the vehicle;
[0031] Headway h based on the ideal target for controlling traffic conditions ideal Distance h between the front and the clearance vehicle Cmin The headway h of the vehicle in state C is obtained. C As shown in the following formula:
[0032] h C =max(h ideal ,h Cmin );
[0033] Based on the headway h obtained from state C C Based on the basic traffic flow relationship expression, the velocity v in state C is calculated. C Density k C and traffic q C The basic relationship expression of traffic flow is as follows:
[0034]
[0035] q C =k C ·v C ;
[0036] Where T represents the safe headway, l is the vehicle length, v0 is the free flow velocity, and s0 is the minimum safe stopping distance, all of which are known quantities.
[0037] In conjunction with the first aspect, the specific method for constructing the optimization model is as follows:
[0038] The optimization model is established with the objective of minimizing the delay D of all vehicles passing through the merging point as the optimization objective:
[0039]
[0040] Where, ω main and ω ramp These represent the weights of vehicle delays on the main line and on the ramps, respectively. 'm' refers to the number of vehicles on the main line to be planned, and 'n' refers to the number of vehicles merging onto the ramps that need to be coordinated. This refers to the delay of the i-th mainline vehicle, where i = 1, 2, 3, ..., m. The delay refers to the delay of the j-th vehicle on the ramp, where j = 1, 2, 3, ..., n;
[0041] The objective function for optimizing the model is:
[0042]
[0043] Where, x AC t represents the distance required for the convoy to form gaps; ACThis indicates the time required for the current traffic flow state on the main line to transition to the target controlled traffic state; h r This indicates that when there are no vehicles in the main lane, vehicles on the ramp travel at the average ramp speed v. r Maintaining a constant speed and keeping a constant headway; h C Headway, representing the time difference between vehicles in the target traffic control state; h A Headway indicates the current traffic flow status on the main line;
[0044] The optimization model must satisfy the following constraints: only connected autonomous vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point is greater than or equal to the distance x required for the convoy to form gaps. AC That is, 0 <x AC ≤d.
[0045] In conjunction with the first aspect, the specific method for obtaining the optimal formation scheme is as follows:
[0046] Determining the optimal formation scheme using a genetic algorithm includes the following steps:
[0047] Based on the vehicle types and their location distribution, possible formation schemes are generated using a finite exhaustive search method, forming the initial population P = {P1, P2, ..., P} for the genetic algorithm. z}, each individual P z This represents a specific formation scheme;
[0048] For each individual P z Substituting the values into the optimization model, we obtain the fitness value f(P). z );
[0049]
[0050] From the initial population P = {P1, P2, ..., P...} z Select b individuals with high fitness as parents; randomly pair two parent individuals with crossover probability p. cross Determine whether to perform a crossover operation; if so, use a multi-point crossover method, randomly select multiple crossover points, and alternately exchange gene segments between these points.
[0051] For each offspring individual, with mutation probability p mutation Determine whether to perform a mutation operation; if so, use a random mutation method to randomly select a gene location and change its value to a randomly generated valid value.
[0052] Based on the new population generated by crossover and mutation operations, recalculate the fitness; determine whether the maximum number of iterations has been reached or the change in population fitness is less than a set threshold. If so, terminate the iteration and select the individual P with the highest fitness.fmax If the optimal formation scheme is not found, then return to the optimal formation scheme for each individual P. z Substituting the values into the optimization model, we obtain the fitness value f(P). z The population is then iteratively updated again using the following steps.
[0053] In conjunction with the first aspect, the specific steps for generating real-time control time and control commands for cooperative vehicles are as follows:
[0054] The speed v of traffic flow is controlled using the basic traffic flow relationship expression and a defined objective. C speed v C This is a deceleration control command. Based on the principle of traffic wave propagation, when the cooperating vehicle receives the command, it decelerates to v. C The time t required for the transition from state A to state C after receiving the instruction. AC ;
[0055]
[0056] Among them, v AC This represents the wave speed during the transition from state A to state C;
[0057]
[0058] The distance x required for the cooperative vehicles to form a gap AC ;
[0059]
[0060] The distance d between the cooperating vehicle and the merging point is detected by the detector on the main line side. The control time t for sending the deceleration control command to the cooperating vehicle is then determined. control As shown in the following formula;
[0061]
[0062] When control time t control Upon arrival, a deceleration control command is sent to the cooperating vehicles to ensure that the gap formed meets the merging needs of the ramp vehicles, while not disrupting the overall stability of the mainline traffic flow.
[0063] In conjunction with the first aspect, the specific method for planning the vehicle speed on ramps is as follows:
[0064] Based on the relative position and speed difference between the ramp vehicle and the vehicle in front, as well as the gap information on the main line, the optimal speed of the ramp vehicle is determined by a pre-built multi-agent reinforcement learning model, thus obtaining the speed command.
[0065] The speed command is issued to the intelligent connected vehicles on the ramp using the road test processor, and the ramp vehicles travel at the optimal speed.
[0066] Secondly, this invention proposes a traffic merging control system, comprising:
[0067] The target control traffic flow state determination module is configured to collect traffic flow data of the mainline and ramps through roadside detectors and on-board terminals; and determine the target control traffic flow state based on the traffic flow data.
[0068] The optimal formation scheme module is configured to determine the intelligent connected vehicles that need to participate in the coordination based on the number of vehicles, vehicle types, and vehicle location distribution in the main line to be planned, combined with the traffic flow parameters of the target controlled traffic flow state, through a pre-built optimization model. The determined intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, thereby obtaining the optimal formation scheme.
[0069] The instruction generation module is configured to generate control time and control instructions for cooperating vehicles in real time based on the target control traffic flow state and optimal formation scheme; the control instructions are issued to the cooperating vehicles through the communication system of the road test processor, instructing the cooperating vehicles to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line;
[0070] The reinforcement learning module is configured to plan the speed of vehicles on the ramp through a multi-agent reinforcement learning model and generate speed commands; after receiving the speed commands, the vehicles on the ramp adjust their speed to merge into the main line.
[0071] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the traffic merging control method described above.
[0072] Fourthly, the present invention provides a computer device comprising:
[0073] Memory, used to store computer programs;
[0074] A processor is used to execute the computer program to implement the steps of the traffic merging control method described above.
[0075] Fifthly, a computer program product includes a computer program that, when executed by a processor, implements the steps of the traffic merging control method described above.
[0076] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0077] (1) This invention uses a basic traffic flow diagram to guide cooperating vehicles to adjust their speed and position in real time, reserve appropriate merging gaps for ramp vehicles, reduce the risk of collisions and scrapes during merging, and improve driving safety in mixed traffic environments.
[0078] (2) Through active control and intelligent collaboration, the present invention enables vehicles to better cope with complex and ever-changing traffic conditions, reduce the decision-making burden of drivers, and improve the stability of autonomous vehicles in mixed traffic environments.
[0079] (3) This invention uses a multi-agent reinforcement learning method to dynamically adjust the speed of vehicles on the ramp in real time, ensuring that vehicles on the ramp merge smoothly and safely into the main line, improving the overall traffic flow efficiency, effectively reducing delays and energy consumption in ramp merging, avoiding frequent braking and acceleration and other unstable behaviors in traditional ramp merging, and achieving a greener and more environmentally friendly mode of transportation.
[0080] (4) This invention has broad adaptability and can be applied to mixed traffic scenarios where manually driven vehicles and autonomous vehicles coexist. Through multi-agent reinforcement learning strategies, it can adapt to different traffic densities, traffic flow and vehicle types, showing strong flexibility and practicality.
[0081] (5) This invention is used for coordinated ramp merging control and has important application prospects. Attached Figure Description
[0082] Figure 1 This is a schematic flowchart of the traffic merging control method according to Embodiment 1 of the present invention;
[0083] Figure 2 This is a schematic diagram of information collection in Embodiment 1 of the present invention;
[0084] Figure 3 This is a schematic diagram of the target control traffic flow state in Embodiment 1 of the present invention, wherein... Figure 3 a is a schematic diagram of the mainline gaps and the spatiotemporal information of the vehicle platoon after formation. Figure 3 b is a schematic diagram illustrating the conversion into a spatiotemporal sequence to aid in speed planning for vehicles on the ramp;
[0085] Figure 4 This is a schematic diagram of the reinforcement learning training process in Embodiment 1 of the present invention. Detailed Implementation
[0086] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0087] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0088] Example 1
[0089] like Figure 1 As shown in the figure, this embodiment introduces a traffic merging control method, including the following steps:
[0090] Step S1: Collect traffic flow data for the mainline and ramps using roadside detectors and onboard terminals. Utilize this traffic flow data in conjunction with basic traffic flow graph theory to determine the target control traffic flow state. Traffic flow data includes vehicle type, location, speed, headway, and distance between vehicles. The current traffic flow state on the mainline is referred to as state A, and its traffic flow parameters include flow rate, density, speed, headway, and distance between vehicles, each q. A k A v A h A and s A The target controlled traffic flow state is called state C, and its traffic flow parameters include flow rate, density, speed, and headway, which are respectively represented by q. c k c v C and h C .
[0091] Step S2: Based on the number of vehicles, vehicle types, and vehicle location distribution in the main line to be planned, and combined with the traffic flow parameters of the target control traffic flow state C, the intelligent connected vehicles that need to participate in the coordination are determined through a pre-built optimization model. The intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, thereby obtaining the optimal formation scheme.
[0092] Step S3: Based on the target traffic flow state C and the optimal platooning scheme, generate control times and control instructions for cooperating vehicles in real time. The control instructions are issued to the cooperating vehicles through the communication system of the road test processor, guiding them to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line.
[0093] Step S4: The multi-agent reinforcement learning model pre-built in the road test processor plans the driving speed of each ramp vehicle based on its actual position and speed, thus obtaining a speed command. After receiving the speed command from the road test processor, the ramp vehicle adjusts its speed to ensure that the ramp vehicle safely merges into the main line within a predetermined time window (i.e., the duration G of the gap).
[0094] The multi-agent reinforcement learning model of this invention dynamically adjusts the routes of each ramp vehicle based on its actual position and speed, ensuring the safety and smoothness of the merging process. This process is optimized through a real-time feedback mechanism to ensure that ramp vehicles merge safely within a predetermined time window (i.e., the duration G of the gap).
[0095] In one specific implementation of this embodiment, step S1, the method for determining the target controlled traffic flow state includes the following steps:
[0096] Step S11: Based on the number of vehicles n in the main convoy p State A: Headway h A Obtain the period length C c As shown in the following formula:
[0097] C c =h A ·n p
[0098] Among them, the number of vehicles in the main line convoy is n p State A: Headway h A It can be obtained directly from roadside detectors and is a known quantity.
[0099] The duration G of the gap is:
[0100] G = n p ·(h A -h C )
[0101] Among them, h C G represents the headway of the vehicles in the target traffic flow state C, and G represents the duration of the gap.
[0102] The number of ramp vehicles that can merge into within the gap, n G for:
[0103]
[0104] Where, n G The number of vehicles that can merge into the ramp within the gap, g m Let C be the minimum headway required for a vehicle to merge from a ramp onto the main line. Then, the cycle length C for a vehicle merging from a ramp is... c Traffic flow q of the ramps that can be accessed ramp As shown in the following formula:
[0105]
[0106] Therefore, the headway h of the ideal target traffic control state can be obtained. ideal As shown below:
[0107] h ideal =h A -q ramp ·h A ·g m
[0108] It should be noted that: the period length C c This can be compared to a traffic light cycle. For vehicles on ramps, the duration G of the gap represents a green light, and they can proceed; the time outside of the duration G represents a red light, and vehicles on the main line can proceed, while ramp vehicles cannot.
[0109] Here n G to q ramp The transformation refers to: assuming G = 10s, C c =25s, n G This refers to the number of vehicles that can merge within these 10 seconds, which is actually the same as the number of vehicles that can merge within these 25 seconds (because traffic cannot pass during the 15 seconds of the red light phase). This is converted to traffic flow q. ramp It means passing "25s through n" G "Vehicle" converted to "passing through q within 1 hour" ramp A vehicle.
[0110] Step S12: Assuming the total number of vehicles in the main line to be planned is m, in order to form an effective gap between vehicles in the main line to be planned, the following formula should be satisfied:
[0111]
[0112] Where [·] represents rounding up, np min This represents the minimum convoy size required to create sufficient clearance. Based on the above formula, np can be obtained. min Given the range of values for , the headway h in state C can be determined by the following formula. C The minimum value limit h Cmin .
[0113]
[0114] Among them, h Cmin The headway h, referred to as state C C The distance between the front and rear of the vehicle.
[0115] Step S13: Based on the ideal target traffic state C obtained in step S11, determine the headway h. ideal The clearance h obtained in step S12 Cmin The headway h of the target traffic state C is obtained. C As shown in the following formula:
[0116] h C=max(h ideal ,h Cmin )
[0117] Step S14: h obtained from step S13 C The basic traffic flow graph derived from the IDM model is used to describe the basic relationships of traffic flow, and the speed v of the target controlled traffic flow state C is calculated. C density k C Traffic q C The basic relationship expression of traffic flow is as follows:
[0118]
[0119] Where T represents the safe headway, l is the vehicle length, v0 is the free flow velocity, and s0 is the minimum safe stopping distance, all of which are known quantities in this invention.
[0120] In one specific implementation of this embodiment, step S2, solving for the optimal formation scheme through the optimization model, includes the following steps:
[0121] Step S21: Identify vehicles waiting to be planned in the vehicle detection area of the main line, and convert the detection results into binary strings using unsigned binary integers, where connected autonomous vehicles (CAVs) that can be used as cooperative vehicles are represented as "1" and manually driven vehicles are represented as "0", thereby obtaining a mathematical expression of the types of vehicles to be planned and the distribution of vehicle positions; the vehicle detection area is the total coverage area of several detectors set on the side of the main line.
[0122] Step S22: With minimizing the delay D of all vehicles passing through the merging point as the optimization objective, the optimization model is established as follows:
[0123]
[0124] Where, ω main and ω ramp These represent the weights of vehicle delays on the main line and ramps, respectively. 'm' refers to the number of vehicles on the main line to be planned, and 'n' refers to the number of vehicles merging from the ramps that need to be coordinated. This refers to the delay of the i-th mainline vehicle, where i = 1, 2, 3, ..., m. The delay refers to the delay of the j-th vehicle on the ramp, where j = 1, 2, 3, ..., n;
[0125] The delay of the i-th mainline vehicle is defined as the excess time spent in mainline vehicle platooning, calculated by the following formula:
[0126]
[0127] in, It is the arrival time of the i-th mainline vehicle when it arrives at the merging point in platooning mode. This refers to the arrival time of the i-th mainline vehicle when it arrives at the merging point without being convoyed.
[0128] Assume the first vehicle to be planned in this process is a cooperative vehicle, i.e., i=1, and its distance from the merging point is d; upon receiving the signal, it decelerates to v. C After receiving the instruction, the time t required for the transition from state A to state C is... AC The expression is:
[0129]
[0130] Among them, v AC The wave speed v represents the transition from state A to state C. A This represents the vehicle speed in state A; s A Indicates the headway of the vehicles in state A:
[0131]
[0132] Therefore, the distance x required for the cooperating vehicles to form a gap AC It can be represented as:
[0133]
[0134] When the cooperating vehicles are at a constant speed v A Driving dx AC When deceleration begins after a certain distance, the actual time it takes for the last vehicle in the main line to reach the merging point in platooning configuration. Expressed as:
[0135]
[0136] For the sake of generality, the actual time for the i-th vehicle in the mainline vehicles to reach the merging point in the platooning scenario is represented as:
[0137]
[0138] If no fleet control operation is performed, the mainline vehicles will remain in state A at a constant speed v. A While driving and maintaining a headway h A In this case, the arrival time of the i-th mainline vehicle at the merging point without platooning. Represented as:
[0139]
[0140] Delay of the jth ramp vehicle Defined as the arrival time of the j-th ramp vehicle when it arrives at the merging point in platooning conditions. Arrival time of the j-th ramp vehicle arriving at the merging point without platooning The difference is expressed as:
[0141]
[0142] For the first vehicle in the convoy, i.e., i=1, its arrival time at the merging point is the end time of the gap, which is the arrival time of the last ramp vehicle j=n at the merging point; therefore, we have:
[0143]
[0144] When there are no vehicles on the main lane, vehicles on the ramp travel at the ramp average speed v. r Maintain a constant speed and keep a headway h between the front and rear of the vehicle. r Based on existing research, the arrival time of the j-th ramp vehicle at the merging point is defined as:
[0145]
[0146] According to the motion plan, the vehicles on the ramp arrive at the merging point in state C, with a headway of h. C Therefore, the actual arrival time of vehicles on the ramp at the merging point is:
[0147]
[0148] Based on the above derivation, the objective function of the optimization model can be written as:
[0149]
[0150] In addition, the optimized model must also meet the following constraints: only connected autonomous vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point must be greater than or equal to the distance x required for the convoy to form gaps. AC That is, 0 <x AC ≤d.
[0151] Step S23: Based on the mathematical expression of the vehicle type and vehicle position distribution obtained in step S21, use a genetic algorithm to determine the optimal formation scheme.
[0152] In one specific implementation of this embodiment, step S23, which uses a genetic algorithm to determine the optimal formation scheme, includes the following process:
[0153] Step S231: Based on the mathematical expressions of the vehicle types and vehicle position distribution obtained in Step S21, generate possible formation schemes using the finite exhaustive search method to form the initial population P = {P1, P2, ..., P...} of the genetic algorithm. z}, each individual P z This represents a specific formation scheme.
[0154] Step S232: For each individual P z Substituting the values into the optimization model established in step S22, the fitness value f(P) is obtained. z )
[0155]
[0156] Step S233: From the initial population P = {P1, P2, ..., P...} z Select b individuals with high fitness as parents. Randomly pair two parent individuals with a crossover probability p. cross Determine whether to perform a crossover operation. If so, use a multi-point crossover method, randomly select multiple crossover points, and alternately exchange gene segments between these points.
[0157] Step S234: For each offspring individual, with mutation probability p mutation Determine whether to perform the mutation operation. If so, use a random mutation method to randomly select a gene location and change its value to a randomly generated valid value.
[0158] Step S235: Recalculate the fitness of the new population generated by the crossover and mutation operations. Determine if the maximum number of iterations has been reached or if the change in population fitness is less than a set threshold. If so, terminate the iteration and select the individual P with the highest fitness. fmax If the optimal formation scheme is not found, return to step S232 and restart the iterative population update.
[0159] In one specific implementation of this embodiment, in step S3, the road test processor is a key component of the present invention. It receives vehicle dynamic data and traffic environment information in real time, runs the control method and system proposed in the present invention, and issues control commands to the networked autonomous driving vehicle.
[0160] In one specific implementation of this embodiment, step S3 involves generating the control time and control commands for the cooperative vehicles in real time, specifically including the following process:
[0161] Step S31: Using the basic traffic flow relationship expression described in step S13 and the target determined in step S14, control the speed v of traffic flow state C. C speed v C This is a deceleration control command. Based on the principle of traffic wave propagation, when the cooperating vehicle receives the command, it decelerates to v. C The time t required for the transition from state A to state C after receiving the instruction. AC .
[0162]
[0163] Among them, v ACThis represents the wave speed at which the state transitions from state A to state C.
[0164]
[0165] The distance x required for the cooperative vehicles to form a gap AC .
[0166]
[0167] The distance d between the cooperating vehicle and the merging point is detected by the detector on the main line side. The control time t for sending the deceleration control command to the cooperating vehicle is then determined. control As shown in the following formula.
[0168]
[0169] Step S32: When the control time t control Upon arrival, a deceleration control command is sent to the cooperating vehicles to ensure that the gap formed meets the merging needs of the ramp vehicles, while not disrupting the overall stability of the mainline traffic flow.
[0170] In one specific implementation of this embodiment, step S4, the method for planning the vehicle speed on the ramp, specifically includes the following steps:
[0171] Step S41: Based on the relative position and speed difference between the ramp vehicle and the vehicle in front, as well as the gap information on the main line, the optimal speed of the ramp vehicle is determined using a pre-built multi-agent reinforcement learning model, thus obtaining the speed command.
[0172] Step S42: The speed command is issued to the intelligent connected vehicles on the ramp using the road test processor described in step S3, and the ramp vehicles travel at the optimal speed.
[0173] It should be noted that the multi-agent reinforcement learning model in this invention is an existing multi-agent reinforcement learning model. This invention defines the input, output and reward functions based on the existing multi-agent reinforcement learning model. The specific process of defining the input, output and reward functions is shown in the specific steps of step S41 below.
[0174] like Figure 4 As shown, step S41 specifically includes the following steps:
[0175] Step S411: Define the simulation environment for the traffic merging scenario and initialize the intelligent agent. For the intelligent connected vehicle e on the ramp, utilize the communication function of the intelligent connected vehicle to obtain its state information.
[0176]
[0177] Where, xe v represents the distance from the intelligent connected vehicle e to the merging point. e x represents the speed of the intelligent connected vehicle e. lea d er v represents the distance from the vehicle preceding the connected vehicle e to the merging point. lea d er G represents the speed of the vehicle in front. info This indicates the start and end times of the gap. The first four quantities can be acquired through sensors configured in intelligent connected vehicles. info This information can be obtained through interaction between intelligent connected vehicles and the road test processor described in step S3.
[0178] Define the action space of the intelligent connected vehicle e:
[0179]
[0180] Among them, a e This represents the acceleration of vehicle e. This represents the action space of the intelligent connected vehicle e.
[0181] Define the reward function for the intelligent connected vehicle e:
[0182] r e =w comfort r comfort +w gap r gap +w safety r safety +w effiency r effiency
[0183] Among them, w comfort ,w gap ,w safety ,w effiency These correspond to the comfort assessment value r. comfort Gap assessment value r gap Safety assessment value r safety Efficiency evaluation value r effiency The weight, r e This represents the reward value. The five evaluation values are calculated using the following method:
[0184] (1) Comfort assessment value r comfort Defined as the negative absolute value of the agility of intelligent connected vehicles, the calculation formula is:
[0185]
[0186] Where, Δa e Δt represents the acceleration difference between two adjacent moments of the intelligent connected vehicle e, and Δt represents the time interval.
[0187] (2) If a connected vehicle merges into the gap, then r gap =10, otherwise r gap =0.
[0188] (3) If the intelligent connected vehicle e collides, then r safety =-10, otherwise r safety =0.
[0189] (4) Efficiency evaluation value r effiency Defined as the negative absolute difference between the vehicle speed and the desired speed, the calculation formula is:
[0190] r efficiency =-|v e -v desired |
[0191] Among them, v e For the speed of the intelligent connected vehicle e, v desired The expected speed of vehicles on the ramp.
[0192] In step S412, at each time step, the agent obtains the current state information from the environment. And select the action space based on the current state information. Action Space This includes adjusting the vehicle speed.
[0193] Step S413, the agent executes the selected action space. When the environment changes, the feedback reward value r e .
[0194] In step S414, the agent updates its strategy based on the rewards obtained and the new state, in order to obtain higher cumulative rewards in the future.
[0195] Step S415: Determine whether the new state update strategy has converged or achieved the expected performance. If yes, end the training; otherwise, return to step S412.
[0196] One specific implementation method in this embodiment is as follows: Figure 2 As shown, this invention employs several detectors placed on the roadside to detect traffic flow data on the mainline and ramps. Specifically, ramp detector A1 is used to monitor traffic flow on the entrance ramps, and two detectors, detector B1 and detector B2, are installed on the mainline roadside. Detector B1 is used to measure the mainline traffic flow and record the corresponding macroscopic state B, while detector B2 is used to identify the type and location of vehicles to be planned.
[0197] One specific implementation method in this embodiment is as follows: Figure 3As shown, step S3, based on the target traffic flow state and optimal platooning scheme, further clarifies the control time of cooperating vehicles and the start and end times of the gaps they form, and sends a deceleration command to the cooperating vehicles when the control time arrives. The traffic flow on the main line forms a segmented platoon and gaps available for use by vehicles on the ramp. When a cooperating vehicle decelerates, a platoon is formed within the gap-forming area on the main line. This process is iteratively applied to control the deceleration of other cooperating vehicles, cyclically generating gaps for use by vehicles on the entrance ramp. Therefore, as shown in Figure (3a), the spatiotemporal information on the main line is divided into two categories: occupied spatiotemporal areas (i.e., areas occupied by platoons) and unoccupied spatiotemporal areas (i.e., areas occupied by gaps). This information can be analogized to the phases of traffic lights, with occupied spatiotemporal areas representing the red phase and unoccupied spatiotemporal areas representing the green phase. Next, the control center captures these two types of information and encodes them into a spatiotemporal sequence, and then transmits it to all vehicles on the entrance ramp, as shown in Figure (3b).
[0198] Example 2
[0199] Based on the same inventive concept as Embodiment 1, this embodiment introduces a traffic merging control system, including:
[0200] The target control traffic flow state determination module is configured to collect traffic flow data of the mainline and ramps through roadside detectors and on-board terminals; and determine the target control traffic flow state based on the traffic flow data.
[0201] The optimal formation scheme module is configured to determine the intelligent connected vehicles that need to participate in the collaboration based on the number, type, and location distribution of vehicles in the main line to be planned, combined with the traffic flow parameters of the target control traffic flow state, through a pre-built optimization model. The intelligent connected vehicles that need to participate in the collaboration are called cooperative vehicles, thereby obtaining the optimal formation scheme.
[0202] The instruction generation module is configured to generate control time and control instructions for cooperating vehicles in real time based on the target traffic flow status and optimal platooning scheme. The control instructions are issued to the cooperating vehicles through the communication system of the road test processor, guiding the cooperating vehicles to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line.
[0203] The reinforcement learning module is configured to plan the speed of vehicles on the ramp through a multi-agent reinforcement learning model and generate speed commands; after receiving the speed commands, the vehicles on the ramp adjust their speed to merge into the main line.
[0204] Example 3
[0205] Based on the same inventive concept as other embodiments, this embodiment describes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the traffic merging control method described above.
[0206] Example 4
[0207] Based on the same inventive concept as other embodiments, this embodiment introduces a computer device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the traffic merging control method described above.
[0208] Example 5
[0209] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including a computer program that, when executed by a processor, implements the steps of the traffic merging control method described above.
[0210] This invention utilizes target-based traffic flow control to guide cooperating vehicles to proactively reserve appropriate merging gaps. Through a multi-agent reinforcement learning strategy, it dynamically adjusts the driving behavior of ramp vehicles in real time, ensuring a smooth and safe merging of ramp vehicles into the main line. This method effectively reduces delays and energy consumption during ramp merging, avoids frequent braking and acceleration, and improves the efficiency, safety, and environmental friendliness of the traffic system.
[0211] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0212] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0213] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0214] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0215] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of the present invention, and these modifications are all within the protection scope of the present invention.
Claims
1. A traffic merging control method, characterized in that, Includes the following steps: Traffic flow data for the mainline and ramps is collected using roadside detectors and on-board terminals; the target traffic flow state is determined based on the traffic flow data. Based on the number of vehicles, vehicle types, and vehicle location distribution in the main line to be planned, and combined with the traffic flow parameters of the target controlled traffic flow state, the intelligent connected vehicles that need to participate in the coordination are determined through a pre-built optimization model. The determined intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, thereby obtaining the optimal formation scheme. Based on the target traffic flow status and optimal platooning scheme, control time and control instructions for cooperating vehicles are generated in real time. The control instructions are issued to the cooperating vehicles through the communication system of the road test processor, guiding the cooperating vehicles to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line. The speed of vehicles on the ramp is planned by a multi-agent reinforcement learning model, and a speed command is obtained. After receiving the speed command, vehicles on the ramp adjust their speed to merge into the main line.
2. The traffic merging control method according to claim 1, characterized in that: The specific steps for determining the target control traffic flow state based on the traffic flow data are as follows: Using traffic flow data combined with basic traffic flow graph theory, the target control traffic flow state is determined. The traffic flow data includes vehicle type, location, speed, headway, and spacing. The traffic flow parameters for the current mainline traffic flow state include flow rate, density, speed, headway, and spacing, each represented by q. A k A v A h A and s A The traffic flow parameters for controlling traffic flow state include flow rate, density, speed, and headway, which are respectively represented as q. c k c v C and h C The current traffic flow state of the main line is referred to as state A, and the target controlled traffic flow state is referred to as state C. Based on the number of vehicles n in the mainline convoy p The headway h of vehicle in state A A Obtain the period length C c As shown in the following formula: C c =h A ·n p ; Among them, the number of vehicles in the main line convoy is n p State A: Headway h A The data is obtained directly from the roadside detector and is a known quantity. The duration G of the gap is: G=n p ·(h A -h C ); Among them, h C G represents the headway in state C, and G represents the duration of the gap. The number of ramp vehicles that can merge into within the gap, n G for: Where, n C The number of vehicles that can merge into the ramp within the gap, g m Let C be the minimum headway required for a vehicle to merge onto the main line from a ramp; then the cycle length C of a vehicle merging onto a ramp is... c Traffic flow q of the ramp passing through ramp As shown in the following formula: Therefore, the headway h of the ideal target traffic control state is obtained. ideal As shown below: h ideal =h A -q ramp ·h A ·g m ; Assuming the total number of vehicles on the planned main line is m, in order to create effective gaps between vehicles on the planned main line, the following formula should be satisfied: Where [·] represents rounding up, np min This represents the minimum convoy size required to form sufficient gaps; np can be obtained from the above formula. min The range of values for h is then determined by the following formula for the headway h in state C. C The minimum value limit; Among them, h Cmin The headway h represents the distance between the front and rear of the vehicle in state C. C The clearance between the front and rear of the vehicle; Headway h based on the ideal target for controlling traffic conditions ideal Distance h between the front and the clearance vehicle Cmin The headway h of the vehicle in state C is obtained. C As shown in the following formula: h C =max(h ideal ,h Cmin ); Based on the headway h obtained from state C C Based on the basic traffic flow relationship expression, the velocity v in state C is calculated. C Density k C and traffic q C The basic relationship expression of traffic flow is as follows: q C =k C ·v C ; Where T represents the safe headway, l is the vehicle length, v0 is the free flow velocity, and s0 is the minimum safe stopping distance, all of which are known quantities.
3. The traffic merging control method according to claim 1, characterized in that: The specific method for constructing the optimization model is as follows: The optimization model is established with the objective of minimizing the delay D of all vehicles passing through the merging point as the optimization objective: Where, ω main and ω ramp These represent the weights of vehicle delays on the main line and on the ramps, respectively. 'm' refers to the number of vehicles on the main line to be planned, and 'n' refers to the number of vehicles merging onto the ramps that need to be coordinated. This refers to the delay of the i-th mainline vehicle, where i = 1, 2, 3, ..., m. The delay refers to the delay of the j-th vehicle on the ramp, where j = 1, 2, 3, ..., n; The objective function for optimizing the model is: Where, x AC t represents the distance required for the convoy to form gaps; AC This indicates the time required for the current traffic flow state on the main line to transition to the target controlled traffic state; h r This indicates that when there are no vehicles in the main lane, vehicles on the ramp travel at the average ramp speed v. r Maintaining a constant speed and keeping a constant headway; h C Headway, representing the time difference between vehicles in the target traffic control state; h A Headway indicates the current traffic flow status on the main line; The optimization model must satisfy the following constraints: only connected autonomous vehicles are selected as cooperative vehicles; the distance d between the cooperative vehicle and the merging point is greater than or equal to the distance x required for the convoy to form gaps. AC That is, 0 <x AC ≤d.
4. The traffic merging control method according to claim 3, characterized in that: The specific method for obtaining the optimal formation scheme is as follows: Determining the optimal formation scheme using a genetic algorithm includes the following steps: Based on the vehicle types and their location distribution, possible formation schemes are generated using a finite exhaustive search method, forming the initial population P = {P1, P2, ..., P} for the genetic algorithm. z }, each individual P z This represents a specific formation scheme; For each individual P z Substituting the values into the optimization model, we obtain the fitness value f(P). z ); From the initial population P = {P1, P2, ..., P...} z Select b individuals with high fitness as parents; randomly pair two parent individuals with crossover probability p. cross Determine whether to perform a crossover operation; if so, use a multi-point crossover method, randomly select multiple crossover points, and alternately exchange gene segments between these points. For each offspring individual, with mutation probability p mutation Determine whether to perform a mutation operation; if so, use a random mutation method to randomly select a gene location and change its value to a randomly generated valid value. Based on the new population generated by crossover and mutation operations, recalculate the fitness; determine whether the maximum number of iterations has been reached or the change in population fitness is less than a set threshold. If so, terminate the iteration and select the individual P with the highest fitness. fmax If the optimal formation scheme is not found, then return to the optimal formation scheme for each individual P. z Substituting the values into the optimization model, we obtain the fitness value f(P). z The population is then iteratively updated again using the following steps.
5. The traffic merging control method according to claim 2, characterized in that: The specific steps for generating real-time control time and control commands for cooperative vehicles are as follows: The speed v of traffic flow is controlled using the basic traffic flow relationship expression and a defined objective. C speed v C This is a deceleration control command. Based on the principle of traffic wave propagation, when the cooperating vehicle receives the command, it decelerates to v. C The time t required for the transition from state A to state C after receiving the instruction. AC ; Among them, v AC This represents the wave speed during the transition from state A to state C. The distance x required for the cooperative vehicles to form a gap AC ; The distance d between the cooperating vehicle and the merging point is detected by the detector on the main line side. The control time t for sending the deceleration control command to the cooperating vehicle is then determined. control As shown in the following formula; When control time t control Upon arrival, a deceleration control command is sent to the cooperating vehicles.
6. The traffic merging control method according to claim 2, characterized in that: The specific method for planning vehicle speeds on ramps is as follows: Based on the relative position and speed difference between the ramp vehicle and the vehicle in front, as well as the gap information on the main line, the optimal speed of the ramp vehicle is determined by a pre-built multi-agent reinforcement learning model, thus obtaining the speed command. The speed command is issued to the intelligent connected vehicles on the ramp using the road test processor, and the ramp vehicles travel at the optimal speed.
7. A traffic merging control system, characterized in that, include: The target control traffic flow state determination module is configured to collect traffic flow data of the mainline and ramps via roadside detectors and on-board terminals; The target control traffic flow state is determined based on the traffic flow data; The optimal formation scheme module is configured to determine the intelligent connected vehicles that need to participate in the coordination based on the number of vehicles, vehicle types, and vehicle location distribution in the main line to be planned, combined with the traffic flow parameters of the target controlled traffic flow state, through a pre-built optimization model. The determined intelligent connected vehicles that need to participate in the coordination are called cooperative vehicles, thereby obtaining the optimal formation scheme. The instruction generation module is configured to generate control time and control instructions for cooperating vehicles in real time based on the target control traffic flow state and optimal formation scheme; the control instructions are issued to the cooperating vehicles through the communication system of the road test processor, instructing the cooperating vehicles to decelerate within the control time, thereby creating sufficient gaps for ramp vehicles to merge into the main line; The reinforcement learning module is configured to plan the speed of vehicles on the ramp through a multi-agent reinforcement learning model and generate speed commands; after receiving the speed commands, the vehicles on the ramp adjust their speed to merge into the main line.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the traffic merging control method according to any one of claims 1 to 6.
9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the traffic merging control method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of the traffic merging control method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Vehicle variable speed limit control optimization method, system, medium and equipment
CN115862322A
Dynamic cooperative confluence control method and system suitable for mixed traffic flow
CN117409564A