Methods, devices, equipment, and storage media for multi-vehicle cooperative control in ramp merging zones

By dividing the ramp merging zone into a preparation zone and a merging zone, and using the MADDPG algorithm to train a multi-agent distributed control model, the problem of high computational complexity in multi-vehicle cooperative control in ramp merging zones is solved, and efficient and safe multi-vehicle cooperative control is achieved.

CN117116040BActive Publication Date: 2026-07-31CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2023-08-18
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies have high computational complexity and poor real-time performance in multi-vehicle cooperative control in ramp merging areas, making it difficult to guarantee traffic efficiency and multi-vehicle cooperative performance.

Method used

By dividing the ramp merging area into a preparation area and a merging area, a multi-agent distributed control model (MADDPG) is used for offline training to generate multi-vehicle cooperative control commands. This includes training the multi-agent distributed control models in the preparation area and the merging area, and optimizing the MADDPG algorithm to improve real-time performance and cooperative performance.

Benefits of technology

It effectively improves the traffic efficiency of the ramp merging area, reduces the computational complexity, and enables vehicles to merge into the main ramp efficiently and safely.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117116040B_ABST
    Figure CN117116040B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, device, and storage medium for multi-vehicle cooperative control in a ramp merging area. The method includes: during offline training, dividing the ramp merging area into a preparation zone and a merging zone, setting different control objectives for each zone, optimizing the MADDPG algorithm, and sequentially training multi-agent distributed control models for the preparation and merging zones using the MADDPG algorithm; then, in the online decision-making phase, directly utilizing the trained multi-agent distributed control models for the preparation and merging zones to achieve efficient and safe merging of the target vehicle cluster on the main ramp. This solves the problems of existing technologies in ramp merging scenarios, such as the increasing multi-vehicle cooperative control time with the increase in the number of controlled vehicles, high computational complexity, poor real-time performance, and difficulty in guaranteeing traffic efficiency and multi-vehicle cooperative performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent connected vehicle technology, and in particular to a method, apparatus, device and storage medium for multi-vehicle cooperative control in a ramp merging zone. Background Technology

[0002] With rapid socio-economic development and a continuously increasing number of vehicles, intercity transportation pressure has surged, posing significant challenges to the efficient and safe operation of highways. Among the various scenarios on highways, ramp merging areas are high-risk zones for congestion and accidents. Benefiting from the development of intelligent connected vehicle technology in recent years, the real-time information transmission and precise vehicle control of intelligent connected vehicles have brought new directions to solving the ramp merging control problem.

[0003] In a connected environment, vehicles can directly receive real-time information commands from the control center to accurately adjust their operating status and avoid conflicts between mainline vehicles and ramp vehicles when they merge. However, the commonly used multi-vehicle cooperative control methods for ramp merging areas are mainly based on optimization methods such as dynamic programming and mixed integer nonlinear programming. As the number of vehicles increases in the ramp merging scenario, the computational complexity increases exponentially, making it difficult to meet the real-time requirements of the cloud control system.

[0004] However, in existing technologies, the duration of multi-vehicle cooperative control in the merging zone of ramps tends to increase with the number of vehicles controlled, resulting in high computational complexity, poor real-time performance, and difficulty in ensuring traffic efficiency and multi-vehicle cooperative performance, which urgently needs to be addressed. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for multi-vehicle cooperative control in ramp merging areas to solve the problems in the prior art where the multi-vehicle cooperative control time tends to increase with the number of vehicles controlled, the computational complexity is high, the real-time performance is poor, and it is difficult to guarantee traffic efficiency and multi-vehicle cooperative performance in ramp merging scenarios.

[0006] The first aspect of this application provides a method for multi-vehicle cooperative control in a ramp merging area, applied in the offline training phase. The method includes the following steps: dividing the target ramp merging area; determining the preparation area and merging area of ​​the target ramp merging area; determining the training vehicle clusters in the preparation area and the merging area; and formulating a multi-vehicle cooperative control target for the training vehicle clusters; optimizing a preset MADDPG algorithm using the multi-vehicle cooperative control target to generate a MADDPG control model; training a multi-agent distributed control model in the preparation area based on the MADDPG control model to obtain a multi-agent distributed control model for the preparation area under continuous traffic flow; and training the multi-agent distributed control model for the merging area based on the MADDPG control model to obtain a multi-agent distributed control model for the merging area under a fixed number of vehicles, so as to output multi-vehicle cooperative control commands for the target ramp merging area using the multi-agent distributed control model for the preparation area under continuous traffic flow and the multi-agent distributed control model for the merging area under a fixed number of vehicles.

[0007] Optionally, in one embodiment of this application, the step of setting a multi-vehicle collaborative control target for the training vehicle cluster includes: setting a first control target for each vehicle in the training vehicle cluster in the preparation area to control the time interval between each vehicle arriving at the end of the preparation area to meet a preset interval requirement; and setting a second control target for each vehicle in the merging area to control the training vehicle cluster to meet a preset safe merging requirement according to the second control target.

[0008] Optionally, in one embodiment of this application, training the multi-agent distributed control model of the preparation area includes: determining lane information, a target vehicle, and multiple reference vehicles in the preparation area based on the MADDPG control model; calculating the distance and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area; obtaining the preparation area state space of the target vehicle based on the distance, the first relative position, and the second relative position; determining the preparation area action space of the target vehicle according to the multi-vehicle cooperative control objective of the preparation area, and setting the preparation area reward function of the multi-agent distributed control model; and training the multi-agent distributed control model of the preparation area using the preparation area state space, the preparation area action space, and the preparation area reward function.

[0009] Optionally, in one embodiment of this application, training the multi-agent distributed control model of the merging zone includes: obtaining the state variables of the target vehicle based on the MADDPG control model, and determining the merging zone state space of the target vehicle based on the state variables; determining the action space of the target vehicle in the merging zone based on the multi-vehicle cooperative control objective of the merging zone, and setting the merging zone reward function of the multi-agent distributed control model; and training the multi-agent distributed control model of the merging zone using the merging zone state space, the merging zone action space, and the merging zone reward function.

[0010] A second aspect of this application provides a method for multi-vehicle cooperative control in a ramp merging zone, applied in the online decision-making stage. The method includes the following steps: dividing the target ramp merging zone into a preparation area, a merging area, and a target vehicle cluster; performing merging planning for each vehicle in the target vehicle cluster in the preparation area based on a pre-trained multi-agent distributed control model of the preparation area, obtaining a first decision result for each vehicle; performing merging planning for each vehicle in the target vehicle cluster in the merging area based on a pre-trained multi-agent distributed control model of the merging area, obtaining a second decision result for each vehicle, wherein both the merging area multi-agent distributed control model and the preparation area multi-agent distributed control model are trained using a MADDPG control model; and generating multi-vehicle cooperative control instructions for the continuous traffic flow in the target ramp merging zone based on the first and second decision results.

[0011] A third aspect of this application provides an apparatus for multi-vehicle cooperative control in a ramp merging area, applied in an offline training phase, comprising: a partitioning module for partitioning a target ramp merging area, determining a preparation area and a merging area of ​​the target ramp merging area, determining a training vehicle cluster in the preparation area and the merging area, and formulating a multi-vehicle cooperative control target for the training vehicle cluster; an optimization module for optimizing a preset MADDPG algorithm using the multi-vehicle cooperative control target to generate a MADDPG control model; a second training module for training a multi-agent distributed control model of the preparation area based on the MADDPG control model to obtain a multi-agent distributed control model of the preparation area under continuous traffic flow; and a third training module for training the multi-agent distributed control model of the merging area based on the MADDPG control model to obtain a multi-agent distributed control model of the merging area under a fixed number of vehicles, so as to output multi-vehicle cooperative control commands for the target ramp merging area using the multi-agent distributed control model of the preparation area under continuous traffic flow and the multi-agent distributed control model of the merging area under a fixed number of vehicles.

[0012] Optionally, in one embodiment of this application, the partitioning module includes: a first control unit, configured to set a first control target for each vehicle in the training vehicle cluster in the preparation area, so as to control the time interval between each vehicle arriving at the end of the preparation area to meet a preset interval requirement; and a second control unit, configured to set a second control target for each vehicle in the merging area, so as to control the training vehicle cluster to reach a preset safe merging requirement according to the second control target.

[0013] Optionally, in one embodiment of this application, the second training module includes: a determination unit, configured to determine lane information, a target vehicle, and multiple reference vehicles in the preparation area based on the MADDPG control model; a calculation unit, configured to calculate the spacing and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area; a first acquisition unit, configured to obtain the preparation area state space of the target vehicle based on the spacing, the first relative position, and the second relative position; a first setting unit, configured to determine the preparation area action space of the target vehicle according to the multi-vehicle cooperative control objective of the preparation area, and set the preparation area reward function of the multi-agent distributed control model; and a first execution unit, configured to train the multi-agent distributed control model of the preparation area using the preparation area state space, the preparation area action space, and the preparation area reward function.

[0014] Optionally, in one embodiment of this application, the third training module includes: a second acquisition unit, configured to acquire the state variables of the target vehicle based on the MADDPG control model, and determine the merging zone state space of the target vehicle according to the state variables; a second setting unit, configured to determine the action space of the target vehicle in the merging zone according to the multi-vehicle cooperative control objective of the merging zone, and set the merging zone reward function of the multi-agent distributed control model; and a second execution unit, configured to train the multi-agent distributed control model of the merging zone using the merging zone state space, the merging zone action space, and the merging zone reward function.

[0015] A fourth aspect of this application provides an apparatus for multi-vehicle cooperative control in a ramp merging area, applied in the online decision-making stage, comprising: a partitioning module for dividing a target ramp merging area and determining a preparation area, a merging area, and a target vehicle cluster within the target ramp merging area; a first planning module for performing merging planning for each vehicle in the target vehicle cluster in the preparation area based on a pre-trained multi-agent distributed control model of the preparation area, obtaining a first decision result for each vehicle; a second planning module for performing merging planning for each vehicle in the target vehicle cluster in the merging area based on a pre-trained multi-agent distributed control model of the merging area, obtaining a second decision result for each vehicle, wherein the multi-agent distributed control model of the merging area and the multi-agent distributed control model of the preparation area are both trained using a MADDPG control model; and a control module for generating multi-vehicle cooperative control commands for the continuous traffic flow in the target ramp merging area based on the first and second decision results.

[0016] A fifth aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for multi-vehicle cooperative control in a ramp merging zone as described in the above embodiments.

[0017] A sixth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for multi-vehicle cooperative control in a ramp merging zone.

[0018] Therefore, the embodiments of this application have the following beneficial effects: The embodiments of this application can divide the merging area of ​​the target ramp into a preparation area, a merging area, and a target vehicle cluster. Based on a pre-trained multi-agent distributed control model in the preparation area, merging planning is performed for each vehicle in the target vehicle cluster in the preparation area, resulting in a first decision result for each vehicle. Based on a pre-trained multi-agent distributed control model in the merging area, merging planning is performed for each vehicle in the target vehicle cluster in the merging area, resulting in a second decision result for each vehicle. Both the merging area multi-agent distributed control model and the preparation area multi-agent distributed control model are trained using the MADDPG control model. Based on the first and second decision results, multi-vehicle cooperative control instructions for the continuous traffic flow in the merging area of ​​the target ramp are generated. Thus, this application can effectively improve traffic efficiency while reducing the computational complexity of multi-vehicle cooperative control, improving multi-vehicle cooperative performance, and achieving efficient and safe merging of vehicles from the main ramp. This solves the problems of existing technologies in ramp merging scenarios, such as the multi-vehicle collaborative control time easily increasing with the number of vehicles controlled, high computational complexity, poor real-time performance, and difficulty in ensuring traffic efficiency and multi-vehicle collaborative performance.

[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a method for multi-vehicle cooperative control in a ramp merging zone during the offline training phase, according to an embodiment of this application. Figure 2 A schematic diagram of a ramp merging scenario provided for one embodiment of this application; Figure 3 A schematic diagram of a distributed multi-agent reinforcement learning architecture is provided for one embodiment of this application; Figure 4 A schematic diagram of a reference vehicle provided for one embodiment of this application; Figure 5 A schematic diagram of a vehicle training scenario in a merging area is provided as an embodiment of this application; Figure 6 This is a flowchart illustrating a method for multi-vehicle cooperative control in the merging zone of a ramp, applied during the online decision-making stage, according to an embodiment of this application. Figure 7 This is an example diagram of a device for multi-vehicle cooperative control in a ramp merging zone during the offline training phase, according to an embodiment of this application. Figure 8This is an example diagram of a device for multi-vehicle cooperative control in a ramp merging zone during the online decision-making stage, according to an embodiment of this application. Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0021] Among them, 10-device for multi-vehicle cooperative control of ramp merging area applied in offline training stage, 20-device for multi-vehicle cooperative control of ramp merging area applied in online decision-making stage; 100-partitioning module, 200-optimization module, 300-second training module, 400-third training module; 500-partitioning module, 600-first planning module, 700-second planning module, 800-control module; 901-memory, 902-processor, 903-communication interface. Detailed Implementation

[0022] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0023] The following description, with reference to the accompanying drawings, describes a method, apparatus, device, and storage medium for multi-vehicle cooperative control in the merging zone of a ramp, according to embodiments of this application. To address the problems mentioned in the background art, this application provides a method for multi-vehicle cooperative control in a ramp merging area. In this method, the target ramp merging area is divided into a preparation area, a merging area, and a target vehicle cluster. Based on a pre-trained multi-agent distributed control model in the preparation area, merging planning is performed for each vehicle in the target vehicle cluster in the preparation area, yielding a first decision result for each vehicle. Based on a pre-trained multi-agent distributed control model in the merging area, merging planning is performed for each vehicle in the target vehicle cluster in the merging area, yielding a second decision result for each vehicle. Both the merging area multi-agent distributed control model and the preparation area multi-agent distributed control model are trained using a MADDPG control model. Based on the first and second decision results, multi-vehicle cooperative control instructions for the continuous traffic flow in the target ramp merging area are generated. Therefore, this application can effectively improve traffic efficiency while reducing the computational complexity of multi-vehicle cooperative control, improving multi-vehicle cooperative performance, and achieving efficient and safe merging of vehicles from the main ramp. This solves the problems of existing technologies in ramp merging scenarios, such as the multi-vehicle collaborative control time easily increasing with the number of vehicles controlled, high computational complexity, poor real-time performance, and difficulty in ensuring traffic efficiency and multi-vehicle collaborative performance.

[0024] Specifically, Figure 1This is a flowchart illustrating a method for multi-vehicle cooperative control in the merging zone of a ramp, provided as an embodiment of this application, and applied during the offline training phase.

[0025] like Figure 1 As shown, the method for multi-vehicle cooperative control in the merging zone of this ramp includes the following steps: In step S101, the target ramp merging area is divided, the preparation area and merging area of ​​the target ramp merging area are determined, the training vehicle clusters in the preparation area and merging area are determined, and multi-vehicle cooperative control targets are formulated for the training vehicle clusters.

[0026] The embodiments of this application take the scenario of highway entrance ramp merging as an example to illustrate and introduce the method of multi-vehicle cooperative control in the ramp merging area of ​​this application.

[0027] First, the embodiments of this application can define the interaction environment of reinforcement learning agents; specifically, in the embodiments of this application, the ramp merging area scenario is a single-lane merging, and the merging area includes an acceleration lane, where ramp vehicles can change lanes to merge into the main road traffic flow, such as... Figure 2 As shown, all vehicles are intelligent connected vehicles that exchange information with the roadside unit of the cloud control system through V2I devices.

[0028] Furthermore, embodiments of this application can divide the merging ramp scene area into a preparation area before the acceleration lane and a merging area containing the acceleration lane area, and formulate different control objectives and training schemes according to different areas of the merging ramp area; wherein, the preparation area is the area 800m before the acceleration lane, and the merging area is the acceleration lane area with a length of 200m.

[0029] Therefore, the embodiments of this application define a reinforcement learning agent interaction environment by exchanging information flow in a preset ramp merging scenario, dividing the preparation area and the merging area, and formulating corresponding control objectives and training schemes for different areas, thereby providing data basis for the training of subsequent control models.

[0030] Optionally, in one embodiment of this application, a multi-vehicle cooperative control objective is formulated for the training vehicle cluster, including: formulating a first control objective for each vehicle in the training vehicle cluster in the preparation area to control the time interval between each vehicle arriving at the end of the preparation area to meet a preset interval requirement; and formulating a second control objective for each vehicle in the merging area to control the training vehicle cluster to meet a preset safe merging requirement according to the second control objective.

[0031] It should be noted that, after dividing the merging ramp scenario area into a preparation area before the acceleration lane and a merging area containing the acceleration lane, the embodiments of this application also need to formulate control objectives for the multi-vehicle coordination method in the preparation area and the merging area respectively: In the preparation area, vehicles adjust their speed to meet the merging conditions and satisfy the requirements in the merging area. This ensures that the time interval between vehicles on the main ramp arriving at the end of the preparation area is appropriate, and that vehicles reach their maximum speed. The specific requirements for the time interval are as follows:

[0032] in, For an acceptable and safe merging time interval, 2.5 seconds is taken. For a safe following distance, a value of 1.6 seconds is used. , For vehicles With the vehicle in front +1 arrival time at the preparation area finish line.

[0033] In the merging zone, the control objective of this application embodiment is to enable vehicles to merge safely as quickly as possible while ensuring safety and avoiding collisions.

[0034] During the training process, the second-order kinematic vehicle model used in the embodiments of this application is as follows:

[0035]

[0036] in, and Representative vehicle Position and velocity, The vehicle's acceleration is the input control quantity. For time.

[0037] For each vehicle in the cooperative control area, its control input variables and speed should meet the following conditions:

[0038]

[0039] in, and These represent the minimum and maximum speeds allowed for vehicles within the controlled area, respectively. and These are the minimum acceleration and the maximum acceleration, respectively.

[0040] In addition, to avoid the phenomenon of vehicles stopping and starting and to enable the intelligent agent to find the correct strategy as soon as possible, the embodiments of this application also need to limit the minimum speed of the vehicle to above 0.

[0041] Therefore, the embodiments of this application provide reliable theoretical guidance and data support for the safe merging of vehicles by setting corresponding control objectives for the preparation area and the merging area, thereby enabling subsequent model training.

[0042] In step S102, the preset MADDPG algorithm is optimized using the multi-vehicle cooperative control objective to generate the MADDPG control model.

[0043] Those skilled in the art should understand that, with the development of deep learning and neural networks in recent years, reinforcement learning (RL) based methods can achieve offline training and online decision-making, effectively solving problems such as perception and decision-making of intelligent agents in complex state spaces. Moreover, the computational cost of the trained network is shorter than that of real-time planning optimization algorithms, resulting in better real-time performance.

[0044] Therefore, after dividing the target ramp merging area and formulating multi-vehicle cooperative control objectives for the training vehicle cluster, the embodiments of this application can further construct the MADDPG algorithm structure to describe the information flow between the roadside unit and the vehicle agent, based on the distributed multi-agent reinforcement learning architecture proposed for the ramp merging area scenario.

[0045] Specifically, in the embodiments of this application, the reinforcement learning method involves a training process in which the agent continuously interacts with the environment and updates its policy after receiving rewards from the environment. The agent obtains the optimal policy during the learning process to maximize the reward value. The embodiments of this application take multiple agents as the research object, training multiple agents to cooperate to maximize traffic efficiency, and using a distributed training architecture based on the MADDPG algorithm to train and plan the optimal speed trajectory.

[0046] For multi-agent research scenarios, the embodiments of this application can utilize a multi-agent Markov decision process (MAMDP), the mathematical expression of which can be represented as:

[0047] in, N The number of CAVs in the control area; For the set of vehicle state variables, This is the set of vehicle actions, where the set of vehicle state variables and the set of vehicle actions can be modified according to the training process in the preparation and merging areas; } represents the set of state transition probabilities for each vehicle. Defined as by arrive The transition probability, For the set of vehicle reward functions, k This represents the step size in discrete time.

[0048] The distributed multi-agent reinforcement learning framework of this application embodiment is as follows: Figure 3 As shown, by Figure 3It can be seen that vehicles obtain state quantities such as speed and relative position of surrounding vehicles through their own vehicle sensors and V2V communication equipment, while the roadside unit provides the state quantities of vehicles that need to cooperate in different lanes.

[0049] In the embodiments of this application, each vehicle has a complete DDPG agent network, including a target value network, an online value network, a target policy network, and an online policy network; wherein the input of the online policy network is environmental state variables provided by the roadside unit and vehicle sensors of the cloud control system, and the output is vehicle action variables, including vehicle acceleration and lane change commands; the update rule of the online policy network aims to maximize the action value function, and the update gradient of its network parameters is:

[0050] in, For vehicles exist k Time-state quantity, For vehicles exist k Momentary action For online policy networks, For policy networks Online value network of time.

[0051] The Online Value Network (OWN) outputs a state value function that evaluates the quality of the policy network based on the current state and all agent actions, with the goal of maximizing rewards. The OWN update rule aims to revise the evaluation of the policy network according to the reward function, minimizing the online action value of its own output. With target value The difference between them, therefore the online value network update loss function is:

[0052]

[0053] in, For vehicles exist k Time-state quantity, For vehicles exist k Momentary action For vehicles exist k Rewards at all times For online policy networks, For policy networks Online value network of time, This is a discount factor, which takes a value based on the number of training rounds. This is the target value of the value network.

[0054] In the embodiments of this application, the data during the training process is in { The system stores data in an experience pool, randomly samples data during training, and inputs it into the network for training. After the training process is completed, the vehicle only needs to retain the online policy network as the vehicle controller. The roadside unit and sensors provide state variables for the controller, and the controller outputs its own acceleration and lane-changing commands.

[0055] Therefore, the embodiments of this application provide theoretical support for the training of multi-agent distributed control models in the subsequent preparation and merging areas by constructing and utilizing the MADDPG control model. Furthermore, the MADDPG-based reinforcement learning control model is of great significance for alleviating traffic congestion in highway merging areas and improving driving safety.

[0056] In step S103, based on the MADDPG control model, a multi-agent distributed control model for the preparation area is trained to obtain a multi-agent distributed control model for the preparation area under continuous traffic flow.

[0057] After constructing the MADDPG algorithm structure, the embodiments of this application can further train a multi-agent distributed control model in the preparation area using the MADDPG algorithm.

[0058] Optionally, in one embodiment of this application, training the multi-agent distributed control model of the preparation area includes: determining lane information, the target vehicle, and multiple reference vehicles in the preparation area based on the MADDPG control model; calculating the distance and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area; obtaining the preparation area state space of the target vehicle based on the distance, the first relative position, and the second relative position; determining the preparation area action space of the target vehicle according to the multi-vehicle cooperative control objective of the preparation area, and setting the preparation area reward function of the multi-agent distributed control model; and training the multi-agent distributed control model of the preparation area using the preparation area state space, the preparation area action space, and the preparation area reward function.

[0059] It should be noted that after completing the above MADDPG algorithm architecture, the embodiments of this application also need to train the multi-agent environment in the preparation area scenario.

[0060] In practice, the training environment of the merging zone is limited by the number of agents that can be trained simultaneously using the MADDPG algorithm, allowing only six vehicle agents to be controlled at the same time. However, the environment of the merging zone often consists of a large number of vehicles and continuous traffic flow. For a single agent, the observations of other agents are constantly changing. Simultaneously, when the policies of other agents are updated, the environment changes, contradicting the common assumption of a static environment in reinforcement learning algorithm training. From the perspective of a static environment, it can exist globally or locally. Therefore, to construct a relatively static environment, vehicle agents can choose surrounding vehicles to build such an environment. The update gradient of the policy network is then:

[0061] in, , They represent those from vehicles. and reference vehicles The set of observed states and the set of actions, vehicles of The reference vehicle was determined based on the vehicle's initial state at the start of training.

[0062] In addition, in the training scenario, to simulate continuous traffic flow, several IDM (Integrated Driver Model) follow-up models were deployed around the six intelligent agent vehicles to supplement the agents' state information; after completing the training process, based on the expected arrival time... The vehicles are grouped in sequence, and the traffic flow is controlled in groups of 6 vehicles. The resulting policy network is then distributed to each intelligent vehicle in the preparation area.

[0063] In the embodiments of this application, the MADDPG algorithm can be used to solve the preparation area scene and its control objective. This mainly includes calculating the state space, action space, and reward function of the preparation area scene. The specific process is as follows: 1. State space: To better enable multi-vehicle collaboration, the state variables of the intelligent agent can be divided into two parts: the state of the surrounding vehicles and the state of the collaborating vehicles. The state of the surrounding vehicles is obtained by the vehicle's own sensors, which acquire the vehicle's position, speed, and the relative positions and speeds of vehicles in front and behind in the same lane.

[0064] The goal of multi-vehicle control in the preparation area is to ensure that the time interval between vehicles on the main ramp arriving at the end of the preparation area is greater than or equal to the safe merging time distance, and that the vehicles reach their maximum speed. Since the speeds on the main ramps are different, the expected arrival time at the end of the preparation area is used as the initial arrival order, and a reference agent for interaction with the vehicle is selected accordingly.

[0065] To enable the intelligent agent to achieve basic collision avoidance with vehicles in the same lane, the vehicles in front and behind in the same lane are selected as reference vehicles.

[0066] The expected time to reach the end of the preparation area is:

[0067] in, For vehicles Distance to the starting point of the acceleration lane The maximum speed is the target speed at which the vehicle reaches the starting point of the acceleration lane. For vehicles The initial velocity. This represents the vehicle. The time taken to reach the starting point of the acceleration lane with uniform acceleration.

[0068] By expected arrival time A preliminary plan is drawn up regarding the expected interactions between vehicles, such as... Figure 4 As shown, according to the vehicle At the expected arrival timeline, vehicles on the main road maintain speeds near their maximum speed, therefore... The value is often further forward than vehicles on ramps at the same longitudinal position; vehicles according to The order of values ​​is selected from the distance in the other lane. Recently, and the gap between vehicles formed by vehicles in another lane. , Its corresponding relative position , Therefore, in this embodiment of the application, the absolute position of the other lane reference vehicle selected is... Figure 4 middle item.

[0069] The relative positions are:

[0070] The relative positions of vehicles in the same lane are , for:

[0071] Therefore, vehicles The state variables are:

[0072]

[0073] in, Lane information is represented by 1 for main lanes and 0 for ramps; vehicle status variables include both the vehicle's own status variables and those of reference vehicles. State variables.

[0074] 2. Movement space action For online policy networks The output represents the vehicle intelligent agent. The longitudinal decision-making quantities are incorporated into the model training in the preparation area, and the actions... For vehicles The acceleration, and the action space A are:

[0075] 3. Reward Function It is understandable that a reward represents the immediate feedback after completing an action, and it defines the learning progress target. Based on the control target, the single-step reward function proposed in this application embodiment is:

[0076]

[0077]

[0078]

[0079]

[0080] in, Indicates speed loss; Represents the weighting coefficient, for For example, the coefficient for vehicles on the main road will be set higher to avoid excessive speed adjustments on the main road, which would affect the traffic efficiency of the main road and thus reflect the priority of vehicles on the main road over vehicles on the ramps. This indicates a penalty for collisions in the same lane. If a collision occurs in the same lane, a penalty will be imposed, and the following vehicle will ignore the agent's actions and be forced to slow down. This is a comfort index, representing the goal of minimizing changes in vehicle acceleration.

[0081] In addition, the final state reward at the end of the round is:

[0082] in, Indicates the weighting coefficient. Indicates the vehicle delay time. Indicates the time when the vehicle arrives at the end of the preparation area. Indicates vehicle The time when the previous vehicle arrives at the finish line of the preparation area. This indicates the safe merging time of the vehicle.

[0083] Therefore, the embodiments of this application, through the constructed MADDPG algorithm structure, train the multi-agent distributed control model of the preparation area, thereby realizing the simultaneous training of multiple agents, enhancing the collaborative performance of multiple agents, and effectively improving the passage efficiency.

[0084] In step S104, based on the MADDPG control model, a multi-agent distributed control model for the merging area is trained to obtain a multi-agent distributed control model for the merging area with a fixed number of vehicles. This model is then used to output multi-vehicle cooperative control commands for the target ramp merging area using the multi-agent distributed control model for the preparation area under continuous traffic flow and the multi-agent distributed control model for the merging area with a fixed number of vehicles.

[0085] After training the multi-agent distributed control model in the preparation area using the MADDPG algorithm structure, embodiments of this application can further train the multi-agent distributed control model in the merging area using the MADDPG algorithm structure.

[0086] Optionally, in one embodiment of this application, training the multi-agent distributed control model of the merging zone includes: obtaining the state variables of the target vehicle based on the MADDPG control model, and determining the merging zone state space of the target vehicle based on the state variables; determining the action space of the target vehicle in the merging zone based on the multi-vehicle cooperative control objective of the merging zone, and setting the merging zone reward function of the multi-agent distributed control model; and training the multi-agent distributed control model of the merging zone using the merging zone state space, the merging zone action space, and the merging zone reward function.

[0087] It should be noted that, similar to the training process of the multi-agent distributed control model in the preparation area described above, the embodiments of this application can describe the multi-vehicle merging problem in the merging area as a multi-agent Markov decision process and perform training to solve it.

[0088] In the embodiments of this application, the MADDPG algorithm can also be used to solve the merging zone scene and its control objective. This mainly includes calculating the state space, action space, and reward function of the merging zone scene. The specific process is as follows: 1. State space S Because the merging zone is short and can only accommodate a limited number of vehicles, a training method with a fixed number of vehicles merging is used; that is, only 6 vehicles are used in the training scenario. Figure 5 As shown, the merging process is completed on a 200m long acceleration lane. The goal of multi-vehicle control is for vehicles to merge safely as quickly as possible without collisions.

[0089] Therefore, vehicles entering the merging area The state variables are:

[0090] in, This is lane change information; it is 1 if a lane change is in progress, and 0 if a lane change is not in progress. in

[0091] The lane-changing process uses a fixed fifth-order polynomial curve:

[0092]

[0093] in, m As a parameter, it is limited by the maximum acceleration / deceleration, initial velocity, and maximum velocity; here, it can be taken as -1.08. For lane changing time, Lane width, It is time to change lanes.

[0094] 2. Movement space Actions in the merging area It is a two-dimensional vector, which includes the vehicle acceleration It also includes lane-changing instructions. The action space A is:

[0095] in, When lane change instruction When the value is 1, the vehicle will execute the above-mentioned fifth polynomial lane-changing curve, and the control quantity will be... It will become invalid.

[0096] 3. Reward Function Similar to the preparation area, the inflow area reward can be modified to some extent according to the control objective. In this embodiment, the single-step reward function for the inflow area is:

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] in, Similar to the reward settings in the preparation area, the speed on the main ramps in this section is approximately the same, therefore the speed on the main ramps... The items are consistent, and new ones have been added. The item represents the lane-changing bonus. If a lane change is successful without collision, a bonus is awarded. The smaller the speed loss during the lane change, the greater the bonus.

[0103] The final state reward at the end of the round is:

[0104] in, Indicates the weighting coefficient. This indicates the vehicle's delay time.

[0105] It is understood that the embodiments of this application train the multi-agent distributed control model of the merging area through the MADDPG control model, which can obtain the distributed multi-agent reinforcement learning merging model of the ramp merging area. This model is then used to output multi-vehicle cooperative control commands for the target ramp merging area by utilizing the multi-agent distributed control model of the preparation area under continuous traffic flow and the multi-agent distributed control model of the merging area under a fixed number of vehicles. This achieves cooperative ramp merging control of intelligent connected vehicles for highway scenarios, enabling efficient, safe and comfortable driving in the ramp area.

[0106] Therefore, the ramp merging strategy based on the MADDPG method in this application adopts different strategy networks in different regions, thereby effectively shortening the control time while ensuring high traffic efficiency, and it does not increase with the increase of the number of vehicles controlled, thus improving the multi-vehicle cooperative performance.

[0107] The method for multi-vehicle cooperative control of ramp merging areas proposed in this application for offline training phase involves dividing the target ramp merging area, determining the preparation area and merging area of ​​the target ramp merging area, identifying the training vehicle clusters in the preparation area and merging area, and formulating multi-vehicle cooperative control objectives for the training vehicle clusters. The method optimizes the preset MADDPG algorithm using the multi-vehicle cooperative control objectives to generate a MADDPG control model. Based on the MADDPG control model, a multi-agent distributed control model for the preparation area is trained to obtain a multi-agent distributed control model for the preparation area under continuous traffic flow. Based on the MADDPG control model, a multi-agent distributed control model for the merging area is trained to obtain a multi-agent distributed control model for the merging area under a fixed number of vehicles. The method then uses the multi-agent distributed control models for the preparation area under continuous traffic flow and the merging area under a fixed number of vehicles to output multi-vehicle cooperative control commands for the target ramp merging area. This simultaneously trains multiple agents to interact, enhancing the cooperative performance of multiple agents, effectively reducing computational complexity, and improving the real-time performance of the model.

[0108] Figure 6 This is a flowchart illustrating a method for multi-vehicle cooperative control in the merging zone of a ramp, provided as an embodiment of this application, applied during the online decision-making stage.

[0109] like Figure 6 As shown, the method for multi-vehicle cooperative control in the merging zone of this ramp includes the following steps: In step S601, the target ramp merging area is divided into preparation area, merging area and target vehicle cluster of the target ramp merging area.

[0110] In step S602, based on the pre-trained multi-agent distributed control model of the preparation area, the merging planning is performed on each vehicle in the target vehicle cluster of the preparation area to obtain the first decision result of each vehicle.

[0111] In step S603, based on the pre-trained multi-agent distributed control model of the merging area, merging planning is performed on each vehicle in the target vehicle cluster of the merging area to obtain the second decision result of each vehicle. The multi-agent distributed control model of the merging area and the multi-agent distributed control model of the prepared area are both trained by the MADDPG control model.

[0112] In step S604, based on the first decision result and the second decision result, a multi-vehicle cooperative control command for the continuous traffic flow in the target ramp merging zone is generated.

[0113] The method for multi-vehicle cooperative control of ramp merging areas proposed in this application, applied to the online decision-making stage, involves dividing the target ramp merging area to determine the preparation area, merging area, and target vehicle cluster. Based on a pre-trained multi-agent distributed control model for the preparation area, merging planning is performed for each vehicle in the target vehicle cluster in the preparation area, yielding a first decision result for each vehicle. Based on a pre-trained multi-agent distributed control model for the merging area, merging planning is performed for each vehicle in the target vehicle cluster in the merging area, yielding a second decision result for each vehicle. Both the merging area multi-agent distributed control model and the preparation area multi-agent distributed control model are trained using the MADDPG control model. Based on the first and second decision results, multi-vehicle cooperative control instructions for the continuous traffic flow in the target ramp merging area are generated. This effectively improves traffic efficiency while reducing the computational complexity of multi-vehicle cooperative control, enhancing multi-vehicle cooperative performance, and enabling efficient and safe merging of vehicles from the main ramp.

[0114] Secondly, with reference to the accompanying drawings, a device for multi-vehicle cooperative control in the merging zone of a ramp, according to an embodiment of this application, is described.

[0115] Figure 7 This is a block diagram of a device for multi-vehicle cooperative control of ramp merging areas applied in the offline training phase, according to an embodiment of this application.

[0116] like Figure 7As shown, the device 10 for multi-vehicle cooperative control in the merging zone of ramps during the offline training phase includes: a division module 100, an optimization module 200, a second training module 300, and a third training module 400.

[0117] The segmentation module 100 is used to segment the target ramp merging area, determine the preparation area and merging area of ​​the target ramp merging area, determine the training vehicle clusters in the preparation area and merging area, and formulate multi-vehicle collaborative control targets for the training vehicle clusters.

[0118] The optimization module 200 is used to optimize the preset MADDPG algorithm using the multi-vehicle cooperative control target to generate a MADDPG control model.

[0119] The second training module 300 is used to train the multi-agent distributed control model of the preparation area based on the MADDPG control model, so as to obtain the multi-agent distributed control model of the preparation area under continuous traffic flow.

[0120] The third training module 400 is used to train the multi-agent distributed control model of the merging area based on the MADDPG control model, so as to obtain the multi-agent distributed control model of the merging area with a fixed number of vehicles, and to output the multi-vehicle cooperative control command of the target ramp merging area using the multi-agent distributed control model of the preparation area under continuous traffic flow and the multi-agent distributed control model of the merging area with a fixed number of vehicles.

[0121] Optionally, in one embodiment of this application, the partitioning module 100 includes: a first control unit and a second control unit.

[0122] The first control unit is used to set a first control target for each vehicle in the training vehicle cluster in the preparation area, so as to control the time interval between each vehicle arriving at the end of the preparation area to meet the preset interval requirements.

[0123] The second control unit is used to set a second control target for each vehicle in the merging area, so as to control the training vehicle cluster to achieve the preset safe merging requirements according to the second control target.

[0124] Optionally, in one embodiment of this application, the second training module 300 includes: a determination unit, a calculation unit, a first acquisition unit, a first setting unit, and a first execution unit.

[0125] The determination unit is used to determine lane information, target vehicles, and multiple reference vehicles in the preparation area based on the MADDPG control model.

[0126] The calculation unit is used to calculate the distance and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area.

[0127] The first acquisition unit is used to obtain the state space of the preparation area of ​​the target vehicle based on the spacing, the first relative position, and the second relative position.

[0128] The first setting unit is used to determine the action space of the target vehicle in the preparation area based on the multi-vehicle cooperative control objective in the preparation area, and to set the reward function of the preparation area of ​​the multi-agent distributed control model.

[0129] The first execution unit is used to train the multi-agent distributed control model of the preparation zone using the state space, action space, and reward function of the preparation zone.

[0130] Optionally, in one embodiment of this application, the third training module 400 includes: a second acquisition unit, a second setting unit, and a second execution unit.

[0131] The second acquisition unit is used to acquire the state variables of the target vehicle based on the MADDPG control model, and determine the merging zone state space of the target vehicle based on the state variables.

[0132] The second setting unit is used to determine the action space of the target vehicle in the merging zone based on the multi-vehicle cooperative control objective of the merging zone, and to set the merging zone reward function of the multi-agent distributed control model.

[0133] The second execution unit is used to train the multi-agent distributed control model of the merging region using the merging region state space, merging region action space, and merging region reward function.

[0134] The apparatus for multi-vehicle cooperative control of ramp merging areas in the offline training phase, as proposed in the embodiments of this application, includes a partitioning module for partitioning the target ramp merging area, determining the preparation area and merging area of ​​the target ramp merging area, determining the training vehicle clusters in the preparation area and merging area, and formulating multi-vehicle cooperative control objectives for the training vehicle clusters; an optimization module for optimizing a preset MADDPG algorithm using the multi-vehicle cooperative control objectives to generate a MADDPG control model; a second training module for training a multi-agent distributed control model of the preparation area based on the MADDPG control model to obtain a multi-agent distributed control model of the preparation area under continuous traffic flow; and a third training module for training a multi-agent distributed control model of the merging area based on the MADDPG control model to obtain a multi-agent distributed control model of the merging area under a fixed number of vehicles. The apparatus outputs multi-vehicle cooperative control commands for the target ramp merging area using the multi-agent distributed control models of the preparation area under continuous traffic flow and the merging area under a fixed number of vehicles. This enhances the cooperative performance of multiple agents by simultaneously training their interaction, effectively reducing computational complexity and improving the real-time performance of the model.

[0135] Figure 8 This is a block diagram of a device for multi-vehicle cooperative control in the merging zone of a ramp, applied in the online decision-making stage according to an embodiment of this application.

[0136] like Figure 8 As shown, the device 20 for multi-vehicle cooperative control of ramp merging areas applied in the online decision-making stage includes: a partitioning module 500, a first planning module 600, a second planning module 700, and a control module 800.

[0137] The partitioning module 500 is used to divide the target ramp merging area and determine the preparation area, merging area and target vehicle cluster of the target ramp merging area.

[0138] The first planning module 600 is used to perform inflow planning for each vehicle in the target vehicle cluster in the preparation area based on a pre-trained multi-agent distributed control model in the preparation area, and obtain the first decision result for each vehicle.

[0139] The second planning module 700 is used to perform merging planning for each vehicle in the target vehicle cluster in the merging area based on the pre-trained multi-agent distributed control model of the merging area, and to obtain the second decision result for each vehicle. The multi-agent distributed control model of the merging area and the multi-agent distributed control model of the preparation area are both trained by the MADDPG control model.

[0140] The control module 800 is used to generate multi-vehicle coordinated control commands for the continuous traffic flow in the target ramp merging zone based on the first decision result and the second decision result.

[0141] It should be noted that the explanation of the above-mentioned method embodiment for multi-vehicle cooperative control in the ramp merging area also applies to the device for multi-vehicle cooperative control in the ramp merging area of ​​this embodiment, and will not be repeated here.

[0142] The device for multi-vehicle cooperative control of ramp merging areas in the online decision-making stage, according to an embodiment of this application, includes a partitioning module for dividing the target ramp merging area and determining the preparation area, merging area, and target vehicle cluster of the target ramp merging area; a first planning module for performing merging planning for each vehicle in the target vehicle cluster of the preparation area based on a pre-trained multi-agent distributed control model of the preparation area, obtaining a first decision result for each vehicle; a second planning module for performing merging planning for each vehicle in the target vehicle cluster of the merging area based on a pre-trained multi-agent distributed control model of the merging area, obtaining a second decision result for each vehicle, wherein both the merging area multi-agent distributed control model and the preparation area multi-agent distributed control model are trained by the MADDPG control model; and a control module for generating multi-vehicle cooperative control commands for continuous traffic flow in the target ramp merging area based on the first and second decision results, thereby effectively improving traffic efficiency while reducing the computational complexity of multi-vehicle cooperative control, improving multi-vehicle cooperative performance, and achieving efficient and safe merging of vehicles from the main ramp.

[0143] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 901, the processor 902, and the computer program stored on the memory 901 and capable of running on the processor 902.

[0144] When the processor 902 executes the program, it implements the method for multi-vehicle cooperative control in the ramp merging area provided in the above embodiments.

[0145] Furthermore, electronic devices also include: Communication interface 903 is used for communication between memory 901 and processor 902.

[0146] The memory 901 is used to store computer programs that can run on the processor 902.

[0147] The memory 901 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0148] If the memory 901, processor 902, and communication interface 903 are implemented independently, then the communication interface 903, memory 901, and processor 902 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0149] Optionally, in a specific implementation, if the memory 901, processor 902, and communication interface 903 are integrated on a single chip, then the memory 901, processor 902, and communication interface 903 can communicate with each other through an internal interface.

[0150] The processor 902 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0151] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for multi-vehicle cooperative control in a ramp merging zone.

[0152] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0153] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0154] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0155] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0156] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0157] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0159] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for multi-vehicle cooperative control in a ramp merging zone, characterized in that, When applied to the offline training phase, the method includes the following steps: The target ramp merging area is divided, the preparation area and the merging area of ​​the target ramp merging area are determined, the training vehicle clusters in the preparation area and the merging area are determined, and a multi-vehicle cooperative control target is formulated for the training vehicle clusters. The preset MADDPG algorithm is optimized using the multi-vehicle cooperative control objective to generate a MADDPG control model; Based on the MADDPG control model, a multi-agent distributed control model for the preparation area is trained to obtain a multi-agent distributed control model for the preparation area under continuous traffic flow. Based on the MADDPG control model, the multi-agent distributed control model of the merging area is trained to obtain the multi-agent distributed control model of the merging area with a fixed number of vehicles. The multi-agent distributed control model of the preparation area under continuous traffic flow and the multi-agent distributed control model of the merging area with a fixed number of vehicles are used to output the multi-vehicle cooperative control command of the target ramp merging area.

2. The method according to claim 1, characterized in that, The step of setting a multi-vehicle cooperative control objective for the training vehicle cluster includes: A first control objective is set for each vehicle in the training vehicle cluster in the preparation area, so as to control the time interval between each vehicle arriving at the end of the preparation area to meet the preset interval requirement. A second control target is set for each vehicle in the merging area, so as to control the training vehicle cluster to achieve the preset safe merging requirements according to the second control target.

3. The method according to claim 1, characterized in that, The training of the multi-agent distributed control model in the preparation area includes: Based on the MADDPG control model, lane information, target vehicle, and multiple reference vehicles in the preparation area are determined; Calculate the distance and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area; Based on the spacing, the first relative position, and the second relative position, the preparation area state space of the target vehicle is obtained; Based on the multi-vehicle cooperative control objective of the preparation area, the action space of the target vehicle in the preparation area is determined, and the reward function of the preparation area of ​​the multi-agent distributed control model is set. The multi-agent distributed control model of the preparation zone is trained using the state space of the preparation zone, the action space of the preparation zone, and the reward function of the preparation zone. The mathematical expression for the first relative position is: The mathematical expression for the second relative position is: in, Indicated by arrive The transition probability, For vehicles exist k The state quantity at any given moment.

4. The method according to claim 1, characterized in that, The training of the multi-agent distributed control model of the merging region includes: Based on the MADDPG control model, the state variables of the target vehicle are obtained, and the merging zone state space of the target vehicle is determined according to the state variables. Based on the multi-vehicle cooperative control objective of the merging zone, the action space of the target vehicle in the merging zone is determined, and the merging zone reward function of the multi-agent distributed control model is set. The multi-agent distributed control model of the inflow region is trained using the inflow region state space, the inflow region action space, and the inflow region reward function.

5. A method for multi-vehicle cooperative control in a ramp merging zone, characterized in that, Applied to the online decision-making stage, the method includes the following steps: The target ramp merging area is divided into sections, and the preparation area, merging area, and target vehicle cluster of the target ramp merging area are determined. Based on a pre-trained multi-agent distributed control model for the preparation area, an inflow planning is performed on each vehicle in the target vehicle cluster in the preparation area to obtain the first decision result for each vehicle. Based on the pre-trained multi-agent distributed control model of the merging area, merging planning is performed on each vehicle in the target vehicle cluster of the merging area to obtain the second decision result of each vehicle. The multi-agent distributed control model of the merging area and the multi-agent distributed control model of the preparation area are both trained by the MADDPG control model. Based on the first decision result and the second decision result, a multi-vehicle coordinated control command is generated for the continuous traffic flow in the target ramp merging zone.

6. A device for multi-vehicle coordinated control in a ramp merging zone, characterized in that, Applied to the offline training phase, including: The segmentation module is used to segment the target ramp merging area, determine the preparation area and merging area of ​​the target ramp merging area, determine the training vehicle clusters in the preparation area and the merging area, and formulate multi-vehicle cooperative control targets for the training vehicle clusters. The optimization module is used to optimize the preset MADDPG algorithm using the multi-vehicle cooperative control target to generate a MADDPG control model; The second training module is used to train the multi-agent distributed control model of the preparation area based on the MADDPG control model, so as to obtain the multi-agent distributed control model of the preparation area under continuous traffic flow. The third training module is used to train the multi-agent distributed control model of the merging area based on the MADDPG control model, to obtain the multi-agent distributed control model of the merging area with a fixed number of vehicles, so as to output the multi-vehicle cooperative control command of the target ramp merging area using the multi-agent distributed control model of the preparation area under continuous traffic flow and the multi-agent distributed control model of the merging area with a fixed number of vehicles.

7. The apparatus according to claim 6, characterized in that, The partitioning module includes: The first control unit is used to set a first control target for each vehicle in the training vehicle cluster in the preparation area, so as to control the time interval between each vehicle arriving at the end of the preparation area to meet the preset interval requirement. The second control unit is used to formulate a second control target for each vehicle in the merging area, so as to control the training vehicle cluster to achieve a preset safe merging requirement according to the second control target.

8. The apparatus according to claim 6, characterized in that, The second training module includes: The determining unit is used to determine the lane information, target vehicle, and multiple reference vehicles in the preparation area based on the MADDPG control model. The calculation unit is used to calculate the distance and first relative position of adjacent reference vehicles in different lanes in the preparation area, and the second relative position of each reference vehicle in the same lane in the preparation area. The first acquisition unit is used to obtain the preparation area state space of the target vehicle based on the spacing, the first relative position and the second relative position; The first setting unit is used to determine the action space of the preparation area of ​​the target vehicle according to the multi-vehicle cooperative control target of the preparation area, and to set the preparation area reward function of the multi-agent distributed control model. The first execution unit is used to train the multi-agent distributed control model of the preparation area using the state space of the preparation area, the action space of the preparation area, and the reward function of the preparation area. The mathematical expression for the first relative position is: The mathematical expression for the second relative position is: in, Indicated by arrive The transition probability, For vehicles exist k The state quantity at any given moment.

9. The apparatus according to claim 6, characterized in that, The third training module includes: The second acquisition unit is used to acquire the state variables of the target vehicle based on the MADDPG control model, and determine the merging zone state space of the target vehicle based on the state variables. The second setting unit is used to determine the action space of the target vehicle in the merging area according to the multi-vehicle cooperative control target in the merging area, and set the merging area reward function of the multi-agent distributed control model. The second execution unit is used to train the multi-agent distributed control model of the merging region using the merging region state space, the merging region action space, and the merging region reward function.

10. A device for multi-vehicle coordinated control in a ramp merging zone, characterized in that, Applied to the online decision-making stage, including: The partitioning module is used to divide the target ramp merging area and determine the preparation area, merging area and target vehicle cluster of the target ramp merging area; The first planning module is used to perform inflow planning for each vehicle in the target vehicle cluster in the preparation area based on a pre-trained multi-agent distributed control model in the preparation area, and to obtain the first decision result for each vehicle. The second planning module is used to perform merging planning for each vehicle in the target vehicle cluster in the merging area based on a pre-trained multi-agent distributed control model of the merging area, and to obtain the second decision result for each vehicle. The multi-agent distributed control model of the merging area and the multi-agent distributed control model of the preparation area are both trained by the MADDPG control model. The control module is used to generate multi-vehicle coordinated control commands for the continuous traffic flow in the target ramp merging zone based on the first decision result and the second decision result.

11. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method for multi-vehicle cooperative control in a ramp merging area as described in any one of claims 1-5.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method for multi-vehicle cooperative control in the merging zone of a ramp as described in any one of claims 1-5.