An Intelligent Connected Vehicle Cooperative Decision-making Method Based on Dual Interaction Perception
By introducing a dual interactive perception mechanism and a multi-agent reinforcement learning framework in the collaborative decision-making method of intelligent connected vehicles, the limitations of improving efficiency of hybrid transportation systems in the existing technology are solved, and collaborative decision-making between intelligent connected vehicles and artificially driven vehicles is realized, and overall traffic efficiency is improved.
Patent Information
- Application Number
- CN202411524399.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-29
AI Technical Summary
The prior art has limitations in improving the efficiency of hybrid transportation systems, including high computational costs based on rules or optimization methods and high training costs of single agent reinforcement learning methods and difficulty in dealing with the problem of coordination among multiple vehicles.
Using the intelligent connected vehicle collaborative decision-making method based on dual interaction perception, the dual mutual perception collaborative control strategy DIACC is designed through a multi-agent reinforcement learning framework, and the distributed interactive adaptive decision-making module D-IADM and centralized interaction enhancement evaluator module C-IEC are used to learn the interaction characteristics and traffic environment information with surrounding vehicles, and capture the relationship between global vehicle interaction and traffic dynamics.
The coordinated decision-making of intelligent connected vehicles under different types of vehicles and driving styles has been realized, the overall traffic efficiency has been improved, the conflicts and lane change behaviors between vehicles have been reduced, and the stability and efficiency of traffic flow have been improved.
Smart Images

Figure CN119428755B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to an intelligent connected vehicle collaborative decision-making method based on dual interaction perception. Background Art
[0002] With the rapid increase in the global vehicle ownership, traffic congestion has become one of the restrictive factors hindering the rapid development of cities. Therefore, improving traffic efficiency has become an urgent and highly concerned research topic. Research shows that intelligent connected vehicles have great potential in improving traffic performance, and human-driven vehicles will coexist with intelligent connected vehicles in the foreseeable future. Human drivers often make decisions dynamically and randomly based on the reactions of surrounding vehicles and generally lack a sense of cooperation. Therefore, human-driven vehicles usually exhibit different driving styles, which pose significant challenges to intelligent connected vehicles. They must adapt to these uncertain and diverse behaviors and make collaborative decisions that contribute to overall traffic optimization.
[0003] In the prior art, the solutions proposed to improve the efficiency of the mixed traffic system can be roughly divided into two categories. The first category includes rule-based or optimization-based methods, and the second category is single-agent reinforcement learning-based methods. The existing research has the following limitations: 1. Rule-based or optimization-based methods rely on the modeling of traffic environments and vehicle behaviors, and the centralized optimization methods that can obtain optimal solutions have a relatively high computational cost; 2. Single-agent reinforcement learning methods are more suitable for solving local control problems but face challenges in global traffic optimization because they require a large amount of interaction experience to learn effective decision-making strategies, making the training cost high and large-scale deployment difficult; 3. Single-agent reinforcement learning methods are difficult to handle the coordination problems between multiple vehicles and usually require additional rule guidance or manually designed coordinated interaction settings. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent connected vehicle collaborative decision-making method based on dual interaction perception, which learns the interaction characteristics with surrounding vehicles and traffic environment information to adapt to different types of vehicles and driving styles; and better understands the impact of vehicle interaction on traffic evolution from a global perspective, thereby guiding the update of the intelligent connected vehicle collaborative strategy.
[0005] To achieve the above purpose, the present invention provides an intelligent connected vehicle collaborative decision-making method based on dual interaction perception, including the following steps:
[0006] Design a dual mutual perception collaborative control strategy DIACC using a centralized training and distributed execution multi-agent reinforcement learning framework;
[0007] Use a distributed interactive adaptive decision-making module D-IADM to learn the interaction characteristics with surrounding vehicles and traffic environment information;
[0008] The centralized interactive enhanced evaluator module C-IEC is used to capture the relationship between global vehicle interactions and global traffic dynamics.
[0009] Preferably, the dual mutual perception collaborative control strategy DIACC includes a centralized critic network and several distributed actor networks;
[0010] The centralized critic network assists in the training phase by accessing the global information of the mixed traffic system, and the distributed actor networks use the global information provided by the centralized critic network to maximize their individual rewards.
[0011] Preferably, the state space, action space, and reward function in the multi-agent reinforcement learning framework are designed according to the evolution characteristics of the mixed traffic system and vehicle interaction requirements;
[0012] The designed state space in the multi-agent reinforcement learning framework includes the motion state of each vehicle, the traffic flow statistical information of each lane, and the static road structure information; the discrete quantities of the designed action space include: maintaining, changing lanes to the left, changing lanes to the right, accelerating, and decelerating; the designed reward function is:
[0013]
[0014] where r e , r g and r tp represent the ego-vehicle reward, global reward, and inactivity penalty, or time penalty respectively, and the coefficients w e =ρ, w g =1 - ρ, and w tp =n - 1 + ρ are the weights of each reward component related to the penetration rate of connected and automated vehicles; r done is the completion reward, and the ego-vehicle reward r e contains several sub-rewards:
[0015] r e =w e,v r e,v +w e,w r e,w +w e,c r e,c ;
[0016]
[0017]
[0018]
[0019] where r e,v is the speed reward;e,w is the warning distance penalty; r e,c is the collision penalty, w e,v , w e,w and w e,c are the weight parameters of r e,v , r e,w and r e,c respectively; v max is the maximum speed allowed by road traffic regulations; the set represents the vehicles within the warning range; the set represents the vehicles within the actual collision range; is the relative longitudinal distance between vehicles; is the relative lateral distance between vehicles, j is or an element in the set; and 's distance thresholds are d th,w and d th,c respectively; p e,v is the speed reward adjustment parameter; p e,c is the collision penalty adjustment parameter.
[0020] The global reward r g reflects the overall performance of all connected and automated vehicles in terms of average speed, specifically:
[0021]
[0022] where, p g,v is the global speed reward adjustment parameter;
[0023] r tp reflects the impact of the speed efficiency of connected and automated vehicles on the overall traffic flow:
[0024] r tp = 10 × sigmoid(v i - v th );
[0025] where, v th is the speed threshold used to adjust the penalty level at different speeds, and sigmoid(·) is the Sigmoid function.
[0026] Preferably, D-IADM includes a Trajectory-Aware Interaction Encoder (TAIE). Through TAIE, the interaction features between the ego vehicle and surrounding vehicles are obtained. Then, it combines the previous decision output with a Gated Recurrent Unit (GRU) and a decision layer to generate a learnable action output, and inputs the action output into an Active Safety-based Action Filter (PSAF).
[0027] Preferably, an interaction graph of the host vehicle and surrounding human-driven vehicles is established, taking the current state information and historical trajectory information of the host vehicle and surrounding human-driven vehicles as inputs, and using a multi-head graph attention mechanism to extract interaction features with surrounding human-driven vehicles;
[0028] An interaction graph of the host vehicle and surrounding connected vehicles is established, taking the current state information of the host vehicle and surrounding connected vehicles and the decision output of the previous moment as inputs, and using a multi-head graph attention mechanism to extract interaction features with surrounding connected vehicles;
[0029] Taking the traffic flow statistical information of adjacent lanes and static road structure information as inputs, a multi-layer perceptron is used to extract traffic environment context information.
[0030] Preferably, the interaction features of surrounding human-driven vehicles, the interaction features of surrounding connected vehicles, the traffic environment context information, and the decision instruction of the previous moment are fused to obtain comprehensive features, and the comprehensive features are input into a gated recurrent unit GRU and a decision layer to obtain a learnable action output;
[0031] The proactive safety action filter PSAF uses the time to collision TTC as the proactive safety evaluation index, optimizes the actions with TTC lower than the set threshold, and outputs the longitudinal decision and lateral decision of the connected vehicle.
[0032] Preferably, the centralized interaction enhancement evaluator module C-IEC receives all state information from the environment. The C-IEC includes an integrated traffic dynamics representation module ITDR, which captures traffic dynamics and vehicle interaction features and outputs a global state value function.
[0033] Preferably, a multi-layer perceptron is used to obtain the traffic flow characteristics of each lane, the movement characteristics of human-driven vehicles, and the movement characteristics of connected vehicles in the mixed traffic system, and then combined with the static road structure information to obtain comprehensive traffic flow dynamic characteristics.
[0034] Preferably, an interaction graph of all vehicles in the mixed traffic system is established, taking the current movement state of each vehicle as an input, and using a graph attention mechanism to obtain global vehicle interaction features.
[0035] Preferably, using a multi-head cross-attention mechanism, taking the comprehensive traffic flow dynamic characteristics as the query object, and taking the global vehicle interaction features as the key and value, to capture the characteristics of traffic dynamics and vehicle interactions.
[0036] Therefore, the present invention adopts the above-mentioned collaborative decision-making method for connected vehicles based on dual interaction perception, and the beneficial effects are as follows:
[0037] (1) The present invention utilizes a distributed interactive adaptive module D-IADM to learn the interaction characteristics with surrounding vehicles and traffic environment information, enabling intelligent connected vehicles to make reasonable collaborative decisions, plan merging routes in advance, and yield to surrounding vehicles at appropriate times, thereby improving the overall traffic efficiency.
[0038] (2) The present invention utilizes a centralized interactive enhanced evaluator module C-IEC to better understand the impact of vehicle interactions on traffic evolution from a global perspective, thereby guiding the update of the collaborative strategies of intelligent connected vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the overall scheme framework diagram of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0040] Figure 2 is a schematic diagram of a mixed traffic scenario of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention, where (a) is a schematic diagram of a pure human-driven vehicle traffic scenario, and (b) is a schematic diagram of a mixed traffic scenario of intelligent connected vehicles and human-driven vehicles adopting collaborative strategies;
[0041] Figure 3 is a schematic diagram of the static road structure division of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0042] Figure 4 is the network framework diagram of a trajectory perception interaction encoder TAIE of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0043] Figure 5 is a schematic diagram of the collision time adjustment decision output of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention, where (a) is a schematic diagram of lane-changing safety, and (b) is a schematic diagram of following vehicle safety judgment;
[0044] Figure 6 is the network framework diagram of an integrated traffic dynamic representation module of an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0045] Figure 7 is a schematic diagram of the average speed comparison of 15 vehicles merging through a bottleneck area in an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0046] Figure 8 is a schematic diagram of the average speed comparison of 25 vehicles merging through a bottleneck area in an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception of the present invention;
[0047] Figure 9 It is a schematic diagram comparing vehicle trajectories generated by the confluence of 15 vehicles in a bottleneck area in an embodiment of a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception according to the present invention. Among them, (a) is a schematic diagram of a mixed traffic flow using a dual interaction perception collaborative control strategy with an intelligent connected vehicle penetration rate of 0.4; (b) is a mixed traffic flow using a multi-agent reinforcement learning algorithm with an intelligent connected vehicle penetration rate of 0.4; (c) is a pure manual driving traffic flow generated by the SUMO simulator. Detailed implementation manners
[0048] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0049] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0050] Embodiment
[0051] As Figure 1 shown, a collaborative decision-making method for intelligent connected vehicles based on dual interaction perception includes the following steps:
[0052] A multi-agent reinforcement learning framework of centralized training and distributed execution is designed, and a dual interaction perception collaborative control strategy DIACC is designed; the dual interaction perception collaborative control strategy DIACC includes a centralized critic network and several distributed executor networks.
[0053] The centralized critic network assists the training phase by accessing the global information of the mixed traffic system, and the distributed executor networks maximize their individual rewards using the global observations provided by the centralized critic network.
[0054] As Figures 2 - 3 shown, first, the optimization problem of the mixed traffic system is modeled and characterized as a distributed partially observable Markov decision process. The multi-agent reinforcement learning model can be represented as a seven-tuple {N, S, A, O, R, P, γ}, where N is the set of intelligent connected vehicles in the mixed traffic environment, involving n intelligent connected vehicles; S represents the global state space of the environment; A is the joint action space of all intelligent connected vehicles; O represents the joint observation space of all intelligent connected vehicles; R is the joint reward function; P is the state transition probability function; γ is the discount factor, which is used to balance the importance of immediate rewards and future rewards.
[0055] In the setting, each intelligent connected vehicle i observes its local state at any time t and selects a control decision instruction according to its policy This strategy is characterized by the parameter θ. The collective goal of all connected and automated vehicles is to optimize the collective policy to maximize the expected cumulative reward within L time steps.
[0056] Design the state space, action space, and reward function in the multi-agent reinforcement learning framework according to the evolution characteristics of the mixed traffic system and the vehicle interaction requirements.
[0057] Among them, the state space includes the position (x, y), speed v, heading angle of each vehicle lane index l, and vehicle type c. Based on this, the roadside unit can obtain information about specific lanes, including the number of vehicles n l , density D l , average speed , and the penetration rate ρ of connected and automated vehicles l . In addition, the high-precision map provides detailed road structure information, and this data constitutes the state information that the centralized enhanced evaluator module can read from the environment.
[0058] Divide the bottleneck scenario into three regions: multi-lane region, pre-merging region, and post-merging region. In the multi-lane region and the post-merging region, vehicles are evenly distributed in each lane and the speed is stable; in the pre-merging region, vehicles concentrate on the central lane, and lateral conflicts increase, resulting in more frequent lane changes and greater speed fluctuations. The behavioral differences are due to the reduction in the number of lanes and traffic capacity in the bottleneck merging region, leading to increased vehicle competition and affecting driving behavior. The above changes are driven by the static road structure. Therefore, the designed local observation state of connected and automated vehicles includes: its own vehicle state information, neighboring vehicle states, neighboring lane traffic flow statistics information, and static road structure information.
[0059] The action space is discrete decision-making behaviors, involving the lateral and longitudinal control of the vehicle. Specifically, there are five decision-making behaviors: maintaining the current lane and speed (denoted as "maintain"), changing lanes to the left or right, accelerating or decelerating.
[0060] The reward function is:
[0061]
[0062] Among them, r e , r g , and r tp represent the self-vehicle reward, global reward, and inactivity penalty (or called time penalty) respectively. The coefficients w e = ρ, w g = 1 - ρ, and w tp = n - 1 + ρ are the weights of each reward component related to the penetration rate of connected and automated vehicles. r doneTo complete the reward and encourage the intelligent connected vehicle to quickly pass through the bottleneck area, thereby improving the traffic capacity. The reward for the vehicle itself includes several sub-rewards:
[0063] r e = w e,v r e,v + w e,w r e,w + w e,c r e,c ;
[0064]
[0065]
[0066]
[0067] Among them, r e,v is the speed reward; r e,w is the warning distance penalty; r e,c is the collision penalty, w e,v , w e,w and w e,c are the weight parameters of r e,v , r e,w and r e,c respectively; v max is the maximum speed allowed by the road traffic regulations; the set represents the vehicles within the warning range; the set represents the vehicles within the actual collision range; is the relative longitudinal distance between vehicles; is the relative lateral distance between vehicles, and j is or an element in the set; and The distance thresholds of are d th,w and d th,c respectively; p e,v is the speed reward adjustment parameter; p e,c is the collision penalty adjustment parameter.
[0068] The global reward r g reflects the overall performance of all intelligent connected vehicles in terms of average speed, and its specific form is:
[0069]
[0070] Among them, p g,v is the global speed reward adjustment parameter.
[0071] r tp Emphasizes the impact of the speed efficiency of intelligent connected vehicles on the overall traffic flow:
[0072] r tp = 10 × sigmoid(v i - v th );
[0073] where v th is the speed threshold for adjusting the penalty level at different speeds, and sigmoid(·) is the Sigmoid function. When dealing with the global traffic coordination problem, some connected and automated vehicles (CAVs) choose to decelerate or even stop in multi-lane areas, which greatly increases the possibility of conflicts and congestion in the merging area. This design can avoid the inert behavior of CAVs when facing congestion.
[0074] Utilize the distributed interactive adaptive decision-making module D-IADM to learn the interaction characteristics with surrounding vehicles and traffic environment information.
[0075] As Figure 4 shown, a CAV obtains local observations from the environment at each moment and outputs its control actions through the distributed interactive adaptive decision-making module D-IADM. D-IADM first obtains the interaction characteristics between its own vehicle and surrounding vehicles through the trajectory-aware interaction encoder TAIE.
[0076] The local observation state of a CAV includes its own state information, neighboring vehicle movement information, lane statistics information, and static road structure information. The set of neighboring vehicles includes the six nearest vehicles in the same and adjacent lanes. Considering the behavioral differences between human-driven vehicles (HDVs) and CAVs, we classify the neighboring vehicles into the set of human-driven vehicles and the set of connected and automated vehicles where H represents a human-driven vehicle (HDV); C represents a connected and automated vehicle (CAV).
[0077] Taking the extraction of the interaction characteristics of human-driven vehicles as an example, an interaction graph is constructed between a CAV and surrounding human-driven vehicles, and each CAV is connected to the neighboring human-driven vehicles within its specified range. The interaction graph of human-driven vehicles is represented as where represents nodes, is the edge connecting these nodes. It should be noted that these edges only connect a CAV to its neighboring human-driven vehicles, and there is no connection between human-driven vehicles.
[0078] The original input consists of the current and historical states of a CAV and surrounding human-driven vehicles: where represents the number of surrounding human-driven vehicles, and T is the control period. The state information m of each vehicle includes position, speed, and heading angle. are the state information of the connected and autonomous vehicle i at different past moments; are respectively the state information of adjacent human-driven vehicles at different past moments; is the state information of the connected and autonomous vehicle i at the current moment; are the state information of adjacent human-driven vehicles at the current moment. First, connect the historical information of each vehicle. First, connect the historical information of each vehicle, and then embed it into the node features:
[0079]
[0080] where h is the motion feature corresponding to each node; S H is the size of the hidden layer of the network part that processes the interaction with human-driven vehicles. Then, input these features into a multi-head graph attention network (MH-GAT) to extract the interaction features of human-driven vehicles:
[0081]
[0082] where are the interaction features corresponding to each node; is the interaction feature of human-driven vehicles at the current moment.
[0083] A similar process is used to generate the interaction graph of connected and autonomous vehicles. Different from the interaction graph of human-driven vehicles, the input of this layer includes the current state and the decision information of the connected and autonomous vehicle i and its surrounding connected and autonomous vehicles at the previous moment:
[0084]
[0085] where represents the number of surrounding connected and autonomous vehicles; is the decision information of the connected and autonomous vehicle at the previous moment. Finally, generate the interaction features of the connected and autonomous vehicle:
[0086]
[0087] where are the interaction features of connected and autonomous vehicles at the current moment; S C is the size of the hidden layer that processes the interaction part of connected and autonomous vehicles.
[0088] The self-attention mechanism quantifies the relative importance of each surrounding human-driven vehicle to the connected and autonomous vehicle i. These processes generate rich node feature representations, enhancing the decision-making ability of the connected and autonomous vehicle i.
[0089] Extract local traffic context information. The present invention uses traffic flow statistical information from its own vehicle lane and adjacent lanes, and the information of the lane where the vehicle is located is represented as The information of the left lane is represented as The information of the right lane is represented as and the static road structure information F static . These information are processed by a multi-layer perceptron (MLP) to obtain local traffic context features
[0090] Furthermore, the interaction features of the human-driven vehicle, the interaction features of the connected and autonomous vehicle, and the local traffic context features are combined to form a comprehensive interaction environment. All the embedded features are then integrated in a fusion layer based on the multi-layer perceptron to form the comprehensive observation embedding features at the time step
[0091] Then, the comprehensive observation embedding features and the previous control decision instruction are input into a gated recurrent unit (GRU) and a decision layer to generate a learnable decision instruction The GRU captures the temporal dependence of the observation features between different moments, enhancing the continuity of the decision-making Subsequently, it is input into an active safety-based action filter (PSAF). This filter evaluates the safety of actions based on the time-to-collision (TTC) criterion, filters out actions with a high collision risk, and provides a safer decision under the current conditions. At the same time, the PSAF maps the decision output to specific lateral and longitudinal control decisions. The lateral control decision involves a lane-changing action, which is represented as
[0092]
[0093] The longitudinal control decision involves a speed action, which is represented as where v max is the maximum speed allowed by the road traffic regulations
[0094] As Figure 5 shown, when is a lane-changing action, the PSAF first evaluates the vehicles in front of and behind in the target lane. The relevant indicators of the vehicle in front in the target lane are represented as where is the time headway between the host vehicle and the vehicle in front; is the relative distance between the host vehicle and the vehicle in front. The indicators of the vehicle behind are represented as is the time headway between the host vehicle and the vehicle behind; This is the relative distance between the host vehicle and the following vehicle. As shown in (a), if the TTC or relative distance of the vehicle in the target lane is lower than a certain threshold, i.e., the area within the blue wireframe, PSAF will reject the lane change action and keep the vehicle in the current lane. If the TTC or relative distance from the vehicle ahead is lower than a certain threshold, i.e., the area within the orange wireframe, PSAF will recommend decelerating for a lane change, and the deceleration calculation formula is:
[0095]
[0096] where is the relative speed between the intelligent connected vehicle and the vehicle ahead. In other cases, the vehicle will change lanes at a constant speed. When is for the lane keeping action, PSAF evaluates the TTC and relative distance between the current lane and the vehicle ahead. As shown in (b), if the TTC or relative distance is lower than a certain threshold (dark blue area), PSAF will output a deceleration action, and its calculation method is similar to the lane change scenario. The light blue area indicates that the vehicle can maintain the current speed; the gray area indicates that the vehicle follows the decision instruction The acceleration or deceleration is set to 2 m / s². Based on the current speed of the intelligent connected vehicle and the calculated acceleration or deceleration, the next speed control action is determined, and this action becomes the control decision instruction output by the D-IADM module for controlling the intelligent connected vehicle.
[0097] The centralized interactive enhanced evaluator module C-IEC is used to capture the relationship between global vehicle interactions and global traffic dynamics.
[0098] As Figure 6 shown, the C-IEC module first obtains global traffic dynamic features through the integrated traffic dynamic representation module (ITDR). In the ITDR module, traffic information is divided into three layers: static road structure information, dynamic traffic information, and vehicle interaction information. The information of each layer is processed separately. Among them, the static road structure information and dynamic traffic information are processed by a multi-layer perceptron (MLP). The static road information is represented as F static , and the dynamic traffic information includes global lane statistics information, the status information of all intelligent connected vehicles, and the status information of all human-driven vehicles. After being processed by the MLP, the corresponding embedded features are obtained, namely the traffic flow feature the intelligent connected vehicle motion feature and the human-driven vehicle motion feature Then these features are concatenated to form the dynamic embedded feature
[0099] The vehicle interaction information is processed by a graph attention network (GAT) to extract the global vehicle interaction feature. Here, a global vehicle interaction graph is constructed, denoted as Gglobal =(V,E), where V represents all vehicle nodes and E is the edge connecting these nodes, where only adjacent vehicles are connected, i.e. The information of each vehicle node includes position, speed and heading angle. These data are first passed through the embedding layer to obtain node features. Among them, h represents the characteristics corresponding to each node; N represents the total number of vehicles in the mixed traffic environment; S N is the hidden layer size of the network part that handles all vehicle interactions. The node features are then fed into GAT to extract interaction features is the interaction feature corresponding to each node.
[0100] The present invention uses dynamic traffic information embedding features as query objects, global vehicle interaction features as key values, and uses a multi-head cross attention mechanism to calculate global vehicle interaction features. Key Projection Matrix Sum Projection Matrix The query matrix Q can be obtained h , key matrix K h , value matrix V h :
[0101]
[0102] Then, the scaled dot-product attention is calculated, and the concatenation and final linear transformation are applied. By using multi-head criss-cross attention, we are able to capture and map the relationship between traffic dynamics and vehicle interactions, thus gaining a comprehensive understanding of traffic evolution.
[0103] These features are then fed into the GRU module to capture the temporal dependencies between global states. Finally, the fully connected layer outputs the global state value function and advantage function evaluation at the current time step t to assess the overall state quality and guide the parameter update in D-IADM.
[0104] In order to evaluate the effectiveness and superiority of the method proposed in the present invention, a mixed traffic system simulation environment was built based on the SUMO simulator and the effect of the method of the present invention was tested. In the test scenario, the penetration rate of intelligent connected vehicles is 0.4. At the same time, three different driving styles of manually driven vehicles are set based on the vehicle control model provided by SUMO: aggressive, normal and cautious. Aggressive driving is characterized by shorter following distances and more frequent lane changes, while cautious driving is characterized by longer following distances and fewer lane changes. Normal driving uses the default parameters of the SUMO model.
[0105] We evaluated the average vehicle speed in the 15-vehicle and 25-vehicle scenarios using three settings: pure manual driving traffic using the IDM+LC2013 model (SUMO’s built-in vehicle control model); mixed traffic using the MAPPO-IADM model (a multi-agent reinforcement learning algorithm that integrates the interactive adaptive decision module D-IADM designed in this paper) when the ICV penetration rate is 0.4; and mixed traffic using the DIACC model when the ICV penetration rate is 0.4.
[0106] like Figures 7 - 8 As shown in the figure, the performance of all vehicles in three areas is evaluated: multi-lane area, pre-merge area, and post-merge area. The mixed traffic scenarios using MAPPO-IADM and DIACC models show higher average speeds than the pure human-driven car scenario using the SUMO default model. The initial speed of all vehicles is 10m / s. In the multi-lane area, the vehicles mainly accelerate, and the performance of the three traffic settings is similar. However, in the pre-merge area, the pure human-driven car traffic shows a lower average speed, while the mixed traffic using MAPPO-IADM and DIACC models maintains a higher average speed, especially in the 25-vehicle scenario. In addition, the distribution of vehicle average speeds is more concentrated in the mixed traffic scenarios using these two models, while the distribution of vehicle average speeds is wider in the pure human-driven car traffic scenario. This shows that when using MAPPO-IADM and DIACC models, the intelligent connected vehicles better coordinate the traffic flow in the pre-merge area, which can not only achieve its own efficient driving, but also guide the surrounding human-driven cars to maintain a higher speed, thereby improving the overall traffic efficiency. In addition, the mixed traffic using the MAPPO-IADM model has a wider speed distribution in the pre-merge area and the post-merge area, while the mixed traffic using the DIACC model has a more concentrated speed distribution.
[0107] In addition, if Figure 9 As shown, the present invention evaluates vehicle trajectories using different models in a 15-vehicle scenario. In the pure manually driven car traffic scenario, congestion occurred in the pre-merge area, which was manifested by intensified vehicle conflicts and multiple lane changes. In addition, there were two vehicles stationary in this area, and the resulting waiting time undoubtedly increased the completion time of the vehicle merge. In contrast, the mixed traffic system using the MAPPO-IADM model performed better overall, and the intelligent connected vehicles began to adjust their behavior in the multi-lane area. In the pre-merge area, vehicle conflicts were significantly reduced, and lane changes were also reduced, although one vehicle still experienced waiting time. The mixed traffic system using the DIACC model performed best, with reduced lane changes in the pre-merge area, no vehicles experiencing waiting time, smoother vehicle trajectories, and higher behavioral consistency.
[0108] Therefore, the present invention adopts the above-mentioned intelligent connected vehicle collaborative decision-making method based on dual interaction perception, making the intelligent connected vehicle play a crucial role in coordinating traffic in bottleneck scenarios. On the one hand, the intelligent connected vehicle actively adjusts its behavior in advance, reducing the lateral conflict between the pre-merge area and other vehicles. On the other hand, the behavior of the intelligent connected vehicle actively guides and coordinates human-driven vehicles, thus improving the overall traffic performance.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A collaborative decision-making method for intelligent connected vehicles based on dual interactive perception, characterized in that: The following steps are involved: Using the centralized training and distributed execution multi-agent reinforcement learning framework, a dual interactive perception collaborative control strategy DIACC is designed; Use the distributed interactive adaptive decision module D-IADM to learn the interaction characteristics and traffic environment information with surrounding vehicles; The centralized interaction enhancement evaluator module C-IEC is used to capture the relationship between global vehicle interactions and global traffic dynamics; The dual interactive perception collaborative control strategy DIACC includes a centralized critic network and several distributed actor networks; The centralized critic network assists the training phase by accessing the global information of the mixed traffic system, and the distributed actor network maximizes its individual rewards by leveraging the global information provided by the centralized critic network. Design the state space, action space and reward function in the multi-agent reinforcement learning framework according to the evolution characteristics of the hybrid traffic system and the vehicle interaction requirements; Designing the state space under the multi-agent reinforcement learning framework includes the motion state of each vehicle, traffic flow statistics of each lane, and static road structure information; The discrete quantities in the designed action space include: hold, left lane change, right lane change, acceleration, and deceleration; the designed reward function is: ; in, , and Represent the vehicle reward, global reward and lazy behavior penalty, or time penalty, respectively. The coefficient , and is the weight of each reward component related to the penetration rate of intelligent connected vehicles; To complete the reward, the vehicle reward Contains several sub-rewards: ; ; ; ; in, Reward for speed; For warning distance penalties; Penalty for collision, , and They are , and The weight parameter of is the maximum speed allowed by road traffic regulations; Indicates vehicles within the warning range; assembly Indicates the vehicles within the actual collision range; is the relative longitudinal distance between vehicles; is the relative lateral distance between vehicles, for or Elements in a collection; and The distance thresholds are and ; Adjust parameters for speed bonus; Adjust parameters for collision penalties; Global Rewards Reflects the overall performance of all intelligent connected vehicles in terms of average speed, which is as follows: ; in, Adjust parameters for global speed bonus; Reflects the impact of the speed efficiency of intelligent connected vehicles on the overall traffic flow: ; in, is the speed threshold used to adjust the penalty level at different speeds, for Sigmoid function.
2. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 1 is characterized by: D-IADM contains a trajectory-aware interaction encoder TAIE, which obtains the interaction features between the vehicle itself and surrounding vehicles. Then, it combines the previous decision output with the gated recurrent unit GRU and the decision layer to generate a learnable action output, which is then input into the active safety-based action filter PSAF.
3. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 2 is characterized by: Establish an interaction graph between the vehicle and surrounding human-driven cars, use the current state information and historical trajectory information of the vehicle and surrounding human-driven cars as input, and use the multi-head graph attention mechanism to extract the interaction features with surrounding human-driven cars; Establish an interaction graph between the vehicle and surrounding intelligent connected vehicles, use the current state information of the vehicle and surrounding intelligent connected vehicles and the decision output at the last moment as input, and use the multi-head graph attention mechanism to extract the interaction features with surrounding intelligent connected vehicles; Taking the traffic flow statistics of adjacent lanes and the static road structure information as input, a multi-layer perceptron is used to extract the traffic environment context information.
4. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 2 is characterized by: The interaction features of surrounding human-driven cars, the interaction features of surrounding intelligent connected cars, the traffic environment context information and the decision instructions at the previous moment are integrated to obtain comprehensive features, which are then input into the gated recurrent unit GRU and the decision layer to obtain learnable action outputs; The active safety action filter PSAF uses the collision time TTC as an active safety evaluation indicator, optimizes actions when the TTC is lower than the set threshold, and outputs the longitudinal and lateral decisions of the intelligent connected vehicle.
5. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 1 is characterized by: The centralized interaction enhancement evaluator module C-IEC receives all state information from the environment. C-IEC includes an integrated traffic dynamic representation module ITDR, which captures traffic dynamics and vehicle interaction characteristics and outputs a global state value function.
6. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 5 is characterized by: The multi-layer perceptron is used to obtain the traffic flow characteristics of each lane in the mixed traffic system, the movement characteristics of manually driven cars, and the movement characteristics of intelligent connected cars, and then the dynamic characteristics of the comprehensive traffic flow are obtained by combining the static road structure information.
7. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 6 is characterized by: An interaction graph of all vehicles in a mixed traffic system is established. The current motion state of each vehicle is used as input, and the graph attention mechanism is used to obtain the global vehicle interaction features.
8. The intelligent connected vehicle collaborative decision-making method based on dual interactive perception according to claim 7 is characterized by: The multi-head cross-attention mechanism is used to capture the characteristics of traffic dynamics and vehicle interactions by taking comprehensive traffic flow dynamic features as query objects and global vehicle interaction features as keys and values.
Citation Information
Patent Citations
Multi-vehicle formation decision-making method and system based on communication and multi-agent reinforcement learning
CN117539254A
Interconnection automatic driving decision-making method based on collaborative awareness and adaptive information fusion
CN117922612A