Mixed traffic flow cooperative platoon control method and device, system and storage medium

By constructing a dynamic spatiotemporal directed graph and graph attention mechanism, and combining the CTDE architecture and multi-agent near-end policy optimization network, the problems of limited perception topology and insufficient global coordination capability in hybrid vehicle fleet control are solved, and efficient, safe and environmentally friendly collaborative control of hybrid traffic flow is achieved.

CN122116619APending Publication Date: 2026-05-29GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610325878.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for hybrid vehicle fleet control suffer from problems such as the perception topology being affected by human-driven vehicles, insufficient global coordination capabilities, and low efficiency in environmental modeling and learning for multi-objective optimization. These issues lead to a trade-off between the accuracy of collaborative control and fuel economy in time-varying hybrid traffic flows.

Method used

A multidimensional state matrix and dynamic spatiotemporal directed graph of hybrid traffic flow are constructed. The graph adjacency matrix is ​​dynamically evolved based on dual decision rules. The graph attention mechanism is used to extract the forward look-ahead feature. A multi-agent near-end policy optimization network based on CTDE architecture is constructed. A multimodal objective reward function is designed. Combined with a one-step kinematic expectation-based underlying safety supervision mechanism, the coordinated control of global stability and local safety is achieved.

Benefits of technology

It achieves forward line-of-sight dynamic topology perception in time-varying mixed traffic flows, overcomes the local optima of "selfish learning", ensures system-level environmental safety and global stability, and improves fuel economy and traffic flow stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116619A_ABST
    Figure CN122116619A_ABST
Patent Text Reader

Abstract

The application discloses a mixed traffic flow cooperative formation control method and device, system and storage medium, comprising: constructing a mixed traffic flow multi-dimensional state matrix and a dynamic space-time directed graph; generating a graph adjacency matrix dynamic evolution based on a double judgment rule according to the dynamic space-time directed graph; extracting a lead sight distance feature based on a graph attention mechanism according to the graph adjacency matrix dynamic evolution; constructing a multi-agent near-end strategy optimization network based on a CTDE architecture according to the lead sight distance feature; and performing bottom-layer safety supervision mechanism intervention based on one-step kinematic expectation according to the action output of the multi-agent near-end strategy optimization network. The technical scheme of the application solves the shortcomings and deficiencies of the prior art, such as the influence of human-driven vehicles (HDVs) on the perception topology of a mixed vehicle fleet, insufficient global cooperation capability, and low efficiency of environment modeling and learning of multi-objective optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation technology, specifically relating to a method, device, system, and storage medium for cooperative formation control of mixed traffic flow. Background Technology

[0002] Hybrid platoon control is a core application in the transition of intelligent transportation systems to fully autonomous driving. It organizes connected vehicles (CAVs) and human-driven vehicles (HDVs) into compact platoons with consistent speeds and spacing, effectively reducing vehicle spacing and aerodynamic drag. This improves road efficiency while reducing fuel consumption and emissions. The most similar existing technologies to this patent mainly follow a path from classical control theory to deep reinforcement learning, and can be summarized into the following three categories: 1. A fleet control method based on classical analytical models and mathematical optimization Early fleet control strategies were mainly divided into state feedback controllers and optimization-based controllers, both of which relied on identifying accurate mathematical models (such as linear or nonlinear vehicle dynamics). State feedback controllers have analytical and closed-form formulas, are easy to implement, and have low computational costs. As a typical representative of optimization controllers, model predictive controllers (MPCs) can predict future vehicle motions and can handle multi-objective problems with state / control constraints.

[0003] The disadvantages of this technical solution are: first, the state feedback controller has difficulty handling multiple explicit control objectives and collision-free constraints at the same time; second, the MPC method is computationally intensive and very time-consuming, which is not conducive to real-time applications.

[0004] 2. A fleet control method based on deep reinforcement learning With the development of artificial intelligence technology, deep reinforcement learning (DRL), as a model-free and learning-based method, is increasingly being applied to fleet control due to its ability to capture stochastic and complex system characteristics. After sufficient offline training, DRL-based controllers can handle complex scenarios under online computational loads.

[0005] The drawback of this technical approach is that current research primarily focuses on single, long convoys under stable and congested traffic conditions, neglecting the complexities of time-varying traffic flow. In time-varying and mixed convoy systems, if connected vehicles (CAVs) attempt to catch up with their preceding vehicles across considerable distances, traditional following strategies may result in excessive speed and acceleration. This can exacerbate fuel consumption and increase the risk of collisions.

[0006] 3. A Hybrid Fleet Control Method Based on Decentralized Reinforcement Learning To address the heterogeneity of mixed traffic flows involving connected vehicles (CAVs) and human-driven vehicles (HDVs), some existing studies classify mixed fleets into multiple sub-fleets and utilize decentralized or centralized control methods to mitigate the uncertainties of HDVs. These methods typically use HDVs lacking communication capabilities as separators between fleets, constructing reinforcement learning environments and implementing cooperative control within truncated local units.

[0007] This technical solution is the closest existing technology to this patent, but it still has the following objective shortcomings and deficiencies: 1. Perception topology is affected by human-driven vehicles (HDVs), resulting in insufficient global coordination capabilities: In mixed traffic, the random and uncertain driving behavior of HDVs easily triggers traffic oscillations, and the amplitude of these oscillations propagates downstream, thus hindering traffic stability. Existing methods for separating convoys based on HDVs lack the ability to generalize to large-scale mixed convoys with a more general intelligent connected vehicle-human-driven vehicle (CAV-HDV) topology. Due to its strict reliance on the preceding vehicle for state transmission, the perception range of intelligent connected vehicles (CAVs) is affected by HDVs, making it impossible to achieve forward-looking information interaction and easily triggering chain-reaction braking.

[0008] 2. Inefficient Environmental Modeling and Learning for Multi-Objective Optimization: The time-varying and mixed traffic flows in the real world greatly complicate the modeling of deep reinforcement learning (DRL) learning environments. Existing decentralized strategies often get stuck in local optima of "selfish learning," lacking an evaluation of the global state of the platoon. In such more general real-world scenarios, the multi-objective optimization platooning problem that aims for both environmental protection and safety remains unsolved.

[0009] In summary, existing technical solutions generally fail to simultaneously coordinate local dynamic topology perception with global environmental protection and safety, resulting in a contradiction between the accuracy of coordinated control and fuel economy in time-varying mixed traffic flows. Summary of the Invention

[0010] To address the problems existing in the prior art, this invention provides a method, device, system, and storage medium for cooperative platooning control of mixed traffic flows, which solves the shortcomings and deficiencies of the prior art, such as the influence of human-driven vehicles (HDVs) on the perception topology of mixed platoons, insufficient global coordination capabilities, and low efficiency in environmental modeling and learning for multi-objective optimization.

[0011] To achieve the above objectives, the present invention provides the following solution: A method for cooperative formation control of hybrid traffic flows includes: Construct a multidimensional state matrix and dynamic spatiotemporal directed graph for hybrid traffic flow; Based on the dynamic spatiotemporal directed graph, the graph adjacency matrix is ​​dynamically evolved using a dual-determination rule. Based on the dynamic evolution of the graph adjacency matrix, the forward look-ahead feature is extracted using a graph attention mechanism; Based on the forward line-of-sight characteristics, a multi-agent near-end policy optimization network based on the CTDE architecture is constructed. The network's action output is optimized based on the multi-agent proximal strategy, and intervention is carried out through a low-level safety supervision mechanism based on one-step kinematic prediction.

[0012] As a preferred approach, in multi-agent proximal policy optimization networks, a multimodal objective reward function that integrates global stability and local security is designed. ,Right now, ,in, This represents the collaboration preference weighting coefficient.

[0013] As a preferred embodiment, the method also includes: acquiring the transfer experience data stream from each data batch and storing it in a shared experience replay pool; limiting the ratio difference between the old and new strategy distributions by replacing the objective function with importance sampling and pruning; and using the backpropagation algorithm to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix. And the gradient partial derivatives of the weights of the Critic fully connected layer.

[0014] It also includes a hybrid traffic flow cooperative formation control device, comprising: The first processing module is used to construct a multi-dimensional state matrix and a dynamic spatiotemporal directed graph of mixed traffic flow. The second processing module is used to generate a dynamic evolution of the graph adjacency matrix based on a dual-determination rule according to the dynamic spatiotemporal directed graph. The third processing module is used to dynamically evolve the graph adjacency matrix and extract the forward look-ahead feature based on the graph attention mechanism. The fourth processing module is used to construct a multi-agent near-end policy optimization network based on the CTDE architecture according to the forward line-of-sight characteristics. The fifth processing module is used to optimize the network's action output based on the multi-agent proximal strategy and to intervene with a low-level safety supervision mechanism based on one-step kinematic prediction.

[0015] As a preferred approach, in multi-agent proximal policy optimization networks, a multimodal objective reward function that integrates global stability and local security is designed. ,Right now, ,in, This represents the collaboration preference weighting coefficient.

[0016] Preferably, the system also includes: a sixth processing module, used to acquire the transfer experience data stream from each data batch and store it in a shared experience replay pool; to limit the ratio difference between the old and new strategy distributions by replacing the objective function with importance sampling and pruning; and to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix using the backpropagation algorithm. And the gradient partial derivatives of the weights of the Critic fully connected layer.

[0017] The present invention also provides a mixed traffic flow cooperative formation control system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a mixed traffic flow cooperative formation control method when executed by the processor.

[0018] The present invention also provides a storage medium storing a computer program that executes a hybrid traffic flow cooperative formation control method during runtime.

[0019] The purpose of this invention is to overcome the limitations of existing technical solutions and achieve the following objectives: First, a global dynamic V2V connectivity graph topology is constructed to replace the traditional local sub-platform segmentation method based on the forced segmentation of human-driven vehicles (HDVs) by their physical locations. The aim is to adaptively capture the dynamic interaction relationships between vehicles in mixed traffic flows by fusing physical radar sensing edges with intelligent connected vehicle (CAV) network communication edges. It not only focuses on the local motion state of vehicles immediately ahead but also obtains advanced warning information from distant CAVs, solving the problems of following delays and downstream propagation of traffic oscillations caused by HDV isolation of CAV platoons.

[0020] Second, a multi-agent and global hybrid reward mechanism based on centralized training and decentralized execution (CTDE) is designed to achieve global optimization of fuel economy and safety stability while ensuring absolute safety for individual vehicles. This aims to address the problem of existing decentralized reinforcement learning methods falling into local optima of "selfish learning." By introducing a global acceleration variance penalty and a global average energy consumption reward covering all vehicles in the entire map (including HDVs with random behavior), the design encourages connected vehicles (CAVs) to adopt a cooperative strategy that actively smooths traffic waves, indirectly guiding and constraining the driving trajectory of HDVs in blind spots.

[0021] Third, an end-to-end collaborative control model is provided, integrating the aforementioned dynamic graph perception structure, efficient multi-agent policy optimization, and underlying kinematic safety supervision mechanism, namely the Graph Multi-Agent Reinforcement Learning (G-MAPPO) architecture. Combined with one-step prediction-based hard-constraint interception, this comprehensively improves the safety, stability, and fuel economy of large-scale hybrid fleets under different penetration rates and time-varying traffic flows, providing reliable technical support for the efficient and smooth management of autonomous driving fleets in future complex mixed traffic environments. Attached Figure Description

[0022] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of the hybrid traffic flow cooperative formation control method according to an embodiment of the present invention; Figure 2 A flowchart for the kinematic prediction and motion interception process of a safety monitor. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] Example 1 Before describing the embodiments in detail, the English abbreviations and industry-standard codes appearing in this specification are explained as follows: CAV (Connected and Automated Vehicle): Intelligent connected vehicles refer to automobiles equipped with advanced onboard sensors, controllers, actuators, and other devices, and integrating modern communication and network technologies to achieve intelligent information exchange and sharing between the vehicle and X (vehicles, roads, people, cloud, etc.).

[0027] HDV (Human-Driven Vehicle): A human-driven vehicle, referring to a traditional vehicle that does not have V2V communication capabilities and is entirely controlled by a human driver based on visual observation.

[0028] V2V (Vehicle-to-Vehicle): Vehicle-to-vehicle communication refers to the direct exchange of data between vehicles via wireless communication technology.

[0029] GAT (GraphAttentionNetwork): A graph neural network based on the spatial domain, which dynamically calculates the feature fusion weights between adjacent nodes in the graph through an attention mechanism.

[0030] MAPPO (Multi-Agent Proximal Policy Optimization): A mainstream reinforcement learning algorithm suitable for multi-agent continuous action control.

[0031] CTDE (Centralized Training with Decentralized Execution): A multi-agent system architecture that utilizes global information during training, while each agent makes independent decisions based solely on local observations during execution.

[0032] TTC (Time-To-Collision): The time required for two vehicles to remain in motion at their current relative distance and speed until a collision occurs.

[0033] like Figure 1 , 2 As shown, the present invention provides a method for cooperative formation control of mixed traffic flow, comprising: Overcoming the shortcomings of existing technologies that rely on static physical isolation rules to cut off fleet topology and the difficulty of achieving system-level fuel optimization through the "selfish learning" of a single agent, this invention reconstructs the dynamic perception topology and cooperative strategy of mixed traffic flow through a fully data-driven approach and centralized multi-agent game theory. The complete execution steps of this invention are as follows: Step S1: Construct the multidimensional state matrix and dynamic spatiotemporal directed graph of the mixed traffic flow. In order for the control model to accurately capture the complex interactions across physical occlusions, the mixed traffic flow within the perception and communication range is first abstracted into a structured dynamic temporal graph. The specific definition is as follows: Represents the time step The monitored set of vehicle nodes includes all connected vehicles (CAVs) and human-driven vehicles (HDVs). Each node... This corresponds to a real physical vehicle. The node feature matrix is ​​defined as follows: The independent state vector of each node is: Specifically, this includes the vehicle's longitudinal position. ,speed acceleration and vehicle type identifier Indicates CAV, (Indicates HDV). The set of edges representing the connectivity of information exchange between vehicles reflects the dynamic topology of mixed traffic flows. It is a time-varying dynamic adjacency matrix constructed based on relative states. The ultimate goal of this step is to break through the traditional limitations of forcibly cutting off convoys based on HDV physical location and to provide the model with graph signal features that include beyond-line-of-sight information.

[0034] Step S2: Dynamic evolution of the graph adjacency matrix based on the dual-determination rule Because the behavior of HDVs in mixed convoys is highly random, and vehicles constantly weave and change lanes, static following topologies are prone to failure. To endow the network with the ability to perceive physical occlusions, this step introduces two independent edge determination rules at the input to dynamically update the adjacency matrix. Physical perception edge generation: Simulates vehicle radar line-of-sight. Extracts nodes with direct physical car-following relationships. If a node... For nodes The vehicle directly in front, and the relative distance Less than the set maximum radar range Then a physical perception edge is generated, that is Otherwise, it is 0.

[0035] V2V communication edge generation: Simulate network communication. Extract the connectivity of nodes with communication capabilities. If nodes... and nodes All are intelligent connected vehicles (CAVs) (i.e., identifiers) and Furthermore, the spatial distance between the two is within the effective communication range. Inside, it ignores the intermediate HDV obstruction in its physical straight path and forcibly establishes a communication connection, that is... Otherwise, it is 0. A logical OR operation is performed on the two types of connected states to obtain the final dynamic adjacency matrix: Based on this, any CAV is determined The dynamic effective neighbor set at the current moment .

[0036] Step S3: Forward-looking feature extraction based on graph attention mechanism (GAT) To address the issue that traditional methods rely solely on the state of the vehicle immediately in front, neglecting the potential threat posed by sudden braking from vehicles further ahead, this step introduces a graph attention network to efficiently remap node features. A latent feature vector is established as an intermediate medium to map discrete relative motion states to a high-dimensional feature space. Its evolutionary process is defined as follows: S3.1 Relative Feature Fusion and Mapping: This involves fusing and mapping the features of the vehicle... with neighbors The state vectors are concatenated and fused to calculate the relative state features: in This is a vector concatenation operation. This is the feature mapping parameter matrix.

[0037] S3.2 Dynamic Threat Attention Allocation: Calculating Neighbor Nodes Using a Single-Layer Feedforward Neural Network For nodes raw attention score Subsequently, the Softmax function is used in the neighbor set. Normalize the input and output attention weights. This weight matrix can measure the degree of threat posed to the vehicle's safety by any neighbor under different motion states.

[0038] S3.3 Neighborhood Message Passing and Hidden State Update: By substituting the weighted feature matrix into the aggregation equation and supplementing it with a nonlinear activation function... The final output captures the high-order hidden state tensor of advanced macroscopic traffic fluctuations: in This is the characteristic transformation matrix.

[0039] Step S4: Construct a Multi-Agent Proximal Policy Optimization (MAPPO) network based on the CTDE architecture Traditional decentralized reinforcement learning is prone to falling into selfish game theory. This step uses a centralized training with decentralized execution (CTDE) framework to reconstruct action mapping.

[0040] S4.1 Decentralized Actor Branching: Each CAV Configure and run the Actor policy network independently The individual high-order graph features output from step S3 are fed into a fully connected neural network layer. After nonlinear activation, the mean and variance of the Gaussian distribution in the continuous action space are output, and the final pre-selected value of longitudinal acceleration is obtained by sampling. (limited to) Within the interval).

[0041] S4.2 Centralized Critic Branch Evaluation: During the training phase, a unified Critic network is constructed. The concatenated vector containing features of all nodes in the entire graph (including HDV) will be generated. The data is fed back into the Critic network to calculate the global state value baseline for the current traffic flow. Subsequently, the advantage function is obtained by calculating the generalized advantage estimate (GAE). The parameters that guide all Actor networks Implement a strategy gradient improvement to break the limitations of single-vehicle vision.

[0042] Step S5: Construct a multimodal objective reward function that integrates global stability and local security. To guide model parameter updates and cultivate the "altruism" of intelligent connected vehicles (CAVs) in actively smoothing traffic waves, this step designs a joint reward function that combines local individual rewards with global shared rewards. ( (This refers to the collaboration preference weighting coefficient).

[0043] S5.1 Local Reward: Measures the vehicle's collision avoidance and following distance error. Introduces the minimum deceleration required to avoid a collision. If the estimated required deceleration is greater than the rated comfort deceleration. This triggers a dynamic penalty mechanism based on the tangent hyperbolic function; if a collision occurs (spacing) An extreme value penalty is then applied. This, combined with a quadratic penalty based on the expected distance and speed error, constitutes a local reward. .

[0044] S5.2 Global System Coordination and Ecological Reward: This measure the energy consumption and oscillations of the overall hybrid vehicle fleet. A global ecological reward term is constructed: based on the VT-Micro macroscopic energy model, the instantaneous speed and acceleration of all vehicles within the entire map are extracted to calculate the fuel consumption rate, and the summation is performed with a negative exponential mapping. Simultaneously, a global string stability penalty term is constructed: the L2 norm decay gradient of acceleration for all vehicles from start to finish is calculated; if the absolute acceleration of a following vehicle is greater than that of the preceding vehicle (i.e., a rearward amplification effect of oscillation occurs), a severe penalty is applied. The two terms combined constitute the global reward. .

[0045] Step S6: Intervention based on a one-step kinematic prediction-based underlying safety supervisor mechanism (Safety Supervisor) After completing the network action output, given the uncertainty of the output in the early stages of reinforcement learning exploration, a one-step safety intervention module based on a physics model is connected in series to ensure absolute system safety. First, the longitudinal acceleration output of the Actor network is received. Based on the model of uniformly accelerated linear motion Predict the coordinates of the vehicle at the next time step, simultaneously predict the position of the vehicle directly ahead, and calculate the prediction distance. Subsequently, a minimum absolute safety distance was introduced. With minimum safe headway threshold If the prediction result satisfies or If this occurs, a safety boundary violation alarm is triggered, the system strips the reinforcement learning network of control, and forcibly overwrites the actual acceleration with the calibrated maximum safe deceleration. If not triggered, the original output action will proceed.

[0046] Step S7: Evolution of Multi-Agent Joint Policy and Parameter Optimization Obtain the transfer experience data stream from each data batch. The data is stored in a shared experience replay pool. Importance sampling and clipped surrogate objective function are used to limit the ratio difference between the old and new policy distributions, preventing policy collapse caused by excessively large single-step update magnitudes. Backpropagation is used to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix. And the gradient partial derivatives of the weights in the Critic fully connected layer. This invention uses the Adam optimization algorithm to dynamically update the network parameters. In a preferred embodiment, the learning rate is set to... Furthermore, a Curriculum Learning mechanism based on the doubling of fleet size is introduced to improve convergence efficiency. The mean square error of the joint advantage estimation is continuously reduced until the model achieves zero collision and minimizes traffic wave amplitude in simulation tests covering extreme stop-and-go conditions. The optimal model weight parameters are then fixed and deployed to the real vehicle cooperative controller.

[0047] Compared with the prior art, the present invention has the following significant advantages and beneficial effects: 1. Prevent convoys from becoming isolated due to HDV separation, and achieve dynamic topology perception with forward line of sight. The shortcomings of existing technologies: Current hybrid vehicle fleet control methods typically utilize human-driven vehicles (HDVs) lacking communication capabilities to forcibly break down the fleet into multiple isolated sub-platoons for control. This "physical isolation" severs information exchange between connected vehicles (CAVs) across HDVs. Because connected vehicles (CAVs) can only passively respond to their immediate neighbors, they lose the ability to anticipate downstream traffic waves. When faced with the random and unpredictable driving behavior of HDVs, they are highly susceptible to triggering traffic oscillations and amplifying them downstream.

[0048] Advantages of this invention: This solution innovatively designs a module for generating an adaptive dynamic V2V connected graph topology. By logically fusing the physical perception edges limited by radar with the V2V communication edges that ignore physical obstructions from HDVs, a time-varying directed graph structure is constructed. Combined with a graph attention network (GAT) to dynamically calculate the threat weights of neighboring vehicles, this structure is entirely data-driven, allowing CAVs to "see" and respond to warning information from other connected vehicles (CAVs) ahead, even when separated by multiple human-driven vehicles (HDVs), thus completely solving the reaction delay problem caused by line-of-sight and communication obstructions.

[0049] 2. Overcome the local optima of "selfish learning" to achieve system-level environmental safety and overall stability. The shortcomings of existing technologies: When dealing with time-varying and hybrid convoy systems, traditional decentralized deep reinforcement learning (DRL) algorithms primarily focus on optimizing the local following distance and safety of individual CAVs. This "selfish learning," lacking global collaboration, often leads to a situation where the acceleration behavior of a single connected vehicle (CAV) for its own smoothness causes significant traffic oscillations in the human-driven vehicle (HDV) group behind it, failing to achieve multi-objective global optimization that balances environmental protection and safety in more general real-world scenarios.

[0050] Advantages of this invention: This scheme introduces a multi-agent proximal policy optimization (MAPPO) architecture with centralized training and decentralized execution (CTDE) at the decision control layer. It innovatively designs a hybrid multi-objective reward function that includes global string stability penalty (global acceleration variance) and global ecological benefits (global average fuel consumption based on the VT-Micro model). This mechanism endows connected vehicles (CAVs) with "altruistic" characteristics, prompting them to act as mobile traffic stabilizers, actively absorbing speed fluctuations ahead, thereby indirectly smoothing the following trajectory of human-driven vehicles (HDVs) behind, minimizing the overall energy consumption and oscillations of large-scale convoys.

[0051] 3. Ensures absolute bottom-layer security and exhibits high topological robustness in the face of extremely low penetration rates. The limitations of existing technologies are that random actions in the early stages of reinforcement learning can easily lead to collisions and disrupt the training process. Even the optimal actions output by a well-trained network cannot guarantee 100% collision avoidance, especially in large-scale mixed fleet scenarios involving a large number of uncertain human-driven vehicle (HDV) behaviors. Furthermore, statically segmented fleet topologies become severely fragmented with extremely low CAV penetration rates, leading to a precipitous drop in cooperative performance.

[0052] Advantages of this invention: This solution incorporates a OneStep Safety Supervisor at the network output. By predicting potential collisions based on kinematic rules and forcibly replacing dangerous actions with minimum safe acceleration, it completely eliminates the trial-and-error costs in the early stages of DRL training, ensuring absolute zero collisions under stringent stop-and-go conditions. Simultaneously, the highly flexible dynamic graph topology can adapt to frequent incursions by human-driven vehicles (HDVs), seamlessly reconstructing communication links even with low CAV penetration rates, maintaining the high robustness of the control strategy.

[0053] The hybrid traffic flow cooperative platooning control method of the present invention is physically based on a CAV equipped with specific electrical and control components. In a specific embodiment, the hardware system of a single CAV mainly includes: ①Environmental perception module: including millimeter-wave radar and vehicle-mounted camera, used to collect static and dynamic physical data (position, speed, acceleration) of itself and vehicles in front.

[0054] ②V2V wireless communication module: Utilizing DSRC or C-V2X communication terminals, responsible for transmitting data to the communication radius (e.g., ...). Other CAVs within the vehicle broadcast their own vehicle status and receive node data from other CAVs. The onboard central computing unit includes a CPU and an AI acceleration chip (such as a GPU / NPU) dedicated to neural network inference. It is electrically connected to the perception module and communication module, and is used to run the G-MAPPO control algorithm and safety monitor of this invention.

[0055] ③ Drive-by-wire chassis actuation module: This includes the electromechanical braking system (EMB) and the drive-by-wire system. The output of the onboard central computing unit is electrically connected to the actuation module via the onboard CAN / Ethernet bus, transmitting the calculated target acceleration. This is converted into actual braking pressure or motor torque.

[0056] The specific implementation steps of the hybrid traffic flow cooperative formation control method of the present invention are as follows: During vehicle operation, the onboard central computing unit operates at a preset time step (e.g., ...). Repeat the following steps: Step S1: Data Acquisition for Topology Generation of Dynamic Spatiotemporal Directed Graph: At time step CAV Using sensors to detect vehicles immediately ahead state vector Simultaneously, it receives state vectors from other CAVs within the communication range via the V2V module. Dynamic adjacency matrix construction: Calculate the radar's physical sensing edge: If the preceding vehicle... With this vehicle distance (Set to 50m), then the physical sensing edge ; Calculate V2V communication edges: If node Send communication handshake and distance (Set to 150m), and identification. Then, regardless of whether there is physical occlusion of HDV between the two, the communication edge is directly constructed. Matrix fusion: Performs a logical OR operation. Determine the dynamic effective neighbor set of this vehicle. .

[0057] Step S2: Graph Attention Feature Extraction (Ahead-of-Look-Ahead Perception) The current vehicle's state is concatenated with the states of each vehicle in the neighboring set. This concatenation is then fed into the pre-trained and fixed GAT network layer. Internally, the network uses formulas... Calculate the original threat score. Obtain the attention weights after normalization using the Softmax function. For example, when the distant CAV When a sudden braking event causes a sharp increase in the relative speed difference, the GAT network automatically assigns that node a very high weight. The output aggregates high-order hidden state features that anticipate macroscopic traffic fluctuations. .

[0058] Step S3: Multi-agent Actor network decision computation incorporates hidden state features The input is fed into a fully connected Actor neural network. After forward propagation, the network outputs the mean and variance parameters of a Gaussian distribution, from which the pre-selected value of the longitudinal acceleration that the vehicle is expected to execute at the next moment is sampled. (Limited to) ).

[0059] Step S4: One-step kinematic safety supervision intervention (Safety Supervisor) Input the underlying security monitor. Based on the physical model of uniformly accelerated linear motion, predict the time step. The position of this car And similarly predict the position of the vehicle immediately in front. Calculate and predict the headway. Set the minimum absolute safety distance. Minimum safe headway Action replacement judgment: If or This indicates that the output of the reinforcement learning network has a collision risk. The system immediately intercepts the action. And force it to be replaced with the maximum safe deceleration. Otherwise, retain the original action. Send the final, safety-verified acceleration command to the drive-by-wire chassis execution module.

[0060] Experimental Data and Comparative Data (Performance Verification): To verify the effectiveness of the proposed method (G-MAPPO), it was deployed on a reinforcement learning traffic simulation platform jointly built with SUMO (Simulation of Urban Mobility) and Python. Experimental Scenario: A large-scale single-lane mixed convoy was constructed, consisting of one lead vehicle (set to perform extreme "acceleration-sudden braking-idle" behavior to inject traffic oscillations) and 16 following vehicles. Heterogeneous Injection: Vehicle trajectories from a real-world traffic flow public dataset (NGSIM dataset) were injected into HDV nodes in the simulation environment to realistically reproduce the random disturbances and slow reaction characteristics of human drivers. CAV penetration was set to 50%. Comparative Baseline: SFC-CTH (Comparison Method 1): A traditional state feedback control algorithm based on constant headway; DPPO (Comparison Method 2): An existing decentralized reinforcement learning algorithm based on physical segmentation (communication is cut off upon encountering HDVs), as shown in Table 1.

[0061] Table 1 Data performance analysis: Fuel economy and traffic efficiency: Compared with Comparative Example 2 (DPPO), the method of this invention reduces the average fuel consumption rate by approximately 8.9%, while increasing the average vehicle speed by approximately 14.4%. This is because the dynamic V2V connectivity graph constructed by this invention breaks down physical obstructions, and the CAV utilizes forward visibility to predict and smoothly accelerate and decelerate, avoiding sudden braking and ineffective idling caused by obstructed vision, thereby significantly reducing energy consumption.

[0062] Traffic oscillation suppression (string stability): Under extreme start-stop disturbances applied by the lead vehicle, the damping ratio of the traditional SFC-CTH is as high as 0.91, which is almost unable to suppress the propagation of oscillations downstream; the DPPO has a damping ratio of 0.78 due to the lack of global coordination; while the present invention relies on a centralized Critic network and a global string stability reward function to give the CAV the characteristics of "altruism", actively absorbing disturbances ahead and reducing the damping ratio to 0.51, showing an extremely excellent traffic flow stabilization and control effect.

[0063] In summary, the method proposed in this invention significantly surpasses existing technologies in terms of system-level fuel economy, traffic efficiency, and absolute safety.

[0064] Example 2 The present invention also provides a hybrid traffic flow cooperative formation control device, comprising: The first processing module is used to construct a multi-dimensional state matrix and a dynamic spatiotemporal directed graph of mixed traffic flow. The second processing module is used to generate a dynamic evolution of the graph adjacency matrix based on a dual-determination rule according to the dynamic spatiotemporal directed graph. The third processing module is used to dynamically evolve the graph adjacency matrix and extract the forward look-ahead feature based on the graph attention mechanism. The fourth processing module is used to construct a multi-agent near-end policy optimization network based on the CTDE architecture according to the forward line-of-sight characteristics. The fifth processing module is used to optimize the network's action output based on the multi-agent proximal strategy and to intervene with a low-level safety supervision mechanism based on one-step kinematic prediction.

[0065] As one embodiment of the present invention, in a multi-agent proximal policy optimization network, a multimodal objective reward function that integrates global stability and local security is designed. ,Right now, ,in, This represents the collaboration preference weighting coefficient.

[0066] As one embodiment of the present invention, it further includes: a sixth processing module, used to acquire the transfer experience data stream in each data batch and store it in a shared experience replay pool; to limit the ratio difference between the distribution of the new and old strategies by replacing the objective function with importance sampling and pruning; and to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix using the backpropagation algorithm. And the gradient partial derivatives of the weights of the Critic fully connected layer.

[0067] Example 3 The present invention also provides a mixed traffic flow cooperative formation control system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a mixed traffic flow cooperative formation control method when executed by the processor.

[0068] Example 4 The present invention also provides a storage medium storing a computer program that executes a hybrid traffic flow cooperative formation control method during runtime.

[0069] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for cooperative formation control of hybrid traffic flow, characterized in that, include: Construct a multidimensional state matrix and dynamic spatiotemporal directed graph for hybrid traffic flow; Based on the dynamic spatiotemporal directed graph, the graph adjacency matrix is ​​dynamically evolved using a dual-determination rule. Based on the dynamic evolution of the graph adjacency matrix, the forward look-ahead feature is extracted using a graph attention mechanism; Based on the forward line-of-sight characteristics, a multi-agent near-end policy optimization network based on the CTDE architecture is constructed. The network's action output is optimized based on the multi-agent proximal strategy, and intervention is carried out through a low-level safety supervision mechanism based on one-step kinematic prediction.

2. The hybrid traffic flow cooperative formation control method as described in claim 1, characterized in that, In multi-agent proximal policy optimization networks, a multimodal objective reward function that integrates global stability and local security is designed. ,Right now, ,in, This represents the collaboration preference weighting coefficient.

3. The hybrid traffic flow cooperative formation control method as described in claim 2, characterized in that, Also includes: Acquire the transfer experience data stream from each data batch and store it in the shared experience replay pool; By replacing the objective function with importance sampling and pruning, the ratio difference between the old and new strategies can be limited; The backpropagation algorithm is used to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix. And the gradient partial derivatives of the weights of the Critic fully connected layer.

4. A hybrid traffic flow cooperative formation control device, characterized in that, include: The first processing module is used to construct a multi-dimensional state matrix and a dynamic spatiotemporal directed graph of mixed traffic flow. The second processing module is used to generate a dynamic evolution of the graph adjacency matrix based on a dual-determination rule according to the dynamic spatiotemporal directed graph. The third processing module is used to dynamically evolve the graph adjacency matrix and extract the forward look-ahead feature based on the graph attention mechanism. The fourth processing module is used to construct a multi-agent near-end policy optimization network based on the CTDE architecture according to the forward line-of-sight characteristics. The fifth processing module is used to optimize the network's action output based on the multi-agent proximal strategy and to intervene with a low-level safety supervision mechanism based on one-step kinematic prediction.

5. The hybrid traffic flow cooperative formation control device as described in claim 4, characterized in that, In multi-agent proximal policy optimization networks, a multimodal objective reward function that integrates global stability and local security is designed. ,Right now, ,in, This represents the collaboration preference weighting coefficient.

6. The hybrid traffic flow cooperative formation control device as described in claim 5, characterized in that, Also includes: The sixth processing module is used to acquire the transfer experience data stream from each data batch and store it in the shared experience playback pool. The objective function is replaced by importance sampling and pruning to limit the ratio difference between the old and new strategy distributions; the backpropagation algorithm is used to calculate the projection parameters of the objective function onto the Actor network and the GAT weight matrix. And the gradient partial derivatives of the weights of the Critic fully connected layer.

7. A hybrid traffic flow cooperative formation control system, characterized in that, include: A memory and a processor, wherein the memory stores a computer program executed by the processor, the computer program performing the hybrid traffic flow cooperative formation control method as described in any one of claims 1-3 when executed by the processor.

8. A storage medium, characterized in that, The storage medium stores a computer program that, when executed, performs the hybrid traffic flow cooperative formation control method as described in any one of claims 1-3.