Collaborative decision-making method and device for multiple automatic driving vehicles

By combining Transformer and GCN to generate a scene graph structure, DQN is used to optimize the collaborative decision-making of multi-autonomous vehicles, the problem of low efficiency and poor reliability of multi-vehicle collaborative decision-making in complex traffic environments in the existing technology is solved, and efficient and safe collaborative decision-making of multi-vehicles is achieved.

CN120540306APending Publication Date: 2025-08-26SHANDONG TOP ELECTRONIC TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510655135.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

It is difficult for existing autonomous driving systems to make effective multi-vehicle collaborative decisions in complex traffic environments. Traditional methods rely on manual engineering rules and deep learning methods have problems such as insufficient scenario representation and vehicle interaction modeling, resulting in low decision efficiency and poor reliability.

Method used

Using a method combining the first neural network and the second neural network, efficient representation and precise modeling of traffic scenes are performed, scene graph structures are generated through Transformer and GCN, and the Markov decision-making process is solved using the deep reinforcement learning algorithm DQN to optimize the collaborative decision-making process of multiple autonomous vehicles.

Benefits of technology

The coordinated decision-making ability of autonomous driving systems in complex traffic environments has been improved, and efficient and safe multi-vehicle collaborative decision-making has been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540306A_ABST
    Figure CN120540306A_ABST
Patent Text Reader

Abstract

The invention provides a collaborative decision-making method and device for multiple automatic driving vehicles, and the method comprises the steps: obtaining and analyzing a traffic scene, so as to obtain the scene features of the traffic scene; based on a first neural network and the scene feature, encoding the scene feature to obtain scene feature encoding information; generating a scene graph structure according to the motion information of the plurality of vehicles and the interaction relationship among the plurality of vehicles; based on a second neural network and the scene graph structure, determining spatial interaction information among the plurality of vehicles, the second neural network being a graph neural network GCN; and determining a cooperative driving strategy among the plurality of vehicles according to the scene feature coding information and the space interaction information. According to the scheme, the collaborative decision-making process of the multiple automatic driving automobiles is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving, and in particular to a method, device, electronic device, and computer-readable storage medium for collaborative decision-making among multiple autonomous driving vehicles. Background Art

[0002] Currently, research on multi-vehicle collaborative decision-making (MCDM) primarily focuses on leveraging mutual information and interactions between vehicles to improve the performance of autonomous driving systems. For example, existing approaches primarily rely on hand-engineered heuristic rules, such as finite state machines. These approaches are typically designed for a specific set of use cases and require extremely tedious manual maintenance of the rule database to ensure safety. For another example, existing deep learning-based CDM approaches, while achieving breakthroughs in some areas, still suffer from low decision-making efficiency and poor reliability due to deficiencies in scenario representation and vehicle interaction modeling.

[0003] Therefore, in order to further improve the collaborative decision-making capabilities of autonomous driving systems in complex traffic environments, it is necessary to solve problems such as interaction modeling between vehicles in dynamic environments, scenario representation, and decision optimization. Summary of the Invention

[0004] In view of this, the present application provides a method, apparatus, device and computer-readable storage medium for collaborative decision-making of multiple autonomous vehicles, which optimizes the collaborative decision-making process of multiple autonomous vehicles by efficiently representing and accurately modeling complex traffic scenarios.

[0005] The present application is introduced below from multiple aspects, and the implementation methods and beneficial effects of the following multiple aspects can be referenced to each other.

[0006] In a first aspect, the present application provides a collaborative decision-making method for multiple autonomous driving vehicles, including: acquiring and analyzing a traffic scene to obtain scene features of the traffic scene, the scene features including multiple vehicles in the traffic scene and a grid map corresponding to each of the multiple vehicles, the grid map including a grid occupied by the vehicle and surrounding vehicles constructed with the vehicle as the center; encoding the scene features based on a first neural network and the scene features to obtain scene feature encoding information; generating a scene graph structure based on the motion information of the multiple vehicles and the interaction relationship between the multiple vehicles; determining the spatial interaction information between the multiple vehicles based on a second neural network and the scene graph structure, the second neural network being a graph neural network GCN; determining the collaborative driving strategy between the multiple vehicles based on the scene feature encoding information and the spatial interaction information.

[0007] According to an embodiment of the present application, by combining the above-mentioned first neural network and second neural network, complex traffic scenarios are efficiently represented and accurately modeled, thereby optimizing the collaborative decision-making process of multiple autonomous vehicles.

[0008] In a possible implementation of the first aspect above, determining the collaborative driving strategy between the multiple vehicles includes: modeling the collaborative decision-making problem between the multiple vehicles as a Markov decision process MDP, wherein the state space of the MDP includes the scene feature encoding information and the motion information of the multiple vehicles, and the action space of the MDP includes the control actions taken by the multiple vehicles; and using a deep reinforcement learning algorithm DQN to solve the MDP to obtain the collaborative driving strategy.

[0009] In a possible implementation of the first aspect above, the reward function of the MDP is determined based on a speed reward, a collision penalty, and a mission intention reward, wherein the speed reward is the ratio of the vehicle's current speed to its maximum speed, the collision penalty is the negative value of the number of collisions, and the mission intention reward is used to instruct the vehicle to take correct actions when approaching the target position.

[0010] In a possible implementation of the first aspect above, the use of the deep reinforcement learning algorithm DQN to solve the MDP includes: initializing the Q network to estimate the Q value of the state space and action space pair; updating the Q network parameters through the Bellman equation to maximize the cumulative reward; using the experience replay cache to store the state, action, reward and next state; and selecting the optimal action based on the updated Q network to generate a collaborative driving decision.

[0011] In a possible implementation of the first aspect above, encoding the scene feature includes: splicing the raster map into a scene representation matrix by row; inputting the scene representation matrix into the embedding layer of the first neural network to obtain an embedding vector; and performing block processing on the embedding vector based on the first neural network to obtain the scene feature encoding information.

[0012] In a possible implementation of the first aspect above, the first neural network is a Transformer neural network.

[0013] In a possible implementation of the first aspect above, determining the spatial interaction information between the multiple vehicles includes: obtaining a feature matrix of the multiple vehicles based on motion information of the multiple vehicles, the motion information including the longitudinal position, speed, lane, category, front vehicle distance and rear vehicle distance of the multiple vehicles; obtaining an adjacency matrix of the multiple vehicles based on the interaction relationship between the multiple vehicles, the interaction relationship including communication and interaction between autonomous driving vehicles among the multiple vehicles and interaction between each autonomous driving vehicle and a manned vehicle within the perception range; obtaining a mask matrix, the mask matrix being used to filter the features of the autonomous driving vehicle and remove information of manned vehicles; and determining the spatial interaction information between the multiple vehicles based on the scene graph structure, the feature matrix, the adjacency matrix, the mask matrix and the graph neural network.

[0014] In a second aspect, the present application provides a device for collaborative decision-making of multiple autonomous driving vehicles, including: an acquisition unit, used to: acquire and analyze an input traffic scene to obtain scene features of the traffic scene, the scene features including multiple vehicles in the traffic scene and a grid map corresponding to each of the multiple vehicles, the grid map including a grid occupied by the vehicle and surrounding vehicles constructed with the vehicle as the center; a processing unit, used to: encode the scene features based on a first neural network and the scene features to obtain scene feature encoding information; generate a scene graph structure based on the motion information of the multiple vehicles and the interaction relationship between the multiple vehicles; determine the spatial interaction information between the multiple vehicles based on a second neural network and the scene graph structure, the second neural network being a graph neural network GCN; determine the collaborative driving strategy between the multiple vehicles based on the scene feature encoding information and the spatial interaction information.

[0015] In a third aspect, the present application provides a device for collaborative decision-making of multiple autonomous vehicles, the device comprising: a memory for storing instructions executed by one or more processors of the device, and a processor, which is one of the processors of the device, for executing the driving behavior classification method based on multimodal fusion disclosed in any aspect of the first aspect above.

[0016] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the driving behavior classification method based on multimodal fusion according to the first aspect.

[0017] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is caused to perform the steps of multimodal fusion-based driving behavior classification according to the second aspect. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer is caused to perform the steps of multimodal fusion-based driving behavior classification according to the second aspect.

[0018] The possible implementation methods and technical effects obtained in the above-mentioned second to fifth aspects are similar to the corresponding technical means and technical effects obtained in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a schematic diagram of the process of a method for collaborative decision-making of multiple autonomous driving vehicles provided in an embodiment of the present application;

[0020] Figure 2 This is a flow chart of another method for collaborative decision-making among multiple autonomous driving vehicles provided in an embodiment of the present application;

[0021] Figure 3 This is a schematic diagram of a scene feature representation provided by an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a scene graph structure provided by an embodiment of the present application;

[0023] Figure 5 This is a schematic diagram of a graph neural network computing framework provided in an embodiment of the present application;

[0024] Figure 6 Schematic diagram of a deep Q-network (DQN) algorithm for solving a Markov decision process (MDP) provided in an embodiment of the present application;

[0025] Figure 7 It is a structural schematic diagram of a device provided in an embodiment of the present application;

[0026] Figure 8 It is a structural diagram of a system on chip (SoC) provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] First, the prior art and existing technical problems involved in this application are introduced.

[0029] With the rapid development of autonomous driving technology, self-driving cars have gradually entered the practical application stage. However, existing autonomous driving systems rely primarily on the perception and decision-making of a single vehicle for path planning, and are unable to effectively address the needs of multi-vehicle collaborative decision-making in complex traffic environments. This is especially true in mixed traffic environments, where autonomous vehicles must engage in complex interactions with conventional vehicles, pedestrians, and other traffic participants. Traditional single-vehicle decision-making models cannot meet the requirements of multi-vehicle collaboration, resulting in poor coordination between vehicles and difficulty in accurately predicting and responding to complex interactions.

[0030] At present, the research on multi-vehicle collaborative decision-making mainly focuses on how to utilize the mutual information and interactive behaviors between vehicles to improve the performance of autonomous driving systems. However, traditional planning methods are mainly based on manually engineered heuristic rules, such as finite state machines. These methods are usually designed only for a specific set of use cases and require extremely tedious manpower to maintain the rule database to ensure safety. In addition, as the number of rules increases, new problems may arise. For example, old scenarios will be forgotten when solving new scenarios, and it is difficult to balance the cost function in countless scenarios that are difficult to model and have conflicting goals (such as safety and efficiency). It is difficult to handle complex traffic scenarios and changing environmental factors. In addition, due to the shortcomings of scene representation and vehicle interaction modeling, existing collaborative decision-making methods based on deep learning still have problems such as low decision-making efficiency and poor reliability.

[0031] Therefore, in order to further improve the collaborative decision-making capabilities of autonomous driving systems in complex traffic environments, it is necessary to solve problems such as interaction modeling between vehicles in dynamic environments, scenario representation, and decision optimization.

[0032] In order to solve the above technical problems, the present application proposes a method 100 for collaborative decision-making of multiple autonomous driving vehicles. Figure 1 1 is a flow chart of a method 100 for collaborative decision-making of multiple autonomous driving vehicles, as shown in FIG. Figure 1 As shown, the method includes steps 110 to 150. The method 100 optimizes the collaborative decision-making process of multiple autonomous vehicles by efficiently representing and accurately modeling complex traffic scenarios, thereby solving the problem in the prior art that multiple autonomous vehicles have difficulty making safe and efficient decisions in complex traffic environments.

[0033] Step 110: Acquire and analyze the traffic scene to obtain scene features of the traffic scene.

[0034] Specifically, the scene feature includes multiple vehicles in the traffic scene and a grid map corresponding to each of the multiple vehicles. The grid map includes a grid constructed with the vehicle as the center and occupied by the vehicle and surrounding vehicles.

[0035] For example, in an embodiment of the present application, after acquiring the traffic scene, the dynamic driving environment features in the traffic scene are extracted, and the traffic scene is rasterized. For example, in the present application, assuming that each autonomous driving vehicle can perceive the traffic environment within a few meters in front and behind the vehicle, a local traffic scene can be constructed with the autonomous driving vehicle as the central coordinate, thereby obtaining a local map centered on the autonomous driving vehicle. Furthermore, based on the lane in which the vehicle is located and the relative positions and lanes of other vehicles within the perception range, the grids occupied by these vehicles in the grid map are marked, and the unoccupied grids are marked as 0. In this way, the traffic scene can be represented more accurately. The following embodiment section will introduce this step 110 in detail.

[0036] Step 120: Encode the scene feature based on the first neural network and the scene feature to obtain scene feature encoding information.

[0037] Exemplarily, the first neural network can be a neural network such as a transformer, which can encode the scene features obtained in step 110 to capture key information about the scene. For example, in an embodiment of the present application, after splicing the scene features into a scene feature representation matrix, the scene feature representation matrix can be encoded using an embedding layer of the first neural network, such as by passing the scene feature representation matrix through the embedding layer to obtain an output embedding vector. The following embodiment section will describe the encoding process of the scene feature representation matrix using a transformer neural network.

[0038] Step 130: Generate a scene graph structure based on the motion information of the multiple vehicles and the interaction relationship between the multiple vehicles.

[0039] Specifically, dynamic traffic scenes also exhibit spatial interactions, namely the spatial distribution and relationships of vehicles and their behaviors. These factors collectively influence the dynamic changes in traffic flow. To effectively represent this characteristic, in step 130, after extracting vehicle motion information and the interactions between vehicles, the present application constructs the dynamic traffic scene into a graph structure, namely, the scene graph structure in step 130.

[0040] Among them, the scene graph structure includes multiple nodes and connection relationships between the multiple nodes, each of the multiple nodes is used to indicate a single vehicle and its movement information, and the connection relationship is used to indicate the interaction relationship between the two vehicles corresponding to the nodes at both ends of the connection relationship.

[0041] Optionally, the vehicle's motion information may include the vehicle's longitudinal position, speed, lane, category, distance between vehicles in front and behind, etc. The interaction relationship between vehicles may include communication and interaction between autonomous driving vehicles, and interaction between autonomous driving vehicles and manned vehicles within the perception range.

[0042] Step 140: Determine spatial interaction information between the multiple vehicles based on the second neural network and the scene graph structure.

[0043] Specifically, the second neural network is a graph neural network (GNN) such as a graph convolutional network (GCN), ChebNet, etc. In step 140, a feature matrix related to vehicle motion information, an adjacency matrix related to interaction relationships, and a mask matrix that makes the output dimension of the second neural network consistent with that of the first neural network can be extracted based on the scene graph structure. Then, based on the scene graph structure and the multiple matrices extracted above, the second neural network is used to model the scene graph structure to obtain spatial interaction information between the multiple vehicles, and the spatial interaction information is used to indicate the interaction characteristics between the multiple vehicles in space.

[0044] Step 150: Determine a cooperative driving strategy among the multiple vehicles based on the scene feature coding information and the spatial interaction information.

[0045] For example, in an embodiment of the present application, the collaborative decision-making problem can be modeled as a Markov decision process (MDP) and solved using a deep reinforcement learning algorithm (deep Q-network, DQN) to output a collaborative driving strategy. The scene feature encoding information and the spatial interaction information can be used to continuously optimize the collaborative driving strategy during the solution process. In addition, the embodiment of the present application can also design reward functions related to traffic efficiency, task completion, and safety to significantly improve the collaborative decision-making capabilities of multiple autonomous vehicles in a mixed traffic environment. The following embodiments will introduce step 150 in detail, and this application will not go into details here.

[0046] Method 100 optimizes the collaborative decision-making process of multiple autonomous vehicles by combining the above-mentioned first neural network (such as transformer) and the second neural network (such as GCN) to efficiently represent and accurately model complex traffic scenarios.

[0047] The following combination Figures 2 to 6 An example of an embodiment corresponding to method 100 is introduced. Figure 2 1 is a flow chart of another method 200 for collaborative decision-making of multiple autonomous vehicles provided in an embodiment of the present application, that is, a schematic diagram of an example method of the method 100; Figure 3 is a schematic diagram of a scene feature representation provided by an embodiment of the present application, i.e., an example diagram of step 110; Figure 4is a schematic diagram of a scene graph structure provided in an embodiment of the present application, i.e., an example diagram of step 130; Figure 5 is a schematic diagram of a graph neural network computing framework provided in an embodiment of the present application, i.e., an example diagram of step 140; Figure 6 This is a schematic diagram of a deep Q-network (DQN) algorithm for solving a Markov decision process (MDP) provided in an embodiment of the present application, that is, an example diagram of step 150.

[0048] First combine Figure 2 Method 200 includes steps 210 to 240. This method extracts dynamic driving environment features and encodes scene information through a Transformer module, capturing key information using a multi-head attention mechanism. The traffic scene is then modeled as a graph structure, and spatial interaction features are extracted using a GNN. Finally, the collaborative decision-making problem is modeled as an MDP, solved using a DQN, and outputting a collaborative driving policy. This enables efficient and safe collaborative decision-making among multiple autonomous vehicles in complex traffic environments.

[0049] Step 210: Encode the scene feature based on the transformer and the scene feature to capture key information in the scene feature.

[0050] Corresponding to steps 110 and 120, in step 210, Figure 3 The input traffic scene is rasterized to extract scene features. In the embodiment of this application, it is assumed that each autonomous driving vehicle can perceive the traffic environment within 50 meters in front and behind the vehicle, and a local traffic scene is constructed with each autonomous driving vehicle as the center coordinate, thereby obtaining a local map of the i-th autonomous driving vehicle. Afterwards, the grid occupied by each autonomous vehicle is marked as map according to the lane j where the vehicle is located. i [j,centric], and according to the relative position and lane information of other vehicles within the perception range, the grids occupied by other vehicles are also marked accordingly, such as Figure 3 The dark grid shown in the scene feature extraction section. All unoccupied local grids are marked as 0, as shown in Figure 3 The white grid shown in the scene feature extraction section.

[0051] Afterwards, the embodiment of the present application splices the grid map obtained by the i-th autonomous driving vehicle row by row to obtain the scene representation matrix Then, based on all m autonomous driving vehicles in the scene, the model input SR = {SR1; SR2; ...; SR m}.

[0052] Furthermore, each SR i The model input SR is encoded into an embedding vector of the same length that can be further processed by the transformer using the transformer's embedding layer. The model, consisting of L layers of transformer blocks, alternates between multi-head attention (MHA) and multilayer perceptron (MLP). A normalization (LN) layer is added before each transformer block, and after each MHA and MLP, the features of the previous layer are added to the output to merge and preserve the original features. The overall calculation process is as follows:

[0053]

[0054] Among them, X0 represents the input embedding vector, which is obtained by passing the scene representation matrix through the embedding layer; X′ l represents the MHA output of layer l; X l represents the output of the lth layer after LN and MLP processing; y represents the final output after processing by all layers.

[0055] Step 220: Spatial interaction information between the multiple vehicles based on the graph neural network and scene graph structure.

[0056] Corresponding to steps 130 and 140, after extracting the motion information of the vehicle, the embodiment of the present application constructs the dynamic traffic scene as follows: Figure 4 The graph structure shown in Figure 2. Figure 4 Each node in represents a single vehicle and includes the vehicle's motion information. The connection relationship between each node is used to indicate the interaction relationship between vehicles. Figure 4 The content involved in the nodes and connection relationships in the network structure.

[0057] For example, the modeled spatial interaction behavior is expressed as G = (N, E). Wherein, N = {n1, n2, ..., n |n|} represents the set of all vehicle features, and E={e1,e2,…,e |ε|} represents the set of interactions between them. |n| constitutes the total number of vehicles in traffic, and |ε| represents the total number of vehicle interactions. The feature matrix N represents the longitudinal position X, speed V, lane L, vehicle category I, lane head H, and lane tail T of each vehicle in the scene. Therefore, the feature matrix can be expressed as:

[0058]

[0059] The motion information of the vehicle in the scene can be expressed as:

[0060]

[0061] Among them, x i_position Indicates the current longitudinal position of the vehicle, x road Indicates the total length of the road; Indicates the current speed of the vehicle, v max Indicates the maximum speed limit of the road; L i Indicates the lane the vehicle is currently in; I i Indicates the type of vehicle (I i = 1 or 2 for autonomous vehicles with different driving tasks, otherwise for human-driven vehicles); H i Includes the distance between the i-th vehicle and the vehicles ahead of it in all lanes. i Includes the distance between the i-th vehicle and the vehicles behind it in all lanes.

[0062] Afterwards, in the context of intelligent networks, this application considers the interaction of multiple vehicles in a space. Specifically, the interaction between the i-th autonomous vehicle and all the j-th vehicles in the scene is denoted as e ij ∈{0,1}, where e ij =1 indicates that there is an interaction between the i-th vehicle and the j-th vehicle, otherwise there is no interaction. To represent spatial interaction behavior, the application conditions of this application are: 1) all autonomous vehicles can communicate and interact with each other; 2) autonomous vehicles can interact with manned vehicles within their perception range; 3) autonomous vehicles can interact with each other. Based on the above assumptions, the adjacency matrix E is obtained:

[0063]

[0064] In order to make the output dimension of GNN consistent with Transformer, a mask matrix M is introduced to retain the information of autonomous vehicles and remove the information of manned vehicles. If M = 1, it means that the current index is an autonomous vehicle, otherwise it is a manned vehicle:

[0065] M=[m1,m2,…,m i ,…,m n ] (5)

[0066] Graph Convolutional Networks (GCN) is an algorithm that performs convolution operations directly on graphs. Figure 5 As shown, when constructing the dynamic traffic scene into Figure 4After extracting the feature matrix, adjacency matrix, and mask matrix, we can directly use the neural network to model the entire graph. Figure 5 In the framework shown, the input layer includes the driving environment and the aforementioned adjacency matrix, feature matrix, and mask matrix. The graph convolutional network processing process includes aggregating the feature information of neighboring nodes through the adjacency matrix to update the current node representation. The calculation formula of the GCN calculation framework is as follows:

[0067]

[0068] in, I represents the identity matrix with the same dimensions as A; D represents the degree matrix, D ii =∑ j A ij ;H (l) Represents the characteristics of each node in the lth layer, H (0) Represents the original node feature, that is, H (0) =N;W (l) Represents the learnable parameters of the lth layer; σ represents the activation function, which uses the ReLU function.

[0069] Furthermore, if Figure 5 As shown in the figure, the fully connected layer can map the high-dimensional features output by the graph convolution to a more discriminative space. The linear rectification function (ReLU) can introduce nonlinearity to enhance the model's expressiveness and prevent gradient vanishing. The hierarchical structure may contain multiple "fully connected + ReLU" modules to gradually abstract features (such as from local interactions to global scene patterns). The feature concatenation process concatenates feature vectors from different layers (such as GCN layers and fully connected layers), fusing multi-scale information. Ultimately, a compact feature representation of each vehicle is generated, which can be used for downstream tasks (such as behavior prediction and trajectory generation).

[0070] Step 230: The vehicle collaborative decision-making problem is modeled as an MDP.

[0071] Corresponding to step 150, the collaborative decision-making problem of multiple autonomous vehicles is modeled as an MDP, which provides a mathematical framework for describing the state, action, reward, and state transition probability in the decision-making process.

[0072] Among them, the state space S of MDP describes all possible states of the autonomous driving vehicle in the traffic environment. In the multi-vehicle trajectory collaborative decision-making, it usually includes scene representation and vehicle motion information. The state s t It can be expressed as:

[0073] s t =[SR t ,N t ] (7)

[0074] Among them, SR t is the scene representation encoded by the transformer module, i.e., the scene feature encoding information in step 120. N t It is the vehicle motion information extracted by GNN.

[0075] The action space A describes all possible actions that the autonomous vehicle can take. In multi-vehicle collaborative decision-making, actions usually include longitudinal and lateral control actions, such as acceleration, deceleration, maintaining speed, changing lanes, etc. Action a t It can be expressed as:

[0076] a t =[a longitudinal ,a lateral ] (8)

[0077] Among them, a longitudinal Including acceleration, maintaining original speed, deceleration, a lateral Including changing lanes to the left, lane keeping, and changing lanes to the right.

[0078] The reward function R of MDP can be used to evaluate the pros and cons of each action, including traffic efficiency, task completion, and safety. For example, the reward function can be expressed as:

[0079] R(s t ,a t )=w1R speed +w2R collision +w3R intention (9)

[0080] Among them, R speed is the speed bonus, calculated as the ratio of the vehicle's current speed to its maximum speed:

[0081]

[0082] Among them, v i is the speed of the i-th vehicle, v max is the maximum speed limit and m is the total number of vehicles.

[0083] R collision is the collision penalty, calculated as the negative of the number of collisions:

[0084] R collision =-N collision (11)

[0085] Among them, N collision is the number of collisions in the current time step.

[0086] R intention is the mission intention reward, which encourages the vehicle to take correct actions when approaching the target location:

[0087]

[0088] Among them, I i is the distance between the i-th vehicle and the target position, L intention is the length of the intention reward region.

[0089] In addition, w1 to w3 are weight coefficients used to balance the importance of different reward components.

[0090] State transition probability P(s t+1 |s t ,a t ) describes the current state s t Next, take action a t Then transfer to the next state s t+1 In autonomous driving scenarios, the state transition probability is usually determined by the vehicle's dynamic model and the dynamic characteristics of the traffic environment.

[0091] Step 240: Use DQN to solve the MDP.

[0092] Corresponding to step 150, the goal of solving the MDP is to find an optimal policy π that maximizes the cumulative reward starting from the initial state. In an embodiment of the present application, a deep reinforcement learning algorithm DQN is used to solve the MDP.

[0093] Figure 6 This paper describes a driving decision model framework based on deep reinforcement learning (DRL), which combines vehicle action scenarios, data flow and training mechanism. Figure 6 As shown in the figure, the data input of the driving action scenario includes possible driving behaviors or decision-making tasks of the vehicle, such as changing lanes to the left, changing lanes to the right, accelerating to overtake, and merging on ramps. These actions correspond to typical scenarios in actual driving. The driving decision model framework can select the optimal action based on the environmental conditions (such as traffic flow and road structure).

[0094] In addition, the Q network structure includes the main network Q θ (s,a) and target network Weight replication is the process of periodically copying the parameters of the main network to the target network to stabilize the target Q value calculation (i.e., fix the target network parameters). The replay buffer is used to store historical experience tuples, and small batches of experience randomly sampled from the replay buffer can be used to train the network parameters.

[0095] In step 240, the Q network Q(s,a) is first initialized to estimate the Q value of the state space-action space pair. Given a policy π, the action value (Q value) of a state-action pair is defined as:

[0096]

[0097] Afterwards, the Bellman equation is used to calculate:

[0098]

[0099] Finally, the optimal Q-value function can be written as:

[0100]

[0101] Among them, γ∈[0,1] weighs the importance of immediate rewards and future rewards.

[0102] In addition, if Figure 6 As shown in Figure 1, during training, the state, action, reward, and next state for each step are stored in the experience replay buffer. A batch of data is randomly sampled from the experience replay buffer and the parameters of the Q network are updated to bring the Q value closer to the target Q value. Based on the updated Q network, the policy π is updated, and the action with the highest Q value is selected as the current cooperative driving policy.

[0103] Now refer to Figure 7 , shown is a block diagram of a device 700 according to one embodiment of the present application. The device 700 may include one or more processors 701 coupled to a controller hub 703. For at least one embodiment, the controller hub 703 communicates with the processor 701 via a multi-drop bus such as a front side bus (FSB), a point-to-point interface such as a quickpath interconnect (QPI), or a similar connection 710. The processor 701 executes instructions that control general types of data processing operations. In one embodiment, the controller hub 703 includes, but is not limited to, a graphics memory controller hub (GMCH) (not shown) and an input / output hub (IOH) (which may be on separate chips) (not shown), wherein the GMCH includes memory and a graphics controller and is coupled to the IOH.

[0104] The device 700 may also include a coprocessor 702 and a memory 704 coupled to the controller hub 703. Alternatively, one or both of the memory and the GMCH may be integrated within the processor, with the memory 704 and the coprocessor 702 directly coupled to the processor 701 and the controller hub 703, with the controller hub 703 and the IOH being in a single chip. The memory 704 may be, for example, a dynamic random access memory (DRAM), a phase change memory (PCM), or a combination of the two. In one embodiment, the coprocessor 702 is a special-purpose processor, such as, for example, a high-throughput MIC processor (many integrated core, MIC), a network or communication processor, a compression engine, a graphics processor, a general purpose computing on GPU (GPGPU), or an embedded processor, etc. The optional nature of the coprocessor 702 is indicated by a dotted line in Figure 7 middle.

[0105] The memory 704, as a computer-readable storage medium, may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. For example, the memory 704 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as one or more hard disk drives (HDD(s)), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives.

[0106] In one embodiment, device 700 may further include a network interface controller (NIC) 706. NIC 706 may include a transceiver for providing a radio interface for device 700, thereby communicating with any other suitable devices (e.g., a front-end module, an antenna, etc.). In various embodiments, NIC 706 may be integrated with other components of device 700. NIC 706 may implement the functionality of the communication unit in the above-described embodiments.

[0107] Device 700 may further include input / output (I / O) devices 705. I / O 705 may include: a user interface designed to enable a user to interact with device 700; a peripheral component interface designed to enable peripheral components to interact with device 700; and / or sensors designed to determine environmental conditions and / or location information related to device 700.

[0108] It is worth noting that Figure 7 This is for illustrative purposes only. Figure 7 It is shown that the device 700 includes multiple components such as a processor 701, a controller hub 703, a memory 704, etc. However, in actual applications, the device using the methods of the present application may only include a part of the components of the device 700, for example, it may only include the processor 701 and the NIC 706. Figure 7 The properties of the optional components are shown with dashed lines. According to some embodiments of the present application, the memory 704 as a computer-readable storage medium stores instructions that, when executed on a computer, cause the device 700 to perform the attention training method according to the above-described embodiment. For details, please refer to the method of the above-described embodiment and will not be repeated here.

[0109] Now refer to Figure 8 , which is a block diagram of a system on chip (SoC) 800 according to an embodiment of the present application. Figure 8 In FIG, similar components have the same reference numerals. In addition, the dashed boxes are optional features of more advanced SoCs. Figure 8 In the embodiment, SoC 800 includes: an interconnect unit 850 coupled to an application processor 810; a system agent unit 880; a bus controller unit 890; an integrated memory controller unit 840; a set of one or more coprocessors 820, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 830; and a direct memory access (DMA) unit 860. In one embodiment, the coprocessors 820 include specialized processors, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, or an embedded processor.

[0110] The static random access memory (SRAM) unit 830 may include one or more computer-readable media for storing data and / or instructions. The computer-readable storage medium may store instructions, specifically, temporary and permanent copies of the instructions. The instructions may include: when executed by at least one unit in the processor, causing the Soc 800 to perform the attention training method according to the above embodiment. For details, please refer to the method of the above embodiment, which will not be repeated here.

[0111] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0112] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0113] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0114] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to floppy disks, optical disks, optical discs, compact disc read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random-access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).

[0115] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0116] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.

[0117] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0118] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the application.

Claims

1. A method for collaborative decision-making among multiple autonomous driving vehicles, characterized in that: include: Acquiring and analyzing a traffic scene to obtain scene features of the traffic scene, the scene features including a plurality of vehicles in the traffic scene and a grid map corresponding to each of the plurality of vehicles, the grid map including a grid centered on the vehicle and occupied by the vehicle and surrounding vehicles; encoding the scene features based on the first neural network and the scene features to obtain scene feature encoding information; generating a scene graph structure based on the motion information of the multiple vehicles and the interaction relationship between the multiple vehicles, the scene graph structure including a plurality of nodes and a connection relationship between the multiple nodes, each of the multiple nodes being used to indicate a single vehicle and its motion information, and the connection relationship being used to indicate the interaction relationship between two vehicles corresponding to the nodes at both ends of the connection relationship; Determining spatial interaction information between the multiple vehicles based on a second neural network and the scene graph structure, where the second neural network is a graph neural network (GCN), and the spatial interaction information is used to indicate spatial interaction characteristics between the multiple vehicles; A cooperative driving strategy among the multiple vehicles is determined based on the scene feature coding information and the spatial interaction information.

2. The method according to claim 1, characterized in that Determining the cooperative driving strategy among the multiple vehicles includes: Modeling the collaborative decision-making problem among the multiple vehicles as a Markov decision process (MDP), wherein a state space of the MDP includes the scene feature encoding information and the motion information of the multiple vehicles, and an action space of the MDP includes control actions taken by the multiple vehicles; The deep reinforcement learning algorithm DQN is used to solve the MDP to obtain the cooperative driving strategy.

3. The method according to claim 2, characterized in that The reward function of the MDP is determined based on a speed reward, a collision penalty, and a mission intention reward. The speed reward is the ratio of the vehicle's current speed to its maximum speed, the collision penalty is the negative of the number of collisions, and the mission intention reward is used to instruct the vehicle to take the correct action when it reaches the target location.

4. The method according to claim 2 or 3, characterized in that Solving the MDP using the deep reinforcement learning algorithm DQN includes: Initialize the Q-network to estimate the Q-values ​​of state-space and action-space pairs; Update the Q network parameters through the Bellman equation to maximize the cumulative reward; Use the experience replay cache to store states, actions, rewards, and next states; The optimal action is selected based on the updated Q-network to generate cooperative driving decisions.

5. The method according to claim 1 or 2, characterized in that The encoding of the scene features comprises: Splicing the grid map into a scene representation matrix by row; Inputting the scene representation matrix into an embedding layer of the first neural network to obtain an embedding vector; Block processing is performed on the embedding vector based on the first neural network to obtain the scene feature encoding information.

6. The method according to claim 5, characterized in that The first neural network is a Transformer neural network.

7. The method according to claim 1 or 2, characterized in that Determining the spatial interaction information between the multiple vehicles includes: Acquire a feature matrix of the plurality of vehicles based on motion information of the plurality of vehicles, the motion information including longitudinal positions, speeds, lanes, categories, front vehicle spacing, and rear vehicle spacing of the plurality of vehicles; Obtaining an adjacency matrix of the plurality of vehicles based on interaction relationships between the plurality of vehicles, the interaction relationships including communications and interactions between autonomous driving vehicles in the plurality of vehicles and interactions between each autonomous driving vehicle and a manned vehicle within a perception range; Obtaining a mask matrix, where the mask matrix is ​​used to filter features of the autonomous driving vehicle and remove information about manned vehicles; Spatial interaction information between the multiple vehicles is determined based on the scene graph structure, the feature matrix, the adjacency matrix, the mask matrix and the graph neural network.

8. A device for collaborative decision-making among multiple autonomous driving vehicles, characterized in that: include: an acquisition unit, configured to acquire and analyze an input traffic scene to obtain scene features of the traffic scene, wherein the scene features include a plurality of vehicles in the traffic scene and a grid map corresponding to each of the plurality of vehicles, wherein the grid map includes a grid centered on the vehicle and occupied by the vehicle and surrounding vehicles; a processing unit configured to: encode the scene features based on the first neural network and the scene features to obtain scene feature encoding information; and generate a scene graph structure based on the motion information of the multiple vehicles and the interaction relationship between the multiple vehicles; Based on a second neural network and the scene graph structure, the spatial interaction information between the multiple vehicles is determined, and the second neural network is a graph neural network GCN; according to the scene feature coding information and the spatial interaction information, the collaborative driving strategy between the multiple vehicles is determined.

9. A device for collaborative decision-making of multiple autonomous driving vehicles, characterized in that: include: a memory for storing instructions executable by the processor; A processor, wherein the processor is configured to implement the method according to any one of claims 1 to 7 when executing the instructions.

10. A computer-readable storage medium storing instructions, characterized in that: When the instruction is executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent driving decision-making method and device based on multilevel risk perception graph neural network

    CN120932206A

  • An intelligent driving decision-making method and device based on a multi-level risk perception graph neural network

    CN120932206B

  • Road ramp opening behavior planning method and system and computer readable storage medium

    CN122337017A