Multi-vehicle Decision-making Method for Autonomous Driving Based on Hierarchical Reinforcement Learning of Multi-dimensional Weighted Graph

CN118092438BActive Publication Date: 2025-07-25BEIJING INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410196496.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2025-07-25
Estimated Expiration
2044-02-22

AI Technical Summary

Technical Problem

The existing deep reinforcement learning methods lack vehicle interaction information modeling in complex traffic scenarios with dynamic interaction, making it difficult to effectively distinguish vehicle decisions under different subtasks, resulting in insufficient efficiency and accuracy in handling dynamic uncertainty in traffic environments.

Method used

A multi-dimensional weighted graph hierarchical reinforcement learning method is adopted to characterize traffic scenes, build a horizontal and vertical decision model, and combine discrete and continuous deep reinforcement learning algorithms to perform hierarchical deep reinforcement learning of horizontal and vertical decision models, and finally the results are coupled and allocated to unmanned vehicles to realize multi-vehicle decision-making in autonomous driving.

Benefits of technology

It improves the efficiency and accuracy of unmanned driving systems in dealing with dynamic traffic environments, can better deal with changing road conditions and traffic scenarios, and enhances the flexibility and accuracy of the decision-making framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118092438B_ABST
    Figure CN118092438B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-vehicle decision-making method for autonomous driving based on hierarchical reinforcement learning of multi-dimensional weighted graphs, comprising the following steps: performing graph representation on a traffic scene, performing temporal representation on the represented graph, and obtaining a node feature matrix of the represented graph; by introducing expert knowledge, obtaining a sub-adjacency matrix of a lateral and longitudinal dimensional graph according to the vehicle type, lateral and longitudinal relative positions, and lateral and longitudinal relative speeds of vehicles in the node feature matrix; constructing a lateral decision-making model and a longitudinal decision-making model based on the sub-adjacency matrix of the lateral and longitudinal dimensional graph; respectively performing hierarchical deep reinforcement learning on the lateral decision-making model and the longitudinal decision-making model by using discrete and continuous deep reinforcement learning algorithms; coupling the hierarchical deep reinforcement learning results of the lateral decision-making model and the hierarchical deep reinforcement learning results of the longitudinal decision-making model and allocating them to corresponding driverless vehicles to achieve multi-vehicle decision-making for autonomous driving. The present invention enables the driverless system to more efficiently and accurately handle dynamic uncertainties in the traffic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of vehicle networking and autonomous driving, and particularly relates to a multi-vehicle decision-making method for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning. Background Art

[0002] In the field of autonomous driving, reinforcement learning algorithms have important application values. On the one hand, autonomous vehicles need to continuously make decisions to ensure safety and efficiency; on the other hand, the complex environments and dynamic interaction scenarios involved make traditional rule-based methods difficult to meet the actual needs. Therefore, reinforcement learning, as an adaptive decision-making method based on reward and punishment signals, can better meet the real-time, flexibility, and robustness of autonomous driving systems.

[0003] A graph neural network (GNN) is an artificial intelligence model based on a graph data structure that can effectively process data in complex non-Euclidean spaces. In the field of autonomous driving, the combination of GNN technology and reinforcement learning can better solve problems such as vehicle control and path planning, improving the efficiency and accuracy of autonomous driving decisions.

[0004] The introduction of graphs can model and learn the relationships between different nodes. In the field of autonomous driving, the combination of graph representation and hierarchical reinforcement learning can better solve decision-making problems in complex road conditions. In autonomous driving, hierarchical reinforcement learning can decompose the entire decision-making process into multiple subtasks, each subtask responsible for completing a specific decision-making goal. Graph representation can be used to construct a hierarchical decision-making model, modeling and learning through the allocation of weights between different nodes, so that the decision-making system can make more appropriate actions according to the strength of the interaction between nodes.

[0005] Existing deep reinforcement learning decision-making methods classify decision-making tasks and then use multi-objective reward functions to explore optimal actions. However, in complex traffic scenarios with dynamic interactions, existing methods lack vehicle interaction information modeling and lack discrimination for different-purpose vehicles under different subtask constraints. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, the present invention proposes a multi-vehicle decision-making method for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning. Based on the combination of multi-dimensional weighted graphs and hierarchical reinforcement learning, it not only considers the physical relationships between vehicles but also the dynamic interactions of vehicles, enabling the autonomous driving system to more efficiently and accurately handle dynamic uncertainties in the traffic environment.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] An autonomous driving multi-vehicle decision-making method based on hierarchical reinforcement learning of a multi-dimensional weighted graph, comprising the following steps:

[0009] Perform a graphical representation of the traffic scene, perform a temporal representation of the representation graph, and obtain a node feature matrix of the representation graph;

[0010] By introducing expert knowledge, obtain a sub-adjacency matrix of the lateral and longitudinal dimension graphs according to the vehicle type, lateral and longitudinal relative positions, and lateral and longitudinal relative speeds in the node feature matrix;

[0011] Construct a lateral decision-making model and a longitudinal decision-making model based on the sub-adjacency matrix of the lateral and longitudinal dimension graphs;

[0012] Adopt discrete and continuous deep reinforcement learning algorithms to perform hierarchical deep reinforcement learning on the lateral decision-making model and the longitudinal decision-making model respectively;

[0013] Couple the hierarchical deep reinforcement learning results of the lateral decision-making model and the longitudinal decision-making model and allocate them to the corresponding driverless vehicles to achieve autonomous driving multi-vehicle decision-making.

[0014] Preferably, the method for performing a graphical representation of the traffic scene includes:

[0015] Define that there are N vehicles in the traffic environment, where M driverless vehicles interact with the vehicles in the environment, and define the traffic environment as a weighted multi-dimensional directed graph Where represents the dimension of the graph; V(G) is the set of graph nodes, and each vehicle is regarded as a node of the graph E(G) is the set of graph edges, and the information interaction between vehicles is represented as an edge between vehicles For the lateral and longitudinal decision-making models in the traffic scene, according to the graph dimension Decouple the weighted multi-dimensional directed graph into a lateral representation graph G H (V H ,E H ) and a longitudinal representation graph G L (V L ,E L ).

[0016] Preferably, the method for performing a temporal representation of the representation graph includes:

[0017] At the t-th time step, it is represented as a node feature matrix and an adjacency matrix where F is the total number of features of each vehicle. For the node feature matrix of the lateral and longitudinal dimension graphs Extract sub-node feature matrices respectively In the formula

[0018] Node feature matrix For the i-th vehicle, it is expressed as the lateral speed V of its own vehicle Hi , longitudinal speed V Li ; lateral position X i ; longitudinal position Y i ; the relative collision time T of a total of six vehicles, namely the front and rear vehicles in the lane where the driverless vehicle is located and its adjacent lanes ij =Δs / Δv, j = 1, 2, 3, 4, 5, 6, where Δs represents the relative displacement between two vehicles along the lane traveling direction, and Δv represents the relative speed between two vehicles along the lane traveling direction; the section R where the own vehicle is located i ; the lane L where the own vehicle is located i ; the category I to which the own vehicle belongs i ; the specific expression is:

[0019]

[0020] Extract the sub-node feature matrix from the lateral dimension graph Among them, only consider the four vehicles T ij , j = 1, 3, 4, 6; extract the sub-node feature matrix from the longitudinal dimension graph Among them, only consider the two vehicles T ij , j = 2, 5.

[0021] Preferably, the sub-adjacency matrix of the lateral dimension graph is expressed as:

[0022]

[0023] Among them, for the j-th vehicle located in the adjacent lane of the driverless vehicle i, the distance along the lane direction is Δd ij , the distance threshold is X, and its weight function is defined as

[0024] Preferably, the sub-adjacency matrix of the longitudinal dimension graph is expressed as:

[0025]

[0026] Among them, for the j-th vehicle located in the lane where the driverless vehicle i is located, the distance along the lane direction is Δd ij , the distance threshold is Y, and its weight function is defined as For the j-th vehicle located in the lane where the driverless vehicle i is located, the speed difference along the lane direction is Δv ij , the speed difference threshold is ΔV; λ d and λ v are the proportionality coefficients of distance and speed difference respectively.

[0027] Preferably, the state space of the lateral decision-making model adopts the lateral dimension graph G of a weighted multi-dimensional directed graph H (V H , E H ), including the node feature matrix of the lateral dimension graph and the sub-adjacency matrix A t H ; the action space is where -1 represents activating a left lane change, 0 represents activating maintaining the current lane, and 1 represents activating a right lane change; the reward function is set as a function of the number of frequent lane changes L change , the time to collision with the vehicles in front and behind after a lane change and the lane position L in the task area Task :

[0028] Preferably, the state space of the longitudinal decision-making model adopts the longitudinal dimension graph G of a weighted multi-dimensional directed graph L (V L , E L ), including the node feature matrix of the longitudinal dimension graph and the sub-adjacency matrix A t L ; the action space is where a dec represents the maximum deceleration, and a acc represents the maximum acceleration; the reward function is set as a function of the current vehicle speed v t , the time to collision T with the vehicles in front and behind L and the current acceleration a t : R L = f L (v t , T L , a t ).

[0029] Preferably, a method for coupling the hierarchical deep reinforcement learning results of the lateral decision-making model and the hierarchical deep reinforcement learning results of the longitudinal decision-making model and allocating them to the corresponding driverless vehicles to achieve multi-vehicle decision-making for autonomous driving includes:

[0030] Transmitting specific instructions to the controller of the driverless vehicle to interact with the environment, forming a quadruple for storage and performing offline DRL training until the specified number of rounds ends or the required performance is achieved.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] 1) The multi-dimensional weighted graph hierarchical reinforcement learning framework proposed by the present invention is a significant improvement to the current autonomous driving decision-making system. The integration of the multi-dimensional weighted graph and hierarchical reinforcement learning not only considers the physical relationships between vehicles but also takes into account the dynamic interactions of vehicles, enabling the unmanned driving system to more efficiently and accurately handle the dynamic uncertainties in the traffic environment.

[0033] 2) The present invention uses a multi-dimensional weighted graph to represent the traffic environment and the interactions between vehicles, and designs dynamic adjacency matrices for the lateral and longitudinal decision-making modules respectively based on the relative positions and speed differences between vehicles. This method can effectively capture the complex interaction relationships between vehicles, which is particularly important in dynamic and uncertain traffic scenarios and is difficult to achieve in traditional deep reinforcement learning methods.

[0034] 3) The present invention decomposes the decision-making into lateral and longitudinal decision-making modules, and designs feature matrices and reward functions according to the tasks of lateral and longitudinal decision-making respectively, which can more flexibly handle changing road conditions and traffic scenarios. By taking the output of the lateral decision-making module as the input of the longitudinal decision-making module, the present invention realizes the effective coupling of the two decision-making modules. This design improves the flexibility and accuracy of the decision-making framework.

[0035] 4) The present invention applies discrete and continuous deep reinforcement learning algorithms such as DQN and its variants, DDPG, SAC, PPO, etc. in the lateral and longitudinal decision-making models respectively, which enables the model to better adapt to different types of decision-making tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0037] Figure 1 It is a schematic flowchart of the multi-vehicle decision-making method for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning in the embodiments of the present invention;

[0038] Figure 2 It is a schematic diagram of the framework of the multi-vehicle decision-making system for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning in the embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Embodiment 1

[0042] As Figure 1 shown, the present invention provides an autonomous driving multi-vehicle decision-making method based on hierarchical reinforcement learning of multi-dimensional weighted graphs, including the following steps:

[0043] Perform graph representation on the traffic scene, perform temporal representation on the represented graph, and obtain the node feature matrix of the represented graph;

[0044] By introducing expert knowledge, obtain the sub-adjacency matrix of the horizontal and vertical dimension graphs according to the types of vehicles, horizontal and vertical relative positions, and horizontal and vertical relative speeds in the node feature matrix;

[0045] Construct a horizontal decision-making model and a vertical decision-making model based on the sub-adjacency matrix of the horizontal and vertical dimension graphs;

[0046] Adopt discrete and continuous deep reinforcement learning algorithms to perform hierarchical deep reinforcement learning on the horizontal decision-making model and the vertical decision-making model respectively;

[0047] Couple the hierarchical deep reinforcement learning results of the horizontal decision-making model and the hierarchical deep reinforcement learning results of the vertical decision-making model and allocate them to the corresponding driverless vehicles to achieve autonomous driving multi-vehicle decision-making.

[0048] First, perform graph representation on the traffic scene. It is defined that there are N vehicles in the traffic environment, and M (M≤N) driverless vehicles interact with the vehicles in the environment, and the traffic environment is defined as a weighted multi-dimensional directed graph where represents the dimension of the graph; V(G) is the set of graph nodes, and each vehicle is regarded as a node of the graph E(G) is the set of graph edges, and the information interaction between vehicles is represented as the edges between vehicles For the horizontal and vertical decision-making models in the traffic scene, according to the graph dimension decouple the weighted multi-dimensional directed graph into a horizontal representation graph G H (V H , E H ) and a vertical representation graph GL (V L , E L ).

[0049] Furthermore, for the temporal representation of the graph, at the t-th time step, it is represented as the node feature matrix and the adjacency matrix where F is the total number of features of each vehicle. For the horizontal and vertical dimension graphs respectively extract the sub-node feature matrix In the formula

[0050] In the node feature matrix for the i-th vehicle, it is represented as its own vehicle's lateral speed V Hi , longitudinal speed V Li ; lateral position X i ; longitudinal position Y i ; the relative collision times T ij of the vehicles in front and behind (a total of six vehicles) in the lane where the unmanned vehicle is located and its adjacent lanes = Δs / Δv, j = 1, 2, 3, 4, 5, 6, where Δs represents the relative displacement between two vehicles along the lane traveling direction, and Δv represents the relative speed between two vehicles along the lane traveling direction; the section R i where the vehicle itself is located; the lane L i where the vehicle itself is located; the category I i to which the vehicle itself belongs; the specific expression is:

[0051]

[0052] When extracting the sub-node feature matrix from the horizontal dimension graph, only consider the four vehicles T ij , j = 1, 3, 4, 6; when extracting the sub-node feature matrix from the vertical dimension graph, only consider the two vehicles T ij , j = 2, 5; in addition, take the output Α M of the horizontal decision-making model as the sub-node feature matrix of the vertical decision-making model to couple the two networks.

[0053] Even further, by introducing expert knowledge, according to information such as the type of vehicle, horizontal and vertical relative positions, and horizontal and vertical relative speeds in the node feature matrix, the sub-adjacency matrices of the horizontal and vertical dimension graphs can be expressed as The calculation of the two sub-adjacency matrices is based on four assumptions: a) All unmanned vehicles within the specified range can share information in the traffic scenario. For example, the sharing of information between the i-th unmanned vehicle and the j-th vehicle within the sensing range is expressed as a ij= 1; b) Information cannot be shared between human-driven vehicles; c) All driverless vehicles can share the information of background vehicles within their perception range; d) A vehicle can share information with itself, denoted as a ii = 1.

[0054] Specifically, a sub-adjacency matrix of the lateral dimension graph is constructed based on the distance and lane between vehicles, and its weight is defined as a function of a distance, and the function value decreases as the distance increases, ranging from 0 to 1. The specific formula is defined as follows

[0055] For the j-th vehicle in the adjacent lane of the driverless vehicle i, the distance between them along the lane direction is Δd ij , the distance threshold is X, and its weight function is defined as where X is the distance threshold, with a value of 50 m. When Δd ij approaches the distance threshold X, the weight approaches 0. When Δd ij is very small, that is, the vehicles are very close, the weight approaches 1. Therefore, the sub-adjacency matrix of the lateral dimension graph is defined as:

[0056]

[0057] Specifically, for the distance and vehicle speed of two vehicles in the same lane, a sub-adjacency matrix of the longitudinal dimension graph is constructed, and its weight is defined as a function of a distance and a speed difference. The function value decreases as the distance increases and decreases as the speed difference increases, ranging from 0 to 1. The specific formula is defined as follows

[0058] For the j-th vehicle in the lane where the driverless vehicle i is located, the distance between them along the lane direction is Δd ij , the distance threshold is Y, and its weight function is defined as

[0059]

[0060] where the distance threshold Y is 50 m. When Δd ij is greater than the distance limit ratio λ d = 20 m, the weight approaches 0. When Δd ij is very small, that is, the vehicles are very close, the weight approaches 1; when Δv ij is greater than the speed limit ratio λ v = 5 m / s, the weight approaches 0. When Δv ij is very large, that is, the vehicle speed difference is very large, the weight approaches 1. Therefore, the sub-adjacency matrix of the longitudinal dimension graph is defined as:

[0061]

[0062] Furthermore, a hierarchical deep reinforcement learning model is built for the lateral decision-making model. The state space of the lateral decision-making model adopts the lateral dimension graph G of a weighted multi-dimensional directed graph H (V H , E H ), including the node feature matrix of the lateral dimension graph and the adjacency matrix A t H . The action space is where -1 represents activating a left lane change, 0 represents activating maintaining the current lane, and 1 represents activating a right lane change. Its reward function is set as a function of the number of frequent lane changes L change , the time to collision with the vehicles in front and behind after a lane change and the lane position L in the task area Task . Its sub-reward function is defined as follows:

[0063]

[0064]

[0065] If L change ≥ 2 within 5s,

[0066] Based on the above design of the lateral decision-making model, the discrete DRL algorithm D3QN algorithm and the GAT network model are used for training.

[0067] Furthermore, a hierarchical deep reinforcement learning model is built for the longitudinal decision-making model. The state space of the longitudinal decision-making model adopts the longitudinal dimension graph G of a weighted multi-dimensional directed graph L (V L , E L ), including the node feature matrix of the longitudinal dimension graph and the adjacency matrix A t L . The action space is Its reward function is set as a function of the current vehicle speed v t , the time to collision T with the vehicles in front and behind L and the current acceleration a t :

[0068]

[0069] R(T L ) = -5 / T L , T L ≤ 5s,

[0070] R(v t ) = V 限 - v t , if v t≤Low-speed limit V 限

[0071] Based on the above design of the longitudinal decision-making model, the continuous DRL algorithm DDPG and the GAT network model are used for training.

[0072] Based on the above design, at each time step, the actions of the lateral and longitudinal decision-making modules are transmitted to the centralized control system, and the specific instructions are transmitted to the controller of the unmanned vehicle for interaction with the environment. A quadruple is formed for storage for offline DRL training until the specified round ends or the required performance is achieved.

[0073] Embodiment 2

[0074] The present invention also provides an autonomous driving multi-vehicle decision-making system based on multi-dimensional weighted graph hierarchical reinforcement learning, the framework of which is as Figure 2 shown, including: the decision-making framework includes multi-dimensional weighted graph representation, lateral and longitudinal decision-making modules, and hierarchical reward functions. In the lateral decision-making module and the longitudinal decision-making module, a graph representation matrix space and a reward function are designed for the two modules respectively; the output of the lateral decision-making module is used as the input of the longitudinal decision-making module, and the two decision-making modules are coupled; the final action is allocated by the centralized control system to the corresponding unmanned vehicle, forming an autonomous driving multi-vehicle decision-making method based on multi-dimensional weighted graph hierarchical reinforcement learning.

[0075] First, the traffic scene is graphically represented. It is defined that there are N vehicles in the traffic environment, and among them, M (M≤N) unmanned vehicles interact with the vehicles in the environment, and the traffic environment is defined as a weighted multi-dimensional directed graph where represents the dimension of the graph; V(G) is the set of graph nodes, and each vehicle is regarded as a node of the graph E(G) is the set of graph edges, and the information interaction between vehicles is represented as the edges between vehicles For the lateral and longitudinal decision-making models in the traffic scene, according to the graph dimension the weighted multi-dimensional directed graph is decoupled into a lateral representation graph G H (V H ,E H ) and a longitudinal representation graph G L (V L ,E L ).

[0076] Furthermore, for the temporal representation of the graph, at the t-th time step, it is represented as the node feature matrix and the adjacency matrix where F is the total number of features of each vehicle. For the lateral and longitudinal dimension graphs respectively extract the sub-node feature matrix In the formula

[0077] The node feature matrix For the i-th vehicle, it is expressed as the lateral speed V of its own vehicle Hi , and the longitudinal speed V Li ; the lateral position X i ; the longitudinal position Y i ; the relative collision time T of the vehicles in front and behind (a total of six vehicles) in the lane where the driverless vehicle is located and its adjacent lanes ij =Δs / Δv, j = 1, 2, 3, 4, 5, 6, where Δs represents the relative displacement along the lane direction between two vehicles, and Δv represents the relative speed along the lane direction between two vehicles; the section R where the own vehicle is located i ; the lane L where the own vehicle is located i ; the category I to which the own vehicle belongs i ; the specific expression is:

[0078]

[0079] Extract the sub-node feature matrix from the lateral dimension graph Among them, only consider the four vehicles T in the adjacent lanes ij , j = 1, 3, 4, 6; extract the sub-node feature matrix from the longitudinal dimension graph Among them, only consider the two vehicles T in front and behind in the same lane ij , j = 2, 5; in addition, take the output Α M of the lateral decision-making model as the sub-node feature matrix of the longitudinal decision-making model to couple the two networks.

[0080] Furthermore, by introducing expert knowledge, according to information such as the type, horizontal and vertical relative positions, and horizontal and vertical relative speeds of the vehicles in the node feature matrix, the sub-adjacency matrix of the horizontal and vertical dimension graphs can be expressed as The calculation of the two sub-adjacency matrices is based on four assumptions: a) All driverless vehicles within a specified range can share information in the traffic scenario. For example, the sharing of information between the i-th driverless vehicle and the j-th vehicle within the sensing range is expressed as a ij = 1; b) Information cannot be shared between human-driven vehicles; c) All driverless vehicles can share the information of background vehicles within their sensing range; d) A vehicle can share information with itself, expressed as a ii = 1.

[0081] Specifically, based on the distance between vehicles and the lane, construct the sub-adjacency matrix of the lateral dimension graph, and its weight is defined as a function of distance, and the function value decreases with the increase of distance, ranging from 0 to 1. The specific formula is defined as follows

[0082] For the j-th vehicle in the adjacent lane of the driverless vehicle i, the distance between them along the lane direction is Δd ij , the distance threshold is X, and its weight function is defined as When Δd ij approaches the distance threshold X, the weight approaches 0. When Δd ij is very small, that is, the vehicles are very close, the weight approaches 1. Therefore, the sub-adjacency matrix of the lateral dimension graph is defined as:

[0083]

[0084] Specifically, for the distance and vehicle speed of two vehicles in the same lane, the sub-adjacency matrix of the longitudinal dimension graph is constructed, and its weight is defined as a function of the distance and speed difference. The function value decreases with the increase of the distance and decreases with the increase of the speed difference, ranging from 0 to 1. The specific formula is defined as follows

[0085] For the j-th vehicle in the lane where the driverless vehicle i is located, the distance between them along the lane direction is Δd ij , the distance threshold is Y, and its weight function is defined as

[0086]

[0087] When Δd ij approaches the distance threshold X, the weight approaches 0. When Δd ij is very small, that is, the vehicles are very close, the weight approaches 1. Therefore, the sub-adjacency matrix of the longitudinal dimension graph is defined as:

[0088]

[0089] Furthermore, a hierarchical deep reinforcement learning model is built for the lateral decision-making model. The state space of the lateral decision-making model uses the lateral dimension graph G H (V H , E H ) of the weighted multi-dimensional directed graph, including the node feature matrix of the lateral dimension graph and the adjacency matrix A t H . The action space is where -1 represents activating a left lane change, 0 represents activating maintaining the current lane, and 1 represents activating a right lane change. Its reward function is set as a function of the number of frequent lane changes L change , the collision time with the front and rear vehicles after the lane change and the lane position L Task in the task area:

[0090]

[0091] Based on the above design of the lateral decision-making model, a discrete DRL algorithm is adopted, including DQN and its variants, the DiscreteAC architecture algorithm, and various graph neural networks, including network models such as GNN and GAT, for training.

[0092] Furthermore, a hierarchical deep reinforcement learning model is built for the longitudinal decision-making model. The state space of the longitudinal decision-making model adopts the longitudinal dimension graph G of a weighted multi-dimensional directed graph L (V L , E L ), including the node feature matrix of the longitudinal dimension graph and the adjacency matrix A t L . The action space is where a dec represents the maximum deceleration, and a acc represents the maximum acceleration. Its reward function is set as a function of the current vehicle speed v t , the collision time T between the front and rear vehicles L and the current acceleration a t :

[0093] R L = f L (v t , T L , a t ).

[0094] Based on the above design of the longitudinal decision-making model, a continuous DRL algorithm is adopted, including DDPG and its variants, SAC, PPO algorithms, and various graph neural networks, including network models such as GNN and GAT, for training.

[0095] Based on the above design, at each time step, the actions of the lateral and longitudinal decision-making modules are transmitted to the centralized control system, and specific instructions are transmitted to the controller of the driverless vehicle for interaction with the environment. A quadruple is formed and stored for offline DRL training until the specified round ends or the required performance is achieved.

[0096] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An autonomous driving multi-vehicle decision-making method based on multi-dimensional weighted graph hierarchical reinforcement learning, characterized in that, Including the following steps: Perform graphical representation on the traffic scenario, perform temporal representation on the representation graph, and obtain the node feature matrix of the representation graph; By introducing expert knowledge, obtain the sub-adjacency matrix of the horizontal and vertical dimension graphs according to the vehicle type, horizontal and vertical relative positions, and horizontal and vertical relative speeds in the node feature matrix; Construct a horizontal decision-making model and a vertical decision-making model based on the sub-adjacency matrix of the horizontal and vertical dimension graphs; Use discrete and continuous deep reinforcement learning algorithms to perform hierarchical deep reinforcement learning on the horizontal decision-making model and the vertical decision-making model respectively; Couple the hierarchical deep reinforcement learning results of the horizontal decision-making model and the vertical decision-making model and allocate them to the corresponding unmanned vehicles to achieve multi-vehicle decision-making for autonomous driving; The method for performing graphical representation on the traffic scenario includes: It is defined that there are N vehicles in the traffic environment, among which M autonomous vehicles interact with the vehicles in the environment, and the traffic environment is defined as a weighted multi-dimensional directed graph G(T, V, E); where T ∈ (H, L) represents the dimension of the graph; V(G) is the set of graph nodes, and each vehicle is regarded as a node u ∈ V T ; E(G) is the set of graph edges, and the information interaction between vehicles is represented as an edge (u, v) ∈ E between vehicles T ; For the horizontal and longitudinal decision-making models in the traffic scenario, according to the graph dimension T ∈ (H, L), the weighted multi-dimensional directed graph G(T, V, E) is decoupled into a horizontal representation graph G H (V H , E H ) and a longitudinal representation graph G L (V L , E L ); The method for performing temporal representation on the representation graph includes: Is represented as a node feature matrix at the t-th time step And the adjacency matrix Where F is the total number of features of each vehicle, for the node feature matrices of the horizontal and vertical dimension graphs Respectively extract the sub-node feature matrices In the formula F T ≤ F; Node feature matrix For the i-th vehicle, it is expressed as the lateral speed V of its own vehicle Hi , longitudinal speed V Li ; lateral position X i ; longitudinal position Y i ; the relative collision time T of a total of six vehicles, namely the front and rear vehicles in the lane where the driverless vehicle is located and its adjacent lanes, is T ij =Δs / Δv, j = 1, 2, 3, 4, 5, 6, where Δs represents the relative displacement between two vehicles along the lane traveling direction, and Δv represents the relative speed between two vehicles along the lane traveling direction; the section R where the own vehicle is located i ; the lane L where the own vehicle is located i ; the category I to which the own vehicle belongs i ; the specific expression is: Extract the sub-node feature matrix for the horizontal dimension graph Among them, only consider four vehicles T in adjacent lanes ij , j = 1, 3, 4, 6; Extract the sub-node feature matrix for the vertical dimension graph Among them, only consider two vehicles T in the front and back of the same lane ij , j = 2, 5; The sub-adjacency matrix of the horizontal dimension graph is expressed as: Among them, for the distance along the lane direction between the j-th vehicle in the adjacent lane of the driverless vehicle i is Δd ij , the distance threshold is X, and its weight function is defined as The sub-adjacency matrix of the vertical dimension graph is expressed as: Among them, the distance along the lane direction between the j-th vehicle in the lane where the driverless vehicle i is located is Δd ij , the distance threshold is Y, and its weight function is defined as The speed difference along the lane direction between the j-th vehicle in the lane where the driverless vehicle i is located is Δv ij , the speed difference threshold is ΔV, l d and l v are the proportionality coefficients of the distance and the speed difference respectively.

2. The multi-vehicle decision-making method for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning according to claim 1, wherein The state space of the horizontal decision-making model adopts the horizontal dimension graph G of the weighted multi-dimensional directed graph H (V H , E H ), including the node feature matrix of the horizontal dimension graph and the sub-adjacency matrix A t H ; The action space is where -1 represents activating a left lane change, 0 represents activating maintaining the current lane, and 1 represents activating a right lane change; The reward function is set as a function of the number of frequent lane changes L change , the time to collision with the vehicles in front and behind after lane changing and the lane position L in the mission area Task :

3. The multi-vehicle decision-making method for autonomous driving based on multi-dimensional weighted graph hierarchical reinforcement learning according to claim 1, characterized in that, The state space of the longitudinal decision-making model adopts the longitudinal dimension graph G of the weighted multi-dimensional directed graph L (V L , E L ), including the node feature matrix of the longitudinal dimension graph and the sub-adjacency matrix A t L ; the action space is where a dec represents the maximum deceleration, and a acc represents the maximum acceleration; The reward function is set as a function of the current vehicle speed v t , the collision time T between the front and rear vehicles L and the current acceleration a t : R L = f L (v t , T L , a t ).

4. The multi-vehicle decision-making method for autonomous driving based on hierarchical reinforcement learning of a multi-dimensional weighted graph according to claim 1, wherein The method for coupling the hierarchical deep reinforcement learning results of the horizontal decision-making model and the vertical decision-making model and allocating them to the corresponding unmanned vehicles to achieve multi-vehicle decision-making for autonomous driving includes: Transmit specific instructions to the controller of the unmanned vehicle to interact with the environment, form a quadruple for storage and perform offline DRL training until the specified round ends or the required performance is achieved.

Citation Information

Patent Citations

  • Unmanned bus cluster decision-making method based on graph neural network reinforcement learning

    CN115731690A

  • Automatic driving vehicle behavior decision-making and model building method and device and vehicle

    CN116306264A

  • Automatic driving centralized decision-making method based on series-parallel hierarchical reinforcement learning

    CN116502703A