Non-signalized intersection end-to-end vehicle motion control method based on DRL
Through the end-to-end vehicle motion control method based on DRL, the information screening module and the DRL module extract the space-time interaction characteristics, the problem of complex interaction relationships between vehicles in signal-free intersections is solved, and the safe and stable passage of autonomous vehicles is achieved.
Patent Information
- Application Number
- CN202510765068.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-10
AI Technical Summary
It is difficult for autonomous vehicles to effectively handle the complex interaction between vehicles in complex traffic scenarios without signal intersections, resulting in difficulty in ensuring safety and stability.
The end-to-end vehicle motion control method based on DRL is adopted, and the space-time interaction state and low-dimensional environment information are extracted through the information screening module, and the driving state characteristics are extracted in combination with the Actor network and the Critic network, and the decision output is made through the smoothing processing module to achieve safe passage of the vehicle.
It improves the driving safety and stability of autonomous vehicles in a signal-free intersection environment, can accurately identify traffic conflict relationships and make the optimal strategic choices.
Smart Images

Figure CN120353172A_ABST
Abstract
Description
Technical Field:
[0001] The present invention belongs to the field of intelligent driving, and more particularly, to an end-to-end vehicle motion control method based on DRL for signal-free intersections. Background Art:
[0002] With the rapid development of computer science and communication technology, autonomous driving technology has become a frontier research field. It shows broad application prospects in improving traffic safety, optimizing traffic efficiency, and promoting the development of intelligent transportation, and has become an important technical path to solve traditional traffic problems.
[0003] Currently, commercial applications of autonomous driving have been realized in the motion control of simple autonomous driving scenarios such as ACC and AEB. However, in complex traffic scenarios, there are still safety hazards in autonomous driving. Especially in the key scenario of signal-free intersections in urban traffic, due to the lack of effective traffic signal guidance, the traffic flow shows high dynamics and randomness. At this time, it is difficult for autonomous vehicles to effectively handle the complex interaction relationships between vehicles and cannot complete the motion control task while ensuring safety.
[0004] As a data-driven learning method, reinforcement learning provides new possibilities for solving the challenges of autonomous driving in complex traffic scenarios. The reinforcement learning method does not rely on pre-programmed rules, but guides the agent to learn through a reward mechanism, enabling the vehicle to master how to make optimal decisions in a changing environment through continuous trial and error.
[0005] Currently, the research on autonomous driving technology mainly focuses on modular methods and end-to-end methods. The traditional modular method decomposes the autonomous driving task into multiple independent modules and conducts information transmission through explicit rules. This method has good interpretability and causality, facilitating problem tracing and module replacement. However, due to the error accumulation in the information transmission process between modules, the overall performance of the system may be limited. Especially when dealing with complex traffic environments, it is difficult to achieve sufficient flexibility and adaptability. The end-to-end method directly maps sensor data to trajectories or control signals through a neural network, with strong adaptive and global optimization capabilities, and can show unique advantages when dealing with complex tasks in high-dynamic scenarios such as signal-free intersections.
[0006] In the study of motion control tasks in the scene of unsignalized intersections, a variety of methods have been proposed, each with its own advantages but also limitations. Patent CN118155429A proposes a traffic adaptive control method based on multi-agent reinforcement learning, which introduces a priority experience playback mechanism and multi-agents, enhances the sample utilization efficiency in the learning process, and accelerates the convergence speed of the algorithm. However, this method does not consider the complex interaction relationship between vehicles in unsignalized intersections, which reduces driving safety. Patent CN108932840A proposes a method for unmanned vehicle traffic in urban intersections based on reinforcement learning, which uses a camera method to collect continuous operation status information and location information of the vehicle, provides rich data support for decision-making, and optimizes the traffic strategy. However, the reinforcement learning neural network described in this method uses a simple fully connected layer, and has certain limitations in understanding the complex interaction relationship between vehicles in the scene of unsignalized intersections, and it is difficult to deal with existing traffic conflicts. Although the above method can improve the efficiency of vehicle cooperative traffic in an unsignalized intersection environment, it is difficult to ensure the safety and stability of smart car driving. Summary of the invention:
[0007] In order to solve the problems existing in the above technical background, the present invention provides an end-to-end vehicle motion control method for unsignalized intersections based on DRL. The method adopts an end-to-end reinforcement learning algorithm, fully utilizes the end-to-end ability to directly perform global optimization from input to output, makes decisions in real time and efficiently, and realizes the acquisition of the optimal safe passage strategy in an unsignalized intersection environment. In addition, the method introduces the spatiotemporal interaction state based on graph structure data in the basic state space, and fully considers the complex interaction relationship between vehicles, so that the intelligent car has the ability to infer the behavioral intention of traffic conflict objects and improves driving safety. At the same time, a control signal smoother is designed to prevent the safety hazards caused by the oscillation of the Actor network control signal.
[0008] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0009] The present invention relates to an end-to-end vehicle motion control method based on DRL, which includes an environment, an information screening module, a DRL module, and a smoothing processing module. First, the information screening module obtains state information from the environment, refines it into low-dimensional environment information and spatio-temporal interaction states, constructs a state space, and outputs a driving state from the state space to the DRL module. Secondly, the Actor network in the DRL module receives the driving state, extracts spatial interaction features in the driving state using a graph attention layer, captures temporal features in the driving state through a multi-head attention layer, and uses a fully connected layer to splice the spatial interaction features, temporal features, and low-dimensional environment information features, further mapping them to the mean and standard deviation of an action, and outputs a vehicle control amount through Gaussian distribution reparameterization sampling. After the Critic network receives the driving state, it extracts spatial interaction features and temporal features in the driving state using a graph convolutional layer and a long short-term memory layer respectively, and splices the spatial interaction features, temporal features, and low-dimensional environment information features through a fully connected layer. Finally, the smoothing processing module smooths the vehicle control amount, and after smoothing, it is used as an action to be transmitted to the autonomous vehicle in the environment for motion control, and the state information of the environment is updated.
[0010] The method includes the following steps:
[0011] Step 1, Reinforcement learning model design:
[0012] Step 1.1, State space design:
[0013] The present invention designs a state space, which includes spatio-temporal interaction states and low-dimensional environment information. It enables the intelligent agent to obtain spatio-temporal interaction information between traffic participants and itself. The state space S is as shown in Equation (1):
[0014] S=(S st ,S ld ) (1)
[0015] In the formula, S st is the spatio-temporal interaction state, and S ld is the low-dimensional environment information.
[0016] Step 1.2, Action space design:
[0017] In the present invention, the vehicle control amount u is used as the action space a, and the lateral control amount and longitudinal control amount in the vehicle control amount are limited within a threshold range to prevent the vehicle from performing overly dangerous actions. Therefore, the action space a is defined as shown in Equation (2):
[0018] a = u = [α, δ], α0 ≤ α ≤ α1; δ0 ≤ δ ≤ δ1 (2)
[0019] In the formula, δ is the lateral control variable, and α is the longitudinal control variable. α0, α1, δ0, and δ1 are the set constraint thresholds.
[0020] Step 1.3, Reward function design:
[0021] In the present invention, for the intelligent vehicle motion control task at an intersection without signals, the present invention comprehensively considers the safety and comfort, traffic efficiency, path keeping ability of the intelligent vehicle, and traffic rules to design the reward function. And at different stages of the task, different weights are given to different reward items. The reward function is mainly set as in Formulas (3) to (5):
[0022]
[0023]
[0024]
[0025] In the formula, r target is the target reward item, r track is the path deviation reward item, r rule is the rule reward item, r v is the speed reward item, r safe is the safety reward item. ω1, ω2, ω3, ω4, and ω5 are the reward distribution weights, which are different at different task stages. x ego , y ego are the horizontal and vertical coordinates of the vehicle itself, x end and y end are the horizontal and vertical coordinates of the end point of the vehicle itself, b1, b2, b3, c1, c2, c3, and c4 are constant terms. Δ path is the path deviation, n is the total number of steps, n step is the number of steps the vehicle has traveled. v ego is the speed of the vehicle itself, v E is the desired speed, v0 is the initial speed of the vehicle itself, ttc is the time to collision, t1 is the maximum reaction time, s0 is the relative distance between the vehicle itself and traffic participants, s E is the desired safety distance. x A,j and y A,j are the horizontal and vertical coordinates of the j-th point of the predicted path, x path,j and y path,j are the horizontal and vertical coordinates of the corresponding reference path point, and are the yaw angles of the j-th point in the predicted path and the yaw angle of the corresponding reference path point respectively, c Δ is a constant term.
[0026] Step 2, Construction of the information screening module:
[0027] The information screening module extracts the spatio-temporal interaction state and low-dimensional environment information from the status information obtained from the environment, which constitutes the state space, and outputs the driving state to the DRL module from the state space. The spatio-temporal interaction state includes spatial interaction information and time information. Among them, the spatial interaction information is obtained in the form of graph structure data. Specifically, the traffic participants at each moment are modeled as nodes in the graph structure data, including four eigenvalue, namely the abscissa, ordinate, speed and yaw angle of the traffic participant, which is expressed as the node feature matrix X t ; the interaction relationship between traffic participants is represented by the edge between nodes, which is expressed as the adjacency matrix A; the node feature matrix X t and the adjacency matrix A are as shown in equations (6) and (7):
[0028]
[0029]
[0030] In the formula, and are the abscissa and ordinate of the ego vehicle at time t, is the speed of the ego vehicle at time t, is the yaw angle of the ego vehicle at time t. and are the abscissa and ordinate of the first surrounding vehicle at time t, is the speed of the first surrounding vehicle at time t, is the heading angle of the first surrounding vehicle at time t. and are the abscissa and ordinate of the nth surrounding vehicle at time t, is the speed of the nth surrounding vehicle at time t, is the heading angle of the nth surrounding vehicle at time t. The adjacency matrix A is an n×n matrix, and n represents the number of vehicles participating in the interaction. Among them, if there is an interaction relationship between node i and node j, the element value in the matrix is 1, otherwise it is 0.
[0031] Secondly, a prediction equation is established with reference to the kinematic equation to obtain the time series of the motion state and determine the time information in the spatio-temporal interaction state. Specifically, the node feature matrix X t provides the spatial interaction information required for the prediction equation at time t, and takes the motion states of the surrounding vehicles at the previous j moments to predict the motion states of the ego vehicle at the next k moments, thereby obtaining the time series of the motion state; the prediction equation is as shown in equation (8):
[0032]
[0033] In the formula, and are the abscissa and ordinate of the ego vehicle at time t, and are the components of the ego vehicle's speed in the x and y directions at time t, and are the components of the ego vehicle's acceleration in the x and y directions at time t, respectively. d t is the prediction interval.
[0034] In summary, the spatio-temporal interaction state S is determined based on the spatial interaction information and the time information st , expressed as Equation (9):
[0035] S st = ((X t-j ; A)…(X t ; A)…(X t+k ; A)) (9)
[0036] In the formula, X t-j , X t and X t+k are the node feature matrices at the historical j time instants, time t, and the future k time instants, respectively. A is the adjacency matrix representing the interaction relationship between the ego vehicle and the surrounding vehicles.
[0037] After obtaining the spatio-temporal interaction state, the low-dimensional environmental information S ld is refined by combining the relative distance between the ego vehicle and the surrounding vehicles and the starting and ending points of the task, as shown in Equation (10):
[0038] S ld = (l1…l n , x start , y start , x end , y end ) (10)
[0039] In the formula, l1…l n represent the relative distances between the ego vehicle and each surrounding vehicle, with a total of n. x start and y start are the starting position coordinates of the task, and x end and y end are the ending position coordinates of the task.
[0040] Finally, the spatio-temporal interaction state S st and the low-dimensional environmental information S ld are integrated into the state space S, as shown in Equation (11):
[0041] S = (S st , S ld ) (11)
[0042] Step 3: Construction of the DRL module:
[0043] The DRL module includes an Actor network and a Critic network. Among them, the Actor network combines the driving state and outputs the vehicle control quantity. The Critic network outputs the action state value according to the state characteristics of the driving state and the actions output by the smoothing processing module. The Actor network updates its parameters according to the action state value output by the Critic network, while the Critic network updates its parameters according to the rewards feedback from the environment.
[0044] The Actor network in the DRL module includes a graph attention layer, a multi-head attention layer, and a fully connected layer. Among them, the graph attention layer calculates the attention weights between nodes, assigns different weights to adjacent nodes, and then extracts the spatial interaction features of the driving state. The calculation of the attention weights is as shown in Equation (12):
[0045]
[0046] In the formula, a ij is the attention weight between the i-th node h i and the j-th adjacent node h j . LeakyReLU is the activation function, a T is the attention vector, ω i and ω j are the linear transformation parameter matrices, || represents vector concatenation. h i ' is the vehicle node feature after the attention weight update, and ω is the weight matrix.
[0047] The multi-head attention layer in the Actor network receives the spatial interaction features extracted by the graph attention layer, arranges the spatial interaction features in chronological order to form a time feature sequence, and performs a linear transformation on the time feature sequence to construct the query Q, key K, and value V matrices. Subsequently, the attention weights are calculated in parallel by multiple attention heads, and the attention weights calculated by each attention head are integrated to obtain the final multi-head attention output vector, thereby capturing the time features of the driving state; the linear transformation, attention weight calculation, and information integration are as shown in Equation (13):
[0048]
[0049] In the formula, Q, K, and V are the query, key, and value matrices respectively, ω Q , ω K , ω V and ω o are the weight matrices, x is the time feature sequence, Attention(Q, K, V) is the calculation of a single attention head, softmax is the activation function, QK T represents the dot product of Q and K, and d kis the dimension of the key vector K. MultiHead(Q, K, V) is the multi-head attention mechanism, and head h is the h-th attention head. Concat(head1, head2,..., head h ) represents concatenating the outputs of multiple attention heads.
[0050] The fully connected layer in the Actor network concatenates the spatial interaction features output by the graph attention layer, the temporal features output by the multi-head attention layer, and the low-dimensional environmental information features in the driving state to construct state features, and maps the state features to the mean and standard deviation of actions. Subsequently, the vehicle control quantity is output through Gaussian distribution reparameterization sampling.
[0051] The Critic network in the DRL module includes a graph convolutional layer, a long short-term memory layer, and a fully connected layer. Among them, the graph convolutional layer adds the identity matrix I N to construct a self-connected adjacency matrix Secondly, through graph convolutional operations, using the weight matrix W (l) to linearly transform the node feature matrix H (l) at layer l and combine it with the self-connected adjacency matrix for calculation to achieve the propagation and fusion of neighbor node information. Finally, through non-linear processing, the node feature matrix H (l) at layer l is updated to the node feature matrix H (l+1) at layer l + 1 containing its own and neighbor node features, thereby extracting the spatial interaction features of the driving state. This process is formalized as Equation (14):
[0052]
[0053] In the formula, is the adjacency matrix with self-connection added, A is the adjacency matrix, and I N is the identity matrix, H (l) W (l) is the linear transformation process, is the process of propagation and fusion of neighbor node information, σ(·) is the process of introducing non-linearity into the node feature matrix, H (l+1) and H (l) are the node feature matrices at layer l + 1 and layer l respectively, σ is the activation function, and W (l) is the weight matrix.
[0054] The long short-term memory layer in the Critic network receives the spatial interaction features extracted by the graph convolutional layer, arranges the spatial interaction features in chronological order to form a time feature sequence. This sequence passes through the gating mechanism of the long short-term memory network at each time step, jointly with the forget gate f t and the input gate it and the output gate o t jointly control the cell state c t and the hidden state h t at the current time step, thereby realizing the selective retention and integration of information. The hidden state h t at each time step of the output can be used as the time feature of that time step. Finally, to simplify the representation and aggregate historical information, the hidden state at the last time step in the sequence is selected as the time feature of the driving state. The state update process in the long short-term memory network is as shown in Equation (15):
[0055]
[0056] where h t-1 is the hidden state of the previous time step, x t is the time feature sequence, σ is the activation function, tanh represents the hyperbolic tangent function, ω f , ω i , ω c and ω k are weight matrices, b f , b i , b c and b o are bias vectors. c t-1 is the cell state of the previous time step, f t is the forget gate, i t is the input gate, is the candidate cell state, c t is the cell state at the current time step, o t is the output gate, h t represents the hidden state at the current time step.
[0057] The fully connected layer in the Critic network concatenates the spatial interaction features output by the graph convolutional layer, the time features output by the long short-term memory layer, and the low-dimensional environmental information features in the driving state to determine the state features. After the action is passed to the Critic network in the smoothing processing module, the action is evaluated in combination with the state features of the driving state, and the action state value is output.
[0058] Step 4, Construction of the smoothing processing module:
[0059] The smoothing processing module uses the moving average method to smooth the vehicle control quantity output by the DRL module, and transfers the smoothed action to the autonomous vehicle in the environment for motion control, updating the state information of the environment. At the same time, it is passed to the Critic network in the DRL module, so that it evaluates the action according to the state features of the driving state and outputs the action state value. The smoothing processing by the moving average method is as shown in Equation (16):
[0060]
[0061] Wherein, and are new longitudinal and lateral control quantities, α t +α t-1 +…+α t-n+1 and δ t +δ t-1 +…+δ t-n+1 are the longitudinal and lateral control quantities of the Actor network at the past n moments.
[0062] Step 5. Update network parameters in the DRL module:
[0063] Step 5.1. Selection of reinforcement learning algorithm:
[0064] In order to extract spatio-temporal interaction features of the driving state in the neural network and complete the motion control task of the intelligent vehicle in the environment of an unsignalized intersection, the present invention uses the Soft Actor-Critic algorithm as the reinforcement learning algorithm for the intelligent vehicle.
[0065] Step 5.2. Extraction of experience samples:
[0066] After one round of loop, the Actor and Critic networks in the DRL module need to update their network parameters to optimize the policy. The Actor network is updated according to the action-state value output by the Critic network, and the Critic network is updated according to the reward feedback by the environment after the selected action. For the update, a batch of experience samples need to be randomly extracted from the experience pool and assigned to each network.
[0067] Step 5.3. Experience replay:
[0068] Split the quadruple [s, a, r, s - given in the randomly extracted experience samples, and respectively form vectors of the state s, the action a, the reward r, and the next state s - . Transmit them into the Actor network and the Critic network for experience replay, and calculate the loss functions Loss of the Actor network and the Critic network according to the quadruple, so as to update the parameters by gradient descent backpropagation, thereby learning a better safety control policy after the update.
[0069] The beneficial effects of the present invention are as follows: The present invention relates to an end-to-end vehicle motion control method based on DRL, which includes an information screening module, a DRL module, and a smoothing processing module. The information screening module well converts the complex interaction relationship between the host vehicle and surrounding vehicles in the environment into spatio-temporal interaction states and low-dimensional environmental information for expression. In the DRL module, the Actor network uses graph attention layers and multi-head attention layers, while the Critic network uses graph convolutional layers and long short-term memory layers to fully extract the feature information of the driving state, extract spatio-temporal features, and endow the network with the ability to analyze the spatio-temporal interaction features of the scene in real time. Finally, the smoothing processing module smooths the determined actions and transmits them back to the autonomous vehicle in the environment for motion control, and updates the state information of the environment. This method enables intelligent vehicles to accurately identify the complex traffic conflict relationships between the host vehicle and surrounding vehicles, and between surrounding vehicles and surrounding vehicles in the signal-free intersection scenario with complex traffic flow, make the current optimal strategy selection to complete the motion control task while ensuring safety, and improve the driving safety. Description of the Drawings:
[0070] Figure 1 It is the flowchart of the method of the present invention.
[0071] Figure 2 It is the application scenario diagram of the present invention.
[0072] Figure 3 It is the information screening module diagram of the present invention. Detailed Embodiments:
[0073] The present invention will be described in detail below with reference to the drawings.
[0074] The present invention relates to an end-to-end vehicle motion control method based on DRL for signal-free intersections. The method includes an environment, an information screening module, a DRL module, and a smoothing processing module. First, the information screening module obtains state information from the environment, refines it into low-dimensional environmental information considering the relative distance between the ego vehicle and surrounding vehicles and the start and end points of the task, and a spatio-temporal interaction state considering the spatio-temporal interaction information between the ego vehicle and surrounding vehicles in the form of a graph structure, forms a state space, and outputs the driving state from the state space to the DRL module. Secondly, the Actor network in the DRL module receives the driving state, extracts the spatial interaction features in the driving state using a graph attention layer, captures the temporal features in the driving state through a multi-head attention layer, uses a fully connected layer to splice the spatial interaction features, temporal features, and low-dimensional environmental information features, further maps the spliced state features to the mean and standard deviation of the action, and outputs the vehicle control quantity through Gaussian distribution reparameterization sampling. After the Critic network receives the driving state, it extracts the spatial interaction features and temporal features in the driving state using a graph convolutional layer and a long short-term memory layer respectively, and splices the spatial interaction features, temporal features, and low-dimensional environmental information features through a fully connected layer. Finally, the smoothing processing module smooths the vehicle control quantity, and the smoothed result is used as an action to be transmitted to the autonomous vehicle in the environment for motion control, and the state information of the environment is updated. Referring to Figure 1 the schematic diagram, it specifically includes the following steps:
[0075] Step 1. Construction of the environment:
[0076] In the present invention, the environment selects a two-way four-lane crossroads position on the map in CARLA, and the vehicle learns to turn left. As Figure 2 shown, the solid line in the figure is the reference trajectory of the ego vehicle, and the dotted lines are the reference trajectories of other traffic participants. The initial speed of the ego vehicle v0 = 15 km / h, and the initial speeds of other traffic participants are 15 - 18 km / h.
[0077] Step 2. Design of the reinforcement learning model:
[0078] Step 2.1. Design of the state space:
[0079] In a signal-free intersection, the behavior intentions of traffic participants from different directions are complex and changeable. An intelligent vehicle may interact with multiple traffic conflict objects simultaneously. Therefore, it needs to extract the state information of traffic conflict objects from the complex traffic environment and comprehensively consider the behavior intentions of all traffic conflict objects to dynamically adjust the motion control strategy to ensure traffic safety. Therefore, the present invention designs a state space containing spatio-temporal interaction states, enabling the intelligent agent to obtain the spatio-temporal interaction information between traffic participants and itself. In addition, a low-dimensional environmental information is set in it, considering the Euclidean distance between the ego vehicle and surrounding vehicles and the start and end points of motion control to supplement the environmental state. The state space is as shown in Equation (17):
[0080] S=(S st , S ld ) (17)
[0081] This state space contains the high-dimensional spatio-temporal interaction state S st and the low-dimensional environmental information S ld .
[0082] Step 2.2, Action space design:
[0083] In the present invention, the action space design of the intelligent vehicle refers to the performance parameters of the Tesla Model 3 standard version. The acceleration range is about -12 m / s 2 ~4.5 m / s 2 , the front wheel steering angle range is about -0.6 rad to 0.6 rad. The performance parameter range of the Tesla Model 3 standard version is mapped to the vehicle control quantity u, and u is set as the action space a. The lateral control quantity and the longitudinal control quantity in the vehicle control quantity are restricted within a threshold range to prevent the vehicle from performing overly dangerous actions. Therefore, the action space in reinforcement learning is defined as in Equation (18):
[0084] a = u = [α, δ], α0 ≤ α ≤ α1; δ0 ≤ δ ≤ δ1 (18)
[0085] Among them, α represents the longitudinal control quantity, which controls the stepping ratio of the accelerator and brake pedals. δ is the lateral control quantity, corresponding to the normalized control ratio of the front wheel steering angle. The control quantities are all subject to upper and lower limit constraints. Among them, α0 and α1 are respectively the minimum and maximum values of the longitudinal control, and the value range is [-1, 1]; δ0 and δ1 are the minimum and maximum values of the lateral control, and also take the values of [-1, 1].
[0086] Step 2.3, Reward function design:
[0087] In order to address the intelligent vehicle motion control task at unsignalized intersections, the reward function design comprehensively considers the safety and comfort of the intelligent vehicle, the traffic efficiency, the path keeping ability, and the road traffic rules. And at different stages of the task, different importance is given to different reward items. Therefore, a variable-weight reward function is set to emphasize the reward items that have a greater impact on that stage at different stages of the task. According to the environment in Step 1, the main settings of the reward function are as in Equation (19):
[0088] R = ω1r target + ω2r track + ω3r rule + ω4r v + ω5r safe (19)
[0089] Among them, rtarget is the target reward item, r track is the path deviation reward item, r rule is the rule reward item, r v is the speed reward item, r safe is the safety reward item.
[0090] From the perspective of vehicle interaction, the intelligent vehicle motion control task can be divided into three stages: before interaction, during interaction, and after interaction. In the pre-interaction stage, the intelligent vehicle does not need to pay too much attention to the behaviors of surrounding vehicles. Its primary task is the motion control to drive towards the target position. During this process, the intelligent vehicle should try to keep the vehicle trajectory consistent with the desired path, and the speed of the intelligent vehicle should also be maintained around the desired speed. At this time, the reward function should assign similar weights to r target , r track , r v for these three rewards. Therefore, ω1 = 0.3, ω2 = 0.4, ω3 = 1, ω4 = 0.3, ω5 = 0. In the interaction stage, the intelligent vehicle should give priority to ensuring driving safety. Therefore, it needs to pay more attention to r v , r safe . So, ω1 = 0, ω2 = 0, ω3 = 1, ω4 = 0.5, ω5 = 0.5. In the post-interaction stage, the intelligent vehicle needs to accelerate within the shortest time and return to the desired path faster. Therefore, the largest weight is given to r v , while r target , r track are assigned the same weight. So, ω1 = 0.3, ω2 = 0.3, ω3 = 1, ω4 = 0.4, ω5 = 0.
[0091] r target It is expected that the position of the intelligent vehicle continuously approaches the end position. Therefore, the sum of the horizontal and vertical coordinate deviations between the position of the intelligent vehicle and the end position is set as the deviation term between the intelligent vehicle and the end point. r target The specific expression of is as shown in Equation (20):
[0092] r target = b1 - c1(|x ego - x end | + |y ego - y end |) (20)
[0093] Among them, x ego , y ego are the coordinates of the host vehicle, x end and y end are the coordinates of the end point, and b1, c1 are constant terms.
[0094] r vThe purpose is to minimize the deviation between the ego-vehicle speed and the desired speed, and the desired speed is determined according to the scenario risk coefficient ε. The specific expressions are as shown in Eqs. (21) and (22):
[0095] r v = b2 - c2|v ego - v E | (21)
[0096]
[0097] where, v ego is the ego-vehicle speed, v E is the desired speed, v0 is the initial speed of the ego-vehicle, b2, c2 are constant terms, ttc is the time to collision, t1 is the maximum reaction time, which is set to 2.6 seconds.
[0098] r safe is the safety reward term, which is used to guide the intelligent vehicle to maintain a relatively safe distance from the traffic conflict object as much as possible. The specific expression is as shown in Eq. (23):
[0099] r safe = b3 - c3(s0 - s E ) (23)
[0100] where, s0 is the relative distance between the ego-vehicle and the traffic participant, s E is the desired safety distance. b3, c4 are constant terms.
[0101] r track is the path deviation reward term. The purpose of this reward is to minimize the deviation between the future trajectory of the intelligent vehicle and the reference trajectory. The specific expressions are as shown in Eqs. (24) and (25):
[0102]
[0103] r track = b2 - c4(Δ path ) (25)
[0104] where, Δ path is the path deviation, x A,j and y A,j are the horizontal and vertical coordinates of the j-th point on the predicted path, x path,j and y path,j are the horizontal and vertical coordinates of the corresponding reference path point, and are the yaw angles of the j-th point on the predicted path and the corresponding reference path point respectively, b2, c4 and c Δ are constant terms. The reference path is a series of coordinate points and the desired yaw angle corresponding to each horizontal and vertical coordinate. The ego-vehicle reference path is calculated as shown in Eq. (26):
[0105]
[0106] Among them, v x , v y is the component of the current vehicle speed in the x and y coordinates; a x , a y It is the component of the current vehicle acceleration in the x and y coordinates.
[0107] r rule is a rule reward item. When the smart car violates traffic rules such as collision, reverse driving, or driving out of the lane, the reward is triggered. The more steps the smart car takes when the event occurs, the greater the reward. The specific expression is as follows:
[0108] r rule =-(nn step ) (27)
[0109] Where n is the total number of steps, set to 400, n step is the number of steps taken by the car.
[0110] Step 3: Construction of information screening module:
[0111] The information screening module extracts the spatiotemporal interaction state and low-dimensional environmental information from the state information obtained from the environment to form a state space, and outputs the driving state from the state space to the DRL module. The spatiotemporal interaction state includes spatial interaction information and time information. The spatial interaction information is obtained in the form of graph structure data. Specifically, the traffic participants at each moment are modeled as nodes in the graph structure data, which contain four eigenvalues, namely, the horizontal coordinate, the vertical coordinate, the speed and the yaw angle, which are expressed as the node feature matrix X t ; The interaction between traffic participants is represented by the edges between nodes, which is represented as the adjacency matrix A. Then the node feature matrix X t And the adjacency matrix A is as shown in equations (28) and (29):
[0112]
[0113]
[0114] in, and is the coordinate of the vehicle at time t, is the vehicle speed at time t, is the yaw angle of the vehicle at time t. and is the coordinate of the first cycle at time t, is the speed of the first cycle car at time t, is the heading angle of the first cycle vehicle at time t. and are the coordinates of the nth cycle vehicle at time t, is the speed of the nth cycle vehicle at time t, is the heading angle of the nth cycle vehicle at time t. The adjacency matrix A is an n×n matrix, where n represents the number of vehicles participating in the interaction and is set to 5. Among them, if there is an interaction relationship between node i and node j, the element A ij in the matrix is 1, otherwise it is 0.
[0115] Secondly, a prediction equation is established with reference to the kinematic equation to obtain the time series of the motion state and determine the time information in the spatio-temporal interaction state, as follows: The node feature matrix X t provides the spatial interaction information required by the prediction equation, and takes the motion states of the cycle vehicle at the previous j moments to predict the motion state of the ego vehicle at the next k moments. Finally, the historical state of the cycle vehicle and the future state of the ego vehicle are concatenated to construct a complete time series of the motion state. The prediction equation is as shown in Equation (30):
[0116]
[0117] In the formula, and are the horizontal and vertical coordinates of the ego vehicle at time t, and are the components of the speed of the ego vehicle at time t in the x and y directions respectively, and are the components of the acceleration of the ego vehicle at time t in the x and y directions respectively, and d t is the prediction interval. In the previous j moments, j is set to 20, and in the next k moments, k is set to 10.
[0118] In summary, the spatio-temporal interaction state S st is determined according to the spatial interaction information and time information, expressed as Equation (31):
[0119] S st = ((X t-j ; A) … (X t ; A) … (X t+k ; A)) (31)
[0120] In the formula, X t-j , X t and X t+k are the node feature matrices at the previous j moments, time t, and the next k moments respectively; A is the adjacency matrix representing the interaction relationship between the ego vehicle and the cycle vehicle.
[0121] After obtaining the spatio-temporal interaction state, combined with the relative distance between the ego vehicle and the cycle vehicle and the start and end points of the task, the low-dimensional environmental information S is refinedld , such as Equation (32):
[0122] S ld = (l1…l n , x start , y start , x end , y end ) (32)
[0123] In the formula, l1…l n represents the relative distance between the host vehicle and each surrounding vehicle, with a total of n, and n is set to 4. x start and y start are the starting position coordinates of the task, and x end and y end are the ending position coordinates of the task.
[0124] Finally, the spatio-temporal interaction state S st and the low-dimensional environmental information S ld are integrated into the state space S, such as Equation (33):
[0125] S = (S st , S ld ) (33)
[0126] Step 4: Construction of the DRL module:
[0127] The DRL module includes an Actor network and a Critic network. Among them, the Actor network combines the driving state and outputs the vehicle control amount; the Critic network outputs the action state value according to the state characteristics of the driving state and the actions output by the smoothing processing module. The Actor network updates its parameters according to the action state value output by the Critic network, and the Critic network updates its parameters according to the rewards feedback by the environment.
[0128] The Actor network in the DRL module includes a graph attention layer, a multi-head attention layer, and a fully connected layer. Among them, the graph attention layer calculates the attention weights between nodes, assigns different weights to adjacent nodes, and then extracts the spatial interaction features of the driving state. The calculation of the attention weights is as shown in Equation (34):
[0129]
[0130] Among them, a ij is the attention weight between the i-th node h i and the j-th adjacent node h j , LeakyReLU is the activation function, a T is the attention vector, ω i and ω jis the linear transformation parameter matrix, and || represents vector concatenation. h i ' is the vehicle node feature after the attention weight update, and ω is the weight matrix.
[0131] The multi-head attention layer in the Actor network receives the spatial interaction features extracted by the graph attention layer, arranges the spatial interaction features in chronological order to form a time feature sequence, constructs query (Q), key (K), and value (V) matrices through linear transformation of the time feature sequence, then calculates the attention weights in parallel through multiple attention heads, and integrates the information of the attention weights calculated by each attention head to obtain the final multi-head attention output vector, thereby capturing the time features of the driving state; the linear transformation, attention weight calculation, and information integration are as shown in Equation (35):
[0132]
[0133] Among them, Q, K, and V are the query, key, and value matrices respectively, ω Q , ω K , ω V and ω o are the weight matrices, x is the time feature sequence, Attention(Q, K, V) is the calculation of a single attention head, softmax is the activation function, QK T represents the dot product of Q and K, d k is the dimension of the key vector K. MultiHead(Q, K, V) is the multi-head attention mechanism, head h is the h-th attention head, Concat(head1, head2,..., head h ) represents concatenating the outputs of multiple attention heads.
[0134] The fully connected layer in the Actor network concatenates the spatial interaction features output by the graph attention layer, the time features output by the multi-head attention layer, and the low-dimensional environmental information features in the driving state to construct state features, and maps the state features to the mean and standard deviation of the actions, and then outputs the vehicle control quantity through Gaussian distribution reparameterization sampling.
[0135] In the DRL module, the Critic network has relatively lower requirements for spatio-temporal feature extraction ability compared to the Actor network, and in order to save computing power costs and avoid unnecessary waste, the Critic network uses a graph convolutional layer and a long short-term memory layer to extract the spatial interaction features and time features of the driving state.
[0136] The Critic network in the DRL module includes a graph convolutional layer, a long short-term memory layer, and a fully connected layer. Among them, the graph convolutional layer constructs a self-connected adjacency matrix by adding the identity matrix I N to the adjacency matrix A Secondly, through graph convolution operation, the weight matrix W is utilized (l) to perform a linear transformation on the node feature matrix H of layer l (l) and combine it with the self-connected adjacency matrix for calculation to achieve the propagation and fusion of neighbor node information. Finally, through non-linear processing, the node feature matrix H of layer l (l) is updated to the node feature matrix H of layer l+1 containing its own and neighbor node features (l+1) , thereby extracting the spatial interaction features of the driving state. This process is formalized as Equation (36):
[0137]
[0138] where is the adjacency matrix with self-connection added, A is the adjacency matrix, I N is the identity matrix, H (l) W (l) is the linear transformation process, is the process of propagation and fusion of neighbor node information, σ(·) is the process of introducing non-linearity to the node feature matrix, H (l+1) and H (l) are the node feature matrices of the (l+1)-th and l-th layers, σ is the activation function, and W (l) is the weight matrix
[0139] The long short-term memory layer in the Critic network receives the spatial interaction features extracted by the graph convolution layer, arranges the spatial interaction features in chronological order to form a time feature sequence. This sequence passes through the gating mechanism of the long short-term memory network at each time step, jointly with the forget gate f t , input gate i t and output gate o t , to jointly control the update of the cell state c t and hidden state h t at the current time step, thereby achieving selective retention and integration of information. The hidden state h t output at each time step can be used as the time feature at that time step. Finally, to simplify the representation and aggregate historical information, the hidden state at the last time step in the sequence is selected as the time feature of the driving state. The specific state update process of the long short-term memory network is carried out according to the following steps
[0140] First, the time feature sequence x t at time t and the hidden state h t-1 at the previous time step are concatenated into a sequence, and the degree of information forgetting is calculated in the forget gate, as shown in Equation (37):
[0141] f t =σ(ω f [ht-1 , x t + b f ) (37)
[0142] Among them, f t is the forget gate, σ is the activation function, ω f is the weight matrix, h t-1 is the hidden state of the previous time step, x t is the time feature sequence, b f is the bias vector.
[0143] Secondly, calculate the information update degree in the input gate to generate the candidate information value, as shown in Equation (38):
[0144]
[0145] Among them, i t is the input gate, is the candidate cell state, σ is the activation function, tanh represents the hyperbolic tangent function, h t-1 is the hidden state of the previous time step, x t is the time feature sequence, b i , b c is the bias vector, ω i , ω c is the weight matrix.
[0146] Update the cell state according to the information forgetting degree and the information update degree, as shown in Equation (39):
[0147]
[0148] Among them, c t is the cell state of the current time step, f t is the forget gate, c t-1 is the cell state of the previous time step, i t is the input gate, is the candidate cell state.
[0149] Finally, the output gate outputs the hidden state according to the cell state, as shown in Equation (40):
[0150]
[0151] Among them, o t is the output gate, h t represents the hidden state of the current time step. h t-1 is the hidden state of the previous time step, x t is the time feature sequence, σ is the activation function, tanh represents the hyperbolic tangent function, ω k is the weight matrix, b ois the bias vector, c t is the cell state at the current time step.
[0152] The fully connected layer in the Critic network concatenates the spatial interaction features output by the graph convolutional layer, the temporal features output by the long short-term memory layer, and the low-dimensional environmental information features in the driving state to determine the state features. After the action is passed to the Critic network in the smoothing processing module, the action is evaluated in combination with the state features of the driving state, and the action state value is output.
[0153] Step 5: Construction of the smoothing processing module:
[0154] The smoothing processing module uses the moving average method to smooth the vehicle control amount output by the DRL module, and transfers the smoothed action to the autonomous vehicle in the environment for motion control to update the state information of the environment. At the same time, it is passed to the Critic network in the DRL module, so that it evaluates the action according to the state features of the driving state and outputs the action state value. The moving average method is as shown in Equation (41):
[0155]
[0156] where and are the new longitudinal and lateral control amounts, α t + α t-1 + … + α t-n+1 and δ t + δ t-1 + … + δ t-n+1 are the longitudinal and lateral control amounts of the Actor network at the past n time steps, and n is set to 10.
[0157] Step 6: Update of network parameters in the DRL module:
[0158] Step 6.1: Selection of the reinforcement learning algorithm:
[0159] In order to extract the spatio-temporal interaction features of the driving state in the neural network and complete the motion control task of the intelligent vehicle in the environment without signal intersections, the present invention uses the Soft Actor-Critic algorithm as the reinforcement learning algorithm for the intelligent vehicle. The Soft Actor-Critic algorithm can adapt to the continuous action space and high sample efficiency. By maximizing the entropy and the expected reward, the policy learned by the agent can be better and more random, which helps to improve the generalization of the control policy.
[0160] Step 6.2: Sampling of experience samples:
[0161] After one round of loop, the Actor and Critic networks in the DRL module update their network parameters to optimize the policy. The Actor network is updated based on the action-state value output by the Critic network, and the Critic network is updated based on the reward feedback from the environment after the action is selected. The update requires randomly sampling a batch of experience samples from the experience pool and distributing them to each network.
[0162] Step 6.3, Experience Replay:
[0163] Split the quadruple [s, a, r, s - given in the randomly sampled experience samples, and respectively form vectors of state s, action a, reward r, and next state s - and transmit them into the Actor network and the Critic network for experience replay. Calculate the loss function Loss of each network through the quadruple, and use the gradient descent method for backpropagation to update the parameters, so as to learn the updated safety control policy and better complete the motion control task. The loss function of the Actor network and gradient descent are as shown in Equations (42) and (43): φ ← φ + η▽ φ J π (φ) (43)
[0164] where J π (φ) is the loss function of the Actor network, φ is the network parameter; E is the expected value, S ∼ D means that the driving state S is sampled from the state distribution D. a ∼ π φ means that the action a is sampled from the driving state S according to the policy π φ . α is the weight of entropy, log(π φ (a|s)) is the logarithmic probability of the policy π φ selecting the action a in the driving state S. is the value function estimate of the Critic network, and the minimum value of the two Critic networks is taken as the estimate of the value function. represents the expected return of taking the action a in the driving state S. η is the learning rate, and ▽ φ is the gradient vector of the parameter φ.
[0165] The loss function of the Critic network and gradient descent are as shown in Equations (44) and (45): θ ← θ - α θ ▽ θ J Q (θ)(45)
[0166] Among them, J Q (θ) is the loss function of the Critic network, and θ is the network parameter; E is the expected value, and (s,a)~D means that the driving state S and the action a are sampled from the state distribution D. is the expected return for taking the action a in the driving state S. r is the immediate reward, and γ is the discount factor, which is used to balance the importance of the immediate reward and the future reward. is to select the minimum Q value from the two Critic networks to reduce the variance of the estimation, and θ j ' is the parameter of the target Critic network. α is the weight of the entropy, and logπ(a - |s - ) is the log probability that the Actor network selects the next action a - in the next driving state s - . α θ is the learning rate of the Critic network, and ▽ θ is the gradient vector of the parameter θ.
[0167] In summary: The present invention proposes an end-to-end vehicle motion control method based on DRL, which can effectively extract the spatio-temporal features of the driving state, and has the ability to infer the behavior intentions of traffic conflict objects from a complex traffic environment, handle the complex interaction relationships between vehicles, and then dynamically adjust the motion control strategy to ensure traffic safety; it improves the generalization and safety of the motion control strategy in the environment of signal-free intersections.
Claims
1. An end-to-end vehicle motion control method for signal-free intersections based on DRL, characterized in that: The method includes an environment, an information screening module, a DRL module, and a smoothing processing module; first, the information screening module obtains state information from the environment and refines it into low-dimensional environmental information and spatio-temporal interaction states, forms a state space, and outputs the driving state from the state space to the DRL module; second, the Actor network in the DRL module receives the driving state, extracts the spatial interaction features in the driving state using a graph attention layer, captures the temporal features in the driving state through a multi-head attention layer, uses a fully connected layer to splice the spatial interaction features, temporal features, and low-dimensional environmental information features, and further maps them to the mean and standard deviation of the action, and outputs the vehicle control quantity through Gaussian distribution reparameterization sampling; after the Critic network receives the driving state, it extracts the spatial interaction features and temporal features in the driving state using a graph convolutional layer and a long short-term memory layer respectively, and splices the spatial interaction features, temporal features, and low-dimensional environmental information features through a fully connected layer; finally, the smoothing processing module smooths the vehicle control quantity, transfers the action to the autonomous vehicle in the environment for motion control, and updates the state information of the environment at the same time.
2. The end-to-end vehicle motion control method based on DRL for signal-free intersections according to claim 1, wherein The described information screening module extracts spatio-temporal interaction states and low-dimensional environmental information from the state information obtained from the environment to form a state space. The described spatio-temporal interaction state includes spatial interaction information and time information; among them, the spatial interaction information is obtained in the form of graph structure data, specifically as follows: traffic participants at each moment are modeled as nodes in the graph structure data, including four eigenvalue, namely the abscissa, ordinate, speed and yaw angle of the traffic participant, which are expressed as the node feature matrix X t ; the interaction relationship between traffic participants is represented by the edges between nodes, which is expressed as the adjacency matrix A; the node feature matrix X t , the adjacency matrix A are as shown in formulas (1) and (2): Among them, and are the horizontal and vertical coordinates of the host vehicle at time t, is the speed of the host vehicle at time t, is the yaw angle of the host vehicle at time t; and are the horizontal and vertical coordinates of the first surrounding vehicle at time t, is the speed of the first surrounding vehicle at time t, is the heading angle of the first surrounding vehicle at time t; and are the horizontal and vertical coordinates of the nth surrounding vehicle at time t, is the speed of the nth surrounding vehicle at time t, is the heading angle of the nth surrounding vehicle at time t; The adjacency matrix A is an n×n matrix, where n represents the number of vehicles participating in the interaction; among them, if there is an interaction relationship between node i and node j, the element value in the matrix is 1, otherwise it is 0; Secondly, a prediction equation is established with reference to the kinematic equation to obtain the time series of the motion state and determine the time information in the spatio-temporal interaction state, which is specifically as follows: the node feature matrix X t Provide the spatial interaction information required for the prediction equation at time t, and take the motion states of the ego vehicle at the previous j moments of the surrounding vehicles to predict the motion states of the ego vehicle at the next k moments, thereby obtaining the time series of the motion state; the prediction equation is as shown in Equation (3): Among them, and are the horizontal and vertical coordinates of the host vehicle at time t, and are the components of the host vehicle's velocity in the x and y directions at time t, respectively, and are the components of the host vehicle's acceleration in the x and y directions at time t, respectively, and d t is the prediction interval; In summary, the spatio-temporal interaction state S is determined based on the spatial interaction information and the time information st , which is expressed as Equation (4): S st = ((X t-j ; A)…(X t ; A)…(X t+k ; A)) (4) Among them, X t-j , X t and X t+k are the node feature matrices at the historical j moments, the t moment, and the future k moments respectively, and A is the adjacency matrix representing the interaction relationship between the host vehicle and the surrounding vehicles; After obtaining the spatio-temporal interaction state, combined with the relative distance between the host vehicle and surrounding vehicles, as well as the task start and end points, the low-dimensional environmental information S is refined ld , as shown in Equation (5): S ld =(l1…l n ,x start ,y start ,x end ,y end ) (5) Among them, l1…l n represent the relative distances between the host vehicle and each surrounding vehicle, with a total of n; x start and y start are the abscissa and ordinate of the starting position of the task, x end and y end are the abscissa and ordinate of the ending position of the task; Finally, the spatio-temporal interaction state and the low-dimensional environmental information are integrated into the state space S, as shown in Equation (6): S = (S st , S ld ) (6) Among them, S st is the spatio-temporal interaction state, and S ld is the low-dimensional environmental information.
3. The end-to-end vehicle motion control method based on DRL for signal-free intersections according to claim 1, wherein, The described DRL module includes an Actor network and a Critic network; among them, the Actor network combines the driving state and outputs the vehicle control quantity; the Critic network outputs the action state value according to the state characteristics of the driving state and the action output by the smoothing processing module; the Actor network updates the parameters according to the action state value output by the Critic network, and the Critic network updates the parameters according to the reward feedback from the environment. The Actor network in the described DRL module includes a graph attention layer, a multi-head attention layer, and a fully connected layer; among them, the graph attention layer calculates the attention weights between nodes, assigns different weights to adjacent nodes, and extracts the spatial interaction features of the driving state. The attention weight calculation is as shown in Equation (7): where a ij is the attention weight between the i-th node h i and the j-th adjacent node h j , LeakyReLU is the activation function, a T is the attention vector, ω i and ω j are the linear transformation parameter matrices, || represents vector concatenation; h i ' is the vehicle node feature after the attention weight update, and ω is the weight matrix; The multi-head attention layer in the Actor network receives the spatial interaction features extracted by the graph attention layer, arranges the spatial interaction features in chronological order to form a time feature sequence, performs a linear transformation on the time feature sequence, constructs query Q, key K, and value V matrices, then calculates the attention weights in parallel through multiple attention heads, and integrates the attention weights calculated by each attention head to obtain the final multi-head attention output vector, thereby capturing the temporal features of the driving state; the linear transformation, attention weight calculation, and information integration are as shown in Equation (8): where Q, K, and V are the query, key, and value matrices respectively, ω Q , ω K , ω V , and ω o are weight matrices, x is the time feature sequence, Attention(Q, K, V) is the calculation of a single attention head, softmax is the activation function, QK T represents the dot product of Q and K, d k is the dimension of the key vector K; MultiHead(Q, K, V) is the multi-head attention mechanism, head h is the h-th attention head, Concat(head1, head2,..., head h ) represents concatenating the outputs of multiple attention heads; The fully connected layer in the Actor network splices the spatial interaction features output by the graph attention layer, the temporal features output by the multi-head attention layer, and the low-dimensional environmental information features in the driving state to construct state features, maps the state features to the mean and standard deviation of the action, and then outputs the vehicle control quantity through Gaussian distribution reparameterization sampling. The Critic network in the DRL module described above includes a graph convolutional layer, a long short-term memory layer, and a fully connected layer; among them, the graph convolutional layer adds the identity matrix I through the adjacency matrix A N Construct a self-connected adjacency matrix Secondly, through graph convolutional operations, using the weight matrix W (l) For the node feature matrix H of layer l (l) Perform a linear transformation and combine it with the self-connected adjacency matrix Perform calculations to achieve the propagation and fusion of neighbor node information, and finally update the node feature matrix H of layer l through non-linear processing (l) To the node feature matrix H of layer l+1 containing its own and neighbor node features (l+1) , thereby extracting the spatial interaction features of the driving state. This process is formalized as Equation (9): Among them, is the adjacency matrix with self-connection added, A is the adjacency matrix, and I N is the identity matrix, and H (l) W (l) is the linear transformation process, is the process of propagation and fusion of neighbor node information, σ(·) is the process of introducing nonlinearity into the node feature matrix, and H (l+1) and H (l) are the node feature matrices of the (l + 1)-th and l-th layers, σ is the activation function, and W (l) is the weight matrix; The long short-term memory layer in the Critic network receives the spatial interaction features extracted by the graph convolutional layer, arranges the spatial interaction features in chronological order to form a time feature sequence; at each time step of this sequence, through the gating mechanism of the long short-term memory network, jointly with the forget gate f t , input gate i t and output gate o t , jointly control the update of the cell state c t and hidden state h t , so as to achieve selective retention and integration of information; the hidden state h t at each time step of the output can be used as the time feature of that time step; finally, to simplify the representation and aggregate historical information, the hidden state of the last time step in the sequence is selected as the time feature of the driving state; the state update process in the long short-term memory network is as shown in Equation (10): where h t-1 is the hidden state of the previous time step, x t is the time feature sequence, σ is the activation function, tanh represents the hyperbolic tangent function, ω f 、ω i 、ω c and ω k are weight matrices, b f 、b i 、b c and b o are bias vectors; c t-1 is the cell state of the previous time step, f t is the forget gate, i t is the input gate, is the candidate cell state, c t is the cell state of the current time step, o t is the output gate, h t represents the hidden state of the current time step; The fully connected layer in the Critic network concatenates the spatial interaction features output by the graph convolutional layer, the temporal features output by the long short-term memory layer, and the low-dimensional environmental information features in the driving state to determine the state features; after the action is passed to the Critic network in the smoothing processing module, the action is evaluated in combination with the state features of the driving state, and the action state value is output. After one round of loop, the Actor network updates its network parameters according to the action state value output by the Critic network, and the Critic network updates its network parameters according to the reward feedback from the environment; for the intelligent vehicle motion control task at the unsignalized intersection, the present invention comprehensively considers the safety and comfort, traffic efficiency, path keeping ability of the intelligent vehicle and the road traffic rules to design the reward function, as shown in Equations (11) to (13): where r target is the target reward item, r track is the path deviation reward item, r rule is the rule reward item, r v is the speed reward item, r safe is the safety reward item, ω1, ω2, ω3, ω4, and ω5 are reward distribution weights; x ego , y ego are the horizontal and vertical coordinates of the ego vehicle, x end and y end are the horizontal and vertical coordinates of the end point of the ego vehicle, b1, b2, b3, c1, c2, c3, and c4 are constant terms; Δ path is the path deviation, n is the total number of steps, n step is the number of steps the vehicle has traveled; v ego is the speed of the ego vehicle, v E is the desired speed, v0 is the initial speed of the ego vehicle, ttc is the time to collision, t1 is the maximum reaction time, s0 is the relative distance between the ego vehicle and traffic participants, s E is the desired safety distance; x A,j and y A,j are the horizontal and vertical coordinates of the j-th point on the predicted path, x path,j and y path,j are the horizontal and vertical coordinates of its corresponding reference path point, and are the yaw angles of the j-th point on the predicted path and the yaw angle of its corresponding reference path point respectively, c Δ is a constant term; The smoothing processing module uses the moving average method to smooth the vehicle control quantity output by the DRL module, and transfers the smoothed action to the autonomous vehicle in the environment for motion control to update the state information of the environment; at the same time, it is transferred to the Critic network in the DRL module, so that it evaluates the action according to the state features of the driving state and outputs the action state value; the moving average method is used for smoothing processing, and the action space is as shown in Equations (14) and (15): a = u = [α, δ], α0 ≤ α ≤ α1; δ0 ≤ δ ≤ δ1 (15) Among them, and are new longitudinal and lateral control quantities, α t +α t-1 +…+α t-n+1 and δ t +δ t-1 +…+δ t-n+1 are the longitudinal and lateral control quantities of the Actor network at the past n moments; a is the action space, u is the vehicle control quantity, where δ is the lateral control quantity; α is the longitudinal control quantity; α0, α1, δ0, δ1 are the set constraint thresholds.
Citation Information
Patent Citations
Intensive learning based urban intersection passing method for driverless vehicle
CN108932840A
Traffic adaptive control method based on multi-agent reinforcement learning
CN118155429A
Unmanned vehicle driving decision-making method based on attention model and deep reinforcement learning
CN112965499A
Non-signalized intersection vehicle passing decision planning method, system and equipment
CN118212808A
Visual end-to-end vehicle following method based on deep reinforcement learning
CN118348988A