Automatic driving behavior decision-making method fusing time sequence Transform and safety constraint

By integrating timing Transformer and safety constraints, an autonomous vehicle intelligent body is built and a regularization mechanism is introduced, which solves the safety and stability problems of reinforcement learning algorithms in the autonomous driving system, and realizes safe and stable decisions in complex traffic environments.

CN120348305APending Publication Date: 2025-07-22SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510758429.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing reinforcement learning algorithms have problems such as safety risks and low training efficiency in autonomous driving systems, especially in complex traffic environments, which are difficult to ensure the safety and stability of decisions.

Method used

The autonomous driving behavior decision-making method that integrates timing Transformer and safety constraints is used to construct a hybrid traffic scenario, use IDM and MOBIL models to simulate human driving behavior, design the state space, action space, reward function and cost function of the autonomous driving vehicle intelligent body, and introduce the action regularization mechanism to establish an overall optimization goal, use the Transformer model for training, and deploy the reinforcement learning model for decision-making.

Benefits of technology

It improves the safety of the autonomous driving system in complex environments, enhances the stability under uncertain conditions, reduces decision-making errors, and improves information capture capabilities and decision-making effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120348305A_ABST
    Figure CN120348305A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving behavior decision-making method fusing time sequence Transform and safety constraints, and relates to the technical field of automatic driving decision-making. The method comprises the following steps: constructing an automatic driving mixed traffic scene, wherein the mixed traffic scene comprises an automatic driving vehicle and a man-driven vehicle; a man-driven vehicle model is constructed for a man-driven vehicle in a mixed traffic scene, an IDM model is adopted to control a longitudinal decision, and an MOBIL model is adopted to control a transverse decision. According to the invention, the cost constraint is introduced, so that the automatic driving system can avoid potential dangerous behaviors in a complex environment, and the safety is improved; an action regularization item is introduced into a strategy loss function, so that the stability of the system under an uncertain condition is enhanced, and decision errors caused by external interference are reduced; and meanwhile, a transform time sequence model is introduced to design a network structure, so that the information capturing capability of the model is enhanced, and the decision effectiveness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving decision-making, and specifically to an autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints. Background Technique

[0002] With the rapid development of artificial intelligence technology, autonomous driving technology has gradually become an important research direction in the field of transportation. The core technologies of an autonomous driving system include perception, decision-making, and control, among which the decision-making module is responsible for making driving decisions based on perception data. Currently, reinforcement learning (RL), as an effective decision-making method, has been widely applied in the field of autonomous driving. Reinforcement learning continuously optimizes the decision-making strategy through interaction with the environment, enabling the autonomous driving system to autonomously learn and make optimal decisions.

[0003] However, in practical applications, existing reinforcement learning algorithms have certain limitations. First, the training process of reinforcement learning usually requires a large amount of interaction data. In the context of autonomous driving, the cost of data collection is extremely high, and during the training process, wrong decisions may lead to serious safety risks. Second, existing reinforcement learning algorithms mainly focus on decision-making efficiency and performance optimization, while ignoring the safety of the system. Especially when facing complex traffic environments and unforeseen emergencies, how to ensure the safety and stability of decisions is an urgent problem to be solved. Therefore, the present invention proposes an autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints. Summary of the Invention

[0004] The purpose of the present invention is to provide an autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints, and solve the problems of safety risks and low training efficiency existing in the application of existing reinforcement learning methods in autonomous driving systems.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints, including:

[0006] Construct an autonomous driving mixed traffic scenario, which includes autonomous driving vehicles and human-driven vehicles;

[0007] Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making, and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior;

[0008] Build an autonomous vehicle agent for autonomous vehicles in a mixed traffic scenario, and design the state space, action space, reward function, and cost function of the autonomous vehicle agent;

[0009] Based on the designed state space, action space, reward function, and cost function of the autonomous vehicle agent, further design the overall optimization goal, introduce an action regularization mechanism, and establish the overall objective equation;

[0010] Build a reinforcement learning model based on the Transformer model, learn the overall optimization goal set by the overall objective equation, and use it to train the reinforcement learning model to obtain the trained reinforcement learning model;

[0011] Deploy the trained reinforcement learning model to the vehicle to achieve autonomous driving behavior decision-making.

[0012] Furthermore, the autonomous driving mixed traffic scenario is built using the highway-env toolkit, and the autonomous driving mixed traffic scenario is specifically the highway off-ramp merging scenario.

[0013] Furthermore, build a human-driven vehicle model for human-driven vehicles in a mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, specifically including:

[0014] (31) Control the longitudinal decision-making through the IDM model, and the control formula is as follows:

[0015]

[0016] In the formula is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle in front, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle in front, and T is the safe time headway of the vehicle;

[0017] (32) Control the lateral decision-making through the MOBIL model, and the control formula is as follows:

[0018]

[0019] Where is the acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n is the acceleration of this vehicle before lane change, is the predicted acceleration of the current vehicle after lane change, a c is the acceleration of the current vehicle before lane change, is the acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o is the acceleration of the vehicle before lane change, Δa th is the lane change threshold, b safe is the safe deceleration threshold.

[0020] Furthermore, the state space and action space of the autonomous vehicle agent are as follows:

[0021] (41) The state space of the autonomous vehicle agent at time step t is s t ={s t,1 , s t,2 , s t,3 ,…, s t,N}, where N is the number of vehicles in the autonomous driving scenario, and the state of each vehicle is a time series Each time series state unit where present represents the perceivable state of the i-th vehicle k steps ago, and represent the x and y coordinates of the position, v x , v y represent the x and y coordinates of the speed;

[0022] (42) The action space is A = {a1, a2, a3, a4, a5}, where a1, a2, a3, a4, a5 represent 5 discrete actions of turning left, turning right, staying the same, accelerating, and decelerating respectively.

[0023] Furthermore, the reward function of the autonomous vehicle agent is as follows:

[0024] R(s, a) = w v ·r v + w cl ·r cl + w sc ·r sc

[0025] where r v is the speed reward, r cl is the successful lane change reward, r sc is the reward for successfully reaching the end, w v , w cl , w sc are their corresponding weight parameters, and the specific formulas are as follows:

[0026]

[0027] where v ego is the current speed of the autonomous vehicle, is the average speed of surrounding vehicles.

[0028] Furthermore, the cost function of the autonomous vehicle agent is specifically as follows:

[0029] C(s,a) = w cr ·c cr +w wa ·c wa +w pr ·c pr

[0030] where c cr is the collision cost, c wa is the cost of an inappropriate action, c pr is the prediction cost, w cr , w wa , w pr are their corresponding weight parameters, and the specific formulas are as follows:

[0031]

[0032] where d min is the distance between the autonomous vehicle and the nearest surrounding vehicle, d safe is the safety distance, is the distance cost between the ego vehicle and surrounding vehicles after executing an action, is the speed difference cost between the ego vehicle and surrounding vehicles after executing an action, c pc is the cost of not changing lanes for a long time, w cr , w vr , w pc are their corresponding weight parameters, The specific design of c pc is as follows:

[0033]

[0034] where v thr is the safety speed threshold.

[0035] Furthermore, based on the state space, action space, reward function, and cost function of the autonomous vehicle agent designed above, the overall optimization goal is further designed, and an action regularization mechanism is introduced to establish the overall objective equation, which is specifically as follows:

[0036] (71) The overall objective equation introduces constraints on the basis of traditional reinforcement learning, and the specific equation is:

[0037]

[0038] where R(s t ,a t ) is the agent reward function, C(st , a t ) is the agent cost function, s t is the state observation corresponding to time step t, λ is the adjustment factor, γ is the discount factor with a value range of 0 to 1, used to balance current and future rewards and costs, and the overall goal is to maximize the obtained rewards while minimizing the driving cost;

[0039] (72) The overall goal equation introduces an action regularization mechanism on the basis of traditional reinforcement learning, and the specific expression is:

[0040]

[0041] where π φ is the agent policy network, represents the introduced noise with a mean of 0 and a variance of σ, and s represents the current state observation of the autonomous vehicle.

[0042] Further, the reinforcement learning model based on the Transformer model includes a policy network, two evaluation networks, two evaluation target networks, a cost network, and a cost target network, a total of 7 components based on the sequential Transformer:

[0043] Among them, the input of the policy network is the sequential observation state sequence, which captures sequential dependencies through the self-attention mechanism and outputs the action probability distribution at the current moment; the evaluation network adopts a dual-network structure to alleviate overestimation, inputs the joint embedding of the sequential state and action, and predicts the Q value of the state-action pair; the evaluation target network is a copy of the evaluation network with delayed update, providing a stable Q value calculation target; the cost network inputs the joint representation of the sequential state and action, predicts the action safety cost, and constrains dangerous behaviors; the cost target network is a copy of the cost network with delayed update, used for calculating the target value of the cost function;

[0044] The loss functions in the reinforcement learning model are as follows respectively:

[0045]

[0046] where L π is the loss function of the policy network, π φ is the current policy to be updated, H(π φ ) is the entropy of the policy, s is the observed state sequence of the current autonomous vehicle, a is the decision made by the current policy, Q(s, a) is the minimum value of the two evaluation networks Q1(s, a) and Q2(s, a), R reg (a) is the action regularization term, (0, σ 2 ) is the Gaussian noise distribution, is the loss function of the evaluation network, R nis the reward TD error obtained through the target network, L C is the loss function of the cost network, C n is the cost TD error obtained through the target network, L λ is the loss function of the cost adjustment factor of, C max is the maximum cost threshold.

[0047] Furthermore, the training process of the autonomous driving vehicle agent is as follows:

[0048] (91) Initialize the policy network π of the reinforcement learning model φ , the evaluation networks Q1(s,a) and Q2(s,a) and their corresponding target networks, and the cost network C ψ (s,a) and their corresponding target networks and environmental parameters;

[0049] (92) At each time step t, perform state observation to obtain a state sequence, make an action selection, obtain action feedback, and store data;

[0050] (93) Determine whether the network update standard is reached, that is, a certain amount of data for update has been stored. If so, execute step (94); otherwise, execute step (92);

[0051] (94) Calculate the value and cost of each state-action pair through the evaluation network and the cost network, and update the policy network according to the formula Update the policy network;

[0052] (95) Update the evaluation network and the cost network according to the rewards and costs in the stored data according to the formula

[0053] Update the evaluation network and the cost network;

[0054] (96) Update the cost adjustment factor through the estimation of the cost network;

[0055] (97) Determine whether the reinforcement learning model has converged, that is, the rewards and costs of the network tend to be smooth and stable. If so, end the training process; otherwise, execute (92).

[0056] According to the second aspect of the present invention, the present invention provides an autonomous driving behavior decision-making system integrating a temporal Transformer and safety constraints, for implementing the autonomous driving behavior decision-making method integrating a temporal Transformer and safety constraints, including:

[0057] A scenario construction module for constructing an autonomous driving mixed traffic scenario, where the mixed traffic scenario includes autonomous driving vehicles and human-driven vehicles;

[0058] The human-driven vehicle model construction module is used to construct a human-driven vehicle model for human-driven vehicles in a mixed traffic scenario, where the IDM model is adopted to control its longitudinal decision-making, and the MOBIL model is adopted to control its lateral decision-making;

[0059] The agent construction module is used to construct an autonomous driving vehicle agent for autonomous driving vehicles in a mixed traffic scenario, and design the state space, action space, reward function and cost function of the autonomous driving vehicle agent;

[0060] The design module is used to further design the overall optimization goal based on the designed state space, action space, reward function and cost function of the autonomous driving vehicle agent, introduce an action regularization mechanism, and establish an overall objective equation;

[0061] The training module is used to construct a reinforcement learning model based on the Transformer model, learn the overall optimization goal set by the overall objective equation, and train the reinforcement learning model to obtain a trained reinforcement learning model;

[0062] The decision-making module is used to deploy the trained reinforcement learning model to the vehicle to achieve autonomous driving behavior decision-making.

[0063] The present invention at least has the following beneficial effects:

[0064] 1. By introducing cost constraints, the present invention enables the autonomous driving system to avoid potential dangerous behaviors in complex environments and improve safety; and by introducing an action regularization term into the policy loss function, the stability of the system under uncertain conditions is enhanced, and decision-making errors caused by external interference are reduced; at the same time, by introducing a transformer time-series model to design the network structure, the model's ability to capture information is enhanced, and the effectiveness of decision-making is improved.

[0065] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a flowchart of the decision-making method described in the present invention;

[0067] Figure 2 It is a structural diagram of the decision-making method described in the present invention;

[0068] Figure 3 It is a three-dimensional schematic diagram of the structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0070] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints, including the following steps:

[0071] S1. Construct an autonomous driving mixed traffic scenario, which includes autonomous vehicles and human-driven vehicles;

[0072] For the technical solution of this embodiment, this embodiment is built using the highway-env toolkit, and a mixed traffic scenario is established, specifically the on-ramp merging scenario on a highway;

[0073] S2. Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making to simulate human driving behavior, specifically including:

[0074] (S21) Control the longitudinal decision-making through the IDM model, and the control formula is as follows:

[0075]

[0076] In the formula is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle in front, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle in front, and T is the safe time headway of the vehicle;

[0077] (S22) Control the lateral decision-making through the MOBIL model, and the control formula is as follows:

[0078]

[0079] Where is the acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n is the acceleration of this vehicle before lane change, is the predicted acceleration of the current vehicle after lane change, a c is the acceleration of the current vehicle before lane change, The acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o The acceleration of the vehicle before lane change, Δa th The lane change threshold, b safe Is the safe deceleration threshold;

[0080] S3. Construct an autonomous vehicle agent for autonomous vehicles in a mixed traffic scenario, and design the state space, action space, reward function, and cost function of the autonomous vehicle agent;

[0081] (S31) The state space of the autonomous vehicle agent at time step t is s t ={s t,1 , s t,2 , s t,3 ,…, s t,N} where N is the number of vehicles in the autonomous driving scenario, and the state of each vehicle is a time series Each time series state unit where present represents the perceivable state of the i-th vehicle k steps ago, and represent the x and y coordinates of the position, v x , v y represent the x and y coordinates of the speed;

[0082] (S32) The action space is A = {a1, a2, a3, a4, a5}, where a1, a2, a3, a4, a5 represent the 5 discrete actions of turning left, turning right, staying the same, accelerating, and decelerating respectively;

[0083] (S33) The reward function of the autonomous vehicle agent is as follows:

[0084] R(s, a) = w v ·r v + w cl ·r cl + w sc ·r sc

[0085] where r v is the speed reward, r cl is the successful lane change reward, r sc is the reward for successfully reaching the end point, w v , w cl , w sc are their corresponding weight parameters, and the specific formulas are as follows:

[0086]

[0087] where vego is the current speed of the autonomous vehicle, is the average speed of surrounding vehicles;

[0088] (S34) The cost function of the autonomous vehicle agent is as follows:

[0089] C(s,a) = w cr ·c cr +w wa ·c wa +w pr ·c pr

[0090] where c cr is the collision cost, c wa is the cost of inappropriate actions, c pr is the prediction cost, w cr , w wa , w pr are their corresponding weight parameters, and the specific formulas are as follows:

[0091]

[0092]

[0093] where d min is the distance between the autonomous vehicle and the nearest surrounding vehicle, d safe is the safety distance, is the distance cost between the autonomous vehicle (ego vehicle) and surrounding vehicles after executing an action, is the speed difference cost between the autonomous vehicle (ego vehicle) and surrounding vehicles after executing an action, c pc is the cost of not changing lanes for a long time, w cr , w vr , w pc are their corresponding weight parameters, The specific design of c pc is as follows:

[0094]

[0095] where v thr is the safety speed threshold;

[0096] S4. Based on the designed state space, action space, reward function, and cost function of the autonomous vehicle agent, further design the overall optimization goal, introduce an action regularization mechanism, and establish the overall objective equation;

[0097] (S41) The overall objective equation introduces constraints based on traditional reinforcement learning, and the specific equation is:

[0098]

[0099] where \(R(s t , a t )\) is the agent reward function, \(C(s t , a t )\) is the agent cost function, \(s t \) is the state observation corresponding to time step \(t\), \(\lambda\) is the adjustment factor, \(\gamma\) is the discount factor with a value range from \(0\) to \(1\), which is used to balance current and future rewards and costs. The overall objective is to maximize the obtained rewards while minimizing the driving cost;

[0100] (S42) The overall objective equation introduces an action regularization mechanism on the basis of traditional reinforcement learning, and the specific expression is:

[0101]

[0102] where \(\pi φ \) is the agent policy network, \) represents the introduced noise with a mean of \(0\) and a variance of \(\sigma\), and \(s\) represents the current state observation of the autonomous vehicle;

[0103] S5. Construct a reinforcement learning model based on the Transformer model to learn the overall optimization objective set by the overall objective equation, and use it to train the reinforcement learning model to obtain the trained reinforcement learning model;

[0104] (S51) The reinforcement learning model includes a policy network, two evaluation networks, two evaluation target networks, a cost network, and a cost target network, a total of 7 components based on the temporal Transformer:

[0105] Among them, the input of the policy network is the temporal observation state sequence, which captures temporal dependencies through the self-attention mechanism and outputs the action probability distribution at the current moment; the evaluation network adopts a dual-network structure to alleviate overestimation, inputs the joint embedding of the temporal state and action, and predicts the state-action pair \(Q\) value; the evaluation target network is a copy of the evaluation network with delayed update, providing a stable \(Q\) value calculation target; the cost network inputs the joint representation of the temporal state and action, predicts the action safety cost, and restricts dangerous behaviors; the cost target network is a copy of the cost network with delayed update, which is used for the target value calculation of the cost function.

[0106] Furthermore, the loss functions of each network in the reinforcement learning model are as follows:

[0107]

[0108] where \(L π \) is the loss function of the policy network, \(\pi φ \) is the current policy to be updated, \(H(\piφ ) is the entropy of the policy, s is the observed state sequence of the current autonomous vehicle, a is the decision made by the current policy, Q(s,a) is the minimum value of the two evaluation networks Q1(s,a) and Q2(s,a), and R reg (a) is the action regularization term, (0, σ 2 ) is the Gaussian noise distribution, is the loss function of the evaluation network, R n is the reward TD error obtained through the target network, L C is the loss function of the cost network, C n is the cost TD error obtained through the target network, L λ is the loss function of the cost adjustment factor of C max is the maximum cost threshold.

[0109] (S52) The training process of the reinforcement learning model is as follows:

[0110] (S52.1) Initialize the policy network π of the reinforcement learning model φ , the evaluation networks Q1(s,a) and Q2(s,a) and their corresponding target networks, and the cost network C ψ (s,a) and its corresponding target network and environmental parameters;

[0111] (S52.2) At each time step t, perform state observation to obtain the state sequence, make action selection, obtain action feedback, and store data;

[0112] (S52.3) Determine whether the network update standard has been reached, that is, a certain amount of data for update has been stored. If so, execute step (S52.4); otherwise, execute step (S52.2);

[0113] (S52.4) Calculate the value and cost of each state-action pair through the evaluation network and the cost network, and update the policy network according to the formula ;

[0114] (S52.5) Update the evaluation network and the cost network according to the rewards and costs in the stored data according to the formula ;

[0115] (S52.6) Update the cost adjustment factor through the estimation of the cost network;

[0116] (S52.7) Determine whether the reinforcement learning model has converged, that is, the rewards and costs of the network tend to be smooth and stable. If so, end the training process; otherwise, execute (S52.2);

[0117] Specifically, as Figure 3 shown, the above training process trained a total of 3,600 rounds. It can be seen that the algorithm began to converge at the 1,500th round, indicating that the algorithm could then safely and stably guide the decision-making behavior of the autonomous vehicle;

[0118] S6. Deploy the trained reinforcement learning model to the vehicle to implement autonomous driving behavior decision-making.

[0119] In summary, by introducing cost constraints, the present invention enables the autonomous driving system to avoid potential dangerous behaviors in complex environments and improve safety; and by introducing an action regularization term into the policy loss function, the stability of the system under uncertain conditions is enhanced, and decision-making errors caused by external interference are reduced; at the same time, by introducing a transformer time-series model to design the network structure, the model's ability to capture information is enhanced, and the effectiveness of decision-making is improved.

[0120] Embodiment 2:

[0121] This embodiment provides an autonomous driving behavior decision-making system integrating a time-series Transformer and safety constraints for implementing the above-mentioned autonomous driving behavior decision-making method integrating a time-series Transformer and safety constraints, including:

[0122] A scenario construction module for constructing an autonomous driving mixed traffic scenario, which includes autonomous vehicles and human-driven vehicles;

[0123] A human-driven vehicle model construction module for constructing a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making to simulate human driving behavior;

[0124] An agent construction module for constructing an autonomous driving vehicle agent for the autonomous driving vehicles in the mixed traffic scenario and designing the state space, action space, reward function, and cost function of the autonomous driving vehicle agent;

[0125] A design module for further designing the overall optimization objective based on the designed state space, action space, reward function, and cost function of the autonomous driving vehicle agent, introducing an action regularization mechanism, and establishing an overall objective equation;

[0126] A training module for constructing a reinforcement learning model based on the Transformer model, learning the overall optimization objective set by the overall objective equation, training the reinforcement learning model, and obtaining the trained reinforcement learning model;

[0127] A decision-making module, which is used to deploy the trained reinforcement learning model to a vehicle for making autonomous driving behavior decisions.

[0128] Specifically, the above-mentioned scenario construction module, human-driven vehicle model construction module, agent construction module, design module, training module and decision-making module can be embedded in a computer processing system. The computer, according to the autonomous driving behavior decision-making method integrating the temporal Transformer and safety constraints provided above, calls the above-mentioned modules to complete the task of making autonomous driving behavior decisions; the above-mentioned scenario construction module, human-driven vehicle model construction module, agent construction module, design module, training module and decision-making module can perform operations according to the specific steps given by the autonomous driving behavior decision-making method integrating the temporal Transformer and safety constraints.

[0129] It should be noted that it should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by processing elements and partially implemented in the form of hardware. For example, the traffic scenario construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above signal processing module is called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the hardware of the processor element or the instruction in the form of software.

[0130] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above methods. For example: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain above-mentioned module is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0131] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0132] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.

[0133] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

[0134] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

Claims

1. An autonomous driving behavior decision-making method that integrates a temporal Transformer and safety constraints, characterized in that, Including the following steps: Construct an autonomous driving mixed traffic scenario, which includes autonomous driving vehicles and human-driven vehicles; Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior; Construct an autonomous driving vehicle agent for the autonomous driving vehicles in the mixed traffic scenario, and design the state space, action space, reward function, and cost function of the autonomous driving vehicle agent; Based on the designed state space, action space, reward function, and cost function of the autonomous driving vehicle agent, further design the overall optimization objective, introduce an action regularization mechanism, and establish the overall objective equation; Construct a reinforcement learning model based on the Transformer model, learn the overall optimization objective set by the overall objective equation, for training the reinforcement learning model to obtain the trained reinforcement learning model; Deploy the trained reinforcement learning model to the vehicle for realizing autonomous driving behavior decision-making.

2. The method for making an autonomous driving behavior decision by fusing a temporal Transformer and safety constraints according to claim 1, wherein: The autonomous driving mixed traffic scenario is built using the highway-env toolkit, and the autonomous driving mixed traffic scenario is specifically the highway ramp merging scenario.

3. The method for autonomous driving behavior decision-making integrating temporal sequence Transformer and safety constraints according to claim 2, characterized in that: Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, specifically including: (31) Control the longitudinal decision-making through the IDM model, and the control formula is as follows: where is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle ahead, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle ahead, and T is the safe time headway of the vehicle; (32) Control the lateral decision-making through the MOBIL model, and the control formula is as follows: Among them is the acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n is the acceleration of this vehicle before lane change is the predicted acceleration of the current vehicle after lane change, a c is the acceleration of the current vehicle before lane change is the acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o is the acceleration of this vehicle before lane change, Δa th is the lane change threshold, b safe is the safe deceleration threshold 4. The autonomous driving behavior decision-making method integrating a temporal transformer and safety constraints according to claim 1, characterized in that, The state space and action space of the autonomous driving vehicle agent are specifically as follows: (41) The state space of the autonomous vehicle agent at time step t is s t = {s t,1 , s t,2 , s t,3 , …, s t,N}, where N is the number of vehicles in the autonomous driving scenario, and the state of each vehicle is a time series Each time series state unit where present represents the perceivable state of the i-th vehicle k steps ago, and represent the x and y coordinates of the position, v x 、v y represent the x and y coordinates of the speed; (42) The action space is A = {a1, a2, a3, a4, a5}, where a1, a2, a3, a4, a5 respectively represent 5 discrete actions of turning left, turning right, staying unchanged, accelerating, and decelerating.

5. The method for autonomous driving behavior decision-making integrating temporal sequence Transformer and safety constraints according to claim 4, wherein: The reward function of the autonomous driving vehicle agent is specifically as follows: R(s,a) = w v ·r v +w cl ·r cl +w sc ·r sc where r v is the speed reward, r cl is the successful lane change reward, r sc is the reward for successfully reaching the end point, w v , w cl , w sc are their corresponding weight parameters respectively, and the specific formula is as follows: where v ego is the current speed of the autonomous vehicle, and is the average speed of surrounding vehicles.

6. The decision-making method for autonomous driving behavior integrating temporal sequence Transformer and safety constraints according to claim 5, wherein: The cost function of the autonomous driving vehicle agent is specifically as follows: C(s,a) = w cr ·c cr +w wa ·c wa +w pr ·c pr where c cr is the collision cost, c wa is the cost of inappropriate actions, c pr is the prediction cost, w cr , w wa , w pr are their corresponding weight parameters respectively, and the specific formula is as follows: where d min is the distance between the autonomous vehicle and the nearest vehicle around it, and d safe is the safety distance, is the distance cost between the host vehicle and the surrounding vehicles after performing an action, is the speed difference cost between the host vehicle and the surrounding vehicles after performing an action, and c pc is the cost of not changing lanes for a long time, and w cr , w vr , w pc are the corresponding weight parameters respectively, The specific design of c pc is as follows: where v thr is the safety speed threshold value.

7. The method for autonomous driving behavior decision-making integrating temporal sequence Transformer and safety constraints according to claim 1, characterized in that, Based on the designed state space, action space, reward function, and cost function of the autonomous driving vehicle agent, further design the overall optimization objective, introduce an action regularization mechanism, and establish the overall objective equation, specifically as follows: (71) The overall objective equation introduces constraints on the basis of traditional reinforcement learning, and the specific equation is: where \(R(s t , a t )\) is the agent reward function, \(C(s t , a t )\) is the agent cost function, \(s t \) is the state observation corresponding to time step \(t\), \(\lambda\) is the adjustment factor, \(\gamma\) is the discount factor with a value range from \(0\) to \(1\), which is used to balance the current and future rewards and costs. The overall goal is to maximize the obtained rewards while minimizing the driving cost; (72) The overall objective equation introduces an action regularization mechanism on the basis of traditional reinforcement learning, and the specific expression is: where π φ is the agent policy network, represents the introduced noise with a mean of 0 and a variance of σ, and s represents the current state observation of the autonomous vehicle.

8. The method for autonomous driving behavior decision-making integrating temporal sequence Transformer and safety constraints according to claim 1, wherein: The reinforcement learning model based on the Transformer model includes a policy network, two evaluation networks, two evaluation target networks, a cost network, and a cost target network, a total of 7 components based on the temporal Transformer: Among them, the input of the policy network is the sequential observation state sequence, which captures the sequential dependence through the self-attention mechanism and outputs the action probability distribution at the current moment; the evaluation network adopts a dual-network structure to alleviate overestimation, inputs the joint embedding of the sequential state and action, and predicts the Q value of the state-action pair; the evaluation target network is a copy of the evaluation network with delayed update, providing a stable Q value calculation target; the cost network inputs the joint representation of the sequential state and action, predicts the action safety cost, and constrains dangerous behaviors; the cost target network is a copy of the cost network with delayed update, which is used for calculating the target value of the cost function; The loss functions in the reinforcement learning model are as follows respectively: where L π is the loss function of the policy network, π φ is the current policy to be updated, H(π φ ) is the entropy of the policy, s is the observed state sequence of the current autonomous vehicle, a is the decision made by the current policy, Q(s,a) is the minimum of the two evaluation networks Q1(s,a) and Q2(s,a), R reg (a) is the action regularization term, (0, σ 2 ) is the Gaussian noise distribution, is the loss function of the evaluation network, R n is the reward TD error obtained through the target network, L C is the loss function of the cost network, C n is the cost TD error obtained through the target network, L λ is the loss function of the cost adjustment factor of, C max is the maximum cost threshold.

9. The method for autonomous driving behavior decision-making integrating temporal transformers and safety constraints according to claim 1, wherein: The training process of the reinforcement learning model is specifically as follows: (91) Initialize the policy network π of the reinforcement learning model φ , the evaluation networks Q1(s,a) and Q2(s,a) and their corresponding target networks, and the cost network C ψ (s,a) and its corresponding target network and environmental parameters; (92) At each time step t, perform state observation to obtain a state sequence, make an action selection, obtain action feedback, and store data; (93) Determine whether the network update standard is reached, that is, a certain amount of data available for update has been stored. If so, execute step (94), otherwise, execute step (92); (94) Calculate the value and cost of each state-action pair through the evaluation network and the cost network, and update the policy network according to the formula Update the policy network; (95)According to the formula based on the rewards and costs in the stored data Update the evaluation network and the cost network; (96) Update the cost adjustment factor through the estimation of the cost network; (97) Determine whether the reinforcement learning model has converged, that is, the rewards and costs of the network tend to be smooth and stable. If so, end the training process, otherwise, execute (92).

10. An autonomous driving behavior decision-making system integrating a temporal Transformer and safety constraints, which is used to implement the autonomous driving behavior decision-making method integrating a temporal Transformer and safety constraints according to any one of claims 1 to 9, characterized in that, Including: A scenario construction module, which is used to construct an autonomous driving mixed traffic scenario, and the mixed traffic scenario includes autonomous driving vehicles and human-driven vehicles; A human-driven vehicle model construction module, which is used to construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making, and the MOBIL model is used to control its lateral decision-making, for simulating human driving behaviors; An agent construction module, which is used to construct an autonomous driving vehicle agent for the autonomous driving vehicles in the mixed traffic scenario, and design the state space, action space, reward function, and cost function of the autonomous driving vehicle agent; A design module, which is used to further design the overall optimization target based on the designed state space, action space, reward function, and cost function of the autonomous driving vehicle agent, introduce an action regularization mechanism, and establish an overall objective equation; A training module, which is used to construct a reinforcement learning model based on the Transformer model, learn the overall optimization target set by the overall objective equation, and is used to train the reinforcement learning model to obtain a trained reinforcement learning model; A decision-making module, which is used to deploy the trained reinforcement learning model to the vehicle to realize autonomous driving behavior decision-making.

Citation Information

Cited By

  • Mobile robot navigation safety reinforcement learning method

    CN121089755A