Causal reinforcement learning system for vehicles in safety-critical scene with causal confusion
By introducing state causal models and reward causal models into the causal reinforcement learning system of autonomous driving vehicles, combining gradient-based causal discovery algorithms and Gumbel-Softmax sampling technology, the causal confusion problem is solved, significantly improving the robustness and safety of autonomous driving vehicles in safety-critical scenarios.
Patent Information
- Application Number
- CN202411830142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-16
AI Technical Summary
When autonomous vehicles face safety critical scenarios with causal confusion, existing robust reinforcement learning algorithms are difficult to identify real causal relationships, resulting in the impact of decision robustness and safety.
A causal reinforcement learning system is adopted, which includes a strategy generation network module and a causal graph module, and uses a state causal model and a reward causal model to identify the true causal relationship to prevent causal confusion and feature misleading. The system adaptively learns and refines causal relationships through gradient-based causal discovery algorithms and Gumbel-Softmax sampling technology.
It significantly enhances the robustness and safety of autonomous vehicles in safety-critical scenarios with causal confusion, and can maintain good performance in normal and safety-critical causal confusion scenarios.
Smart Images

Figure CN120012838A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion. Background Art
[0002] In recent years, the rapid development of autonomous driving technology is expected to significantly reduce traffic accidents and improve traffic efficiency. In autonomous driving systems, the decision-making system is crucial, acting as the "brain" of the vehicle to control its actions. However, there are still many challenges to be solved before fully autonomous driving can be deployed in complex real-world environments.
[0003] As an important component of artificial intelligence, reinforcement learning (RL) has shown promising results in the decision-making process of autonomous driving. However, ensuring that autonomous vehicles are sufficiently robust in certain scenarios remains a significant challenge, especially in rare but safety-critical scenarios such as emergency braking. To enhance the robustness of autonomous vehicles (AVs), methods such as removing non-safety-critical states and training environments based on random processes have been explored. However, in a specific type of safety-critical scenario, which this application refers to as a "safety-critical scenario with causal confusion", the robustness challenge may become more significant. Causal confusion occurs when a system incorrectly identifies a cause, leading to decisions based on misleading rather than true causal factors. In such a safety-critical scenario, causal confusion may undermine decision-making and lead to inappropriate responses, which further impairs robustness and increases the risk of accidents. For example, if an AV mistakenly attributes emergency braking to the fact that the vehicle is a small car (e.g., assuming that small cars are more likely to brake suddenly) rather than an obstacle, it may respond inappropriately.
[0004] Solving the causal confusion problem in such safety-critical scenarios is crucial to enhancing the robustness and safety of autonomous vehicles. Commonly used robust RL algorithms have difficulty addressing these challenges due to their lack of understanding of causal relationships. Causal discovery is a fundamental aspect of human cognition and has been integrated into RL to help agents resolve causal confusion and make more robust decisions. Causal models, as a key tool in the process, reveal the causal relationships between variables. Specifically, researchers use both structural and functional causal models to identify causal relationships between state and action features, thereby resolving causal confusion issues associated with these features. However, accurately identifying causal relationships associated with rewards remains underexplored, especially in the case of delayed rewards. Inherent reward delays can exacerbate causal confusion and increase the risk of agents misidentifying actions that lead to conflicts. Longer reward delays make it more difficult for agents to correctly associate rewards with correct actions. Therefore, the work of this application addresses this gap by focusing on causal confusion of state-action features and reward delays, enhancing the robustness and safety of RL algorithms in AVs.
[0005] Causal graphs are commonly used to represent and exploit causal relationships, and researchers have adopted various methods to learn these graphs, such as independence tests, structural causal models, and exhaustive sampling from all possible graphs. However, as the number of nodes in the causal graph increases, the computational requirements also increase significantly. This increase makes previous causal discovery methods less effective in high-dimensional autonomous driving environments. Therefore, a more adaptable causal discovery method is needed to cope with the complexity of these environments. Summary of the invention
[0006] An embodiment of the present application provides a causal reinforcement learning system for a vehicle in a safety-critical scenario with causal confusion, so as to enhance the robustness of an autonomous driving vehicle in a safety-critical scenario with causal confusion.
[0007] In order to solve the above technical problems, an embodiment of the present application provides a causal reinforcement learning system for vehicles in a safety-critical scenario with causal confusion, including: a strategy generation network module and a causal graph module; the causal graph module is composed of a causal model; the causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; the causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; the causal graph is represented by a directed binary adjacency matrix, and the matrix contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action features; the rows of the state causal graph represent state features, and the columns represent cascaded state and action features.
[0008] In some exemplary embodiments, the causal reinforcement learning system further includes a verification module for verifying the robustness of causal reinforcement learning; the verification process of the verification module is performed in three different scenarios, each of which includes normal and safety-critical scenarios, and there is causal confusion between the normal and safety-critical scenarios.
[0009] In some exemplary embodiments, the scenarios include a normal car-following scenario, a safety-critical car-following scenario, a normal lane-changing scenario, a safety-critical lane-changing scenario, a normal pedestrian crossing scenario, and a safety-critical pedestrian crossing scenario.
[0010] In some exemplary embodiments, the strategy generation network module is an actor-critic network; the actor-critic network is configured with two hidden layers, each hidden layer consists of 256 neurons and uses a ReLU activation function; the actor-critic network includes a critic network and an actor network; the critic network outputs a single Q value without the need for an activation function; the output layer of the actor network has a dimension equal to that of the action space and uses a Tanh activation function to generate an action value.
[0011] In some exemplary embodiments, the causal graph module uses trajectories extracted from the replay buffer for learning; the Gumbel-Softmax sampling method is used to generate the causal graph, so that the causal model can adaptively learn and refine causal relationships across various dimensions, and maintain robustness even in high-dimensional and dynamic environments; each edge of the causal graph is denoted as G ij , sampling is performed according to the Gumbel-Softmax distribution; during the training process, Gumbel-Softmax sampling uses soft sampling to generate a causal graph based on the probability distribution; during the testing process, Gumbel-Softmax sampling uses hard sampling to select the elements with the highest probability to form a causal graph.
[0012] In some exemplary embodiments, the Gumbel-Softmax distribution is as shown in formula (1):
[0013]
[0014] Among them, p ij represents the probability that there is a causal relationship between node i and node j; based on this probability, edge values are sampled, g ik and g ij Both are Gumbel noise; the temperature parameter τ in the Gumbel-Softmax distribution controls the smoothness of the logits-softmax curve, and a lower value makes the output closer to a discrete binary value.
[0015] In some exemplary embodiments, in a directed binary adjacency matrix, each node of the causal graph corresponds to a variable, and each directed edge reflects a causal relationship from one node to another node; in the directed binary adjacency matrix, a value of 1 indicates the existence of a causal relationship, and a value of 0 indicates the absence of a causal relationship; in a causal model, the input is a concatenated vector of state and action feature vectors; the concatenated vector is Hadamard-producted with each column of the causal graph, and then the result is concatenated into a single vector and input into a fully connected layer for predicting the next state feature or reward; the causal graph module captures the causal relationship between state, action and reward variables through binary masks and Hadamard products; the state feature is calculated by formula (2);
[0016]
[0017] Where s(t), a(t) and r(t) represent the state features, action features and rewards at time step t, respectively, and s(t+1) represents the state features at the next time step t+1; Respectively represent the causal relationship between the state, action features and the i-th state feature at the next moment, G s→r , G a→r Respectively represent the causal relationship between state, action features and rewards; ⊙ is the Hadamard product, i = (1, 2, ..., |s|); ∈ s,t and ∈ s,t is random noise.
[0018] In some exemplary embodiments, the state causal model and the reward causal model are expressed using formula (3):
[0019]
[0020] in, and represents the causal structure and strength of the edge, and Represents a set of basis functions.
[0021] In some exemplary embodiments, the strategy generation network module is used to design an algorithm based on the parameter design of the network structure; when designing the algorithm, the consistency between the behavior of the autonomous driving vehicle and the expected safety and efficiency goals is achieved through a reward function; the reward function includes a collision reward and a speed reward, as shown in formula (4):
[0022] r=r collsion +r velocit y (4)
[0023] Among them, r collosion 、r velocityThey represent collision reward and speed reward respectively; the negative reward given for a collision that occurs in the simulation is -10.
[0024] In some exemplary embodiments, a speed reward function is designed to keep the vehicle at its optimal speed at all times; the speed reward function is calculated based on the speed of the vehicle at each step; the speed reward consists of a normalized function of the vehicle speed and a reward coefficient, as shown in formula (5):
[0025]
[0026] Among them, r max velocity is the reward coefficient, v is the speed of the vehicle, v max and v min Represent the maximum speed and minimum speed in the scene respectively.
[0027] The technical solution provided by the embodiments of the present application has at least the following advantages:
[0028] An embodiment of the present application provides a causal reinforcement learning system for a vehicle in a safety-critical scenario with causal confusion, including: a strategy generation network module and a causal graph module; the causal graph module is composed of a causal model; the causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; the causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; the causal graph is represented by a directed binary adjacency matrix, and the matrix contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action features; the rows of the state causal graph represent state features, and the columns represent cascaded state and action features.
[0029] This application develops a causal reinforcement learning (CRL) paradigm that includes an actor-critic network and a causal graph module to enhance the robustness of autonomous vehicles (AVs) in safety-critical scenarios with causal confusion. Moreover, in order to reveal the causal relationship between causal models, this application also proposes a gradient-based causal discovery method that uses Gumbel-Softmax sampling technology, which can adapt to causal relationship learning across multiple dimensions, including safety and efficiency. In order to verify the robustness of CRL, this application designs car-following, lane changing, and pedestrian crossing scenarios, each of which contains normal and safety-critical causal confusion scenarios. The results show that CRL has good robustness in both normal and safety-critical causal confusion scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplifications do not constitute limitations on the embodiments. Unless otherwise stated, the pictures in the drawings do not constitute proportional limitations.
[0031] Figure 1 A schematic diagram of the results of a causal reinforcement learning system for a vehicle in a safety-critical scenario with causal confusion provided in one embodiment of the present application.
[0032] Figure 2 This is a diagram of a normal scenario and a safety-critical scenario with causal confusion provided by an embodiment of the present application.
[0033] Figure 3 This is a graph showing changes in rewards during training under normal circumstances provided by an embodiment of the present application.
[0034] Figure 4 This is a learning causal diagram provided by an embodiment of the present application.
[0035] Figure 5 This is an algorithm comparison diagram of event 1 in a pedestrian crossing scenario provided by an embodiment of the present application.
[0036] Figure 6 This is an algorithm comparison diagram of event 2 in a crosswalk scenario provided by an embodiment of the present application.
[0037] Figure 7 This is an algorithm comparison diagram of event 3 in a pedestrian crossing scenario provided in an embodiment of the present application.
[0038] Figure 8 This is an algorithm comparison diagram of event 4 in a pedestrian crossing scenario provided by an embodiment of the present application.
[0039] Fig. 9 This is an algorithm comparison diagram of event 5 in a pedestrian crossing scenario provided in an embodiment of the present application.
[0040] Fig.10 This is an algorithm comparison diagram of event 6 in a pedestrian crossing scenario provided by an embodiment of the present application. DETAILED DESCRIPTION
[0041] As can be seen from the background technology, in safety-critical scenarios, autonomous driving vehicles may encounter robustness issues due to causal confusion, leading to poor decision-making and further increasing the risk of accidents.
[0042] Among traditional robust reinforcement learning methods, RL algorithms have been shown to achieve human-level performance on gaming tasks, such as playing Atari games and the board game Go. Based on these successes in controlled environments, some researchers have attempted to introduce RL algorithms to the more complex and unpredictable domain of autonomous driving. For example, RL has been applied to improve lane-changing decisions, merging maneuvers, and navigation in complex traffic environments. Techniques such as deep Q-networks (DQNs) and proximal policy optimization (PPO) have made significant progress in handling dynamic traffic scenarios and optimizing vehicle control policies.
[0043] Despite their success, traditional RL algorithms often exhibit poor robustness in the face of environmental changes. Therefore, researchers have adopted a variety of approaches to develop RL-based robust autonomous driving systems. One effective approach is adversarial RL, which uses simulated adversarial perturbations to train the observed states to help the agent become robust to unexpected environmental changes. Another widely adopted approach is domain randomization, which enhances real-world generalization by training RL algorithms under widely randomly varying environmental conditions, thereby improving its ability to handle previously unseen scenarios. In addition, some researchers have explored hybrid models that combine RL with rule-based systems. These hybrid models allow the agent to switch between different models depending on the scenario, thereby enhancing the robustness of the AV in various scenarios.
[0044] While these approaches improve the robustness of RL algorithms in dynamic environments, they often fail to address a more complex challenge, namely causal confusion. Causal confusion occurs when an algorithm incorrectly identifies the relationship between cause and effect, leading to decisions based on spurious correlations rather than true causal factors. This misidentification can significantly weaken the robustness of the system, especially leading to poor decisions in emergency situations.
[0045] Humans are born with the ability to understand causal relationships, which enables the application to make informed decisions. In this context, causal discovery is introduced into causal reinforcement learning to help agents better understand their environment by identifying causal relationships, enabling them to make more effective decisions. Researchers usually use structured causal models to determine causal relationships. Structured causal models provide a systematic approach to formalize and infer causal relationships between variables. A structured causal model consists of endogenous variables, exogenous variables, structural equations, and the joint probability distribution of exogenous variables.
[0046] Research on causal RL mainly addresses issues such as generalizability, causal confusion, and sample efficiency. The deceptive relationships between variables caused by causal confusion pose a major challenge to RL. The deceptive relationships between variables caused by causal confusion pose a major challenge to RL. To address these challenges, researchers have proposed various innovative methods. For example, a related technique develops a parameterized strategy based on a causal graph, takes the causal graph as a hidden variable, and iteratively optimizes it during the policy learning process. Another related technique solves the causal confusion problem in reinforcement learning by using robust exploration to distinguish true causal relationships from spurious correlations. However, most of these studies focus on causal confusion related to state-action features. Inherent reward delays can also lead to reward causal confusion because they make it difficult for agents to identify the true causal relationship behind the reward. Therefore, solving the causal confusion caused by state-action features and the causal confusion caused by reward delays is crucial to enhancing the robustness and safety of RL algorithms in autonomous vehicles.
[0047] A causal graph is a graphical representation of a structured causal model that uses nodes and directed edges to intuitively represent variables and their causal relationships. To effectively identify these relationships, researchers have used various methods to learn causal graphs, such as independence tests and sampling from all possible causal graphs. Therefore, a causal discovery method that can effectively adapt to high-dimensional environments is needed. In the study of causal reinforcement learning for safety-critical scenarios with causal confusion, normal scenarios are defined as common driving situations that are typically encountered during daily driving. Safety-critical scenarios with causal confusion are a special class of safety-critical scenarios in which the AV incorrectly identifies the causal relationships between key features in the scenario. This causal confusion causes the autonomous vehicle to make decisions based on incorrect or misleading factors instead of accurately identifying the true causes of critical events. As a result, the autonomous vehicle may respond inappropriately to the scenario, compromising its robustness and significantly increasing the risk of accidents. A major focus of this application is the robustness of the model in these scenarios.
[0048] To solve the above technical problems, the embodiment of the present application provides a causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion, including: a strategy generation network module and a causal graph module; the causal graph module is composed of a causal model; the causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; the causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; the causal graph is represented by a directed binary adjacency matrix, and the matrix contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action features; the rows of the state causal graph represent state features, and the columns represent cascaded state and action features. The present application provides a causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion to enhance the robustness of autonomous driving vehicles in safety-critical scenarios with causal confusion.
[0049] The following will describe the various embodiments of the present application in detail with reference to the accompanying drawings. However, it will be appreciated by those skilled in the art that in the various embodiments of the present application, many technical details are provided in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solution claimed in the present application can be implemented.
[0050] See also Figure 1 , an embodiment of the present application provides a causal reinforcement learning system for vehicles in a safety-critical scenario with causal confusion, including: a strategy generation network module and a causal graph module; the causal graph module is composed of a causal model; the causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; the causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; the causal graph is represented by a directed binary adjacency matrix, and the matrix contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action features; the rows of the state causal graph represent state features, and the columns represent cascaded state and action features.
[0051] The framework of the causal reinforcement learning (CRL) paradigm proposed in this application is as follows Figure 1As shown in , it includes a strategy generation network module and a causal graph module; specifically, the strategy generation network adopts an actor-critic network, and the causal graph module is composed of a causal model. The strategy generation network can use the state causal model and the reward causal model to solve two types of causal confusion, namely, confusion caused by state-action features and confusion caused by reward delay, such as Figure 1 The causal graph module is trained using the data in the replay buffer, where s(t), a(t), and r(t) represent the state features, action features, and rewards at time step t, respectively; s(t+1) represents the state features at the next time step t+1; in this application, the ego vehicle represents an autonomous driving vehicle.
[0052] This application defines three normal scenarios with causal confusion and their corresponding safety-critical scenarios in SUMO (with a simulation time step of 0.5s) to verify the robustness of the algorithm in this paper (see Figure 2 The first scenario involves vehicles of different types (see Figure 2 a and Figure 2 b) Following scenario. Both vehicles are prohibited from changing lanes in the simulation environment. In normal scenarios, when the leading vehicle is a small car, it may show sudden braking or acceleration. On the contrary, when the leading vehicle is a bus, it usually accelerates and decelerates smoothly due to the presence of passengers. However, in safety-critical scenarios, buses may also show sudden braking or acceleration. In other cases, the speed control of the leading vehicle is controlled by the SUMO-based IDM. In this safety-critical scenario, the vehicle type will be a causal confounding feature.
[0053] Figure 2 The following are normal and safety-critical scenarios with causal confusion. (a) indicates that a small car may brake suddenly. (b) indicates that a bus may brake suddenly. (c) indicates that the lane changes with a turn signal. (d) indicates that a sudden lane change occurs without a turn signal. (e) indicates that a pedestrian crosses the street at a crosswalk. (f) indicates that a pedestrian unexpectedly crosses the street without a crosswalk. The second scenario is a lane change scenario with different vehicle states. In the normal scenario, the leading vehicle uses the turn signal for four time steps before changing lanes (see Figure 2 c). In contrast, in safety-critical scenarios, the vehicle ahead changes lanes suddenly without using a turn signal (see Figure 2 d). At other times, the vehicle ahead is prohibited from changing lanes. In this safety-critical scenario, the turn signal will be a causal confusion feature.
[0054] The third scenario that has attracted the attention of some researchers involves an AV passing through a median crosswalk where pedestrians are present. In this scenario, the AV’s field of view is blocked by another vehicle, making it challenging for the AV to detect pedestrians ( Figure 2 e). Figure 2 The edge of the line of sight, as shown by the black dashed line in e, is determined by connecting the AV’s perception sensors to the edges of other vehicles. However, pedestrians sometimes cross the street unexpectedly. Therefore, the safety-critical scenario is a simulation scenario where a pedestrian crosses the street at a location without a crosswalk ( Figure 2 f) Such scenarios pose greater challenges to AVs due to occlusion by other vehicles.
[0055] In the causal diagram module (see Figure 1 ), the state and reward causal model of the present application is designed as a combination of a causal graph and a fully connected (FC) layer. The causal graph is represented as a directed binary adjacency matrix, which contains both the reward causal graph G(r) and the state causal graph G(s). The rows of the reward causal graph represent rewards, while the columns represent the cascaded state and action features. Similarly, the rows of the state causal graph represent state features, and the columns represent the cascaded state and action features. The reward causal model is used to redistribute rewards based on the causal relationship between rewards and other features, while the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features. In the causal model, the input is a cascade of state and action feature vectors. The concatenated value is Hadamard-producted with each column of the causal graph, and the results are then concatenated into a single vector and input into the FC layer.
[0056] The causal model of this application integrates causal graphs with fully connected (FC) layers, aiming to combine the interpretability of causal inference with the nonlinear mapping capabilities of deep learning. The design uses Hadamard products to ensure that the model only focuses on causally relevant features, enhancing its robustness when processing complex, high-dimensional data. This combination not only retains the accuracy of causal information, but also improves the performance of the model in dynamic environments. The theoretical basis of the causal model of this application is structural causal models (SCMs). SCMs are constructed by using the form of X i =f i (PA i , U i ) represents a causal relationship, where X i is an endogenous variable, PA i Their parent variable, U i are exogenous variables; these models utilize directed acyclic graphs (DAGs) to describe causal dependencies, allowing do-calculus to be applied to determine the effects of interventions. For example, the average causal effect (ACE) of X on Y is given by ACE = E[Y|X = x] - E[Y|X = x′], which helps distinguish between correlation and causation and facilitates precise causal inference. In SCMs, if two variables use a common confounding variable, they may be correlated even if there is no direct causal relationship. This confounding effect may lead to causal confusion if it is not properly handled.
[0057] In order to resolve causal confusion and accurately model causal relationships, in some embodiments, the causal models of the causal graph module of the present application are learned using trajectories extracted from the replay buffer, and these models are built using gradient-based causal discovery techniques. The Gumbel-Softmax sampling method is used to generate the causal graph, so that the causal model can adaptively learn and refine causal relationships across various dimensions, and maintain robustness even in high-dimensional and dynamic environments; each edge of the causal graph is denoted as G ij , sampling is performed according to the Gumbel-Softmax distribution; during the training process, Gumbel-Softmax sampling uses soft sampling to generate a causal graph based on the probability distribution; during the testing process, Gumbel-Softmax sampling uses hard sampling to select the elements with the highest probability to form a causal graph.
[0058] In some embodiments, the Gumbel-Softmax distribution is as shown in formula (1):
[0059]
[0060] Among them, p ij represents the probability that there is a causal relationship between node i and node j; based on this probability, edge values are sampled, g ik and g ij Both are Gumbel noise; the temperature parameter τ in the Gumbel-Softmax distribution controls the smoothness of the logits-softmax curve, and a lower value makes the output closer to a discrete binary value.
[0061] In some embodiments, in a directed binary adjacency matrix, each node of the causal graph corresponds to a variable, and each directed edge reflects a causal relationship from one node to another; in the directed binary adjacency matrix, a value of 1 indicates the existence of a causal relationship, and a value of 0 indicates the absence of a causal relationship; in a causal model, the input is a concatenated vector of state and action feature vectors; the concatenated vector is Hadamard-producted with each column of the causal graph, and then the result is concatenated into a single vector and input into a fully connected layer for predicting the next state feature or reward; the causal graph module captures the causal relationship between state, action, and reward variables through binary masks and Hadamard products; the state feature is calculated by formula (2);
[0062]
[0063] Where s(t), a(t) and r(t) represent the state features, action features and rewards at time step t, respectively, and s(t+1) represents the state features at the next time step t+1; Respectively represent the causal relationship between the state, action features and the i-th state feature at the next moment, G s→r , G a→r Respectively represent the causal relationship between state, action features and rewards; ⊙ is the Hadamard product, i = (1, 2, ..., |s|); ∈ s,t and ∈ r,t is random noise.
[0064] In order to obtain the parameters and causal graph of the causal model, this application designs a gradient-based causal discovery method for causal discovery, which effectively explores the causal structure in high-dimensional data, uses differentiability to accelerate convergence, and adaptively adjusts to dynamic environments. It is worth noting that the methods for discovering causal relationships between state and reward causal models are different. The following is a proof that causal relationships can be obtained from causal models.
[0065] This application assumes that ∈s ,t和 ∈ r,t ∈ Formula (2) is iid additive noise. From the perspective of the weight space of the Gaussian process, equivalently, the state causal model and the reward causal model can be expressed using Formula (3):
[0066]
[0067] in, and represents the causal structure and strength of the edge, and Represents a set of basis functions.
[0068] Proposition 1: Causation, G s→s , G a→r and the transition dynamics f can be identified from the state causal model.
[0069] Proof: In order to obtain The closed-form solution can be used to calculate the covariate (X t Indicates the current state s t and action a t features) and the response variable s t+1 The covariate can be used to model the dependency between the current state s and the current state s, both of which are continuous. t and action a t The closed solution can be expressed as the loss function shown below.
[0070]
[0071] In the formula, represents the predicted state characteristics, λ||G state||1 is used to adjust the sparsity of the learned causal graph, where λ represents the regularization parameter and n is the number of state transition samples from the replay buffer. A state transition is represented by a five-tuple (s t ,a t ,r t ,s t+1 , completed). The input of the state causal model consists of state features and corresponding action features modified by the heuristic algorithm, and the purpose is to predict the state features at the next moment.
[0072] By taking the derivative of the cost function and setting it to zero, the present application can obtain a closed-form solution.
[0073]
[0074] Therefore, W can be identified from the data in the replay buffer. f This conclusion can be applied to all dimensions of the state. Therefore, the parent node of the i-th dimension of the state and the causal edge strength f are identifiable. In summary, through the state causal model, it can be found that G s→s , G a→s This completes the proof.
[0075] Finally, during the training process, a heuristic algorithm is designed to help destroy causal confusion in the state causal model as shown in the following formula. This algorithm helps to provide guidance for the search space of state features.
[0076]
[0077] Among them, s jk is the state eigenvector s j The kth feature of . Where k is randomly selected. ik Represents another vector s i The kth feature of , where i≠j. m represents the number of state feature vectors.
[0078] Proposition 2: Based on the entire trajectory return R, the causal relationship G can be identified from the state causal model s→r , G a→r and the transfer dynamics τ.
[0079] Proof: Similarly, the training data for the reward causal model is sampled from the replay buffer. The input consists of the state features changed by the heuristic algorithm and the corresponding action features. The output is the predicted reward. To achieve reward redistribution, the loss function is based on the return of the entire trajectory. The output is the predicted reward. To achieve reward redistribution, the loss function is based on the return of the entire trajectory. A loss function based on the total return of the entire trajectory can be used for reward redistribution because it ensures that rewards are distributed in a way that reflects the long-term impact of actions. By considering the entire trajectory, the algorithm can better understand which actions contribute to the overall success or failure, thereby achieving more accurate and effective reward distribution. This approach helps to solve problems such as delayed rewards and ensures that the agent learns to make decisions that maximize long-term benefits.
[0080] Based on the above assumptions, the function for calculating trajectory benefits can be rewritten as follows.
[0081]
[0082] Then, by modeling the covariate X τ The dependence between the response variable R and The closed-form solution of , where R is also continuous. The predicted trajectory gain is defined as The name of the discount factor is γ. represents the predicted reward at time step t. The solution needs to minimize the loss function and include a weight decay regularizer to prevent overfitting. The loss function of the reward causal model is shown below.
[0083]
[0084] in represents the replay buffer, and τ represents the sampling trajectory. A trajectory is a complete episode that records all data from the beginning to the end. t is the reward at time step t. λ||G reward ||1 is used to adjust the sparsity of the learned causal graph. The true trajectory gain is defined as
[0085] By substituting the derivative of the valence function and setting it to zero, a closed-form solution can be obtained, as shown in the following equation.
[0086]
[0087] Therefore, W can be identified from the data in the replay buffer. g This conclusion can be applied to all dimensions of the state. In addition, g represents the parent node of the state i dimension and the strength of the causal edge, which is identifiable. In summary, through the state causal model, it can be found that G s→r , Ga→r The reward causal model can be used to redistribute the reward r to solve the problem of delayed rewards. This completes the proof.
[0088] Algorithm 1 provides an overview of the design of the causal discovery algorithm in this application. The algorithm takes all the data in the replay buffer as input and outputs modified state transitions and rewards. These modified state transitions and rewards are then added back to the replay buffer to train the RL policy.
[0089]
[0090]
[0091] The CRL provided by this application is an innovative framework that integrates the state and reward causal models of this application to solve two specific types of causal confusion. The state causal model solves the causal confusion problem caused by state-action characteristics and reward delays. The core of CRL is the Causal MDP (CMDP), which is built on the traditional Markov Decision Process (MDP). MDP provides a mathematical foundation for RL problems by defining a framework for finding optimal policies. CMDP extends this paradigm by incorporating a causal graph module, enriching the standard MDP with causal relationships between state variables, actions, and rewards.
[0092] CMDP is a robust MDP enhanced by a causal graph module. CMDP can be defined by a six-tuple [S, A, p, r, γ, G]. S is a set of state features, called the state space. The action space A includes all potential actions that AVs can perform, such as acceleration, braking, etc. The transition probability p defines the probability of transitioning from the current state s to the next state s′, r is the reward function based on the current state, and γ∈(0, 1) represents the discount factor. G represents the causal graph module.
[0093] CMDP uses state features and rewards to optimize the decision-making behavior of AV. This method can learn robust driving behaviors with causal relationships. CMDP is used to solve the following problems:
[0094]
[0095] This application develops a new policy iteration paradigm called causal policy iteration to solve the CMDP problem. The paradigm consists of two main parts: causal policy evaluation and causal policy improvement, which are updated in an alternating manner. This iterative process continues until it reaches a convergence state, at which point the obtained policy is the optimal policy.
[0096] The role of causal policy evaluation is to optimize the critic (state-action value) network. The state-action value function Q G (s, a) is a basic concept in critic networks, which is used to estimate the expected total benefit of the causal graph module of causality when taking a specific action a in a given state s. Its main role is to help the agent choose the best action in a given state. G (s, a) is redistributed by the reward causal model. Based on the given policy π, the state-action value function is recursively updated using the Bellman backup operator T.
[0097]
[0098] Among them, s represents the state characteristics at the current moment, s′ G represents the state features of the next time step, which are modified by the causal component to reflect the causal relationship. a and a′ are based on s and s′ G Obtained by sampling.
[0099] In order to stabilize the training process, two k The state-action value function is k∈[0,1]. These two state-action value functions can be optimized by minimizing the objective function (Formula (12)).
[0100]
[0101] in, Replay buffer The sampled trajectory, y k is the target value of the state-action function. The data in the replay buffer is periodically changed by the causal graph module. In addition, during the training of the critic network, smaller state-action values are used to reduce overestimation. Therefore, y k It can be calculated by formula (13).
[0102]
[0103] Among them, Q′(s, a; θ k ) is the parameter θ k The target state-action value.
[0104] The parameters of the critic network can be updated by the Polyak averaging method, as shown in formula (14).
[0105] θ′ k ←βθ′ k +(1-β)θ k (14)
[0106] Among them, β is a smoothing factor close to 1, usually ranging from 0.95 to 0.99 to ensure progressive updating.
[0107] Causal policy improvement aims to improve the state-action value function based on the Optimizing actor (policy) networks. The calculation of takes into account the causal effect of action a on the immediate reward and subsequent state characteristics when action a is taken in state s. Therefore, the problem to be solved in formula (10) can be rewritten as:
[0108]
[0109] The optimal policy should choose the action that maximizes the expected reward in each state. This is usually done by gradually approximating the optimal policy π * This is done through an iterative approach that ensures that the strategy is gradually improved, consistent with the goal of achieving the highest possible overall return under the given circumstances.
[0110]
[0111] Since two Q-value networks are used, the policy model parameter a can be learned through the following objective function associated with the actor network, as shown in formula (17).
[0112]
[0113] Among them, J π (α) represents the objective function for optimizing the actor network parameter a. The parameter is updated using the gradient ascent method.
[0114] To fully understand the causal reinforcement learning (CRL) paradigm for robust autonomous driving, Algorithm 2 outlines the step-by-step operational flow. The parameters of the policy network and state-action value network are initialized from random distributions. The parameters of the target state-action value network are initialized by equating it to the state-action value network. The agent interacts with the environment, samples actions from the policy network, and stores the state transitions in the replay buffer. The “done” signal indicates that the agent terminated the current episode at time step t due to reasons such as collision. Algorithm 1 is used to modify the sampled state transitions to ensure that there is a true causal relationship between the state and the reward. The state-action value, target state-action value, and parameters of the policy network are updated at each gradient update step.
[0115]
[0116]
[0117] To facilitate the study, the state features of these three scenarios are standardized. The state space is a set of all possible state features. The state space contains information such as the coordinates, speed, angle, and distance of vehicles and pedestrians. Specifically, the state spaces δ1, δ2, and δ3 in the three scenarios are shown in formula (18). In these three scenarios, the origin of the coordinate axis is located in the upper left corner of the simulation scene. The X-axis represents the horizontal direction, and the Y-axis represents the vertical direction.
[0118] δ1={x ego ,y ego , x other ,y other , v ego ,v other ,d,θ ego ,θ other ,T other ,C}
[0119] δ2={x ego ,y ego , x other ,y other ,y ego ,v other ,d x ,d y ,θ ego ,θ other , T other ,C}
[0120] δ3={x ego ,y ego , x other ,y other , x ped ,y ped , v ego , v ped ,d,θ ego ,θ ped , T crossuaik ,C} (18)
[0121] Among them, x ego ,y ego 、x other ,y other 、x ped and ped Represent the coordinates of the vehicle, other vehicles, and pedestrians, respectively. Similarly, v ego 、v other and v egoRepresent the speed of the vehicle, other vehicles and pedestrians respectively. C represents the collision situation of the vehicle, which is a Boolean value. T represents the causal confusion feature, which includes information such as vehicle type, turn signal and crosswalk. d in δ1 represents the lateral distance between the vehicle and other vehicles. d in δ2 x and d y The lateral and longitudinal distances between the ego vehicle and other vehicles are represented respectively. The d in δ3 represents the lateral distance between the ego vehicle and the target position. This feature is used to motivate the ego vehicle to reach the target position instead of staying at other positions.
[0122] The action space A is defined as a one-dimensional continuous action: the longitudinal acceleration a∈[a min , a max ]. In order to simulate the sudden change of vehicle speed, each vehicle’s a min and a max They are set to -5m / s respectively. 2 and 5m / s 2 .
[0123] In some embodiments, the strategy generation network module is used to design the algorithm according to the parameter design of the network structure; when designing the algorithm, the consistency of the behavior of the autonomous vehicle with the expected safety and efficiency goals is achieved through the reward function. The role of the reward function is to effectively align the behavior of the ego vehicle with the desired goal. In general, the three scenarios share the same goals: safety and efficiency. The reward function in this application is composed of the collision reward r collision and speed reward r velocity A collision involving the ego car results in a negative reward of -10.
[0124] The reward function includes collision reward and speed reward, as shown in formula (19):
[0125] r=r collision +r velocity (19)
[0126] Among them, r collision 、r velocity They represent collision reward and speed reward respectively; the negative reward given for a collision that occurs in the simulation is -10.
[0127] In some embodiments, a speed reward function is designed to keep the vehicle at its optimal speed at all times; the speed reward function is calculated based on the speed of the vehicle at each step; the speed reward is composed of a normalized function of the vehicle speed and a reward coefficient, as shown in formula (5):
[0128]
[0129] Among them, r max velocityis the reward coefficient, v is the speed of the vehicle, v max and v min Represent the maximum speed and minimum speed in the response scenario. Specifically, in the first and second scenarios, v max is 27.7m / s, v min is 10m / s. In the third scenario, v max is 15m / s, v min The reward coefficient is designed to be 0.5.
[0130] In some embodiments, the strategy generation network module is an actor-critic network; the actor-critic network is configured with two hidden layers, each hidden layer consists of 256 neurons and uses a ReLU activation function; the actor-critic network includes a critic network and an actor network; the critic network outputs a single Q value without the need for an activation function; the output layer of the actor network has a dimension equal to that of the action space and uses a Tanh activation function to generate action values.
[0131] The correct configuration of the algorithm is crucial to its performance. The actor-critic network is configured with two hidden layers, each consisting of 256 neurons and using the ReLU activation function. The critic network outputs a single Q value without the need for an activation function. The output layer of the actor network has the same dimension as the action space and uses the Tanh activation function to produce action values. In the causal graph module, the fully connected (FC) layer network includes an input layer, a fully connected layer, and an output layer. The fully connected layer contains 256 neurons and uses the ReLU activation function. The noise of the causal graph is generated from a uniform distribution and transformed using the Gumbel formula g = -log(-log(U)). Other hyperparameters of the proposed paradigm are detailed in Table 1.
[0132] Table 1 Main hyperparameters of the proposed paradigm
[0133]
[0134]
[0135] The rewards are used to evaluate the comprehensive performance of the algorithm in three different scenarios, each of which includes normal and safety-critical scenarios, where there is causal confusion. In addition, the average speed and collision rate are used to evaluate the driving efficiency and traffic safety of the autonomous vehicle. The collision rate is calculated by dividing the number of collisions by the total number of occurrences in all episodes. An occurrence is defined as rapid acceleration or deceleration in the first scenario, lane change of the leading vehicle in the second scenario, and interaction between the vehicle and pedestrians in the third scenario. In order to improve the performance of the autonomous vehicle under extreme conditions, the events in the first and second scenarios are designed to occur sequentially as a single event in each scenario. This means that multiple events occur in a single scenario. In contrast, the third scenario only occurs once per episode. Finally, 6 specific events are designed from each of the three scenarios to intuitively demonstrate the performance of the algorithm.
[0136] In order to benchmark the proposed CRL paradigm for autonomous driving, this application experimentally compares with state-of-the-art methods in three scenarios. Among them, Soft Actor Critic (SAC): This is a baseline algorithm that does not consider robustness. Robust Reinforcement Learning with Gaussian Noise (RRL-G): This algorithm achieves robustness by adding Gaussian noise to the observations at each time step before they are processed by the policy function, thereby smoothing the policy. Alternating Training with Learning Adversaries (ATLA): Adversarial reinforcement learning improves robustness by training an adversary online together with the agent. The adversary learns to optimally perturb state observations, forcing the agent to adapt and improve its robustness to stronger adversarial attacks. Fully Connected Robust Learning (FCRL): The only difference between this algorithm and the algorithm proposed in this application is that in the causal model, only the FC layer network is used, and the causal graph is not included.
[0137] In order to avoid the impact of environmental differences on training results, this application uses the same environment seed.
[0138] The following shows the test results and performance evaluation of the model through model training. In the model training section, the training results in normal scenarios and safety-critical scenarios with causal confusion are given. First, the five algorithms are trained in normal scenarios to obtain the initial autonomous driving policy models. Then, these policy models are further trained in safety-critical scenarios with causal confusion to enhance their robustness and ability to handle high-risk situations. This two-step training process ensures that the model not only performs well in normal situations, but also can effectively handle safety-critical scenarios with causal confusion. The reward value represents the overall performance of the algorithm during the training process. Each scenario contains a specific number of training sets, each containing 50 time steps to ensure the full convergence of the algorithm: the first scenario contains 500 sets, the second scenario contains 350 sets, and the third scenario contains 450 sets. Figure 3The reward curves of these algorithms in normal and safety-critical scenarios with causal confusion are shown. Each reward curve combines the results of 10 training runs, the solid line represents the average, and the shaded area represents the 95% confidence interval.
[0139] During training under normal scenarios, CRL achieved comparable performance to the baseline algorithm (see Figure 3 ). The introduction of the causal graph module does not significantly reduce the convergence speed of the algorithm of this application. This shows that CRL improves the performance in safety-critical scenarios with causal confusion without compromising its performance in normal scenarios.
[0140] In safety-critical scenarios with causal confusion, the introduction of the causal graph module leads to a significant improvement of CRL over the baseline algorithm. This shows that the robustness of CRL is enhanced in safety-critical scenarios with causal confusion. It is worth noting that the lead of CRL gradually increases from the first scenario to the third scenario, indicating that the impact of the causal graph module becomes more obvious in more complex scenarios. In addition, the performance of FCRL is very unstable. For example, Figure 3 The performance of FCRL in b is significantly worse than that of other baseline algorithms. Figure 3 In Figure 3, FCRL slightly outperforms other baseline algorithms. This shows that the causal graph plays a vital role in the causal graph module.
[0141] In causal relationship analysis, the causal graph is visualized as a heat map, such as Figure 4 As shown. Dark blue indicates that there is a causal relationship between two features, while light yellow indicates that there is no causal relationship. For ease of analysis, the state causal graph and reward causal graph are combined into one graph. The features on the vertical axis represent the current state and action features, while the features on the horizontal axis represent the state features and reward features of the next time step. Overall, the reward and state causal model proposed in this paper accurately identifies the causal relationship between different features. Specifically, Figure 4 a shows that in the first scenario, the confusion feature T other There is no significant causal relationship between the velocity vector of the ego vehicle and the number of collisions C. This helps the algorithm of the present application avoid being misled by confusing features. In addition, the y coordinate y of the other vehicle other or angle θ other There is no significant causal relationship with the reward. Accurately identifying the causal relationship of the reward can help CRL redistribute the reward more efficiently, thereby avoiding delays in reward distribution.
[0142] In the second and third scenarios, similarly, the confusing feature T other / T crossThere is also no significant causal relationship between the ego vehicle’s speed, collision, or reward. In addition, there are other significant causal relationships. For example, the ego vehicle’s speed is related to the longitudinal distance d y There is a causal relationship (see Figure 4 b). This is because the longitudinal distance reflects the position of the ego vehicle within the lane, which directly affects the ego vehicle’s speed. Finally, the speed of the ego vehicle shows a significant causal relationship with the positions of other vehicles (see Figure 4 c). The reason for the causal relationship is that other vehicles can block the field of view of the vehicle, potentially leading to collisions with pedestrians. These causal relationships can also help RL learn the correct strategy in complex environments.
[0143] During the model verification process, policy models trained in safety-critical scenarios with causal confusion are used to test their performance in all scenarios. Compared with the other five algorithms (see Table 2), CRL shows balanced performance. Specifically, CRL has the lowest collision rate in three scenarios. The average speed of CRL in the first and second scenarios is second only to the best results of RRL-G and FCRL, respectively. In the third scenario, the average speed of CRL is second only to RRL-G and FCRL. These examples show that under normal circumstances, the five algorithms perform almost the same in terms of average speed and collision rate.
[0144] Table 2 Performance of five algorithms in three normal scenarios
[0145]
[0146]
[0147] Table 3 Performance of five algorithms in three safety-critical scenarios with causal confusion
[0148]
[0149] CRL shows advantages in almost all safety-critical scenarios with causal confusion (see Table 3). Specifically, the average speed of CRL in the first scenario is only 6.43% less than the optimal result of FCRL. In addition, CRL has the lowest collision rate of 0. In the second scenario, CRL obtains the highest average speed, which is 8.14% higher than other algorithms on average. Its collision rate is also the best, which is 84.03% lower than the average collision rate of other algorithms. Similarly, in the third scenario, CRL obtains the highest average speed, which is 4.78% higher than other algorithms on average, and has the lowest collision rate of 0.
[0150] Six different events are designed to visually demonstrate the performance differences between the algorithms and compare the algorithms under the same event. Event 1 and Event 2 correspond to the normal scenario and the safety-critical scenario of the first scenario with causal confusion, respectively. The speed distribution of the leading vehicle is randomly generated by the speed control function in the first scenario for a total of 200 time steps. In Event 1 and Event 2, the speed distribution is the same, differing only in the type of leading vehicle: a car in Event 1 and a bus in Event 2. The results show that all algorithms perform similarly in Event 1, except that RRL-G has a collision (e.g. Figure 5 shown).
[0151] In event 2, the CRL algorithm is significantly better than other algorithms. The speed curve of CRL not only closely follows the speed of the preceding vehicle, but also shows a higher responsiveness (see Figure 6 ).
[0152] Event 3 and Event 4 correspond to the normal scenario and safety-critical scenario of causal confusion in the second scenario, respectively, and both include lane-changing events. Event 3 involves a vehicle changing lanes after using the turn signal for four time steps, and completing the lane change in the next four time steps. Event 4 simulates a vehicle suddenly changing lanes without using a turn signal, and completing the lane change after four time steps. The results show that CRL responds best to lane-changing vehicles. After observing the lane-changing signal, the vehicle maintains the highest speed to overtake the lane-changing vehicle (see Figure 7 ). In event 4 (see Figure 8 ), when the vehicle suddenly changed lanes, the CRL effectively adjusted the speed to maintain safety. Overall, these results show that the CRL balances safety and speed well in dynamic and unexpected situations.
[0153] Event 5 and event 6 correspond to the normal and safety-critical scenarios of causal confusion in scenario 3, respectively. Event 5 involves pedestrians crossing the street at a crosswalk, while event 6 involves pedestrians crossing the street without a crosswalk (running a red light). The speed distribution of pedestrians in events 5 and 6 is consistent. The results show that in event 5, all algorithms can effectively reduce the vehicle's position (see Fig. 9 ) speed. In event 6, CRL adjusts the speed at a more appropriate position compared to other algorithms. This approach avoids the overly conservative behavior of other algorithms, thereby minimizing the impact on vehicle efficiency (see Fig.10 ). These results highlight the superior performance of CRL in balancing safety and efficiency in pedestrian crossing scenarios.
[0154] Based on the above technical scheme, an embodiment of the present application provides a causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion, including: a strategy generation network module and a causal graph module; the causal graph module is composed of a causal model; the causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; the causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; the causal graph is represented by a directed binary adjacency matrix, and the matrix contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action features; the rows of the state causal graph represent state features, and the columns represent cascaded state and action features.
[0155] This application develops a causal reinforcement learning (CRL) paradigm that includes an actor-critic network and a causal graph module to enhance the robustness of autonomous vehicles (AVs) in safety-critical scenarios with causal confusion. Moreover, in order to reveal the causal relationship between causal models, this application also proposes a gradient-based causal discovery method that uses Gumbel-Softmax sampling technology, which can adapt to causal relationship learning across multiple dimensions, including safety and efficiency. In order to verify the robustness of CRL, this application designs car-following, lane changing, and pedestrian crossing scenarios, each of which contains normal and safety-critical causal confusion scenarios. The results show that CRL has good robustness in both normal and safety-critical causal confusion scenarios.
[0156] This application proposes a new CRL paradigm to enhance the robustness of RL strategies in safety-critical scenarios with causal confusion. The paradigm includes an actor-critic network and a causal graph module, which aims to address two types of causal confusion. The state causal model addresses the causal confusion problem caused by state-action features, while the reward causal model addresses the reward delay problem. These causal models are learned using the gradient-based method based on the Gumbel-Softmax technique proposed in this application. Three safety-critical scenarios with causal confusion, namely car following, lane changing, and pedestrian crossing, are designed in SUMO for robustness testing.
[0157] The training results show that CRL significantly improves robustness, which manifests as higher reward values in safety-critical scenarios with causal confusion. This improvement is largely due to the causal model's ability to accurately identify significant causal relationships. Based on these training insights, the test results further validate the robustness of CRL. Specifically, the results show that in normal scenarios, CRL achieves balanced performance compared to the other five algorithms. However, in safety-critical scenarios with causal confusion, CRL achieves superior performance by maintaining a low collision rate and a high average speed. In the first scenario, CRL's average speed is only slightly lower than the best. In the second and third scenarios, its average speed exceeds the other algorithms by an average of 8.14% and 4.78%, respectively. CRL's collision rate in these three scenarios is 94.68% lower than that of other algorithms on average. Finally, the results of six events designed according to the three scenarios consistently support these conclusions, demonstrating CRL's superior ability to balance safety and efficiency. Overall, this study provides an effective method to improve the robustness of RL algorithms in safety-critical scenarios with causal confusion.
[0158] Those skilled in the art can understand that the above-mentioned embodiments are specific examples for implementing the present application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application, so the scope of protection of the present application shall be based on the scope defined in the claims.
Claims
1. A causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion, characterized in that include: A strategy generates a network module and a causal graph module; the causal graph module is composed of a causal model; The causal model includes a state causal model and a reward causal model; the reward causal model is used to redistribute rewards according to the causal relationship between rewards and other features, and the state causal model identifies the true causal relationship to prevent the algorithm from being misled by causal confusion features; The causal model is a combination of a causal graph and a fully connected layer; the causal model is constructed using a gradient-based causal discovery algorithm; The causal graph is represented by a directed binary adjacency matrix, which contains both a reward causal graph and a state causal graph; the rows of the reward causal graph represent rewards, and the columns represent cascaded state and action characteristics; the rows of the state causal graph represent state characteristics, and the columns represent cascaded state and action characteristics.
2. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: It also includes a validation module for verifying the robustness of causal reinforcement learning; The verification process of the verification module is performed in three different scenarios, each of which includes normal and safety-critical scenarios, and there is causal confusion between the normal and safety-critical scenarios.
3. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 2, characterized in that: The scenarios include a normal car-following scenario, a safety-critical car-following scenario, a normal lane-changing scenario, a safety-critical lane-changing scenario, a normal pedestrian crossing scenario, and a safety-critical pedestrian crossing scenario.
4. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: The strategy generation network module is an actor-critic network; the actor-critic network is configured with two hidden layers, each hidden layer consists of 256 neurons and uses a ReLU activation function; The actor-critic network includes a critic network and an actor network; the critic network outputs a single Q value without an activation function; the output layer of the actor network has a dimension equal to the action space and uses a Tanh activation function to generate an action value.
5. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: The causal graph module is learned using trajectories extracted from the replay buffer; The Gumbel-Softmax sampling method is used to generate causal graphs, which enables the causal model to adaptively learn and refine causal relationships across various dimensions and remain robust even in high-dimensional and dynamic environments; Let each edge of the causal graph be G ij , sampling is performed according to the Gumbel-Softmax distribution; During training, Gumbel-Softmax sampling uses soft sampling to generate causal graphs based on probability distributions; During testing, Gumbel-Softmax sampling uses hard sampling to select the elements with the highest probability to form a causal graph.
6. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 5, characterized in that: The Gumbel-Softmax distribution is shown in formula (1): Among them, p ij represents the probability that there is a causal relationship between node i and node j; based on this probability, edge values are sampled, g ik and g ij Both are Gumbel noise; the temperature parameter τ in the Gumbel-Softmax distribution controls the smoothness of the logits-softmax curve, and a lower value makes the output closer to a discrete binary value.
7. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: In a directed binary adjacency matrix, each node of the causal graph corresponds to a variable, and each directed edge reflects the causal relationship from one node to another; In the directed binary adjacency matrix, a value of 1 indicates the existence of a causal relationship, and a value of 0 indicates the absence of a causal relationship; In the causal model, the input is a concatenated vector of state and action feature vectors; the concatenated vector is Hadamard-producted with each column of the causal graph, and the result is then concatenated into a single vector and input into a fully connected layer for predicting the next state feature or reward; The causal graph module captures the causal relationship between state, action and reward variables through binary masks and Hadamard products; calculates state features through formula (2); Where s(t), a(t) and r(t) represent the state features, action features and rewards at time step t, respectively, and s(t+1) represents the state features at the next time step t+1; Respectively represent the causal relationship between the state, action features and the i-th state feature at the next moment, G s→r , G a→r Respectively represent the causal relationship between state, action features and rewards; ⊙ is the Hadamard product, i = (1, 2, ..., |s|); ∈ s,t and ∈ r,t is random noise.
8. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: The state causal model and the reward causal model are expressed using formula (3): in, and represents the causal structure and strength of the edge, and Represents a set of basis functions.
9. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 1, characterized in that: The strategy generation network module is used to design an algorithm based on the parameter design of the network structure; when designing the algorithm, the consistency between the behavior of the autonomous driving vehicle and the expected safety and efficiency goals is achieved through the reward function; The reward function includes a collision reward and a speed reward, as shown in formula (4): r=r collision +r velocity (4) Among them, r collision 、r velocity They represent collision reward and speed reward respectively; the negative reward given for a collision that occurs in the simulation is -10.
10. The causal reinforcement learning system for vehicles in safety-critical scenarios with causal confusion according to claim 9, characterized in that: By designing a speed reward function, the vehicle is always kept at its optimal speed; the speed reward function is calculated based on the speed of the vehicle at each step; the speed reward consists of a normalized function of the vehicle speed and a reward coefficient, as shown in formula (5): Among them, r max velocity is the reward coefficient, v is the speed of the vehicle, v max and v min Represent the maximum speed and minimum speed in the scene respectively.
Citation Information
Cited By
Electromechanical system fault diagnosis system based on deep learning
CN120163069A
A mechatronic system fault diagnosis system based on deep learning
CN120163069B
Event deduction and early warning method based on causal inference
CN120258155A
Accident scene generation method based on scene knowledge graph and considering accident causes
CN120496333A
A method for generating accident scenarios based on scene knowledge graphs and considering accident causes.
CN120496333B