A risk prediction method based on Granger causality test and interpretable reinforcement learning
The interpretable reinforcement learning method based on Granger causality test solves the problems of lack of human factors and insufficient interpretability in the reward function in autonomous driving, realizes efficient prediction and causal tracing of autonomous driving risks, and improves the safety and interpretability of autonomous driving.
Patent Information
- Application Number
- CN202211723861.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-12-30
AI Technical Summary
Existing reinforcement learning methods lack explainability in autonomous driving, making it difficult to track the reasons why low-risk goals are not achieved, and the reward function does not fully consider human factors, resulting in insufficient applicability.
An interpretable reinforcement learning method based on Granger causality test is adopted to construct the driving risk fitness function and reward function through semantic analysis and fuzzy set representation. Combined with Granger causality analysis, potential risk factors are screened, and a reinforcement learning model is constructed to achieve risk prediction and causal tracing.
It improves the interpretability and applicability of reinforcement learning, and can quickly and accurately predict and avoid potential road risks, achieving low-risk or even zero-risk autonomous driving.
Smart Images

Figure CN115841158B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to an autonomous driving risk association discovery method based on uncertainty causal reasoning using reinforcement learning. Background Art
[0002] Intelligent agent problems often require balancing interests among all participants and tracking positive or negative impacts. Effective adjustments are needed to benefit the current traffic situation, and tracking accident factors is also valuable for maintaining the safety of all participants. To avoid obstacles and accidents, autonomous vehicles need to account for all participants and their conditions, including changes in the speed of the autonomous vehicle, the sudden appearance of pedestrians, and unexpected traffic light breakdowns. Multi-agent reinforcement learning can operate effectively in multi-agent environments. It can find well-designed strategies to contribute to decision-making through specific reward scoring strategies. Incidental association discovery helps identify causal pairs and essential factors. Multi-agent problems can then be optimized and tracked simultaneously.
[0003] Causation has become a crucial component of autonomous driving technology. If the causes of car accidents can be predicted, preventative measures and assisted driving recommendations can be implemented, thereby improving autonomous driving safety. Many scientists, both domestically and internationally, have dedicated themselves to discovering these relationships. For example, Professor Judea Pearl conducts research on causal reasoning, and Professor Wang Peizhuang uses fuzzy sets to construct a factor space to identify the contributing factors of causal relationships. Furthermore, some scholars have introduced association mining methods to identify potential causal relationships, conducting experiments using synthetic data with a complete causal factor space.
[0004] As a repetitive self-learning method, reinforcement learning (RL) can automatically identify actions corresponding to low driving risks. There are two technical issues with predicting autonomous driving risks through reinforcement learning: First, RL usually requires multiple comparisons of various actions of autonomous vehicles to select the most appropriate action. In order to achieve autonomous learning, it is necessary to establish an action reward strategy and find a reward function based on autonomous driving actions. Since current reward functions rarely consider human factors, they lack corresponding user participation, which reduces their applicability. Second, the RL process in the autonomous driving process lacks explainability. Existing RL methods rarely explain why the selected action can lead to the best result. When the RL process does not achieve the expected low-risk goal, these methods are usually unable to track which risk factor has gone wrong. Summary of the Invention
[0005] In order to overcome the two problems mentioned above for reinforcement learning (RL), the present invention studies the interpretability of reinforcement learning (RL) and performs causal verification. The present invention starts with a basic but very typical case to study interpretability: reinforcement learning of autonomous driving risk based on time data, where the evolution of the operating state of a motor vehicle can be dynamically learned. The research goal is to improve the interpretability of the following instructions: (1) propose an action reward strategy; (2) evaluate the effectiveness of the decision; (3) predict its risk level based on the time series data of autonomous driving risk and assist in decision-making and analysis of the current driving behavior. The present invention is applied to risk prediction and risk factor analysis of autonomous driving. The method proposed in the present invention can indicate potential road condition risks, plan the current vehicle's driving route, and avoid potential risks ahead. Through the method proposed in the present invention, autonomous driving vehicles can perform self-road condition detection and simple risk avoidance with less human intervention.
[0006] A risk prediction method based on Granger causality test and interpretable reinforcement learning is proposed. First, the process of strengthening the fitness function and reward function based on the semantics of autonomous driving is studied. Then, a fuzzy correlation analysis method based on Granger causality is studied. Then, based on the reinforcement learning results, the Granger causality analysis is transformed. Finally, a reinforcement learning model for autonomous driving risk prediction based on Granger causality is proposed, and risk factor tracing based on time series data is realized. Specifically:
[0007] A risk prediction method based on Granger causality test and interpretable reinforcement learning is applied to autonomous driving of motor vehicles, comprising the following steps:
[0008] S1. Perform semantic analysis on autonomous driving risk data. After corpus-based semantic analysis, use fuzzy sets to represent user-defined goals. Then, perform preliminary quantification on the goals to obtain a driving risk fitness function based on fuzzy sets.
[0009] S2. According to the driving risk fitness function based on fuzzy sets obtained in step S1, combined with the relationship between the motor vehicle agent and the strategy set π, it is converted into the corresponding risk-related reward function R(s,a) in reinforcement learning, where:
[0010] s is a variable in the vehicle driving state set S, s∈S,
[0011] a is a variable in the driving action set A, a∈A;
[0012] For the final risk-related reward function R(s,a), the higher the function value, the safer the current vehicle state and action; conversely, the more dangerous the current driving state;
[0013] S3. Based on the theory of fuzzy sets, the autonomous driving factor space based on the reward function is constructed, and the driving risk factors are stored in the potential cause set to obtain the initial driving risk potential cause set X = {x1, x2, ..., x m}, where m is the total number of initial potential causes;
[0014] S4. Based on the initial set of potential driving risk causes X obtained in step S3, a driving risk fuzzy correlation analysis method based on time series is used, and the user defines the fuzzy support and membership. Then, a Granger causality hypothesis test method is used to obtain a preliminary screened set of potential driving risk causes X′={x′1, x′2,…, x′ m′}, where m′ is the total number of potential causes of driving risk after initial screening;
[0015] S5. Based on the positive environmental rewards resulting from the driving behavior strategy of each motor vehicle agent on the road, find the initial risk-related reward function R(s,a). This rule must satisfy the following: if a driving behavior strategy results in a positive reward in a low-risk environment, then the motor vehicle agent will be more likely to adopt this behavior strategy in the future. The goal of the motor vehicle agent is to find the optimal strategy in each discrete state to minimize the risk of autonomous driving.
[0016] S6, the initial risk-related reward function R(s,a) constructed in step S5 and the potential causes of driving risk X′={x′1, x′2,…, x′ m′}, initially build a reinforcement learning model for autonomous driving risk prediction;
[0017] S7. The autonomous driving risk prediction reinforcement learning model constructed in step S6 matches the potential driving risk causes X′ after preliminary screening with the fitness function driving risk Y according to the relationship between the initial risk-related reward function R(s,a) and the fitness function driving risk Y, and combines the potential driving risk causes X′ after preliminary screening with {x′1, x′2, …, x′ m′}, building a reinforcement learning model based on Granger causality of potential autonomous driving risk factors;
[0018] S8. Based on the Granger causality obtained by reinforcement learning, Granger causality judgment is performed on the time series data set, and the Granger cause X of the autonomous driving risk that is conducive to the fitness function driving risk Y is found;
[0019] S9. Predicting the driving risk Y of the fitness function based on the Granger cause X' of the autonomous driving risk found in step S8, and guiding and assisting the current motor vehicle to perform the safest driving behavior by increasing the proportion of factors in the Granger cause X'.
[0020] In step S7, the specific implementation steps of the reinforcement learning model based on the Granger causality of potential autonomous driving risk factors are as follows:
[0021] S7a: Screening reinforcement learning input feature dimensions based on the Granger causal factor library;
[0022] S7b: Calculate the value function v of each state based on the current state and strategy π;
[0023] S7c: Find the current optimal strategy according to the optimization algorithm and update the strategy π;
[0024] S7d: Iterate the policy evaluation process, repeating steps S7b-S7c. During this process, the value function v of the current state has been proven to be iteratively convergent. The iteration termination condition is to meet a user-defined value function difference, or the policy function does not change;
[0025] Step S7e: combining the association rules obtained from the fuzzy association analysis model based on Granger causality with the optimization strategy of reinforcement learning to perform corresponding analysis on the reinforcement learning results;
[0026] Step S7f: Assist decision making based on time series data according to the analysis of reinforcement learning.
[0027] In step S7d, the user-defined value function difference is set to [0.01-0.001].
[0028] The optimization algorithm in step S7c is a greedy algorithm.
[0029] This paper presents a risk prediction method based on Granger causality testing and interpretable reinforcement learning. It finds the corresponding Granger risk factors, breaks through the black-box approach, proposes a possible explanation, and predicts road condition risks in the short term, thereby significantly reducing road condition risks. This method has the following main advantages:
[0030] First, this invention can help identify reward functions based on uncertain driving actions and driving risk in autonomous driving. Based on time series data of driving uncertainty, it develops a reward function based on fuzzy set representation, which can be adjusted at any time during each optimization step. Its versatility and dynamism allow for faster convergence to the goal of low road risk.
[0031] The present invention then proposes an explanation for reinforcement learning through a driving risk association discovery method. The proposed method is applicable to reinforcement learning and deep learning methods that can quantify fitness functions. It is particularly applicable to time series data for autonomous driving risk prediction, particularly data with relatively stable distributions and linear relationships. It can initially address the interpretability issue of reinforcement learning and improve algorithmic efficiency during Granger screening.
[0032] (1) Efficiency: Granger causality analysis can be used to mine association rules based on Granger causality while ensuring a certain degree of accuracy. This method can be used to perform preliminary input feature screening before conducting driving risk reinforcement learning, reducing the training dimension of the reinforcement learning model and improving algorithm efficiency.
[0033] (2) Real-time: The time series data used to train the model can be updated in real time. According to the Granger time series data hypothesis testing method, the time-based autonomous driving data can be used to update the model at any time, thereby mining a more accurate and effective risk prediction reinforcement model, thereby assisting in achieving low road risk or even zero risk autonomous driving.
[0034] (3) Feasibility: The proposed method for constructing an enhanced model for autonomous driving risk prediction based on Granger causality analysis has good feasibility. The interpretability of the overall algorithm is achieved through association rules, and its reliability can be further tested through Granger causality hypothesis testing.
[0035] In this driving risk reinforcement learning method, fuzzy sets are used to assist semantic analysis to find association rules based on Granger causality and construct a corresponding reinforcement learning model, thereby realizing the construction of an autonomous driving risk reinforcement learning model based on Granger causality. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 4 is a flowchart of the risk prediction method of interpretable reinforcement learning based on Granger causality test described in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.
[0038] A risk prediction method based on interpretable reinforcement learning using Granger causality test is based on the following two points: first, the fitness function and reward function reinforcement process based on autonomous driving semantics are studied, then the fuzzy association analysis method based on Granger causality is studied, and then the Granger causality analysis is transformed according to the reinforcement learning results of autonomous driving risk prediction. Finally, an algorithm for the driving risk prediction reinforcement model based on Granger causality is proposed to achieve low-risk or even zero-risk autonomous driving technology.
[0039] This paper utilizes an association rule approach based on Granger causality analysis to address the unexplainable limitations of reinforcement learning algorithms. It then uses the discovered Granger causality to assist in decision-making, resulting in results that are more favorable to the user's predefined goals. Because the interpretability of the reinforcement learning process can greatly expand the algorithm's applicability, and its reward function also presents difficulties in defining, this paper is valuable in exploring solutions to address the unexplainable nature of reinforcement learning processes, the uncertainty of time series data-assisted decision-making, the quantification of reward functions, and the further tracing of black-box algorithms.
[0040] The present invention utilizes the temporal order of time series data and combines it with the hypothesis testing method of Granger causality analysis to process time series data with quantifiable targets. Through hypothesis testing, the causal relationship between the set of potential factors and the target is determined, which solves the problem of difficult tracing of reinforcement learning and realizes the establishment of an interpretable reinforcement learning model based on time series data.
[0041] The reinforcement learning modeling method can quickly and effectively establish a model to achieve target optimization based on user-defined goals and quantified reward functions, preparing for the prediction of possible future decision-making outcomes. By combining the results of Granger causality analysis and adjusting the input factors, it is possible to find the Granger causes closely related to the decision-making results while ensuring the advantage of the decision-making results.
[0042] In summary, the association rule method based on Granger causality analysis and the reinforcement learning process can be used to perform autonomous driving risk prediction modeling analysis.
[0043] The following describes the modeling process of the association rule method of Granger causality analysis, that is, how to further screen the Granger cause set according to the target-based association rules based on the Granger causality assumption. The details are as follows.
[0044] According to the potential cause, let X={x1,x2,…,x m} and the quantitative target Y. Assume that,
[0045]
[0046]
[0047] in, and Independent of each other.
[0048] At this time, two null hypotheses are made: H0: α1=α2=…=α q =0 and H′0: δ1=δ2=…=δ q =0.
[0049] Then, regress all the variables before the time series of y and calculate the residual sum of squares (RSS) y ; Similarly, the square difference corresponding to x is recorded as RSS x .
[0050] Through the F test, Here n is the total number of samples, q is the number of regression training (to be evaluated), and k is the number of regression training (to be evaluated) corresponding to y.
[0051] In this case, if F>F(q,(nk)), then the null hypothesis is rejected, and it can be said that x is the Granger cause of y. Note that Granger cause is not necessarily a strong causal relationship.
[0052] The analytical method of the present invention is as follows:
[0053] 1. Construct a fitness function and a reward function. Perform semantic analysis on user-defined goals, fuzzify and quantize time series data, and obtain the fitness function driving risk Y and the risk-related reward function R(s,a), and find their corresponding relationship.
[0054] 2. Establish a fuzzy correlation analysis model based on Granger causality. Using a fuzzy correlation analysis method based on time series, we construct a correlation model based on Granger causality analysis, screen the target cause candidate library, and establish the relationship between factors and targets.
[0055] 3. Build a reinforcement learning model based on Granger causality. Build a preliminary association rule mining model. Process multi-source data separately, calculate their data distributions, and then build a preliminary association rule mining model.
[0056] The specific implementation steps of the Granger causality-based autonomous driving risk prediction reinforcement learning model are as follows:
[0057] Step a: Screen reinforcement learning input feature dimensions based on the Granger causal factor library.
[0058] Step b: Calculate the value function v of each state based on the current state and strategy π.
[0059] Step c: Find the current optimal strategy based on the greedy algorithm and update π.
[0060] Step d: Iterate the policy evaluation process, repeating steps b-c until the value function v of the current state has been proven to converge. The iteration terminates when a small difference in the value function is satisfied, or when the policy function remains unchanged.
[0061] Step e: Combine the association rules obtained from the fuzzy association analysis model based on Granger causality with the optimization strategy of reinforcement learning to perform corresponding analysis on the reinforcement learning results.
[0062] Step f: Make auxiliary decisions based on time series data according to the analysis of reinforcement learning.
[0063] This paper proposes a risk prediction method based on Granger causality test and interpretable reinforcement learning. First, the method studies the semantic fitness function and reward function reinforcement process, then studies the fuzzy association analysis method based on Granger causality. Then, based on the reinforcement learning results, the Granger causality analysis is transformed. Finally, an algorithm based on Granger causality reinforcement learning model is proposed to realize auxiliary decision-making of time series data. Specifically:
[0064] (1) In view of the uncertainty of the transformation of fitness function and reward function in the reinforcement learning process, the research supports the transformation method of fitness function and reward function under the conditions of semantic ambiguity and target uncertainty, so that it can find the fitness function and reward function that are closest to the actual target and its needs based on the actual needs while retaining more known information. Through the quantitative transformation process of fuzzy sets, the uncertainty problem of target and reward function in reinforcement learning is solved.
[0065] (2) In view of the difficulty of reinforcement learning analysis, the association rule discovery process is added to make it possible to perform reverse interpretation by combining existing association rules with reinforcement learning results. This method makes the interpretation of reinforcement learning simpler and highlights the advantages of fuzzy set semantic analysis.
[0066] (3) In response to the causal requirements of reinforcement learning analysis, a Granger causality hypothesis testing process is added to conduct Granger causality analysis on all potential analytical causes in advance. At the same time, based on the analysis results, the input feature range of some reinforcement learning is further screened and strengthened, reducing the reinforcement learning dimension. This increases theoretical support for its analytic behavior on the basis of improving learning efficiency.
[0067] The above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed in the present invention should be included in the protection scope recorded in the claims.
Claims
1. A risk prediction method based on Granger causality test and interpretable reinforcement learning, applied to autonomous driving of motor vehicles, characterized by: The following steps are involved: S1. Perform semantic analysis on autonomous driving risk data. After corpus-based semantic analysis, use fuzzy sets to represent user-defined goals. Then, perform preliminary quantification on the goals to obtain a driving risk fitness function based on fuzzy sets. S2, according to the driving risk fitness function based on fuzzy sets obtained in step S1, combined with the motor vehicle agent and the strategy set The relationship between them is converted into the corresponding risk-related reward function in reinforcement learning. ,in, s is a variable in the vehicle driving state set S, , a is a variable in the driving action set A, ; For the final risk-related reward function Generally speaking, the higher the function value, the safer the current vehicle state and action; conversely, the more dangerous the current driving state; S3. Construct the autonomous driving factor space based on the reward function by using the fuzzy set theory, save the driving risk factors in the potential cause set, and obtain the initial driving risk potential cause set. ,in is the total number of potential causes at the outset; S4: The initial driving risk potential cause set obtained in step S3 Through the fuzzy correlation analysis method of driving risk based on time series, user-defined fuzzy support and membership, and Granger causality hypothesis test method, the potential cause set of driving risk after preliminary screening is obtained. ,in, The total number of potential reasons for driving risk after initial screening; S5. Find the initial risk-related reward function based on the positive reward of the environment caused by the driving behavior strategy of each motor vehicle agent on the road ,The rule needs to be satisfied: if a certain driving behavior strategy leads to a positive reward in a low-risk environment, then the tendency of the motor vehicle agent to produce this behavior strategy in the future will be strengthened. The goal of the motor vehicle agent is to find the optimal strategy in each discrete state to make autonomous driving less risky; S6. The initial risk-related reward function constructed according to step S5 Potential causes of driving risk after the preliminary screening obtained in step S4 , preliminarily build a reinforcement learning model for autonomous driving risk prediction; S7: The autonomous driving risk prediction reinforcement learning model constructed in step S6 is used according to the initial risk-related reward function Driving Risk with Fitness Function relationship, matching the potential causes of driving risk after the initial screening Driving risk with fitness function , combined with the potential causes of driving risk after preliminary screening Construct a reinforcement learning model based on Granger causality of potential autonomous driving risk factors; S8. Based on the Granger causality obtained by reinforcement learning, Granger causality judgment is performed on the time series data set, and the driving risk that is beneficial to the fitness function is found. Granger causes of autonomous driving risks ''; S9. Granger causes of the autonomous driving risk found in step S8 '', driving risk for fitness function Make predictions and by adding Granger causes ''The proportion of these factors guides and assists the current motor vehicle to make the safest driving behavior.
2. The risk prediction method based on Granger causality test and interpretable reinforcement learning according to claim 1, characterized in that: In step S7, the specific implementation steps of the reinforcement learning model based on the Granger causality of potential autonomous driving risk factors are as follows: S7a: Screening reinforcement learning input feature dimensions based on the Granger causal factor library; S7b: Based on current status and strategy , calculate the value function v of each state; S7c: Find the current optimal strategy based on the optimization algorithm and update the strategy ; S7d: Iterate the policy evaluation process, repeating steps S7b-S7c. During this process, the value function v of the current state has been proven to be iteratively convergent. The iteration termination condition is to meet a user-defined value function difference, or the policy function remains unchanged; Step S7e: combining the association rules obtained from the fuzzy association analysis model based on Granger causality with the optimization strategy of reinforcement learning to perform corresponding analysis on the reinforcement learning results; Step S7f: Assist decision making based on time series data according to the analysis of reinforcement learning.
3. The risk prediction method based on Granger causality test and interpretable reinforcement learning according to claim 2, characterized in that: In step S7d, the user-defined value function difference is set to [0.01 - 0.001].
4. The risk prediction method based on Granger causality test and interpretable reinforcement learning according to claim 2, characterized in that: The optimization algorithm in step S7c is a greedy algorithm.
Citation Information
Patent Citations
Causal relationship mining method based on deep learning
CN109993281A
Causal network learning method based on local Granger causal analysis
CN114036736A