A method for intelligent vehicle fault identification based on minimum gray box

By constructing a minimal gray-box variable set and a dynamic Bayesian network, the challenges of data processing and privacy protection in autonomous vehicles are solved, enabling transparent fault tracing and real-time detection, and supporting the safety supervision of high-level autonomous vehicles.

CN120850094BActive Publication Date: 2026-04-03TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for high-level autonomous vehicles face challenges such as difficulties in data collection and processing, privacy protection, fault tracing, and poor interpretability of causal relationships, making it difficult to effectively implement sandbox regulation.

Method used

A minimum gray box-based intelligent vehicle fault identification method is adopted. By constructing a Bayesian network and a fault tree model, the minimum cut subset is determined. Combined with Markov blanket screening and dynamic Bayesian network, the fault state detection and type judgment are realized.

Benefits of technology

It achieves significant data compression and reduces the risk of privacy leaks, makes fault tracing transparent and visible, improves the real-time performance and accuracy of fault identification, and supports accountability for sandbox supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850094B_ABST
    Figure CN120850094B_ABST
Patent Text Reader

Abstract

This invention discloses a method for intelligent vehicle fault identification based on a minimum gray box, comprising: acquiring intelligent vehicle fault injection data with fault classification labels; constructing a Bayesian network using a structural learning method to form a directed acyclic graph (DAG) reflecting the causal dependencies of variables; mapping the DAG to a fault tree model to establish a correspondence structure between fault events and logical relationships; determining the minimum cut subset through the fault tree model; and extracting a minimum and sufficient set of key variables using Markov blanket screening to form a minimum gray box; using the DAG as structural input and combining it with fault injection data for parameter learning to construct a dynamic Bayesian network; and inputting the minimum gray box variable set into the network to achieve the detection of intelligent vehicle fault states and the determination of fault types. This invention solves the problems of large data redundancy, strong black-box nature, and insufficient causal explanation in existing technologies through minimum gray box construction and dynamic Bayesian network inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving technology, specifically relating to an intelligent vehicle fault identification method based on a minimum gray box. Background Technology

[0002] The development of Highly Automated Vehicles (HAVs) is progressing rapidly, demonstrating enormous application potential in intelligent transportation and mobility services. However, because autonomous driving systems rely on multi-source sensor fusion, complex deep learning algorithms, and high-precision control modules, they are prone to potential safety hazards and various malfunctions during operation. How to effectively identify and locate these malfunctions, and ensure the safety of vehicles in open road testing and real-world applications, has become a core issue in current research and regulation.

[0003] To address these challenges, existing technologies have introduced a sandbox regulatory model. This mechanism aims to construct a data closed loop of safety supervision, maintenance, and continuous optimization through two-way data feedback from open road testing and virtual simulation testing, thereby driving a shift in regulatory models from passive to proactive oversight. In this process, regulators and companies form a collaborative regulatory framework to achieve in-depth safety testing of autonomous driving systems and transparent, penetrating fault attribution.

[0004] However, existing technologies still have the following problems in practical applications:

[0005] First, data collection and processing present challenges. Each highly automated vehicle can generate approximately 3-6TB of raw data per hour, making it difficult for companies and regulators to fully retain and transmit this data during storage and transmission. Furthermore, the high-precision fault tracing and liability determination required by sandbox regulation often involve in-depth causal analysis, which not only burdens data asset management but also risks data privacy and intellectual property breaches. Therefore, traditional in-vehicle event data recorders (EDRs) and data storage systems for Automated Driving (DSSADs) are no longer sufficient to meet the needs of sandbox regulation.

[0006] Secondly, the black-box problem makes fault tracing difficult. Autonomous driving systems commonly contain numerous deep learning models and complex modules, whose internal operating mechanisms are highly opaque. Once a vehicle malfunctions, the fault propagation path often traverses multiple black-box stages, making it difficult for regulators to accurately determine the specific source of the fault and the responsible party. This hinders the identification and feedback of technical defects, and impedes companies' timely iteration, optimization, and repair of the system.

[0007] Existing research and patents have made significant progress in causal discovery and fault analysis, but key limitations remain. For example, Chinese patent CN112990467A discloses a method for automobile fault analysis based on generative Bayesian causal networks, with the following steps: 1) Obtain automobile fault data; 2) Establish a representative subset of automobile fault samples; 3) Establish a feature selection model HSICLasso; 4) Construct a generative Bayesian function causal network model; 5) Input the feature vector set v into the generative Bayesian causal network model and output a causal directed graph; 6) Calculate the model evaluation index of the causal directed graph. If the model evaluation index meets the requirements, output the current causal directed graph; otherwise, update the weights and offsets of the generative Bayesian causal network model and return to step 5).

[0008] However, this method has the following limitations: 1. Although it can construct a causal directed graph, it does not propose how to determine a minimum and sufficient set of variables to ensure that fault tracing and liability determination can still be supported while reducing data backhaul and protecting privacy. 2. This patent directly outputs a causal directed graph, lacking interpretable application tools in regulatory scenarios. 3. This patent relies on feature selection methods such as HSIC Lasso, lacking a mathematically provable optimal variable selection mechanism.

[0009] In summary, how to ensure data sufficiency while protecting privacy, and how to build a minimal yet necessary data collection and modeling system, has become a key issue in achieving efficient fault detection, transparent tracing and accountability, and supporting sandbox regulatory tasks. Summary of the Invention

[0010] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a method for intelligent vehicle fault identification based on a minimum gray box.

[0011] The objective of this invention can be achieved through the following technical solutions:

[0012] This invention provides a method for intelligent vehicle fault identification based on a minimum gray box, comprising the following steps:

[0013] Acquire fault injection data for intelligent vehicles with fault classification labels;

[0014] Based on the fault injection data, a Bayesian network is constructed using a structural learning method to obtain a directed acyclic graph representing the causal dependencies of variables;

[0015] The obtained directed acyclic graph is mapped to a fault tree model to obtain the corresponding structure of fault events and logical relationships;

[0016] Based on the fault tree model, the minimum cut subset is determined, and combined with Markov blanket filtering of variables, the minimum gray box variable set is obtained.

[0017] Using a directed acyclic graph as structural input and combining fault injection data for parameter learning, a dynamic Bayesian network is constructed.

[0018] The minimum gray box variable set is input into the constructed dynamic Bayesian network to detect the fault status of intelligent vehicles and determine the fault type.

[0019] Furthermore, the fault injection data includes injection data for the same vehicle under different faults. The injection data includes fault data and non-fault data, and the fault injection data is represented as follows:

[0020]

[0021] in, This indicates non-fault data. Indicates fault data, Injecting data to the fault, including n 1 sample, each sample includes m One variable, m The variables include fault data. Non-fault data Each sample corresponds to a fault classification label, which is used to characterize the fault type or non-fault state to which the sample belongs.

[0022] Furthermore, based on the fault injection data, a Bayesian network is constructed using a structure learning method to obtain a directed acyclic graph representing the causal dependencies of variables, specifically including:

[0023] For each type of variable in the aforementioned fault injection data Initialize its candidate parent-child node set:

[0024]

[0025] in, Indicates the first i Class variables, Indicates the first i The set of candidate parent and child nodes for class variables. Indicates the empty set;

[0026] Inject faults into each type of variable in the data. Mapped to nodes in a Bayesian network :

[0027] For except external candidate variable set Each variable in Based on the variable vector composed of samples in the fault injection data and Estimate marginal and joint probabilities and calculate and Dependence strength ,in, This represents the universal set of variables, that is, the set of all variables in the fault injection dataset. Indicates the first i Class variables, Indicates the first j Class variables, Indicates the first n The first sample i The value of a class variable, Indicates the first n The first sample j The value of the class variable, T represents transpose;

[0028] when At that time, join in ,in, Preset dependency threshold;

[0029] right Each candidate parent-child node in Perform a conditional independence test: determine whether a subset exists. So that in a given under conditions and Conditions are independent, that is, they satisfy:

[0030]

[0031] in, Indicates conditional independence. This represents a subset of variables used for conditional independence testing;

[0032] When the conditions are met, from Remove and reverse prune the nodes; if the condition is not met, add undirected edges to the Bayesian network. The reverse pruning refers to when... When there are undirected edges, Delete the undirected edge between them;

[0033] right all variables After processing, the initial undirected graph is obtained. ;

[0034] Based on the initial undirected graph A directed acyclic graph is constructed using a greedy hill-climbing directional search method based on a scoring function.

[0035] Furthermore, the dependency strength is expressed by the following formula:

[0036]

[0037] in, Representing variables With variables The strength of the dependency between them n Represents the total number of samples. , They represent the first k The first sample i Class variables and the first j The value of a class variable, Indicates the first sample based on all samples. i Class variables in taking values The empirical probability at that time; Indicates the first sample based on all samples. j Class variables in taking values The empirical probability at that time, Indicates the pair of variables based on all sample pairs. In the k Sample values The joint empirical probability.

[0038] Furthermore, the empirical probability is formulated as follows:

[0039]

[0040] in, This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. Indicates the first t The first sample i The value of the class variable;

[0041] The joint empirical probability is given by the following formula:

[0042]

[0043] in, Indicates the pair of variables based on all sample pairs. In the k Sample values The joint empirical probability.

[0044] Furthermore, the initial undirected graph is... A directed acyclic graph is constructed using a greedy hill-climbing directional search method based on a scoring function, specifically including:

[0045] The initial undirected graph The skeleton is defined as the candidate adjacency set, and the skeleton is... The set of all undirected edges in the skeleton, and during the directed search process, structural transformations are only allowed on the adjacencies within the skeleton, and the introduction of new edges is prohibited. New adjacency outside the skeleton;

[0046] The initial undirected graph As the current network structure ;

[0047] For the current network structure Perform a greedy hill-climbing directed search: for any undirected edge in the candidate adjacency set... Perform one of the following random operations:

[0048] Introducing edges: at nodes and Add a directed edge between ;

[0049] Delete edge: Delete a directed edge from the current network structure. ;

[0050] Reverse edge: Reverse a directed edge in the current network structure Reversed to ;

[0051] Calculate the Bayesian Information Criterion (BIC) score for the graph after each network structure transformation;

[0052] If the BIC score improves compared to the current network structure, the current transformation is accepted and used as the current network structure for the next iteration; if the score does not improve, the transformation is discarded and the current network structure remains unchanged.

[0053] Repeat the iterations until the BIC score of the scoring function no longer improves in consecutive iterations or the preset number of iterations is reached, thus obtaining a directed acyclic graph. .

[0054] Furthermore, the scoring function is defined by the following formula:

[0055]

[0056] in, This represents the BIC score, This indicates fault injection data. Indicates the current network structure. Indicates the corresponding current network structure The set of parameters, This indicates the total number of parameters in the parameter set. Represents the total number of samples. Indicates the current network structure and parameter set Below, fault injection data The log-likelihood.

[0057] Furthermore, mapping the obtained directed acyclic graph to a fault tree model to obtain the corresponding structure of fault events and logical relationships specifically includes:

[0058] In the directed acyclic graph In this context, a node is defined as a root node if it has only a directed edge pointing to other nodes and no directed edge pointing to itself; a leaf node if it has only a directed edge pointing to itself and no directed edge pointing to other nodes; and an intermediate node if it has both directed edges pointing to itself and directed edges pointing to other nodes.

[0059] Map the root node to the base event in the fault tree, the intermediate nodes to intermediate events, and the leaf nodes to the top event, where:

[0060] The base event corresponding to the root node serves as the basic triggering factor for generating intermediate or top events in the fault tree.

[0061] The intermediate events corresponding to intermediate nodes are triggered by the base events or intermediate events of all parent nodes that directly point to that node;

[0062] The top event corresponding to a leaf node is triggered by the base event or intermediate event of all parent nodes that directly point to that node;

[0063] The directed edges in the directed acyclic graph are mapped to logic gates in the fault tree, and the logical relationships are determined by calculating the trigger probability of base events on intermediate or top events.

[0064]

[0065] in, Represents the base event set The probability of both triggering an intermediate event or a top event; Indicates data injection during fault diagnosis D Lower base event set The expected probability of triggering the event; This represents the total number of subnetworks sampled from a directed acyclic graph. Indicates the sampling number b The set of base events for each subnetwork; Indicates the sampling number b Sub-network structure; Indicates data injection during fault diagnosis D The first sampling b The probability of each subnetwork;

[0066] when At that time, the base event set is considered Together, they trigger the occurrence of corresponding intermediate or top events, constructing the corresponding logic gates in the fault tree.

[0067] Furthermore, based on the fault tree model, determining the minimum cut subset and combining it with Markov blanket filtering of variables to obtain the minimum gray box variable set specifically includes:

[0068] The universal set of variables X Each variable in In a given set of variables and Under the given conditions, perform Markov blanket screening to determine whether the conditional independence is satisfied:

[0069]

[0070] in, Indicates the first i Class variables, Represents the set of candidate Markov blanket variables; To express conditional independence, it means that given... Under the conditions, Except for the whole set and Other variables are independent of each other;

[0071] If conditional independence holds, then As a variable Initial screening variable set ;

[0072] Based on the initial screening variable set Represent all paths in the fault tree that trigger the top event as a minimum cut subset:

[0073]

[0074] in, Represents the set of triggering chains for the top event; This represents the total number of minimum cut subsets. Indicates the first k The set of indices in the minimum cut subset Indicates the first i The initial selection set of variables for class variables; This represents a logical AND operation, indicating that the top event is triggered by the combined action of all variables within the minimum cut subset. Represents a logical OR operation, indicating that any set of variables from different minimum cut subsets triggers a top event;

[0075] The sets of variables in all minimum cut subsets are merged to form the final minimum gray box variable set. :

[0076]

[0077] in, This represents the smallest set of gray box variables.

[0078] Furthermore, the construction of a dynamic Bayesian network by using a directed acyclic graph as structural input and combining it with fault injection data for parameter learning specifically includes:

[0079] Using a directed acyclic graph as the structural input to a dynamic Bayesian network, parameter estimation is performed on the fault injection data D, and the optimal parameters of the dynamic Bayesian network are obtained using maximum likelihood estimation.

[0080]

[0081] in, n Represents the total number of samples. Indicates the first i All variable vectors of the sample, Represents the sample mean vector. This represents the state mean parameter estimate for each variable in a dynamic Bayesian network; This represents the estimation of the covariance parameter between variables in a dynamic Bayesian network.

[0082] Compared with the prior art, the present invention has the following advantages:

[0083] (1) In the prior art, high-level autonomous vehicles can generate 3-6 TB of raw data per hour during operation, which is difficult for enterprises and regulators to store and transmit completely, causing data asset management pressure and potentially leading to privacy leakage risks. This invention proposes a method for constructing a minimum gray box variable set, combined with fault tree model and Markov blanket screening, to extract only the minimum and sufficient set of key variables, avoiding the transmission and storage of redundant data. This solves the problem of balancing data processing and privacy protection in the prior art, and achieves the technical effect of significantly compressing data volume and reducing the risk of privacy leakage.

[0084] (2) In the prior art, there are a large number of deep learning black box modules inside the autonomous driving system. The fault transmission path often spans multiple links, making it difficult for regulators to accurately locate the fault source, resulting in difficulties in tracing the source and assigning responsibility. This invention maps the directed acyclic graph obtained by structure learning into a fault tree model, and further combines it with minimum cut subset analysis to form a corresponding structure of fault events and logic gates, intuitively showing the fault triggering link, solving the problem of the difficulty in explaining causal relationships in the prior art, and realizing the transparency and visualization of fault tracing.

[0085] (3) In the prior art, existing studies rely on feature selection methods such as HSIC Lasso, but lack a theoretically provable optimal variable selection mechanism, making it difficult to guarantee the sufficiency and minimumity of the selected variable set. This invention selects variables by using conditional independence tests based on Markov blankets, and strictly determines the minimum set of variables most relevant to the causal relationship of the fault according to the theory of probabilistic graphical models. This solves the problems of arbitrary and unreliable variable selection in the prior art, and achieves the technical effect of significantly reducing redundant variables while ensuring the sufficiency of the analysis.

[0086] (4) In the prior art, generative causal network methods are mostly static models that do not consider the temporal dependencies in the operation of intelligent vehicles, making it difficult to achieve dynamic detection and prediction of faults. This invention constructs a dynamic Bayesian network by learning parameters based on a Bayesian network structure and inputting the minimum gray box variable set into it, which solves the problem of lack of temporal fault modeling in the prior art, realizes dynamic detection and type judgment of fault states, and improves the real-time performance and accuracy of fault identification in intelligent vehicles.

[0087] (5) In the prior art, the generation process of causal directed graphs relies on generative training, and model evaluation metrics (such as AUPR, SHD, SID) are difficult to directly convert into causal paths or responsibility links required by regulators, resulting in results that are difficult to meet the requirements of sandbox regulation. This invention, by using fault injection data as a basis and combining Bayesian network structure learning and fault tree logic mapping, can not only obtain causal dependencies, but also clarify the minimum cut subset triggering path, solving the problem of the disconnect between causal model results and regulatory tasks in the prior art, and achieving the technical effect of efficiently supporting responsibility tracing and sandbox regulation. Attached Figure Description

[0088] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;

[0089] Figure 2 This is a flowchart illustrating the intelligent vehicle fault detection method according to an embodiment of the present invention;

[0090] Figure 3 A schematic diagram illustrating the construction of the minimum gray box according to an embodiment of the present invention;

[0091] Figure 4 This is a schematic diagram of a fault tree according to an embodiment of the present invention. Detailed Implementation

[0092] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0093] Example 1:

[0094] This embodiment provides a method for intelligent vehicle fault identification based on a minimum gray box, such as... Figure 1 As shown, it includes the following steps:

[0095] Step S1: Obtain intelligent vehicle fault injection data with fault classification labels;

[0096] Step S2: Based on the fault injection data, a Bayesian network is constructed using a structural learning method to obtain a directed acyclic graph representing the causal dependencies of variables;

[0097] Step S3: Map the obtained directed acyclic graph to a fault tree model to obtain the corresponding structure of fault events and logical relationships;

[0098] Step S4: Based on the fault tree model, determine the minimum cut subset and combine it with Markov blanket filtering of variables to obtain the minimum gray box variable set;

[0099] Step S5: Using a directed acyclic graph as the structural input, combine it with fault injection data to learn parameters and construct a dynamic Bayesian network;

[0100] Step S6: Input the minimum gray box variable set into the constructed dynamic Bayesian network to detect the fault status of the intelligent vehicle and determine the fault type.

[0101] Example 2:

[0102] like Figure 2 As shown, this embodiment provides a minimum gray box-based intelligent vehicle fault identification method. Through three core steps—fault injection data acquisition, minimum gray box construction, and fault detection based on a minimum gray box-dynamic Bayesian network—it achieves gray box-based sandbox supervision of black-box autonomous driving systems. The specific process is as follows:

[0103] 1. Fault Injection Data Acquisition

[0104] The fault injection data collection step in this embodiment first constructs a joint simulation framework of Apollo and VTD based on the concept of fault injection: taking the Apollo autonomous driving system as the monitored object and VTD as the traffic simulation environment, traffic scenarios generated using the HighD and Intersction datasets are used to inject multi-dimensional faults. Based on the aforementioned multi-dimensional faults, the fault space needs to be defined as follows:

[0105]

[0106] In the formula, It is the fault space. It is the location of the fault (covering Apollo's five modules: perception, planning, localization, prediction, and control). This refers to the duration of the fault (lasting 1-6 frames). It is the fault injection time (any moment in the entire test cycle). It is a fault injection value (deviating from the normal value).

[0107] 2. Minimum Gray Box Construction

[0108] like Figure 3 As shown, the construction of the minimum gray box follows a three-level filtering mechanism: first, irrelevant variables that have no causal dependency on the core fault are eliminated; then, privacy variables are embedded in the causal chain as intermediate nodes; finally, sufficient variables are selected to form the main body of the minimum gray box, while the remaining observable variables are included in the intermediate nodes of the causal chain, thus completing the minimum gray box architecture.

[0109] The process of constructing a minimal gray box includes the following:

[0110] Step A1: Construct a directed acyclic graph:

[0111] Bayesian networks employ a directed acyclic graph-based approach. A network structure, in which each vertex Corresponding to having parameters random variables X i It has parameters. global probability distribution X Based on the arc set in the graph a ij ∈A It can be decomposed into several local probability distributions. Given a vertex set... V (correspond{ X i}), the figure G By arc set A Uniquely determined. Arc set A The directed arc in the diagram represents the dependency relationship between vertices, where v i →vj express v i for v j The parent node, v j These are child vertices. In the causal structure network to be constructed, each vertex represents a variable, and each directed arc represents a conditional dependency. These directed arcs define the conditional probability distribution. P(X i ) Thus, the global distribution is derived. P(X) :

[0112]

[0113] In the formula ΠX i Represents X i The parent node group.

[0114] To train and obtain a directed acyclic graph (DAG), fault-injected data and prior probabilities are used to establish relationships between variables. This task employs the Max-Min Hill-Climbing algorithm, which combines the Max-Min Parents and Children algorithm to significantly reduce the search space and efficiently obtain results, achieving faster optimization within a finite candidate space. The Max-Min Hill-Climbing algorithm process includes the following:

[0115] Step A1-1: Use the Max-Min Parents and Children (MMPC) method for skeleton learning. After initializing each variable, output the variable with the strongest dependency relationship in the forward direction, and then prune in the reverse direction to eliminate spurious dependency variables. Finally, after ensuring consistency, output an undirected acyclic graph.

[0116] Step A1-2: Select the BIC criterion as the scoring function:

[0117]

[0118] in, This represents the BIC score, This indicates fault injection data. Indicates the current network structure. Indicates the corresponding current network structure The set of parameters, This indicates the total number of parameters in the parameter set. Represents the total number of samples. Indicates the current network structure and parameter set Below, fault injection data The log-likelihood.

[0119] Step A1-3: The Hill-Climbing (HC) algorithm is used for directional optimization. By greedily searching, the algorithm attempts to introduce, delete, and reverse edges to obtain the optimal situation of the scoring function, thus obtaining a directed acyclic graph.

[0120] Step A2: Convert the directed acyclic graph into a fault tree

[0121] The conversion process includes the following steps:

[0122] Step A2-1: Map the root node of the directed acyclic graph (DAG) to a base event in the fault tree, map the intermediate nodes of the DAG to intermediate events in the fault tree, and map the leaf nodes of the DAG to the top events in the fault tree. If the root node is connected to multiple leaf nodes, represent them separately in the fault tree using the corresponding base events.

[0123] Step A2-2: Map the edges of the directed acyclic graph to the logic gates of the fault tree, for each node connected to its child nodes. b=1,2……B || {e i ∈E b } Sure e i If an edge belongs to all its child nodes, assign it a value of 1; otherwise, assign it a value of 0. D Indicates source data, G The formula for calculating the relationship strength of a directed acyclic graph is as follows:

[0124]

[0125] in, Represents the base event set The probability of both triggering an intermediate event or a top event; Indicates data injection during fault diagnosis D Lower base event set The expected probability of triggering the event; This represents the total number of subnetworks sampled from a directed acyclic graph. Indicates the sampling number b The set of base events for each subnetwork; Indicates the sampling number b Sub-network structure; Indicates data injection during fault diagnosis D The first sampling b The probability of each subnetwork;

[0126] If the relationship strength is greater than 0.5, it can be considered that the basic variables jointly triggered the occurrence of the intermediate event.

[0127] Step A3: Solve for the minimum cut subset to obtain the minimum gray box.

[0128] Based on the fault tree obtained above, as follows Figure 4 The process for finding the minimum cut subset includes the following:

[0129] Step A3-1: Select a Markov Blanket (MB) for each variable to filter different variable structures within the complete set of random variables. U In the context of a given variable X∈U and variable set MB∈U(X∉MB) If the following formula exists:

[0130]

[0131] in, Indicates the first i Class variables, Represents the set of candidate Markov blanket variables; To express conditional independence, it means that given... Under the conditions, Except for the whole set and Other variables are independent of each other;

[0132] The smallest set of variables, MB, that satisfies the above conditions is called the initial screening variable.

[0133] Step A3-2: Select the minimum cut subset from the initial screening variables. It is defined as the smallest set of base events that can cause the top event to occur. Finally, use it as the minimum gray box variable.

[0134] Third, implement fault detection based on minimum gray box-dynamic Bayesian network.

[0135] A dynamic Bayesian network is chosen for fault detection. By introducing time dependencies, it extends the theoretical framework of Bayesian networks, enabling the systematic construction of network structures from time-series data. This extension can model the dynamic evolution patterns between system variables, capturing three key dimensions: 1) causal relationships between system components; 2) temporal evolution of variable states; and 3) conditional dependencies across time steps. The process of constructing a dynamic Bayesian network based on the minimum gray box and directed acyclic graph is as follows:

[0136] Step A4-1: Use the directed acyclic graph obtained in step A2 as the output of structure learning.

[0137] Step A4-2: Obtain the optimal parameters of the dynamic Bayesian network using maximum likelihood estimation.

[0138]

[0139] in, n Represents the total number of samples. Indicates the first i All variable vectors of the sample, Represents the sample mean vector. This represents the state mean parameter estimate for each variable in a dynamic Bayesian network; This represents the estimation of the covariance parameter between variables in a dynamic Bayesian network.

[0140] Step A4-3: Use the smallest gray box as the input to the dynamic Bayesian network to predict the fault condition of the next frame of data and detect the fault type.

[0141] Specifically, 30% of the sample set was selected as the training set. The construction process included: removing duplicate records, filtering out Apollo system crash data, excluding non-startup state data, and selecting samples with a simulation duration exceeding one minute. For each module, an independent directed acyclic graph was constructed based on its corresponding training set. The resulting fault tree can completely represent the causal chain. The specific fault tree structure is as follows: Figure 4 As shown.

[0142] The minimal cut subset method aims to construct a minimal subsystem reflecting the system state, which is the minimum combination of basic events leading to the top event of the fault tree, i.e., a minimal gray box. Each minimal cut subset represents a potential path triggering the top event. Based on the calculated relationship strength, the relationship strength between each basic event can be calculated. If the propagation path from the basic event BE to the top event TE is BE→A→B→C→…→TE, then The calculation formula is as follows:

[0143]

[0144] After obtaining the trigger probability of the top event for each basic event, the probability of the top event being triggered by the n minimum cut subsets. The calculation formula is as follows:

[0145]

[0146] This example uses three metrics—Criticality (CRIT), Risk Increase Weight (RAW), and Risk Reduction Weight (RRW)—to quantify the impact of each basic event on system reliability under potential failure conditions. Criticality (CRIT) characterizes the probability that a specific event is the root cause of system failure; Risk Reduction Value (RAW) measures the hypothetical components... i The increase in system risk when in a faulty state; Risk Achievement Value (RRW) measures the increase in risk when a component is in a faulty state.i The potential for reducing system risk when considered 100% reliable. The specific calculation formulas for the three are as follows:

[0147]

[0148]

[0149]

[0150] In the formula, i For the first i A basic event, j Take from 1 N , N For the first i The number of minimal cut subsets to which each basic event belongs.

[0151] In this example, fault tree analysis was used to identify the defects in Apollo, as shown in Table 1. The analysis is as follows:

[0152] (1) In terms of Risk Increase Weight (RAW) assessment, the three variables—autonomous vehicle throttle control signal (ego_control_throttle, RAW=1.109), background vehicle position prediction value (car_1_predict_x, RAW=1.009), and background vehicle presence probability (car_1_predict_probability, RAW=0.525)—showed high risk sensitivity. Among them, two variables were directly related to the background vehicle perception module, indicating that perception system anomalies were the main cause of failure, and the failure of related variables would significantly increase the system risk level.

[0153] (2) In the Risk Reduction Potential (RRW) analysis, the background vehicle acceleration parameter (car_1_acc_x) showed an astonishing RRW value of 75.673. Since the background vehicle acceleration not only directly affects the path planning decision of the master vehicle, but also makes prediction difficult due to its randomness, this parameter should be set as the first optimization target.

[0154] (3) Based on the comprehensive evaluation of various indicators, the autonomous vehicle's planned heading (ego_plan_heading) and the predicted location of background vehicles (car_1_predict_x) both performed exceptionally well across multiple risk indicators. The former determines the direction of the vehicle's journey in the future, while the latter provides dynamic information about the surrounding vehicles. Together, they form the basis for the planning and decision-making of the autonomous driving system. Ensuring the stability of these two parameters can significantly enhance the overall reliability of the system.

[0155] Table 1 Reliability Analysis of Minimum Gray Box Variables

[0156]

[0157] In summary, this embodiment completes the minimum gray box modeling of the Apollo system through fault tree analysis, achieving dual optimization of variable dimensions and data storage: the original 79 variables are simplified to 18 core variables, and 61 redundant variables are eliminated; the data storage volume is compressed from an average of 1.9TB per vehicle per day (full storage of floating-point numbers) to 0.6TB, a reduction of 68.4%.

[0158] Specifically, the obtained directed acyclic graph is used as the structural input, and the parameter learning is further completed using the training set. Only the smallest gray box is used as the input variable to train a dynamic Bayesian network to achieve fault detection.

[0159] Four classic classification models—XGB, SVM, KNN, and DT—were selected for comparison with the Dynamic Bayesian Network (DBN). The fault detection results obtained using the remaining data as the test set are shown in Table 2. The data indicates that among the 20 core indicators of the five modules of the Apollo system, DBN performed best in 15 individual evaluations, highlighting the accuracy and robustness of this method in fault detection.

[0160] Table 2 Comparison of Fault Detection Results

[0161]

[0162] In summary, this intelligent vehicle sandbox regulatory method based on the minimum gray box balances the sufficiency and privacy of high-dimensional data, transforms the black-box state of autonomous driving systems into a transparent and governable model from a data perspective, and achieves module-level fault tracing based on the minimum gray box, making responsibility tracing quantifiable and verifiable. Overall, it constructs a data-closed-loop sandbox regulatory paradigm of safety supervision, maintenance and protection, and continuous optimization.

[0163] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0164] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for intelligent vehicle fault identification based on a minimum gray box, characterized in that, Includes the following steps: Acquire fault injection data for intelligent vehicles with fault classification labels; Based on the fault injection data, a Bayesian network is constructed using a structural learning method to obtain a directed acyclic graph representing the causal dependencies of variables; The obtained directed acyclic graph is mapped to a fault tree model to obtain the corresponding structure of fault events and logical relationships; Based on the fault tree model, the minimum cut subset is determined, and combined with Markov blanket filtering of variables, the minimum gray box variable set is obtained. Using a directed acyclic graph as structural input and combining fault injection data for parameter learning, a dynamic Bayesian network is constructed. The minimum gray box variable set is input into the constructed dynamic Bayesian network to detect the fault state of the intelligent vehicle and determine the fault type. Based on the fault injection data, a Bayesian network is constructed using a structure learning method to obtain a directed acyclic graph representing the causal dependencies of variables, specifically including: For each type of variable in the aforementioned fault injection data Initialize its candidate parent-child node set: in, Indicates the first i Class variables, Indicates the first i The set of candidate parent and child nodes for class variables. Represents the empty set; Inject faults into each type of variable in the data. Mapped to nodes in a Bayesian network : For except external candidate variable set Each variable in Based on the variable vector composed of samples in the fault injection data and Estimate marginal and joint probabilities and calculate and Dependence strength ,in, This represents the universal set of variables, that is, the set of all variables in the fault injection dataset. Indicates the first i Class variables, Indicates the first j Class variables, Indicates the first n The first sample i The value of a class variable, Indicates the first n The first sample j The value of the class variable, T represents transpose; when At that time, join in ,in, Preset dependency threshold; right Each candidate parent-child node in Perform a conditional independence test: determine whether a subset exists. So that in a given under conditions and Conditions are independent, that is, they satisfy: in, Indicates conditional independence. This represents a subset of variables used for conditional independence testing; When the conditions are met, from Remove and reverse prune the nodes; if the condition is not met, add undirected edges to the Bayesian network. The reverse pruning refers to when... When there are undirected edges, Delete the undirected edge between them; right all variables After processing, the initial undirected graph is obtained. ; Based on the initial undirected graph A directed acyclic graph is constructed using a greedy hill-climbing directional search method based on a scoring function.

2. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, The fault injection data includes injection data of the same vehicle under different faults. The injection data includes fault data and non-fault data. The fault injection data is represented as follows: in, This indicates non-fault data. Indicates fault data, Injecting data to the fault, including n 1 sample, each sample includes m One variable, m The variables include fault data. Non-fault data Each sample corresponds to a fault classification label, which is used to characterize the fault type or non-fault state to which the sample belongs.

3. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, The dependency strength is calculated using the following formula: in, Representing variables With variables The strength of the dependency between them n Represents the total number of samples. , They represent the first k The first sample i Class variables and the first j The value of a class variable, Indicates the first sample based on all samples. i Class variables in taking values The empirical probability at that time; Indicates the first sample based on all samples. j Class variables in taking values The empirical probability at that time, Indicates the pair of variables based on all sample pairs. In the k Sample values The joint empirical probability.

4. The intelligent vehicle fault identification method based on a minimum gray box according to claim 3, characterized in that, The empirical probability is given by the formula: in, This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. Indicates the first t The first sample i The value of the class variable; The joint empirical probability is given by the following formula: in, Indicates the pair of variables based on all sample pairs. In the k Sample values The joint empirical probability.

5. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, The initial undirected graph A directed acyclic graph is constructed using a greedy hill-climbing directional search method based on a scoring function, specifically including: The initial undirected graph The skeleton is defined as the candidate adjacency set, and the skeleton is... The set of all undirected edges in the skeleton, and during the directed search, structural transformations are only allowed on the adjacencies within the skeleton, and the introduction of other structures is prohibited. New adjacency outside the skeleton; The initial undirected graph As the current network structure ; For the current network structure Perform a greedy hill-climbing directed search: for any undirected edge in the candidate adjacency set... Perform one of the following random operations: Introducing edges: at nodes and Add a directed edge between ; Delete edge: Delete a directed edge from the current network structure. ; Reverse edge: Reverse a directed edge in the current network structure Reversed to ; Calculate the Bayesian Information Criterion (BIC) score for the graph after each network structure transformation; If the BIC score improves compared to the current network structure, the current transformation is accepted and used as the current network structure for the next iteration; if the score does not improve, the transformation is discarded and the current network structure remains unchanged. Repeat the iterations until the BIC score of the scoring function no longer improves in consecutive iterations or the preset number of iterations is reached, thus obtaining a directed acyclic graph. .

6. The intelligent vehicle fault identification method based on a minimum gray box according to claim 5, characterized in that, The scoring function is defined as follows: in, This represents the BIC score, This indicates fault injection data. Indicates the current network structure. Indicates the corresponding current network structure The set of parameters, This indicates the total number of parameters in the parameter set. Represents the total number of samples. Indicates the current network structure and parameter set Below, fault injection data The log-likelihood.

7. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, The process of mapping the obtained directed acyclic graph to a fault tree model to obtain the corresponding structure of fault events and logical relationships specifically includes: In the directed acyclic graph In this context, a node is defined as a root node if it has only a directed edge pointing to other nodes and no directed edge pointing to itself; a leaf node if it has only a directed edge pointing to itself and no directed edge pointing to other nodes; and an intermediate node if it has both directed edges pointing to itself and directed edges pointing to other nodes. Map the root node to the base event in the fault tree, the intermediate nodes to intermediate events, and the leaf nodes to the top event, where: The base event corresponding to the root node serves as the basic triggering factor for generating intermediate or top events in the fault tree. The intermediate events corresponding to intermediate nodes are triggered by the base events or intermediate events of all parent nodes that directly point to that node; The top event corresponding to a leaf node is triggered by the base event or intermediate event of all parent nodes that directly point to that node; The directed edges in the directed acyclic graph are mapped to logic gates in the fault tree, and the logical relationships are determined by calculating the trigger probability of base events on intermediate or top events. in, Represents the base event set The probability of both triggering intermediate or top events; Indicates data injection during fault diagnosis D Lower base event set The expected probability of triggering the event; This represents the total number of subnetworks sampled from a directed acyclic graph. Indicates the sampling number b The set of base events for each subnetwork; Indicates the sampling number b Sub-network structure; Indicates data injection during fault. D The first sampling b The probability of each subnetwork; when At that time, the base event set is considered Together, they trigger the occurrence of corresponding intermediate or top events, constructing the corresponding logic gates in the fault tree.

8. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, Based on the fault tree model, the minimum cut subset is determined, and combined with Markov blanket filtering of variables, the minimum gray box variable set is obtained, specifically including: The universal set of variables X Each variable in In a given set of variables and Under the given conditions, perform Markov blanket screening to determine whether the conditional independence is satisfied: in, Indicates the first i Class variables, Represents the set of candidate Markov blanket variables; To express conditional independence, it means that given... Under the conditions, Except for the whole set and Other variables are independent of each other; If conditional independence holds, then As a variable Initial screening variable set ; Based on the initial screening variable set Represent all paths in the fault tree that trigger the top event as a minimum cut subset: in, Represents the set of triggering chains for the top event; This represents the total number of minimum cut subsets. Indicates the first k The set of indices in the minimum cut subset Indicates the first i The initial selection set of variables for class variables; This represents a logical AND operation, indicating that the top event is triggered by the combined action of all variables within the minimum cut subset. Represents a logical OR operation, indicating that any set of variables from different minimum cut subsets triggers a top event; The sets of variables in all minimum cut subsets are merged to form the final minimum gray box variable set. : in, This represents the smallest set of gray box variables.

9. The intelligent vehicle fault identification method based on a minimum gray box according to claim 1, characterized in that, The process of using a directed acyclic graph as structural input and combining it with fault injection data for parameter learning to construct a dynamic Bayesian network specifically includes: Using a directed acyclic graph as the structural input to a dynamic Bayesian network, parameter estimation is performed on the fault injection data D, and the optimal parameters of the dynamic Bayesian network are obtained using maximum likelihood estimation. in, n Represents the total number of samples. Indicates the first i All variable vectors of the sample, Represents the sample mean vector. This represents the state mean parameter estimate for each variable in a dynamic Bayesian network. This represents the estimation of the covariance parameter between variables in a dynamic Bayesian network.

Citation Information

Patent Citations

  • Automobile fault analysis method based on generative Bayesian causal network

    CN112990467A