A fault root cause tracing method and apparatus
By constructing and updating the initial causal graph, and combining the Dowhy framework and fused feature vectors, the real-time and interpretability issues of fault diagnosis in gear processing are solved, and the root cause of the fault is accurately traced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NEW GENERATION IND INTELLIGENT TECHNOLOGY (TIANJIN) CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-17
AI Technical Summary
In existing gear manufacturing processes, traditional fault diagnosis methods are unable to achieve real-time and accurate process monitoring and fault early warning, and cannot explain the underlying causal mechanism of fault occurrence, resulting in unreliable and poorly interpretable diagnostic results.
An initial causal graph is constructed, nodes and causal relationships are obtained through a gear domain knowledge base, raw data is collected to generate fused feature vectors, the causal graph is updated using the Dowhy framework, and the root cause nodes of the fault are determined by tracing the causal graph.
It enables precise and interpretable root cause localization of faults during gear machining, improves the accuracy and interpretability of causal inference, and dynamically adapts to complex and changing working conditions.
Smart Images

Figure CN121413739B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gear technology, and in particular to a method and apparatus for tracing the root cause of a failure. Background Technology
[0002] As a core component of mechanical transmission systems, gears' machining accuracy and surface quality directly determine the transmission efficiency, operational stability, noise level, and service life of the equipment. With the rapid development of high-end manufacturing, the requirements for gear precision and consistency are becoming increasingly stringent. Modern gear machining widely employs automated processes such as CNC hobbing, gear shaping, and gear grinding, and is equipped with numerous sensors and data acquisition systems, achieving digitalization and informatization of the machining process. However, complex physical interactions during machining, such as dynamic changes in cutting forces, tool wear, machine tool thermal deformation, and workpiece clamping vibration, can easily lead to quality defects such as tooth profile errors, tooth direction errors, and surface burns. Traditional quality control relies on offline sampling inspections and simple threshold-based alarms, making it difficult to achieve real-time, accurate process monitoring and fault early warning. Currently, fault diagnosis methods in gear machining mainly include association rule mining, heuristic search, and machine learning.
[0003] However, association rule mining is essentially a co-occurrence analysis, which can only discover statistical associations between variables and cannot determine the causal direction or eliminate confounding factors, leading to unreliable diagnostic conclusions. Heuristic search relies entirely on the model structure pre-set by experts and cannot be dynamically updated based on actual operating data or discover unknown fault modes. When the processing technology changes, equipment ages, or new complex faults occur, the pre-set rule base often fails. The models used by machine learning are like "black boxes" and cannot reveal the internal causal mechanism of fault occurrence, resulting in low reliability and poor interpretability of diagnostic results, making it difficult for engineers to locate the root cause and optimize the process. Summary of the Invention
[0004] This invention provides a method and apparatus for tracing the root causes of faults, so as to achieve accurate and interpretable root cause localization of faults in the gear machining process.
[0005] According to a first aspect of the present invention, a method for tracing the root cause of a fault is provided, comprising: constructing an initial causal graph based on a gear domain knowledge base, wherein the initial causal graph includes observable nodes in the gear machining process and edges representing causal relationships between nodes;
[0006] Collect raw data during gear machining, and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data;
[0007] The initial causal graph is updated based on the fused feature vector to obtain an updated causal graph, wherein the updated causal graph includes the weights between each node;
[0008] When it is determined that there is a fault node in the updated cause-effect graph, the root cause node of the fault node is determined by tracing the fault node in the updated cause-effect graph.
[0009] According to another aspect of the present invention, a fault root cause tracing device is provided. The device includes: an initial causal graph construction module, used to construct an initial causal graph based on a gear domain knowledge base, wherein the initial causal graph includes observable nodes in the gear processing process and edges representing causal relationships between nodes.
[0010] The fusion feature vector determination module is used to collect raw data during the gear processing and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data.
[0011] The causal graph update module is used to update the initial causal graph according to the fused feature vector to obtain an updated causal graph, wherein the updated causal graph includes the weights between each node;
[0012] The fault root cause node determination module is used to determine the fault root cause node by tracing the fault node in the updated cause-effect graph when it is determined that there is a fault node in the updated cause-effect graph.
[0013] According to another aspect of the present invention, a terminal device is provided, the terminal device comprising: one or more processors;
[0014] Storage device for storing one or more programs.
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a storage medium for computer-executable instructions is provided, on which a computer program is stored, which, when executed by a processor, implements the method as described in any of the embodiments of the present invention.
[0017] The technical solution of this invention constructs an initial causal graph through a knowledge base in the gear field, providing a clear and interpretable causal structure framework for the entire fault tracing. The raw data collected during the gear processing is categorized according to the semantic correlation between the nodes in the causal graph, generating an informative fusion feature vector for each node in the causal graph. The weights between nodes are determined based on the fusion feature vectors, achieving accurate quantification of the causal effects between nodes in the causal graph. Based on the quantified weighted causal graph, accurate, interpretable, and dynamically adaptive root cause localization of abnormal events in gear processing is achieved.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of a fault root cause tracing method provided in Embodiment 1 of the present invention;
[0021] Figure 2 This is a flowchart of a fault root cause tracing method provided in Embodiment 2 of the present invention;
[0022] Figure 3 This is a schematic diagram of the structure of a fault root cause tracing device provided in Embodiment 3 of the present invention;
[0023] Figure 4 This invention provides a structural block diagram of a terminal device. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or terminal device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or terminal devices.
[0026] Example 1
[0027] Figure 1 This is a flowchart illustrating a fault root cause tracing method provided in an embodiment of the present invention. This embodiment is applicable to situations where faults are traced to their root causes during gear machining. The method can be executed by a fault root cause tracing device, which can be implemented in hardware and / or software, and can be integrated into a terminal device. Figure 1 As shown, the method includes:
[0028] Step S101: Construct an initial causal graph based on the gear domain knowledge base.
[0029] Optionally, an initial causal graph is constructed based on a gear domain knowledge base, including: obtaining a gear domain knowledge base, which includes entities, attributes, relations, and knowledge triples; extracting observable entities or attributes in each knowledge triple as nodes, and using causal relations related to entities or attributes in each triple as edges; and constructing an initial causal graph based on the extracted nodes and edges.
[0030] Specifically, this implementation constructs a preliminary and interpretable initial causal graph based on a gear domain knowledge base. This causal graph serves as a strong prior for subsequent data-driven causal discovery, guiding, verifying, and correcting learned causal relationships. First, the domain knowledge related to gear machining is systematically represented and extracted. This domain knowledge can originate from expert interviews, process manuals, equipment manuals, fault repair records, and gear production lines, etc., and is used to construct a gear domain knowledge base H=(EARP), where E represents the entity set, A represents the attribute set, R represents the relation set, and P is the set of knowledge triples connected by relations. The entities in the entity set refer to specific objects in the machining process, such as machine tools, cutting tools, workpieces, coolant, sensors, and fault modes. Attributes refer to the characteristics or states of entities, such as tool wear state, spindle speed, feed rate, depth of cut, surface roughness, and vibration signals. Relations refer to the interactions between entities or attributes, especially causal relationships, compositional relationships, and influence relationships. For example, "tool wear leads to increased cutting force" or "insufficient coolant flow causes workpiece burns." Of course, this embodiment is only an example and does not limit the specific content of entities, attributes, and relations. Furthermore, a knowledge triple takes the form of (Head Entity / Attribute, Relation, Tail Entity / Attribute), i.e., (H, R, T).
[0031] In this embodiment, knowledge triples are extracted from the gear domain knowledge base, focusing on constructing preliminary causal chains. The causal relationships identified in this embodiment mainly include parameter-parameter causality (e.g., increasing feed rate may lead to increased cutting force); parameter-state causality (e.g., excessive depth of cut may lead to increased tool wear); state-state causality (e.g., tool wear may lead to increased cutting vibration); state-failure causality (e.g., increased cutting vibration may lead to excessive surface roughness); and failure-failure causality (e.g., excessive surface roughness may further induce fatigue failure). The identified causal chains are represented as a series of directed edges, and a preliminary causal graph G driven by domain knowledge is constructed based on the identified causal chains. K = (V K E K ), where V K This represents a set of nodes, signifying key variables in the gear machining process. These key variables can be attributes or important failure modes, such as V. K= {v1,v2,...v n}, where v i This could include "tool wear condition", "cutting force", "surface roughness", "coolant temperature", etc. E K Let v represent a set of directed edges, representing causal relationships determined by domain knowledge, if variable v i It is the variable v j The direct cause, then, lies in a path from v i Point to v j Directed edge .
[0032] In a specific implementation, after obtaining the gear domain knowledge base H=(EARP), each knowledge triple (H, R, T) in H is extracted. A knowledge triple can be read as "there is a relationship R (relation) between H (head) and T (tail)". For example, "tool wear will lead to increased cutting force", where H = "tool wear" (an attribute / state), R = "leads to" (a causal relationship), and T = "increased cutting force" (an attribute / state). The corresponding causal edge is: tool wear. The cutting force increases. Of course, this embodiment is merely illustrative and does not limit the specific content of H, R, and T. Furthermore, when it is determined that H and T in the extracted knowledge triple represent observable entities or properties during gear machining, and in V... K If H and T are not currently present in V, then H and T will be added to V. KIn the example, if R is a clear causal relationship, such as "cause" or "lead to", then add a directed edge. To E K If R is an indirect causal relationship or association, it needs to be further transformed into a direct causal relationship through expert confirmation or heuristic rules, and then added to... Therefore, the initial causal graph constructed includes observable nodes during gear machining, as well as edges representing causal relationships between nodes. In this embodiment, observable nodes include entities or properties observable on the gear machining tool, such as shaft vibration, shaft speed, motor position, and workpiece position. Subsequent analysis and processing can only be performed on these observable entities or properties. Of course, this embodiment is merely illustrative and does not limit the specific construction process of the initial causal graph.
[0033] Step S102: Collect raw data during gear processing and determine the fusion feature vectors associated with each node in the initial causal graph based on the raw data.
[0034] Optionally, the fusion feature vectors associated with each node in the initial causal graph are determined based on the original data, including: preprocessing the original data to obtain preprocessed original data, wherein the preprocessing operations include time alignment, missing value imputation, outlier removal, and normalization; filtering the data subsets associated with each node from the preprocessed original data, and identifying the different modes of data contained in the data subsets, wherein the modes include vibration data, structured time series data, and discrete parameters; extracting feature vectors from the data of different modes respectively, and fusing the extracted feature vectors to obtain the fusion feature vectors associated with each node.
[0035] Specifically, this embodiment collects raw data involved in gear machining. This raw data includes equipment status data, machining process parameters, sensor data, quality data, and machine tool fault data. Equipment status data includes spindle speed, feed rate, position feedback of each motion axis, motor current, torque signal, and tool life counter values, reflecting the macroscopic operating state of the machine tool. Machining process parameters include set depth of cut, feed rate, tool compensation value, coolant pressure, flow rate, and temperature; these are key input variables directly affecting cutting force and thermal effects. Sensor data includes vibration sensors measuring the vibration acceleration of key machine tool components to reflect the dynamic stability of the machining process. Quality data and machine tool fault data include fault types and quality defect types. This embodiment is merely illustrative and does not limit the specific types of raw data collected. Because the collected raw data comes from diverse sources and exhibits significant heterogeneity and non-ideality, including different modes, sampling rates, and data formats, as well as potential missing values, noise, and drift, data preprocessing is necessary to ensure the effectiveness of subsequent analysis. First, time alignment is achieved by precisely matching the timestamps of each data source based on a common time reference, such as the spindle rotation period, to ensure comparability between variables. Second, multiple interpolation is used to fill in missing values, especially for high-dimensional time series signals to preserve their complex structural features. Next, outliers are removed to avoid interfering with the training process. Finally, all variables are normalized or standardized to the same numerical range to eliminate dimensional differences and accelerate algorithm convergence. Of course, this implementation is only an example and does not limit the specific methods of preprocessing.
[0036] It is worth mentioning that in this embodiment, a subset of data related to each node is selected from the preprocessed raw data. This subset includes data of different patterns, such as vibration data, structured time-series data, and discrete parameters. Of course, this embodiment is merely illustrative and does not limit the data models contained in the subsets. For each node v in the causal graph... i Its corresponding original data subset Dv iIt may be multimodal and high-dimensional, requiring targeted feature extraction and fusion. In this embodiment, the feature extraction methods used for different modes of data are also different. For example, for vibration data, continuous wavelet transform (CWT) is performed, using time-domain-frequency domain transformation and feature extraction. Specifically, continuous wavelet transform (CWT) is applied to the original vibration data. CWT can convert the vibration signal from a single time domain to a time-frequency two-dimensional plane, generating a "wavelet coefficient map" containing rich transient features. The wavelet coefficient map intuitively reveals the changing trend of different frequency components over time, effectively capturing key fault features such as periodic impacts and frequency shifts caused by tool wear, broken teeth, or workpiece defects. Specifically, the following formula (1) can be used to perform wavelet transform on the vibration data to obtain the wavelet coefficient map:
[0037]
[0038] in, Representing vibration data, This represents a wavelet coefficient diagram, where 'a' is the scaling parameter, which is frequency-dependent, and 'b' is the translation parameter, which is time-dependent. It is the mother wavelet function. yes . This forms a two-dimensional matrix representing the energy distribution of the vibration signal at different times and frequencies, based on the wavelet coefficient diagram. Feature extraction is performed using a 2D-CNN. The CNN network consists of convolutional layers, pooling layers, and fully connected layers. Convolutional operations extract local features from the input data. After the convolutional layers, the pooling layers downsample the feature maps to reduce their size while preserving key information. The fully connected layer at the end of the network fuses the high-level features extracted from the preceding layers, ultimately generating the output result: a high-level, discriminative vibration feature vector. Furthermore, for structured time-series data such as spindle speed, current, and temperature, time-series feature engineering is performed. This includes calculating statistics, trend terms, and correlation indices with other variables within the sliding window, thereby extracting low-dimensional feature vectors that characterize the steady-state performance and potential degradation trends of the equipment. For discrete parameters, equipment states are directly used as features or embedded encoding is employed. Therefore, the feature vector extraction methods for structured time-series data and discrete parameters are relatively simpler than those for vibration data. In this embodiment, all feature vectors extracted from the data subset are concatenated to form a representative node v. i The fused feature vector at a specific time step Of course, this embodiment is only an example and does not limit the specific method of obtaining the fusion feature vector of each node.
[0039] Step S103: Update the initial causal graph based on the fused feature vector to obtain the updated causal graph.
[0040] Optionally, the initial causal graph is updated based on the fused feature vector to obtain an updated causal graph, including: obtaining the fused feature vector of directly connected nodes, and calculating the causal effect value between directly connected nodes based on the fused feature vector; using the causal effect value as the weight of the edge between directly connected nodes, and adding the weight to the initial causal graph to obtain the updated causal graph.
[0041] Specifically, in this embodiment, the initial causal graph obtained above is used as the prior causal structure. Combined with the fused feature vectors of each node, the Dowhy framework is used to identify, estimate, and verify causal effects. Dowhy is a Python library that provides a systematic set of causal inference steps: identification, estimation, and verification. Specifically, Dowhy involves causal graph input, defining the causal problem, and identifying causal effects when constructing and identifying the causal model. Causal graph input refers to inputting the initial causal graph G... K As input to the Dowhy framework, a latent causal structure is defined, which provides the graph information needed for the causal identification phase. Defining a causal problem refers to defining a problem for G... K Each causal edge in We define it as a causal problem, for example, processing variable T as node v i The resulting variable Y is node v j The confounding variable W is a variable that affects both the treatment variable T and the outcome variable Y. If left uncontrolled, it can create a false causal relationship, according to G... K Based on its structure, Dowhy automatically identifies the set of confounding variables that need to be controlled. This is Dowhy's key advantage, avoiding the bias of manually selecting confounding variables. When identifying causal effects, the Dowhy framework uses the input G... K For causal problems with structure and definition, the system automatically applies causal calculus rules to identify unbiased estimation expressions for causal effects.
[0042] In a specific implementation, since Dowhy supports multiple causal effect estimators, which can be selected based on data characteristics and problem complexity, this implementation chooses the Causal Forests estimator. For each edge v in the initial causal graph... i —>v jThe fused feature vectors of directly connected nodes are obtained, and the node v is calculated based on the two fused feature vectors using the Causal Forests estimator selected above. i and v j The causal effect value between the two sides is used, and in this embodiment, the causal effect value can be the average causal effect (ACE) or the conditional average causal effect (CACE). This embodiment does not limit the specific type of causal effect value, and the obtained causal effect value can be directly used as the edge. weight w ij The weights of each edge are then added to the initial causal graph, thus constructing a weighted updated causal graph G. W Of course, this embodiment is only an example and does not limit the way the weights of each edge in the updated causal graph are obtained.
[0043] Step S104: When it is determined that there is a fault node in the updated cause-effect graph, the root cause node of the fault node is determined by tracing the root cause of the fault in the updated cause-effect graph.
[0044] Optionally, when it is determined that there is a fault node in the updated causal graph, the root cause of the fault node is determined by tracing the fault node in the updated causal graph, including: identifying the node whose fused feature vector in the updated causal graph exceeds a specified range as the fault node; determining the candidate causal path related to the fault node from the updated causal graph; determining the target causal path from the candidate causal path; and determining the root cause of the fault based on the target causal path.
[0045] Optionally, the target causal path is determined from the candidate causal paths, including: calculating the weight product of each edge on the candidate causal path and using the weight product as the causal strength of the candidate causal path; and selecting the candidate causal path with the highest causal strength as the target causal path.
[0046] Optionally, the root cause node of the fault can be determined based on the target causal path, including calculating the causal contribution of each node on the target causal path to the fault node; and the node with the largest causal contribution is taken as the root cause node of the fault.
[0047] Specifically, in this embodiment, when a fault is detected in the gear machining process in real time, such as a quality indicator exceeding tolerance or a fusion characteristic of a certain node, When the behavior deviates from the normal range, root cause analysis is triggered, and the faulty node v is identified from the updated cause-effect graph. fault And starting from the faulty node, along G WBacktracking the reverse causal paths with higher weights identifies candidate causal paths related to the faulty node. For each candidate causal path... The system calculates the causal strength, specifically by multiplying the weights of each side as the causal strength of the candidate causal path. Of course, this embodiment is just an example and does not limit the specific calculation method of the causal strength. Specifically, the candidate causal path with the highest causal strength is taken as the target causal path. Since the target causal path contains all the associated nodes related to the fault node, the fault diagnosis results are interpretable.
[0048] Optionally, the causal contribution of each node on the target causal path to the faulty node is calculated, including: obtaining the connection path between the node and the faulty node for each node on the target causal path; obtaining the weight product result of each edge on the connection path and the path length; and calculating the causal contribution of the node to the faulty node based on the weight product result and the path length.
[0049] In this embodiment, after determining the target causal path, the causal contribution of each node to the faulty node can be calculated from the target causal path. The causal contribution reflects the intensity of the direct or indirect influence of the node on the occurrence of the fault, and the node with the largest contribution is taken as the root cause node of the fault related to the faulty node, that is, the biggest factor affecting the occurrence of the fault.
[0050] For example, when the target causal path is determined to be At that time, v will be calculated separately. roat and v a For v fault The causal contribution of v roat In this context, the connection path between it and the faulty node is the target causal path itself, and v roat With v a The weight is A. The weight of the path is B, the product of the weights of each edge on the connecting path is A×B, and the length of the connecting path is 2. Since the longer the path, the lower the influence, the causal contribution is inversely proportional to the path length. Therefore, we can obtain v. roat For v fault The causal contribution is (A×B) / 2, and similarly, v can be obtained. a For v fault The causal contribution is B, and the node with the largest causal contribution is taken as the root cause node of the fault. Of course, this embodiment is only an example and does not limit the specific calculation method of causal sharing.
[0051] It is worth mentioning that this implementation uses a coarse-grained causal graph constructed from a gear domain knowledge base as a strong prior, providing a clear and interpretable causal structural framework for the entire fault tracing framework. This "structure-first, data-enhanced" fusion paradigm effectively avoids the pseudo-causal problem that is easily generated by pure data-driven causal discovery, while also compensating for the limitations of pure expert systems in adapting to complex and variable working conditions, significantly improving the accuracy and interpretability of causal inference. In addition, by classifying heterogeneous data such as raw vibration data, current, temperature, and process parameters according to their semantic relevance to the nodes of the causal graph, and using a deep learning model to extract features and perform high-dimensional fusion of the fine-grained data corresponding to each node, a rich and informative fused feature vector that can be used as a numerical input is finally generated for each abstract node in the causal graph. This solves the technical problem of directly inputting complex raw data into the causal inference model, providing high-quality input for subsequent causal quantitative analysis. Furthermore, the Dowhy causal inference framework is applied to achieve accurate quantification of the causal effects between variables in the causal graph structure. By using an expert causal graph as input to Dowhy and combining it with extracted node fusion feature vectors, the Dowhy framework can automatically perform the identification, estimation, and verification of causal effects. Finally, based on the quantized and weighted updated causal graph, it can achieve accurate, multi-path tracing from the fault phenomenon to the deepest root cause. By reasoning along the reverse causal path in the causal graph and combining the weight accumulation of each causal edge, the system can quantify the total contribution of each potential root cause to the occurrence of the fault, thereby identifying the most critical root cause.
[0052] The technical solution of this invention constructs an initial causal graph through a knowledge base in the gear field, providing a clear and interpretable causal structure framework for the entire fault tracing. The raw data collected during the gear processing is categorized according to the semantic correlation between the nodes in the causal graph, generating an informative fusion feature vector for each node in the causal graph. The weights between nodes are determined based on the fusion feature vectors, achieving accurate quantification of the causal effects between nodes in the causal graph. Based on the quantified weighted causal graph, accurate, interpretable, and dynamically adaptive root cause localization of abnormal events in gear processing is achieved.
[0053] Example 2
[0054] Figure 2 This is a flowchart of a fault root cause tracing method provided by an embodiment of the present invention. Based on the above embodiments, after constructing an initial causal graph based on the extracted nodes and edges, this embodiment further includes: identifying the initial causal graph to obtain specified edges that need adjustment, wherein the specified edges include duplicate edges, conflicting edges, and redundant edges; determining the adjustment method matching the specified edges; and adjusting the matching specified edges according to the adjustment method to obtain an adjusted initial causal graph that closely resembles real-world causal relationships and has a simplified structure. Figure 2 As shown, the method includes:
[0055] Step S201: Construct an initial causal graph based on the gear domain knowledge base.
[0056] Step S202: Identify the specified edges that need to be adjusted in the initial cause-effect graph, determine the adjustment method that matches the specified edges, and adjust the matched specified edges according to the adjustment method to obtain the adjusted initial cause-effect graph.
[0057] Specifically, in this embodiment, after obtaining the initial causal graph G... K Next, the initial cause-effect graph will undergo structural optimization. During optimization, the initial cause-effect graph will be identified to obtain specified edges that require adjustment. These specified edges include duplicate edges, conflicting edges, and redundant edges. Adjustments will then be made using methods matching these specified edges to obtain the adjusted initial cause-effect graph. Duplicate edges can refer to edges containing two identical... An edge with conflict refers to an edge that exists. and The contradictory causal relationship, redundant edges refer to the existence of... and At the same time, there is also Of course, this embodiment is merely an example and does not limit the specific types of duplicate edges, conflicting edges, and redundant edges. In this embodiment, different adjustment methods are pre-set for different types of specified edges, and the specified edges are adjusted according to the corresponding adjustment methods. For duplicate edges, the adjustment method is to delete the duplicate edges. For conflicting edges, the adjustment method is to first manually judge the conflict and then correct it based on the manual judgment. For redundant edges, the adjustment method is to delete redundant edges that directly connect multiple nodes and retain edges that connect multiple nodes in segments. In other words, the principle behind the adjustment methods for specified edges is to closely approximate the real-world causal relationships, making the initial causal graph structure more concise. For example, for duplicate edges, they are removed from E... K Remove from the process; for conflicting edges, such as those that exist... and When there are contradictory causal relationships, human experts are needed to make judgments and corrections; for redundant edges, such as... and At the same time, there is also , When experts believe Through When transferring, it can be directly retained. and and remove Redundant edges, of course Except where there is an independent direct causal relationship.
[0058] It should be noted that, in this embodiment, by adjusting the repeated edges, conflicting edges, and redundant edges of the initial causal graph according to the matching adjustment method, the initial causal graph can be greatly simplified and its interpretability improved, which facilitates the accuracy of subsequent root cause tracing. Of course, this embodiment is only an example and does not limit the structural optimization method of the initial causal graph. As long as the accuracy of the initial causal graph can be improved, it is within the protection scope of this application.
[0059] Step S203: Collect raw data during the gear processing and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data.
[0060] Optionally, the fusion feature vectors associated with each node in the initial causal graph are determined based on the original data, including: preprocessing the original data to obtain preprocessed original data, wherein the preprocessing operations include time alignment, missing value imputation, outlier removal, and normalization; filtering the data subsets associated with each node from the preprocessed original data, and identifying the different modes of data contained in the data subsets, wherein the modes include vibration data, structured time series data, and discrete parameters; extracting feature vectors from the data of different modes respectively, and fusing the extracted feature vectors to obtain the fusion feature vectors associated with each node.
[0061] Step S204: Update the initial causal graph based on the fused feature vector to obtain the updated causal graph.
[0062] Optionally, the initial causal graph is updated based on the fused feature vector to obtain an updated causal graph, including: obtaining the fused feature vector of directly connected nodes, and calculating the causal effect value between directly connected nodes based on the fused feature vector; using the causal effect value as the weight of the edge between directly connected nodes, and adding the weight to the initial causal graph to obtain the updated causal graph.
[0063] Step S205: When it is determined that there is a fault node in the updated cause-effect graph, the root cause node of the fault node is determined by tracing the root cause of the fault in the updated cause-effect graph.
[0064] Optionally, when it is determined that there is a fault node in the updated causal graph, the root cause of the fault node is determined by tracing the fault node in the updated causal graph, including: identifying the node whose fused feature vector in the updated causal graph exceeds a specified range as the fault node; determining the candidate causal path related to the fault node from the updated causal graph; determining the target causal path from the candidate causal path; and determining the root cause of the fault based on the target causal path.
[0065] Optionally, the target causal path is determined from the candidate causal paths, including: calculating the weight product of each edge on the candidate causal path and using the weight product as the causal strength of the candidate causal path; and selecting the candidate causal path with the highest causal strength as the target causal path.
[0066] Optionally, the root cause node of the fault can be determined based on the target causal path, including calculating the causal contribution of each node on the target causal path to the fault node; and the node with the largest causal contribution is taken as the root cause node of the fault.
[0067] The technical solution of this invention constructs an initial causal graph through a knowledge base in the gear field, providing a clear and interpretable causal structure framework for the entire fault tracing. The raw data collected during the gear processing is categorized according to the semantic correlation between the nodes in the causal graph, generating an informative fusion feature vector for each node in the causal graph. The weights between nodes are determined based on the fusion feature vectors, achieving accurate quantification of the causal effects between nodes in the causal graph. Based on the quantified weighted causal graph, accurate, interpretable, and dynamically adaptive root cause localization of abnormal events in gear processing is achieved.
[0068] Example 3
[0069] Figure 3 This is a schematic diagram of a fault root cause tracing device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the device includes:
[0070] The initial causal graph construction module 310 is used to construct an initial causal graph based on the gear domain knowledge base. The initial causal graph includes observable nodes in the gear processing process and edges representing causal relationships between nodes.
[0071] The fusion feature vector determination module 320 is used to collect raw data during the gear processing and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data.
[0072] The causal graph update module 330 is used to update the initial causal graph based on the fused feature vector to obtain an updated causal graph, wherein the updated causal graph includes the weights between each node;
[0073] The fault root cause node determination module 340 is used to determine the fault root cause node by tracing the fault node in the updated cause-effect graph when it is determined that there is a fault node in the updated cause-effect graph.
[0074] Optionally, the initial cause-effect graph construction module 310 is used to obtain a gear domain knowledge base, wherein the gear domain knowledge base includes entities, attributes, relations and knowledge triples;
[0075] Extract the observable entities or attributes in the gear processing process contained in each knowledge triple as nodes, and use the causal relationships related to the entities or attributes in each triple as edges;
[0076] An initial causal graph is constructed based on the extracted nodes and edges.
[0077] Optionally, the device further includes an initial causal graph adjustment module for identifying the initial causal graph and obtaining specified edges that need to be adjusted, wherein the specified edges include duplicate edges, conflicting edges, and redundant edges.
[0078] Determine the adjustment method that matches the specified edge, and adjust the matched specified edge according to the adjustment method to obtain an adjusted initial causal graph that closely resembles the real-world causal relationship and has a simplified structure.
[0079] Optionally, a feature vector fusion determination module is used to perform preprocessing operations on the original data to obtain preprocessed original data. The preprocessing operations include temporal alignment, missing value imputation, outlier removal, and normalization.
[0080] From the preprocessed raw data, a subset of data related to each node is selected, and different patterns of data contained in the subset are identified. The patterns include vibration data, structured time series data, and discrete parameters.
[0081] Feature vectors are extracted from data of different modes, and the extracted feature vectors are fused to obtain the fused feature vectors related to each node.
[0082] Optionally, a causal graph update module is used to obtain the fused feature vectors of directly connected nodes and calculate the causal effect values between directly connected nodes based on the fused feature vectors.
[0083] The causal effect value is used as the weight of the edge between directly connected nodes, and the weight is added to the initial causal graph to obtain the updated causal graph.
[0084] Optionally, a fault root cause node determination module is used to identify nodes whose fused feature vectors in the updated causal graph exceed a specified range as fault nodes.
[0085] From the updated cause-effect graph, candidate causal paths related to the faulty node are identified, and the target causal path is determined from the candidate causal paths.
[0086] The root cause node of the failure is determined based on the target causal path.
[0087] Optionally, the root cause node determination module includes a target causal path determination unit, which is used to calculate the weight product of each edge on the candidate causal path and use the weight product as the causal strength of the candidate causal path.
[0088] The candidate causal path with the strongest causal strength is selected as the target causal path.
[0089] Optionally, the root cause node determination module includes a root cause node determination unit, which is used to calculate the causal contribution of each node on the target causal path to the fault node.
[0090] The node with the largest causal contribution is taken as the root cause node of the failure.
[0091] The fault root cause tracing device provided in the embodiments of the present invention can execute a fault root cause tracing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0092] Example 4
[0093] Figure 4 A schematic diagram of a terminal device 10 that can be used to implement embodiments of the present invention is shown. The terminal device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The terminal device can also represent various forms of mobile devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0094] The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0095] like Figure 4 As shown, the terminal device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the terminal device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0096] Multiple components in terminal device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows terminal device 10 to exchange information / data with other terminal devices through computer networks such as the Internet and / or various telecommunications networks.
[0097] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as root cause analysis methods.
[0098] In some embodiments, the root cause analysis method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on terminal device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the root cause analysis method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the root cause analysis method by any other suitable means (e.g., by means of firmware).
[0099] Various embodiments of the apparatuses and techniques described above herein can be implemented in digital electronic circuit devices, integrated circuit devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), device-on-a-chip (SoCs), complex programmable logic terminal devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable device including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage device, at least one input device, and at least one output device, and transmitting data and instructions to the storage device, the at least one input device, and the at least one output device.
[0100] Computer programs used to implement the fault root cause tracing method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other business-uninterrupted data migration device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, or as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0101] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution apparatus, device, or terminal device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage terminal devices, magnetic storage terminal devices, or any suitable combination thereof.
[0102] To provide interaction with a user, the apparatus and techniques described herein can be implemented on a terminal device having: a display device (e.g., a touchscreen) for displaying information to the user; and buttons through which the user can provide input to the terminal device. Other types of apparatus can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including voice input, speech input, or haptic input).
[0103] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.
[0104] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for tracing the root cause of a fault, characterized in that, The method includes: An initial causal graph is constructed based on a knowledge base in the gear domain. The initial causal graph includes observable nodes in the gear manufacturing process and edges representing causal relationships between nodes. Collect raw data during gear machining, and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data; The initial causal graph is updated based on the fused feature vector to obtain an updated causal graph, wherein the updated causal graph includes the weights between each node; When it is determined that there is a fault node in the updated cause-effect graph, the root cause node of the fault node is determined by tracing the fault node in the updated cause-effect graph. The step of constructing an initial causal graph based on a gear domain knowledge base includes: acquiring a gear domain knowledge base, wherein the gear domain knowledge base includes entities, attributes, relationships, and knowledge triples, and the gear domain knowledge base is derived from expert interviews, process manuals, equipment manuals, and fault repair records; Each knowledge triple contains observable entities or attributes during gear processing, which are extracted as nodes. The causal relationships related to the entities or attributes in each knowledge triple are used as edges. The causal relationships include parameter-parameter causality, parameter-state causality, state-fault causality, and fault-fault causality. The initial causal graph is constructed based on the extracted nodes and edges; The step of updating the initial causal graph based on the fused feature vector to obtain the updated causal graph includes: obtaining the fused feature vector of directly connected nodes, and calculating the causal effect value between directly connected nodes based on the fused feature vector; The causal effect value is used as the weight of the edge between the directly connected nodes, and the weight is added to the initial causal graph to obtain the updated causal graph.
2. The method according to claim 1, characterized in that, After constructing the initial causal graph based on the extracted nodes and edges, the process further includes: The initial causal graph is identified to obtain the specified edges that need to be adjusted, wherein the specified edges include repeated edges, conflicting edges, and redundant edges; Determine the adjustment method that matches the specified edge, and adjust the matched specified edge according to the adjustment method to obtain an adjusted initial causal graph that closely resembles real-world causal relationships and has a simplified structure.
3. The method according to claim 1, characterized in that, The step of determining the fusion feature vectors associated with each node in the initial causal graph based on the original data includes: The original data is preprocessed to obtain the preprocessed original data. The preprocessing operations include time alignment, missing value imputation, outlier removal, and normalization. From the preprocessed raw data, a subset of data related to each node is selected, and different patterns of data contained in the subset are identified, wherein the patterns include vibration data, structured time series data, and discrete parameters; Feature vectors are extracted from data of different modes, and the extracted feature vectors are fused to obtain the fused feature vectors related to each node.
4. The method according to claim 1, characterized in that, When it is determined that a faulty node exists in the updated cause-effect graph, the root cause node of the faulty node is determined by tracing its origin in the updated cause-effect graph, including: The nodes whose fused feature vectors in the updated causal graph exceed the specified range are designated as faulty nodes. From the updated causal graph, candidate causal paths related to the faulty node are determined, and from the candidate causal paths, the target causal path is determined. The root cause node of the failure is determined based on the target causal path.
5. The method according to claim 4, characterized in that, Determining the target causal path from the candidate causal paths includes: Calculate the weight product of each edge on the candidate causal path, and use the weight product as the causal strength of the candidate causal path; The candidate causal path with the highest causal strength is selected as the target causal path.
6. The method according to claim 4, characterized in that, The step of determining the root cause node of the fault based on the target causal path includes: Calculate the causal contribution of each node on the target causal path to the faulty node; The node with the largest causal contribution is designated as the root cause node of the fault.
7. The method according to claim 6, characterized in that, The calculation of the causal contribution of each node on the target causal path to the faulty node includes: For each node on the target causal path, obtain the connection path between the node and the faulty node; Obtain the weight product of each edge on the connection path and the path length; The causal contribution of the node to the faulty node is calculated based on the weighted product result and the path length.
8. A fault root cause tracing device, characterized in that, The device includes: An initial causal graph construction module is used to construct an initial causal graph based on a gear domain knowledge base. The initial causal graph includes observable nodes in the gear processing process and edges representing causal relationships between nodes. The fusion feature vector determination module is used to collect raw data during the gear processing and determine the fusion feature vectors related to each node in the initial causal graph based on the raw data. The causal graph update module is used to update the initial causal graph according to the fused feature vector to obtain an updated causal graph, wherein the updated causal graph includes the weights between each node; The fault root cause node determination module is used to determine the fault root cause node by tracing the fault node in the updated cause-effect graph when it is determined that there is a fault node in the updated cause-effect graph. The initial cause-effect graph construction module is used to obtain a gear domain knowledge base, wherein the gear domain knowledge base includes entities, attributes, relationships and knowledge triples, and the gear domain knowledge base is derived from expert interviews, process manuals, equipment manuals and fault repair records; Each knowledge triple contains observable entities or attributes during gear processing, which are extracted as nodes. The causal relationships related to the entities or attributes in each knowledge triple are used as edges. The causal relationships include parameter-parameter causality, parameter-state causality, state-fault causality, and fault-fault causality. The initial causal graph is constructed based on the extracted nodes and edges; The causal graph update module is used to obtain the fusion feature vector of directly connected nodes and calculate the causal effect value between directly connected nodes based on the fusion feature vector. The causal effect value is used as the weight of the edge between the directly connected nodes, and the weight is added to the initial causal graph to obtain the updated causal graph.
Citation Information
Patent Citations
Wafer processing quality backtracking evaluation method driven by cross-modal knowledge graph
CN120873974A
Circuit board production yield root cause tracing method
CN121212766A