Industrial fault diagnosis method and system based on intelligent causal correction

Through intelligent causal correction methods that integrate large language models, graph neural networks and expert knowledge bases, the problems of causal relationship fuzzy and knowledge gaps in traditional fault diagnosis are solved, and more efficient and accurate fault diagnosis and system optimization are achieved.

CN120217262AActive Publication Date: 2025-06-27YANTAI UNIV

Patent Information

Application Number
CN202510677149.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

When dealing with complex power equipment failures, traditional industrial fault diagnosis methods face problems such as fuzzy causal relationships, knowledge gaps, and complex multivariate interactions, resulting in insufficient diagnostic accuracy and response speed.

Method used

Using intelligent causal correction method, the causal relationship diagram is automatically constructed, dynamically optimized and real-time correction is improved by integrating large language model (LLM), graph neural network (GNN) and expert knowledge base, thereby improving the robustness, interpretability and adaptability of the diagnosis.

Benefits of technology

It significantly improves the accuracy and efficiency of fault diagnosis, enhances the interpretability of the causal graph and the anti-interference ability of the system, and realizes real-time knowledge update and closed-loop optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217262A_ABST
    Figure CN120217262A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fault detection, in particular to an industrial fault diagnosis method and system based on intelligent causal correction. The method comprises the following steps: acquiring monitoring data and user task requirements; constructing task structure information based on the obtained user task demand; processing the monitoring data through a three-layer cascaded framework to generate metadata; constructing an initial causal graph structure by utilizing the task structure information and the metadata; performing deep optimization on the initial causal graph structure through a graph neural network to obtain a causal composite relation graph; fault diagnosis and information retrieval are carried out by utilizing the causal composite relation graph and combining a knowledge base; and generating a structured diagnostic report. Through an innovative structure integrating the LLM and the graph neural network, the system can dynamically optimize a causal relationship model, automatically learn and correct fault association, enhance the interpretability of a causal graph by using semantic reasoning of the LLM and a graph attention mechanism, and significantly improve the robustness and accuracy of diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault detection, and in particular, to an industrial fault diagnosis method and system with intelligent causal correction. Background Art

[0002] With the rapid development of Industry 4.0 and artificial intelligence technologies, industrial fault diagnosis systems play a crucial role in enhancing equipment reliability and production efficiency, especially in the field of power equipment, such as the maintenance of transformers, generators, and transmission lines. However, fault diagnosis of power equipment faces significant challenges, mainly due to complex working environments, multiple fault modes, and data uncertainty. These problems include the ambiguity of fault causality, knowledge gaps (such as unknown variable associations and potential fault chains), and complex interactions among multiple variables, resulting in reduced diagnostic accuracy and response delays. For example, in a power system, traditional methods may not be able to effectively capture the dynamic association of "voltage fluctuations causing temperature rise and then leading to insulation failure", especially when data variations or new fault modes occur.

[0003] Traditional fault diagnosis methods usually rely on predefined thresholds or static models, unable to dynamically adapt to changing environments, ignoring potential fault chains and cross - impacts, and thus performing poorly when dealing with unknown faults or data variations. This limitation not only reduces the interpretability and reliability of diagnosis but also increases maintenance costs and risks. Specifically, it is manifested in the following aspects: Power equipment monitoring data generally has noise interference, data missing, and abnormal fluctuations caused by harsh environments, seriously affecting data reliability; at the same time, key diagnostic information is scattered in multi - source heterogeneous data, and traditional methods lack effective mechanisms for in - depth fusion and high - quality processing, resulting in the failure to fully exploit the data value and limiting diagnostic accuracy and robustness; When existing systems are based on statistical methods for causal relationship modeling, they lack robustness, are easily limited by data noise, outliers, and task information limitations, resulting in unstable or unreliable causal graph construction, and thus reducing the overall performance and anti - interference ability of fault diagnosis; Traditional methods based on preset rules or single statistical models are difficult to dynamically capture the complex, non - linear, time - delayed fault causal chains generated by power equipment in a high - dimensional, time - varying operating environment. When existing systems are based on neural network methods for causal relationship modeling, the interpretability is not high, and they cannot fully utilize the semantic reasoning of LLM and the deep learning ability of graph neural networks, resulting in incomplete and lack of transparency of the causal graph, affecting the reliability of diagnostic results and user credibility; The power equipment fault diagnosis system lacks a real-time knowledge update and feedback optimization mechanism, and the knowledge base cannot be dynamically expanded and corrected according to new data or user feedback. As a result, the system has poor adaptability when facing new fault modes or environmental changes, and the diagnostic results may be lagged or inaccurate, unable to achieve the closed-loop optimization of continuous learning. Traditional diagnostic results are often simple fault labels or alarms, lacking clear and logical explanations for the causes, propagation paths, and potential impacts of faults. At the same time, the confidence level of the diagnostic conclusion and the uncertainty of the prediction result are not effectively quantified, reducing the user's trust in the diagnostic suggestions and decision-making efficiency. Summary of the Invention

[0004] To solve the above-mentioned problems, the present invention provides an industrial fault diagnosis method and system with intelligent causal correction. By integrating a large language model (LLM), a graph neural network (GNN), and an expert knowledge base, the automatic construction, dynamic optimization, and real-time correction of causal relationships are realized, thus significantly improving the robustness, interpretability, and adaptability of diagnosis.

[0005] In the first aspect, an industrial fault diagnosis method with intelligent causal correction provided by the present invention adopts the following technical solutions: An industrial fault diagnosis method with intelligent causal correction includes: Obtain monitoring data and user task requirements; Construct task structure information based on the obtained user task requirements; Process the monitoring data through a three-level cascaded framework to generate metadata; Construct an initial causal graph structure using the task structure information and metadata; Deeply optimize the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph; Use the causal composite relationship graph and combine it with the knowledge base for fault diagnosis and information retrieval; Generate a structured diagnostic report.

[0006] Further, the constructing of the task structure information based on the obtained user task requirements includes, for the current task T1, calculating the similarity with one or more historical tasks T2 in the historical task knowledge base, and through a hybrid similarity mechanism to comprehensively consider the similarity of the semantic content of the task and the set of structured variables, expressed as: Where emb ( T ) represents the key entity embedding vector of task T , var ( T ) represents the standardized set of monitoring variables in task T , α1 denotes the weight parameter, Cosine denotes the cosine similarity between vectors, reflecting the proximity of tasks in the semantic space, Jaccard denotes the proportion of shared variables between sets, reflecting the degree of structural overlap of tasks in key variables.

[0007] Furthermore, the generation of metadata by processing the monitoring data through a three - level cascaded framework includes first performing feature normalization on the monitoring data using the Robust Scaling method, scaling based on the quartiles and median of the data, and for the missing regions in the time - series feature matrix, interpolating using the weighted average of several latest valid observations before the missing points; then estimating the noise level in the monitoring data through the median absolute deviation, decomposing the detected high - frequency components using wavelet transform for multi - scale decomposition, and decomposing the data sequence into low - frequency approximation coefficients and high - frequency detail coefficients; for low - frequency noise, non - Gaussian distributed noise or structural noise, suppressing it using a GAN - based denoising model.

[0008] Furthermore, the construction of the initial causal graph structure using the task structure information and metadata includes first constructing a causal graph using the Granger causality test algorithm G Granger , generating a directed graph by analyzing the time - series prediction relationship between variables, expressed as: where, RSS resticted is the sum of squared residuals fitted using only the historical data of B , RSS unresticted is the sum of squared residuals fitted using the historical data of i and j , k denotes the lag order, dynamically selected based on autocorrelation function analysis, p denotes the autoregressive order, and the algorithm assumes that if the past values of variable j can significantly improve the prediction of variable i , that is, F is higher than the set threshold, then there exists a causal relationship of j → i .

[0009] Furthermore, the construction of the initial causal graph structure using the task structure information and metadata also includes using the PC algorithm based on conditional independence testing, assuming that there may be potential causal relationships between all variables, starting from a fully - connected graph and gradually increasing the conditional set k , testing variables i and jConditional independence; by iteratively increasing k, the algorithm gradually removes redundant edges, constructs a sparse undirected skeleton graph, and infers the edge directions in combination with the Meek rule to construct a directed graph GPC. Among them, for the initial condition set k, the correlation coefficient of each variable pair (i, j) is calculated r ij∣k , and its conditional independence score is tested Z : Among them, r ij∣k is the partial correlation coefficient of variables i and j under the given condition set k , n is the sample size. The larger the absolute value of the z-value, the stronger the correlation, thus rejecting the conditional independence hypothesis.

[0010] Furthermore, the construction of the initial causal graph structure using task structure information and metadata also includes assuming no direct causal connection between variables using the GES algorithm, and starting from an empty graph, a greedy search strategy is adopted to construct the graph structure G GES , adding edges through forward search and removing edges through backward search, maximizing a global scoring function. Among them, starting from an empty graph, the scoring gain of all edges is calculated. In the forward stage, the edge with the highest score is added until no operation can optimize the scoring function. In the backward stage, redundant edges are removed to further optimize the scoring function to ensure that the local optimal solution approximates the global optimal solution, expressed as: Among them, is the probability of variable i given its parent node , | E | is the number of edges, λ is the regularization parameter.

[0011] Furthermore, the construction of the initial causal graph structure using task structure information and metadata also includes using LLM and industrial knowledge graphs to preliminarily optimize the four candidate causal graphs. Among them, using the high-dimensional text embedding model BERT to map each variable of interest in the knowledge prior to a high-dimensional continuous vector space emb , as the basic input for semantic evaluation by LLM; then combining the metadata vector features X total with emb to generate an embedding vector X, and calculate the comprehensive confidence by combining the confidence of multi-source information: fuse multiple candidate causal graphs optimized by the LLM to generate a preliminary comprehensive causal graph G. The fusion process is based on the Bayesian network framework, combining data-driven confidence, LLM semantic confidence, and domain knowledge prior to ensure the connectivity and acyclicity of the graph structure.

[0012] Furthermore, the initial causal graph structure is deeply optimized through the graph neural network to obtain a causal composite relationship graph, including taking the preliminarily constructed causal graph and mixed features as inputs, and successively realizing graph embedding representation, edge weight optimization, and causal path enhancement based on the graph neural network, and using the LLM to perform semantic completion and inference correction on the potential causal structure. Among them, the optimized graph of the current causal graph and all relevant historical causal graphs in the historical causal graph library maintained by the system generate graph-level global node embeddings and edge embeddings, and an asymmetric graph similarity evaluation function S ( G opt , G hist ) is used to quantify the current optimized graph G opt and the historical causal graph G hist in terms of the dual structural and semantic similarity, which is expressed as: where represents the Jaccard similarity of the node set, represents the normalized Hamming distance of the edge set, represents the cosine similarity of the core topological feature vectors of the graph spectrum.

[0013] Furthermore, the deep optimization of the initial causal graph structure through the graph neural network to obtain a causal composite relationship graph also includes integrating the LLM into the training loop of the GNN to form an interactive feedback mechanism. After the GNN is trained for several rounds, the current graph structure and node and edge representations are fed back to the LLM. The LLM combines the semantic verification process to evaluate the rationality of the edge or path and outputs a correction signal. The formula is the joint loss update: where L GNN represents the GNN prediction loss, represents the weight parameter, L LLM_feedback represents the supervised loss generated by the LLM, and the formula is as follows: where C LLM ( e ) represents the edge eLLM semantic confidence p ( e ) represents the GNN predicted edge probability entropy ( G ) represents the graph entropy represents the weight parameter

[0014] Furthermore, the utilization of the causal composite relationship graph and the combination with the knowledge base for fault diagnosis and information retrieval include retrieving the historical fault instances most similar to the current diagnosis result in the local knowledge base through keyword extraction and graph matching technology, calculating the similarity using the joint criterion of semantic vector embedding and causal structure matching. When the similarity is higher than the preset threshold, the system extracts the key information of the matching case and generates a natural language diagnosis report. If the retrieval fails, relevant entries are extracted using uncertainty sampling and the LLM is used to generate structured metadata, and the new knowledge is incrementally integrated into the local knowledge base and fed back to the causal graph to expand the coverage range of the causal graph, recalculating the similarity and generating the report. The similarity is expressed as: where R is the current diagnosis result C is the historical fault case Sim emb ( R , C ) represents the semantic similarity of the fault description Sim topo ( R , C ) represents the structural similarity of the fault causal link represents the weight parameter

[0015] In a second aspect, an industrial fault diagnosis system with intelligent causal correction includes: A data acquisition module configured to acquire monitoring data and user task requirements A task structure module configured to construct task structure information based on the acquired user task requirements A metadata module configured to generate metadata by processing the monitoring data through a three - level cascaded framework A causal graph module configured to construct an initial causal graph structure using the task structure information and metadata A deep optimization module configured to deeply optimize the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph A fault diagnosis module configured to utilize the causal composite relationship graph and combine with the knowledge base for fault diagnosis and information retrieval A report module configured to generate a structured diagnosis report

[0016] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions adapted to be loaded and executed by a processor of a terminal device to implement the industrial fault diagnosis method with intelligent causal correction as described above.

[0017] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium. The processor is configured to implement each instruction, and the computer-readable storage medium is configured to store multiple instructions adapted to be loaded and executed by the processor to implement the industrial fault diagnosis method with intelligent causal correction as described above.

[0018] In summary, the present invention has the following beneficial technical effects: Compared with the prior art, the industrial fault diagnosis system with intelligent causal correction of the present application has the following beneficial effects: Through the innovative structure integrating the LLM and the graph neural network, the system can dynamically optimize the causal relationship model, automatically learn and correct fault associations, and enhance the interpretability of the causal graph by using the semantic reasoning of the LLM and the graph attention mechanism, significantly improving the robustness and accuracy of diagnosis; In addition, the closed-loop optimization mechanism of the system realizes real-time knowledge update and task scheduling, flexibly coping with multivariable interactions and dynamic environmental changes, improving diagnosis efficiency and reliability, and providing a solid foundation for subsequent decision-making, maintenance, and optimization of power equipment fault diagnosis.

[0019] Aiming at the limitations of traditional causal graph construction methods, the system further uses a large language model to perform semantic reasoning and knowledge completion, and combines a graph neural network for feature extraction and fault association mining, enabling the causal graph to more comprehensively adapt to multivariable interactions and dynamic fault scenarios. Based on the enhanced causal graph, the system generates a structured diagnostic report, supports real-time knowledge update and decision assistance, and realizes the intelligence and closed-loop optimization of power equipment fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of an industrial fault diagnosis method with intelligent causal correction according to Embodiment 1 of the present invention; Figure 2 is an example diagram of a causal graph according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Embodiment 1 Referring to Figure 1 , an industrial fault diagnosis method with intelligent causal correction in this embodiment includes: Obtaining monitoring data and user task requirements; Constructing task structure information based on the obtained user task requirements; Generate metadata by processing monitoring data through a three - level cascaded framework; Construct an initial causal graph structure using task structure information and metadata; Deeply optimize the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph; Use the causal composite relationship graph and combine it with a knowledge base for fault diagnosis and information retrieval; Generate a structured diagnostic report.

[0023] Specifically: Step S1, Task Definition: For the task information input by the user in natural language form, the system parses the input text through the Prompt Engineering strategy, constructs three specific prompt templates: diagnostic objectives, variables of interest, and output requirements, identifies and extracts key entities in the text, and maps the extracted entity names and descriptions to a predefined standardized term set and variable list. After being processed by the natural language processing and structured engine, the unstructured task information input by the user is transformed into preliminary structured data and stored in the task information file.

[0024] The system maintains a historical task knowledge base T which stores the structured task definition files that have been successfully processed in the past, their related processing configurations, intermediate results, or even the optimized processing flow parameters. For the current task T1, the system calculates its similarity with one or more historical tasks T2 in the historical task knowledge base, using a hybrid similarity mechanism to comprehensively consider the similarity of the semantic content of the tasks and the set of structured variables: where, emb ( T ) represents T the key entity embedding vector of var ( T ) represents the set of standardized monitoring variables in task T , α 1 represents the weight parameter, Cosine represents the cosine similarity between vectors, reflecting the degree of closeness of tasks in the semantic space, Jaccard represents the proportion of shared variables between sets, reflecting the degree of structural overlap of tasks in key variables.

[0025] If the system detects that there is one or more historical tasks T 2 , and its hybrid similarity with T 1 is higher than the preset similarity threshold of 0.95, then the system determines that T 1With T 2 shared variable dependencies and core diagnostic patterns, secondly, the system triggers the historical resource loading mechanism to directly load the processing configuration and model parameters related to T 2 including the preliminary causal graph structure skeleton constructed by S3 and the GNN model weights trained in step S4.

[0026] Step S2, Metadata Generation: The core function of step S2 is to construct a highly robust and high-quality time series feature matrix as a reliable input for subsequent causal modeling and fault diagnosis modules. The system uses a three-level cascaded framework to process the monitoring device data, covering three sub-modules: data format unification and completion, adaptive noise suppression, and data trend prediction.

[0027] Among them, the format unification and completion sub-module reads multi-source heterogeneous data through an automated script of the pandas library, maps it to a standard CSV structure, and adds a time stamp field in the UTC standard format to achieve time alignment. A two-dimensional feature matrix is constructed with the time stamp field as the main index X , set the monitoring variable dimension to d , and the number of sampling points is n .

[0028] Considering the order of magnitude differences caused by different physical units, dimensions, or measurement ranges in industrial data, there may be extreme values or outliers. The Robust Scaling method is used to perform feature normalization processing, scaling based on the quartiles and median of the data: Among them, median ( X ) represents the median of the feature, IQR ( X ) = Q 3 ( X ) - Q 1 ( X ) represents the interquartile range, Q 3 ( X ) represents the first quartile, Q 1 ( X ) represents the third quartile. Move the center of the data to 0 and scale it according to its interquartile range to make the processed data distribution more concentrated and maintain the consistency of the monitoring variables in the structured data matrix X on the numerical scale.

[0029] For the time series feature matrix XMissing regions commonly existing in [the data] due to reasons such as transmission interruptions and sensor failures. Following the principles of inherent time continuity and local correlation of industrial time-series data, interpolation is performed using the weighted average of several latest valid observed values before the missing point, and the weights decay exponentially with time: Among them, Controls the decay of historical weights, x t Represents the missing point to be interpolated, x t-k Is the moment t - k Of the valid observed value, K Is the selected historical window length. The parameter α 2 Is determined according to the autocorrelation characteristics of the data to retain the local data fluctuations and trends within a short time.

[0030] As a further implementation manner, The adaptive noise suppression sub-module in the present invention aims to process the significant noise problems in the power equipment monitoring data caused by high-voltage environmental electromagnetic interference and sensor accuracy limitations. These noise characteristics usually manifest as high-frequency pulse interference, low-frequency drift, or non-Gaussian distributed noise, which are prone to masking the true signal characteristics and amplifying the risk of false anomalies. The sub-module first estimates the noise level σ By the median absolute deviation: Among them, x i Represents the i Th sampling point in the data sequence, median Represents the median operation. Based on the high-frequency components detected by the estimated noise level, wavelet transform is used for multi-scale decomposition, and the data sequence is decomposed into low-frequency approximation coefficients and high-frequency detail coefficients. The decomposition process can be expressed as: Among them, Represents the wavelet coefficient, s > 0 Represents the scale parameter, which is used to control the fineness of decomposition. A smaller s corresponds to the high-frequency detail coefficient, and a larger s corresponds to the low-frequency approximation coefficient. Represents the time translation parameter, Represents the conjugate of the wavelet function. The low-frequency approximation coefficients remain unchanged, and soft threshold filtering is applied to the high-frequency detail coefficients to adaptively remove the noise components: Among them, λ Based on the noise standard deviation σCalculation. Perform inverse wavelet transform on the filtered wavelet coefficients to reconstruct the signal suppressing high-frequency noise: Among them, C ψ is the constant of the wavelet function, and the reconstructed signal f denoised ( t ) is used as the denoising output to replace the original signal sequence to reduce high-frequency noise interference.

[0031] For low-frequency noise, non-Gaussian distributed noise, or structural noise that is difficult to effectively process by wavelet transform, the system adopts a GAN-based denoising model. The generator of GAN receives the noisy data X noisy , and attempts to generate a denoising output X clean approximating the original pure data : The discriminator D is trained to distinguish between the real pure data X clean and the denoised data generated by the generator. The training objective function is the KL divergence between and X clean . The system optimizes the model on a historical dataset of power equipment containing typical noise patterns, and the training objective function adopts the standard minimax adversarial loss: Among them, P data represents the pure data distribution, and P noisy is the noisy data distribution. The optimized generator regenerates the denoised data X denoised for the input noisy data, and combines it with the wavelet transform output f denoised ( t ) for fusion, and finally outputs the comprehensively suppressed signal X .

[0032] As a further implementation method, The data trend prediction module addresses the problem of delayed diagnosis caused by strong periodicity and sudden change characteristics (such as load fluctuations, seasonal electricity consumption patterns, or fault precursor signals) in power equipment monitoring data. It adopts a hybrid prediction framework of ARIMA and LSTM networks to achieve adaptive prediction of future data trends, improving the prediction accuracy and the reliability of anomaly detection.

[0033] Specifically, the ARIMA model captures linear trends and seasonality through parameters θ =( p , d , q ), where p represents the autoregressive order, d represents the differencing order used to eliminate non-stationarity, q represents the moving average order, and the selection of parameter θ is based on information: where, θ =( p , d , q ) is the parameter vector, p ( x t ∣ x t -1, θ ) is the conditional probability density of the given historical data and parameters. Feature prediction is performed based on the obtained parameters: where, and θ j are the autoregressive and moving average coefficients respectively, is the white noise residual. Parameter estimation uses maximum likelihood estimation, and historical data is fitted through numerical optimization to generate a preliminary linear prediction sequence. To address the limitations of ARIMA in dealing with non-linear sudden change events, an LSTM network is introduced for supplementation, and the outputs of ARIMA and LSTM are combined using a weighted fusion framework: where, horizon represents the prediction step size, w ARIMA and w LSTM are the adaptive weights, satisfying w ARIMA + w LSTM= 1, calculate the prediction errors of ARIMA and LSTM on historical data to dynamically update the weights, ensuring that the ARIMA weight is higher during the stationary period and the LSTM weight increases during mutation events. Generate prediction data X pred and confidence level P , combined with the original detection data X to form the total input X total = X , X pred .

[0034] Upload the structured task file and X total to the system. For the key targets in the task file, the system uses a fuzzy matching algorithm to perform field matching and consistency verification on the X total set of column names. All successfully matched data columns are bound to the task information to generate a complete structured task metadata structure.

[0035] Step S3, Causal relationship graph construction: The said step S3 includes three components: a data-driven causal hypothesis generation component, an LLM semantic verification and preliminary correction component, and a multi-source information integration component, aiming to infer the causal dependencies between variables from power equipment monitoring data to support subsequent fault diagnosis and prediction. The overall process starts with purely statistical-driven hypothesis generation and gradually integrates semantic and domain knowledge to achieve the evolution from candidate graphs to robust comprehensive causal graphs.

[0036] Among them, the data-driven causal hypothesis generation component, based on the set of variables of interest in the metadata, uses four complementary statistical causal structure learning algorithms to perform batch tests on all variable pairs to generate diverse candidate causal graphs.

[0037] First, use the Granger causality test algorithm to construct a causal graph G Granger , and generate a directed graph by analyzing the time series prediction relationship between variables: Among them, RSS resticted is the sum of squared residuals fitted using only the B historical data, RSS unresticted is the sum of squared residuals fitted using the i and j historical data, k represents the lag order, dynamically selected based on autocorrelation function analysis, pdenotes the autoregressive order. The algorithm assumes that if the past values of variable j can significantly improve the prediction of variable i , that is, F the value is higher than the set threshold, then there exists j → i causal relationship.

[0038] Using the PC algorithm based on conditional independence test, it is assumed that there may be potential causal relationships between all variables. Starting from a fully connected graph, the conditional set k is gradually increased, and the conditional independence of variables i and j is tested. For the initial conditional set k, the correlation coefficient r ij∣k of each variable pair (i, j) is calculated, and its conditional independence score Z is tested: where r ij∣k is the partial correlation coefficient of variables i and j under the given conditional set k , n is the sample size. The larger the absolute value of the z-value, the stronger the correlation, thus rejecting the conditional independence hypothesis. By iteratively increasing k , the algorithm gradually removes redundant edges (edges with z-value lower than the threshold), constructs a sparse undirected skeleton graph, and infers the edge direction using Meek's rules to construct a directed graph G PC .

[0039] Using the GES algorithm, it is assumed that there is no direct causal connection between variables. Starting from an empty graph, a greedy search strategy is adopted to construct the graph structure G GES , adding edges through forward search and removing edges through backward search to maximize a global scoring function: where is the probability of variable i given its parent nodes , | E | is the number of edges, λ is the regularization parameter. The algorithm starts from an empty graph, calculates the scoring gain of all edges, adds the edge with the highest score in the forward stage until no operation can optimize the scoring function, and removes redundant edges in the backward stage to further optimize the scoring function to ensure that the local optimal solution approximates the global optimal.

[0040] Aiming at the lag effect and autocorrelation characteristics of power time series data, the PCMCI+ algorithm is used to construct the graph structure starting from a fully connected directed time series graph. For the current variablex , calculate the mutual information of relevant variables in the set of variables of interest: Among them, y represents a single variable in the subset of variables composed of the past time step y t-i and the current time step y t or y t or y t-i , p ( x , y ) is the joint probability distribution, p ( x ) and p ( y ) are the marginal probability distributions. If the mutual information I is lower than the threshold, it does not conform to the causal hypothesis, and the non-causal edges are removed. Finally, the constructed graph G PCMCI is a sparse directed graph, the nodes include variable and time information, and the edges contain time delay information.

[0041] However, pure statistical methods may be limited by the limitations of task information and cannot fully explore the deep causal relationships between variables or capture implicit physical mechanisms. Further, the LLM semantic verification and preliminary correction component introduces the LLM and the industrial knowledge graph to optimize the four candidate causal graphs. Domain knowledge priors and the LLM perform semantic reasoning and probability evaluation to perform semantic rationality verification, conflict identification, and structural adjustment on the graph edges.

[0042] To enable the LLM to understand and operate on the variables of interest and their domain knowledge in its internal semantic space, the system first uses the high-dimensional text embedding model BERT pre-trained specifically for the industrial domain or the general domain to map the detailed domain knowledge description text of each variable of interest in the knowledge prior to the high-dimensional continuous vector space emb , as the basic input for the LLM to perform semantic evaluation ( S semantic function). Combine the metadata vector features X total with emb to generate the embedding vector X . For each candidate edge existing in the candidate graph, calculate the comprehensive confidence by combining the multi-source information confidence: Among them, e represents the edge that actually exists in the candidate graph, x and y respectively represent the edge​e The starting point and the ending point, C LLM ( e ) represents an edge e of the comprehensive confidence level, α 3 、 α 4 、 α 5 are adjustable weight parameters. S semantic ( x , y ) represents semantic similarity, represents the node confidence level in the semantic space, and the calculation formula is as follows: Among them, emb x and emb y respectively represent the BERT embedding vectors of nodes x and y , that is, semantic representations, to quantify the knowledge association degree of variables x and y in the knowledge graph, S semantic ( x , y ) The larger the value, the higher the semantic correlation. Attention ( x , y ) represents the Transformer-based attention, represents the node confidence level in the LLM inference space: Among them, d k is the vector dimension, and the attention weight is calculated by the dot product of nodes x and y . By calculating the attention weight between the semantic embedding vectors of variables x and y to reflect the association strength under a more complex non-linear mapping. R counterfactual Adopts the embedding vector X to evaluate the causal relationship strength, represents the causal counterfactual rationality evaluation score: Among them, do( X = x ) represents the causal intervention operation, E Y |do( X = x )] is to force X ​The value taken is x After intervention Y The expected value. By causal intervention, distinguish correlation and causation to identify the true causal path. The C LLM ( e ) calculated by the LLM is fused with the original weights of the candidate graph to generate an optimized candidate graph. Specifically, consider three cases: ① If the C LLM ( e ) calculated by the LLM is in the same direction as that in the candidate graph, add the weights to enhance the confidence, and limit the final weight value within [0, 1].

[0043] ② If there is x → y in the candidate graph, but C LLM ( e ) infers y → x , the new weight is calculated by the following formula: Where γ is the attenuation factor, and if the direction conflict is reduced by subtraction to handle the weights.

[0044] ③ If C LLM ( e ) > θ but there is no edge e in the candidate graph, then directly add a new edge, and the initial weight W new ( e ) = C LLM ( e ), and generate new nodes through counterfactual reasoning.

[0045] As a further implementation, The multi-source information integration component fuses multiple candidate causal graphs optimized by the LLM to generate a preliminary comprehensive causal graph G. The fusion process is based on the Bayesian network framework, combining data-driven confidence, LLM semantic confidence, and domain knowledge prior to ensure the connectivity and acyclicity of the graph structure.

[0046] First, construct a set containing all potential causal edges, including all candidate causal graphs ( G Granger ,G PC , G GES ,G PCMCI+) The existing edges in it. For any edge e , its final confidence P ( e |all) is calculated by combining the evidence probabilities from different sources: Among them, P ( Data | e ) is the edge existence probability based on a data-driven algorithm, calculated based on feature similarity, P ( LLM | e ) is the confidence obtained based on LLM semantic verification, and the value is derived from C LLM ( e ), P ( Knowledge | e ) is the probability of determining the edge existence based on the domain knowledge graph, obtained through Bayesian inference. By calculating the probability values for all existing edges and normalizing, the final comprehensive confidence P ( e | Data, LLM, Knowledge ) ∈ [0, 1].

[0047] Execute directed cycle checking based on depth-first search DFS on the fused graph structure. If a directed cycle is detected, the system identifies all the edges that form the cycle and removes the edge with the lowest comprehensive confidence in the cycle. Repeat the process of detecting and removing the lowest-confidence edge until the graph no longer contains any directed cycles. The optimized causal graph G, as a directed acyclic graph, has nodes as the variables of interest, and the edges carry the comprehensive confidence, providing prior information for the graph neural network training in step S4.

[0048] Step S4: Causal graph optimization based on graph neural network: To support the deep learning and modeling of causal relationships by GNN, the system constructs a mixed feature representation H of the nodes, X time which consists of temporal features X semantic , semantic features E , edge features X structure and structural features X time are used to capture the dynamic statistical metrics of variables, such as mean, variance, rate of change, autocorrelation, and historical volatility patterns; semantic features X semantic are generated by encoding variable names, descriptions, and related knowledge entries based on the LLM and BERT models, with dimensions ranging from 128 to 512; edge featuresE including the type of the edge and the initial confidence obtained through the LLM semantic verification in S4 C LLM ( e ); Structural features X structure Based on the calculation of the topological attributes of the causal graph G output by S4, its key indicators include the degree of nodes ki , betweenness centrality CB ( v ). Betweenness centrality CB ( v ) measures the bridging role of a node in the shortest path, that is, how many shortest paths pass through this node. The calculation formula is as follows: where, σ st represents the total number of shortest paths of the node pair ( s , t ), σ st ( v ) represents the number of shortest paths passing through the node v . Mixed features H Through attention weighted fusion: where, represents the weight based on feature correlation.

[0049] As a further implementation manner, The causal graph enhancement and optimization training module in the present invention includes three parts: a causal graph embedding layer, a graph structure optimization layer, and a causal relationship reconstruction layer. This module takes the preliminarily constructed causal graph and mixed features as inputs, realizes graph embedding representation, edge weight optimization, and causal path enhancement based on graph neural networks, and uses the LLM to perform semantic completion and inference correction on potential causal structures.

[0050] Among them, the causal graph embedding layer receives the causal graph G =( V , E ) and its corresponding node feature matrix X ∈ R n*d , and first performs self-looping and normalization processing on the causal graph to obtain a normalized adjacency matrix: where, A is the original adjacency matrix, I is the identity matrix, D is the node degree matrix, α 6Self-loop adjustment factor. Introduce a multi-scale receptive field aggregation mechanism: Among them, W (l) is the trainable weight matrix of the l th layer, ReLU represents the activation function, FFN represents the feed-forward neural network, || represents vector concatenation, k represents the hop distance on the graph, K is the maximum considered hop count to cover the multi-level cascading fault propagation range in the power system, represents node i 's neighbor set, α ij represents the attention weight, calculated through the learnable power flow sensitivity matrix: Among them, P ij represents the physical prior term derived from the power flow sensitivity matrix. To retain the initial feature information, introduce a residual connection mechanism: Among them, represents the node representation after applying the residual connection, represents the node-specific residual weight factor, represents the residual transformation matrix, l represents the number of model layers. Each layer of the model performs multi-scale receptive field aggregation and residual connection. Through the l -layer model, node features are obtained. The edge weights are adjusted according to the embedding similarity of the two ends of each edge to improve the rationality and distinguishability of the causal structure. Let the embedding vectors of the node pair ( i , j ) be h i , h j Define the edge weight using a high-order similarity function: Among them, represents the estimation of the feature covariance matrix, represents the mean vector of all node embeddings. Combining the known causal paths annotated in the training set, the edge weights are learned and updated through a weighted edge classification loss function: Among them, y ij∈{0,1} is the edge label, which is derived from the semantic reasoning of the domain knowledge graph and the LLM. The GNN iteratively updates the edge weights, adds or deletes edges according to w ij and outputs the optimized causal graph G opt .

[0051] As a further implementation, The causal relationship reconstruction layer aims to transcend the limitations of a single causal graph through cross-graph analysis, deeply analyze the association patterns between multiple causal graphs accumulated historically and related to the current diagnosis task, and construct a similarity network between causal graphs to mine composite causal links across time or scenarios. To compare and associate different causal graphs in a unified vector space, the current optimized graph G opt and all relevant historical causal graphs in the historical causal graph library maintained by the system G hist generate global node embeddings and edge embeddings at the graph level: Establish an asymmetric graph similarity evaluation function S ( G opt , G hist ) to quantify the structural and semantic double similarity between the current optimized graph G opt and the historical causal graph G hist : Among them, represents the Jaccard similarity of the node set, represents the normalized Hamming distance of the edge set, represents the cosine similarity of the core topological feature vectors of the graph spectrum. By setting an adaptive threshold τ sim (G opt ): Among them, density ( G opt ) represents the edge density of the current graph, μ and η are adjustable parameters. Screen high-similarity historical graphs to form a knowledge transfer source set G sim = { G hist ∣ S( G opt , Ghist ) > τ sim (G opt )。

[0052] For G sim each graph G hist in G opt , standardize the matching of all variables with the set of variables of interest in the current graph , and analyze the consistency of the pre - causal chain of the matching nodes to determine whether to migrate and update the structure of Gopt. The pre - causal chain is defined as a sequence of paths recursively extended in the in - degree direction from the target variable. For each matching variable node v, construct the pre - causal chain on both graphs and traverse recursively as follows: starting from the current variable, extend along the in - degree edge. If the weight of the edge is greater than the threshold , then include the node and repeat until the weight is below the threshold or the maximum depth of 3 is reached to prevent infinite recursion: denotes the union operation, Chainpre(v) denotes the pre - causal chain of node v v denotes the set of nodes of the pre - causal chain of node v for nodes u that meet the requirements. If the chain sequences match exactly (i.e., the number of nodes and the edge directions are the same), then perform weighted average of the edge weights; if they are inconsistent, prefer the longer chain and directly add and update the non - repeating part to the corresponding parts of G opt and G hist . The confidence of the newly added edge is obtained based on the mutual information of historical data.

[0053] On this basis, the system performs multi - graph collaborative inference and optimizes through the following formula to achieve unified integration of cross - graph knowledge: where L align is the inter - graph alignment loss based on KL - divergence, is the trade - off parameter. Based on the constructed causal graph similarity network, the causal relationship reconstruction layer performs a data - driven causal knowledge transfer and multi - graph collaborative inference process. Different from inferring only relying on local or global information of the current graph, this method regards the current graph as a node in the network and uses the causal graph with high similarity to it as an external knowledge source to assist in the reconstruction and optimization of the current graph and construct a "composite causal chain" to identify multi - hop paths such as "fan failure → temperature increase → motor overload".

[0054] To achieve intelligent causal optimization, the system integrates the LLM into the training loop of the GNN to form an interactive feedback mechanism: after the GNN is trained for several rounds, the current graph structure and node / edge representations are fed back to the LLM. The LLM evaluates the rationality of edges or paths in combination with the semantic verification process of S4 and outputs a correction signal, and the formula is the joint loss update: where, L GNN represents the GNN prediction loss, represents the weight parameter, L LLM_feedback represents the supervised loss generated by the LLM, and the formula is as follows: where, C LLM ( e ) represents the LLM semantic confidence of edge e , p ( e ) represents the GNN predicted edge probability, entropy ( G ) represents the graph entropy, represents the weight parameter. The output of the LLM is converted into a supervised signal for the GNN, such as adding a loss term to penalize conflicting edges or strengthening high-confidence paths.

[0055] The final optimized causal graph G opt includes enhanced node / edge representations, dynamic weights, and cross-graph association information as the input for S5 fault diagnosis.

[0056] Step S5, Fault Diagnosis and Knowledge Retrieval: Since the causal graph output by S4 only provides structured potential causal paths, lacking specific historical context, risk assessment, and natural language expression, and cannot directly generate a reliable diagnostic report, the system needs to combine the local knowledge base to generate a fault identification result. Specifically, the system retrieves the historical fault instances most similar to the current diagnosis result in the local knowledge base through keyword extraction and graph matching technology. The similarity calculation adopts a joint criterion of semantic vector embedding and causal structure matching, and the following similarity function is defined: where, R is the current diagnosis result, C is the historical fault case, Sim emb ( R , C ) represents the semantic similarity of the fault description, Sim topo ( R ,C ) represents the structural similarity of the fault causal link represents the weight parameter. When the similarity is higher than the preset threshold, the system extracts the key information of the matching case and generates a natural language diagnosis report, including the fault location, potential causes, handling suggestions, risk level, and alarm suggestions.

[0057] If the retrieval fails, the system triggers an external knowledge completion mechanism to handle novel fault modes. First, the system queries multi-source external knowledge sources online, extracts relevant entries using uncertainty sampling, and generates structured metadata with the LLM. The new knowledge is incrementally integrated into the local knowledge base and fed back to the causal graph generation module in S3 to expand the coverage of the causal graph. The update is limited to the relevant part of the current task to avoid global recomputation. Based on the updated knowledge base, the similarity is recalculated and the report is generated again.

[0058] Finally, the module pushes the diagnosis results in both structured data and natural language forms to the user interface and the upper-layer control system, providing an accurate decision-making basis for fault response.

[0059] Step S6, System Verification and Feedback Optimization: The system actively obtains the feedback annotation based on the user's professional judgment. For positive feedback, the system will solidify the processing path, model configuration, and key parameters of this task as an optimization reference example to enhance the performance and efficiency of future similar tasks. For negative feedback, the system starts a reverse analysis process, associates the user feedback information with the diagnosis process log to form a tagged historical task sample library, and constructs a closed-loop optimization system in combination with the active learning strategy and LLM-assisted feedback explanation. First, causal deviation identification is performed, and the deviation index is calculated for the GNN model in S4. Calculate the deviation index: If the deviation exceeds the threshold, the system identifies the potential problem edges or subgraphs and triggers the correction process. After the GNN model parameters are corrected based on the feedback data, a more user-cognitive and optimized causal graph will be generated. The system uses this improved causal graph to re-execute the diagnosis process in step S5 until the user gives positive feedback on the regenerated diagnosis report or reaches the preset maximum number of iterations. In addition, the system introduces an active learning mechanism. When the diagnosis confidence is low or there is a knowledge conflict, the historical knowledge base is preferentially updated. The active learning mechanism is a global optimization triggered by feedback analysis, and the external knowledge completion mechanism targets the local knowledge gap of the current task.

[0060] As Figure 2 shown, it describes multiple causal paths that cause overload faults. Figure 2As shown, the overload fault is directly caused by the increase in winding temperature, which may be due to two factors: the decrease in heat dissipation capacity or the excessive input voltage. Among them, the decrease in heat dissipation capacity can be caused by abnormal fans, forming a complex causal chain.

[0061] Embodiment 2 This embodiment provides an industrial fault diagnosis system with intelligent causal correction. The system is designed with a modular architecture and includes six core functional modules: a task definition and standardization module, a metadata generation module, a causal relationship graph construction module, a causal graph optimization module based on graph neural networks, a fault diagnosis and knowledge retrieval module, and a system verification and feedback optimization module. Data interaction and process coordination are achieved between modules through standardized interfaces, and a complete fault diagnosis closed-loop is formed through seven consecutive processing steps: Step S1, first receive and standardize the user's task requirements; Step S2, comprehensively preprocess and improve the quality of the uploaded monitoring data, and generate structured metadata as the unified entry for subsequent processing; Step S3, on this basis, use multi-source algorithms and LLM to construct an initial causal relationship graph; Step S4, deeply optimize and correct the causal structure through graph neural networks, and construct a causal composite relationship graph; Step S5, implement accurate fault diagnosis and information retrieval in combination with the knowledge base; Step S6, finally, achieve system verification and continuous optimization through user feedback. This process design ensures that the system has the ability of adaptive learning and continuous evolution characteristics, and can effectively handle complex fault scenarios in the industrial environment.

[0062] A computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor of a terminal device to perform the industrial fault diagnosis method with intelligent causal correction.

[0063] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to perform the industrial fault diagnosis method with intelligent causal correction.

[0064] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. An industrial fault diagnosis method with intelligent causal correction, characterized in that, Including: Obtain monitoring data and user task requirements; Construct task structure information based on the obtained user task requirements; Process the monitoring data through a three-level cascaded framework to generate metadata; Construct an initial causal graph structure using the task structure information and metadata; Deeply optimize the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph; Perform fault diagnosis and information retrieval using the causal composite relationship graph in combination with a knowledge base; Generate a structured diagnostic report.

2. The industrial fault diagnosis method with intelligent causal correction according to claim 1, characterized in that The process of processing the monitoring data through a three-level cascaded framework to generate metadata includes first performing feature normalization on the monitoring data using the Robust Scaling method, scaling based on the quartiles and median of the data, and for the missing regions in the time series feature matrix, using the weighted average of several latest valid observed values before the missing point for interpolation; then estimating the noise level in the monitoring data through the median absolute deviation, and for the high-frequency components detected based on the estimated noise level, performing multi-scale decomposition using wavelet transform to decompose the data sequence into low-frequency approximation coefficients and high-frequency detail coefficients; for low-frequency noise, non-Gaussian distributed noise, or structural noise, using a GAN-based denoising model for suppression.

3. An industrial fault diagnosis method with intelligent causal correction according to claim 2, characterized in that The construction of the initial causal graph structure using task structure information and metadata includes first constructing a causal graph G using the Granger causality test algorithm Granger , generating a directed graph by analyzing the time series prediction relationship between variables, expressed as: Among them, RSS resticted is the sum of squared residuals fitted using only the historical data of B, RSS unresticted is the sum of squared residuals fitted using the historical data of i and j. k represents the lag order, which is dynamically selected based on autocorrelation function analysis. p represents the autoregressive order. The algorithm assumes that if the past values of variable j can significantly improve the prediction of variable i, that is, the F value is higher than the set threshold, then there is a causal relationship of j→i.

4. An industrial fault diagnosis method with intelligent causal correction according to claim 3, characterized in that The construction of the initial causal graph structure using task structure information and metadata further includes using the PC algorithm based on conditional independence testing to assume that there may be potential causal relationships between all variables, starting from a fully connected graph and gradually increasing the conditional set k, and testing the conditional independence of variables i and j; by iteratively increasing k, the algorithm gradually removes redundant edges to construct a sparse undirected skeleton graph, and combines the Meek rule to infer the edge direction to construct a directed graph GPC. Among them, for the initial conditional set k, calculate the correlation coefficient r of each variable pair (i, j) ij∣k , and test its conditional independence score Z: where r ij∣k is the partial correlation coefficient of variables i and j under the given condition set k, n is the sample size, and the larger the absolute value of the z-value, the stronger the correlation, thus rejecting the conditional independence hypothesis.

5. An industrial fault diagnosis method with intelligent causal correction according to claim 4, characterized in that, The construction of the initial causal graph structure using task structure information and metadata further includes assuming no direct causal connection between variables using the GES algorithm, and constructing the graph structure G starting from an empty graph using a greedy search strategy GES , adding edges through forward search and removing edges through backward search to maximize a global scoring function. Among them, the scoring gain of all edges is calculated starting from the empty graph. In the forward stage, the edge with the highest score is added until no operation can optimize the scoring function. In the backward stage, redundant edges are removed to further optimize the scoring function to ensure that the local optimal solution approximates the global optimal solution, expressed as: Among them, is the probability that variable i gives its parent node , |E| is the number of edges, and λ is the regularization parameter.

6. The industrial fault diagnosis method with intelligent causal correction according to claim 5, characterized in that, The construction of the initial causal graph structure using task structure information and metadata also includes preliminarily optimizing the four candidate causal graphs using an LLM and an industrial knowledge graph. Among them, the high-dimensional text embedding model BERT is used to map each variable of interest in the knowledge prior to a high-dimensional continuous vector space emb, which serves as the basic input for the LLM to perform semantic evaluation. Then, the metadata vector feature X total is combined with emb to generate an embedding vector X, and the comprehensive confidence is calculated by combining the multi-source information confidence: multiple candidate causal graphs optimized by the LLM are fused to generate a preliminary comprehensive causal graph G. The fusion process is based on the Bayesian network framework, combining data-driven confidence, LLM semantic confidence, and domain knowledge prior to ensure the connectivity and acyclicity of the graph structure.

7. An industrial fault diagnosis method with intelligent causal correction according to claim 6, characterized in that, The deep optimization of the initial causal graph structure through the graph neural network to obtain the causal composite relationship graph includes using the preliminarily constructed causal graph and mixed features as inputs, successively implementing graph embedding representation, edge weight optimization, and causal path enhancement based on the graph neural network, and using the LLM to perform semantic completion and inference correction on the potential causal structure. Among them, the optimized graph of the current causal graph and all relevant historical causal graphs in the historical causal graph library maintained by the system are used to generate graph-level global node embeddings and edge embeddings, and an asymmetric graph similarity evaluation function S(G opt ,G hist ) is established to quantify the structural and semantic double similarity between the current optimized graph G opt and the historical causal graph G hist , which is expressed as: Among them, represents the Jaccard similarity of the node set, represents the normalized Hamming distance of the edge set, represents the cosine similarity of the spectral core topological feature vectors.

8. An industrial fault diagnosis method with intelligent causal correction according to claim 7, characterized in that, The process of deeply optimizing the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph further includes integrating the LLM into the training loop of the GNN to form an interactive feedback mechanism. After the GNN is trained for several rounds, the current graph structure and node and edge representations are fed back to the LLM. The LLM combines the semantic verification process to evaluate the rationality of the edges or paths and outputs a correction signal, and the formula is for joint loss update: Among them, L GNN represents the GNN prediction loss, represents the weight parameter, and L LLM_feedback represents the supervised loss generated by the LLM. The formula is as follows: Among them, C LLM (e) represents the LLM semantic confidence of edge e, p(e) represents the GNN predicted edge probability, and entropy(G) represents the graph entropy, represents the weight parameter.

9. An industrial fault diagnosis method with intelligent causal correction according to claim 8, characterized in that, The process of performing fault diagnosis and information retrieval using the causal composite relationship graph in combination with a knowledge base includes, through keyword extraction and graph matching technology, retrieving the historical fault instances most similar to the current diagnosis result in the local knowledge base, calculating the similarity using the joint criterion of semantic vector embedding and causal structure matching. When the similarity is higher than the preset threshold, the system extracts the key information of the matching case and generates a natural language diagnostic report. If the retrieval fails, relevant entries are extracted using uncertainty sampling and the LLM generates structured metadata, and the new knowledge is incrementally integrated into the local knowledge base and fed back to the causal graph to expand the coverage range of the causal graph, recalculating the similarity and generating a report. The similarity is expressed as: Among them, R is the current diagnosis result, C is the historical failure case, Sim emb (R, C) represents the semantic similarity of the fault description, Sim topo (R, C) represents the structural similarity of the fault causal link, represents the weight parameter.

10. An industrial fault diagnosis system with intelligent causal correction, characterized in that, Including: A data acquisition module configured to obtain monitoring data and user task requirements; A task structure module configured to construct task structure information based on the obtained user task requirements; A metadata module configured to process the monitoring data through a three-level cascaded framework to generate metadata; A causal graph module configured to construct an initial causal graph structure using the task structure information and metadata; A deep optimization module configured to deeply optimize the initial causal graph structure through a graph neural network to obtain a causal composite relationship graph; A fault diagnosis module configured to perform fault diagnosis and information retrieval using the causal composite relationship graph in combination with a knowledge base; A report module configured to generate a structured diagnostic report.

Citation Information

Patent Citations

  • Industrial equipment running state evaluation system for multi-parameter coupling

    CN118859868A

  • Intelligent auxiliary diagnosis and maintenance method and system based on multi-path recall

    CN119357787A

  • Abnormal behavior detection method and device for networked device, and computer program product

    CN119416130A

  • Interdependent causal networks for root cause localization

    US20230069074A1

Cited By

  • Fault diagnosis method and device, medium and product

    CN120492503A

  • Fault maintenance scheme automatic recommendation method based on large language model

    CN120509415A

  • Graph comparison learning method for solving graph combination optimization problem

    CN120523870A

  • Disease monitoring and diagnosing system for allergic rhinitis and asthma syndrome treated by traditional Chinese medicine

    CN120600217A

  • Deep neural network delay method based on graph attention mechanism and large language model

    CN120611799A