A method for causal tracing of faults in heterogeneous time series data

Through maximum and minimum value standardization and context adaptive excitation Gaussian core embedding technology, the causal analysis problem of mixed-type time series data in industrial systems is solved, and the accurate identification of causal relationships and the traceability of fault root causes is achieved.

CN119538992BActive Publication Date: 2025-05-13ZHEJIANG UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510109009.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Existing causal analysis methods are difficult to effectively deal with mixed-type time series data in industrial systems, especially the information asymmetry problem between continuous and discrete variables, which leads to difficulty in accurately identifying causal relationships.

Method used

A method of causal traceability for heterogeneous time sequence data failure is proposed, continuous variables are standardized through maximum and minimum values, and context adaptive excitation Gaussian kernel embedding to restore the implicit continuity of discrete variables, and a unified causal model is constructed.

Benefits of technology

It significantly improves the accuracy of causal relationship extraction, enhances the model's ability to recover the potential continuity of discrete variables, and realizes accurate identification of causal relationships and traces the root cause of failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538992B_ABST
    Figure CN119538992B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for causal tracing of heterogeneous time series data faults, which converts discrete variables into latent continuous variables with higher information granularity through context-adaptive excitation Gaussian kernel embedding, thereby realizing causal discovery in a unified continuous space; in the potential continuity recovery stage, the guiding information of continuous variables is introduced by designing prediction tasks, and the parameters of context-adaptive excitation Gaussian kernel embedding are adjusted in a self-supervisory manner to enhance its ability to restore potential continuity, and the reversibility of the recovery process is ensured through a reconstruction mechanism; in the causal structure learning stage, key causal relationships are screened through sparsity regularization constraints, and an overall causal graph is constructed to provide support for fault analysis. The present invention can accurately determine the time series causal relationship in the heterogeneous variable scenario, provide important and reliable information for tracing the root cause of process failures, and help accurately diagnose the source of the failure, thereby effectively ensuring the safety of actual production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault causal tracing in process industries, and in particular to a method for tracing the fault causal tracing of heterogeneous time series data. Background Art

[0002] Fault causal tracing is of great significance in industrial production. Its core goal is to reveal the causal relationship between equipment status, process parameters and key quality indicators, so as to achieve accurate fault analysis, diagnosis and traceability. In the traditional fault handling mode, tracing work usually relies on manual experience and preset rules, but this method is difficult to effectively cope with the challenges brought by multivariate interaction and nonlinear characteristics in complex industrial environments. Causal analysis can mine potential causal relationships from large-scale historical data by constructing a systematic causal model, accurately identify the source of faults and their propagation paths, provide scientific explanations and clear basis, so as to quickly locate problems and formulate targeted repair strategies. In actual production, causal analysis can not only improve the accuracy and timeliness of fault detection, but also effectively identify the key factors that cause abnormalities, providing strong support for real-time monitoring and intelligent early warning. By quantifying the contribution of different process variables to abnormal behavior, the root cause of the problem can be accurately located, quality defects and equipment failures can be fundamentally solved, production interruptions and resource waste can be avoided, and maintenance costs and operational risks can be significantly reduced. In addition, causal analysis can also reveal the deep relationship between process parameters and product performance, provide scientific guidance for process optimization and quality control, thereby improving product qualification rate and enhancing process stability. With the accelerated development of industry and intelligent manufacturing, data-driven causal analysis is gradually becoming a core means to improve the level of intelligence in industrial systems. Through causal analysis, industrial enterprises can achieve predictive maintenance, accurate fault diagnosis and real-time optimization, and greatly improve equipment health management capabilities. This technology has laid a solid foundation for the transformation of industrial production to intelligence and automation, and has provided strong support for improving production efficiency and market competitiveness.

[0003] Currently, most causal analysis methods are based on a core assumption that time series data consists entirely of continuous value data. However, in many practical application scenarios, the collected data is usually of mixed types, including discrete variables (DVs) and continuous variables (CVs). In industrial systems, some process variables (such as temperature and flow) are usually measured in continuous form, while alarm signals are recorded in discrete form. In microservice systems, continuous key performance indicators (KPIs) (such as CPU and memory usage) are often stored together with discrete event sequences to support system stability monitoring. However, causal analysis of mixed data is still an emerging and challenging task. Specifically, the difference in information granularity and distribution type between DVs and CVs leads to information asymmetry (InfoA) problems, making it difficult to accurately identify the causal relationship between the two. When dealing with mixed data, common strategies include ignoring DVs, applying traditional causal analysis methods only to CVs to generate partial causal graphs, or discretizing CVs to infer discrete causal networks. However, mixed time series in actual industrial systems often have complex nonlinear and high-dimensional characteristics, and traditional methods cannot accurately model the causal relationship between the two variables. In addition, due to the limitations of measurement accuracy and storage efficiency, many signals that are continuous in nature are usually recorded in discrete form. For example, whether certain continuous pressure variables exceed the control limit may be recorded in the form of a binary alarm signal. These DVs may originate from latent continuous variables (LCVs), while the implicit continuity of DVs, that is, the intrinsic information of LCVs, cannot be directly observed. The discretization process not only leads to the loss of continuous information of LCVs, but also changes its distribution characteristics, further exacerbating the difficulty of identifying the causal relationship between DVs and CVs. Based on this phenomenon, recovering the LCVs behind the DVs and restoring their original distribution type is a feasible solution to the InfoA problem. By enriching the information content of DVs, information differences can be compensated and the accuracy of causal relationship modeling can be enhanced.

[0004] In order to solve the variable heterogeneity problem of mixed time series data in causal analysis models in industrial processes, a method integrating implicit continuity recovery and fault causal tracing is urgently needed to provide an effective solution for causal analysis of complex industrial systems. Summary of the invention

[0005] The purpose of the present invention is to provide a method for tracing the cause and effect of heterogeneous time series data faults in view of the deficiencies of the prior art. The present invention can extract cause and effect relationships from heterogeneous industrial process data to locate and explain the root causes of faults, and ultimately guide fault diagnosis and maintenance decisions.

[0006] The objective of the present invention is to achieve the following technical solution: A method for tracing the cause and effect of a heterogeneous time series data fault, comprising the following steps:

[0007] (1) Collect heterogeneous time series fault data, including continuous variables and discrete variables, and use the maximum and minimum value standardization method to process the continuous variables to construct a training set;

[0008] (2) Use context-adaptive excitation Gaussian kernel embedding to process discrete variables in the training set and recover the corresponding latent continuous variables;

[0009] (3) For each continuous variable and discrete variable, a corresponding continuous variable predictor and discrete variable decoder are constructed, and the training set constructed in step (1) and the latent continuous variable recovered in step (2) are used for training. The total loss function of the latent continuous learning phase is minimized as the optimization objective, and the parameters of the continuous variable predictor, discrete variable decoder, and context-adaptive excitation Gaussian kernel embedding are optimized to obtain the trained discrete variable decoder and context-adaptive excitation Gaussian kernel embedding;

[0010] (4) Using the discrete variable decoder and context-adaptive excitation Gaussian kernel embedding trained in step (3), using the data in the training set to perform causal structure learning, taking minimizing the total loss function of the causal structure learning stage as the optimization goal, optimizing the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor, so as to obtain the trained latent continuous variable predictor, the continuous variable predictor and the context-adaptive excitation Gaussian kernel embedding; and obtaining the overall causal graph based on the parameters of the trained latent continuous variable predictor and the continuous variable predictor;

[0011] (5) The context-adaptive excitation Gaussian kernel embedding trained in step (4) is used to restore the discrete variables in the online heterogeneous time series fault data into latent continuous variables, and the variables are input together with the continuous variables in the online heterogeneous time series fault data into the latent continuous variable predictor and the continuous variable predictor trained in step (4) to obtain the overall causal graph for process fault analysis.

[0012] Furthermore, the collecting of heterogeneous time series fault data specifically includes:

[0013] Collecting a collection of heterogeneous time series fault data ,in represents the pth variable, P represents the total number of variables, and T represents the length of the variable time series; the set Contains PN continuous variables and N discrete variables. The set of PN continuous variables is expressed as , the set of N discrete variables is expressed as ,in represents the i-th continuous variable, represents the jth discrete variable, , Represents the number of discrete states.

[0014] Furthermore, the context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables in the training set to obtain the corresponding hidden continuous variables, specifically: for the jth discrete variable time series , input it into the corresponding context-adaptive excitation Gaussian kernel embedding to obtain the discrete variable time series The corresponding latent continuous variable , expressed as:

[0015] ;

[0016] in, Represents a discrete variable time series The corresponding latent continuous variable is represents the j-th discrete variable time series, ; Represents the context-adaptive excitation Gaussian kernel embedding corresponding to the j-th discrete variable; represents the softmax activation function; represents the weight coefficient of the multi-scale excitation Gaussian kernel, represents the weight coefficient of the excitation Gaussian kernel of the rth scale, and d is the number of multi-scale excitation Gaussian kernels; Trainable weight coefficient parameters for context-adaptive excitation Gaussian kernel embedding, Represents the trainable weight coefficient parameters of the excitation Gaussian kernel of the rth scale; the bandwidth of the Gaussian kernel representing the multiscale temporal patterns, represents the bandwidth of the Gaussian kernel of the rth scale; and The trainable bandwidth parameters of the context-adaptive excitation Gaussian kernel embedding are used to control the upper and lower limits of the bandwidth of the Gaussian kernel. The other bandwidths are determined according to the following formula:

[0017] .

[0018] Furthermore, the context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables to obtain the corresponding continuous value time series, and multi-scale adaptive weighted aggregation is performed to integrate the multi-scale time information of the discrete variables to restore the corresponding hidden continuous variables, which specifically includes the following process:

[0019] (2.1) The Gaussian kernel in the context-adaptive excitation Gaussian kernel embedding has a smooth property and can structurally characterize the temporal adjacency density, which is expressed as:

[0020] ;

[0021] in, represents the Gaussian kernel, and is any time point in the discrete variable time series, represents the bandwidth of the Gaussian kernel;

[0022] (2.2) Use the Gaussian kernel determined in step (2.1) to process discrete variables. Specifically, for a given j-th discrete variable time series , consider the time point with non-zero value as the excitation time point, assuming is an excitation time point in the time series, and a bandwidth of The Gaussian kernel of is normalized and its expression is:

[0023] ;

[0024] in, represents the Gaussian kernel, represents the normalization factor, Represents any point in time and incentive timing The distance between them; by superimposing the discrete variable time series in the manner shown in the above formula Gaussian kernel function corresponding to all excitation time points on the image to obtain the discrete variable time series A sequence of continuous values ​​at time point t, expressed as:

[0025] ;

[0026] in, Represents a discrete variable time series A sequence of consecutive values ​​at time point t, Represents the jth discrete variable time series The value at the tth time point in ;

[0027] (2.3) Based on the discrete variable time series obtained in step (2.2) At the continuous value sequence of time point t, by aggregating the excitation Gaussian kernel functions of various scales according to their corresponding weights, the multi-scale time information of discrete variables is integrated to realize the restoration of implicit continuity and obtain the implicit continuous variable corresponding to time point t , expressed as:

[0028] ;

[0029] in, is the noise term that follows the normal distribution, that is ;

[0030] (2.4) Based on the discrete variable time series obtained in step (2.3) Latent continuous variable at time point t Obtaining discrete variable time series The corresponding latent continuous variable , expressed as .

[0031] Furthermore, the continuous variable predictor and discrete variable decoder are both implemented using a multi-layer fully connected neural network, and prediction and decoding are achieved through multi-layer linear transformations and nonlinear activation functions, wherein the network parameters of each layer are randomly initialized, and the network parameters of the continuous variable predictor and the discrete variable decoder are different.

[0032] Furthermore, the continuous variable predictor and the discrete variable decoder are trained using the training set constructed in step (1) and the hidden continuous variables recovered in step (2), with minimizing the total loss function of the potential continuity learning stage as the optimization goal, and the parameters of the continuous variable predictor, the discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding are optimized to obtain the trained discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding, which specifically includes the following process:

[0033] (3.1) The continuous variables in the training set The historical values ​​and the recovered hidden continuous variables The historical values ​​of are input into the continuous variable predictor to obtain the continuous variable Predicted value at a future time , expressed as:

[0034] ;

[0035] in, Represents a continuous variable The predicted value at time t; represents the continuous variable predictor corresponding to the i-th continuous variable; is the specified history window length; Represents a continuous variable from Time has come The historical value of the moment, represents a latent continuous variable from Time has come Historical value at a moment;

[0036] (3.2) Calculation of continuous variables The mean square error between the predicted value at the future moment and its corresponding true continuous variable value is used as the prediction error loss function of the continuous variable predictor;

[0037] (3.3) will recover the obtained latent continuous variable The historical values ​​of are input into the discrete variable decoder to reconstruct the original discrete variable , get the reconstructed discrete variables , expressed as:

[0038] ;

[0039] in, represents the reconstructed discrete variable at time t, represents the discrete variable decoder corresponding to the jth discrete variable;

[0040] (3.4) Calculate the cross entropy loss between the reconstructed discrete variable and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder;

[0041] (3.5) Calculate the total loss function of the latent continuous learning stage based on the prediction error loss function of the continuous variable predictor and the reconstruction error loss function of the discrete variable decoder; optimize the parameters of the continuous variable predictor, discrete variable decoder and context-adaptive excitation Gaussian kernel embedding with the optimization goal of minimizing the total loss function of the latent continuous learning stage to obtain the trained discrete variable decoder and context-adaptive excitation Gaussian kernel embedding.

[0042] Furthermore, the calculation formula of the prediction error loss function of the continuous variable predictor is:

[0043] ;

[0044] in, represents the prediction error loss function for continuous variable predictors, represents the mean square error, Represents a continuous variable The predicted value at time t is Represents a continuous variable The true continuous variable value at time t, represents the vector two norm;

[0045] The calculation formula of the reconstruction error loss function of the discrete variable decoder is:

[0046] ;

[0047] in, represents the reconstruction error loss function of the discrete variable decoder, represents the cross entropy loss, represents the reconstructed discrete variable at time t, Represents the original discrete variable The real discrete variable at time t, Represents the reconstructed discrete variable The value corresponding to the cth discrete state of Represents a real discrete variable The value corresponding to the cth discrete state of represents the natural logarithm;

[0048] The calculation formula of the total loss function in the potential continuity learning stage is:

[0049] ;

[0050] in, represents the total loss function of the potential continuous learning stage, and They are and The weight coefficient of .

[0051] Furthermore, the step (4) includes the following sub-steps:

[0052] (4.1) Fix the network parameters of the trained discrete variable decoder; fix the bandwidth parameters of the trained context-adaptive excitation Gaussian kernel embedding, and keep its weight coefficient parameters trainable; build a hidden continuous variable predictor, whose network parameters are randomly initialized; use the continuous variable predictor built in step (3), whose network parameters are randomly initialized; wherein, the hidden continuous variable predictor is implemented using a multi-layer fully connected neural network, and prediction is achieved through multi-layer linear transformation and nonlinear activation function;

[0053] (4.2) Input the discrete variables in the training set into the context-adaptive excitation Gaussian kernel embedding trained in step (3) to obtain the corresponding hidden continuous variables ;

[0054] (4.3) The latent continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the continuous variable predictor constructed in step (3) to obtain the continuous variable Predicted value at a future time ; The latent continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the latent continuous variable predictor constructed in step (4.1) to obtain the latent continuous variable Predicted value at a future time ; To complete the prediction task under the Granger causality paradigm;

[0055] (4.4) Calculation of continuous variables The mean square error between the predicted value at the future moment and its corresponding true continuous variable value and the hidden continuous variable The mean square error between the predicted value at the future moment and its corresponding true latent continuous variable value, and the sum of the two mean square errors is used as the total prediction error loss function;

[0056] (4.5) The latent continuous variable obtained in step (4.2) The historical values ​​of are input into the discrete variable decoder trained in step (3) to reconstruct the original discrete variable , get the reconstructed discrete variables ;

[0057] (4.6) Calculate the cross entropy loss between the reconstructed discrete variable obtained in step (4.5) and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder ;

[0058] (4.7) applying a sparsity-inducing constraint, i.e., a hierarchical group minimum absolute shrinkage and selection operator, on each latent continuous variable predictor and the first layer of each continuous variable predictor to obtain the corresponding sparsity-constrained loss function;

[0059] (4.8) Calculate the total loss function of the causal structure learning stage according to the total prediction error loss function obtained in step (4.4), the reconstruction error loss function of the discrete variable decoder obtained in step (4.6), and the sparse constraint loss function obtained in step (4.7); Taking minimizing the total loss function of the causal structure learning stage as the optimization goal, the proximal gradient descent method is used to optimize the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor, so as to obtain the optimal context-adaptive excitation Gaussian kernel embedding weight coefficient parameters, the network parameters of the latent continuous variable predictor and the continuous variable predictor, and obtain the final trained context-adaptive excitation Gaussian kernel embedding, latent continuous variable predictor and continuous variable predictor;

[0060] (4.9) Based on the optimal implicit continuous variable predictor and continuous variable predictor network parameters obtained in step (4.8), determine the variable to be predicted The corresponding latent continuous variable predictor or continuous variable predictor weight matrix ; According to the weight matrix Perform Granger causality judgment to obtain the variables to be predicted The corresponding Granger causality judgment results; obtain the overall causal diagram based on the Granger causality judgment results of all variables to be predicted.

[0061] Furthermore, the calculation formula of the total prediction error loss function is:

[0062] ;

[0063] in, represents the total prediction error loss function, Represents a continuous variable The predicted value at time t is Represents a continuous variable The true value of the continuous variable at time t; represents a latent continuous variable The predicted value at time t is represents a latent continuous variable The true value of the latent continuous variable at time t;

[0064] The calculation formula of the sparse constraint loss function is:

[0065] ;

[0066] in, represents the sparse constraint loss function, Indicates the pth variable to be predicted The corresponding hidden continuous variable predictor or the first layer weight vector of the continuous variable predictor, is a weight matrix with normalized shape, the weight matrix The element value in reflects the variable at lag time t Whether it can help predict the variables to be predicted The future value of

[0067] The calculation formula of the total loss function in the causal structure learning stage is:

[0068] ;

[0069] in, represents the total loss function of the causal structure learning phase, , and They are , and The weight coefficient of .

[0070] Furthermore, according to the weight matrix Granger causality judgment includes:

[0071] If the weight matrix If the element value in is zero at all t, then the variable Not a predictor variable Granger causality, then the Granger causality judgment result ;in, is an element in the overall causal matrix A;

[0072] If the weight matrix The element value in is non-zero at any t, which indicates that the variable is the variable to be predicted The Granger cause is the Granger causal judgment result.

[0073] Compared with the prior art, the present invention has the following beneficial effects: the present invention proposes a method for causal tracing of heterogeneous time series data faults in response to the problem of variable heterogeneity in causal analysis models; the present invention proposes a context-adaptive excitation Gaussian kernel embedding technology in response to the information asymmetry problem between discrete variables and continuous variables, which uses a multi-scale Gaussian kernel to aggregate discrete variable time series context information, refine the information granularity of the variables, restore their potential continuity, achieve the unification of variable distribution types, and significantly improve the accuracy of causal relationship extraction; the present invention designs prediction tasks, adopts self-supervised learning to restore potential continuity, and guides implicit continuity by introducing supervisory information of continuous variables. Recovery enhances the model's ability to recover the potential continuity of discrete variables; the present invention uses a reconstruction mechanism to ensure the reversibility of the recovery process, and adopts sparse regularization constraints to learn the causal structure, accurately screens the key causal relationships between variables, and constructs an overall causal graph for fault causal tracing, effectively ensuring the accuracy and interpretability of causal relationship extraction; different from traditional causal analysis methods that simply process heterogeneous time series data, the present invention mines the potential continuity of discrete variables through implicit continuity recovery, and performs causal discovery of heterogeneous variables in a unified continuous value space, which improves the accuracy of the model in identifying causal relationships while having better fault root cause tracing capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a flow chart of the method for tracing the cause and effect of heterogeneous time series data failures of the present invention;

[0075] Figure 2 The flowchart of the context-adaptive excitation Gaussian kernel embedding of the present invention is shown in FIG.

[0076] Figure 3 It is a process architecture diagram of the potential continuity learning stage and the causal structure learning stage of the present invention;

[0077] Figure 4 This is the root cause tracing result of a failure of a biomass fuel processing plant in Example 1 of the present invention; wherein, Figure 4 (a) is a cause-and-effect diagram of the root cause tracing results of a failure in a biomass fuel processing plant. Figure 4(b) is the causal adjacency matrix of the root cause tracing results of a biomass fuel processing plant;

[0078] Figure 5 is the visualization result of the potential continuity recovery of discrete variables in Example 1 of the present invention; wherein, Figure 5 (a) in is the original discrete variable, Figure 5 (b) in the figure is the restored latent continuous variable;

[0079] Figure 6 is the causal analysis result of a certain atmospheric dynamics system in Example 2 of the present invention; wherein, Figure 6 (a) in the equation is the real cause and effect of a certain atmospheric dynamics system. Figure 6 (b) is the estimated causal diagram of an atmospheric dynamics system;

[0080] Figure 7 is the visualization result of the potential continuity recovery of discrete variables in Example 2 of the present invention; wherein, Figure 7 (a) in is the original discrete variable, Figure 7 (b) in the figure is the real latent continuous variable and the restored latent continuous variable;

[0081] Figure 8 This is the result of tracing the root cause of a fault in a real chemical plant in Example 3 of the present invention; wherein, Figure 8 (a) is a cause-and-effect diagram of the root cause tracing result of a failure in a real chemical plant. Figure 8 (b) is the causal adjacency matrix of the root cause tracing results of a real chemical plant;

[0082] Fig. 9 is the visualization result of the potential continuity recovery of discrete variables in Example 3 of the present invention; wherein, Fig. 9 (a) in is the original discrete variable, Fig. 9 (b) in the figure is the true latent continuous variable and the restored latent continuous variable. DETAILED DESCRIPTION

[0083] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0084] The present invention proposes a new method for causal tracing of faults in heterogeneous time series data, which realizes causal discovery in a unified continuous space by converting discrete value variables into latent continuous variables with higher information granularity. Specifically, a context-adaptive excitation Gaussian kernel embedding is proposed. By adopting a multi-scale learnable excitation Gaussian kernel, combined with different bandwidths and weighted aggregation strategies, the discrete variable time series context information is aggregated, thereby refining the information granularity of the variable while restoring its potential continuity. Afterwards, it also includes two interdependent model training stages, namely the potential continuity recovery stage and the causal structure learning stage. In the potential continuity recovery stage, the guiding information of continuous variables is introduced by designing prediction tasks, and the Gaussian kernel embedding parameters are adjusted in a self-supervised manner to enhance the model's ability to restore potential continuity, and the reversibility of the recovery process is ensured by the reconstruction mechanism. In the causal structure learning stage, key causal relationships are screened through sparsity regularization constraints, and an overall causal graph is constructed to provide support for fault analysis.

[0085] See also Figure 1 The heterogeneous time series data fault causal tracing method of the present invention specifically includes the following steps:

[0086] (1) Heterogeneous time series fault data, including continuous variables and discrete variables, are collected, and the maximum and minimum value normalization method is used to process the continuous variables to construct a training set.

[0087] Specifically, a collection of heterogeneous time series fault data is collected ,in represents the pth variable, P represents the total number of variables, and T represents the length of the variable time series; the set Contains PN continuous variables and N discrete variables. The set of PN continuous variables is expressed as , the set of N discrete variables is expressed as ,in represents the i-th continuous variable, represents the jth discrete variable, , Represents the number of discrete states.

[0088] It should be understood that different types of discrete variables may have multiple different discrete states. For example, a valve switch signal may have two discrete states, open / closed, i.e., state 1 / 0, corresponding to .

[0089] In this embodiment, Set to 2, that is, all discrete variables in the mixed time series are set to binary discrete variables.

[0090] It should be understood that the maximum and minimum value normalization method is a commonly used data preprocessing method, which aims to convert data to a specific range, usually between [0,1], to eliminate the dimensional differences of variables.

[0091] (2) Use contextual adaptive Gaussian kernel embedding (CAGKE) to process the discrete variables in the training set and recover the corresponding latent continuous variables.

[0092] Furthermore, context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables in the training set to obtain the corresponding hidden continuous variables. Specifically, for the jth discrete variable time series , input it into the corresponding context-adaptive excitation Gaussian kernel embedding to obtain the discrete variable time series The corresponding latent continuous variable , expressed as:

[0093] (1)

[0094] in, Represents a discrete variable time series The corresponding latent continuous variable is represents the j-th discrete variable time series, ; Represents the context-adaptive excitation Gaussian kernel embedding corresponding to the j-th discrete variable; represents the softmax activation function; represents the weight coefficient of the multi-scale excitation Gaussian kernel, represents the weight coefficient of the excitation Gaussian kernel of the rth scale, and d is the number of multi-scale excitation Gaussian kernels; Trainable weight coefficient parameters for context-adaptive excitation Gaussian kernel embedding, Represents the trainable weight coefficient parameters of the excitation Gaussian kernel of the rth scale; the bandwidth of the Gaussian kernel representing the multiscale temporal patterns, represents the bandwidth of the Gaussian kernel of the rth scale; the bandwidths of the Gaussian kernels are arranged from large to small as , and The trainable bandwidth parameters of the context-adaptive excitation Gaussian kernel embedding are used to control the upper and lower limits of the bandwidth of the Gaussian kernel. Other bandwidth parameters are determined according to formula (2):

[0095] (2)

[0096] It should be understood that different discrete variables correspond to different context-adaptive excitation Gaussian kernel embeddings, and also correspond to different weight coefficient parameters and bandwidth parameters.

[0097] Furthermore, context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables to obtain the corresponding continuous value time series, and multi-scale adaptive weighted aggregation is performed to integrate the multi-scale time information of discrete variables to recover the corresponding hidden continuous variables, such as Figure 2 As shown, the specific process includes the following:

[0098] (2.1) Since the Gaussian kernel in the context-adaptive Gaussian kernel embedding has a smooth property, the use of context-adaptive Gaussian kernel embedding to process discrete variables can structurally characterize the temporal adjacency density, which is expressed as:

[0099] (3)

[0100] in, represents the Gaussian kernel, and is any time point in the discrete variable time series, represents the bandwidth of the Gaussian kernel. The Gaussian kernel will change with time and distance. The increase of Nearby time points show similar values.

[0101] (2.2) Use the Gaussian kernel determined in step (2.1) to process discrete variables. Specifically, for a given j-th discrete variable time series , consider the time point with non-zero value as the excitation time point, assuming is an excitation time point in the time series, and a bandwidth of The Gaussian kernel of is normalized and its expression is:

[0102] (4)

[0103] in, represents the Gaussian kernel, represents the normalization factor, Represents any point in time and incentive timing According to the above method, considering the discrete variable time series By superimposing the Gaussian kernel function corresponding to each excitation time point, the discrete variable time series can be obtained. A sequence of continuous values ​​at time point t, expressed as:

[0104] (5)

[0105] in, Represents a discrete variable time series A sequence of consecutive values ​​at time point t, Represents the jth discrete variable time series The value at the tth time point in , .

[0106] It should be understood that the excitation time point refers to the time point at which the discrete variable takes a non-zero value. T in formula (3) represents the length of the variable time series, and the excitation time point It refers to the discrete variable time series The time point with a non-zero value. If a time point is not a stimulus time point, then its corresponding If the value is zero, no , only the incentive time point will produce .

[0107] (2.3) Based on the discrete variable time series obtained in step (2.2) At the continuous value sequence of time point t, by aggregating the excitation Gaussian kernel functions of various scales according to their corresponding weights, the multi-scale time information of discrete variables is integrated to realize the restoration of implicit continuity and obtain the implicit continuous variable corresponding to time point t , expressed as:

[0108] (6)

[0109] in, is the noise term that follows the normal distribution, that is .

[0110] (2.4) Based on the discrete variable time series obtained in step (2.3) Latent continuous variable at time point t Obtaining discrete variable time series The corresponding latent continuous variable , expressed as .

[0111] In summary, the process from step (2.1) to step (2.4) is the specific implementation process of context-adaptive excitation Gaussian kernel embedding.

[0112] (3) Build a corresponding continuous variable predictor and discrete variable decoder for each continuous variable and discrete variable respectively, and use the training set constructed in step (1) and the latent continuous variables recovered in step (2) to train the continuous variable predictor and discrete variable decoder. Take minimizing the total loss function of the latent continuous learning stage as the optimization goal, optimize the parameters of the continuous variable predictor, discrete variable decoder and context-adaptive excitation Gaussian kernel embedding, and obtain the trained discrete variable decoder and context-adaptive excitation Gaussian kernel embedding.

[0113] In this embodiment, the continuous variable predictor and the discrete variable decoder are both implemented using a multi-layer fully connected neural network, and prediction and decoding are achieved through multi-layer linear transformation and nonlinear activation functions, wherein the network parameters of each layer are randomly initialized, and the network parameters of the continuous variable predictor and the discrete variable decoder are different.

[0114] In this embodiment, the continuous variable predictor and the discrete variable decoder are trained using the training set constructed in step (1) and the latent continuous variables recovered in step (2). The total loss function of the latent continuous learning stage is minimized as the optimization goal, and the parameters of the continuous variable predictor, the discrete variable decoder, and the context-adaptive excitation Gaussian kernel embedding are optimized to obtain the trained discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding, such as Figure 3 As shown, the specific process includes the following:

[0115] (3.1) The continuous variables in the training set The historical values ​​and the recovered hidden continuous variables The historical values ​​of are input into the continuous variable predictor to obtain the continuous variable Predicted value at a future time , expressed as:

[0116] (7)

[0117] in, Represents a continuous variable The predicted value at time t; represents the continuous variable predictor corresponding to the i-th continuous variable, , represents the dimension of the input and output of the continuous variable predictor, which means that The input vector of dimension is mapped to an output The predicted value of the dimension; is the specified history window length; Represents a continuous variable from Time has come The historical value of the moment, represents a latent continuous variable from Time has come The historical value at each moment can be obtained using continuous variables and latent continuous variables Forecast continuous variables using historical values ​​of a specified historical window length The data at time t.

[0118] (3.2) Calculation of continuous variables The mean square error (MSE) between the predicted value at the future moment and its corresponding true continuous variable value is used as the prediction error loss function of the continuous variable predictor.

[0119] Furthermore, the prediction error loss function of the continuous variable predictor is calculated as:

[0120] (8)

[0121] in, represents the prediction error loss function for continuous variable predictors, represents the mean square error, Represents a continuous variable The predicted value at time t is Represents a continuous variable The true continuous variable value at time t is determined by the corresponding true continuous variable in the training set. Represents the vector two-norm.

[0122] (3.3) will recover the obtained latent continuous variable The historical values ​​of are input into the discrete variable decoder to reconstruct the original discrete variable , get the reconstructed discrete variables , expressed as:

[0123] (9)

[0124] in, represents the reconstructed discrete variable at time t, Represents the discrete variable decoder corresponding to the j-th discrete variable.

[0125] It should be understood that formula (7) uses the data at historical moments (i.e., moment t-1 and before) to predict the data at future moments (i.e., moment t); whereas formula (8) is decoding reconstruction, so the data at moment t can be used.

[0126] (3.4) Calculate the cross entropy loss (CE) between the reconstructed discrete variable and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder.

[0127] Furthermore, the calculation formula of the reconstruction error loss function of the discrete variable decoder is:

[0128] (10)

[0129] in, represents the reconstruction error loss function of the discrete variable decoder, represents the cross entropy loss, represents the reconstructed discrete variable at time t, Represents the original discrete variable The true discrete variable at time t is determined by the corresponding true discrete variable in the training set. Represents the reconstructed discrete variable The value corresponding to the cth discrete state of Represents a real discrete variable The value corresponding to the cth discrete state of Represents the natural logarithm.

[0130] (3.5) Calculate the total loss function of the latent continuous learning stage based on the prediction error loss function of the continuous variable predictor and the reconstruction error loss function of the discrete variable decoder; optimize the parameters of the continuous variable predictor, discrete variable decoder and context-adaptive excitation Gaussian kernel embedding with the optimization goal of minimizing the total loss function of the latent continuous learning stage to obtain the trained discrete variable decoder and context-adaptive excitation Gaussian kernel embedding.

[0131] Furthermore, the total loss function in the potential continuity learning phase is calculated as:

[0132] (11)

[0133] in, represents the total loss function of the potential continuous learning stage, for The weight coefficient of for The weight coefficient of .

[0134] (4) Using the discrete variable decoder and context-adaptive excitation Gaussian kernel embedding trained in step (3), use the data in the training set to perform causal structure learning, and take minimizing the total loss function of the causal structure learning stage as the optimization goal, optimize the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor to obtain the trained latent continuous variable predictor, the continuous variable predictor and the context-adaptive excitation Gaussian kernel embedding; and obtain the overall causal graph based on the parameters of the trained latent continuous variable predictor and the continuous variable predictor, as shown in the following example. Figure 3 shown.

[0135] (4.1) Fix the network parameters of the trained discrete variable decoder; fix the bandwidth parameters of the trained context-adaptive excitation Gaussian kernel embedding, and keep its weight coefficient parameters trainable; build a hidden continuous variable predictor, and randomly initialize its network parameters; use the continuous variable predictor built in step (3), and randomly initialize its network parameters. The hidden continuous variable predictor is implemented using a multi-layer fully connected neural network, and prediction is achieved through multi-layer linear transformation and nonlinear activation function.

[0136] It should be understood that although both the latent continuous variable predictor and the continuous variable predictor are implemented using multi-layer fully connected neural networks, and prediction is achieved through multi-layer linear transformations and nonlinear activation functions, the specific structures of the latent continuous variable predictor and the continuous variable predictor are different, for example, the number of linear transformation layers is different; the number of latent continuous variable predictors and the continuous variable predictor is also different, there are N latent continuous variable predictors, corresponding to N latent continuous variables, and there are PN continuous variable predictors, corresponding to PN continuous variables. The essence of the continuous variable predictor is to predict the future state of the continuous variable, while the essence of the latent continuous variable predictor is to predict the future state of the latent continuous variable generated by the discrete variable. The latent continuous variables in the input of the latent continuous variable predictor and the continuous variable predictor will change with the training process of the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, and considering the fairness of the latent continuous variable predictor and the continuous variable predictor in causal discovery, the network parameters of both must be randomly initialized for training.

[0137] (4.2) Input the discrete variables in the training set into the context-adaptive excitation Gaussian kernel embedding trained in step (3) to obtain the corresponding hidden continuous variables .

[0138] It should be understood that by fixing the bandwidth parameters of the trained context-adaptive excitation Gaussian kernel embedding, the trainability of its weight coefficient parameters is maintained, so the hidden continuous variable The context-adaptive excitation Gaussian kernel embedding changes continuously during training.

[0139] (4.3) The latent continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the continuous variable predictor constructed in step (3) to obtain the continuous variable Predicted value at a future time ; The latent continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the latent continuous variable predictor constructed in step (4.1) to obtain the latent continuous variable Predicted value at a future time ; To complete the prediction task under the Granger Causality (GC) paradigm, they are respectively denoted as:

[0140] (12)

[0141] (13)

[0142] in, represents a latent continuous variable The predicted value at time t is represents the latent continuous variable predictor corresponding to the jth latent continuous variable, , represents the dimension of the input and output of the latent continuous variable predictor, which means that The input vector of dimension is mapped to an output The predicted value of the dimension.

[0143] (4.4) Calculation of continuous variables The mean square error between the predicted value at the future moment and its corresponding true continuous variable value and the hidden continuous variable The mean square error between the predicted value at the future moment and its corresponding true latent continuous variable value, and the sum of the two mean square errors is used as the total prediction error loss function.

[0144] Furthermore, the calculation formula of the total prediction error loss function is:

[0145] (14)

[0146] in, represents the total prediction error loss function, Represents a continuous variable The predicted value at time t is Represents a continuous variable The true continuous variable value at time t is determined by the corresponding true continuous variable in the training set; represents a latent continuous variable The predicted value at time t is represents a latent continuous variable The true latent continuous variable value at time t is the true latent continuous variable obtained by step (4.2) Sure.

[0147] (4.5) The latent continuous variable obtained in step (4.2) The historical values ​​of are input into the discrete variable decoder trained in step (3) to reconstruct the original discrete variable , get the reconstructed discrete variables , whose expression is shown in formula (9).

[0148] (4.6) Calculate the cross entropy loss between the reconstructed discrete variable obtained in step (4.5) and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder , and its calculation formula is shown in formula (10).

[0149] (4.7) In order to obtain sparse and important causal structures, a sparsity-inducing constraint, namely, the hierarchical group least absolute shrinkage and selection operator (Lasso), is imposed on each latent continuous variable predictor and the first layer of each continuous variable predictor to obtain the corresponding sparse constrained loss function.

[0150] Furthermore, the calculation formula of the sparse constraint loss function is:

[0151] (15)

[0152] in, represents the sparse constraint loss function, Indicates the pth variable to be predicted The corresponding hidden continuous variable predictor or the first layer weight vector of the continuous variable predictor, is a weight matrix with normalized shape, the weight matrix The element value in reflects the variable at lag time t Whether it can help predict the variables to be predicted The future value of .

[0153] (4.8) Calculate the total loss function of the causal structure learning stage according to the total prediction error loss function obtained in step (4.4), the reconstruction error loss function of the discrete variable decoder obtained in step (4.6), and the sparse constraint loss function obtained in step (4.7); Taking minimizing the total loss function of the causal structure learning stage as the optimization goal, the proximal gradient descent (PGD) method is used to optimize the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor, so as to obtain the optimal weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor, and obtain the final trained context-adaptive excitation Gaussian kernel embedding, latent continuous variable predictor and continuous variable predictor.

[0154] It should be understood that the use of the proximal gradient descent method to adjust parameters can ensure the sparsity of the optimization results.

[0155] Furthermore, the calculation formula of the total loss function in the causal structure learning stage is:

[0156] (16)

[0157] in, represents the total loss function of the causal structure learning phase, , and They are , and The weight coefficient of .

[0158] (4.9) Based on the optimal implicit continuous variable predictor and continuous variable predictor network parameters obtained in step (4.8), determine the variable to be predicted The corresponding latent continuous variable predictor or continuous variable predictor weight matrix ; According to the weight matrix Perform Granger causality judgment to obtain the variables to be predicted The corresponding Granger causality judgment results; the overall causal diagram is obtained based on the Granger causality judgment results of all variables to be predicted, which can be used for process failure analysis.

[0159] It should be understood that since the total loss function of the causal structure learning phase is minimized, the variables to be predicted are included. The corresponding sparse constraint loss function, so the optimal network parameters of the hidden continuous variable predictor and the continuous variable predictor obtained after the training in step (4.8) are the variables to be predicted The corresponding hidden continuous variable predictor or continuous variable predictor can obtain the corresponding weight matrix according to its network parameters .

[0160] Furthermore, according to the weight matrix Make Granger causal judgment, including: if the weight matrix If the element value in is zero at all t, then the variable Not a predictor variable Granger causality, then the Granger causality judgment result ; If the weight matrix The element value in is non-zero at any t, which indicates that the variable is the variable to be predicted Granger causality, then the Granger causality judgment result .in, are the elements in the overall causal matrix A, and the overall causal diagram can be obtained based on these elements.

[0161] (5) In the application stage, the context-adaptive excitation Gaussian kernel embedding trained in step (4) is used to restore the discrete variables in the online heterogeneous time series fault data into latent continuous variables, and the variables are input together with the continuous variables in the online heterogeneous time series fault data into the latent continuous variable predictor and continuous variable predictor trained in step (4) to obtain the overall causal graph for process fault analysis.

[0162] It should be noted that when using the trained latent continuous variable predictor and continuous variable predictor to obtain the overall causal graph, the weight matrix at different lag times t The element values ​​in change as the specific values ​​of the online heterogeneous timing fault data change.

[0163] The embodiment of the present invention lists three typical examples, namely: root cause tracing of a biomass fuel processing process, causal analysis of an atmospheric dynamics system, and root cause tracing of a chemical process. The heterogeneous time series data fault causal tracing method described in the present invention can cover any causal analysis task with heterogeneous variables in multiple fields including but not limited to energy, chemical industry, metallurgy, meteorology, biomedicine, etc.

[0164] Example 1: Cause and effect tracing of heterogeneous time series data failures in biomass fuel processing

[0165] This embodiment takes the causal tracing task of heterogeneous time series fault data measured in the process of a biomass fuel processing plant as an example to illustrate and verify the effectiveness of the heterogeneous time series data fault causal tracing method described in the present invention, but is not limited to this. The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments in the energy field.

[0166] In the energy field, the processing and manufacturing process of biomass fuels (such as wood, straw, etc.) is one of the key links that determine the quality and combustion characteristics of fuel particles, and has an important impact on the final power generation efficiency and environmental emissions. The process mainly includes three parts: raw material processing, particle forming and combustion regulation. This section focuses on the particle forming link, that is, the process of processing biomass fuel raw materials and converting them into particles suitable for combustion. The process involves two core equipment: a heating humidifier and a drum dryer. Among them, the heating humidifier provides suitable environmental conditions for the molding process of biomass fuel through precise temperature and humidity regulation; the drum dryer uses saturated steam to heat the raw materials and reduce their moisture content from 20% to 12%, ensuring the combustibility of the fuel and the high efficiency of combustion. The entire process contains a total of 23 measurement variables with a sampling interval of 10 seconds.

[0167] In order to evaluate the performance of the method of the present invention in fault root cause diagnosis, this example selects an actual industrial fault case in biomass fuel processing technology. In this case, the increase in the hot air speed set point of the drum dryer causes the hot air temperature to change, and through its coupling relationship with other variables, it further causes the abnormality of related variables. Combining the fault mechanism analysis and expert knowledge, the 23 measured variables are analyzed and the main fault variables are confirmed to be , and its specific variable information is shown in Table 1.

[0168] Table 1: Key fault variables and their descriptions for biomass fuel processing failures

[0169]

[0170] The specific mechanism of the failure process is described as follows: When the hot air speed of the drum dryer (i.e., variable 18) increases, the speed of the fuel raw material passing through the dryer is accelerated, resulting in a decrease in the evaporation of the water in the raw material, which fails to dry to the expected moisture content, thereby reducing the water removal quality of the dryer (i.e., variable 19). Correspondingly, more water remains in the raw material, resulting in the moisture content of the dried raw material (i.e., variable 20) being higher than the normal value. In addition, due to the shortened residence time of the raw material in the dryer, the drying temperature of the dryer (i.e., variable 21) is also lower than the reference value. Subsequently, the raw material output from the drum dryer is transported to the shake cooler. In the cooler, the cooling temperature (i.e., variable 23) is higher than the normal value, while the moisture content after cooling (i.e., variable 22) is lower than the normal level. It should be noted that since the fault interference occurs in the dryer equipment, which is located after the heating and humidifying machine processing unit, the fault will not affect the heating and humidifying machine processing unit. In this embodiment, 400 samples after the fault occurs are used for causal analysis and root cause tracing.

[0171] In this embodiment, the root cause tracing result of a biomass fuel processing plant is finally obtained as follows: Figure 4 As shown, Figure 4 The causal analysis results based on the method of the present invention are shown in FIG. Figure 4 As shown in (a) in the figure, the nodes represent the main variables and the arrows represent the causal relationship between the variables. For example, right and has a direct impact, and and The causal effects on other variables are also intuitively presented through arrows. In addition, the direction of the arrows reflects the causal transmission path, which is gradually transmitted from the root dependent variable to the midstream and downstream variables. Figure 4 As shown in (b) in the figure, the causal strength between variables is further quantified in the form of a heat map, where the rows of the matrix represent the affected variables, the columns represent the influencing variables, and the dark areas indicate stronger causal relationships. right , The causal strength is high, which is consistent with the arrow path in the causal graph. Through the analysis of the above results, the method can clearly reveal the propagation path of the fault and clearly identify the is the root dependent variable, and is the midstream variable, , , It intuitively, quantitatively and accurately presents the cause-effect chain of failures in the fuel processing process for downstream variables, providing strong support for accurate location and diagnosis of root causes. and The corresponding implicit continuous recovery result is as follows Figure 5 As shown, Figure 5 The visualization results of the potential continuity recovery of two discrete variables are shown, where the original discrete variables are Figure 5 As shown in (a) in FIG. 1 , the hidden continuous variable recovered by the method of the present invention is as follows: Figure 5 As shown in (b) in . Figure 5 The visualization results shown demonstrate that, with the support of potential continuity recovery, the context-adaptive excitation Gaussian kernel embedding described in the present invention can focus on the intrinsic continuity of latent continuous variables implied in discrete variables and their fine-grained temporal dynamic characteristics, effectively promoting the learning and discovery of causal structures.

[0172] In addition, in order to verify the universality and effectiveness of the method described in the present invention in causal analysis tasks in other fields, the following two brief embodiments are given.

[0173] Example 2: Causal tracing analysis of heterogeneous time series data of atmospheric dynamics system

[0174] In a certain atmospheric dynamics system, such as the Lorenz-96 system in the field of meteorology, this is a common simplified atmospheric dynamics model that aims to study the predictability of atmospheric circulation. As a typical chaotic dynamic system, this model has been widely used in modeling and prediction tasks in climate science. The system consists of multiple slow variables, representing large-scale meteorological factors, such as surface air temperature; each slow variable is coupled with multiple fast variables, which represent small-scale meteorological factors, such as the path of liquid water in clouds. The causal relationship in the system is mainly reflected in the mutual feedback between slow variables and fast variables, which in turn affects the dynamic behavior of the system, such as the conservation and dissipation of energy. Studying this causal relationship is of great significance for a deeper understanding of atmospheric dynamics, improving the prediction accuracy of climate models, and improving data assimilation technology. It helps to better understand the complexity of atmospheric circulation and provide theoretical support for exploring methods to improve climate prediction.

[0175] In this embodiment, 1000 samples generated by the system are used for causal analysis, and the total number of slow variables and fast variables is 20, including 4 slow variables and 16 fast variables. Eight variables are randomly selected and the variable average value is regarded as the threshold for binary discretization to simulate heterogeneous scenarios, and the method of the present invention is executed to obtain a causal graph. Figure 6 The real causal graph of the Lorenz-96 system and the causal graph estimated by the method of the present invention are shown, wherein the real causal graph is as follows: Figure 6 As shown in (a) of FIG. 1 , the causal graph estimated by the method of the present invention is as follows: Figure 6 As shown in (b) in the figure. Figure 6 The visualization results shown demonstrate that the method of the present invention has good causal discovery capabilities, can accurately estimate and reveal the true causal structure based on the actual situation, and accurately capture the interactions between heterogeneous variables. Figure 7 The visualization results of the potential continuity recovery of three discrete variables are shown, where the original discrete variables are Figure 7 As shown in (a) of FIG. 1 , the real hidden continuous variables and the restored hidden continuous variables of the method of the present invention are as follows: Figure 7 As shown in (b), Figure 7 The visualization results shown indicate that the context-adaptive excitation Gaussian kernel embedding described in the present invention has good implicit continuity recovery capability and can more accurately reproduce the underlying potential continuity of discrete variables, proving the effectiveness and reliability of the implicit continuity recovery idea.

[0176] Example 3: Cause-effect tracing analysis of failures in heterogeneous time series data of chemical processes

[0177] In the chemical industry, the Tennessee Eastman Process (TEP) is a realistic process for evaluating process monitoring and fault diagnosis methods. The process includes five main units: reactor, condenser, gas-liquid separator, circulating compressor and stripper, and includes eight components A, B, C, D, E, F, G and H. The process has a total of 41 process variables (PV1-41) and 12 operating variables (MV1-11). The data set provided by the real chemical plant includes a total of 21 fault cases with various pre-specified interferences, and each case contains 960 samples.

[0178] In this example, the first typical fault IDV (1) in TEP is selected to verify the effectiveness of the method described in the present invention. In this case, the root cause of the fault is the step change of the feed ratio of stream 4, which leads to a decrease in component A in the recovery stream. A component controller adjusts the corresponding valve opening (MV3) Increase the flow rate of feed A in stream 1 (PV1). This operation causes the reactor level to increase, changing the material residence time, which in turn affects other process variables. Specifically, the flow rate in stream 4 (PV4) is affected by the liquid level controller and changes the stripping tower pressure (PV16) and stripping tower temperature (PV18), which in turn affects the reactor pressure (PV7) and product separator pressure (PV13). Due to the control of the stripping tower temperature, the system adjusts the stripping tower gas valve (MV9) to control the steam flow of the stripping tower (PV19). Combining the above abnormal mechanism and fault isolation, it is confirmed that the fault variable is , and its specific variable information is shown in Table 2. Three continuous variables, reactor pressure, product separator pressure, and stripping tower pressure, are selected, and the average value is regarded as the threshold for binary discretization to simulate heterogeneous scenarios.

[0179] Table 2: Key fault variables and their descriptions for TEP process faults

[0180]

[0181] Executing the method of the present invention can obtain Figure 8 The root cause tracing result of a fault in a real chemical plant is shown in the figure. Figure 8 As shown in (a) in the figure, the causal adjacency matrix is Figure 8 As shown in (b) of FIG. 1 , it can be seen from the figure that the causal graph and adjacency matrix obtained by executing the method described in the present invention successfully reconstruct the path of fault propagation. The key causal relationships, such as and feedback control are correctly mined. In addition, the two root dependent variables and The disease was also clearly diagnosed, demonstrating the accuracy of the method of the present invention in root cause tracing. Fig. 9 The visualization results of the potential continuity recovery of three discrete variables in the embodiment are shown, where the original discrete variables are as follows: Fig. 9 As shown in (a) of FIG. 1 , the real hidden continuous variables and the restored hidden continuous variables of the method of the present invention are as follows: Fig. 9 As shown in (b), Fig. 9 The visualization results shown indicate that, in actual chemical industry scenarios, the context-adaptive excitation Gaussian kernel embedding described in the present invention still maintains good implicit continuous recovery capability and has good robustness and stability.

[0182] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for tracing the cause and effect of heterogeneous time series data failures, characterized in that: The following steps are involved: (1) Collecting heterogeneous time series fault data, specifically including continuous variables and discrete variables, and using the maximum and minimum value normalization method to process the continuous variables to construct a training set; wherein the heterogeneous time series fault data includes feed flow, flow flow, reactor pressure, product separator pressure, stripping tower pressure, stripping tower temperature, stripping tower steam flow, flow valve opening and stripping steam valve opening; the feed flow, flow flow, stripping tower temperature, stripping tower steam flow, flow valve opening and stripping steam valve opening are continuous variables; the reactor pressure, product separator pressure and stripping tower pressure are discrete variables; (2) using context-adaptive excitation Gaussian kernel embedding to process discrete variables in the training set and recover corresponding latent continuous variables; the context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables to obtain corresponding continuous value time series, and multi-scale adaptive weighted aggregation is performed to integrate multi-scale time information of discrete variables to recover corresponding latent continuous variables, specifically including the following steps: (2.1) The Gaussian kernel in context-adaptive excitation Gaussian kernel embedding has a smooth property and can structurally characterize the temporal adjacency density. Its expression is: in, represents the Gaussian kernel, t and t′ are any time points in the discrete variable time series, and σ represents the bandwidth of the Gaussian kernel; (2.2) Use the Gaussian kernel determined in step (2.1) to process discrete variables. Specifically, for a given j-th discrete variable time series The time point with non-zero value is regarded as the excitation time point. Assume that t′ is an excitation time point in the time series, and a Gaussian kernel with bandwidth σ is superimposed around it and normalized. The expression is: Where f(Δt;σ) represents the Gaussian kernel, represents the normalization factor, Δt 2 =(tt′) 2 represents the distance between any time point t and the excitation time point t′; by superimposing the discrete variable time series in the manner shown in the above Gaussian kernel expression Gaussian kernel function corresponding to all excitation time points on the image to obtain the discrete variable time series A sequence of continuous values ​​at time point t, expressed as: in, Represents a discrete variable time series A sequence of consecutive values ​​at time point t, Represents the jth discrete variable time series The value at the tth time point in ; (2.3) Based on the discrete variable time series obtained in step (2.2) At the continuous value sequence of time point t, by aggregating the excitation Gaussian kernel functions of various scales according to their corresponding weights, the multi-scale time information of discrete variables is integrated to realize the restoration of implicit continuity and obtain the implicit continuous variable corresponding to time point t It is expressed as: Among them, ε t is the noise term that follows the normal distribution, that is (2.4) Based on the discrete variable time series obtained in step (2.3) Latent continuous variable at time point t Obtaining discrete variable time series The corresponding latent continuous variable Expressed as (3) For each continuous variable and discrete variable, a corresponding continuous variable predictor and discrete variable decoder are constructed, and the training set constructed in step (1) and the hidden continuous variable recovered in step (2) are used for training. The total loss function of the potential continuous learning phase is minimized as the optimization goal, and the parameters of the continuous variable predictor, discrete variable decoder, and context-adaptive excitation Gaussian kernel embedding are optimized to obtain the trained discrete variable decoder and context-adaptive excitation Gaussian kernel embedding; (4) using the discrete variable decoder and context-adaptive excitation Gaussian kernel embedding trained in step (3), using the data in the training set to perform causal structure learning, taking minimizing the total loss function of the causal structure learning stage as the optimization goal, optimizing the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the latent continuous variable predictor and the continuous variable predictor, so as to obtain the trained latent continuous variable predictor, the continuous variable predictor and the context-adaptive excitation Gaussian kernel embedding; and obtaining the overall causal graph according to the parameters of the trained latent continuous variable predictor and the continuous variable predictor; (5) The context-adaptive excitation Gaussian kernel embedding trained in step (4) is used to restore the discrete variables in the online heterogeneous time series fault data into latent continuous variables, and the variables are input together with the continuous variables in the online heterogeneous time series fault data into the latent continuous variable predictor and the continuous variable predictor trained in step (4) to obtain the overall causal graph for process fault analysis.

2. The method for tracing the cause and effect of heterogeneous time series data failures according to claim 1 is characterized in that: The collecting of heterogeneous time series fault data specifically includes: Collecting a collection of heterogeneous time series fault data where x p represents the pth variable, P represents the total number of variables, and T represents the length of the variable time series; the set x contains PN continuous variables and N discrete variables. The set of PN continuous variables is expressed as The set of N discrete variables is expressed as in represents the i-th continuous variable, represents the jth discrete variable, N DVs Represents the number of discrete states.

3. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 1 is characterized in that: The context-adaptive excitation Gaussian kernel embedding is used to process the time series of discrete variables in the training set to obtain the corresponding hidden continuous variables. Specifically, for the jth discrete variable time series Input it into the corresponding context-adaptive excitation Gaussian kernel embedding to obtain the discrete variable time series The corresponding latent continuous variable It is expressed as: in, Represents a discrete variable time series The corresponding latent continuous variable is represents the j-th discrete variable time series, j = 1, 2, ..., N; CAGKE j () represents the context-adaptive excitation Gaussian kernel embedding corresponding to the j-th discrete variable; softmax() represents the softmax activation function; represents the weight coefficient of the multi-scale excitation Gaussian kernel, represents the weight coefficient of the excitation Gaussian kernel of the rth scale, d is the number of multi-scale excitation Gaussian kernels; (α1,...,α r ,...,α d ) is the trainable weight coefficient parameter of context-adaptive excitation Gaussian kernel embedding, α r Represents the trainable weight coefficient parameters of the excitation Gaussian kernel of the rth scale; {σ1,σ2,...,σ r ,...,σ d } represents the bandwidth of the Gaussian kernel of the multi-scale temporal pattern, σ r represents the bandwidth of the Gaussian kernel of the rth scale; σ min and σ max The trainable bandwidth parameters of the context-adaptive excitation Gaussian kernel embedding are used to control the upper and lower limits of the bandwidth of the Gaussian kernel. The other bandwidths are determined according to the following formula:

4. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 1 is characterized in that: The continuous variable predictor and discrete variable decoder are both implemented using a multi-layer fully connected neural network, and prediction and decoding are achieved through multi-layer linear transformation and nonlinear activation functions, wherein the network parameters of each layer are randomly initialized, and the network parameters of the continuous variable predictor and the discrete variable decoder are different.

5. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 1 is characterized in that: The continuous variable predictor and the discrete variable decoder are trained using the training set constructed in step (1) and the hidden continuous variables recovered in step (2), with minimizing the total loss function of the potential continuity learning stage as the optimization goal, and optimizing the parameters of the continuous variable predictor, the discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding to obtain the trained discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding, specifically comprising the following steps: (3.1) The continuous variables in the training set The historical values ​​and the recovered hidden continuous variables The historical values ​​of are input into the continuous variable predictor to obtain the continuous variable Predicted value at a future time It is expressed as: in, Represents a continuous variable The predicted value at time t; represents the continuous variable predictor corresponding to the i-th continuous variable; τ is the specified historical window length; Represents a continuous variable The historical value from time t-τ to time t-1, represents a latent continuous variable Historical values ​​from time t-τ to time t-1; (3.2) Calculation of continuous variables The mean square error between the predicted value at the future moment and its corresponding true continuous variable value is used as the prediction error loss function of the continuous variable predictor; (3.3) The hidden continuous variables obtained by recovering The historical values ​​of are input into the discrete variable decoder to reconstruct the original discrete variable Get the reconstructed discrete variables It is expressed as: in, Decoder represents the reconstructed discrete variable at time t. j () represents the discrete variable decoder corresponding to the jth discrete variable; (3.4) Calculate the cross entropy loss between the reconstructed discrete variable and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder; (3.5) The total loss function of the latent continuous learning stage is calculated based on the prediction error loss function of the continuous variable predictor and the reconstruction error loss function of the discrete variable decoder; taking minimizing the total loss function of the latent continuous learning stage as the optimization goal, the parameters of the continuous variable predictor, the discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding are optimized to obtain the trained discrete variable decoder and the context-adaptive excitation Gaussian kernel embedding.

6. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 5 is characterized in that: The calculation formula of the prediction error loss function of the continuous variable predictor is: in, represents the prediction error loss function of the continuous variable predictor, MSE() represents the mean square error, Represents a continuous variable The predicted value at time t is Represents a continuous variable The real continuous variable value at time t, ||·||2 represents the vector two norm; The calculation formula of the reconstruction error loss function of the discrete variable decoder is: in, represents the reconstruction error loss function of the discrete variable decoder, CE() represents the cross entropy loss, represents the reconstructed discrete variable at time t, Represents the original discrete variable The real discrete variable at time t, Represents the reconstructed discrete variable The value corresponding to the cth discrete state of Represents a real discrete variable The value corresponding to the c-th discrete state of , log(·) represents the natural logarithm; The calculation formula of the total loss function in the potential continuity learning stage is: in, represents the total loss function of the potential continuous learning stage, λ1 and λ2 are and The weight coefficient of .

7. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 1 is characterized in that: The step (4) comprises the following sub-steps: (4.1) Fix the network parameters of the trained discrete variable decoder; fix the bandwidth parameters of the trained context-adaptive excitation Gaussian kernel embedding, and keep the trainability of its weight coefficient parameters; build a hidden continuous variable predictor, and randomly initialize its network parameters; use the continuous variable predictor built in step (3), and randomly initialize its network parameters; wherein the hidden continuous variable predictor is implemented using a multi-layer fully connected neural network, and prediction is achieved through multi-layer linear transformation and nonlinear activation function; (4.2) Input the discrete variables in the training set into the context-adaptive excitation Gaussian kernel embedding trained in step (3) to obtain the corresponding hidden continuous variables (4.3) The implicit continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the continuous variable predictor constructed in step (3) to obtain the continuous variable Predicted value at a future time The latent continuous variable obtained in step (4.2) The historical values ​​of and continuous variables in the training set The historical values ​​of are input into the latent continuous variable predictor constructed in step (4.1) to obtain the latent continuous variable Predicted value at a future time To complete the prediction task under the Granger causality paradigm; (4.4) Calculation of continuous variables The mean square error between the predicted value at the future moment and its corresponding true continuous variable value and the hidden continuous variable The mean square error between the predicted value at the future moment and its corresponding true latent continuous variable value, and the sum of the two mean square errors is used as the total prediction error loss function; (4.5) The implicit continuous variable obtained in step (4.2) The historical values ​​of are input into the discrete variable decoder trained in step (3) to reconstruct the original discrete variable Get the reconstructed discrete variables (4.6) Calculate the cross entropy loss between the reconstructed discrete variable obtained in step (4.5) and its corresponding true discrete variable as the reconstruction error loss function of the discrete variable decoder (4.7) Apply a sparsity-inducing constraint, i.e., a hierarchical group minimum absolute shrinkage and selection operator, to each latent continuous variable predictor and the first layer of each continuous variable predictor to obtain the corresponding sparsity-constrained loss function; (4.8) Calculate the total loss function of the causal structure learning stage according to the total prediction error loss function obtained in step (4.4), the reconstruction error loss function of the discrete variable decoder obtained in step (4.6), and the sparse constraint loss function obtained in step (4.7); Taking minimizing the total loss function of the causal structure learning stage as the optimization goal, the proximal gradient descent method is used to optimize the weight coefficient parameters of the context-adaptive excitation Gaussian kernel embedding, the network parameters of the hidden continuous variable predictor and the continuous variable predictor, so as to obtain the optimal context-adaptive excitation Gaussian kernel embedding weight coefficient parameters, the network parameters of the hidden continuous variable predictor and the continuous variable predictor, and obtain the finally trained context-adaptive excitation Gaussian kernel embedding, the hidden continuous variable predictor and the continuous variable predictor; (4.9) Based on the optimal implicit continuous variable predictor and continuous variable predictor network parameters obtained in step (4.8), determine the variable to be predicted x p The corresponding hidden continuous variable predictor or continuous variable predictor weight matrix W q,t ; According to the weight matrix W q,t Perform Granger causality judgment to obtain the variable x to be predicted p The corresponding Granger causality judgment results; obtain the overall causal diagram based on the Granger causality judgment results of all variables to be predicted.

8. The method for tracing the cause and effect of heterogeneous time series data failures according to claim 7 is characterized in that: The calculation formula of the total prediction error loss function is: in, represents the total prediction error loss function, Represents a continuous variable The predicted value at time t is Represents a continuous variable The true value of the continuous variable at time t; represents a latent continuous variable The predicted value at time t is represents a latent continuous variable The true value of the latent continuous variable at time t; The calculation formula of the sparse constraint loss function is: in, represents the sparse constraint loss function, represents the pth variable x to be predicted p The corresponding hidden continuous variable predictor or the first layer weight vector of the continuous variable predictor, is a weight matrix with normalized shape, the weight matrix W q,t The element value in reflects the variable x at lag time t q Can it help predict the variable x? p The future value of The calculation formula of the total loss function in the causal structure learning stage is: in, represents the total loss function of the causal structure learning stage, η1, η2 and η3 are and The weight coefficient of .

9. The method for tracing the cause and effect of heterogeneous time series data faults according to claim 7, characterized in that: According to the weight matrix W q,t Conduct Granger causality judgment, including: If the weight matrix W q,t If the element value in is zero at all t, then the variable x q Not the predicted variable x p Granger causality, then the Granger causality judgment result A (p,q) =0; where A (p,q) is an element in the overall causal matrix A; If the weight matrix W q,t The element value in is non-zero at any t, which means that the variable x q is the variable to be predicted x p Granger causality, then the Granger causality judgment result

Citation Information

Patent Citations

  • Causal aggregation and grouping fault root cause tracing method for typical equipment of gas turbine

    CN116663184A