Discovery method and device for nonlinear causal relationship of time series data
By applying kernel independent component analysis and variational autoencoder methods in time series data, the problem of difficulty in discovering nonlinear causal relationships in the prior art is solved, and accurate identification of immediate and lagging causal relationships and effective update of causal graphs are achieved.
Patent Information
- Application Number
- CN202510534445.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively discover nonlinear causal relationships from time series data, especially in identifying immediate and lagging causal relationships.
A method of discovering nonlinear causal relationships in time series data is proposed. By constructing a training data set and using kernel independent component analysis and variational autoencoder techniques, the real-time causal relationship and lagging causal relationships between variables in each time window are identified, and the final causal graph is merged to generate.
The accurate discovery of nonlinear real-time and lagging causal relationships in time series data is achieved, which significantly improves the update efficiency and accuracy of the causal graph structure, and is suitable for complex dynamic systems.
Smart Images

Figure CN120069094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of next-generation Internet applications, artificial intelligence, and cyberspace security, and particularly to a method and device for discovering non-linear causal relationships in time series data. Background Art
[0002] Causal discovery is the process of identifying causal relationships between variables from observational data. Different from merely revealing the correlation between variables in a statistical sense, a causal relationship means that one variable directly causes a change in another variable. Generally, the causal relationships between multiple variables can be represented by a directed acyclic graph, where the directed edges represent the causal directions between variables, and such a directed acyclic graph is called a causal graph. Time series causal discovery is the causal discovery of data with time series characteristics, and its causal relationships not only reflect the immediate effects between variables at the same time point, but also include the lag effects of the change of a certain variable on another variable at subsequent time points. Causal discovery is widely applied in fields such as malicious behavior recognition, security risk analysis, network protocol optimization, network vulnerability identification, and intelligent healthcare. By exploring the causal structures in these fields, the system mechanisms can be better understood and effective interventions can be implemented.
[0003] Intervention is a core concept in structural causal inference, which refers to artificially changing the values of one or more variables in a system to observe the impact of such changes on other variables. In causal inference, intervention is used to distinguish causal relationships from correlation relationships, helping researchers understand the causal dependencies between different variables. For example, in a causal graph, the edges between nodes represent potential causal influences, and the intervention operation tests the impact of modifying a certain node's state on other nodes by modifying the state of a certain node, thereby helping to verify causal relationships.
[0004] A time-expanded graph is an extended graph structure that can represent multiple levels or types of node and edge relationships simultaneously. A time-expanded graph is a multi-layer graph, where each layer of the graph represents a specific type of relationship, and the connections between nodes within a layer and between layers represent the associations of the same type and different type attributes respectively.
[0005] Independent component analysis is a commonly used signal separation technique, whose basic assumption is that the observed signal is linearly composed of multiple independent source signals that satisfy non-Gaussian distributions, and the purpose is to recover these independent source signals. The independent component analysis algorithm optimizes an objective function to maximize the independence of each separated component, thereby achieving the demixing of the mixed signal and the extraction of the source signal. The kernel method is a commonly used feature transformation method in machine learning, which calculates the similarity between input data points using a kernel function to find the mapping of the feature implicit high-dimensional space.
[0006] The variational autoencoder is a deep generative model that learns the latent representation of data through the encoder and decoder network structures. The goal of the variational autoencoder is to approximate the distribution of the latent variables by maximizing the marginal log-likelihood of the data. Different from traditional autoencoders, the variational autoencoder introduces a probabilistic graphical model, treating the latent space as a distribution rather than just a fixed point, making the generative model more flexible and interpretable. During the training process, the variational autoencoder optimizes the lower bound (ELBO) through variational inference, thereby achieving both data reconstruction and regularization of the latent space. Summary of the Invention
[0007] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0008] The present invention proposes a method for discovering the non-linear causal relationship of time series data, aiming to discover the non-linear instantaneous causal relationship and non-linear lag causal relationship from time series data, and to update the causal graph structure.
[0009] Another object of the present invention is to propose a device for discovering the non-linear causal relationship of time series data.
[0010] To achieve the above object, on the one hand, the present invention proposes a method for discovering the non-linear causal relationship of time series data, including: Construct a training data set based on multivariate time series data, and extract window observation data from the training data set according to the determined time window; Map the first original training data extracted from the window observation data to a high-dimensional space, process the mapped feature data, and identify the instantaneous causal relationship between variables within each time window to obtain an instantaneous causal graph; Train a variational autoencoder using the second original training data extracted from the window observation data, and analyze the change in the output distribution based on the latent variable representation of the intervention variable to determine the causal relationship, so as to obtain a lag causal graph; Merge the instantaneous causal graph and the lag causal graph, and analyze the legality of the merged causal graph to generate a final causal graph.
[0011] The method for discovering the non-linear causal relationship of time series data according to the embodiments of the present invention may further have the following additional technical features: In an embodiment of the present invention, constructing a training data set based on multivariate time series data, and extracting window observation data from the training data set according to the determined time window includes: Collect observation variables containing multi-dimensional timestamps and used for causal discovery to obtain multivariate time series data, so as to construct a training data set; Determine the unit time interval, discretize the training data set according to the time stamps, and divide the global time axis into several time intervals of a preset length; Obtain the minimum time stamp and the maximum time stamp of all the observed data, and calculate the number of time intervals and the time interval numbers of each piece of data; Determine the width of the time window according to the unit time interval, and randomly extract all the observed data within a single time window from the training data set based on the width of the time window.
[0012] In an embodiment of the present invention, the first original training data extracted from the window observed data is subjected to high-dimensional space mapping, and the mapped feature data is processed to identify the immediate causal relationships between variables within each time window, so as to obtain an immediate causal graph, including: Define the time interval positions according to the number of time intervals, extract the first original training data from all the observed data, and perform centering and whitening operations on the first original training data to obtain the processed first original training data; Perform high-dimensional space mapping on the processed first original training data using a kernel function and calculate the kernel matrix Perform eigenvalue decomposition on the kernel matrix to obtain a basis matrix; Perform eigenvector transformation on the basis matrix to obtain a feature space data matrix, and perform independent component analysis on the feature space data matrix to obtain a confusion matrix; Perform sorting and pruning processing on the confusion matrix, calculate a residual matrix according to the feature space data matrix and the processed confusion matrix, and obtain the correspondence between the residuals and the first original training data; Perform an independence test on the first original training data and the residuals based on the correspondence to determine the causal direction, and delete the causal graph loop structure to obtain the immediate causal graph of the time layer.
[0013] In an embodiment of the present invention, performing eigenvalue decomposition on the kernel matrix to obtain a basis matrix includes: Perform eigenvalue decomposition on the kernel matrix to obtain eigenvalues and eigenvectors; Stack the eigenvectors by columns to form a basis matrix, where each eigenvector forms a column of the basis matrix.
[0014] In an embodiment of the present invention, performing independent component analysis on the feature space data matrix to obtain a confusion matrix includes: Perform independent component analysis on the feature space data matrix to randomly initialize the confusion matrix and construct an optimization objective as a likelihood function that maximizes the target independent components; Calculate the gradient of the likelihood function with respect to the confusion matrix, and use the gradient ascent method to iteratively update the confusion matrix.
[0015] In one embodiment of the present invention, sorting and pruning the confusion matrix includes: Adjust the order of the rows of the confusion matrix to maximize the sum of the diagonal elements of the confusion matrix, so as to sort the confusion matrix; Prune the sorted confusion matrix to set the element with the minimum confidence in the matrix to .
[0016] In one embodiment of the present invention, performing an independence test on the first original training data and the residuals to determine the causal direction includes: Traversingly extract data pairs from the first original training data; Calculate the mutual information between the residuals and the data pairs to judge the direction of the causal edge.
[0017] In one embodiment of the present invention, training a variational autoencoder using the second original training data extracted from the window observation data, and analyzing the change of the output distribution based on the latent variable representation of the intervention variable to determine the causal relationship, so as to obtain a lagged causal graph, includes: Extract the data of the complete time window from all the observation data and record it as the second original training data; Randomly sort the second original training data and perform normalization processing on all features to obtain the row-normalized second original training data; Initialize a variational autoencoder including an encoder and a decoder, both of which have a multi-layer linear neural network structure, and the output of the encoder is a latent variable; Select a batch of data from the normalized second original training data and input it into the variational autoencoder to obtain the output data to calculate the loss function of the variational autoencoder, and obtain the trained variational autoencoder based on the loss; Traversingly select variables from the second original training data, and add a noise to the latent variable corresponding to the variable to obtain the latent variable representation of the intervention variable; Use the trained variational autoencoder to generate a new time series variable matrix according to the latent variable representation of the intervention variable, and obtain a difference matrix based on the difference between the second original training data and the new time series variable matrix; Compare the absolute value of the difference matrix with the preset threshold to determine the causal relationship between variables according to the numerical comparison result, so as to obtain a lagged causal graph.
[0018] In one embodiment of the present invention, the method further includes: Initialize a variational autoencoder including an encoder and a decoder, where the input dimension of the encoder and the output dimension of the decoder are preset dimensions, both have a multi-layer linear neural network structure, the output of the encoder is a latent variable, which follows a multi-dimensional Gaussian distribution; initialize the parameters using a normal distribution; Select a batch of data from the normalized second original training data and input it into the autoencoder to obtain output data, calculate the loss function of the variational autoencoder, and use the gradient descent method to update the model parameters to train the variational autoencoder. Traversingly select the current variable, and add a noise to the latent variable corresponding to the current variable to obtain the latent variable representation of the intervention variable. According to the latent variable representation of the intervention variable, use the trained variational autoencoder to generate a new time series variable matrix, and obtain a difference matrix based on the difference between the matrix of the second original training data and the new time series variable matrix. If the absolute value of the difference matrix is greater than the preset threshold, and the first moment at which the row element variable is located is less than the second moment at which the row element variable is located, there is a causal relationship between the row element variable and the column element variable, and an edge connecting from the row element variable node to the column element variable node is connected in the causal graph.
[0019] To achieve the above object, on the other hand, the present invention proposes a device for discovering non-linear causal relationships in time series data, including: A data acquisition and processing module, configured to construct a training data set based on multivariate time series data, and extract window observation data from the training data set according to a determined time window. An immediate causal graph construction module, configured to perform high-dimensional space mapping on the first original training data extracted from the window observation data, and process the mapped feature data to identify the immediate causal relationships between variables within each time window to obtain an immediate causal graph. A lagged causal graph construction module, configured to train a variational autoencoder using the second original training data extracted from the window observation data, and analyze the change in the output distribution based on the latent variable representation of the intervention variable to determine the causal relationship to obtain a lagged causal graph. A final causal graph generation module, configured to merge the immediate causal graph and the lagged causal graph, and analyze the legality of the merged causal graph to generate a final causal graph.
[0020] The method and device for discovering non-linear causal relationships in time series data according to the embodiments of the present invention discover non-linear immediate causal relationships and non-linear lagged causal relationships from time series data, and realize the update of the causal graph structure. Compared with the traditional linear causal discovery method, it has fewer constraints, is convenient for modeling non-linear causal relationships, and significantly improves the accuracy and practicality of time series causal relationship discovery.
[0021] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0022] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 is a structural diagram of a method for discovering nonlinear causal relationships in time series data according to an embodiment of the present invention; Figure 2 is a flow chart of a method for discovering nonlinear causal relationships in time series data according to an embodiment of the present invention; Figure 3 4 is a structural diagram of a device for discovering nonlinear causal relationships in time series data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0024] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0025] The following describes a method and device for discovering nonlinear causal relationships in time series data according to an embodiment of the present invention with reference to the accompanying drawings.
[0026] The present invention is mainly aimed at multivariate time series observation data, aiming to simultaneously discover nonlinear immediate causal relationships and lagged causal relationships, and realize automatic updating of causal graph structures. Specifically, the present invention is divided into two stages. In the immediate causal relationship discovery stage, the kernel independent component analysis algorithm is adopted to accurately identify the immediate causal dependencies between variables in each time window by mapping the data to a high-dimensional feature space and iteratively sorting, pruning and independence testing; in the lagged causal relationship discovery stage, the variational autoencoder is used to learn the potential representation of the observed variables, and the causal effects between the variables are evaluated by intervening in the latent space, thereby exploring the nonlinear lagged causal influences between different time layers. In order to ensure that the causal graph satisfies the directed acyclic property and has no reverse time causal edges, the present invention maintains the legitimacy of the graph structure based on the order of time series intervals and iterative deletion of loop structures. Its overall architecture is as follows Figure 1 shown.
[0027] The goal of the present invention is to achieve causal discovery of time series data and automatic generation of causal graph structure.
[0028] Figure 2It is a flowchart of a method for discovering the non - linear causal relationship of time - series data according to an embodiment of the present invention. As Figure 2 shown, the method includes: S1. Construct a training data set based on multivariate time - series data, and extract window observation data from the training data set according to the determined time window.
[0029] It can be understood that this step is for multivariate time - series data acquisition and data processing. Specifically, it is to collect raw data, determine the time window, then extract data according to the time window, and perform data processing as follows: S11. Collect observation variables containing multi - dimensional timestamps and used for causal discovery to obtain multivariate time - series data, so as to construct a training data set; S12. Determine the unit time interval, discretize the training data set according to the timestamps, and divide the global time axis into several time intervals with a preset length; S13. Obtain the minimum timestamp and the maximum timestamp of all the observation data, and calculate the number of time intervals and the time - interval number of each data; S14. Determine the width of the time window according to the unit time interval, and randomly extract all the observation data within a single time window from the training data set based on the width of the time window.
[0030] Among them, the collected multivariate time - series data should contain timestamps and observation variables that can be used for causal discovery. The observation data can come from an open - source data set or be sampled from the task scenario.
[0031] Among them, the way to determine the time window is to first determine the unit time interval , and ensure that within a time period of length , the distribution of each observation variable remains relatively stable.
[0032] Among them, the way of data processing is to discretize the training data according to the timestamps, and divide the global time axis into several time intervals with a fixed length of . Obtain the minimum timestamp and the maximum timestamp of the training data , then the number of time intervals can be calculated by the following formula:
[0033] And number each interval from to . Calculate which time interval each data belongs to one by one. For the data with a timestamp of The data, and the number of its time interval is:
[0034] Among them, according to Determine the width of the time window , to ensure Within the range of [specific number] unit time, all possible lagged causal effects between variables can be captured.
[0035] Among them, the data extraction method is to obtain all the observation data within a single time window randomly from the training data set obtained from S11 .
[0036] S2. Map the first original training data extracted from the window observation data to a high-dimensional space, and process the mapped feature data to identify the immediate causal relationships between variables within each time window, so as to obtain an immediate causal graph.
[0037] It can be understood that the present invention obtains immediate causal relationships. First, define the time interval position , and initialize the current time interval position . Extract the first original training data , and perform centering and whitening operations on the data. Use the kernel function to map the data to a high-dimensional space and calculate the kernel matrix . Perform eigenvalue decomposition on the kernel matrix to obtain the basis matrix . Perform feature transformation on the data to obtain the feature space data matrix. Execute the independent component analysis algorithm on the feature space data matrix to obtain the confusion matrix. Iteratively sort and prune the confusion matrix. Calculate the residual matrix according to the confusion matrix and the feature space data matrix, and obtain the corresponding relationship between the residuals and the original data. Conduct an independence test on the original data and the residuals to determine the causal direction. Delete the causal graph loop structure to obtain The immediate causal graph of the time layer. Continue to obtain the immediate causal relationships of the next time layer until the immediate causal graph is constructed. Among them, the variable relationship at the time interval position P corresponds to the Pth time layer in the causal graph. Specifically, it includes the following steps: S21. Extract the data of the th time interval within the time window from all the observation data in step S13, and record the first original training data, which is the original data matrix , The dimension of .
[0038] S22. The method of performing centering and whitening operations on the data is that all data are subtracted by the feature mean value; the features are standardized to keep the variance as Perform singular value decomposition on the data to ensure the independence of different features, and obtain the processed first original training data.
[0039] S23. Use the kernel function to perform high-dimensional space mapping on the processed first original training data and calculate the kernel matrix , where the polynomial kernel function is selected as the kernel function, which can be expressed as:
[0040] where represent any two pieces of data in the data matrix respectively, represents the inner product, is a hyperparameter. Calculate the kernel matrix according to the kernel function , where each element has a value given by .
[0041] S24. Perform eigenvalue decomposition on the kernel matrix to obtain the eigenvalues and eigenvectors . Stack these eigenvectors column by column to form a basis matrix , where each eigenvector forms a column of the basis matrix.
[0042] S25. Perform feature transformation on the basis matrix by multiplying the basis matrix with the original data matrix to obtain the feature space matrix .
[0043] S26. Perform the independent component analysis algorithm on the feature space vectors, randomly initialize the mixing matrix , and construct the optimization objective as the likelihood function that maximizes the likelihood of the target independent components , which can be expressed as:
[0044] where represents the i-th column of the feature space matrix, is the probability density function. Calculate the gradient of the likelihood function with respect to the mixing matrix , and use the gradient ascent method to iteratively update the matrix until the likelihood function is greater than the threshold .
[0045] S27. The way to sort the mixing matrix is to adjust the order of the rows of the mixing matrix to maximize the sum of the diagonal elements of the matrix. The -th row elements of this matrix correspond to the The contribution of column elements can be regarded as the value at the corresponding position in the matrix. This optimization problem can be transformed into a bipartite graph matching problem, and the Hungarian algorithm is used to obtain the optimal solution. Specifically: Step S27.1: For the confusion matrix , first, find the row minimum for each row and subtract it from all elements in that row to form a new matrix .
[0046] Step S27.2: Based on the matrix , find the column minimum for each column and subtract it from all elements in that column to obtain the matrix .
[0047] Step S27.3: Use the minimum number of horizontal or vertical lines to cover all zero elements in the matrix .
[0048] Step S27.4: Calculate the number of covering lines and denote it as . If is equal to the number of rows of the confusion matrix A, the Hungarian algorithm is completed; otherwise, find the minimum value among all uncovered elements, subtract this value from the uncovered elements, and add this value to the elements covered by the intersection of two lines. Let the updated matrix be , and return to Step S27.2.
[0049] S28, the way to prune the sorted confusion matrix is to set the element with the minimum confidence in the matrix to . Its confidence can be the size of the matrix element itself or the size of the statistical test index of its corresponding row and column, such as the size of the Wald test index.
[0050] S29, return to Step S28 until the confusion matrix after the end of S28 is a lower triangular matrix.
[0051] S210, calculate the residual matrix according to the confusion matrix and the feature space data matrix . The calculation method is
[0052] S211, the corresponding relationship between the residual and the original data is that for the residual vector , record the row of its confusion matrix as , and the feature space data vector corresponding to the column with the same value as it is its corresponding feature space data vector, corresponding to as its original data.
[0053] S212. The way to perform an independence test on the original data and the residuals is to iteratively extract data pairs from the original data , calculate the mutual information between the residuals and the data, and determine the direction of the causal edge.
[0054] S213. The way to calculate the mutual information between the residuals and the data is that for the data pair , calculate and respectively in the mutual information. Calculation method: where is the conditional entropy of under the condition of , and is or . Compare and . If the former is larger, the causal direction is ; otherwise, the causal direction is
[0055] S214. The way to delete the cyclic structure of the causal graph is to iteratively find the cyclic structure in the graph and delete the edge with the smallest edge weight in the cyclic structure until the causal graph forms a directed acyclic graph.
[0056] S215. The way to obtain the instantaneous causal relationship of the next time layer until the construction of the instantaneous causal graph is completed is to set . At this time, if , the construction of the instantaneous causal graph is completed, and go to step S3; otherwise, return to step S21.
[0057] It can be known that the present invention uses a variational autoencoder (VAE) to learn the latent space representation of variables. Specifically, the original variables are mapped to a high-dimensional feature space through a kernel method to capture the complex nonlinear relationships between variables. In the high-dimensional space, the latent space representation of each variable can better reveal its internal structure and potential patterns, and thus provide a richer information basis for subsequent causal relationship analysis.
[0058] It can be known that the present invention further expands the causal graph into a time-expanded graph for modeling causal relationships in time series data. Each layer represents a time unit, where the edges within the layer represent immediate causal effects, and the edges between layers represent lagged causal effects. On this basis, the feature matrix is processed through independent component analysis (ICA) to calculate the residual matrix, and based on the correspondence between the residuals and the original data, it is determined whether there is a causal relationship between variables. Among them, through independent component analysis of the feature matrix, the residual matrix is extracted. This step helps to separate the independent components of each variable in order to more accurately evaluate its impact on the overall system. Among them, according to the correspondence between the residuals and the original data, an independence test is performed to judge the causal direction between variables and construct a causal graph. The cyclic structure in the causal graph is deleted to obtain the final immediate causal graph of the time layer. This process ensures the accuracy and acyclicity of the causal relationship, improving the reliability and interpretability of the model.
[0059] S3. Use the second original training data extracted from the window observation data to train the variational autoencoder, and determine the causal relationship based on the analysis of the output distribution change of the latent variable representation of the intervention variable to obtain the lagged causal graph.
[0060] S31. The way to extract the training data is to randomly extract the data of a complete time window from all the observation data in step S13, which is denoted as the second original training data, that is, recorded as the original data matrix , The dimension of .
[0061] S32. The data processing method is to randomly sort the second original training data and normalize all features.
[0062] S33. The way to initialize the variational autoencoder is to initialize a variational autoencoder including an encoder and a decoder. The input dimension of the encoder and the output dimension of the decoder are preset dimensions. In one embodiment of the present invention, it is ([[]] , and the network structures are all multi-layer linear neural networks. The output of the encoder is the latent variable , which follows a multi-dimensional Gaussian distribution . Initialize the parameters using the normal distribution and use the reparameterization trick to maintain the continuity of the gradient.
[0063] S34. The way to train the variational autoencoder is to select a batch of data from the normalized second original training data and send it into the autoencoder to obtain the output, calculate the loss function of the variational autoencoder, and obtain the trained variational autoencoder based on the loss. The loss function of the variational autoencoder can be expressed as:
[0064] Among them, is the model output, is a hyperparameter, represents the mean squared error loss, represents the KL divergence of the distribution. The model parameters are updated using the gradient descent method. Repeat this step until the model converges.
[0065] S35, variables are iteratively selected from the second original training data , in the corresponding latent variable add a noise to obtain the latent variable representation of the intervened variable, i.e.:
[0066] S36, the way to observe the change in the output distribution is, according to obtained in step S35, use the trained variational autoencoder to generate a new time series variable matrix. Calculate the difference between the original data matrix and the new time series variable matrix to obtain the difference matrix . If the absolute value of the difference matrix is greater than the preset threshold , and the row element variable in the difference matrix is at the time <variable is at the time , that is, corresponding to the first moment and the second moment in the embodiment of the present invention, then there is a causal relationship between the row element variable and the column element variable, and connect the edge from the node to the column element variable node in the causal graph. Return to step S35 until all variables in the time window are traversed.
[0067] It can be known that the present invention intervenes in the latent space representation of features and judges the causal relationship by calculating the difference between the feature matrices before and after the intervention.
[0068] It can be known that the present invention uses a variational autoencoder to learn the latent space representation of variables, intervenes in this latent space representation and calculates the difference of the feature matrix to judge whether there is a causal relationship between variables.
[0069] It can be known that the present invention expands the causal graph into a time-expanded graph for modeling the causal relationship in time series data, where each layer represents a time unit. In this type of causal graph, the edges within the layer are used to represent the immediate causal effect; the edges between layers are used to represent the lagged causal effect.
[0070] S4, merge the immediate causal graph and the lagged causal graph, and analyze the legality of the merged causal graph to generate the final causal graph.
[0071] S41, the causal graphs of step S2 and step S3 are merged by making a one-to-one correspondence between the nodes of the delayed causal graph obtained in step S3 and the nodes of the immediate causal graph in step S2, and taking the union of the edge sets of the two graphs.
[0072] S42, the way to check the legitimacy of the causal graph is to first check whether there is a loop in the merged causal graph. If so, return to step S2; otherwise, check whether there is a reverse time edge in the merged causal graph. A reverse time edge is defined as the time interval value of the starting point of the edge is greater than the time interval value of the end point of the edge. If there is a reverse time edge, return to step S3; otherwise, generate the final causal graph and the algorithm is executed.
[0073] According to the method for discovering nonlinear causal relationships of time series data according to the embodiment of the present invention, it is aimed at multivariate time series observation data, and aims to simultaneously discover nonlinear immediate causal relationships and lagged causal relationships, and realize the automatic update of the causal graph structure. In the immediate causal relationship discovery stage, the present invention adopts the kernel independent component analysis algorithm to accurately identify the immediate causal dependency between variables in each time window by mapping the data to a high-dimensional feature space and iteratively sorting, pruning and independence testing; in the lagged causal relationship discovery stage, the variational autoencoder is used to learn the potential representation of the observed variables, and the causal effects between the variables are evaluated by intervening in the latent space, thereby mining the nonlinear lagged causal influence between different time layers. In order to ensure that the causal graph satisfies the directed acyclic property and has no reverse time causal edges, the present invention maintains the legitimacy of the graph structure based on the order of time series intervals and iterative deletion of loop structures. Compared with other methods, the present invention can adapt to the complex nonlinear characteristics of time series data, and significantly improves the effectiveness, accuracy and practicality of causal discovery of time series data.
[0074] The beneficial effects of the present invention are: The present invention discovers nonlinear immediate causal relationships and nonlinear lagged causal relationships and updates the causal graph structure from time series data. On the one hand, by accurately separating independent components in high-dimensional feature space, and iterative causal edge pruning and direction judgment, the erroneous associations that are prone to occur in traditional linear methods when faced with complex real-world scenarios are reduced; on the other hand, by intervening in latent space to detect lagged causal relationships, flexible mining of nonlinear causal effects at multiple time series levels is achieved. Compared with traditional causal discovery methods that can only handle static scenarios or linear relationships, the present invention can efficiently and accurately model nonlinear time series data causal relationships, is closer to real-world causal application scenarios, and can be used in a variety of complex dynamic systems, such as network protocol optimization, medical diagnosis decisions, etc.
[0075] In order to implement the above embodiment, Figure 3 As shown, this embodiment also provides a device 10 for discovering nonlinear causal relationships of time series data, including: The data acquisition and processing module 100 is used to construct a training data set based on multivariate time series data, and extract window observation data from the training data set according to a determined time window; The instant causal graph construction module 200 is used to perform high-dimensional space mapping on the first original training data extracted from the window observation data, and process the mapped feature data to identify the instant causal relationship between the variables in each time window to obtain an instant causal graph; A lagged causal graph construction module 300 is used to train a variational autoencoder using second original training data extracted from the window observation data, and to determine the causal relationship by analyzing the output distribution change based on the latent variable representation of the intervention variable, so as to obtain a lagged causal graph; The final causal graph generation module 400 is used to merge the immediate causal graph and the delayed causal graph, and analyze the legality of the merged causal graph to generate a final causal graph.
[0076] According to the device for discovering nonlinear causal relationships of time series data according to the embodiment of the present invention, it is aimed at multivariate time series observation data, and aims to simultaneously discover nonlinear immediate causal relationships and lagged causal relationships, and realize the automatic update of the causal graph structure. In the immediate causal relationship discovery stage, the present invention adopts the kernel independent component analysis algorithm, and accurately identifies the immediate causal dependency between variables in each time window by mapping the data to a high-dimensional feature space and iteratively sorting, pruning and independence testing; in the lagged causal relationship discovery stage, the variational autoencoder is used to learn the potential representation of the observed variables, and the causal effect between the variables is evaluated by intervening in the latent space, so as to mine the nonlinear lagged causal influence between different time layers. In order to ensure that the causal graph satisfies the directed acyclic property and has no reverse time causal edges, the present invention maintains the legitimacy of the graph structure based on the time series interval order and iterative deletion of the loop structure. Compared with other methods, the present invention can adapt to the complex nonlinear characteristics of time series data, and significantly improves the effectiveness, accuracy and practicality of causal discovery of time series data.
[0077] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0078] Furthermore, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A method for discovering nonlinear causal relationships in time series data, characterized in that: include: Construct a training data set based on multivariate time series data, and extract window observation data from the training data set according to a determined time window; The first original training data extracted from the window observation data is mapped into a high-dimensional space, and the mapped feature data is processed to identify the immediate causal relationship between the variables in each time window to obtain an immediate causal graph; The variational autoencoder is trained using the second original training data extracted from the window observation data, and the causal relationship is determined by analyzing the output distribution change based on the latent variable representation of the intervention variable to obtain a lagged causal graph; The immediate causal graph and the delayed causal graph are merged, and the legality of the merged causal graph is analyzed to generate a final causal graph.
2. The method according to claim 1, characterized in that A training data set is constructed based on multivariate time series data, and window observation data is extracted from the training data set according to a determined time window, including: Collect observed variables containing multi-dimensional timestamps and used for causal discovery to obtain multivariate time series data to construct a training data set; Determine the unit time interval, discretize the training data set according to the timestamp, and divide the global time axis into several time intervals of preset length; Get the minimum and maximum timestamps of all observed data, and calculate the number of time intervals and the time interval number of each data; The width of the time window is determined according to the unit time interval, and all observations within a single time window are randomly extracted from the training data set based on the width of the time window.
3. The method according to claim 2, characterized in that The first original training data extracted from the window observation data is mapped into a high-dimensional space, and the mapped feature data is processed to identify the immediate causal relationship between the variables in each time window to obtain an immediate causal graph, including: Defining the time interval position according to the number of time intervals, extracting the first original training data from all the observation data, and performing centering and whitening operations on the first original training data to obtain processed first original training data; Use the kernel function to map the processed first original training data into a high-dimensional space and calculate the kernel matrix Perform eigenvalue decomposition on the kernel matrix to obtain the basis matrix; Perform feature transformation on the basis matrix to obtain a feature space data matrix, and perform independent component analysis on the feature space data matrix to obtain a confusion matrix; The confusion matrix is sorted and pruned, and a residual matrix is calculated according to the feature space data matrix and the processed confusion matrix, and a corresponding relationship between the residual and the first original training data is obtained; Based on the corresponding relationship, an independence test is performed on the first original training data and the residual to determine the causal direction, and the causal graph loop structure is deleted to obtain an instant causal graph at the time layer.
4. The method according to claim 3, characterized in that Perform eigenvalue decomposition on the kernel matrix to obtain the basis matrix, including: Perform eigenvalue decomposition on the kernel matrix to obtain eigenvalues and eigenvectors; The eigenvectors are stacked in columns to form a basis matrix, where each eigenvector constitutes a column of the basis matrix.
5. The method according to claim 4, characterized in that Perform independent component analysis on the feature space data matrix to obtain the confusion matrix, including: Perform independent component analysis on the feature space data matrix, randomly initialize the confusion matrix, and construct an optimization goal to maximize the likelihood function of the target independent component; Calculate the gradient of the likelihood function with respect to the confusion matrix and iteratively update the confusion matrix using the gradient ascent method.
6. The method according to claim 5, characterized in that Sorting and pruning of the confusion matrix, including: Adjust the order of the rows of the confusion matrix to maximize the sum of the diagonal elements of the confusion matrix to sort the confusion matrix; The sorted confusion matrix is pruned to set the element with the smallest confidence in the matrix as .
7. The method according to claim 6, characterized in that Independence tests are performed on the first original training data and residuals to determine the causal direction, including: ergodicly extracting data pairs from the first original training data; The mutual information between residual and data pairs is calculated to determine the direction of the causal edge.
8. The method according to claim 1, characterized in that The variational autoencoder is trained using the second original training data extracted from the window observation data, and the causal relationship is determined by analyzing the output distribution change based on the latent variable representation of the intervention variable to obtain a lagged causal graph, including: Extract the data of the complete time window from all the observation data and record them as the second original training data; Randomly sort the second original training data, and normalize all features to obtain the second original training data after row normalization; Initialize a variational autoencoder that includes an encoder and a decoder. The network structure is a multi-layer linear neural network, and the output of the encoder is a hidden variable. Selecting a batch of data from the normalized second original training data and inputting it into the variational autoencoder, obtaining output data to calculate the loss function of the variational autoencoder, and obtaining a trained variational autoencoder based on the loss; ergodicly selecting variables from the second original training data, and adding a noise to the latent variable corresponding to the variable to obtain the latent variable representation of the intervention variable; Generate a new time series variable matrix based on the latent variable representation of the intervention variable using the trained variational autoencoder, and obtain a difference matrix based on the difference between the second original training data and the new time series variable matrix; The absolute value of the difference matrix is compared with the size of the preset threshold to determine the causal relationship between the variables based on the numerical comparison results to obtain a lagged causal diagram.
9. The method according to claim 8, characterized in that The method further comprises: Initialize a variational autoencoder including an encoder and a decoder, where the input dimension of the encoder and the output dimension of the decoder are preset dimensions, the network structure is a multi-layer linear neural network, the output of the encoder is a latent variable, and obeys a multi-dimensional Gaussian distribution; use normal distribution to initialize parameters; Select a batch of data from the normalized second original training data and input it into the autoencoder to obtain output data, calculate the loss function of the variational autoencoder; and use the gradient descent method to update the model parameters to train the variational autoencoder; Erratically select the current variable, add a noise to the latent variable corresponding to the current variable to obtain the latent variable representation of the intervention variable; According to the latent variable representation of the intervention variable, a new time series variable matrix is generated using the trained variational autoencoder, and the difference matrix is obtained based on the difference between the matrix of the second original training data and the new time series variable matrix; if the absolute value of the difference matrix is greater than the preset threshold, and the row element variable of the difference matrix is less than the row element variable at the second moment, then there is a causal relationship between the row element variable and the column element variable, and the edge from the row element variable node to the column element variable node is connected in the causal graph.
10. A device for discovering nonlinear causal relationships in time series data, characterized in that: include: A data acquisition and processing module is used to construct a training data set based on multivariate time series data and extract window observation data from the training data set according to a determined time window; An instant causal graph construction module is used to perform high-dimensional space mapping on the first original training data extracted from the window observation data, and process the mapped feature data to identify the instant causal relationship between the variables in each time window to obtain an instant causal graph; A lagged causal graph construction module is used to train a variational autoencoder using second original training data extracted from the window observation data, and to determine the causal relationship by analyzing the output distribution change based on the latent variable representation of the intervention variable, so as to obtain a lagged causal graph; The final causal graph generation module is used to merge the immediate causal graph and the delayed causal graph, and analyze the legality of the merged causal graph to generate a final causal graph.
Citation Information
Cited By
Aero-engine fault positioning method based on causal discovery
CN120258125A
High-temperature furnace intelligent control system and method based on machine learning
CN121386413A
Spacecraft magnetic moment attitude control method based on non-inverse iterative algorithm
CN122009527A
Spacecraft magnetic torque attitude control method based on inverse-free iterative algorithm
CN122009527B
Abnormality diagnosis method and device for sensor of power equipment monitoring system
CN122130142A