Multivariable time sequence prediction interpretation method and system based on information theory and causal reasoning

By optimizing the causality inference algorithm and dynamic feature weight calculation based on information theory and causal reasoning methods, the time-element two-dimensional heat map is generated, and the time continuity and model dependence problems in multivariate time series prediction explanation are solved, providing an intuitive prediction explanation.

CN120373478AActive Publication Date: 2025-07-25SHANDONG FUTURE NETWORK RES INST (PURPLE MOUNTAIN LAB IND INTERNET INNOVATION APPL BASE)

Patent Information

Application Number
CN202510854544.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Existing artificial intelligence models lack time continuity considerations in multivariate time series prediction interpretation, the explanation results are not intuitive, the model is highly dependent, and it is difficult to provide users with an easy-to-understand explanation of decision-making process.

Method used

Using the method based on information theory and causal reasoning, by obtaining the original multivariable time series data, performing data preprocessing, calculating dynamic feature weights, and generating a time-feature two-dimensional heat map, combining causal information inference and dynamic threshold adjustment, causality inference algorithm is optimized to generate an intuitive and easy-to-understand prediction explanation.

Benefits of technology

It realizes the more accurate exploration of causal relationships in time series prediction, provides intuitive and easy-to-understand interpretation results, reduces model dependence, and improves user comprehension and interpretation reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373478A_ABST
    Figure CN120373478A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence model interpretation, in particular to a multivariable time sequence prediction interpretation method and system based on an information theory and causal reasoning. The method comprises the following steps: acquiring original multivariable time sequence data; performing data preprocessing on the obtained original multivariable time sequence data; performing causal information reasoning on the preprocessed multivariable time sequence data; based on causal information reasoning, calculating a dynamic feature weight of the multivariable time sequence data; performing local fitting and interpretation generation based on the dynamic feature weight; and a time-feature two-dimensional thermodynamic diagram is obtained. Through multi-dimensional feature relationship mining, an optimized causal relationship reasoning algorithm, dynamic weight calculation and an intuitive time-feature two-dimensional thermodynamic diagram generation mechanism, time continuity is effectively considered, prediction interpretation is presented in an intuitive and understandable mode, and a user can quickly understand an artificial intelligence model decision process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence model interpretation, and in particular to a method and system for predicting and interpreting multivariate time series based on information theory and causal reasoning. Background Art

[0002] Artificial intelligence is widely used in time series prediction, covering multiple fields such as healthcare, physics, energy, and sensor data. In these application scenarios, it is crucial to ensure that users can quickly understand the reasons for artificial intelligence decisions, especially in time-sensitive decision-making assistance scenarios. Interpretation methods help explore the factors behind decisions, thereby increasing human trust in the model and promoting the wide and in-depth application of artificial intelligence.

[0003] Currently, there are some significance methods for explaining the importance of input features for prediction, mainly including: (1) perturbation-based methods, such as feature occlusion FO and RISE, which provide explanations by perturbing the input and comparing the results; (2) gradient-based methods, such as integrated gradients, DeepLIFT, and GradSHAP, which use the gradients of feature representations for interpretation. (3) Attention mechanisms are combined with significance methods, and the effectiveness of these methods in explaining black-box models is still controversial. (4) Other methods, such as SHAP using Shapley values and LIME by randomly combining records and their neighborhoods and weighting according to the proximity of the neighborhoods to generate explanations.

[0004] In addition, some studies have begun to consider the time-sensitive characteristics of time series applications. For example: FIT quantifies the importance of observations over time by evaluating the contribution of observations to the prediction output; Dynamask takes into account the time dependence of time series data in its design and perturbs the input by dynamically combining adjacent values in the input.

[0005] The existing methods have the following deficiencies in explaining multivariate time series prediction.

[0006] (1) Using significance methods such as perturbation and gradient to directly explain time series prediction usually produces results that are difficult for humans to understand. Although it is possible to generate a significance map by scoring the feature importance on a fine time scale, the significance score of each feature has an exact value within a certain time range, and the resulting fragmented results are difficult to provide users with meaningful and easy-to-understand explanations.

[0007] (2) Although the FIT method quantifies the feature importance over time, it still scores in numerical form, and the output results are still not intuitive enough for users.

[0008] (3) The Dynamask method can identify the changes in feature importance over time and endeavors to maintain this continuity in the explanations, making it more user-friendly than previous algorithms such as FIT. However, such methods usually evaluate the importance of features through other internal states of the model (such as gradients), resulting in limitations in practical applications. Summary of the Invention

[0009] In view of the problems existing in the existing prediction explanation methods of artificial intelligence models, such as the lack of consideration of time continuity, the unintuitive explanation results, and the strong model dependence, this proposal presents a multivariate time series prediction explanation method based on information theory and causal reasoning. By means of a multivariate time series prediction explanation method based on information theory and causal reasoning that can effectively consider time continuity, has intuitive and easy-to-understand explanation results, and is model-independent, it helps users better understand the decision-making process of artificial intelligence models in time series prediction.

[0010] In a first aspect, a multivariate time series prediction explanation method based on information theory and causal reasoning provided by the present invention adopts the following technical solutions: Obtain the original multivariate time series data; Perform data preprocessing on the obtained original multivariate time series data; Conduct causal information reasoning on the preprocessed multivariate time series data; Calculate the dynamic feature weights of the multivariate time series data based on the causal information reasoning; Perform local fitting and explanation generation based on the dynamic feature weights; Obtain a time-feature two-dimensional heat map.

[0011] Further, the performing data preprocessing on the obtained original multivariate time series data includes normalizing the original multivariate time series data, where represents the number of features, D represents the number of time observations, T is the data point of feature at time d ; for each feature dimension t , normalization is performed through the formula d , where and are the mean and standard deviation of feature d in the time series respectively; meanwhile, the time series is divided into subsequences of length , and the s -th subsequence is denoted as .

[0012] ​Further, the causal information inference for the preprocessed multivariate time series data includes, for each subsequence, calculating the mutual information between different features. For feature and feature , calculate the joint distribution probability and the marginal distribution probability , , and then calculate the mutual information value. By calculating the mutual information values between every two features respectively, construct a mutual information matrix to record the mutual influence between each feature dimension; the calculation of the mutual information value is expressed as: .

[0013] Further, the causal information inference for the preprocessed multivariate time series data also includes, based on the calculated mutual information matrix, using a constraint-based causal discovery algorithm to mine the causal relationships between features. Among them, construct a complete undirected graph , where the nodes in the graph correspond to each feature dimension of the multivariate time series, and the nodes are connected by undirected edges, indicating that there may be associations between features in the initial state. According to the time order of the time series, add time constraint marks to each edge to record the chronological order information of the occurrence of the association between two features; denote the size of the current conditional set used in the iterative process as , set the maximum conditional set size , starting from , for each edge in the graph, conduct a conditional independence test under different conditional sets formed by other nodes.

[0014] Further, in the orientation stage of the causal information inference for the preprocessed multivariate time series data, if there is time constraint information indicating that the change of feature always precedes the change of feature , and is an undirected edge, and there are no conflicts with other orientation rules, preferentially orient the edge as i→j, and repeatedly apply the orientation rules until all edges are oriented, finally constructing a causal relationship graph.

[0015] Further, based on the causal information inference, calculating the dynamic feature weights of the multivariate time series data includes calculating the causal centrality index, the time series change trend coefficient, and the information gain rate of the features of the multivariate time series data respectively, and integrating them into the final weight , where , and are adjustment coefficients, and , where, first, measure the feature The influence scope of other features as a cause , where represents the feature the number of edges pointing to other features, is the total number of nodes. Next, consider the feature to the prediction target the length and strength of the causal path, and the formula is , where represents all directed path sets from the feature to the prediction target , is the product of the weights on the path , is the path length, is the attenuation coefficient. Combining the above indicators, calculate the causal centrality of the feature where , where is the weight coefficient, and .

[0016] Furthermore, the above-mentioned causal information-based reasoning for calculating the dynamic feature weights of multivariate time series data also includes calculating the time series change trend coefficient. Among them, first calculate the average change rate of the feature at adjacent time points within the subsequence , where is the feature at time value, is the time interval. Next, introduce a time decay mechanism to make recent data have higher weights , where is the attenuation coefficient, is the feature mean within the subsequence. Comprehensively calculate the change trend coefficient , where weight coefficient, and .

[0017] Furthermore, the above-mentioned local fitting and interpretation generation based on dynamic feature weights includes, during the prediction process, for a given input subsequence , use the weighted Euclidean distance to calculate the similarity between the input subsequence and each subsequence in the training data, expressed as: , where is the dynamic weight of the feature . By introducing dynamic weights, the similarity calculation focuses on causal key features and dimensions with obvious temporal trends.

[0018] Further, the local fitting and interpretation generation based on dynamic feature weights further includes selecting several subsequences with the smallest distance as the similar subsequence set, and constructing a weighted linear regression model for the selected similar subsequence set: , where is the regression number of the feature , and is solved by minimizing the weighted mean square error: .

[0019] In a second aspect, a multivariate time series prediction and interpretation system based on information theory and causal reasoning includes: A data acquisition module, configured to acquire original multivariate time series data; A preprocessing module, configured to perform data preprocessing on the acquired original multivariate time series data; An information reasoning module, configured to perform causal information reasoning on the preprocessed multivariate time series data; A weight module, configured to calculate the dynamic feature weights of the multivariate time series data based on causal information reasoning; A fitting module, configured to perform local fitting and interpretation generation based on dynamic feature weights; An output module, configured to obtain a time-feature two-dimensional heat map.

[0020] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the method for predicting and interpreting multivariate time series based on information theory and causal reasoning.

[0021] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the method for predicting and interpreting multivariate time series based on information theory and causal reasoning.

[0022] In summary, the present invention has the following beneficial technical effects: The present invention optimizes the traditional PC algorithm for time series prediction and interpretation. In the initialization stage, by introducing time constraint information and closely combining the time order of time series data, the mined causal relationships are more in line with the actual time logic; in the conditional independence test stage, a dynamic threshold adjustment mechanism is introduced, which can adaptively adjust the conditional independence judgment standard according to the data fluctuation degree, effectively avoiding the situation of misdeleting or missing edges caused by fixed thresholds; in the orientation stage, combining time constraint information with traditional orientation rules further improves the accuracy of edge orientation. The algorithm can closely combine the time series characteristics and data dynamic changes, and more accurately mine real and reliable causal relationships.

[0023] The present invention effectively considers the time continuity through multi-dimensional feature relationship mining, an optimized causal relationship reasoning algorithm, dynamic weight calculation, and an intuitive time-feature two-dimensional heat map generation mechanism, presents the prediction explanation in an intuitive and understandable way, and enables users to quickly understand the decision-making process of the artificial intelligence model. The present invention can more accurately mine the causal relationships in multivariate time series data through time constraint marking, dynamic threshold adjustment, and improved orientation rules, provides a more reliable basis for subsequent prediction explanations, and is more accurate and reliable in causal relationship analysis compared with existing methods. Brief Description of the Drawings

[0024] Figure 1 It is a schematic diagram of a method for predicting and interpreting multivariate time series based on information theory and causal reasoning according to Embodiment 1 of the present invention. Detailed Description of the Embodiments

[0025] The present invention will be further described in detail below with reference to the accompanying drawings.

[0026] Embodiment 1 Referring to Figure 1 , a method for predicting and interpreting multivariate time series based on information theory and causal reasoning in this embodiment includes: Obtaining the original multivariate time series data; Performing data preprocessing on the obtained original multivariate time series data; Performing causal information reasoning on the preprocessed multivariate time series data; Calculating the dynamic feature weights of the multivariate time series data based on the causal information reasoning; Performing local fitting and explanation generation based on the dynamic feature weights; Obtaining a time-feature two-dimensional heat map.

[0027] Specifically: 1. Data preprocessing, Normalizing the original multivariate time series data For each feature dimension d , normalizing through the formula , where and are the mean and standard deviation of the feature d in the time series respectively. At the same time, dividing the time series into subsequences with a length of , and each subsequence is represented as .

[0028] 2. Calculating the mutual information between features, The characteristics of multivariate time series data can be obtained in the following ways.

[0029] (1) In the medical record scenario, patient indicators (such as red blood cell count, blood oxygen saturation, heart rate, body temperature) are extracted based on the in-hospital information system and diagnostic equipment as features, and sorted by timestamp to form a time series. Similarly, in scenarios such as financial transactions, key indicators (such as opening price, transaction amount, position size) are extracted from the business system logs as features and sorted by timestamp to form a time series.

[0030] (2) In scenarios such as industrial monitoring and environmental perception, multi-dimensional time series data is collected in real time by deploying multiple types of sensors (such as temperature, pressure, and flow sensors), and each sensor corresponds to a feature dimension d ( d = 1, 2,..., D ) The length of the time series is T .

[0031] For the calculation process of joint distribution probability and marginal distribution probability, including for each subsequence X (s) , calculate the joint distribution probability and marginal distribution probability between different features. Taking feature and feature as an example, the data discretization method is used to calculate the joint distribution probability and the marginal distribution probability , , and the specific steps are as follows.

[0032] (1) Data discretization: The equal-width binning method is used to discretize continuous time series data and map the feature values to a finite number of intervals (bins). For feature , let the set of discrete intervals be , where K is the number of intervals, the interval width is , the k rd interval is , and for feature similarly, let the set of discrete intervals be , where M is the number of intervals, the interval width is , the m th interval is . For example, if the value range of feature is [10, 30] and it is divided into 5 intervals by equal-width binning, then the intervals are [10, 14), [14, 18),..., [26, 30].

[0033] (2) Calculation of marginal distribution probability, Respectively count the features and features The number of samples of the subsequence data of 、 in each interval; the marginal distribution probability is the ratio of the number of interval samples to the total number of samples, .

[0034] (3)Calculation of joint distribution probability, Statistical feature pairs falling within the joint interval The number of samples , and the joint distribution probability is the ratio of the number of joint interval samples to the total number of samples . Calculation of mutual information value, Based on the discretized probability distribution, calculate the mutual information value of feature and feature through the following formula .

[0035] Mutual information reflects the dependence relationship between two features. The higher the mutual information value, the closer the association between the two features. Calculate the mutual information value of each pair of features respectively to construct a mutual information matrix , thereby recording the mutual influence between each feature dimension.

[0036] Different from traditional methods that analyze feature importance from a single perspective, this proposal explores the inherent dependence relationship between features in multivariate time series data from the perspective of information theory. It is difficult for existing technologies to comprehensively capture the complex associations between features, while mutual information calculation can quantify the dependence degree between different feature dimensions, discover potential data patterns and hidden relationships, making the understanding of multivariate time series data deeper, providing rich and accurate basic data for subsequent causal relationship reasoning, and thus realizing more comprehensive and accurate prediction explanations.

[0037] 3. Causal relationship reasoning, Based on the calculated mutual information matrix, use the constraint-based causal discovery algorithm (PC algorithm) to mine the causal relationship between features. When the traditional PC algorithm is applied to time series causal relationship mining, it does not fully consider the temporal characteristics of time series and the characteristics of data dynamic changes, and there are problems such as misjudgment of causal relationships and inaccurate edge orientation. This patent proposal optimizes on the basis of the traditional PC algorithm for time series prediction scenarios, and the specific steps are as follows.

[0038] (1)Initialization stage: Construct a completely undirected graph , the nodes in the figure correspond to the respective feature dimensions of the multivariate time series, and the nodes are connected by undirected edges, indicating that there may be associations between the features in the initial state. At the same time, according to the time order of the time series, a time constraint label is added to each edge to record the time sequence information of the occurrence of the association between two features.

[0039] Construct a complete undirected graph . Among them is the set of nodes, corresponding to the feature dimensions of the multivariate time series ; is the set of undirected edges, indicating that there may be associations between any two features in the initial state. The initial weight of each undirected edge is the mutual information value, reflecting the dependence strength between the features.

[0040] Determine the time sequence relationship through Granger causality test. First, for feature and feature , establish a regression model with a lag of q periods . Second, test the null hypothesis (that is, is not the Granger cause of ). If is rejected and the average lag time , then , indicating that precedes ; conversely, if is the Granger cause of , then ; if the time sequence cannot be determined clearly, such as changing simultaneously, it is marked as . The output of the final initialization stage is: feature dimension V , undirected edges with initial weights , and time constraint labels .

[0041] (2) Conditional independence test stage: Denote the size of the current conditional set used in the iterative process as , set the maximum conditional set size , starting from , for each edge in the graph, under the condition of a different conditional set formed by other nodes, conduct a conditional independence test. The conditional independence test introduces a dynamic threshold adjustment mechanism, and the threshold is dynamically calculated according to the data fluctuation degree of the current subsequence. The formula is as follows: , where is the adjustment coefficient, and are respectively the feature and the feature at time step , and are respectively the mean values of the feature and the feature within the current subsequence, where is the subsequence length.

[0042] Conditional independence is judged by calculating conditional mutual information. For the feature and the feature , the formula for conditional mutual information under the given condition set is:[[]] , If under a certain condition set , is less than the dynamic threshold , it is considered that under this condition, the feature and the feature are conditionally independent, and at this time, the edge is deleted. After each round of deletion operation, according to the connection situation of the remaining edges and the data characteristics, the value of is dynamically adjusted. If the remaining edges are relatively sparse, the growth step of is increased, otherwise it is decreased. Repeat the above process until no more edges can be deleted.

[0043] (3) Orientation stage: Based on the traditional orientation rules, time constraint information is combined for edge orientation. In addition to following the original orientation rules such as the collider structure, if there is time constraint information indicating that the change of the feature always precedes the feature , and is an undirected edge, and there is no conflict with other orientation rules, the edge is preferentially oriented as i→j. Repeat the application of these orientation rules until all edges are oriented, and finally a causal relationship graph is constructed.

[0044] 1) Traditional orientation rules, Collider structure orientation: If there are three nodes satisfying ① i is connected to j, j is connected to k, but i is not connected to k; ② under a certain condition set S, i and k are conditionally independent, and . Then it is forced to be oriented as , forming a collider structure.

[0045] Non-collider structure orientation: If the edge forms a path and with the already oriented edges , then the edge Oriented as to avoid forming a loop.

[0046] 2) Time constraint priority orientation rule, Based on the time constraint markers in the initialization phase define undirected edges according to the following rules.

[0047] Orientation when the time order is clear: If the edge has a time constraint marker , and the undirected edge has not been oriented by the traditional rules, it is preferentially oriented as ; conversely, if , it is oriented as .

[0048] Conflict handling between time constraints and traditional rules: If the time constraint orientation conflicts with the traditional collision structure orientation, the traditional rules shall prevail. For example but the traditional rules require , then the traditional rules shall prevail and the time constraint marker is adjusted to to avoid violating the causal logic.

[0049] 3) Mathematical iterative process of edge orientation, Define the orientation state matrix , where indicates that the edge is oriented as , indicates that the edge is oriented as , indicates an undirected edge or an unoriented edge.

[0050] 4) Orientation iteration steps, First, initialize the orientation matrix (all edges are undirected); second, apply the traditional rules for orientation, traverse all triples , detect the collision structure, and update or if the conditions are met; in addition, apply the time constraint orientation. For an unoriented edge , if , then update: ; when the orientation matrix no longer changes or all edges are oriented, terminate the iteration, and finally generate a directed causal graph, including: feature dimension V , directed edges , where the edge weight (normalized mutual information value, reflecting the causal strength).

[0051] The optimized PC algorithm of this patent introduces time constraint information, closely combines with the time order of time series data, and makes the mined causal relationships more in line with the actual time logic. The dynamic threshold adjustment mechanism can adaptively adjust the conditional independence judgment criterion according to the data fluctuation degree, effectively avoiding the situation of misdeleting or missing edges caused by fixed thresholds. The improved orientation rule further improves the accuracy of edge orientation. In summary, the optimized algorithm can more accurately mine the true and reliable causal relationships from multivariate time series data.

[0052] 4. Dynamic weight calculation, For each feature in the subsequence, dynamically calculate the weight of the feature according to its causal role in the causal relationship graph, the time series change trend, and the information gain rate. The specific calculation is as follows.

[0053] (1)Calculation of causal centrality index, First, measure the influence range of feature as a cause on other features , where represents the number of edges pointing from feature to other features, is the total number of nodes. Secondly, consider the causal path length and strength from feature to the prediction target , and the formula is , where represents the set of all directed paths from feature to the prediction target , is the product of the weights on path , is the length of path , is the attenuation coefficient. Combining the above indexes, calculate the causal centrality of feature , where is the weight coefficient, and .

[0054] (2)Calculation of time series change trend coefficient, First, calculate the average change rate of feature at adjacent time points within the subsequence, where is the value of feature at time , is the time interval. Secondly, introduce the time decay mechanism to make recent data have higher weights , where is the attenuation coefficient, is feature Mean within the subsequence. Comprehensive calculation of the change trend coefficient , where Weight coefficient, and .

[0055] (3) Information gain ratio calculation, Measure the feature For the prediction target Information contribution of , where Is the feature And the prediction target Mutual information of, Is the feature Entropy of, the formula is .

[0056] (4) Dynamic weight integration, Integrate the above indicators into the final weight of the feature , where , , And Are adjustment coefficients, and , which can be optimized according to the characteristics of the actual data.

[0057] The dynamic weight calculation method proposed in this patent, by introducing the causal centrality index, quantifies the core position of features in the causal relationship network from the structural level; combines the time series change trend coefficient to capture the impact of the dynamic changes of features in the time dimension on prediction; introduces the information gain ratio to evaluate the information contribution of features to the prediction target. The combination of the three breaks through the limitation of traditional methods that only consider a single dimension and can more accurately identify key causal features and their importance changes at different time points.

[0058] 5. Local fitting and explanation generation, (1) Similar subsequence retrieval, During the prediction process, for the given input subsequence , the weighted Euclidean distance is used to calculate the similarity between the input subsequence and each subsequence in the training data. The formula is as follows: , Among them, Is the dynamic weight of the feature in step 4 . By introducing the dynamic weight, the similarity calculation pays more attention to the causal key features and the dimensions with obvious time series trends.

[0059] (2) Weighted linear regression fitting, Select several subsequences with the smallest distance as the set of similar subsequences, and construct a weighted linear regression model for the selected set of similar subsequences: , where is the regression number of the feature , which is solved by minimizing the weighted mean square error: .

[0060] Combined with the causal relationship graph constructed in step 3, if the feature is the direct cause of the prediction target (there is a directed edge ), then force its regression coefficient to be positive (positive causal relationship), so as to prevent the model from learning parameters that violate causal logic.

[0061] (3) Dynamic interpretation generation mechanism, Based on the regression coefficient and the dynamic weight , generate a two-dimensional time-feature heat map. The horizontal axis is the time step , the vertical axis is the feature , and the pixel value is , highlighting the "feature-time" combination that has the greatest impact on the prediction result.

[0062] Example 2 This example provides a multivariate time series prediction and interpretation system based on information theory and causal reasoning, including: A data acquisition module, configured to acquire original multivariate time series data; A preprocessing module, configured to perform data preprocessing on the acquired original multivariate time series data; An information reasoning module, configured to perform causal information reasoning on the preprocessed multivariate time series data; A weight module, configured to calculate the dynamic feature weights of the multivariate time series data based on causal information reasoning; A fitting module, configured to perform local fitting and interpretation generation based on the dynamic feature weights; An output module, configured to obtain a two-dimensional time-feature heat map.

[0063] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device to perform the described multivariate time series prediction and interpretation method based on information theory and causal reasoning.

[0064] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to perform the described multivariate time series prediction and interpretation method based on information theory and causal reasoning.

[0065] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A multivariate time series prediction explanation method based on information theory and causal reasoning, characterized in that Including: Obtain the original multivariate time series data; Perform data preprocessing on the obtained original multivariate time series data; Perform causal information inference on the preprocessed multivariate time series data; Based on the causal information inference, calculate the dynamic feature weights of the multivariate time series data; Perform local fitting and interpretation generation based on the dynamic feature weights; Obtain the time-feature two-dimensional heat map.

2. The multivariate time series prediction and interpretation method based on information theory and causal reasoning according to claim 1, characterized in that The data preprocessing of the obtained original multivariate time series data includes the original multivariate time series data is normalized, where D represents the number of features, T represents the number of time observations, is the feature d at time t data point; for each feature dimension d , it is normalized by the formula , where and are the mean and standard deviation of the feature d in the time series respectively; meanwhile, the time series is divided into subsequences of length , and the s -th subsequence is denoted as .

3. A multivariate time series prediction and interpretation method based on information theory and causal inference according to claim 2, characterized in that Performing causal information inference on the preprocessed multi-variable time series data includes, for each subsequence, calculating the mutual information between different features. For feature and feature , calculating the joint distribution probability and marginal distribution probability , , and then calculating the mutual information value. By calculating the mutual information values of every two features respectively, a mutual information matrix is constructed to record the mutual influence between each feature dimension. The calculation of the mutual information value is expressed as: 。 4. A multivariable time series prediction explanation method based on information theory and causal reasoning according to claim 3, characterized in that, Performing causal information inference on the preprocessed multivariate time series data further includes, based on the calculated mutual information matrix, using a constraint-based causal discovery algorithm to mine the causal relationships between features. Among them, a complete undirected graph is constructed. , where the nodes in the graph correspond to the respective feature dimensions of the multivariate time series, and the nodes are connected by undirected edges, indicating that there may be associations between the features in the initial state. According to the chronological order of the time series, time constraint marks are added to each edge to record the chronological order information of the occurrence of the association between two features; denote the size of the current conditional set used in the iterative process as , set the maximum conditional set size , starting from , for each edge in the graph, under the condition of different conditional sets composed of other nodes, perform conditional independence tests.

5. A method for predicting and interpreting multivariate time series based on information theory and causal reasoning according to claim 4, characterized in that Performing causal information inference on the preprocessed multi-variable time series data further includes, in the orientation stage, if there is time constraint information indicating that the change of feature always precedes the change of feature , and is an undirected edge, and in the case where other orientation rules do not conflict, the edge is preferentially oriented as i→j, and the orientation rule is repeatedly applied until all edges are oriented, and finally a causal relationship graph is constructed.

6. A multivariate time series prediction explanation method based on information theory and causal inference according to claim 5, characterized in that The above-mentioned causal information-based reasoning calculates the dynamic feature weights of multivariate time series data, including calculating the causal centrality index, the time series change trend coefficient, and the information gain rate of the features of multivariate time series data respectively, and integrating them into the final weights. , where , and are adjustment coefficients, and . First, measure the influence range of feature as a cause on other features , where represents the number of edges pointing from feature to other features, is the total number of nodes. Secondly, consider the causal path length and strength from feature to the prediction target . The formula is , where represents the set of all directed paths from feature to the prediction target , is the product of the weights on path , is the length of path , is the attenuation coefficient. Combining the above indicators, calculate the causal centrality of feature , where is the weight coefficient, and .

7. A method for predicting and interpreting multivariate time series based on information theory and causal reasoning according to claim 6, characterized in that The above-mentioned causal information-based reasoning for calculating the dynamic feature weights of multivariate time series data further includes the calculation of the time series change trend coefficient. Among them, first, calculate the feature average rate of change of adjacent time points within the subsequence , where is the feature at time value, is the time interval. Secondly, introduce a time decay mechanism to make recent data have higher weights , where is the decay coefficient, is the mean value of the feature within the subsequence. Comprehensively calculate the change trend coefficient , where weight coefficient, and .

8. A method for explaining multivariate time series prediction based on information theory and causal reasoning according to claim 7, characterized in that, The local fitting and interpretation generation based on dynamic feature weights includes, during the prediction process, for a given input subsequence , calculating the similarity between the input subsequence and each subsequence in the training data using the weighted Euclidean distance, expressed as: , Among them, is the dynamic weight of the feature By introducing the dynamic weight, the similarity calculation focuses on the causal key features and the dimensions with obvious temporal trends.

9. A method for predicting and interpreting multivariate time series based on information theory and causal reasoning according to claim 8, characterized in that, The local fitting and interpretation generation based on dynamic feature weights further includes selecting several subsequences with the smallest distance as the set of similar subsequences, and constructing a weighted linear regression model for the selected set of similar subsequences: , where is the regression number of the feature , and is solved by minimizing the weighted mean square error: .

10. A multivariate time series prediction and interpretation system based on information theory and causal inference, characterized in that, Including: A data acquisition module configured to obtain the original multivariate time series data; A preprocessing module configured to perform data preprocessing on the obtained original multivariate time series data; An information inference module configured to perform causal information inference on the preprocessed multivariate time series data; A weight module configured to calculate the dynamic feature weights of the multivariate time series data based on the causal information inference; A fitting module configured to perform local fitting and interpretation generation based on the dynamic feature weights; An output module configured to obtain the time-feature two-dimensional heat map.

Citation Information

Patent Citations

  • Analysis and prediction method for high-dimensional time series data based on feature extraction

    CN108399434A

  • Networked data prediction method based on causal Transform

    CN116777068A

  • Multi-dimensional time sequence interpretable prediction method based on multi-attention collaborative network

    CN117764178A

  • Time sequence prediction system based on causal relationship

    CN118607689A

  • Transform-based method and device for constructing sample efficient world model

    CN119271974A

Cited By

  • Geological disaster prediction method and device integrating space-time sequence analysis and causal reasoning

    CN121279530A