Conditional granger causality analysis method based on variable selection and reverse time lag feature selection for complex systems such as meteorology

By employing conditional Granger causality analysis with variable selection and inverse time lag feature selection, the problem of identifying variable relationships in high-dimensional time series was solved, enabling efficient causal relationship exploration and dynamic information display, and improving the accuracy and efficiency of meteorological data forecasting.

CN116881798BActive Publication Date: 2026-03-24DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional Granger causality analysis methods suffer from structural limitations in high-dimensional time series, making it difficult to accurately identify relationships between variables. This leads to increased difficulty in model prediction and longer training time, thus affecting prediction accuracy.

Method used

A conditional Granger causal analysis method based on variable selection and inverse time lag feature selection is adopted. Through data preprocessing, outlier detection using the isolated forest algorithm, selection of conditional variables using the maximum correlation minimum redundancy algorithm, and selection of the optimal explanatory vector by combining the Bayesian information criterion, a conditional Granger causal analysis model is established and a causal relationship heatmap is drawn.

Benefits of technology

It enables quantitative causal relationship analysis of high-dimensional meteorological data, improves prediction accuracy, provides detailed information on dynamic causal relationships between variables, and reduces the complexity of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881798B_ABST
    Figure CN116881798B_ABST
Patent Text Reader

Abstract

The application provides a conditional Granger causality analysis method based on variable selection and reverse time lag feature selection for a complex system such as weather, and belongs to the technical field of data mining.The collected data is pretreated, missing data is supplemented by an average value interpolation method, and the data is subjected to stationarity inspection and treatment to meet the assumption of model establishment. Then, the data is normalized to eliminate the influence of different variable dimensions. Finally, a conditional Granger causality analysis model based on variable selection and reverse time lag feature selection is established to realize the purpose of accurately exploring the causality between variables, and the causality index between different variables is displayed to quantitatively and accurately analyze the causality between variables in the system. The application can realize accurate modeling of a complex system, and aims to expand the original method to be applicable to high dimensions and display more dynamic information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data mining, and relates to a conditional Granger causality analysis method based on variable selection and reverse time lag feature selection applied to a complex system such as meteorology, and aims to explore the relationship between variables in high-dimensional data in the field of meteorology. BACKGROUND

[0002] Multivariate time series exist widely in many fields, such as industry, meteorology, medicine and finance. Time series analysis is an analysis method for predicting time series data by using mathematical statistics, aiming to reveal the development law in the future, and is one of important tools for studying system mechanism and system modeling. For example, in the field of meteorological research, in recent years, with the rapid development of China's industry and the increase of transportation tools, the concentration of harmful substances in the atmosphere due to the combustion of energy materials has increased significantly, causing air quality to decline and the occurrence of haze weather and other phenomena. Air pollutants not only cause poor atmospheric visibility, thereby causing environmental problems such as acid rain, but also are inhaled into the body, thereby causing harm to human health.

[0003] At present, meteorological time series data presents the characteristics of multi-dimension and large scale. The increase of data dimension leads to the mutual relationship between variables becoming more complex and variable, so that there are irrelevant variables and abnormal variables with adverse effects on the prediction target in the data. These irrelevant variables and abnormal variables not only increase the difficulty of establishing a prediction model, but also prolong the training time of the system, which has an adverse effect on the prediction of the model. Therefore, for air pollutants, it is important to establish a model, depict the causal relationship between air quality indicators, and realize the prediction of the possible phenomenon in the next step according to the current and historical time state of each variable, thereby providing theoretical support for air pollution prevention and control. Therefore, it is of great practical significance to establish an effective multivariate causal analysis model.

[0004] At present, there are many methods that can identify the mutual relationship between variables in a multivariate system, and causal relationship analysis is one of the most widely used and most perfect methods. Causal relationship analysis can not only effectively identify variables without causal relationship and redundant variables in a multivariate time series, but also can well explain the causal relationship between variables in the modeling and prediction analysis of a multivariate time series. Thus, the purpose of accurate establishment of a prediction model and improvement of prediction accuracy is achieved.

[0005] Granger causality model is an analysis method for revealing the correlation between time series variables by establishing a linear autoregressive model (Vector Autoregressive, VAR). Since its inception, it has attracted the attention of many scholars. However, due to its applicability to only bivariate linear causal relationship analysis, it has great limitations in multivariate time series analysis. Therefore, in order to solve the above limitations, domestic and foreign scholars have proposed a large number of improved models. In the establishment of the model of multivariate time series and the analysis of causal relationship, Geweke et al. introduced the concept of conditional variables in the paper "Geweke J. Measurement of linear dependence and feedback between multiple timeseries[J]. Journal of the American statistical association, 1982, 77(378): 304-313." to establish a conditional Granger causality model, which can effectively identify the direct and indirect causal relationship between variables. However, due to the conditional Granger causality in the establishment of the model has a very large number of parameter estimation and calculation, so the estimation of large-scale multivariate data set exists considerable difficulty. Dimitris et al. in the paper "Siggiridou E, Kugiumtzis D. Granger causality in multivariate time series using a time-ordered restricted vector autoregressive model[J]. IEEE Transactions on Signal Processing, 2016, 64(7): 1759-1773." proposed a scheme to combine the selection of lag variables with the conditional Granger causality index (CGCI) by using the recently developed reverse time lag variable selection method, and then realize accurate prediction modeling. These methods extend the classic Granger causality analysis to multivariate time series and are suitable for causal relationship exploration of high-dimensional data.

[0006] The present application is aimed at the causal relationship analysis problem of complex high-dimensional system, and proposes a conditional Granger causality analysis method based on variable selection and reverse time lag feature selection, which is used for causal analysis modeling between variables in the research field of meteorological pollution and the like. Compared with the traditional prediction analysis model, the prediction accuracy of the model of the present application is higher. SUMMARY

[0007] The technical problem solved by the present application is that the traditional Granger causality analysis method cannot be applied to high-dimensional time series due to structural limitations, and cannot accurately identify the variable relationship between variables in the process. The original Granger causality analysis model is expanded, a conditional Granger causality analysis method based on variable selection and reverse time lag feature selection is proposed, accurate causal relationship exploration of high-dimensional data is realized, and real-time information in the process is displayed. Further, precise modeling of complex systems is realized. The purpose of the method is to expand the original method to be applicable to high-dimensional and to display more dynamic information.

[0008] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0009] A conditional Granger causality analysis method based on variable selection and reverse time lag feature selection, the present application faces complex systems such as meteorology, and proposes a causal analysis method to explore the causal relationship between each variable in the system and the main pollutant AQI. First, the collected data is preprocessed, and the missing data is completed by using the average value interpolation method. Then, the completed data is tested and processed for stationarity to meet the assumptions for model establishment. After that, the data is normalized to eliminate the influence of different variable dimensions. Finally, a conditional Granger causality analysis method based on variable selection and reverse time lag feature selection is established to accurately explore the causal relationship between variables, and to display the dynamic causal relationship index between different variables, so as to quantitatively and clearly analyze the relationship between each variable in the system and AQI, and realize the analysis of the influencing factors of AQI. The specific steps are as follows:

[0010] Step 1: Obtain air quality index observation data, pre-process multi-dimensional time series data, use unit root test method to test the stationarity of time series data, and then normalize the time series data.

[0011] Step 2: Adopt isolated forest algorithm to detect data outliers, and perform rejection processing according to the score height. For time series represents an N*n dimensional matrix, there are n variables, and the time series length is N. Assuming that there are Z outliers in the time series, the time series processed by the isolated forest algorithm becomes

[0012]

[0013] In the formula, H(k)=ln(k)+ζ,ζ=0.5772156649(Euler constant). Wherein, S(x,n) is the abnormal index of the tree composed of the training data of x samples, S(x,n) takes the value range [0,1], the closer to 0, the normal value, the closer to 1, the abnormal value.

[0014] Step 3: Select the maximum correlation and minimum redundancy algorithm as the appropriate conditional variable for the established conditional Granger causality analysis model, combine the maximum correlation and minimum redundancy between variables, calculate a comprehensive score, select the feature with the highest comprehensive score as the next selected feature, and iterate until the time series X * processed by the isolation forest algorithm are selected as the conditional variables, and the target variable x i is defined as the feature subset The comprehensive score calculation method is as follows:

[0015]

[0016] Where m is the number of feature subsets, S m is the feature subset, x j , x k is the feature subset variable.

[0017] Step 4: Select the optimal explanatory vector of the above processed data set (including target variable, source variable, and conditional variable), specifically:

[0018] Step 4.1: First, set the maximum delay number p max , representing the termination of model selection. The Bayesian Information Criterion (BIC) is the variance of the target variable x i . Before selection, the explanatory vector is empty, and maxlag is defined as an auxiliary array to track the delay number of each variable that has been searched, so for each variable, it is initially set to zero.

[0019] Step 4.2: For each variable of the processed data set, increase the delay number by 1, and add this delayed variable to the current explanatory vector Thus forming the candidate explanatory vector. For the first cycle, the current explanatory vector is empty, and the candidate explanatory vector is [x i,t-1 , x j,t-1 , x 1,t-1 ,..., x m,t-1 ]. If the source variable x j is included in the feature subset S, the candidate explanatory vector is [x i,t-1 , x 1,t-1 ,..., x j,t-1 ,..., x m,t-1 ]. For the l+1th cycle, the explanatory vector has l lag variables, and the delay number of each variable is τ(l m+2 ), which is set to zero if it has not appeared. Then the candidate explanatory vector is or The BIC value of an m+2 order dynamic regression model composed of m+2 candidate explanatory vectors is calculated.

[0020] Step 4.3: The current BIC value is compared with the BIC value of the explanatory vector after the next cycle of increasing the explanatory vector. The explanatory vector with the minimum BIC value is selected to update the current explanatory vector set, and then the next step is entered. If the BIC value after adding the candidate explanatory vector is not smaller than the current BIC value, the current explanatory vector is retained, then the delay number is increased by 1, and then the next step is performed.

[0021] Step 4.4: If the delay number of all variables in the data set reaches the maximum delay number p max (τ(l m+2 )=p max ), the termination is performed, otherwise, step 4.1 is entered. After the loop ends, the algorithm gives the optimal explanatory vector of the model.

[0022] Step 5: A conditional Granger causality analysis model is established, the above explanatory vector is brought into the model, the causality values between variables are obtained, and a causality heat map between variables is drawn by drawing the results of the conditional Granger causality analysis;

[0023] Step 6: According to the causality heat map, the variable with the largest correlation with the target variable is selected. The above correlation variable is brought into the echo state network model to obtain the comparison between the predicted value and the actual value of the target variable and the prediction error. The advantages and disadvantages of the comparison method and the method of the application are compared by comparing the prediction error values.

[0024] Further, the conditional Granger causality analysis model further explores the dynamic characteristics between variables on the basis of the causality analysis results between variables; the conditional Granger causality analysis model includes two VAR models, also known as dynamic regression models:

[0025] The first VAR model is an unrestricted model, which is referred to as an unrestricted model, and the model includes the optimal explanatory vector of all lag variables, which is represented as:

[0026]

[0027] In the formula, a i,k is the parameter of the unrestricted model; m+2 is the model order defining the time lag; u i,k is the prediction error of the unrestricted model.

[0028] The second VAR model is a restricted model, which is obtained by excluding the source variable in the explanatory variable, and is represented as:

[0029]

[0030] where b i,k is the parameter of the restricted model; e i,t is the prediction error of the restricted model.

[0031] Definition of conditional Granger causality index (CGCI) between variables i,j:

[0032]

[0033] where, is the residual variance obtained according to u i,t and e i,t .

[0034] Compared with the prior art, the present application has the following beneficial effects:

[0035] The present application aims at quantitative and accurate analysis of causality between high-dimensional data variables. Firstly, the present application pre-processes its original data, and then establishes a causality analysis algorithm based on variable selection and feature selection method. This algorithm can be distinguished from general causality analysis method based on vector autoregressive framework, and realizes causality analysis of high-dimensional data. Then, the optimal explanatory vector is sent into a conditional Granger causality analysis model to obtain quantitative causality value between variables. In addition, the conditional Granger causality analysis method also obtains a causality heat map between variables, which provides more information for exploring the causality between variables. That is, the present application not only obtains quantitative causality measurement for high-dimensional data, but also obtains causality between variables. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 The Granger causality graph is shown, the node represents a vector, and the arrow represents that there is a causality relationship between vectors;

[0037] Figure 2 The flowchart of the present application is shown;

[0038] Figure 3 The dynamic relationship graph between each variable is shown; Figure 3 (a) is a variable AQI change over time graph; Figure 3 (b) is a variable PM2.5 change over time graph; Figure 3 (c) is a variable PM10 change over time graph; Figure 3 (d) is a variable SO2 change over time graph; Figure 3 (e) is a variable CO change over time graph; Figure 3 (f) is a variable NO2 change over time graph; Figure 3 (g) is a variable O3 change over time graph;

[0039] Figure 4 a causal graph between variables for each method variable; Figure 4 (a) a causal graph between variables for FBLG method; Figure 4 (b) a causal graph between variables for LassoGC method; Figure 4 (c) a causal graph between variables for CopulaGC method; Figure 4 (d) a causal graph between variables for the method of the present application;

[0040] Figure 5(a1) is a comparison chart of target value and predicted value of AQI for FBLG; Figure 5(a2) is a prediction error chart of AQI for FBLG;

[0041] Figure 5(b1) is a comparison chart of target value and predicted value of AQI for LassoGC; Figure 5(b2) is a prediction error chart of AQI for LassoGC;

[0042] Figure 5(c1) is a comparison chart of target value and predicted value of AQI for CopulaGC; Figure 5(c2) is a prediction error chart of AQI for CopulaGC;

[0043] Figure 5(d1) is a comparison chart of target value and predicted value of AQI for the method of the present application; Figure 5(d2) is a prediction error chart of AQI for the method of the present application. DETAILED DESCRIPTION

[0044] The present application will be further described in detail below in combination with specific examples and simulation charts.

[0045] The hardware equipment used in the present application includes one PC machine.

[0046] Figure 2 The conditional Granger causal analysis method flowchart based on variable selection and reverse time lag feature selection provided by the present application specifically includes the following steps:

[0047] Step 1: This subsection selects the Shanghai meteorological data set for simulation experiment. The data set records the daily data from December 2, 2013 to April 13, 2021, a total of 2690 samples, each sample contains a 7-dimensional time series, and the variable meaning number is described in detail in Table 1. Then the missing values of the multi-dimensional AQI and meteorological data set are interpolated and the outliers are analyzed and processed; the unit root test method is used to test the stationarity of the data, and the data is once-differenced and stationary according to the test result; the time series data is normalized;

[0048] Table 1 Shanghai meteorological data number correspondence table

[0049]

[0050] Step 2: Isolation Forest algorithm is used to detect data outliers, and the outliers are removed according to the score. For time series The representative N x n matrix, a total of n-dimensional variables, and the time series length is N. Assuming that there are Z outliers in the time series, the time series processed by the Isolation Forest algorithm becomes Here, the proportion of outliers removed is 10%, and the number of samples after processing is 2421.

[0051]

[0052] In the formula, H(k) = ln(k) + ζ, ζ = 0.5772156649 (Euler's constant). S(x, n) is the abnormal index of the tree formed by the training data recording x samples, and S(x, n) takes a value in the range [0, 1]. The closer to 0, the normal value, the closer to 1, the abnormal value.

[0053] Step 3: The maximum correlation and minimum redundancy algorithm is used to select appropriate conditional variables for the established conditional Granger causality analysis model. The conditional variable selection method is to select 30% of all variables, that is, for the experimental data set with 7 variables, the corresponding conditional variables are 2. The maximum correlation and minimum redundancy between variables are combined to calculate a comprehensive score, and the feature with the highest comprehensive score is selected as the next selected feature. Iterative calculation is performed until the time series X * is processed by the Isolation Forest algorithm, and m variables are selected as the conditional variables of the target variable x i , defined as the feature subset of x The comprehensive score calculation method is:

[0054]

[0055] In the formula, m is the number of feature subsets, S m is the feature subset, x j , and x k is the feature subset variable.

[0056] Step 4: Select the optimal explanatory vector of the data set after the above step processing (including the target variable, source variable, and conditional variable). Specifically:

[0057] Step 4.1: First, set the maximum delay number p max , which represents the model selection termination. The Bayesian Information Criterion (BIC) value is the variance of the target variable x i . Before starting selection, the explanatory vector is empty, and maxlag is defined as an auxiliary array used to track the delay number of each variable that has been searched. Therefore, for each variable, it is initially set to zero.

[0058] Step 4.2: For each variable of the processed data set, increase its delay number by 1, and add this delayed variable to the current explanatory vector Thus the candidate explanatory vector is formed. For the first cycle, the current explanatory vector is empty, and the candidate explanatory vector is [x i,t-1 , x j,t-1 , x 1,t-1 ,..., x m,t-1 ]. If the source variable x j is included in the feature subset S, the candidate explanatory vector is [x i,t-1 , x 1,t-1 ,..., x j,t-1 ,..., x m,t-1 ]. For the (l+1)th cycle, the explanatory vector has l lagged variables, and let the delay number of each variable be τ(l m+2 ). If it has not appeared, set its value to zero. Then the candidate explanatory vector is or Calculate the BIC value of the m+2th order dynamic regression model composed of the m+2 candidate explanatory vectors.

[0059] Step 4.3: Compare the current BIC value with the BIC value of the explanatory vector with increased explanatory vector in the next cycle. Select the explanatory vector with the minimum BIC value to update the current explanatory vector set, and then go to the next step. If the BIC value of the candidate explanatory vector after increasing does not have a smaller value than the current BIC value, keep the current explanatory vector, then increase the delay number by 1, and then go to the next step.

[0060] Step 4.4: If the delay number of all variables in the data set reaches the maximum delay number p max (τ(l m+2 ) = p max ), terminate, otherwise go to step 4.1. After the loop ends, the algorithm gives the optimal explanatory vector of the model.

[0061] Step 5: Establish a conditional Granger causality analysis model, bring the above explanatory vector into the model, obtain the quantitative causal relationship values between variables, draw the conditional Granger causality analysis result to obtain the causal relationship thermodynamic map between variables;

[0062] Step 6: Select the variable with the strongest correlation to the target variable based on the causal relationship heatmap. In this experiment, AQI was selected as the target variable. FBLG selected 3 (PM10) as the relevant variable, LassoGC selected 3 and 5 (PM10, CO) as the relevant variables, CopulaGC selected 7 (O3) as the relevant variable, and FBLG selected 2, 3, 5, and 6 (PM2.5, PM10, CO, NO2) as the relevant variables. These relevant variables were then substituted into the echo state network model to obtain a comparison between the predicted and actual values ​​of the target variable, as well as the prediction error. The advantages and disadvantages of the previous method were compared with the method of this invention by analyzing the prediction error values.

[0063] By plotting the time series of the dataset, we can obtain the following: Figure 3 The dynamic causal relationships between other variables and AQI are shown, providing dynamic information about the time distribution of these variables. Furthermore, averaging the time series results yields quantitative causal values ​​between the variables, providing a quantitative indicator for subsequent determination of causal relationships. Figure 3 The provided dynamic curves illustrate the dynamic relationships between variables over time. For studies focusing on a specific period, more detailed information from the curves is needed. Based on the results of causal analysis, the most relevant variables to the target variable are identified. The causal analysis heatmaps for each method are shown below. Figure 4 As shown. Then, an echo state network is used to build a prediction model to test the causal relationship. The reservoir parameters of the echo state network are set as follows: reservoir dimension 40, sparsity 0.05, spectral radius 0.99, and connection input weight 0.05. A total of 10 independent experiments are conducted, and the mean of the 10 experiments is used as the final result to eliminate the influence of error. The prediction indicators are root mean square error (RMSE), standard root mean square error (NRMSE), mean absolute percentage error (MAPE), and symmetric mean absolute percentage error (SMAPE). The definitions of the four are as follows:

[0064]

[0065]

[0066]

[0067]

[0068] In the formula, x i and These represent the actual value and the predicted value, respectively, and n represents the number of samples.

[0069] The prediction results are shown in Figure 5. As can be seen from the figure, the subset of variables selected by the method of this invention (FS-BTSCG) enables the predicted AQI results to accurately reflect the true trend of change. The errors of 10 independent repeated experiments are shown in Table 2. Further comparison with the forward and backward Lasso Granger causality analysis method (FBLG), the Lasso Granger method (LassoGC), and the Copula Granger method (CopulaGC) shows that the error indices obtained by this invention achieve the best results, further demonstrating the effectiveness of this invention.

[0070] Table 2 Prediction results of various methods

[0071]

[0072] The examples described above are merely illustrative of the embodiments of the invention and should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make corresponding improvements without departing from the concept of the invention, and these improvements all fall within the protection scope of the invention.

Claims

1. A conditional Granger causal analysis method for complex systems such as meteorology, based on variable selection and inverse time-delay characteristic selection, wherein the conditional Granger causal analysis method, for meteorological or other complex systems, proposes a method for exploring the causal relationship between variables in the system and their main pollutant AQI; characterized in that, Includes the following steps: First, the collected data is preprocessed, and missing data is filled in using the average imputation method; Then, the completed data undergoes stationarity testing and processing to meet the assumptions of the model. Next, the data is normalized to eliminate the influence of different variable units. Finally, a conditional Granger causality analysis method based on variable selection and inverse time lag characteristics is established to accurately explore the causal relationships between variables, while simultaneously displaying dynamic causal relationship indices between different variables. This allows for a quantitative and explicit analysis of the relationship between various variables in the system and the AQI, enabling the analysis of factors influencing the AQI. The specific steps are as follows: Step 1: Obtain air quality index observation data, preprocess the multidimensional time series data, perform stationarity test on the time series data, and then normalize the time series data; Step 2: Use the Isolation Forest algorithm to detect outliers and remove them according to their scores; for time series data... ,represent Given an n-dimensional matrix with n variables and a time series of length N; assuming the time series has Z outliers, the time series processed by the Isolation Forest algorithm becomes... ; ; In the formula, , (Euler's constant); where, It records the anomaly index of the tree formed by the training data of x samples. The range of values ​​is The closer the value is to 0, the more normal it is; the closer the value is to 1, the more abnormal it is. Step 3: Select appropriate condition variables for the established conditional Granger causal analysis model using the maximum correlation and minimum redundancy algorithm. Combine the maximum correlation and minimum redundancy among the variables to calculate a comprehensive score. Select the feature with the highest comprehensive score as the next feature to be selected. Iterate until the time series processed by the Isolation Forest algorithm is reached. Select m variables as condition variables and define them as target variables. Feature subset The comprehensive score is calculated as follows: ; In the formula, m is the number of feature subsets. For feature subset, , For feature subset variables; Step 4: Select the optimal explanatory vector for the dataset after the above steps, where the dataset includes the target variable, source variable, and condition variable; Step 5: Establish a conditional Granger causal analysis model, substitute the above explanatory vectors into the model to obtain the causal relationship values ​​between variables, and draw the results of the conditional Granger causal analysis to obtain a causal relationship heatmap between variables; Step 6: Select the variable with the highest correlation to the target variable based on the causal relationship heatmap; substitute the above-mentioned correlated variables into the echo state network model to obtain the comparison between the predicted value and the actual value of the target variable and the prediction error; compare the advantages and disadvantages of the method and the method of the present invention by comparing the prediction error values.

2. The conditional Granger causal analysis method for complex systems such as meteorology based on variable selection and reverse time-delay characteristic selection as described in claim 1, characterized in that, Step 4 is described in detail below: Step 4.1: First, set the maximum delay number p. max The selection of the model terminates; the Bayesian Information Criterion (BIC) value is the target variable. The variance; before starting the selection, interpret the vectors. If empty, maxlag is defined as an auxiliary array to track the delay for each variable that has been searched; therefore, it is initially set to zero for each variable. Step 4.2: For each variable in the processed dataset, increment its delay number by 1 and add this delay variable to the current explanatory vector. This forms the candidate explanatory vector; for the first period, the current explanatory vector is empty, and the candidate explanatory vector is... If the source variable Included in the feature subset In the middle, the candidate explanatory vector is ; For the (l+1)th loop, the explanatory vector There are l lagged variables, and the delay for each variable is 1. If it has not yet appeared, its value is set to zero; then the candidate explanatory vector is... or ; Calculation by candidate Composed of candidate explanatory vectors BIC value of the dynamic regression model; Step 4.3: Compare the current BIC value with the BIC value of the explanatory vector to be added in the next loop; select the explanatory vector with the smallest BIC value to update the current explanatory vector group, and then proceed to the next step; if the BIC value after adding candidate explanatory vectors is not smaller than the current BIC value, retain the current explanatory vector, then increase the delay number by 1, and then proceed to the next step; Step 4.4: If the delay count of all variables in the dataset reaches the maximum delay count p max ( =p max If the condition is met, the algorithm terminates; otherwise, it proceeds to step 4.

1. After the loop ends, the algorithm provides the optimal interpretation vector for the model.

3. The conditional Granger causal analysis method for complex systems such as meteorology, based on variable selection and inverse time-delay characteristic selection, as described in claim 1, is characterized in that... The conditional Granger causality analysis model, based on the causal analysis results between variables, further explores the dynamic characteristics between variables; the conditional Granger causality analysis model includes two VAR models, also known as dynamic regression models: The first VAR model is an unrestricted model, which includes the optimal explanatory vector for all lagged variables, denoted as: ; In the formula, These are the parameters of the unrestricted model; It is the model order that defines the time delay; It is the prediction error of the unrestricted model; The second VAR model is the restricted model, which is obtained by excluding the source variables from the explanatory variables, and is represented as: ; In the formula, These are the parameters of the restricted model; It represents the prediction error of the restricted model; variable Definition of Conditional Granger Causality Index (CGCI): ; In the formula, It is based on and The resulting residual variance.

Citation Information

Patent Citations

  • Zero-time-lag nonlinear extended Granger causality analysis method

    CN111367959A

  • Air quality prediction method based on depth space-time similarity

    CN113077097A