Time sequence causal analysis method and device for environmental change and infectious disease onset risk, and storage medium
Through the time series causal analysis method, combined with convergent cross-mapping, Peter-Clark instantaneous condition independence test and causal forest algorithm, causal analysis of environmental changes and infectious disease risks is solved, and the problem of lack of stable standards and low efficiency of causal analysis in the existing technology is solved, achieving a more reliable causal explanation.
Patent Information
- Application Number
- CN202411762351.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The prior art lacks stable standards in causal analysis of environmental changes and infectious disease risk, is inefficient and prone to errors.
The time series causal analysis method is used to process environmental data and infectious disease incidence through convergent cross-mapping, Peter-Clark instantaneous condition independence test and causal forest algorithm, and a variety of causal relationship information is comprehensively considered to determine the causal relationship.
Through the complementation and strengthening of multiple algorithms, the causal relationship between environmental changes and infectious disease incidence can be more powerfully explained, providing more reliable answers.
Smart Images

Figure CN119943371A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental health technology, and in particular to a time series causal analysis method, a computer device and a storage medium for environmental change and infectious disease risk. Background Art
[0002] There is a causal relationship between the incidence of infectious diseases and environmental changes. For example, as climate change intensifies, the transmission pattern of infectious diseases may change, and host susceptibility, social factors, and environmental conditions may affect the activity of infectious diseases. In-depth understanding of potential environmental drivers has become an important means of preventing infectious diseases. Through time series causal inference, the causal relationship between environmental changes (such as climate change, land use changes, and social personnel environment changes) and the incidence of infectious diseases can be revealed. This understanding helps to identify which environmental factors are the key drivers of increased infectious disease risks, and use historical data to predict future trends in infectious disease incidence, thereby providing a scientific basis for public health interventions. Summary of the invention
[0003] In view of the current technical problems that causal analysis of environmental changes and infectious disease risks lacks stable standards, is inefficient and prone to errors, the purpose of the present invention is to provide a time series causal analysis method, computer device and storage medium for environmental changes and infectious disease risks.
[0004] On the one hand, an embodiment of the present invention includes a time series causal analysis method for environmental changes and infectious disease risk, the time series causal analysis method for environmental changes and infectious disease risk including the following steps:
[0005] Obtain and process environmental data time series and infectious disease incidence time series;
[0006] Performing convergent cross-mapping processing on the environmental data time series and the infectious disease incidence time series, evaluating the causal impact of candidate driving factors on infectious disease activities, and obtaining first causal relationship information;
[0007] Performing Peter-Clark instantaneous conditional independence test on the environmental data time series and the infectious disease incidence time series to obtain visualized lag and contemporaneous causal dependencies and obtain second causal relationship information;
[0008] Processing the environmental data time series and the infectious disease incidence time series with a causal forest algorithm to quantitatively analyze the linear and nonlinear relationships between environmental factors and infectious disease activities, and obtain third causal relationship information;
[0009] The causal relationship between the environmental data time series and the infectious disease incidence time series is determined based on the first causal relationship information, the second causal relationship information and the third causal relationship information.
[0010] Furthermore, the step of performing convergent cross-mapping processing on the environmental data time series and the infectious disease incidence time series to evaluate the causal impact of candidate driving factors on infectious disease activities and obtain first causal relationship information includes:
[0011] Set embedding dimension and lag time;
[0012] Embedding the environmental data time series into an X phase space;
[0013] Embed the infectious disease incidence time series into the Y phase space;
[0014] Reconstructing the state in the Y phase space based on the embedded neighborhood information in the X phase space, and calculating the Pearson correlation coefficient between the predicted Y and the actually observed Y;
[0015] Confirm the direction and strength of the causal relationship based on the Pearson correlation coefficient;
[0016] When the causality test is passed, a linear model is fitted at each point in the X phase space to obtain the fitting coefficient, and a commonly used heat map or the like can be drawn to show the spatiotemporal distribution of the impact of environmental factors analyzed by S-map on infectious disease activities;
[0017] The Pearson correlation coefficient is used as the first causal relationship information.
[0018] Furthermore, the Peter-Clark instantaneous conditional independence test is performed on the environmental data time series and the infectious disease incidence time series to obtain a visualized lag and contemporaneous causal dependency relationship, and obtain the second causal relationship information, including:
[0019] Set the maximum time lag value based on the information criterion;
[0020] In the PC stage, partial correlation analysis is used as an independence test method within the range of the maximum time lag value, that is, a conditional independence test is performed on each pair of variables X and Y. After controlling the influence of other variables, the partial correlation coefficients of X and Y are calculated to identify potential driving variables (parent nodes) and causal relationships. In the MCI stage, based on the preliminary causal relationship identified in the PC stage, the partial correlation coefficient is further calculated, that is, the partial correlation coefficient is calculated for each lag time of the variable pair X→Y, and the strength of the lag causal relationship is analyzed to obtain the causal strength and significance between the environmental data time series and the infectious disease incidence time series corresponding to multiple different lag times;
[0021] Generating a directed acyclic graph according to each of the causal paths;
[0022] Determining the autocorrelation strength and causal path, strength (number of lags) and significance of each variable based on the directed acyclic graph;
[0023] The partial correlation coefficient is used as the second causal relationship information.
[0024] Further, the determining of the autocorrelation strength and causal path, strength (lag number) and significance of each variable according to the directed acyclic graph includes:
[0025] Visualizing the directed acyclic graph to obtain multiple nodes and edges connecting the nodes;
[0026] Determine the autocorrelation strength of each variable according to the color of the node;
[0027] Determine whether it is a lagged causal relationship based on the curvature of the edge;
[0028] The causal strength is determined according to the color of the edge.
[0029] Furthermore, the performing of causal forest algorithm processing on the environmental data time series and the infectious disease incidence time series to obtain third causal relationship information includes:
[0030] Constructing a causal forest model;
[0031] Training the causal forest model using the environmental data time series and the infectious disease incidence time series;
[0032] Calculate the average treatment effect and the conditional average causal effect using the trained causal forest model;
[0033] The conditional average causal effect is used as the third causal relationship information.
[0034] Furthermore, the construction of the causal forest model includes:
[0035] According to the formula
[0036] Y t =f(X t )+∈
[0037] Establish the causal forest model; where Y t is the data in the time series of the number of infectious diseases; X t is the data in the environmental data time series (including atmospheric pollutants, meteorological factors, holiday effects, day of the week effects, lagged variables, etc.); f is the nonlinear relationship function learned through causal forest; ∈ is the random error term.
[0038] Furthermore, determining the causal relationship between the environmental data time series and the infectious disease incidence time series based on the first causal relationship information, the second causal relationship information, and the third causal relationship information includes:
[0039] performing weighted fusion on the first causal relationship information, the second causal relationship information, and the third causal relationship information;
[0040] According to the result of weighted fusion, the causal relationship between the environmental data time series and the infectious disease incidence time series is determined.
[0041] Further, the weighted fusion of the first causal relationship information, the second causal relationship information and the third causal relationship information includes:
[0042] The causal information (Pearson correlation coefficient, partial correlation coefficient, conditional average causal effect) obtained by the above three methods is standardized / normalized, respectively S CCM (X→Y), S PCMCI+ (X→Y), S CF (X→Y);
[0043] According to the applicability or reliability of each method in a specific scenario, a weight W is assigned. CCM , W PCMCI+ , W CF ;
[0044] Weighting is based on: a. Method characteristics (e.g., CCM is suitable for dynamic nonlinear data, PCMCI+ assumes that the independence test of the condition is valid and can control confounding, and the causal forest assumes that the data is randomly distributed and provides quantitative effects, etc.); b. Significance of causal association (P value, more significant method results are given greater weight); c. Consistency of results: relationships with higher consistency between methods are given greater weight; d. Data characteristics (e.g., when the sample size is small, the causal forest may be weaker, while CCM and PCMCI+ are more robust);
[0045] Overall causal strength = S CCM (X→Y)·W CCM +S PCMCI+ (X→Y)·W PCMCI+ +S CF (X→Y)·W CF
[0046] The comprehensive causal strength is taken as the result of weighted fusion.
[0047] On the other hand, an embodiment of the present invention also includes a computer device including a memory and a processor, the memory being used to store at least one program, and the processor being used to load at least one program to execute the time series causal analysis method of environmental changes and infectious disease risks in the embodiment.
[0048] On the other hand, an embodiment of the present invention also includes a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the time series causal analysis method of environmental changes and infectious disease risks in the embodiment.
[0049] The beneficial effect of the present invention is that the time series causal analysis method of environmental changes and infectious disease risks in the embodiment can use the convergent cross-mapping algorithm, the Peter-Clark instantaneous conditional independence test algorithm and the causal forest algorithm to process the environmental data time series and the infectious disease incidence time series respectively, and the obtained first causal relationship information, the second causal relationship information and the third causal relationship information respectively represent the causal relationship between the environmental data time series and the infectious disease incidence time series from different aspects. Finally, the causal relationship between the environmental data time series and the infectious disease incidence time series is determined according to the first causal relationship information, the second causal relationship information and the third causal relationship information, which can make multiple causal relationship algorithms complement each other and strengthen the causal relationship identified by each algorithm. It can more effectively provide a causal explanation for the environmental changes represented by the environmental data time series and the infectious disease incidence time series, so as to provide a more reliable answer to the driving effect of environmental factors on infectious diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the steps of the time series causal analysis method of environmental change and infectious disease risk in the embodiment;
[0051] Figure 2 Schematic diagram of the process of time series causal analysis of environmental change and infectious disease risk in the embodiment. DETAILED DESCRIPTION
[0052] For causal inference assessments of environmental changes and infectious disease incidence, if the focus is on linear models or complex nonlinear models, potential temporal confounding factors may not be adequately controlled. For example, if CCM and PCMCI+ are used but no linear relationship or specific functional form is assumed, the results are usually interpreted in terms of standardized effect measures (such as correlation coefficients or partial correlation coefficients), which emphasize the existence and direction of causal relationships, and it is impossible to derive regression coefficients or risk ratios to quantitatively assess changes in the incidence of infectious diseases caused by changes in environmental factors.
[0053] Causal forest is a nonparametric method based on random forest, which aims to estimate treatment effects and capture potential nonlinear relationships. Causal forest can flexibly handle high-dimensional data, especially suitable for situations with complex interactions and nonlinear effects. Its core idea is to estimate the conditional average causal effect (CATE) by building multiple decision trees.
[0054] Therefore, a dynamic data comprehensive analysis method for time series causal inference of environmental change and infectious disease risk ("CCM"-"PCMCI+"-"CF") can be constructed to solve the problems of causal structure inference and quantitative estimation of impact coefficients. The combination of the three methods can cross-validate the results and provide a more powerful causal explanation, in order to provide a more reliable answer to the driving effect of environmental factors on infectious diseases.
[0055] Based on the above principles, in this embodiment, a time series causal analysis method for environmental changes and infectious disease risks is provided.
[0056] Reference Figure 1 The time series causal analysis method of environmental change and infectious disease risk includes the following steps:
[0057] S1. Obtain time series of environmental data and infectious disease incidence;
[0058] S2. Perform convergent cross-mapping processing on the environmental data time series and the infectious disease incidence time series to evaluate the causal impact of candidate driving factors on infectious disease activities and obtain the first causal relationship information;
[0059] S3. Perform Peter-Clark instantaneous conditional independence test on the environmental data time series and the infectious disease incidence time series to obtain the visualized lag and contemporaneous causal dependency and obtain the second causal relationship information;
[0060] S4. Use the causal forest algorithm to process the environmental data time series and the infectious disease incidence time series, quantitatively analyze the linear and nonlinear relationship between environmental factors and infectious disease activities, and obtain the third causal relationship information;
[0061] S5. Determine the causal relationship between the environmental data time series and the infectious disease incidence time series based on the first causal relationship information, the second causal relationship information and the third causal relationship information.
[0062] In this embodiment, the process of steps S1-S5 is as follows Figure 2 shown.
[0063] In step S1, the environmental data time series to be acquired is data related to the natural environment or social environment, and is arranged according to the time of collection or occurrence. The time unit of the environmental data time series and the infectious disease incidence time series can be daily, weekly, etc.
[0064] For example, we can collect statistics on meteorological data (data related to the natural environment) within a unit time (for example, daily or weekly), pollutant data (data related to the natural environment) within a unit time, and personnel flow (data related to the social environment) within a unit time to obtain a time series of environmental data.
[0065] Taking meteorological data as an example, when executing step S1, meteorological data such as average temperature, relative humidity, rainfall, wind speed, etc. per unit time (for example, daily or weekly) in a regional range can be obtained from data sources such as the National Meteorological Data Center or public atmospheric reanalysis data sets (for example, the European Center for Medium-Term Weather Forecasts reanalysis data set, the U.S. National Center for Environmental Prediction reanalysis data set), thereby forming an environmental data time series composed of multiple average temperatures, an environmental data time series composed of multiple relative humidities, an environmental data time series composed of multiple rainfalls, and an environmental data time series composed of multiple wind speeds.
[0066] Similarly, in step S1, the infectious disease incidence time series is also a time series composed of the incidence of a certain infectious disease in multiple unit times (which can be data in the form of the number of cases or incidence rate, etc.). Specifically, the infectious disease case reporting data per unit time can be obtained through the national or regional Center for Disease Control and Prevention, or the local infectious disease sentinel monitoring system can collect infectious disease etiology monitoring data conducted by sentinel hospitals, and the outpatient and inpatient data per unit time (or medical record homepage data) can be obtained through the regional medical institutions, and then uniformly converted into the infectious disease incidence time series.
[0067] In step S2, the environmental data time series and infectious disease incidence time series obtained in step S1 are processed using the convergent cross-mapping (CCM) method based on state space reconstruction (SSR) to obtain the first causal relationship information obtained by the convergent cross-mapping process. The first causal relationship information can screen the preliminary causal relationship, and the obtained correlation coefficient represents the causal impact of the candidate driving factor (environmental data time series) on the infectious disease activity (infectious disease incidence time series).
[0068] In step S3, the environmental data time series and the infectious disease incidence time series obtained in step S1 are processed using the Peter-Clark Momentary Conditional Independence plus (PCMCI+) test to obtain the second causal relationship information obtained by the Peter-Clark Momentary Conditional Independence test. The second causal relationship information can represent the causal dependency relationship of the lag and contemporaneous relationship between the environmental data time series and the infectious disease incidence time series.
[0069] In step S4, the causal forest algorithm (CF) is used to process the environmental data time series and the infectious disease incidence time series obtained in step S1 to obtain the third causal relationship information obtained by the causal forest algorithm. The third causal relationship information can quantitatively represent the linear and nonlinear relationship between environmental factors (environmental data time series) and infectious disease activities (infectious disease incidence time series).
[0070] In step S5, the causal relationship between the environmental data time series and the infectious disease incidence time series can be determined based on any one, two or all of the first causal relationship information, the second causal relationship information and the third causal relationship information.
[0071] In this embodiment, by executing steps S1-S5, the convergent cross-mapping algorithm, the Peter-Clark instantaneous conditional independence test and the causal forest algorithm can be used to process the environmental data time series and the infectious disease incidence time series respectively. The obtained first causal relationship information, the second causal relationship information and the third causal relationship information respectively represent the causal relationship between the environmental data time series and the infectious disease incidence time series from different aspects. Finally, the causal relationship between the environmental data time series and the infectious disease incidence time series is determined according to the first causal relationship information, the second causal relationship information and the third causal relationship information. This can make multiple causal relationship algorithms complement each other and strengthen the causal relationship identified by each algorithm. It can more effectively provide a causal explanation for the environmental changes represented by the environmental data time series and the infectious disease incidence time series, so as to provide a more reliable answer to the driving effect of environmental factors on infectious diseases.
[0072] In this embodiment, when executing step S2, that is, performing convergence cross-mapping processing on the environmental data time series and the infectious disease incidence time series to obtain the first causal relationship information, the following steps can be specifically performed:
[0073] S201. Setting embedding dimension and lag time;
[0074] S202. Embed the environmental data time series into the X phase space;
[0075] S203. Embed the time series of the number of infectious disease incidence into the Y phase space;
[0076] S204. Reconstruct the state in the Y phase space based on the embedded neighborhood information in the X phase space, and calculate the Pearson correlation coefficient between the predicted Y and the actually observed Y;
[0077] S205. Confirm the direction and strength of the causal relationship according to the Pearson correlation coefficient;
[0078] S206. When the causality test is passed, a linear model is fitted at each point in the X phase space to obtain a fitting coefficient, and a commonly used heat map or the like can be drawn to display the spatiotemporal distribution of the impact of environmental factors analyzed by S-map on infectious disease activities;
[0079] S207. Use the Pearson correlation coefficient as the first causal relationship information.
[0080] Before executing steps S201-S207, data loading and preprocessing can be performed. In the data loading and preprocessing process, the environmental data time series and the infectious disease incidence time series are loaded, and the missing data in the environmental data time series and the infectious disease incidence time series are checked and processed; each data in the environmental data time series and the infectious disease incidence time series is standardized or normalized to eliminate the dimension effect.
[0081] Step S201 is the step of parameter selection: embedding dimension (E) and lag time (τ), the range of E is The lower limit E = 2 ensures that at least one external variable is included in the embedding dimension. The upper limit is determined based on the time series length n of the environmental data time series and the infectious disease incidence time series. Leave-One-Out Cross-Validation (LOO-CV) is used in Test each E within the range, take the univariate predictability (such as correlation coefficient) as the indicator, find the best embedding dimension E value as the value of the embedding dimension to be set in step S201, and determine the lag time τ as the value of the lag time to be set in step S201 through the autocorrelation function or the cross-correlation function, so as to maximize the predictability of the univariate prediction.
[0082] Steps S202-S203 are the steps of constructing the phase space: based on Takens embedding theorem, the delayed coordinate method is used to embed the driving variable environmental factors, i.e., the environmental data time series, into the X phase space, and the response variable, i.e., the infectious disease incidence time series, into the Y phase space.
[0083] Step S204 is a cross-mapping step: this step uses the neighborhood information in one phase space to reconstruct the state in another phase space. The specific steps are:
[0084] a. Select a target point in the X phase space (usually the embedding point corresponding to time t);
[0085] b. Find the k nearest neighbors of the target point selected in step a (usually k = E + 1);
[0086] c. Use the weighted average of the Euclidean distances of the neighboring points found in step b to predict t at the current time point, and map the predicted t at the current time point into the Y phase space; calculate the statistical measure of the correlation between the mapped (predicted) points in phase space Y and the actually observed points in phase space Y (specifically, it can be the Pearson correlation coefficient).
[0087] Step S205 is a causality test step: analyze the results of the cross-mapping in step S204 to confirm the direction and strength of the causal relationship; specifically, if the correlation statistic is close to 1 (for example, the correlation statistic is greater than a threshold such as 0.7), it indicates that phase space X can well predict phase space Y, which indicates that phase space X may have a causal effect on phase space Y, so it can be determined that the causality test has passed, and then steps S206 and other steps can be performed. Because CCM has convergence, the longer the time series used, the smaller the estimation error of the cross-mapping will be, so a longer environmental data time series and an infectious disease incidence time series with a larger n can be used.
[0088] Step S206 is the step of constructing the S-map model: after passing the causality test of step S205, the multivariate S-map is used to fit the linear model at each point in the phase space X, and the fitting coefficient of the fitted linear model can be obtained, and the commonly used heat map and other methods can be drawn to show the spatiotemporal changes of the impact of environmental factors analyzed by the S-map on infectious disease activities. In step S207, the Pearson correlation coefficient confirms the direction and strength of the causal relationship, which can be used as the first causal relationship information.
[0089] In this embodiment, when executing step S3, that is, performing Peter-Clark instantaneous conditional independence test processing on the environmental data time series and the infectious disease incidence time series to obtain the second causal relationship information, the following steps can be specifically performed:
[0090] S301. Set the maximum time lag value L according to the information criterion max ;
[0091] S302. In the PC stage, partial correlation analysis is used as an independence test method within the range of the maximum time lag value, that is, a conditional independence test is performed on each pair of variables X and Y. After controlling the influence of other variables, the partial correlation coefficients of X and Y are calculated to identify potential driving variables (parent nodes) and causal relationships. The MCI stage is based on the preliminary causal relationship identified in the PC stage, by further calculating the partial correlation coefficient, that is, calculating the partial correlation coefficient for each lag time of the variable pair X→Y, analyzing the strength of the lagged causal relationship, and obtaining the causal strength and significance between the environmental data time series and the infectious disease incidence time series corresponding to multiple different lag times;
[0092] S303. Generate a directed acyclic graph according to each causal path;
[0093] S304. Determine the autocorrelation strength and causal path, strength (lag number) and significance of each variable according to the directed acyclic graph;
[0094] S305. Use the partial correlation coefficient as the second causal relationship information.
[0095] Before executing steps S301-S305, data loading and preprocessing can be performed. In the data loading and preprocessing process, the environmental data time series and the infectious disease incidence time series are loaded, and the missing data in the environmental data time series and the infectious disease incidence time series are checked and processed; each data in the environmental data time series and the infectious disease incidence time series is standardized or normalized to eliminate the dimension effect.
[0096] Step S301 is the step of parameter selection: there may be a lag between the environmental data time series and the infectious disease incidence time series, and the maximum time lag value L can be set by the information criterion (AIC, BIC). max。 Maximum time lag value L max It represents the maximum value of the lag time between the environmental data time series and the infectious disease incidence time series, that is, the lag time range is [0, L max ].
[0097] Step S302 is the step of identifying PCMCI+ causal relationships: Step S302 is the PC stage, in which partial correlation analysis is used as an independence test method within the range of the maximum time lag value, that is, a conditional independence test is performed on each pair of variables X and Y, and after controlling the influence of other variables, the partial correlation coefficients of X and Y are calculated to identify potential driving variables (parent nodes) and causal relationships. The MCI stage is based on the preliminary causality identified in the PC stage, by further calculating the partial correlation coefficient, that is, calculating the partial correlation coefficient for each lag time of the variable pair X→Y, analyzing the strength of the lagged causal relationship, and obtaining the causal strength and significance between the environmental data time series and the infectious disease incidence time series corresponding to multiple different lag times;
[0098] Steps S303-S304 are steps for analysis using a directed acyclic graph: a directed acyclic graph can visualize the lagged causal relationship and the concurrent causal relationship between the environmental data time series and the infectious disease incidence time series. Specifically, the curved edges and straight edges between nodes represent the lagged causal relationship and the concurrent causal relationship, respectively. The lag number will be marked on the curved edge. The color of the node represents the autocorrelation strength (i.e., auto-MCI, autocorrelation instantaneous conditional independence), and the color of the edge represents the causal strength estimated by partial correlation (i.e., cross-MCI, cross-instantaneous conditional independence). Therefore, in step S305, the partial correlation coefficient can be used as the second causal relationship information.
[0099] In this embodiment, when executing step S4, that is, performing causal forest algorithm processing on the environmental data time series and the infectious disease incidence time series to obtain the third causal relationship information, the following steps can be specifically performed:
[0100] S401. Construct a causal forest model;
[0101] S402. Use the environmental data time series and the infectious disease incidence time series to train the causal forest model;
[0102] S403. Calculate the average treatment effect and the conditional average causal effect using the trained causal forest model;
[0103] S404. Use the conditional average causal effect as the third causal relationship information.
[0104] Before executing steps S401-S404, data loading and preprocessing can be performed. During the data loading and preprocessing process, the environmental data time series and the infectious disease incidence time series are loaded, and the missing data in the environmental data time series and the infectious disease incidence time series are checked and processed; each data in the environmental data time series and the infectious disease incidence time series is standardized or normalized to eliminate the dimension effect. During the data loading and preprocessing process, the holiday effect, day of the week effect, etc. in the environmental data time series and the infectious disease incidence time series can also be processed as dummy variables, and the infectious disease data of the previous time unit can be introduced in the form of lagged variables to control the potential autocorrelation effect.
[0105] Steps S401-S402 are steps for establishing and training a causal forest model: In step S401, the constructed causal forest model has
[0106] Y t =f(X t )+∈
[0107] The form shown. Among them, Y t is the data of the time series of the number of infectious diseases; X t is the data in the environmental data time series (including atmospheric pollutants, meteorological factors, holiday effects, day of the week effects, lagged variables, etc.), f is the nonlinear relationship function learned through causal forest, and ∈ is the random error term.
[0108] In step S402, the causal forest model is trained using the environmental data time series and the infectious disease incidence time series. Specifically, the nonlinear relationship function f contains multiple parameters, and t=1, 2, 3...n is taken respectively, so that each data Y t and X t Input into the causal forest model, adjust the value of the parameter in the nonlinear relationship function f until the calculated random error term ∈ converges, and complete the training of the causal forest model. The trained causal forest model has the function of calculating the average treatment effect (ATE), that is, using the average treatment effect to measure the overall environmental factors (such as atmospheric particulate matter PM 2.5 ) Increase the impact of unit quantity on the risk of infectious diseases and other functions.
[0109] Step S403 is a statistical analysis step: the trained causal forest model is used to further calculate the conditional average treatment effect (CATE) under different conditions (e.g., different concentration ranges or meteorological conditions). The conditional average causal effect can quantitatively evaluate the causal impact of specific environmental indicators on the risk of infectious diseases. The result analysis can include the 95% confidence interval of ATE and CATE to clarify the statistical significance of the causal effect.
[0110] In step S404, the conditional average causal effect calculated in step S403 is used as the third causal relationship information.
[0111] In this embodiment, if the causal relationship between the infectious disease incidence time series and the environmental data time series represented by the first causal relationship information, the second causal relationship information, and the third causal relationship information is consistent, a more comprehensive and reliable causal relationship conclusion can be drawn, and a more powerful causal explanation can be made; if the causal relationship represented by the three causal relationship information is inconsistent, it does not mean that the causal relationship analysis has failed, and a more comprehensive and reliable causal relationship conclusion can be obtained through analytical method assumptions, data processing, external verification, fusion of multiple results, and biological rationality. Specifically, further analysis can be performed through the following steps:
[0112] a. Assumptions of analytical methods: Consider the advantages and disadvantages and applicability of different methods, analyze the degree of match between data and assumptions, and determine which methods have more reliable results. For example, CCM is applicable to nonlinear and large complex dynamic systems, and clarifies the causal direction between two variables. PCMCI+ is applicable to high-dimensional time series data, assuming that the conditional independence test is valid and can control confounding, making it easy to identify instantaneous and lagged causal relationships. Causal forest assumes that data are randomly distributed, does not provide a clear causal direction, but can provide quantitative causal estimation coefficients;
[0113] b. Check the data: perform sensitivity analysis on data preprocessing and parameter settings to see if the inconsistent results are caused by these factors;
[0114] c. External validation: By using validation sets and other methods to verify the predictive power, robustness, and consistency with known domain knowledge, we can evaluate which method's causal inference is more reliable. If the results of a method are inconsistent with most methods or domain knowledge, we may need to exclude the results of that method.
[0115] d. Result fusion: The causal relationships inferred by the three methods are combined into a union as a candidate set for further verification, which may include common causal relationships or causal relationships unique to certain methods. These can be discussed together, or different weights can be given to the results based on the theoretical basis and applicability of the assumptions of each method for a comprehensive evaluation;
[0116] e. Comprehensive explanation: clearly point out the source of the difference. Some causal relationships are significant in dynamic systems but not captured by the causal forest model; or some complex interaction effects are discovered by the causal forest but are not significant in the Bayesian network, etc.; interpret the results in a hierarchical manner, with multiple methods consistently supporting high-confidence causal relationships, and contradictory results requiring further verification of causal relationships; finally, combine relevant knowledge in the field to explain the contradictions.
[0117] In this embodiment, the specific forms of the first causal relationship information obtained by executing steps S201-S207 are Pearson correlation coefficients and visual heat maps, the specific forms of the second causal relationship information obtained by executing steps S301-S305 are partial correlation coefficients and directed acyclic graphs, and the specific forms of the third causal relationship information obtained by executing steps S401-S404 are conditional average causal effects, which all quantitatively represent the causal relationship between the environmental data time series and the infectious disease incidence time series from their respective aspects. Therefore, the first causal relationship information, the second causal relationship information, and the third causal relationship information can be de-dimensionalized respectively, so that operations such as summation can be performed between the first causal relationship information, the second causal relationship information, and the third causal relationship information.
[0118] In this embodiment, when executing step S5, that is, determining the causal relationship between the environmental data time series and the infectious disease incidence time series based on the first causal relationship information, the second causal relationship information, and the third causal relationship information, the following steps may be specifically performed:
[0119] S501. Data preprocessing and standardization or normalization of causal information output by all methods;
[0120] S502. Determine weight;
[0121] S503. Perform weighted fusion on the first causal relationship information, the second causal relationship information and the third causal relationship information to determine the causal relationship between the environmental data time series and the infectious disease incidence time series;
[0122] S504. Visualization and consistency verification.
[0123] In step S501, the causal directions, causal strength indicators (Pearson correlation coefficient, partial correlation coefficient, conditional average causal effect) and their corresponding significant P values of the three methods, namely CCM, PCMCI+ and causal forest, are collected. The causal strength indicators are standardized or normalized; all P values are converted to -log(P), which can be used as weights later.
[0124] In step S502, the principles of weight allocation can be based on a. method characteristics (e.g., CCM is suitable for dynamic nonlinear data, PCMCI+ assumes that the independence test of the conditional is valid and can control confounding, and the causal forest assumes that the data is randomly distributed and provides quantitative effects, etc.); b. significance of causal association (P value, the method results with higher significance are given greater weight); c. consistency of results: relationships with higher consistency between methods are given greater weight; d. data characteristics (e.g., when the sample size is small, the causal forest may be weaker, while CCM and PCMCI+ are more robust); preliminary allocation suggestions: if the data is highly nonlinear, CCM has a higher weight (W CCM =0.5, W PCMCI+ is 0.3, W CF is 0.2), if the confounding variable is large, the PCMCI+ weight is higher, etc.
[0125] In step S5003, the causal strength of the three methods is calculated as follows: Comprehensive causal strength = S CCM (X→Y)·W CCM +S PCMCI+ (X→Y)·W PCMCI+ +S CF (X→Y)·W CF , confidence intervals were calculated for the weighted results to assess the consistency of results between different methods.
[0126] In step S504, a heat map or a network diagram is generated to visualize the weighted causal relationship strength, and the causal relationship with low confidence is further verified and comprehensively explained in combination with actual data.
[0127] A computer program for executing the time series causal analysis method for environmental changes and infectious disease risks in this embodiment can be written and written into a computer device or storage medium. When the computer program is read out and executed, the time series causal analysis method for environmental changes and infectious disease risks in this embodiment is executed, thereby achieving the same technical effect as the time series causal analysis method for environmental changes and infectious disease risks in the embodiment.
[0128] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. In addition, the descriptions of up, down, left, right, etc. used in the present disclosure are only relative to the relative positional relationship of the components of the present disclosure in the accompanying drawings. The singular forms of "a" and "the" used in the present disclosure are also intended to include the plural forms, unless the context clearly indicates other meanings. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by technicians in this technical field. The terms used in the specification of this embodiment are only for describing specific embodiments, not for limiting the present invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.
[0129] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided by the present embodiment is only intended to better illustrate embodiments of the present invention, and unless otherwise required, the scope of the present invention will not be limited.
[0130] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed dedicated integrated circuit for this purpose.
[0131] In addition, the operations of the process described in this embodiment may be performed in any suitable order, unless otherwise indicated in this embodiment or otherwise clearly contradicted by the context. The process described in this embodiment (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as a code (e.g., executable instructions, one or more computer programs, or one or more applications) executed on one or more processors in common, by hardware or a combination thereof. A computer program includes a plurality of instructions that may be executed by one or more processors.
[0132] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or part thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0133] The computer program can be applied to input data to perform the functions of the present embodiment, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.
[0134] The above are only preferred embodiments of the present invention. The present invention is not limited to the above embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation methods may have various modifications and changes.
Claims
1. A time series causal analysis method for environmental change and infectious disease risk, characterized in that: The time series causal analysis method of environmental change and infectious disease risk includes: Obtain and process environmental data time series and infectious disease incidence time series; Performing convergent cross-mapping processing on the environmental data time series and the infectious disease incidence time series, evaluating the causal impact of candidate driving factors on infectious disease activities, and obtaining first causal relationship information; Performing Peter-Clark instantaneous conditional independence test on the environmental data time series and the infectious disease incidence time series to obtain visualized lag and contemporaneous causal dependencies and obtain second causal relationship information; Processing the environmental data time series and the infectious disease incidence time series with a causal forest algorithm to quantitatively analyze the linear and nonlinear relationships between environmental factors and infectious disease activities, and obtain third causal relationship information; The causal relationship between the environmental data time series and the infectious disease incidence time series is determined based on the first causal relationship information, the second causal relationship information and the third causal relationship information.
2. The time series causal analysis method of environmental change and infectious disease risk according to claim 1 is characterized in that: The step of performing convergent cross-mapping processing on the environmental data time series and the infectious disease incidence time series, evaluating the causal impact of the candidate driving factors on infectious disease activities, and obtaining first causal relationship information includes: Set embedding dimension and lag time; Embedding the environmental data time series into an X phase space; Embed the infectious disease incidence time series into the Y phase space; Reconstructing the state in the Y phase space based on the embedded neighborhood information in the X phase space, and calculating the Pearson correlation coefficient between the predicted Y and the actually observed Y; Confirm the direction and strength of the causal relationship based on the Pearson correlation coefficient; When the causality test is passed, a linear model is fitted at each point in the X phase space to obtain fitting coefficients, showing the spatiotemporal distribution of the impact of environmental factors on infectious disease activities analyzed by S-map; The Pearson correlation coefficient is used as the first causal relationship information.
3. The time series causal analysis method of environmental change and infectious disease risk according to claim 1 is characterized in that: The Peter-Clark instantaneous conditional independence test is performed on the environmental data time series and the infectious disease incidence time series to obtain a visualized lag and contemporaneous causal dependency relationship and obtain the second causal relationship information, including: Set the maximum time lag value based on the information criterion; In the PC stage, within the range of the maximum time lag value, partial correlation analysis is used as an independence test method, that is, a conditional independence test is performed on each pair of variables X and Y. After controlling the influence of other variables, the partial correlation coefficient of X and Y is calculated to identify potential driving variables and causal relationships; Generating a directed acyclic graph according to each of the causal relationships; Determining the autocorrelation strength and causal path, strength and significance of each variable based on the directed acyclic graph; The partial correlation coefficient is used as the second causal relationship information.
4. The time series causal analysis method of environmental change and infectious disease risk according to claim 3 is characterized in that: Determining the autocorrelation strength and causal path, strength and significance of each variable according to the directed acyclic graph includes: Visualizing the directed acyclic graph to obtain multiple nodes and edges connecting the nodes; Determine the autocorrelation strength of each variable according to the color of the node; Determine whether it is a lagged causal relationship based on the curvature of the edge; The causal strength is determined according to the color of the edge.
5. The time series causal analysis method of environmental change and infectious disease risk according to claim 1 is characterized in that: The step of performing causal forest algorithm processing on the environmental data time series and the infectious disease incidence time series to obtain third causal relationship information includes: Constructing a causal forest model; Training the causal forest model using the environmental data time series and the infectious disease incidence time series; Calculate the average treatment effect and the conditional average causal effect using the trained causal forest model; The conditional average causal effect is used as the third causal relationship information.
6. The time series causal analysis method of environmental change and infectious disease risk according to claim 5 is characterized in that: The causal forest model is constructed, comprising: According to the formula Y t =f(X t )+∈ Establish the causal forest model; where Y t is the data in the time series of the number of infectious diseases; X t is the data in the environmental data time series; f is the nonlinear relationship function learned by causal forest; ∈ is the random error term.
7. The time series causal analysis method of environmental change and infectious disease risk according to claim 1 is characterized in that: Determining the causal relationship between the environmental data time series and the infectious disease incidence time series based on the first causal relationship information, the second causal relationship information, and the third causal relationship information includes: performing weighted fusion on the first causal relationship information, the second causal relationship information, and the third causal relationship information; According to the result of weighted fusion, the causal relationship between the environmental data time series and the infectious disease incidence time series is determined.
8. The time series causal analysis method of environmental change and infectious disease risk according to claim 7 is characterized in that: The weighted fusion of the first causal relationship information, the second causal relationship information and the third causal relationship information includes: The Pearson correlation coefficient, partial correlation coefficient and conditional average causal effect are standardized; the standardized Pearson correlation coefficient is S CCM (X→Y), the standardized partial correlation coefficient is S PCMCI+ (X→Y), the standardized conditional average causal effect is S CF (X→Y); Weights are assigned to the Pearson correlation coefficient, partial correlation coefficient, and conditional average causal effect according to their methodological characteristics, significance of causal association, and consistency of results. The weight corresponding to the Pearson correlation coefficient is W. CCM , the weight corresponding to the partial correlation coefficient is W PCMCI+ , the weight corresponding to the conditional average causal effect is W CF ; According to the formula Overall causal strength = S CCM (X→Y)·W CCM +S PCMCI+ (X→Y)·W PCMCI+ +S CF (X→Y)·W CF Calculate the combined causal strength; The comprehensive causal strength is taken as the result of weighted fusion.
9. A computer device, characterized in that: It includes a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the time series causal analysis method of environmental change and infectious disease risk as described in any one of claims 1-8.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to execute the time series causal analysis method of environmental changes and infectious disease risks as described in any one of claims 1 to 8 when executed by the processor.
Citation Information
Patent Citations
Infectious disease prediction method and apparatus, electronic device and computer readable medium
CN109859854A
System and method for prediction of self-similar signals
WO2013113111A1
Cited By
Intelligent agriculture remote supervision method based on Internet of Things
CN121581246A