Dynamic estimation method and device for causal effect, and medium
Time sequence features are extracted through one-dimensional expansion convolution and multi-layer intervention-aware convolution modules, and combined with adversarial learning strategies and regularization terms, the dynamic estimation model of causal effects is optimized, solving the problem of insufficient control of time-varying confounding factors and insufficient long-term cumulative effect characterization, and achieving higher accuracy and generalization performance.
Patent Information
- Application Number
- CN202510257215.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
In the dynamic estimation of causal effects, the problems of insufficient control of time-varying confounding factors, insufficient long-term cumulative effects portrayal, and insufficient balance of loss during training.
One-dimensional expansion convolution is used to jointly process the time-varying covariates and intervention variables, and the preliminary timing feature representation is extracted, and a high-dimensional timing feature representation is generated through nonlinear mapping. Then, the high-dimensional features are input into the multi-layer intervention-aware convolution module for deep extraction to generate a comprehensive feature representation. A predictive model is constructed based on the comprehensive feature representation, and the objective function of the factual loss, adversarial loss and regularization terms are combined to iteratively update the model parameters.
It significantly improves the ability to control time-varying confounders, reduces selective bias, accurately captures the timing characteristics of long-term cumulative effects and dynamic interventions, balances factual losses with counterfactual risks, and improves the trade-off between model deviation and variance.
Smart Images

Figure CN120086535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of causal inference and time series prediction, and specifically to a method, device, and medium for dynamically estimating causal effects. Background Art
[0002] The dynamic estimation of causal effects is one of the important tasks in modern data-driven decision-making, and is widely applied in fields such as policy evaluation, medical research, educational intervention, and social sciences. In practical scenarios, longitudinal observational data often records information such as intervention measures, time-varying covariates, and outcome variables. Analyzing these data is crucial for revealing the short-term and long-term effects of intervention measures. However, traditional causal inference methods face many challenges when dealing with complex longitudinal data.
[0003] Classical causal effect estimation methods, such as propensity score matching and inverse probability weighting, usually rely on linear assumptions and cannot effectively capture the complex interaction relationships between time-varying covariates and intervention variables. At the same time, these methods show obvious limitations when analyzing long-term cumulative effects and optimizing dynamic intervention strategies. In recent years, with the rapid development of deep learning technology in the field of time series modeling, some researchers have tried to introduce it into the field of causal inference, thus proposing a series of deep learning-based causal effect estimation algorithms, such as RMSN (Recurrent Marginal Structural Networks). These methods provide new ideas for causal inference tasks, but still have the following problems:
[0004] Insufficient control of time-varying confounding factors: When dealing with time-varying confounding factors, existing methods fail to effectively model their dynamic changes over time, which may lead to selection bias and thus affect the accurate estimation of causal effects.
[0005] Insufficiently detailed characterization of dynamic intervention effects: Most methods do not fully consider the influence of long-term cumulative effects and intervention timing, and cannot accurately characterize the global effects of dynamic interventions.
[0006] Insufficient loss balance during training: In model training, existing methods fail to balance the factual loss and counterfactual risk, resulting in the model being difficult to achieve a reasonable trade-off between bias and variance. Summary of the Invention
[0007] The present invention provides a method, device, and medium for dynamically estimating causal effects, aiming to solve the problems of insufficient control of time-varying confounding factors, insufficiently fine characterization of long-term cumulative effects, and insufficient loss balance during training in existing methods for dynamic intervention effect estimation.
[0008] To achieve the above object, the first aspect of the present invention provides a method for dynamically estimating causal effects, including the following steps:
[0009] Obtain a longitudinal observational dataset containing time-varying covariates, static covariates, intervention variables, and outcome variables. Align the longitudinal observational dataset according to the time series and organize it into a structured sample to generate input data;
[0010] Use one-dimensional dilated convolution to jointly process the time-varying covariates and intervention variables in the input data to extract preliminary temporal feature representations;
[0011] Perform non-linear mapping processing on the extracted preliminary temporal feature representations to generate high-dimensional temporal feature representations;
[0012] Input the high-dimensional temporal feature representations into a multi-layer intervention-aware convolution module, and perform in-depth extraction and characterization of the high-dimensional temporal features through cascaded multi-layer convolutions to generate comprehensive feature representations;
[0013] Construct a prediction model based on the comprehensive feature representations for predicting the outcome variables and time-varying covariates at the next time step;
[0014] Use the output outcome variables and time-varying covariates, combine with the actual observed values to train the prediction model. Minimize the loss through an optimization algorithm and generate adversarial samples. Construct an objective function containing a factual loss, an adversarial loss, and a regularization term, and iteratively update the prediction model parameters to obtain the trained prediction model;
[0015] Use the trained prediction model to perform counterfactual inference on the input data, generate counterfactual trajectories, and conduct dynamic causal effect evaluation.
[0016] Further, the time-varying covariates are feature variables that change over time, the intervention variables are intervention information of the observed samples over time, the outcome variables are effect evaluation values of the observed samples over time, and the static covariates are fixed feature variables of the observed samples.
[0017] Further, the method of using one-dimensional dilated convolution to jointly process the time-varying covariates and intervention variables in the input data to extract preliminary temporal feature representations includes:
[0018] Jointly process the time-varying covariates and intervention variables in the input data, align them according to the time series to form an input tensor with consistent time step dimensions;
[0019] Use one-dimensional dilated convolution operations to perform convolution calculations on the jointly processed input tensor. During the convolution calculation process, embed the intervention variables into the convolution operation through a non-linear mapping function, and set the dilation factor of the convolution kernel to capture long-term dependencies across time steps to generate a temporal feature map;
[0020] Use the generated temporal feature map as the preliminary temporal feature representation.
[0021] Furthermore, a one-dimensional dilated convolution operation is used to perform convolution calculation on the jointly processed input tensor, and the calculation formula is as follows:
[0022]
[0023] Wherein, is the output at time step t, k is the convolution kernel size, d is the dilation factor used to expand the receptive field, is the i-th convolution weight matrix, φ(a t ) is the non-linear mapping function of the intervention, is the input at time step t-i, is the bias term;
[0024] The convolution output is used as the preliminary temporal feature representation.
[0025] Furthermore, the method for generating the high-dimensional temporal feature representation by performing non-linear mapping processing on the extracted preliminary temporal feature representation includes:
[0026] For the extracted preliminary temporal feature representation, a radial basis function is used for non-linear mapping, and Gaussian radial basis function and multiquadric radial basis function are selected as the basis functions;
[0027] In the Gaussian radial basis function, the non-linear mapping value is calculated based on the input feature and the predefined center point. The calculation formula of the Gaussian radial basis function is:
[0028]
[0029] Wherein, a i is the i-th dimension of the intervention vector, c j is the j-th center point, and ζ is the width parameter of the Gaussian function;
[0030] In the multiquadric radial basis function, the non-linear mapping value is calculated based on the input feature and the predefined center point. The calculation formula of the multiquadric radial basis function is:
[0031]
[0032] Wherein, σ is the shape parameter;
[0033] The calculation results of the Gaussian radial basis function and the multiquadric radial basis function are combined in the form of a weighted sum to generate the final non-linear mapping value, and its formula is:
[0034]
[0035] Wherein, φ(a) is the high-dimensional feature representation of the intervention variable a, w ij is the learnable weight, d a and d care the intervention variable dimension and the number of center points respectively, ψ(a i , c j ) represents the radial basis function, which is used to calculate the non-linear relationship between the i-th dimension of the intervention vector and the j-th center point;
[0036] The non-linear mapping result is used as the high-dimensional time series feature representation.
[0037] Furthermore, the method for generating the comprehensive feature representation includes:
[0038] Input the high-dimensional time series feature representation into the first-layer intervention-aware convolutional module, and use a one-dimensional dilated convolutional kernel to perform a convolutional operation on the input features, where the dilation factor is used to expand the receptive field of the time step to capture long-term dependencies;
[0039] In each layer of the intervention-aware convolutional module, a residual connection method is adopted to add the input features and the output features after the convolutional operation element by element; among them, the input features and the output features are added element by element using the residual connection, and its calculation formula is:
[0040] o = Activation(B(Z) + F(Z))
[0041] where o is the time series feature output, Activation represents the non-linear activation function, B(Z) is the mapping operation of the input features, and when the input and output dimensions are the same, it is adjusted by a 1×1 convolution, and F(Z) is the output feature of the convolutional module;
[0042] Perform a non-linear activation process on the features after the convolutional operation to generate an updated time series feature representation;
[0043] Input the updated time series feature representation into subsequent multiple layers of intervention-aware convolutional modules in sequence, and repeat the above steps to perform in-depth extraction and characterization of the high-dimensional time series features;
[0044] In the last layer of the intervention-aware convolutional module, a comprehensive feature representation is generated.
[0045] Furthermore, a prediction model is constructed based on the comprehensive feature representation to predict the result variable and the time-varying covariate at the next time step, which specifically includes the following steps:
[0046] The prediction model includes an output decoder and a covariate decoder;
[0047] Input the comprehensive feature representation into the output decoder, and map the comprehensive feature representation through a multi-layer feedforward neural network to predict the result variable at the next time step and output the predicted value, and its calculation formula is:
[0048]
[0049] Among them, is the result variable for the predicted next time step, G Y is the output decoder, which adopts a multi-layer feedforward neural network, r t is the comprehensive feature representation at time step t;
[0050] Input the comprehensive feature representation into the covariate decoder and the smoothing factor generation network to respectively predict the covariate base value and the smoothing factor for the next time step. The calculation formula is:
[0051]
[0052] η = Sigmoid(J X (r t ))
[0053] Among them, is the time-varying covariate for the predicted next time step, G X is the covariate decoder, which adopts a feedforward neural network, X t is the observed covariate at time step t, η is the smoothing factor; J X is the feedforward network for generating the smoothing factor, and Sigmoid is used to limit the smoothing factor within the range of [0, 1];
[0054] Use the result variable and the time-varying covariate obtained by the output decoder and the covariate decoder as the output for the next time step.
[0055] Furthermore, use the result variable and the time-varying covariate output by the prediction model, combine with the actual observed values to train the prediction model, minimize the loss through an optimization algorithm and generate adversarial samples, construct an objective function including a factual loss, an adversarial loss and a regularization term, and iteratively update the prediction model parameters, which specifically includes the following steps:
[0056] Construct a loss function, including a factual loss and an adversarial loss. The calculation formula is as follows:
[0057]
[0058] Among them, is the total loss function, N represents the number of samples in the training set, i represents the i-th sample, is the training data set, t is the index of the time step, T is the maximum number of observed time steps, is the prediction loss of the result variable at the t-th time step, is the prediction loss of the time-varying covariate at the t-th time step, θ is the model parameter, ∈ is the adversarial perturbation, ρ is the constraint range of the adversarial perturbation, represents the distribution deviation between the time series feature and the adversarial sample feature, is the regularization term of the model, λ1 , λ 2 , λ 3 are the weights of each part of the loss;
[0059] Using the SAM optimization method, minimize the maximum loss within the ρ-neighborhood of the parameter space, and the calculation formula is as follows:
[0060] min θ max ||∈||≤ρ R S (h θ+∈ )
[0061] where θ is the set of model parameters, ∈ represents the perturbation size of the model parameters θ, ∈ is a vector whose norm is constrained by ||∈|| ≤ ρ, ρ is the size of the SAM neighborhood, and R S is the empirical risk of the loss function on the training data, and h θ+∈ is the prediction model with adversarial perturbations;
[0062] Introduce an adversarial sample generation strategy, and use the fast gradient sign method to generate adversarial samples. The calculation formula is as follows:
[0063]
[0064] where r adv is the adversarial sample, which is obtained by applying a perturbation to the original input data r. r is the original input data and is a sample in the training dataset. is the sign of the gradient of the loss function. The gradient sign represents the sign of the partial derivative direction of the loss function l(f(r), y)) with respect to the original input data r, indicating how to adjust the perturbation size ∈ to maximize the increment of the loss function. l is the loss function, f(r) is the predicted output of the model, and y is the true value. is the gradient of the loss function with respect to the feature, and ∈ is the perturbation size, which is determined by maximizing the loss increment;
[0065] Use the Adam optimizer to iteratively update the model parameters, and stop training when the validation set loss has not improved for multiple consecutive rounds.
[0066] Furthermore, use the trained prediction model for counterfactual reasoning, generate counterfactual trajectories, and conduct dynamic causal effect evaluation, which specifically includes the following steps:
[0067] Input the observed history, and predict the result variable and covariates at the next time step. The calculation formula is as follows:
[0068]
[0069] where is the output value predicted at time step T + 1, GY is the output decoder, is the observation history, including the feature representations, intervention information, and outcome variables at all time steps. Φ(·) is a feature extraction function that extracts effective features from the observation history data ;
[0070] Specify the sequence of future intervention variables: A T+1:T+τ = a 1:τ , and use the rolling prediction method to generate the prediction trajectory. The specific steps are as follows:
[0071] Initialize the historical observations:
[0072]
[0073] Iteratively execute the following steps from t = T + 1 to T + τ:
[0074]
[0075] Output the complete counterfactual trajectory:
[0076]
[0077] where X 1 ,...., X T represents the covariates between time steps 1 and T, Y 1 ,...., Y T represents the outcomes between time steps 1 and T, A 1 ,...., A T represents the interventions between time steps 1 and T, a 1:τ is the intervention, τ represents the number of time steps for prediction, is the updated historical data, which combines the currently predicted covariates Output and the intervention a t into the historical data. F X is the covariate prediction function that predicts the covariates based on the historical data , and are the time series of the predicted outcome variables and time-varying covariates, respectively.
[0078] Furthermore, the non-linear mapping of the radial basis function includes:
[0079] Gaussian radial basis function: generates high-dimensional features with a Gaussian distribution for continuous intervention variables;
[0080] Multiquadric radial basis function: generates high-dimensional features with a multiquadric distribution for discrete intervention variables.
[0081] To achieve the above object, a second aspect of the present invention provides an electronic device, including a processor and a memory. The processor is configured to implement the steps of the causal effect dynamic estimation method when executing the computer program stored in the memory.
[0082] To achieve the above object, a third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. The computer program, when run by a processor, executes the steps of the causal effect dynamic estimation method.
[0083] Advantages of the present invention:
[0084] Compared with the prior art, a causal effect dynamic estimation method, device and medium provided by the present invention significantly improve the ability to control time-varying confounding factors and reduce selection bias by introducing domain generalization theory and adversarial learning strategies; use intervention-aware convolution and radial basis functions to perform fine-grained modeling on intervention information, and can accurately capture the long-term cumulative effect and the temporal characteristics of dynamic intervention; by combining Sharpness-Aware Minimization (SAM) optimization and adversarial sample generation strategies, balance the factual loss and counterfactual risk, so as to achieve a good trade-off between the bias and variance of the model. The overall framework is based on an end-to-end deep learning architecture, which can efficiently perform multi-step counterfactual prediction, providing higher accuracy and generalization performance for the estimation of dynamic causal effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0086] Figure 1 It is a flowchart of a causal effect dynamic estimation method disclosed in an embodiment of the present invention.
[0087] Figure 2 It is a flowchart of an extended convolution disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0088] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0089] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the following methods, in some cases, the steps shown or described can be executed in a different order than here.
[0090] As Figure 1 , Figure 2 shown, the present invention provides a causal effect dynamic estimation method, including the following steps:
[0091] Step S100, obtain a longitudinal observation data set including time-varying covariates, static covariates, intervention variables, and outcome variables, align the longitudinal observation data set according to the time series and organize it into a structured sample to generate input data;
[0092] First, it is necessary to obtain a longitudinal observation data set including time-varying covariates, static covariates, intervention variables, and outcome variables from the research object. Specifically:
[0093] Time-varying covariates: represent features that change dynamically over time, such as a patient's blood pressure, heart rate, or a country's economic indicators (such as GDP, inflation rate).
[0094] Static covariates: represent features that do not change over time, such as a patient's gender, year of birth, or a country's geographical location.
[0095] Intervention variables: refer to the intervention measures applied in the time series, such as a drug treatment plan, a financial assistance policy.
[0096] Outcome variables: the final indicators reflecting the intervention effect, such as the health recovery situation, economic growth rate, etc.
[0097] Since different samples may have different time spans or observation time points, when implementing step S100, it is necessary to align the data in time series. The purpose of time series alignment is to align the data of different samples according to the same time interval to construct a unified input data format. For example, align the data by day, month, or year to ensure the consistency of each sample in the time dimension. This step processes incomplete data points through techniques such as interpolation methods and filling missing values to ensure the continuity and effectiveness of the model input data.
[0098] After completing the time alignment, the data needs to be organized into a structured sample format that can be directly processed by the model. Each sample is represented in the following form:
[0099]
[0100] where represents the time-varying covariate of sample i at time t, Denote the intervention variable of sample \(i\) at time \(t\), Denote the outcome variable of sample \(i\) at time \(t\), Denote the static covariates of sample \(i\) at time \(t\), which are stored in the form of tensors or matrices for subsequent processing and model input.
[0101] Use the aligned and organized structured samples as input data for feature extraction and learning of subsequent models. During this process, techniques such as standardization and normalization can be used to adjust the scale of variables and eliminate the impact of different dimensions on the model. For example, normalize numerical variables to the range \([0, 1]\), or convert categorical variables to one-hot encoding. Finally, the generated input data meets the requirements of model processing.
[0102] Step S200: Jointly process the time-varying covariates and intervention variables in the input data using one-dimensional dilated convolution to extract preliminary temporal feature representations;
[0103] Time-varying covariates refer to variables that change over time, such as a patient's heart rate or an economic indicator of a country. Intervention variables refer to variables that are intervened or processed over time, such as drug treatment, financial assistance, etc. Since the time steps of these data may be different, the task of step S200 is to align the data according to the time series to ensure that different samples have the same observation time points in the time dimension. Through such processing, it is ensured that the time-varying covariates and intervention variables of each sample correspond one by one in the time dimension, forming an input tensor with a consistent time step dimension.
[0104] Next, perform convolution calculations on the jointly processed input tensor using one-dimensional dilated convolution operations. One-dimensional dilated convolution operations slide and calculate the input data through a convolution kernel, and determine the size of the receptive field according to the set dilation factor at each slide. The dilation factor is the number of time steps that the convolution kernel jumps each time, which can capture long-term dependencies spanning multiple time steps. During this process, the time-varying covariates and intervention variables are input into the convolutional layer together, and the convolution operation will extract the temporal features therein to help the model understand the patterns that change over time.
[0105] When performing convolution calculations, the intervention variable will be embedded into the convolution operation through a non-linear mapping function. This non-linear mapping function \(\varphi(a\) t ) can transform the intervention variable \(a\) t into a new feature space, enabling the intervention information to participate more fully in the convolution calculation. In this way, the intervention variable is not only convolved with the time-varying covariates, but also can further enrich its influence on the temporal features through non-linear mapping, enabling the model to better capture the impact of the intervention on the outcome.
[0106] After convolution calculation, the generated output result is a temporal feature map. This temporal feature map contains the convolution results of each time step, and each time step F(t) in it is a feature representation calculated from time-varying covariates, intervention variables, and convolution kernels.
[0107] During the convolution calculation process, the calculation formula of the temporal feature map is as follows:
[0108]
[0109] Among them, is the output at time step t, k is the convolution kernel size, d is the dilation factor used to expand the receptive field, is the i-th convolution weight matrix, φ(a t ) is the non-linear mapping function of the intervention, is the input at time step t - i, is the bias term;
[0110] Through the above convolution operation, the model can extract preliminary temporal features from the input time-varying covariates and intervention variables.
[0111] Step S300, perform non-linear mapping processing on the extracted preliminary temporal feature representation to generate a high-dimensional temporal feature representation;
[0112] Perform non-linear mapping processing on the preliminary temporal feature representation extracted in the previous step to generate a more expressive high-dimensional temporal feature representation. This process is achieved by using the radial basis function (RBF). The radial basis function is a mathematical tool widely used for non-linear transformation, which can help the model capture more complex relationships. Specifically, two types of radial basis functions are selected - Gaussian radial basis function (Gaussian RBF) and multiquadric radial basis function (Multiquadric RBF), which can effectively transform the input features into a high-dimensional space, thereby enhancing the non-linear expression ability of the model.
[0113] In the Gaussian radial basis function, the non-linear mapping value is calculated based on the input features and predefined center points. The difference between each input feature (for example, a certain dimension of the time-varying covariates and intervention variables) and the preset center point determines the size of the mapping value. Specifically, the calculation formula of the Gaussian radial basis function is:
[0114]
[0115] Among them, a i is the i-th dimension of the intervention vector, c j is the j-th center point, and ζ is the width parameter of the Gaussian function;
[0116] This calculation method can convert the distance between the input feature and the center point into a non-linear mapping value. The closer the distance, the larger the mapping value, thus highlighting the influence of the local area.
[0117] In addition to the Gaussian radial basis function, the multi-quadratic radial basis function is also adopted to further enhance the diversity of the mapping. The multi-quadratic radial basis function calculates the non-linear mapping value through the difference between the input feature and the predefined center point. Its calculation formula is:
[0118]
[0119] where σ is the shape parameter used to adjust the shape of the function.
[0120] This multi-quadratic function is smoother when the distance is far, and is suitable for capturing non-linear relationships at different scales.
[0121] To make full use of the advantages of these two radial basis functions, the calculation results of the Gaussian radial basis function and the multi-quadratic radial basis function are combined in the way of weighted sum. Specifically, given the input feature, its final non-linear mapping value can be calculated by the following formula:
[0122]
[0123] where φ(a) is the high-dimensional feature representation of the intervention variable a, w ij is the learnable weight, d a and d c are the dimension of the intervention variable and the number of center points respectively, ψ(a i , c j ) represents the radial basis function (RBF) used to calculate the non-linear relationship between the i-th dimension of the intervention vector and the j-th center point. It can be understood that φ(a) is a non-linear mapping function obtained by weighted summing the radial basis functions of each dimension of the intervention vector and each center point. This function is used to capture the role of treatment information in temporal causal reasoning and enhance the model's ability to model intervention effects.
[0124] Step S400: Input the high-dimensional temporal feature representation into the multi-layer intervention-aware convolution module, and deeply extract and characterize the high-dimensional temporal features through cascading multiple layers of convolution to generate a comprehensive feature representation;
[0125] Input the high-dimensional time series feature representation generated in step S300 into the first-layer Intervention-aware Convolution Module (IFC). The role of this layer is to further extract time series features and help the model understand the long-term dependencies in the time series data. Specifically, the Intervention-aware Convolution Module uses one-dimensional dilated convolutional kernels to perform convolutional operations on the input features. The role of dilated convolution is to expand the receptive field of each convolutional operation by setting the dilation factor, enabling the convolutional kernel to capture long-term dependencies spanning multiple time steps. This processing step helps the model capture the complex time series patterns existing in the time series.
[0126] In each layer of the Intervention-aware Convolution Module, a residual connection is introduced. The key to the residual connection is to add the input features and the output features after the convolutional operation element-wise, thus avoiding the problem of gradient vanishing or gradient explosion in deep networks. In this way, the input features can be effectively combined with the features extracted by the convolutional operation, enabling the model to more easily learn the important information in the time series. The specific calculation formula is:
[0127] o = Activation(B(Z) + F(Z))
[0128] where o is the output of the time series features, Activation represents the non-linear activation function, B(Z) is the mapping operation of the input features, which is adjusted by 1×1 convolution when the input and output dimensions are the same, and F(Z) is the output features of the convolutional module;
[0129] Next, perform non-linear activation processing on the features after the convolutional operation. The role of this activation processing is to perform non-linear transformation on the linear output after the convolutional operation, enabling the model to capture more complex features and patterns. Common non-linear activation functions include ReLU (Rectified Linear Unit) or other types of activation functions, which can introduce non-linearity to the convolutional features, thereby enhancing the representation ability of the model.
[0130] Input the time series feature representation after non-linear activation into the subsequent Intervention-aware Convolution Module. The convolutional operation in step S400 is not performed only once, but is iteratively processed through multiple layers of Intervention-aware Convolution Modules (IFC). In each layer of the convolutional module, the model continues to perform deep extraction and characterization of the features, gradually enhancing the understanding of the time series data. The convolutional operation in each layer will learn more fine-grained time series patterns, making the final feature representation more abundant and capable of capturing time series information at different levels.
[0131] After deep processing through multiple Intervention-aware Convolution Modules, the last layer of the convolutional module will generate the final comprehensive feature representation.
[0132] Step S500: Construct a prediction model based on the comprehensive feature representation for predicting the outcome variable and time-varying covariates at the next time step;
[0133] Based on the obtained comprehensive feature representation, construct a prediction model. This prediction model consists of two main parts: an output decoder and a covariate decoder. The role of the output decoder is to predict the outcome variable at the next time step according to the comprehensive feature representation, while the covariate decoder is used to predict the time-varying covariate values at the next time step. Through these two decoders, the model can simultaneously perform two prediction tasks, providing a more comprehensive prediction of time series data.
[0134] First, the comprehensive feature representation will be input into the output decoder G Y . The output decoder is composed of multiple layers of feedforward neural networks, and its role is to map the comprehensive feature representation to predict the outcome variable at the next time step. Specifically, the output decoder receives the comprehensive feature representation and obtains the prediction result at the next time step through multiple layers of processing of the neural network The formula is as follows:
[0135]
[0136] Among them, is the predicted outcome variable at the next time step, G Y is the output decoder, using multiple layers of feedforward neural networks, r t is the comprehensive feature representation at time step t;
[0137] The comprehensive feature representation will also be input into the covariate decoder G X and the smoothing factor generation network. The task of the covariate decoder is to predict the covariate base value at the next time step (i.e., the predicted value of the time-varying covariate). In addition, the smoothing factor generation network will generate a smoothing factor, and its role is to smooth the difference between the predicted value of the covariate and the current actual value. The specific smoothing prediction formula is as follows:
[0138]
[0139] η = Sigmoid(J X (r t ))
[0140] Among them, is the predicted time-varying covariate at the next time step, G X is the covariate decoder, using a feedforward neural network, X t is the observed covariate at time step t, η is the smoothing factor; J X is the feedforward network for generating the smoothing factor, and Sigmoid is used to limit the smoothing factor within the range of [0,1];
[0141] Through this smoothing mechanism, the model can balance the actual value and the predicted value of the covariates during prediction, thereby improving the stability and accuracy of the prediction.
[0142] Through the combination of the output decoder and the covariate decoder, step S500 finally generates two outputs for the next time step: the predicted result variable The predicted time-varying covariates
[0143] In step S600, the output result variable and the time-varying covariates are used, combined with the actual observed values to train the prediction model. By optimizing the algorithm to minimize the loss and generate adversarial samples, an objective function including the factual loss, the adversarial loss, and the regularization term is constructed, and the parameters of the prediction model are iteratively updated to obtain the trained prediction model;
[0144] The goal of model training is to optimize using the result variable and the time-varying covariates output by the prediction model, combined with the actual observed values. Through this method, the model can be iteratively updated according to the error between the real data and the prediction results, thereby improving the prediction accuracy. During the training process, the optimization algorithm will gradually adjust the parameters of the prediction model by minimizing the loss function to make the prediction results closer to the actual values. To achieve this goal, this method designs a comprehensive loss function including the factual loss, the adversarial loss, and the regularization term to ensure that the model can not only predict accurately but also improve its generalization ability on unseen data.
[0145] To enhance the generalization ability of the model, the present invention introduces the Sharpness-Aware Minimization (SAM) optimization method. The key idea of SAM is to not only focus on minimizing the training loss during the optimization process but also consider the robustness and stability of the model. Specifically, SAM avoids overfitting the training data by minimizing the maximum loss within the ρ-neighborhood of the parameter space. This method is described by the following formula:
[0146] min θ max ||∈||≤ρ R S (h θ+∈ )
[0147] where θ is the set of model parameters, ∈ represents the perturbation size of the model parameters θ, ∈ is a vector whose norm is constrained by ||∈|| ≤ ρ, ρ is the size of the SAM neighborhood, and R S is the empirical risk of the loss function on the training data, and h θ+∈ is the prediction model with adversarial perturbations; it can be understood that this formula means to find the parameter θ that minimizes the risk of the model in the worst case within a neighborhood of the parameter space.
[0148] The upper bound of the factual risk is:
[0149]
[0150] Where: R S (h θ ) represents the risk (loss) of the prediction model h θ on the training set S, represents the maximum risk of the prediction model h θ+∈ on the training set S under the condition that the norm of the parameter perturbation ∈ does not exceed ρ, represents a regularization term used to control the complexity of the model, related to the norm of the model parameter θ and the perturbation range ρ, k represents the number of model parameters, i.e., the dimension of θ, n represents the size of the training set, i.e., the number of training samples, δ represents the probability bound, usually used to control the confidence level. For example, δ = 0.05 represents a 95% confidence level, represents a constant term used to capture high-order terms or terms not explicitly represented.
[0151] During the training process, an adversarial training strategy is adopted to improve the anti-interference ability of the model. Specifically, adversarial samples are generated by the Fast Gradient Sign Method (FGSM) to enhance the robustness of the model in the face of small perturbations. The adversarial samples are generated by the following formula:
[0152]
[0153] Where, r adv is the adversarial sample, which is obtained by applying a perturbation to the original input data r. r is the original sample (original input data), which is a sample in the training dataset, is the sign of the gradient of the loss function. The sign gradient represents the sign of the partial derivative direction of the loss function l(f(r), y)) of the original sample r, indicating how to adjust the perturbation size ∈ to maximize the increment of the loss function. Specifically, this part tells how to change the input data in the sample space to produce the maximum loss. l is the loss function, f(r) is the predicted output of the model, specifically, the predicted value generated based on the input data r, and y is the true value or true label, is the gradient of the loss function with respect to the features, ∈ is the perturbation size, determined by maximizing the loss increment:
[0154] ∈ = argmax ∈ [l(f(r adv ), y)) - l(f(r), y))]
[0155] To organically combine SAM optimization and adversarial training, a comprehensive optimization objective is constructed. This optimization objective consists of three important parts:
[0156] The SAM loss, by maximizing the loss within the SAM neighborhood, enhances the generalization ability of the model.
[0157] The IPM (Integral Probability Metric) of adversarial samples, used to measure the distribution difference between normal samples and adversarial samples, ensuring that the model can not only handle normal data but also resist perturbations.
[0158] The regularization term, used to control the complexity of model parameters and avoid overfitting.
[0159] Therefore, the comprehensive objective function is:
[0160]
[0161] where, is the total loss function, N represents the number of samples in the training set, i represents the i-th sample, is the training data set, t is the index of the time step, T is the maximum number of observed time steps, is the predicted loss of the result variable at the t-th time step, is the predicted loss of the time-varying covariates at the t-th time step, θ is the model parameter, ∈ is the perturbation size, ρ is the constraint range of the adversarial perturbation, represents the distribution deviation between the time series features and the adversarial sample features, is the regularization term of the model, λ 1 、λ 2 、λ 3 are the weights of each part of the loss;
[0162] During the training process, the Adam optimizer is used to optimize the above objective function. The Adam optimizer can adaptively adjust the learning rate and improve the training efficiency. The settings of the optimizer are as follows: the initial learning rate is 0.001, the batch size is 32, the total number of training epochs is 200, and the early stopping criterion is set to stop training when the validation set loss does not improve within 10 epochs.
[0163] Through the cooperation of this optimization algorithm, the model can continuously update its parameters during the training process and gradually converge to an optimal state that can handle both normal data and adversarial samples.
[0164] Step S700: Use the trained prediction model to perform counterfactual reasoning on the input data, generate counterfactual trajectories, and conduct dynamic causal effect evaluation.
[0165] The goal of counterfactual reasoning is to speculate on how the results would change if different intervention measures were implemented in the future based on historical observational data. Specifically, the potential effects of different intervention strategies are evaluated by generating counterfactual trajectories. By simulating different intervention strategies, the long-term impacts of different intervention measures on the outcome variables and covariates can be dynamically estimated, thus providing support for decision-making.
[0166] The first step in counterfactual reasoning is to perform a one-step prediction. In one-step prediction, the model directly uses the current observational history to predict the outcome variable at the next time step. Specifically, by inputting the observational history at the current time point into the trained model, the model predicts the outcome variable at the next time step through the output decoder and the corresponding high-dimensional feature representation. The calculation formula for one-step prediction is:
[0167]
[0168] where, is the output value predicted at time step T+1, G Y is the output decoder, is the observational history, including the feature representations, intervention information, and outcome variables at all time steps, Φ(·) is a feature extraction function that extracts effective features from the historical data ;
[0169] When performing multi-step prediction, a rolling prediction method is used to predict the results at multiple future time steps. The rolling prediction starts from one-step prediction and then iteratively predicts the outcome variables and time-varying covariates at each future time step. The specific steps are as follows:
[0170] First, use the historical observational data: as the initial input, where X 1 ,....,X T represent the covariates (feature variables) from time step 1 to T, Y 1 ,....,Y T represent the outcomes (or observations) from time step 1 to T, and A 1 ,....,A T represent the intervention measures from time step 1 to T. These historical data include time-varying covariates, outcome variables, and intervention measures. At the same time, specify a future intervention sequence A T+1:T+τ = a 1:τ , where a 1:τ is the intervention measure, and τ represents the time step length of the prediction. From t = T+1 to T+τ, use the following steps to iteratively predict the outcomes and covariates at future time steps:
[0171] Process the comprehensive feature representation at the current moment using the output decoder and the covariate decoder to obtain the predicted outcome variable and the time-varying covariate.
[0172] Update the historical observation data: Combine the new predicted results and covariates with the original historical data to form the new historical observation data:
[0173]
[0174] Among them, is the updated historical data, and the currently predicted covariate output and the intervention measure a t are merged into the historical data.
[0175] Through iterative prediction, a complete counterfactual trajectory is finally obtained, that is, within the future time steps from T + 1 to T + τ, based on the specified intervention sequence: A T+1:T+τ The prediction results and covariate sequences are as follows:
[0176]
[0177] This complete trajectory shows the changes in the outcome variable and the covariate within the future period under different intervention measures.
[0178] Through counterfactual reasoning, the possible impacts of different intervention strategies in the future can be simulated, and then the effects of each strategy can be evaluated. For each specified intervention sequence, the model will provide a series of prediction results, including the outcome variable and the time-varying covariate. These prediction results can help decision-makers evaluate the effects of different policies, treatments, or intervention measures, so as to select the optimal decision-making strategy in practical applications. In addition, counterfactual reasoning can also help us understand the roles of different factors in dynamic systems, reveal causal relationships, and provide theoretical support for personalized treatment or policy optimization.
[0179] Take the evaluation of the long-term effect of the economic aid policy as an example. First, collect a set of longitudinal economic and social development data of recipient countries, including:
[0180] Economic indicators: GDP, import and export volume, foreign exchange reserves, inflation rate, etc.
[0181] Social indicators: employment rate, education expenditure, medical expenditure, infrastructure investment, etc.
[0182] Political indicators: government stability, corruption perception index, legal system level, etc.
[0183] Aid records: aid type, aid amount, implementation time, project progress, etc.
[0184] Effect evaluation: Degree of people's livelihood improvement, economic growth rate, debt level, international relations, etc.
[0185] Align the above data according to the timestamp and organize it into a structured sample Among them represents the covariate of the i-th country at time t, represents the aid policy at time t, represents the effect evaluation index at time t.
[0186] Then construct a counterfactual prediction model based on the proposed TACIN framework:
[0187] Construct an Intervention-Aware Function Convolution (IFC) module, and use the historical observation sequence and the aid record as inputs, and extract temporal features through multi-layer convolution. Among them, continuous aid measures (such as aid amount) are encoded using Gaussian RBF, and discrete aid measures (such as project type) are encoded using multi-quadratic RBF to obtain the state representation of the country at time t
[0188] Design a multi-layer IFC module with residual connections for feature extraction and representation learning. When the input and output dimensions are inconsistent, perform dimension adjustment through 1×1 convolution. Construct the output decoder G Y for predicting the aid effect, the covariate predictor G X and the smoothing network J X for predicting economic and social indicators.
[0189] Adopt SAM optimization and FGSM adversarial training to improve the generalization performance of the model. SAM minimizes the maximum loss within the ρ-neighborhood of the parameter space, and FGSM generates adversarial samples to enhance robustness. Construct an overall optimization objective that includes prediction loss, adversarial loss, and regularization terms, and use the Adam optimizer for training. Set the initial learning rate to 0.001, batch size to 16, and the maximum number of training epochs to 300. Early stop when the validation set loss does not improve for 15 epochs.
[0190] Use the trained TACIN model for counterfactual prediction. For the observed historical data, evaluate the short-term effects of different aid programs through single-step prediction. By designing multiple aid sequence programs, use the rolling prediction method to evaluate the long-term effects. Compare and analyze the performance of different programs in dimensions such as economic benefits (GDP growth, industrial upgrading), social benefits (people's livelihood improvement, social stability), and diplomatic benefits (bilateral relations, international influence). Finally, optimize the aid scale, structure, and timing according to the prediction results to improve the aid efficiency and policy accuracy.
[0191] Through the counterfactual prediction ability of the TACIN framework, the long-term effects of different aid policies can be systematically evaluated, the optimal aid strategies and implementation paths can be discovered, and data support can be provided for foreign aid decision-making. This method can also be extended to other diplomatic scenarios that require the evaluation of the long-term effects of policies.
[0192] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the method are implemented.
[0193] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0194] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0195] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0196] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical disks, etc., which can store program codes.
[0197] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A causal effect dynamic estimation method, characterized in that: The steps include: Obtain a longitudinal observational data set containing time-varying covariates, static covariates, intervention variables, and outcome variables, align the longitudinal observational data set according to time series and organize it into a structured sample to generate input data; One-dimensional dilated convolution is used to jointly process the time-varying covariates and intervention variables in the input data to extract preliminary time series feature representation; Perform nonlinear mapping on the extracted preliminary time series feature representation to generate a high-dimensional time series feature representation; The high-dimensional time series feature representation is input into the multi-layer intervention-aware convolution module, and the high-dimensional time series features are deeply extracted and represented by connecting multiple layers of convolution in series to generate a comprehensive feature representation; Building a prediction model based on the comprehensive feature representation to predict the outcome variable and time-varying covariates at the next time step; The output result variables and time-varying covariates are used in combination with actual observations to train the prediction model. The loss is minimized and adversarial samples are generated through the optimization algorithm. An objective function including factual loss, adversarial loss and regularization terms is constructed. The prediction model parameters are iteratively updated to obtain the trained prediction model. The trained prediction model is used to perform counterfactual reasoning on the input data, generate counterfactual trajectories and perform dynamic causal effect evaluation.
2. The causal effect dynamic estimation method according to claim 1, characterized in that: The method of using one-dimensional dilated convolution to jointly process the time-varying covariates and intervention variables in the input data and extracting preliminary time series feature representations includes: The time-varying covariates and intervention variables in the input data are jointly processed and aligned according to the time series to form an input tensor with consistent time step dimensions; The input tensor after joint processing is convolved using a one-dimensional dilated convolution operation. During the convolution operation, the intervention variable is embedded in the convolution operation through a nonlinear mapping function. The dilation factor of the convolution kernel is set to capture the long-term dependency across time steps to generate a time series feature map. The generated time series feature graph is used as a preliminary time series feature representation.
3. The causal effect dynamic estimation method according to claim 2, characterized in that: The one-dimensional dilated convolution operation is used to perform convolution calculation on the input tensor after joint processing. The calculation formula is as follows: in, is the output of time step t, k is the convolution kernel size, and d is the expansion factor used to expand the receptive field. is the i-th convolution weight matrix, φ(a t ) is the nonlinear mapping function of the intervention, is the input of time step ti, is the bias term; The convolution output is used as the preliminary time series feature representation.
4. The causal effect dynamic estimation method according to claim 1, characterized in that: The method of performing nonlinear mapping processing on the extracted preliminary time series feature representation to generate a high-dimensional time series feature representation includes: The extracted preliminary time series feature representation is represented by using radial basis function for nonlinear mapping, and Gaussian radial basis function and multi-quadratic radial basis function are selected as basis functions; In the Gaussian radial basis function, the nonlinear mapping value is calculated based on the input features and the predefined center point, where the calculation formula of the Gaussian radial basis function is: Among them, a i is the i-th dimension of the intervention vector, c j is the jth center point, ζ is the width parameter of the Gaussian function; In the multi-quadratic radial basis function, the nonlinear mapping value is calculated based on the input features and the predefined center point, where the calculation formula of the multi-quadratic radial basis function is: Among them, σ is the shape parameter; The calculation results of Gaussian radial basis function and multi-quadratic radial basis function are combined in the form of weighted sum to generate the final nonlinear mapping value, and the formula is: Among them, φ(a) is the high-dimensional feature representation of the intervention variable a, w ij is the learnable weight, d a and d c are the intervention variable dimension and the number of central points, ψ(a i ,c j ) represents the radial basis function, which is used to calculate the nonlinear relationship between the i-th dimension of the intervention vector and the j-th center point; The nonlinear mapping results are represented as high-dimensional time series features.
5. The causal effect dynamic estimation method according to claim 1, characterized in that: Methods for generating comprehensive feature representations include: The high-dimensional temporal feature representation is input into the first layer of intervention-aware convolutional module, and the input features are convolved using a one-dimensional dilated convolution kernel, where the dilation factor is used to expand the receptive field of the time step to capture long-term dependencies; In each layer of the intervention-aware convolution module, the residual connection method is used to add the input features and the output features after the convolution operation element by element. The residual connection is used to add the input features and the output features element by element, and the calculation formula is: o=Activation(B(Z)+F(Z)) Among them, o is the temporal feature output, Activation represents the nonlinear activation function, B(Z) is the mapping operation of the input feature, when the input and output dimensions are consistent, they are adjusted by 1×1 convolution, and F(Z) is the output feature of the convolution module; Perform nonlinear activation processing on the features after the convolution operation to generate updated temporal feature representation; The updated temporal feature representation is sequentially input into the subsequent multi-layer intervention-aware convolutional module, and the above steps are repeated to perform deep extraction and representation of high-dimensional temporal features; In the last layer of intervention-aware convolutional modules, comprehensive feature representations are generated.
6. The causal effect dynamic estimation method according to claim 1, characterized in that: Based on the comprehensive feature representation, a prediction model is constructed to predict the outcome variable and time-varying covariate at the next time step, which specifically includes the following steps: The prediction model includes an output decoder and a covariate decoder; The comprehensive feature representation is input to the output decoder, and the comprehensive feature representation is mapped through a multi-layer feedforward neural network to predict the result variable of the next time step and output the predicted value. The calculation formula is: in, is the outcome variable for the next time step, G Y As the output decoder, a multi-layer feedforward neural network is used, r t is the comprehensive feature representation of time step t; The comprehensive feature representation is input into the covariate decoder and the smoothing factor generation network to predict the covariate base value and smoothing factor of the next time step respectively. The calculation formula is: η=Sigmoid(J X (r t )) in, is the time-varying covariate of the next time step to be predicted, G X is a covariate decoder using a feedforward neural network, X t is the observed covariate at time step t, η is the smoothing factor; J X For the feedforward network that generates the smoothing factor, Sigmoid is used to limit the smoothing factor to the range of [0,1]; The result variables and time-varying covariates obtained by using the output decoder and covariate decoder are used as the output for the next time step.
7. The causal effect dynamic estimation method according to claim 1, characterized in that: The prediction model is trained using the outcome variables and time-varying covariates output by the prediction model in combination with actual observations. The loss is minimized and adversarial samples are generated through the optimization algorithm. An objective function including factual loss, adversarial loss and regularization terms is constructed, and the prediction model parameters are iteratively updated. Specifically, the following steps are included: Construct a loss function, including fact loss and adversarial loss, and the calculation formula is as follows: in, is the total loss function, N represents the number of samples in the training set, i represents the i-th sample, is the training dataset, t is the index of the time step, T is the maximum number of observed time steps, is the prediction loss of the outcome variable in the tth time step, is the prediction loss of the time-varying covariate in the tth time step, θ is the model parameter, ∈ is the adversarial disturbance, ρ is the constraint range of the adversarial disturbance, Represents the distribution deviation between time series features and adversarial sample features, is the regularization term of the model, λ1, λ2, λ3 are the weights of the losses of each part; Using the SAM optimization method, the maximum loss is minimized in the ρ-neighborhood of the parameter space. The calculation formula is as follows: min θ max ||∈||≤ρ R S (h θ+∈ ) Among them, θ is the parameter set of the model, ∈ represents the perturbation size of the model parameter θ, ∈ is a vector whose norm is subject to the constraint ||∈||≤ρ, ρ is the size of the SAM neighborhood, R S is the empirical risk of the loss function on the training data, h θ+∈ For the prediction model with adversarial perturbations; The adversarial sample generation strategy is introduced, and the adversarial sample is generated using the fast gradient symbol method. The calculation formula is as follows: Among them, r adv is an adversarial sample, which is obtained by applying perturbations to the original input data r, where r is the original input data and is a sample in the training data set. is the gradient sign of the loss function. The gradient sign represents the sign of the partial derivative direction of the loss function l(f(r),y)) with respect to the original input data r, indicating how to adjust the perturbation size ∈ to maximize the increment of the loss function. l is the loss function, f(r) is the predicted output of the model, and y is the true value. is the gradient of the loss function with respect to the feature, ∈ is the perturbation size, which is determined by maximizing the loss increment; The Adam optimizer is used to iteratively update the model parameters. When the validation set loss does not improve in multiple consecutive rounds, the training is stopped.
8. The causal effect dynamic estimation method according to claim 1, characterized in that: The trained prediction model is used to perform counterfactual reasoning, generate counterfactual trajectories and conduct dynamic causal effect evaluation, which specifically includes the following steps: Input the observation history and predict the outcome variable and covariate for the next time step. The calculation formula is as follows: in, is the output value predicted at time step T+1, G Y is the output decoder, is the observation history, including the feature representation of all time steps, intervention information and outcome variables, Φ(·) is a feature extraction function, which is used to extract the historical data from the observation history. Extract effective features from Specify the future sequence of intervention variables: A T+1:T+τ =a 1:τ , the rolling prediction method is used to generate the prediction trajectory. The specific steps are: Initialize historical observations: Iterate from t = T + 1 to T + τ and perform the following steps: Output the complete counterfactual trajectory: Among them, X1,....,X T represents the covariate from time step 1 to T, Y1,...,Y T Represents the results from time step 1 to T, A1,...,A T represents the intervention measures from time step 1 to T, a 1:τ is the intervention measure, τ represents the time step of the forecast, The covariates of the current forecast are updated historical data. Output and interventionsa t Merge into historical data, F X is the covariate prediction function, based on historical data Predicting covariates, and are the time series of the predicted outcome variable and the time-varying covariates, respectively.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor is used to implement the steps of the causal effect dynamic estimation method as claimed in any one of claims 1 to 8 when executing the computer program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the causal effect dynamic estimation method according to any one of claims 1 to 8 are executed.
Citation Information
Cited By
Time sequence prediction method based on cumulative causal effect and application
CN120724098A
Training and process parameter optimization method of digital ray imaging detection prediction model
CN120932034A
CaualVAE-based yield increase measure effect evaluation method, system and equipment
CN121094338A
Electric energy loss evaluation electric energy meter based on causal reasoning driving and judgment method
CN121299220A
A method for evaluating energy loss based on causal reasoning and its determination
CN121299220B