Computer-implemented method, computer program, and computer system (average treatment effect for paired data)
The method addresses limitations in current data collection by formulating causal inference and using propensity score reweighting to enhance data collection efficiency and accuracy for multivariate event datasets, particularly in asynchronous event occurrences.
Patent Information
- Application Number
- JP2022120133
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-28
- Filing Date
- 2022-07-28
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2042-07-28
AI Technical Summary
Current data collection techniques provide limited scope for collecting data for event datasets and struggle to identify resulting repeated occurrences within the event data, assuming discrete time in generating frameworks associated with data models.
A computer-implemented method that formulates causal inference between paired data variables within a multivariate event dataset, dynamically generating a mean-effect framework and using propensity score and inverse propensity score reweighting procedures for better ATE estimation, adjusting for the influence of other event covariates.
Improves the efficiency of data collection and reduces inaccuracy in data predictions by learning and quantifying causal effects between any pair of events in asynchronous, irregularly spaced event occurrences, providing a framework for better ATE estimation.
Smart Images

Figure 0007808403000188 
Figure 0007808403000189 
Figure 0007808403000190
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to the field of data pairing techniques, and more particularly to causal inference data collection techniques. [Background technology]
[0002] Data collection is the process of collecting and measuring information about variables of interest in an established system so that relevant questions can be answered and results can be evaluated. Data collection is a component of research in all fields of study, including the physical and social sciences, humanities, and business. Methods vary by field, but the focus remains on ensuring accurate and fair collection. In general, the goal of all data collection is to obtain high-quality evidence that can lead to the formulation of convincing and reliable answers to the questions posed. Summary of the Invention [Problem to be solved by the invention]
[0003] The goal of all data collection is to obtain high-quality evidence that can lead to the formulation of convincing and reliable answers to the questions posed. [Means for solving the problem]
[0004] According to one aspect of the present invention, there is provided a computer-implemented method comprising: identifying a plurality of data variables in a multivariate event dataset, formulating a causal inference between at least two identified data variables in the multivariate event dataset, generating a structural framework of mean effect values for the multivariate event dataset based on the formulation of the causal inference for the identified data variables, and calculating an inverse propensity score of the generated structural framework of mean effects based on the types of identified variables, a predetermined time associated with the identified variables, and the strength of the causal relationship between the identified variables. [Brief explanation of the drawings]
[0005] Preferred embodiments of the present invention will now be described, by way of example only, with reference to the following drawings:
[0006] [Figure 1] 1 illustrates a block diagram of a computing environment in accordance with an embodiment of the present invention.
[0007] [Figure 2] 1 is a flowchart illustrating operational steps for inferring connection strength between paired events, in accordance with at least one embodiment of the present invention.
[0008] [Figure 3A] 10 illustrates connection strength results in accordance with at least one embodiment of the present invention. [Figure 3B] 10 illustrates connection strength results in accordance with at least one embodiment of the present invention.
[0009] [Figure 4] 1 is a block diagram of an exemplary system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Embodiments of the present invention recognize certain deficiencies with current data collection techniques. Specifically, embodiments of the present invention recognize that data collection typically provides limited scope for collecting data for event datasets and struggles to identify resulting repeated occurrences within the event data. Current data collection techniques generally assume discrete time in generating frameworks associated with data models. Embodiments of the present invention provide a solution that improves the efficiency of data collection and reduces the inaccuracy of data predictions of current data collection techniques by formulating causal inference between paired data variables within a multivariate event dataset and dynamically generating a mean-effect framework for the multivariate event dataset based on the causal inference formulation.
[0011] Specifically, embodiments of the present invention improve on current data collection techniques by learning and quantifying causal effects between any pair of events, assuming only time-stamped, asynchronous, irregularly spaced event occurrence data on a timeline spanning multiple event types as input. Embodiments of the present invention accomplish this by formulating an average treatment effect (ATE) framework for dynamically correlated event datasets, using the framework for ATE on multivariate point processes, and generating propensity score and inverse propensity score reweighting procedures for better ATE estimation that adjust for the influence of other event covariates (i.e., events that may influence other past occurrences or later events).
[0012] Figure 1 is a functional block diagram illustrating a computing environment, generally designated computing environment 100, according to one embodiment of the present invention. Figure 1 provides only an illustration of one implementation and is not intended to imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by one skilled in the art without departing from the scope of the invention as defined in the claims.
[0013] Computing environment 100 includes client computing devices 102 and server computers 108, all interconnected via network 106. Client computing devices 102 and server computers 108 may be standalone computing devices, management servers, web servers, mobile computing devices, or any other electronic devices or computing systems capable of receiving, transmitting, and processing data. In other embodiments, client computing devices 102 and server computers 108 may represent a server computing system utilizing multiple computers as a server system, such as in a cloud computing environment. In another embodiment, client computing devices 102 and server computers 108 may be laptop computers, tablet computers, netbook computers, personal computers (PCs), desktop computers, personal digital assistants (PDAs), smartphones, or any programmable electronic device capable of communicating with various components and other computing devices (not shown) within computing environment 100. In another embodiment, client computing devices 102 and server computers 108 each represent a computing system utilizing clustered computers and components (e.g., database server computers, application server computers, etc.) that function as a single pool of seamless resources when accessed within computing environment 100. In some embodiments, the client computing device 102 and the server computer 108 are a single device. The client computing device 102 and the server computer 108 may include internal and external hardware components capable of executing machine-readable program instructions, as shown and described in further detail with respect to FIG.
[0014] In this embodiment, the client computing device 102 is a user device associated with a user and includes an application 104. The application 104 communicates with a server computer 108 to access an average treatment effect program 110 for dataset information (e.g., using TCP / IP). In this embodiment, the dataset information can be synthetic or real. For example, the synthetic dataset can include PGEM, Hawkes, or a hybrid, or a combination thereof. In some examples, the real dataset can include a diabetes dataset. The application 104 can further communicate with the average treatment effect program 110 to formulate causal inference between paired event variables in a multivariate point process, provide an average treatment effect (ATE) framework for the multivariate point process based on the formulation, propose a propensity score and inverse propensity (IP) score reweighting procedure for better ATE estimation to adjust for the influence of other event covariates, and derive and obtain equivalence of propensity scores and equivalence of ATE in a multivariate point process, as discussed in more detail in FIGS. 2-4 .
[0015] Network 106 may be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of the three, and may include wired, wireless, or fiber optic connections. Network 106 may include one or more wired or wireless networks, or a combination thereof, capable of receiving and transmitting data, audio, or video signals, or a combination thereof, including multimedia signals including voice, data, and video information. In general, network 106 may be any combination of connections and protocols that support communication between client computing device 102 and server computer 108, as well as other computing devices (not shown) in computing environment 100.
[0016] Server computer 108 is a digital device that hosts average treatment effect program 110 and database 112. In this embodiment, average treatment effect program 110 resides on server computer 108. In other embodiments, average treatment effect program 110 may have an instance of the program stored locally on client computing device 102 (not shown). In other embodiments, average treatment effect program 110 may be a standalone program or system that can formulate causal inference between paired event variables in a multivariate point process and provide a framework for average treatment effect (ATE) for multivariate point processes based on the formulation. In still other embodiments, average treatment effect program 110 may be stored on any number of computing devices.
[0017] The average treatment effect program 110 uses the event dataset to infer the strength of the causal relationship between pairwise events between event types. In this embodiment, the event dataset may refer to the occurrence of various events and event types over time. In this embodiment, an event may refer to the occurrence of a particular type of data. For example, with respect to customer transactions, an event may represent a purchase or a sale. Each occurrence may be received as part of a multivariate / marked asynchronous event stream, with each event or occurrence having a timestamp and a compound object that serves as a "mark." As used herein, a "mark" refers to the type of detail associated with an event (e.g., events that indicate a relationship, also known as a "dyadic" relationship, may be associated with an Actor 1). <action>Actor 2, etc.), can be organized hierarchically and may include location information.
[0018] The event dataset may include web logs, customer transactions, network notifications, political events, financial events, insurance claims, health information (e.g., log information), and other medical events. Specifically, in an event dataset related to health information, the event dataset may log event types such as first hospitalization, first home health visit, second hospitalization, first prescription refill, second home health visit, and third home health visit in a chronological timeline. The average treatment effect program 110 may infer the strength of the causal relationship between a pair of events (e.g., first hospitalization and first home health visit).
[0019] In this embodiment, the event dataset is:
number
number
number
[0020] In this embodiment, the average treatment effect program 110 can measure the time between events. For example, from t0 to t 20 In a timeline ranging from , there are three different event types (e.g., A, B, and C, each representing a different event) that can occur at different times. Specifically, event A could occur at time 2, followed by event B at time 3, and event C at time C. A second instance of event B could occur at time 6, a second instance of event A could occur at time 12, a third instance of event B could occur at time 13, and a second instance of event C could occur at time 20. In this example, the average treatment effect program 110 can express the events between event labels Z and X (where (Z≠X)) as follows:
number
number
number
[0021] The average treatment effect program 110 can generate causal inferences (e.g., causal effects) for event datasets without a time series (e.g., to answer how the occurrence of event A in the past affects event B in the future). In other words, the average treatment effect program 110 learns and quantifies the causal effect between any pair of events, given only time-stamped, asynchronous, irregularly spaced event occurrence data on a timeline spanning multiple event types as input. In this embodiment, the average treatment effect program 110 achieves this by formulating an average treatment effect framework for dynamically correlated event datasets and a computational method for quantifying causal effects using balanced propensity scores for dynamically correlated event datasets.
[0022] In this embodiment, the average treatment effect program 110 formulates causal inference between paired event variables in a multivariate point process by providing an average treatment effect (ATE) framework for the multivariate point process based on the formulation. For example, the average treatment effect program 110 can provide the following causal inference framework for a pair of event labels (z, y):
[0023] Treatment variable Z at time t t is a function of past occurrences of Z (i.e., Eq. 1,
number
[0024] Outcome variable Y t is a function of the future occurrence of y at t, i.e., Eq. 2
number
number
[0025] Covariate X t is expressed as Equation 3
number
[0026] In this embodiment, the average treatment program 110 converts the following definition of the average treatment effect (ATE) under the strong ignorability and overlap assumptions into Equation 4:
number
[0027] The average treatment effect program 110 then calculates the following formula, Equation 5
number
[0028] In this embodiment, the average treatment effect program 110 uses the following equation, Equation 6, to adjust for covariates, as discussed in more detail in FIG.
number
[0029] The mean treatment effect program 110 then derives and obtains the equivalence of the propensity scores and balances for the ATE in the multivariate point process, as discussed in more detail in FIG. 2, using the following formulation, Equation 7:
number
[0030] In this embodiment, the average treatment effect program 110 can actually estimate the ATE. Given a parameter W (e.g., given or using prior knowledge),
number
number
number
[0031] In this embodiment, the mean treatment effect program 110 uses the proximal event model, Equation 9
number
[0032] Therefore, in this embodiment, the ATE can be estimated using the following formulas, Equation 10 and Equation 11:
number
number
[0033] Thus, the mean treatment effect program 110 can formulate causal inference between paired event variables in a multivariate point process by generating a framework of mean treatment effects (ATE) on a multivariate point process and generating propensity score and inverse propensity score reweighting procedures for better ATE estimation that adjusts for the influence of other event covariates. Finally, the mean treatment effect program 110 can derive and obtain propensity score and equivalence of ATE in a multivariate point process.
[0034] The database 112 may represent one or more databases that store received information and provide authorized access to the average treatment effect program 110 or publicly available databases. For example, the database 112 can store received datasets or databases, or both. As previously mentioned, the dataset information can be synthetic or real. For example, a synthetic dataset may include PGEM, Hawkes, or a hybrid, or a combination thereof. In some examples, the real dataset may include a diabetes dataset. Generally, the database 112 can be implemented using any non-volatile storage medium known in the art. For example, the database 112 can be implemented using a tape library, an optical library, one or more independent hard disk drives, or multiple hard disk drives in a redundant array of independent disks (RAID). In this embodiment, the database 112 is stored on the server computer 108.
[0035] FIG. 2 is a flowchart 200 illustrating operational steps for inferring connection strength between pairs of events, according to at least one embodiment of the present invention.
[0036] In step 202, the average treatment effect program 110 identifies a number of data variables in a multivariate event dataset. As used herein, a data variable refers to a data point that may be modified in response to changes to the multivariate event dataset that occur over a given time period. A data point may be an event or the occurrence of an event in the multivariate event dataset. For example, the data points may be a number of scheduled events in an estimated data model associated with a user's medical visits. In yet another example, the duration of a car's use, the location of a building associated with a schedule, or the type of data in the dataset may also be used as a data point. In general, a data variable, also referred to as a data point or an event depending on the context, may be any other variable that changes over time.
[0037] In this embodiment, the average treatment effect program 110 identifies a plurality of data variables in the multivariate event dataset by analyzing the multivariate event dataset for data variables based on a plurality of indicator markers, obtaining at least two analyzed data variables based on the plurality of indicator markers, and identifying the obtained data variables. For example, the average treatment effect program 110 may receive a stream of information regarding a user's health information (e.g., a health dataset) according to a user's permission. In this example, the average treatment effect program 110 may identify 13 events or occurrences, classify each event into a different type, and identify the occurrences of each event type. In this example, the average treatment effect program 110 may identify three different event types (e.g., prescription refills, hospitalizations, and home health visits). In this example, the average treatment effect program 110 may identify that three prescription refills occurred, five hospitalizations occurred, and five home health visits also occurred.
[0038] The average treatment effect program 110 can also identify events in a time sequence by noting which event occurred first in a series of events in a given data set. For example, the average treatment effect program 110 can identify that a hospitalization is the first event in a data set, followed by a home health visit, and so on.
[0039] In this embodiment, the average treatment effect program 110 can use indicator markers to distinguish one event and associated event type from another respective event and event type within a dataset. Examples of indicator markers may include the type of data within a dataset, the size of the dataset, the age of the dataset, the origin of the dataset, or the accessibility of the dataset may be indicator markers. The average treatment effect program 110 can also use indicator markers to completely distinguish between different datasets.
[0040] In another embodiment, the average treatment effect program 110 identifies a treatment variable in the multivariate event dataset. As used herein, a treatment variable is defined as a function of past occurrences associated with the identified data points. For example, the treatment variable may be the frequency of occurrence of a building location (i.e., a data variable) in the multivariate event dataset. In this embodiment, the average treatment effect program 110 identifies the treatment variable using the following formula:
number
[0041] In Equation 1, the average treatment effect program 110 calculates Z as the treatment variable at a given time t, which is equal to a function of the treatment variable and a historical formulation based on past occurrences associated with the treatment variable. t In another embodiment, the average treatment effect program 110 identifies an outcome variable in a multivariate event dataset. In this embodiment, the average treatment effect program 110 defines the outcome variable as a prediction of the future occurrence of the identified variable in the multivariate event dataset. In this embodiment, the average treatment effect program 110 defines the outcome variable as a prediction of the future occurrence of the identified variable in the multivariate event dataset. In this embodiment, the average treatment effect program 110 defines the outcome variable according to the following formula:
number
[0042] In Equation 2, the average treatment effect program 110 calculates Y as the outcome variable in a given time period t, which is equal to a function of the outcome variable and the incidence of that outcome variable at a given time. t In another embodiment, the average treatment effect program 110 modifies the incidence of the outcome variable using the treatment variable. In another embodiment, the average treatment effect program 110 identifies covariate variables in a multivariate event dataset. In this embodiment, the average treatment effect program 110 defines the covariate variables as a function of past occurrences of an event label other than the treatment variable. In this embodiment, the average treatment effect program 110 modifies the covariate variables using the following equation:
number
[0043] In Equation 3, the average treatment effect program 110 calculates X as a covariate variable at a given time t, which is equal to a function of any data type other than the treatment variable and a historical formulation based on past occurrences associated with any data type other than the treatment variable. t Define
[0044] In step 204, the average treatment effect program 110 formulates a causal inference between at least two identified data variables in the multivariate event dataset. In this embodiment, the average treatment effect program 110 formulates a causal inference between the identified variables by determining the strength of the causal relationship between the identified data variables in the multivariate event dataset. For example, the average treatment effect program 110 formulates a causal inference between two event data points in the event database (e.g., a binary treatment variable at time t that occurred at least once within a window, and an outcome variable is the incidence rate of an effect label, and formulates an inference between the two variables).
[0045] In another embodiment, the average treatment effect program 110 formulates causal inference between at least two identified data variables by generating a data structure that plots each identified data variable within an estimated data model. The average treatment effect program 110 generates a data structure that plots each identified data variable by estimating multiple treatment effects (e.g., changing the flow of electricity around a damaged capacitor) between a treatment variable associated with a past occurrence (e.g., a decrease in charge) and an outcome variable associated with a different occurrence (e.g., a change in calculated charge).
[0046] In this embodiment, the average treatment effect program 110 formulates causal inferences between the identified data variables using an estimated data model to predict outcomes associated with the collected data based on the formulated causal inferences.
[0047] In step 206, the average treatment effect program 110 generates a framework of average effects for the multivariate event dataset based on the formulation of causal inference. In this embodiment, the average treatment effect program 110 generates a framework (i.e., a data model) of average effect values for the multivariate event dataset by calculating the differences between the identified variables at multiple predetermined times. As used herein, an average effect value is defined as a numerical value associated with the calculated average change in the estimated data model in response to the formulation of causal inference between the identified data variables. In this embodiment, the average treatment effect program 110 calculates the average treatment effect value using the following formula:
number
[0048] In this embodiment, the average treatment effect program 110 formulates causal inferences between the identified data variables using an estimated data model to predict outcomes associated with the collected data based on the formulated causal inferences. In Equation 4, the average treatment effect program 110 defines ATE as the average treatment effect value associated with the identified treatment variables. In this equation, the average treatment effect program 110 defines E as the effect value given a given time. t In this formula, the average treatment effect program 110 is calculated as the incidence of an identified data variable, in this case the outcome variable, taken over two predetermined periods.
number
[0049] In another embodiment, the mean treatment effect program 110 adjusts the mean effect value based on adjustments associated with the identified covariate variables. In another embodiment, the mean treatment effect program 110 formulates a causal inference between at least two identified data variables in a multivariate event dataset, where the dataset is a multivariate time-event dataset, generates a structural data framework for a mean effect value of the multivariate event dataset based on the formulation of the causal inference for the identified data variables, and calculates an inverse propensity score for the generated data framework of the mean effect based on multiple factors. In this embodiment, the mean treatment effect program 110 has a clear definition of estimation in causal inference in the ATE. Specifically, the mean treatment effect program 110 proposes a technique for obtaining the ATE by using a first computable propensity score to adjust for covariates, adjusting for the covariates using the propensity score, and using weights to account for data imbalance.
[0050] In step 208, the mean treatment effect program 110 calculates the inverse propensity score of the generated framework of mean effects based on multiple factors. In this embodiment, the mean treatment effect program 110 calculates the inverse propensity score of the generated framework of mean effects by modifying the identified variables. In this embodiment, the modifications to the identified variables may include changing the type of identified variable, changing the predetermined time associated with the identified variables, and changing the strength of the causal relationship between the identified variables.
[0051] As used herein, an inverse propensity score refers to a standardized calculation that optimizes the generated framework associated with the mean effect by strengthening causal inference between identified variables based on multiple predictive outcomes. In this embodiment, the mean treatment effect program 110 calculates the calculated inverse propensity score using the following formula:
number
[0052] In Equation 5, the mean treatment effect 110 is expressed as an inverse propensity score, which is equal to the effect value of the covariate variable in a given time period within a given window w.
number
[0053] In this embodiment, the mean treatment effect program 110 can adjust for covariates, i.e., other variables that may affect the event using countertrends. In this embodiment, the mean treatment effect program 110 adjusts for the adjustment using the following formula:
number
[0054] In the formula, the average treatment effect program 110 defines D as the difference value associated with the identified variable between a first predetermined time period and first window and a second predetermined time period and second window.
[0055] In step 210, the mean treatment effect program 110 validates the inverse propensity score based on the derivation of the equivalent propensity score. In this embodiment, the mean treatment effect program 110 validates the calculated inverse propensity score by calculating the equivalent propensity score and ensuring that the inverse propensity score and the equivalent propensity score are equal. Conversely, the propensity score refers to a standardized calculation that optimizes the generated framework. In this embodiment, the mean treatment effect program 110 validates the calculated inverse propensity score by using the following formula:
number
[0056] In Equation 7, the mean treatment effect program 110 sets both sides of the equation equal to each other to ensure that the calculation of the inverse propensity value is correct. In this embodiment, the mean treatment effect program 110 validates the inverse propensity score to improve the efficiency of predicting outcomes associated with the generated framework. For example, the mean treatment effect program 110 provides 11 pairs in which an expert determines the likelihood that a cause label will generate an effect label, divides the 11 pairs into a training dataset and a test dataset, identifies the optimal time window based on the training dataset, and then validates the inverse propensity score on a diabetes dataset that is expanded into the test dataset for validation.
[0057] In another embodiment, the average treatment effect program 110 automatically terminates operation of the generated framework in response to the validated inverse propensity score meeting or exceeding a predetermined threshold. In response to the validation of the inverse propensity score, the average treatment effect program 110 automatically terminates operation of the computing device 102 based on the generated framework.
[0058] 3A and 3B show connection strength results according to at least one embodiment of the present invention. Generally, FIGS. 3A and 3B show experimental setups between synthetic and real datasets. In this example, the criteria are set with the following conditional strength score formulas (CIS Formula 1 and CIS Formula 2):
[0059] CIS expression 1:
number
[0060] CIS expression 2:
number
[0061] 3A depicts the results of a synthetic dataset. In this embodiment, the synthetic dataset includes PGEM and Hawkes. In other embodiments, the mean treatment effect program 110 can use any number of synthetic datasets.
[0062] Specifically, Figure 3A shows table 302. Table 302 shows results for a synthetic dataset including PGEM1, 2, and Hawkes1, 2, and 3, respectively, each showing various window sizes (10, 20, 30, 40, 60 for PGEM, and 10, 15, 20, 30 for Hawkes).
[0063] Figure 3B shows the results of a real dataset. In this embodiment, the real dataset used relates to a medical condition. Specifically, Figure 3B shows the results of diabetes, showing various methods and their results. [Other comments and / or embodiments]
[0064] Causal inference and discovery from observational data have been widely studied across multiple fields. However, most previous studies have focused on independent and identically distributed (i.i.d.) data. Some embodiments of the present invention propose a formulation for causal inference between paired event variables in multivariate recurrent event streams by extending Rubin's framework of average treatment effect ("ATE") and propensity scores to multivariate point processes. Similar to joint probability distributions representing i.i.d. data, multivariate point processes represent data that cause asynchronous and irregularly spaced occurrences of various types of events on a common timeline. Some embodiments of the present invention theoretically justify our point process causal framework and demonstrate how to obtain unbiased estimates of the proposed measure. Some embodiments of the present invention conduct experimental investigations using synthetic and real-world event datasets, and the proposed causal inference framework is shown to perform well on a set of criteria pairwise causality scores. [introduction]
[0065] It is widely known that the gold standard for effective causal inference is through the use of intervention data, such as randomized controlled trials, deployed to measure the impact of some treatment on an outcome of interest. However, intervening in a system can often be impractical or even impossible, in which case causal analysis must be performed using only observational data.
[0066] Observational data in the form of multivariate event streams are readily available in several domains, including health, finance, retail, sales, and maintenance. Similar to how observations in an iid dataset can be viewed as samples from a joint distribution over a set of random variables, multivariate event streams can be viewed as samples from a multivariate point process over a set of event labels, capturing the interdependent dynamics of event arrival, with the occurrence rate of an event label dependent on the prior past occurrence of a subset of event labels. Modeling, fitting, and predicting future occurrences given multivariate temporal event streams is an important area of research in statistics, explored in data mining for both pattern extraction and prediction, and more recently in machine learning. Recent research in this area leverages advances in sequential deep learning of event datasets. Other models of temporal point processes include Poisson nets, nonhomogeneous Poisson processes, Poisson cascades, piecewise constant conditional intensity models, and proximal graphical event models.
[0067] There is a vast body of related research in survival analysis, which studies paired events with continuous covariates, including both continuous and discrete variables and short event streams where outcomes typically occur only once. While a general theory of counterfactual inference for point processes has been developed, practical algorithms and models focus on continuous covariates and hazard models. Furthermore, there is related research that avoids event interventions using counterfactual Gaussian processes with behavior. When the covariates are only discrete events, existing models and estimation methods cannot be easily applied to the problem. Furthermore, related research on dynamic treatments crucially motivates later research on continuous-time data, which assumes time as a discrete unit and aims to reduce bias or high variance in the time discretization.
[0068] Some embodiments of the present invention focus on settings where data are in the form of long event streams with multiple occurrences of various types of events, including both treatment and outcome events, and where a version of the pairwise causal inference problem is fixed to a multivariate point process. For example, many people are interested in knowing whether taking a particular medication will increase a patient's chances of recovery, or whether an earthquake in Japan will lead to major market changes in the near future.
[0069] Causal inference involves drawing conclusions about causal relationships between potential causes and effects. Unlike Granger causal graph learning, some embodiments of the present invention focus on causal inference between pairs of event labels observed in event stream data. In particular, some embodiments of the present invention pose the following causal inference problem for multivariate point processes: "How can we meaningfully measure the causal relationship between a cause event label z and an effect event label y?" Such a causal relationship measurement reveals whether the effect label y is amplified, suppressed, or unaffected by the cause label z, while taking into account the potential influence of all other event labels x. This problem deviates from typical causal inference settings in several ways. First, most causal inference methods assume a set of i.d. observations. In some settings, events may be correlated over time, and therefore a certain independence must be assumed to identify the effect of interest. Second, in most causal inference settings, some embodiments of the present invention are interested in the expected value of a single observable outcome, such as mortality rate. However, in some settings, some embodiments of the present invention are interested in repeated occurrences of such outcomes over time and are interested in event frequencies (or, more precisely, instantaneous intensities), and therefore existing methods for estimating causal effects in terms of expected values are not applicable.
[0070] Note that similar causal inference problems are primarily easier to deal with random variables on i.e., i.e., i.d. data, because it is not (theoretically) difficult to calculate the probability of a pair of random variables by marginalizing a set of other random variables. However, because the past occurrence of an event z can have a complex dynamic effect on the occurrence of another event y at any time in a multivariate point process, there are several challenges that must be addressed before the causal effect of z on y can be studied. [background]
[0071] A multivariate event stream (or event dataset) is a series of events,
number
[0072] A multivariate event stream can be regarded as a sample from a multivariate point process that associates each label with a counting process. This is the past stream of event labels up to t, i.e., [Number] the conditional intensity function that represents the occurrence rate of an event of type y at time t when [Number] is used. In a multivariate point process, the probability of observing y as the next event at time t is [Number] is as follows. Regarding Equation 8, t n is the most recent event occurrence time before t.
number
number
[0073] Previous work has proposed the concept of process independence between event labels to characterize the relationship between label counting processes. The basic idea is that the intensity of a type of event does not depend on specific past events, given knowledge of certain other past events. This is an asymmetric concept similar to Granger causality. Informally,
number
[0074] Using a minimal graphical representation, we can define the direct causes of a multivariate point process similar to the direct causes of a causal network: an event label z is the direct cause of label y if z belongs to the minimal set of nodes u, and thus y is a process independent of all other labels given u. [Causal inference in multivariate point processes]
[0075] Some embodiments of the present invention introduce an extension of the Neyman-Rubin potential outcome causal inference framework that includes average treatment effects (ATE) to study how event label z affects event label y in a multivariate point process. In this class of models, treatments are represented as z∈{0,1}, where 0 is considered control and 1 is treatment. For each z, the potential outcome y z is the outcome when treatment z is applied. The main difficulty with causal inference arises from the fact that we only observe outcomes from the treatment administered, and no other outcomes. Therefore, it is sometimes considered a missing data problem.
[0076] Some embodiments of the present invention estimate the treatment effect between a treatment variable associated with past occurrences of z and an outcome variable (or response) associated with occurrences of y, under certain assumptions. The covariate variable x includes past occurrences of labels other than z, i.e., x = L\z. Some embodiments of the present invention assume that all variables are observed and therefore causal sufficiency is satisfied. Some embodiments of the present invention first define the treatment, outcome, and covariate variables in the context of a multivariate point process and underlying assumptions, and then derive a propensity score. Treatment, outcome, and covariate definitions
[0077] There are many possible ways to define treatments, outcomes, and covariates in a multivariate point process. Some embodiments of the invention start with the following general formulation: Some embodiments of the invention use Z to represent the treatment variable at time t. t to distinguish it from the event label z.
number
number
[0078] General formulation of causal inference: Given a pair of event labels (z, y), the treatment variable Z at time t t is a function of past occurrences of z, i.e.,
number
number
number
[0079] Some embodiments of the present invention are
number
[0080] Some embodiments of the present invention may be implemented using the future occurrence of y from time t.
number
number
number
number
number
number
[0081] Recent past formulation of causal inference: Given a pair of event labels (z,y), a binary treatment variable at time t
number
number
number
[0082] Some embodiments of the present invention summarize the assumptions in this formulation of causal inference as follows: 1) Events prior to t−w do not affect the rate of occurrence of y at time t. This allows for memory in time and provides a compact representation of the past. 2) Only occurrences of z within a window affect the rate of y at time t, regardless of the number of occurrences. 3) A particular time of occurrence of z does not further affect the rate of y at time t. Such a model can be robust to outliers or noisy past observations. [Definition of average treatment effect]
[0083] To measure how label y responds to past occurrences of z, the average treatment effect (ATE) can be extended to a multivariate point process formulation. Some embodiments of the present invention provide a method for determining treatment assignment.
number
number
number
number
number
[0084] The average treatment effect (ATE) for the 285 paired events was
number
number
number
number
number
number
number
number
number
[0085] Therefore, the unobserved half of the severity rate, i.e., the fact
number
number
number
number
number
[0086] The above theory also implies that the ATE of (z,y) is 0 if z is not a parent of y in the graphical event model representation of the underlying multivariate point process. To use the ATE as defined in Equation 9, there are several assumptions that must hold in order to mimic a randomized trial to truly establish causality. The ignorability condition is that whether y is 0 or 1 at each time t depends on the probability that y is 0 or 1 at that time.
number
number
number
number
number
number
number
number
[0087] The propensity score has been proposed to mimic randomized studies by resolving differences in covariates in non-randomized experiments. The propensity score is a balanced score, and conditional on a given balanced score, the distribution of observed covariates is similar between the treatment and control groups. The propensity score provides a way to summarize covariate information related to treatment selection, making direct comparisons between treatment and non-treatment groups more meaningful.
[0088] Some embodiments of the present invention provide a set of balanced scores,
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0089] All past covariates at time t
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0090] Some embodiments of the present invention relate the equilibrium score and the propensity score to estimates of causal event pairs. The ATE definition in Equation 10 takes into account the conditional intensity function. However, the treatment
number
number
number
number
number
number
number
number
number
number
number
number
number
[0091] Immediately thereafter, the two sampling processes provide unbiased estimates of the ATE for pairwise matching and subclassification techniques to adjust the propensity score. However, in practice, pairwise matching is difficult to perform given the continuity of time t, where the sampling size becomes infinite. Subclassification is also difficult when the number of covariate event labels is large, leading to a large number of classes and increasing the sample requirements for T. Therefore, we next propose a typical inverse propensity score weighting procedure for calculating the ATE for multivariate point processes. [ATE score estimation]
[0092] To adjust for propensity scores and obtain the ATE, several parameters must be provided or estimated, including the treatment definition window size, the conditional intensity rate of the treatment, and the conditional intensity rate of the outcome.
number
number
number
[0093]
number
number
number
number
number
number
number
number
number
number
number
number
number
[0094] Given a recent past view in the treatment definition, some embodiments of the present invention may use the same recent past formulation to
number
[0095] If the parents U of all nodes V are known, the log-likelihood of a multivariate point process given a PGEM is given by the times and periods in the data, as well as the PGEM
number
number
number
number
number
number
[0096] Some embodiments of the present invention set z as the parent of y in PGEM, and then calculate the intensity when z occurs and does not occur in a given proximal window w:
number
number
number
number
[0097] In a causal inference framework, the idea of weighting samples is simple: if the populations of the treated and control event datasets are different,
number
number
number
number
number
number
number
[0098] Some embodiments of the present invention propose a method for inverse probability weighting of events. Some embodiments of the present invention estimate the propensity score and then weight w for all t. t Some embodiments of the present invention then estimate the factual and counterfactual outcomes.
number
number
number
number
number
[0099] One common problem with IPTW is that the propensity score for some time unit t can be very close to 0, indicating that Z is very unlikely to occur in the window [tw,t). This can occur in any continuous timeline sampling procedure. Therefore, the weights for these t's can become very large, leading to unstable estimation. To address this issue, stabilized IPTW uses the marginal probabilities of treatments to counteract such instability. This is
number
[0100] Evaluating causal inference algorithms is more challenging than evaluating algorithms for forecasting tasks because observational datasets rarely contain ground truth treatment effects. To this end, most experiments in the literature analyze causal models using synthetic datasets where the ground truth is known. Some embodiments of the present invention begin by comparing the ATE estimation performance of the proposed IPTW method on three synthetic event datasets generated using different conditional intensity functions. Following standard practice in the causal inference literature, we measure each method's ATE accuracy performance using the root mean square error ("RMSE"), along with its standard deviation.
[0101] For comparison, the standard for using event rates as outcomes of multivariate point processes is not well established. y Since is not directly observable, a simple adaptation of ATE from the iid case does not work. Therefore, we propose two criteria scores that fit parametric models to intensity rates, CI (conditional intensity), and CIM, each considering a single parent event and setting. For the first criteria method, some embodiments of the present invention consider the association between a pair of events (z, y) and assume that the intensity of y depends only on whether z occurred at least once within a specified time window w. Therefore, some embodiments of the present invention use a conditional intensity score to estimate the causal effect of z on y.
number
number
number
[0102] Some embodiments of the present invention compare CI scores with three versions of ATE estimation based on one proposed IPTW method. First, some embodiments of the present invention calculate the ATE without weighting as in Equation 9, and this approach is abbreviated as IP-NW. Then, some embodiments of the present invention use the proposed IPTW with unstable weights ("IP-NS"), with weights according to Equation 13. Finally, some embodiments of the present invention calculate the ATE using IPTW with stable weights ("IP-Stable").
[0103] Some embodiments of the present invention first generate event data that conforms to the assumption of proximal intensity functions. Some embodiments of the present invention generate three models with different numbers of events, randomly generated graph structures between events, a fixed window size of w=30, T=2000, and random intensities between 0.1 and 0.4. Some embodiments of the present invention use the data and the generated models to calculate the underestimated lambda at a selected time ts.
number
[0104] Some embodiments of the present invention use existing toolboxes to test the approach on synthetic multivariate Hawkes process datasets. Some embodiments of the present invention generate three datasets with 30, 40, and 50 event labels with a ground truth window w=15. Some embodiments of the present invention use a fixed base rate of 0.016, with each parent event contributing an additional spike of 0.06 to the base rate with an exponential decay rate of 0.15. Some embodiments of the present invention generate an event stream with T=2000. Counterfactual rate
number
[0105] Some embodiments of the present invention also generate synthetic hybrid datasets that combine a Hawkes process-like proximal graphical event model and the idea of additive excitation with a constant kernel. IP-NW outperforms the CI score in all but one case, and IP-Stable exhibits the lowest RMSE in all but two cases.
[0106] Some embodiments of the present invention test the proposed method on a diabetes dataset (a real-world dataset that processes events of change in diet, physical activity, insulin dose, and blood glucose measurements of 70 diabetic patients). Some embodiments of the present invention treat the evaluation as ground truth, where experts provided 11 pairs such that the cause label is more likely to generate the effect label. Because the evaluation is only partial and does not provide the true ATE, this experiment measures performance using hits@K among the highest absolute values of the estimated ATE, a common metric in information retrieval. Specifically, some embodiments of the present invention determine how many of the 11 pairs are completed by the method's top-K absolute scores. The dataset is split into a 50% / 50% training / test set, and the optimal window setting is determined on the training set. The dataset is then expanded to the test set for evaluation. During training, days with w={0.1, 0.3, 0.5, 1} were considered for all models. [Conclusion]
[0107] Some embodiments of the present invention propose a framework for pairwise causal relationships in multivariate point processes. Some embodiments of the present invention formulate the problem similar to Rubin's causal inference framework and propose definitions for treatment, outcome, and propensity score. Some embodiments of the present invention estimate the average treatment effect using the proposed propensity score weighting procedure and demonstrate that the average treatment effect achieves the best performance relative to the criterion. Some embodiments of the present invention bridge causal inference using multivariate point processes and show promising performance in estimating pairwise causal relationships between events. Future research could explore efficient estimation techniques for ATE without sampling and more general problem settings defined with various past representations. It would also be interesting to explore other estimators, such as the area under the intensity rate curve over time. [Ethics statement]
[0108] Causal inference is a fundamental research technique for inferring how two variables, in this case two event variables, are causally related to each other. When considering the event stream datasets that are the focus of this paper, this research can be thought of primarily as a machine learning technique for processing event stream datasets. There are many potential applications of this research, such as modeling news events and user behavior patterns over time. However, because it is difficult to further infer the impact on potential downstream applications, we focus on the broader impact only from an algorithmic and theoretical perspective.
[0109] The main contribution of this paper is to advance the modeling and understanding of event-pair relationships through a causal inference framework. It is possible to more accurately estimate the causal influence between two events, and therefore its applications should be broader if the domain conforms to the assumptions inherent in this framework. This improved understanding should bring event stream modeling closer to reality. This work is, in fact, considered a potentially important step toward reducing spurious correlations, bias, and misinterpretation. However, caution is required when applying the proposed model to any application, especially with regard to validating assumptions. Failure to do so may lead to misidentifications that cause misunderstandings and errors, many of which may lead to erroneous conclusions. Further steps to carefully validate the results are necessary to avoid serious downstream impacts.
[0110] 4 illustrates a block diagram of computing system components within computing environment 100 of FIG. 1 in accordance with embodiments of the present invention. It should be understood that FIG. 4 is intended to provide only an illustration of one implementation and is not intended to imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment are possible.
[0111] The programs described herein are identified based on the applications for which they are implemented in particular embodiments of the invention. However, it should be understood that any particular program nomenclature herein is used merely for convenience, and thus the present invention should not be limited to use with only any particular application identified and / or implied by such nomenclature.
[0112] Computer system 400 includes a communications fabric 402 that provides communications between cache 416, memory 406, persistent storage 408, communications unit 412, and input / output (I / O) interface 414. Communications fabric 402 can be implemented with any architecture designed to pass data and / or control information between processors (such as microprocessors, communications and network processors), system memory, peripheral devices, and any other hardware components in the system. For example, communications fabric 402 can be implemented using one or more buses or crossbar switches.
[0113] Memory 406 and persistent storage 408 are computer-readable storage media. In this embodiment, memory 406 includes random access memory (RAM). In general, memory 406 may include any suitable volatile or non-volatile computer-readable storage medium. Cache 416 is a high-speed memory that improves performance of computer processor 404 by holding recently accessed data and data near recently accessed data from memory 406.
[0114] The average treatment effect program 110 (not shown) may be stored in persistent storage 408 and memory 606 for execution by one or more of the respective computer processors 404 via cache 416. In one embodiment, persistent storage 408 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 408 may include a solid-state hard drive, a semiconductor storage device, read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0115] The media used by persistent storage 408 may also be removable. For example, a removable hard drive may be used for persistent storage 408. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage medium that is also part of persistent storage 408.
[0116] In these examples, the communications unit 412 provides for communication with other data processing systems or devices. In these examples, the communications unit 412 includes one or more network interface cards. The communications unit 412 may provide communication using either or both physical and wireless communications links. The average treatment effect program 110 may be downloaded to persistent storage 508 via the communications unit 412.
[0117] The I / O interface 414 allows for the input and output of data with other devices that may be connected to the client computing device, the server computer, or both. For example, the I / O interface 414 may provide connection to external devices 420, such as a keyboard, keypad, touchscreen, or other suitable input device, or a combination thereof. The external devices 420 may also include portable computer-readable storage media, such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to implement embodiments of the present invention, such as the average treatment effect program 110, may be stored on such portable computer-readable storage media and loaded into persistent storage 408 via the I / O interface 414. The I / O interface 414 also connects to a display 422.
[0118] Display 422 provides a mechanism for displaying data to a user and may be, for example, a computer monitor.
[0119] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0120] A computer-readable storage medium may be any tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or groove-embossed structures having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over an electrical wire.
[0121] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0122] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, and traditional procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0123] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored therein comprises a product containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0125] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0126] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, having one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0127] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. The terms used herein have been selected to best explain the principles of the embodiments, practical applications, or technical improvements beyond those found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.< / action>
Claims
1. A computer-implemented method executed by a computer, comprising: identifying, by the computer, a plurality of data variables in a multivariate event dataset; formulating, by the computer, a causal inference between at least two identified data variables in the multivariate event dataset; generating, by the computer, a structural framework of mean effect sizes for the multivariate event dataset based on the formulation of the causal inference of the identified data variables; calculating, by the computer, an inverse propensity score of the generated structural framework of the mean effect size based on the types of identified variables, the predetermined time associated with the identified variables, and the strength of the causal relationships between the identified variables; 1. A computer-implemented method comprising:
2. identifying the plurality of data variables analyzing, by the computer, the multivariate event dataset for a data variable based on a plurality of indicator markers; using a plurality of scanning devices, the computer identifying at least two analyzed data variables based on average treatment effect and trend values; obtaining, by the computer, the at least two analyzed data variables based on a positive agreement rate that meets or exceeds a predetermined threshold of change associated with the plurality of indicator markers, the obtained data variables being useful for formulating causal inference within the multivariate event dataset; The computer-implemented method of claim 1 , comprising:
3. formulating the causal inference between the at least two identified data variables, generating, by the computer, a data structure that plots each identified data variable associated with the multivariate event dataset within an estimated data model; predicting, by the computer, outcomes associated with data collected using the generated data structure by estimating multiple treatment effects between treatment variables associated with past occurrences and outcome variables associated with different occurrences; formulating, by the computer, the causal inference between the at least two identified variables based on estimates of the plurality of treatment effects between the treatment variable and the outcome variable; The computer-implemented method of claim 1 , comprising:
4. generating the structural framework of the mean effect size, said computer calculating a difference between said at least two identified data variables at a plurality of predetermined times; The computer-implemented method of claim 1 , comprising:
5. calculating the inverse propensity score of the generated framework, modifying, by the computer, the at least two identified variables, wherein the modification may change the type of the identified variables, the predetermined time associated with the identified variables, and the strength of the causal relationship between the identified plurality of data variables. The computer-implemented method of claim 1 , comprising:
6. The computer-implemented method of claim 1 , further comprising the computer validating the inverse propensity score based on a derivation of an equivalent propensity score.
7. 7. The computer-implemented method of claim 6, further comprising the computer automatically terminating operation of the generated framework in response to the validated inverse propensity score meeting or exceeding a predetermined threshold.
8. formulating, by the computer, a causal inference between at least two identified data variables in a multivariate event dataset, wherein the dataset is a multivariate time-event dataset; generating, by the computer, a second structural data framework of mean effect sizes for the multivariate event dataset based on the formulation of the causal inference of the identified data variables; calculating, by the computer, an inverse propensity score of the second structural data framework from which the mean effect size was generated based on a plurality of factors; The computer-implemented method of claim 1 , further comprising:
9. On the computer, a procedure for identifying multiple data variables in a multivariate event dataset; formulating a causal inference between at least two identified data variables in the multivariate event dataset; generating a structural framework of mean effect sizes for the multivariate event dataset based on the formulation of the causal inference for the identified data variables; calculating an inverse propensity score of the generated structural framework of the mean effect value based on the type of identified variables, the predetermined time associated with the identified variables, and the strength of the causal relationships between the identified variables; A computer program that executes
10. the step of identifying a plurality of data variables comprises: analyzing the multivariate event dataset for a data variable based on a plurality of indicator markers; identifying at least two analyzed data variables based on average treatment effects and trend values using a plurality of scanning devices; obtaining the at least two analyzed data variables based on a positive agreement rate that meets or exceeds a predetermined threshold of change associated with the plurality of indicator markers, the obtained data variables being useful for formulating causal inference within the multivariate event dataset; 10. The computer program of claim 9, comprising:
11. formulating the causal inference between the at least two identified data variables, generating a data structure that plots each identified data variable associated with the multivariate event dataset within an estimated data model; predicting outcomes associated with data collected using the generated data structure by estimating multiple treatment effects between treatment variables associated with past occurrences and outcome variables associated with different occurrences; formulating the causal inference between the at least two identified variables based on the estimates of the plurality of treatment effects between the treatment variable and the outcome variable; 10. The computer program of claim 9, comprising:
12. generating the structural framework of the average effect size, calculating the difference between said at least two identified data variables at a plurality of predetermined times; 10. The computer program of claim 9, comprising:
13. The step of calculating the inverse propensity score of the generated framework comprises: modifying the at least two identified variables, wherein the modification may change the type of the identified variables, the predetermined time associated with the identified variables, and the strength of the causal relationship between the identified data variables.
10. The computer program of claim 9, comprising:
14. The computer, A procedure for validating the inverse propensity score based on the derivation of an equivalent propensity score.
14. A computer program according to any one of claims 9 to 13, further comprising:
15. The computer, automatically terminating operation of the generated framework in response to the validated inverse propensity score meeting or exceeding a predetermined threshold. The computer program of claim 14 , further comprising:
16. one or more computer processors; one or more computer-readable storage media; program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors; 1. A computer system comprising: program instructions for identifying a plurality of data variables in a multivariate event dataset; program instructions for formulating a causal inference between at least two identified data variables in the multivariate event dataset; program instructions for generating a structural framework of mean effect values for the multivariate event dataset based on the formulation of the causal inference of the identified data variables; and program instructions for calculating an inverse propensity score of the generated structural framework of the mean effect value based on the type of identified variables, a predetermined time associated with the identified variables, and the strength of the causal relationships between the identified variables. A computer system having:
17. the program instructions for identifying the plurality of data variables: program instructions for analyzing the multivariate event dataset for a data variable based on a plurality of indicator markers; program instructions for identifying at least two analyzed data variables based on average treatment effect and trend values using a plurality of scanning devices; and program instructions for obtaining the at least two analyzed data variables based on a positive agreement rate that meets or exceeds a predetermined change threshold associated with the plurality of indicator markers, the obtained data variables being useful for formulating causal inference within the multivariate event dataset.
17. The computer system of claim 16, comprising:
18. the program instructions for formulating the causal inference between the at least two identified data variables, program instructions for generating a data structure that plots each identified data variable associated with the multivariate event dataset within an estimated data model; program instructions for predicting outcomes associated with data collected using the generated data structure by estimating multiple treatment effects between treatment variables associated with past occurrences and outcome variables associated with different occurrences; program instructions for formulating the causal inference between the at least two identified variables based on estimates of the plurality of treatment effects between the treatment variable and the outcome variable; 17. The computer system of claim 16, comprising:
19. the program instructions for generating the structural framework of the average effect value include: program instructions for calculating differences between said at least two identified data variables at a plurality of predetermined times; 17. The computer system of claim 16, comprising:
20. the program instructions for calculating the inverse propensity score of the generated framework include: program instructions for modifying the at least two identified variables, the modification may change the type of the identified variables, the predetermined time associated with the identified variables, and the strength of a causal relationship between the identified data variables; 20. A computer system according to any one of claims 16 to 19, comprising:
Citation Information
Patent Citations
Analysis system and analysis method
JP2014225177A
Method and system for measuring effectiveness of user treatment
US20160055320A1
System and process to determine the causal relationship between advertisement delivery data and sales data
US20200202382A1