Methods and systems for evaluating drug efficacy in observational studies with comorbidities

By calculating the contribution weights of causal parameters for individual samples and the confidence intervals of the average causal parameters of drugs in observational studies, the problem of drug efficacy assessment under the influence of comorbidities is solved, enabling accurate identification of drug efficacy in observational studies and improving assessment efficiency and accuracy.

CN117405845BActive Publication Date: 2026-03-13PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In observational studies, co-occurring events can lead to outcomes of interest that are not well defined or do not reflect the research question. Existing technologies struggle to accurately assess drug efficacy, especially when randomization is not feasible or the characteristics of the subjects do not represent the target population. In such cases, existing methods cannot effectively identify causal effects.

Method used

By acquiring observational study sample data, identifying proxy variables and covariates, calculating the contribution weight of each individual sample to the causal parameters, and using asymptotic variance to calculate the confidence interval of the average causal parameters of the drug based on the average causal parameters of the permanent survivors, the effectiveness of the drug is evaluated.

Benefits of technology

It enables accurate identification of drug efficacy in observational studies, avoids the need for randomized trials, improves the accuracy and efficiency of drug efficacy assessment, provides robust causal indicators, and reduces computation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117405845B_ABST
    Figure CN117405845B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for evaluating drug efficacy in observational studies with comorbidities. The method includes the following steps: acquiring observational study sample data, identifying proxy variables and covariates, and dividing the samples into a treatment group and a control group according to the treatment protocol; calculating the contribution weight of each individual sample to the causal parameter based on the proxy variables and covariates; calculating the mean causal parameter of the drug in the permanent survivors of the control group based on the contribution weight of each individual sample to the causal parameter; calculating the confidence interval of the mean causal parameter of the drug based on the asymptotic variance of the mean causal parameter of the drug; and determining the effect of the drug on the outcome of interest based on the confidence interval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug efficacy assessment technology, and in particular to a method and system for assessing drug efficacy in the presence of comorbidities in observational studies. Background Technology

[0002] In some clinical studies, participants may experience comorbid events, making it difficult to well define the outcome of interest (primary outcome), or the outcome may exist but not accurately reflect the research's focus. For example, if the aim is to investigate the impact of different stem cell transplant types (fully matched transplant, haploidentical transplant) on leukemia relapse rates, but a patient dies after transplantation, researchers cannot define whether that patient will relapse. A significant challenge in statistical analysis is that these outcomes of interest cannot be collected; this phenomenon is called "death cutoff." More generally, the outcome of interest is truncated by comorbid events: because the comorbid event occurs before the participant exhibits the outcome, researchers cannot measure it. It is important to note that death (comorbid event) truncation and missing data are two entirely different concepts. Missing data refers to situations where an individual's outcome actually exists, but the researcher did not record it due to loss to follow-up or other reasons; in contrast, the outcome of an individual with death (comorbid event) truncation is meaningless. Using methods for handling missing data to handle death (concurrent events) cutoff may lead to erroneous conclusions, because patients in each treatment group who have not yet died (concurrent events) at the end of the trial may have different characteristics, and direct comparisons may be affected by these potential characteristics.

[0003] Researchers typically use a potential outcome framework to study drug efficacy. A potential outcome refers to the potential result for each treatment received by each subject: the potential outcome of taking the medication, and the potential outcome of not taking the medication (or taking a placebo). Drug efficacy is the difference between the potential outcomes of taking the medication and not taking the medication; it reflects the causal effect of the drug on the outcome of interest. A suitable approach to handling the death (as a comorbid event) cutoff problem is main stratification.

[0004] Main stratification has been applied to "disrupted randomized trials" with non-compliance, i.e., scenarios where participants may not adhere to their assigned treatment. Main stratification stratifies the population according to potential comorbidities (compliance), including adherents (those who take medication when assigned, and those who do not take medication when assigned), perpetual medication takers (those who take medication regardless of assignment), non-medication takers (those who do not take medication regardless of assignment), and resistant individuals (those who do not take medication when assigned, and those who take medication when assigned, while the causal effect of interest is defined in the adherent population.

[0005] We assume the accompanying event is death. For the death cutoff problem, the main stratification approach divides the population into four layers: permanent survivors (those who would survive regardless of medication), protectors (those who would survive with medication and die without it), harmful individuals (those who would die with medication and survive without it), and permanent dead individuals (those who would die regardless of medication). The causal effect of interest is defined for the permanent survivor population because only for permanent survivors can all potential outcomes be defined; this causal measure is generally called the survivor average causal effect. For other subgroups, either the potential outcome of taking medication or the potential outcome of not taking medication does not exist, so we cannot define causal effects.

[0006] To study the average causal effect of survivors, existing techniques have constructed upper and lower bounds for it. However, these bounds are generally too wide, making inference inconvenient. To identify the average causal effect of survivors, a monotonicity hypothesis is typically introduced, meaning that individuals who would survive in the control group would also survive in the treatment group.

[0007] Furthermore, information about survival status can be used to estimate the average causal effect of survivors, but this method requires a covariate to play a crucial role. Existing techniques in randomized trials consider identifying the average causal effect of survivors based on a baseline covariate, which requires this baseline covariate to contain information from the main layer, but also requires that there be no confounding between the survival process and the outcome process.

[0008] Ethical considerations prevent randomized trials from being readily available in practice. In particular, when the monotonicity assumption—that treatment is better than no treatment (treatment reduces mortality or comorbidities)—is accepted, physicians should not discourage patients from receiving treatment. Even when randomized trials are feasible, they have stringent inclusion criteria, meaning the participants often do not accurately reflect the characteristics of the target population. In contrast, data from observational or retrospective studies are much easier to obtain. In recent years, real-world studies have received increasing attention, particularly in the area of ​​drug safety and efficacy evaluation.

[0009] However, few studies have addressed the use of master stratification to identify or estimate causal effects in observational studies. Furthermore, all the methods described above still assume that there can be no unobserved confounding between treatment and survival, which holds true in randomized trials but not necessarily in observational studies. Methods for estimating drug efficacy in randomized trials are not suitable for observational studies. Summary of the Invention

[0010] Based on the above analysis, the embodiments of the present invention aim to provide a method and system for evaluating drug efficacy in observational studies when comorbid events occur, in order to solve the problem that existing observational studies cannot evaluate drug efficacy when comorbid events occur.

[0011] On one hand, embodiments of the present invention provide a method for evaluating drug efficacy in the presence of comorbidities in observational studies, comprising the following steps:

[0012] Obtain observational study sample data, identify proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan;

[0013] The contribution weight of each sample individual to the causal effect parameter is calculated based on the proxy variable and covariate.

[0014] The average drug causal parameters in the control group of permanent survivors were calculated based on the contribution weight of each individual sample to the causal parameters.

[0015] The confidence interval of the drug's average causal effect parameter is calculated based on the asymptotic variance of the drug's average causal effect parameter, and the effect of the drug on the outcome of interest is determined based on the confidence interval.

[0016] Based on a further improvement to the above method, the contribution weight of each sample individual to the causal effect parameter is calculated based on the proxy variable and covariates, including:

[0017] Using the proxy variables and covariates as explanatory variables and the sample processing scheme as the response variable, the sample data is fitted to calculate the probability that the sample belongs to the processing group;

[0018] Using the proxy variables and covariates as explanatory variables and whether the sample did not experience a co-occurring event as the response variable, the sample data of the treatment group and the control group were fitted respectively, and the probability of no co-occurring event in the treatment group and the control group was calculated respectively.

[0019] Based on the probability that a sample belongs to the treatment group and the probability that no co-occurring event occurs, the contribution weight of each individual sample to the causal effect parameter is calculated.

[0020] Furthermore, based on the probability that a sample belongs to the treatment group and the probability that no co-occurrence event occurs, the contribution weight R(X, V) of each individual sample to the causal parameter is calculated using the following formula:

[0021]

[0022]

[0023] Where X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, e(X, V) represents the probability that the sample belongs to the treatment group, p(V|X) represents the conditional probability density of V given X, and E V|XLet X represent the expected value of V given X. P(Z=0, S=1) represents the proportion of the total sample that is in the control group and has not experienced any co-occurring events. Z represents whether the sample belongs to the treatment group and S represents whether no co-occurring events have occurred.

[0024] Furthermore, based on the contribution weight of each individual sample to the causal parameter, the mean drug causal parameter Δ for the control drug regimen in permanent survivors is calculated according to the following formula. C :

[0025]

[0026] Where R(X, V) represents the contribution weight of each sample individual to the causal effect parameter, X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, m0(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the control group, m1(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the treatment group, Z represents whether the sample belongs to the treatment group, S represents whether no co-occurrence occurs, Y represents the outcome of interest, and P(Z=0, S=1) represents the proportion of samples in the control group without co-occurrence in the total sample.

[0027] Furthermore, using the proxy variables and covariates as explanatory variables and the outcome of interest of the sample as the response variable, the sample data of the treatment group and the control group without the co-occurring event are fitted respectively, and the mean of the outcome of interest of the sample without the co-occurring event in the treatment group and the control group is calculated respectively.

[0028] Furthermore, the method also includes:

[0029] A sensitivity analysis is performed on the calculated results of the average causal effect parameter to determine its robustness.

[0030] Furthermore, a sensitivity analysis is performed on the calculated results of the average causal action parameter to determine its robustness, including:

[0031] Take different sensitivity parameters Obtain the corresponding sensitivity function

[0032] Calculate the contribution weight R(X, V) of each individual sample to the causal effect parameter under different sensitivity parameters.

[0033]

[0034]

[0035] Where X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, e(X, V) represents the probability that the sample belongs to the treatment group, p(V|X) represents the conditional probability density of V given X, and E V|X Let P(Z=0, S=1) represent the expected value of V given X, where P(Z=0, S=1) represents the proportion of the total sample that is in the control group and has not experienced any co-occurring events, Z represents whether the sample belongs to the treatment group, and S represents whether no co-occurring events have occurred.

[0036] Based on the contribution weight of each individual sample to the causal parameter under different sensitivity parameters, the mean causal parameter Δ of the drug in the control drug regimen for permanent survivors under different sensitivity parameters was calculated. C ;

[0037] If Δ under different sensitivity parameters C If the confidence intervals all have the same sign, then the calculation result of the average causal action parameter is robust.

[0038] Furthermore, determining the effect of the drug on the outcome of interest based on the confidence interval includes:

[0039] If the confidence interval includes 0, the drug to be evaluated is ineffective; if both the upper and lower bounds of the confidence interval are positive, the drug to be evaluated has a positive effect on the outcome of interest; if both the upper and lower bounds of the confidence interval are negative, the drug to be evaluated has a negative effect on the outcome of interest.

[0040] On the other hand, embodiments of the present invention provide a drug efficacy assessment system for observational studies in the presence of comorbid events, comprising the following modules:

[0041] The sample acquisition module is used to acquire sample data for observational studies, determine proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan.

[0042] The contribution weight calculation module is used to calculate the contribution weight of each sample individual to the causal effect parameter based on the proxy variable and covariate.

[0043] The causal effect parameter calculation module is used to calculate the average drug causal effect parameter in the control group of permanent survivors based on the contribution weight of each individual sample to the causal effect parameter.

[0044] The drug effect assessment module is used to calculate the confidence interval of the drug's average causal effect parameter based on the asymptotic variance of the drug's average causal effect parameter, and to determine the drug's effect on the outcome of interest based on the confidence interval.

[0045] Based on further improvements to the above system, the causal effect parameter calculation module is used to calculate the average drug causal effect parameter Δ for permanent survivors treated as control drug regimens according to the following formula. C :

[0046]

[0047] Where R(X, V) represents the contribution weight of each sample individual to the causal effect parameter, X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, m0(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the control group, m1(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the treatment group, Z represents whether the sample belongs to the treatment group, S represents whether no co-occurrence occurs, Y represents the outcome of interest, and P(Z=0, S=1) represents the proportion of samples in the control group without co-occurrence in the total sample.

[0048] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0049] 1. To address the limitation of existing technologies that can only assess drug efficacy through randomized trials, this invention proposes a technique for assessing drug efficacy using observational (retrospective) study data. This eliminates the need for randomized trials.

[0050] 2. In view of the problem that the existing technology does not take into account the influence of the occurrence of potential co-occurrence events on the outcome of interest, the present invention proposes a causal effect index with practical significance, namely the average causal effect of the survivors in the control group, to reflect the effectiveness of the drug.

[0051] 3. To address the problem that existing technologies cannot separate the influence of concomitant events when evaluating drug efficacy, this invention proposes a method that uses proxy variables to reflect the occurrence of potential concomitant events, thereby extracting the direct causal effect of the drug on the outcome of interest.

[0052] 4. In view of the problem that existing technologies use parametric models with strong limitations when dealing with co-occurring events, this invention proposes a more general semi-parametric estimation method that does not require the introduction of an overly strong parametric model and is more efficient than existing technologies in estimating the causal effects of drug efficacy.

[0053] 5. To address the problem that existing technologies cannot make inferences (due to the lack of a theoretical understanding of the causal effect estimator), this invention provides an asymptotic variance for estimating the causal effect of drug effectiveness and constructs a confidence interval for the causal effect of drug effectiveness, which significantly reduces computation time and improves the accuracy of drug effectiveness assessment.

[0054] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0055] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0056] Figure 1 This is a flowchart of a method for evaluating drug efficacy in the presence of comorbid events in an observational study according to an embodiment of the present invention;

[0057] Figure 2 This is a block diagram of a drug efficacy assessment system in the presence of comorbid events in an observational study according to an embodiment of the present invention.

[0058] Figure 3 This is a variable relationship diagram of an embodiment of the present invention;

[0059] Figure 4 Box plots showing the deviation distribution of causal action parameters estimated by different methods in embodiments of the present invention. Detailed Implementation

[0060] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0061] First, the terms / variables used in this invention will be explained:

[0062] X is a covariate (optional), which is a number of physical characteristics of the subjects collected at baseline. It generally includes several baseline variables that may affect the allocation of treatment plans, the occurrence of comorbid events, or the outcome of interest.

[0063] V is a proxy variable (required), which is a variable related to the occurrence of comorbidities in the state of not receiving treatment (natural state), and is generally a risk factor for comorbidities.

[0064] Z represents the treatment regimen (required), a binary variable. Generally, Z=1 indicates treatment, and Z=0 indicates control. That is, Z=1 indicates that the treatment regimen is the drug regimen to be evaluated, and Z=0 indicates that the treatment regimen is the control regimen. The control regimen can be placebo, other drug regimens, or no treatment.

[0065] S represents the occurrence status of the accompanying event (which is mandatory). It is a binary variable, where S = 1 indicates that no accompanying event has occurred, and S = 0 indicates that the accompanying event has occurred.

[0066] Y is the outcome of interest (required), which can be binary or continuous. It needs to be specified that Y is only meaningful when S=1, i.e., no accompanying event has occurred.

[0067] One specific embodiment of the present invention discloses a method for evaluating drug efficacy in observational studies when comorbidities occur, such as... Figure 1 As shown, it includes the following steps:

[0068] S1. Obtain observational study sample data, determine proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan;

[0069] S2. Calculate the contribution weight of each sample individual to the causal effect parameter based on the proxy variable and covariate;

[0070] S3. Calculate the average drug causal parameters of the permanent survivors in the control group based on the contribution weight of each individual sample to the causal parameters;

[0071] S4. Calculate the confidence interval of the average causal parameter of the drug using an asymptotic method, and determine the effect of the drug on the outcome of interest based on the confidence interval.

[0072] Compared with existing technologies, the drug efficacy assessment method provided in this embodiment for observational studies with comorbidities accurately identifies drug efficacy in observational studies with comorbidities by using the average causal effect of the drug in the permanent survivors of the control group as the estimation target. It does not require randomization trials, the protocol is simple and easy to implement, and improves the efficiency of drug efficacy assessment.

[0073] Within the potential outcome framework, S(Z) represents the potential co-occurrence state when the treatment is Z, and Y(Z) represents the potential outcome of interest when the treatment is Z.

[0074] It should be noted that, within the potential outcome framework, for each subject, there is a corresponding potential co-occurrence state S(1) when the treatment is Z=1 and a corresponding potential co-occurrence state S(0) when the treatment is Z=0. However, in reality, only the co-occurrence state of one treatment can be observed.

[0075] Within the potential outcome framework, the observed S and Y satisfy the following conditions:

[0076] S=Z*S(1)+(1-Z)*S(0), Y=Z*Y(1)+(1-Z)*Y(0)

[0077] In implementation, a primary stratification method was adopted, defining G = (S(1), S(0)) as the primary stratum, that is, the potential occurrence status of co-occurring events under different treatment schemes as the primary stratum. The primary stratum was divided into: S(1) = S(0) = 1 as the permanent survival group, indicating that no co-occurring events will occur regardless of whether treatment is received, denoted as LL; S(1) = 1, S(0) = 0 as the protection group, indicating that no co-occurring events will occur if treatment is received, and co-occurring events will occur if no treatment is received, denoted as LD; S(1) = 0, S(1) = 1 as the damage group, indicating that co-occurring events will occur if treatment is received, and co-occurring events will not occur if no treatment is received, denoted as DL; S(1) = S(0) = 0 as the permanent death group, indicating that co-occurring events will occur regardless of whether treatment is received, denoted as DD. The primary stratum grouping is shown in Table 1.

[0078] Main layer S(1) S(0) Y(1) Y(0) LL Forever Survival Group 1 1 Defined Defined LD protection group 1 0 Defined Undefined DL damage group 0 1 Undefined Defined DD Forever Death Group 0 0 Undefined Undefined

[0079] Existing techniques typically use Y(1)-Y(0) to measure the causal effect of a drug on the outcome of interest, but Y(1)-Y(0) is only well-defined when S(1) = S(0) = 1. The causal effect parameter for measuring drug effectiveness is...

[0080] Δ=E[Y(1)-Y(0)|S(1)=S(0)=1]

[0081] However, in observational studies, the causal parameter Δ cannot be identified because it is impossible to determine which LL or LD layer the subjects in the treatment group who did not experience comorbidities belong to, and it is impossible to calculate the respective proportions of LL and LD.

[0082] This invention proposes to

[0083] Δ C =E[Y(1)-Y(0)|S(1)=S(0)=1, Z=0]

[0084] As an estimation target, it represents the average causal effect of the drug on permanent survivors (those who would not experience any comorbidity regardless of the treatment received) in the control group. If Δ C ≠0, indicating that the drug has a causal effect on the outcome.

[0085] To estimate Δ C The following conditions are required (in which P represents probability, E represents expectation, and ⊥ represents statistical independence):

[0086] Condition 1: Potentially negligible: Z⊥(Y(1), Y(0))|X, V, S(1), S(0);

[0087] This condition means that, given the baseline covariates (including X and V) and principal components of the subjects, the treatment allocation (i.e., treatment as the evaluated drug regimen or the control regimen) is independent of the potential outcome. This condition is a fundamental assumption of causal inference. This condition cannot be statistically tested and can only be demonstrated in the context of practical application. This condition is generally considered valid when a sufficient number of covariates X are included to explain the allocation mechanism or outcome variable.

[0088] Condition 2: Monotonicity: S(1)≥S(0);

[0089] This condition means that treatment can suppress comorbid events; if a person does not experience comorbid events without treatment, then they will certainly not experience them with treatment either. This condition is usually valid in drug trials; otherwise, ethical issues arise—the drug should not be approved if monotonicity is not met.

[0090] Condition 3: Positive: 0 < P(Z = 1 | X, V) < 1, 0 < P(S(0) = 1 | Z, X, V) < 1;

[0091] This condition means that subjects have a positive probability of being assigned to each treatment regimen (treatment group and control group), and whether or not a subject experiences a comorbid event is not absolute, but rather occurs with a certain probability. This condition can be verified by examining whether both the treatment group and the control group have individuals who experience comorbid events and individuals who do not.

[0092] Condition 4: Substitutional correlation: V⊥S(0)|Z=1, S(1)=1, X;

[0093] The condition means that in the treatment group (Z=1), V is a proxy variable for S(0). This condition can be verified by checking whether V and S are correlated in the treatment group. In practice, a linear regression of S on X and V can be fitted in the treatment group, and the correlation between V and S in the treatment group can be verified by a t-test of the coefficient of V, thus verifying whether the sample data meets the condition. If not, the sample should be increased or the proxy variable should be redefined.

[0094] Condition 5: Indifference substitution: V⊥S(1)|Z=1,S(0)=0,X;

[0095] The meaning of this condition is: in the treatment group (Z=1), if it is known that S(0)=0, that is, if it is known that a comorbidity event will occur if the subject does not receive treatment (i.e., does not receive the treatment regimen), then V no longer contains information about S(1). Whether the subject experiences a comorbidity event under treatment conditions is solely due to the treatment and is unrelated to the proxy variable V. For example, V is a risk factor indicator measured at a baseline, which measures the risk of a comorbidity event occurring in the untreated natural state. If the subject receives treatment, this risk factor is blocked, thereby reducing the probability of a comorbidity event to an even lower level, the degree of reduction being unaffected by V.

[0096] Condition 6: No interaction: E[Y(1)|S(1)=1,S(0)=s,X,V]-E[Y(0)|S(1)=S(0)=1,X,V] is independent of V and holds for both S=1 and 0.

[0097] The meaning of this condition is that the causal effect of V on the potential outcome of interest Y(Z) will not interact with the causal effect of treatment regimen Z on the potential outcome of interest or the causal effect of the main layer G on the potential outcome of interest. In the perpetual survival group, the causal effect of drug effectiveness is independent of V. A special case of condition 6 is the "exclusivity constraint" condition, where V has no direct effect on the potential outcome Y(z), and the effect of V on Y(z) can only be transmitted through the main layer. This condition cannot be statistically tested and can only be demonstrated in practical application contexts. For example, in leukemia research, age is a risk factor for transplant-related mortality, but age is not a risk factor for relapse (i.e., age does not directly affect relapse, but only indirectly affects relapse through potential survival status). Therefore, if we want to study the effect of different transplantation regimens on leukemia relapse, we can take age as a proxy variable V.

[0098] Under conditions 1-6, the relationship between the variables is as follows: Figure 3 As shown (U is an unobserved confounding variable that can simultaneously affect treatment assignment, the occurrence of accompanying events, and the outcome of interest).

[0099] Figure 3 The left-hand plot shows the relationship between the observed variables Z, S, Y, V and the unobserved confounding variable U; the right-hand plot shows the relationship between the observed variables Z, V, the latent variables S(0), S(1), Y(0), Y(1) and the unobserved confounding variable U. The causal action parameter of interest is the average difference between Y(1) and Y(0) under the condition that S(0) = S(1) = 1.

[0100] If all conditions 1-6 above are met, then the contribution weight of each sample individual to the causal effect parameter is calculated based on the proxy variable and covariate. Specifically, step S2 includes:

[0101] S21. Using the proxy variables and covariates as explanatory variables and the sample processing scheme as response variables, fit the sample data to calculate the probability e(X, V) that the sample is the processing group;

[0102] In implementation, logistic regression can be used to fit the sample data. The response variable is the sample treatment plan Z, and the explanatory variables are the covariate X and the proxy variable V. The regression formula is:

[0103]

[0104] The parameters α0 and α are obtained by fitting the sample data. X α V This leads to the probability that the sample belongs to the treatment group.

[0105]

[0106] S22. Using the proxy variable and covariate as explanatory variables and whether the accompanying event did not occur in the sample as the response variable, fit the sample data of the treatment group and the control group respectively, and calculate the probability π of no accompanying event occurring in the treatment group and the control group respectively. Z (X, V);

[0107] In implementation, logistic regression can be used to fit the sample data of the treatment and control groups separately, with whether the co-occurring event S did not occur in the sample as the response variable, and the covariate X and proxy variable V as explanatory variables. The probability of no co-occurring event is denoted as π. Z (X, V), the regression formula is:

[0108]

[0109] The parameter β was obtained by fitting the sample data. Z0 ,β ZX ,β ZV This allows us to obtain the probability that no co-occurring event occurs in different groups:

[0110]

[0111] S23. Based on the probability that the sample is in the treatment group and the probability that no co-occurring event occurs, calculate the contribution weight of each individual sample to the causal effect parameter.

[0112] Specifically, the contribution weight R(X, V) of each sample individual to the causal effect parameter is calculated using the following formula:

[0113]

[0114]

[0115] Where X represents the covariate, V represents the proxy variable, π1(X,V) represents the probability that no co-occurrence occurs in the treatment group, π0(X,V) represents the probability that no co-occurrence occurs in the control group, e(X,V) represents the probability that the sample belongs to the treatment group, and E V|X Let X represent the expected value of V given X. P(Z=0, S=1) represents the proportion of the total sample that is in the control group and has not experienced any co-occurring events. Z represents whether the sample belongs to the treatment group and S represents whether no co-occurring events have occurred.

[0116] In practice, to calculate the conditional expectation involved in the weights R(X, V), the sample weights can be calculated using the following method:

[0117] Iterate through each sample and calculate the covariate X for each of the other samples. j Covariate X to the current sample i Mahalanobis distance:

[0118] d(i,j)=(X) i -X j )′cov(X) -1 (X i -X j )

[0119] Samples whose distance is less than a preset threshold are taken as the neighborhood of the current sample, denoted as set B(i), and there are Ni samples in it. The estimated R(X, V) is as follows:

[0120]

[0121] Where n represents the total number of samples.

[0122] Let Δ 1LL (X, V) = E[Y(1)|G = LL, X, V], representing the expectation of the potential outcome of interest for the permanent survivors in the treatment group;

[0123] Δ 0LL (X, V) = E[Y(0)|G = LL, X, V], representing the expected potential outcome of interest for the permanent survivors in the control group;

[0124] Under condition 6, in the permanent survival group, the causal effect of drug effectiveness is independent of V; that is, in the permanent survival group, the causal effect of drug effectiveness is a function of X. Therefore:

[0125] Δ 1LL (X, V)-Δ 0LL (X,V)=E[Y(1)-Y(0)|G=LL,X,V]

[0126] =E[Y(1)-Y(0)|G=LL,X]=Δ(X)

[0127] Under conditions 1-6,

[0128]

[0129] Where R(X, V) represents the contribution weight of each sample individual to the causal effect parameter, X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, m0(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the control group, m1(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the treatment group, Z represents whether the sample belongs to the treatment group, S represents whether no co-occurrence occurs, Y represents the outcome of interest, and P(Z=0, S=1) represents the proportion of samples in the control group without co-occurrence in the total sample.

[0130] By introducing the contribution weight of R(X,V) sample individuals to the causal effect parameter, the influence of sample individuals with S=1 in the LD layer on the causal effect parameter is eliminated, and only the influence of sample individuals with S=1 in the LL layer on the causal effect parameter is retained, so that the average causal effect of the drug in the permanent survivors in the control group can be identified.

[0131] The mean of the outcome of interest for samples in which no co-occurring events occurred was calculated as follows:

[0132] Using the proxy variables and covariates as explanatory variables and the outcome of interest of the sample as the response variable, the sample data of the treatment group and the control group without the occurrence of the co-occurring event were fitted, and the mean of the outcome of interest of the sample without the occurrence of the co-occurring event in the treatment group and the control group was calculated respectively.

[0133] In implementation, a generalized linear model can be used to fit the sample data of the treatment and control groups that did not experience comorbid events, with covariates X and proxy variables V as explanatory variables and the outcome of interest Y as the response variable. The specific fitting formula is as follows:

[0134] m z (x, v) = δ Z0 +Xδ ZX +Vδ ZV

[0135] The parameter δ was obtained by fitting the sample data. Z0 δ ZX δ ZV This allows us to obtain the mean function of the outcomes of interest for samples from different groups that did not experience the accompanying events.

[0136] The average causal parameter Δ of the drug was calculated. CThen, in step S4, the confidence interval for the average causal parameter of the drug is calculated based on the asymptotic variance, including:

[0137] First, calculate the average causal effect parameter Δ of the drug. C asymptotic variance:

[0138]

[0139] Utilizing asymptotic normality, i.e.

[0140]

[0141] Where N(0,1) is a standard normal distribution, → d This indicates convergence according to the distribution. Indicates Δ C The true value of Δ is obtained. C The 95% confidence interval is:

[0142]

[0143] Step S4, which involves determining the effect of the drug on the outcome of interest based on the confidence interval, includes:

[0144] If the confidence interval includes 0, the drug to be evaluated is ineffective; if both the upper and lower bounds of the confidence interval are positive, the drug to be evaluated has a positive effect on the outcome of interest; if both the upper and lower bounds of the confidence interval are negative, the drug to be evaluated has a negative effect on the outcome of interest.

[0145] During implementation, the p-value can also be calculated using the following formula:

[0146]

[0147] Where Φ(·) is the cumulative distribution function of the standard normal distribution, which can be obtained by looking up a table or by software calculation. In the field of pharmaceutical efficacy evaluation, the p-value is also commonly used to assess drug efficacy.

[0148] When the evaluation results are obtained, a sensitivity analysis can be performed on the calculated results of the average causal effect parameter to determine the robustness of the calculated results of the average causal effect parameter, that is, to determine how far the estimated results may differ from the actual results.

[0149] Specifically, a sensitivity analysis is performed on the calculation results of the average causal action parameter to determine the robustness of the calculation structure of the average causal action parameter, including:

[0150] Take different sensitivity parameters Obtain the corresponding sensitivity function

[0151] Calculate the contribution weight R(X, V) of each individual sample to the causal effect parameter under different sensitivity parameters.

[0152]

[0153] Where X represents the covariate, V represents the proxy variable, π1(X,V) represents the probability that no co-occurrence occurs in the treatment group, π0(X,V) represents the probability that no co-occurrence occurs in the control group, e(X,V) represents the probability that the sample belongs to the treatment group, and E V|X Let P(Z=0, S=1) represent the expected value of V given X, where P(Z=0, S=1) represents the proportion of the total sample that is in the control group and has not experienced any co-occurring events, Z represents whether the sample belongs to the treatment group, and S represents whether no co-occurring events have occurred.

[0154] During implementation, Take several values ​​of the 0 appendix that satisfy the following conditions:

[0155] Based on the contribution weight of each individual sample to the causal parameter under different sensitivity parameters, the mean causal parameter Δ of the drug in the control drug regimen for permanent survivors under different sensitivity parameters was calculated. C ;

[0156] If Δ under different sensitivity parameters C If the confidence intervals all have the same sign, then the calculation result of the average causal action parameter is robust.

[0157] That is, if in the unused sensitive parameters Below, Δ C The confidence intervals all include 0, or Δ C The upper and lower bounds of the confidence interval are both greater than 0, or Δ C If both the upper and lower bounds of the confidence interval are less than 0, the evaluation conclusion is robust and can withstand scenarios where the conditions are violated; otherwise, the conclusion is not robust, and it is recommended that the user increase the sample size or find other proxy variables.

[0158] To demonstrate the effectiveness of this invention, a simulation study example is given below.

[0159] First, generate the data. Define the exponential link function expit(x) = 1 / {1 + exp(-x)}. Let the baseline covariates X1 and X2 follow a bivariate normal distribution with a mean of 1, a variance of 1, and a correlation coefficient of 0.5. Let the baseline covariate X3 follow a uniform distribution on (0, 2). Denote the baseline covariate vector (including the intercept) as X = (1, X1, X2, X3). The proxy variable V follows a normal distribution with a mean of expit(Xg) and a variance of 4. The allocation scheme Z follows a two-point distribution with a probability of expit{(X, V)α}. If Z = 1, the participant is assigned to the treatment group; if Z = 0, the participant is assigned to the control group. In observational studies, the subjects in the treatment and control groups may come from populations with different traits; therefore, the distributions of potential comorbidities and potential outcomes of interest may differ. In the treatment group, S(0) follows a two-point distribution with probability expit{(X,V)β1}; if S(0)=1, let S(1)=1; if S(0)=0, let S(1) follow a two-point distribution with probability expit{Xγ1}. In the control group, S(0) follows a two-point distribution with probability expit{(X,V)β0}; if S(0)=1, let S(1)=1; if S(0)=0, let S(1) follow a two-point distribution with probability expit{Xγ0}. Suppose that the potential outcome of interest Y(1) follows a normal distribution with mean (X,V,S(0),S(1))δ1 and variance 1, and Y(0) follows a normal distribution with mean (X,V,S(0),S(1))δ0 and variance 1.

[0160] The parameters are set as g = (-1, 1, 0, 0), α = (0, -2, 0, 1, 1), β1 = (2, -2, -2, -2, 4), β0 = (-2, 2, 2, 2, 4), γ1 = (-3, 1, 1, 1), γ0 = (-1, -1, 1, 1), δ1 = (-1, 1, 2, 2, -2, 3, 1), δ0 = (0, 1, 2, 2, -2, 0, -2). The generated data satisfies conditions 1 to 6. From δ1 and δ0, we know that:

[0161] E[Y(1)|X,V,S(0),S(1)]=-1+X1+2X2+2X3-2V+3S(0)-2S(1)

[0162] E[Y(0)|X,V,S(0),S(1)]=X1+2X2+2X3-2V-2S(1)

[0163] so

[0164] E[Y(1)-Y(0)|X, V, S(0)=S(1)=1]=5

[0165] Furthermore, the true value of the average causal effect among the permanent survivors of the control group.

[0166]

[0167] The present invention is compared with the prior art below. Prior art 1: directly uses observed survivors (people who did not experience comorbid events) to calculate the mean difference of the outcome of interest between the treatment group and the control group; Prior art 2: based on the principal layer imputation parameter model of linear regression, which is suitable for random number experimental data, but not suitable for observational studies.

[0168] Generate 100 datasets using the data generation method described above, with 200 samples in each dataset. Within each dataset,

[0169] Z is a 200×1 vector (each element takes the value 0 or 1), representing the processing scheme;

[0170] S is a 200×1 vector (each element takes the value 0 or 1), representing the survival state (the state of the accompanying event).

[0171] Y is a 200×1 vector (when an element of S is 0, the corresponding element of Y is set to 0), representing the outcome variable;

[0172] X is a 200×3 matrix, representing covariates;

[0173] V is a 200×1 vector representing a proxy variable.

[0174] Using this data as input, estimate the mean causal effect Δ among permanent survivors of the control group as described above. C , as output.

[0175] Each method yields 100 estimates of the causal interaction parameters. The closer the estimate is to the true value, the higher the accuracy. Each estimate is defined as Δ. C Compared with the true value The difference is termed bias, therefore 100 bias values ​​are calculated for each method. Figure 4 The box plot shows the distribution of 100 bias values ​​for various methods. Here, bias.eff represents the bias of the proposed technique, bias.sc represents the bias of prior art 1, and bias.wzr represents the bias of prior art 2. The box plot shows that the bias of the proposed technique is close to 0, while the bias of the prior art is relatively large. Therefore, the proposed technique can obtain a more accurate estimate of drug efficacy.

[0176] On the other hand, embodiments of the present invention provide a drug efficacy assessment system for observational studies in the presence of comorbidities, such as... Figure 2 As shown, the system includes the following modules:

[0177] The sample acquisition module is used to acquire sample data for observational studies, determine proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan.

[0178] The contribution weight calculation module is used to calculate the contribution weight of each sample individual to the causal effect parameter based on the proxy variable and covariate.

[0179] The causal effect parameter calculation module is used to calculate the average drug causal effect parameter in the control group of permanent survivors based on the contribution weight of each individual sample to the causal effect parameter.

[0180] The drug effect assessment module is used to calculate the confidence interval of the drug's average causal effect parameter based on the asymptotic variance of the drug's average causal effect parameter, and to determine the drug's effect on the outcome of interest based on the confidence interval.

[0181] Preferably, the causal effect parameter calculation module is used to calculate the average drug causal effect parameter Δ of the permanent survivors treated as control drug regimens according to the following formula. C :

[0182]

[0183] Where R(X, V) represents the contribution weight of each sample individual to the causal effect parameter, X represents the covariate, V represents the proxy variable, π1(X, V) represents the probability that no co-occurrence occurs in the treatment group, π0(X, V) represents the probability that no co-occurrence occurs in the control group, m0(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the control group, m1(X, V) represents the mean of the outcome of interest in the samples without co-occurrence in the treatment group, Z represents whether the sample belongs to the treatment group, S represents whether no co-occurrence occurs, Y represents the outcome of interest, and P(Z=0, S=1) represents the proportion of samples in the control group without co-occurrence in the total sample.

[0184] The above-described method and system embodiments are based on the same principles, and their related aspects can be referenced from each other to achieve the same technical effects. For specific implementation processes, please refer to the foregoing embodiments, which will not be repeated here.

[0185] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0186] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for assessing drug efficacy in observational studies when comorbid events occur, characterized in that, Includes the following steps: Obtain observational study sample data, identify proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan; The contribution weight of each sample individual to the causal effect parameter is calculated based on the proxy variable and covariate. The average drug causal parameters in the control group of permanent survivors were calculated based on the contribution weight of each individual sample to the causal parameters. The confidence interval of the drug's average causal effect parameter is calculated based on the asymptotic variance of the drug's average causal effect parameter, and the effect of the drug on the outcome of interest is determined based on the confidence interval. The contribution weight of each sample individual to the causal effect parameter is calculated based on the proxy variable and covariate, including: Using the proxy variables and covariates as explanatory variables and the sample processing scheme as the response variable, the sample data is fitted to calculate the probability that the sample belongs to the processing group; Using the proxy variables and covariates as explanatory variables and whether the sample did not experience a co-occurring event as the response variable, the sample data of the treatment group and the control group were fitted respectively, and the probability of no co-occurring event in the treatment group and the control group was calculated respectively. Based on the probability that the sample is in the treatment group and the probability that no co-occurrence event occurs, the contribution weight of each sample individual to the causal effect parameter is calculated. Based on the probability that a sample belongs to the treatment group and the probability that no co-occurrence event occurs, the contribution weight of each individual sample to the causal parameter is calculated using the following formula. : ; ; in, Representing covariates, Represents a proxy variable. This represents the probability that no accompanying event occurs in the treatment group. This represents the probability that no co-occurring event occurs in the control group. This represents the probability that the sample belongs to the treatment group. Indicates a given Under the conditions The conditional probability density, Indicates a given Under the conditions of Take the expected value. This indicates the proportion of the total sample that was in the control group and did not experience any associated events. Indicates whether a sample belongs to the treatment group. Indicates whether no accompanying event occurred; Determining the effect of a drug on the outcome of interest based on the confidence interval includes: If the confidence interval includes 0, the drug to be evaluated is ineffective; if both the upper and lower bounds of the confidence interval are positive, the drug to be evaluated has a positive effect on the outcome of interest; if both the upper and lower bounds of the confidence interval are negative, the drug to be evaluated has a negative effect on the outcome of interest.

2. The method for evaluating drug efficacy in observational studies with comorbidities as described in claim 1, characterized in that, Based on the contribution weight of each individual sample to the causal parameter, the mean causal parameter of the drug in the permanent survivors treated with the control drug regimen was calculated according to the following formula. : ; in, This represents the weight of each sample individual's contribution to the causal parameter. Representing covariates, Represents a proxy variable. This represents the probability that no accompanying event occurs in the treatment group. This represents the probability that no co-occurring event occurs in the control group. This represents the mean of the outcome of interest in the control group where no comorbid event occurred. This represents the mean of the outcome of interest for the samples in the treatment group where no co-occurring event occurred. Indicates whether a sample belongs to the treatment group. This indicates whether no accompanying event occurred. Indicate interest in the ending. This indicates the proportion of the total sample that was in the control group and did not experience any associated events. This represents the probability that the sample belongs to the treatment group.

3. The method for evaluating drug efficacy in observational studies with comorbidities as described in claim 2, characterized in that, Using the proxy variables and covariates as explanatory variables and the outcome of interest of the sample as the response variable, the sample data of the treatment group and the control group without the occurrence of the co-occurring event were fitted, and the mean of the outcome of interest of the sample without the occurrence of the co-occurring event in the treatment group and the control group was calculated respectively.

4. The method for evaluating drug efficacy in observational studies with comorbidities as described in claim 1, characterized in that, The method further includes: A sensitivity analysis is performed on the calculated results of the average causal effect parameter to determine its robustness.

5. The method for evaluating drug efficacy in observational studies with comorbidities as described in claim 4, characterized in that, Sensitivity analysis is performed on the calculated results of the average causal effect parameter to determine its robustness, including: Take different sensitivity parameters The corresponding sensitivity function is obtained. ; Calculate the contribution weight of each individual sample to the causal effect parameter under different sensitivity parameters. ; ; in, Representing covariates, Represents a proxy variable. This represents the probability that no accompanying event occurs in the treatment group. This represents the probability that no co-occurring event occurs in the control group. This represents the probability that the sample belongs to the treatment group. Indicates a given Under the conditions The conditional probability density, Indicates a given Under the conditions of Take the expected value. This indicates the proportion of the total sample that was in the control group and did not experience any associated events. Indicates whether a sample belongs to the treatment group. Indicates whether no accompanying event occurred; Based on the contribution weight of each individual sample to the causal effect parameter under different sensitivity parameters, the mean causal effect parameter of the drug in the control drug regimen for permanent survivors under different sensitivity parameters was calculated. ; Under different sensitivity parameters If the confidence intervals all have the same sign, then the calculation result of the average causal action parameter is robust.

6. A system for assessing drug efficacy in observational studies with concomitant events, based on the method for assessing drug efficacy in observational studies according to any one of claims 1-5, characterized in that, Includes the following modules: The sample acquisition module is used to acquire sample data for observational studies, determine proxy variables and covariates, and divide the samples into treatment and control groups according to the treatment plan. The contribution weight calculation module is used to calculate the contribution weight of each sample individual to the causal effect parameter based on the proxy variable and covariate. The causal effect parameter calculation module is used to calculate the average drug causal effect parameter in the control group of permanent survivors based on the contribution weight of each individual sample to the causal effect parameter. The drug effect assessment module is used to calculate the confidence interval of the drug's average causal effect parameter based on the asymptotic variance of the drug's average causal effect parameter, and to determine the drug's effect on the outcome of interest based on the confidence interval.

7. The drug efficacy assessment system for observational studies in the presence of comorbidities as described in claim 6, characterized in that, The causal effect parameter calculation module is used to calculate the average drug causal effect parameter for permanent survivors treated with the control drug regimen according to the following formula. : ; in, This represents the weight of each sample individual's contribution to the causal parameter. Representing covariates, Represents a proxy variable. This represents the probability that no accompanying event occurs in the treatment group. This represents the probability that no co-occurring event occurs in the control group. This represents the mean of the outcome of interest in the control group where no comorbid event occurred. This represents the mean of the outcome of interest for the samples in the treatment group where no co-occurring event occurred. Indicates whether a sample belongs to the treatment group. This indicates whether no accompanying event occurred. Indicate interest in the ending. This indicates the proportion of the total sample that was in the control group and did not experience any associated events. This represents the probability that the sample belongs to the treatment group.

Citation Information

Patent Citations

  • Method for providing information about early diagnosis of drug-induced liver injury type

    KR1020150081632A

  • Methods and apparatus to determine a causal effect of observation data without reference data

    US20170193529A1