Method for estimating effectiveness of intermediate treatment
Through neural network models, analyzing clinical covariates and comparing the output of received and untreated patients has been solved, and the problem of difficulty in accurately estimating the effectiveness of medical treatment in the prior art is achieved, and accurate prediction of treatment effects and more effective treatment options are achieved.
Patent Information
- Application Number
- CN202411937141.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-21
- Filing Date
- 2020-02-13
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to accurately estimate the effectiveness of medical treatment, especially when dealing with complex clinical covariates and response indicators.
The therapeutic effect of medical treatment is estimated by analyzing the clinical covariates of patients by applying neural network models and comparing the output of patients who have been treated with medical treatment with those who have not been treated. Specific steps include receiving the data set from the database, training the neural network model, determining the subset of covariate quantities of the patients receiving and not receiving treatment, and comparing the associated probabilities.
Accurate prediction of the effectiveness of medical treatment is achieved, helping the medical system to provide more effective treatment options, reduce side effects, reduce treatment costs, and improve patients' recovery speed.
Smart Images

Figure CN120032773A_ABST
Abstract
Description
This application is a divisional application of the Chinese patent application with application number 202080017760.7 (application date: February 13, 2020, invention name: Method for estimating the effectiveness of intermediate treatment). Background Art
[0001] Responder analysis is performed in clinical data analysis to examine the therapeutic effect of an investigational product or medical practice to determine how patients respond to treatment. When a patient's response to treatment exceeds a threshold, the patient is considered a responder. Summary of the invention
[0002] Implementations of the present disclosure include computer-implemented methods for estimating the therapeutic effect of a medical treatment. These implementations perform this estimation by applying a neural network model to a patient's clinical covariates (e.g., blood pressure, heart rate, body temperature, etc.) and comparing the outputs of patients who have used the medical treatment to the outputs of patients who have not used the medical treatment. The effectiveness of the medical treatment on the patient can be determined by a response indicator (e.g., a biomarker) that indicates how the patient's body responds to the medical treatment.
[0003] In some implementations, the method includes: receiving a data set from a database, the data set including a set of covariate vectors and a set of response indicators, each covariate vector including clinical covariates of a corresponding patient, each response indicator being associated with a corresponding covariate vector, wherein the response indicators in the set of response indicators vary over a range; receiving from the database a plurality of segmentations of the response indicators over the range to define a plurality of response categories for the response indicators; converting each response indicator into a corresponding one-hot encoded vector based on the response categories indicated by the plurality of segmentations; training a neural network model based on each covariate vector and the corresponding one-hot encoded vector associated with the covariate vector, the neural network model utilizing a nonlinear activation function and a loss function function; and estimating a treatment effect of the medical treatment by determining a first subset of covariate vectors for patients who received the medical treatment and a second subset of covariate vectors for patients who did not receive the medical treatment, and comparing (i) a first probability with (ii) a second probability for one or more response categories, wherein the first probability is the probability that a covariate vector in the first subset of covariate vectors is associated with the one or more response categories, and the second probability is the probability that a covariate vector in the second subset of covariate vectors is associated with the one or more response categories, wherein the first probability and the second probability are calculated by using a trained neural network model; and providing the estimated treatment effect for display on a graphical user interface of a computing device. Other implementations include corresponding systems, apparatus, and computer programs configured to perform the actions of the method encoded on a computer storage device.
[0004] In some implementations, the method includes: receiving a first data set from a database, the first data set including a set of covariate vectors and a set of response indicators, each covariate vector including clinical covariates of a corresponding patient, each response indicator being associated with a corresponding covariate vector, wherein the response indicators in the set of response indicators vary over a range; receiving a plurality of segmentations of the range of the response indicators from the database to define a plurality of response categories for the response indicators; training a neural network model (NNM) with the first data set to obtain a first NNM; bootstrapping the first data set n times to obtain n second data sets; training the NNM with each of the n second data sets to obtain n sets of second NNMs; for each covariate vector, obtaining n+1 predicted responses for the covariate vector from the first NNM and the n second NNMs, each predicted response being obtained by combining the first NNM and the n second NNMs The invention relates to a method for providing a method for providing a plurality of medical treatments for a patient who has undergone the medical treatment and a plurality of covariate vectors for a patient who has not ...
[0005] The present disclosure also provides one or more non-transitory computer-readable storage media coupled to one or more processors and having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform operations according to an implementation of the method provided herein.
[0006] The present disclosure further provides a system for implementing the method provided herein. The system includes one or more processors and a computer-readable storage medium coupled to the one or more processors, the computer-readable storage medium having instructions stored thereon, the instructions causing the one or more processors to perform operations according to the implementation of the method provided herein when executed by the one or more processors.
[0007] The method according to the present disclosure may include any combination of aspects and features described herein. That is, the method according to the present disclosure is not limited to the combination of aspects and features specifically described herein, but also includes any combination of aspects and features provided.
[0008] Among other advantages, the present implementation provides the following benefits. The method proposed herein can be used to predict the effectiveness of medical treatment for a specific patient. Accurate prediction of treatment effects can provide significant improvements to medical and health systems, including prescribing a medical treatment that is predicted to be more effective for the specific patient, and excluding a medical treatment that appears to be less effective from the prescription of the specific patient, which can lead to faster recovery of the patient, less side effects that the patient encounters due to taking less effective treatments, and reduced treatment costs (including currency, time, clinical facilities used, and health care providers).
[0009] The present implementation adopts a nonparametric approach by using a neural network model without making a linear assumption on the relationship between the model's predictor variables and the patient's response variables. This method is superior to a generalized linear model that models the relationship between the predictor variables and the response variables. Such a linear model may suffer from power loss and deviations from linear assumptions. For example, the response endpoint may not be linearly related to the covariate of interest, or the patient's characteristics (e.g., covariates) may not be completely balanced. The present disclosure provides two methods to overcome such limitations. In the first method, the implementation discretizes the continuous response variable into a categorical variable, and applies a neural network model to the classification without making a linear assumption, and provides an automatic feature representation capability embedded in a deep learning method. The second method is built on the first method, but does not discretize the continuous response variable. Therefore, the present implementation improves the estimation accuracy on the linear method by avoiding making a linear assumption in the relationship between the predictor variables and the response variables.
[0010] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other features and advantages of the disclosure will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 An example environment is depicted that can be used to perform implementations of the present disclosure.
[0012] Figure 2 Depicted is an example feed-forward neural network model that may be used in implementations of the present disclosure.
[0013] Figure 3 Depicted are example processes that may be performed in accordance with implementations of the present disclosure.
[0014] FIG. 4A to FIG. 4B Depicted are example processes that may be performed in accordance with implementations of the present disclosure.
[0015] Figure 5 is a schematic illustration of an example computer system that can be used to perform implementations of the present disclosure.
[0016] The same reference symbols in different drawings denote the same elements. DETAILED DESCRIPTION
[0017] Implementations of the present disclosure include computer-implemented methods for estimating the therapeutic effect of a medical treatment. These implementations also use deep learning methods to predict the probability that a patient with specific clinical covariates is a responder to the medical treatment.
[0018] Figure 1 An example environment 100 is depicted that can be used to perform implementations of the present disclosure. The environment 100 shows a user 116 who uses a computing device 102 to request an estimate of the effect of a treatment. The computing device 102 communicates with a database 106, for example, via a network 110. The database 106 stores clinical covariate data for a sample patient. The database 106 provides this data to the computing device 102. The computing device 102 uses this data to train a (deep) neural network model (NNM) and provide an estimate of the probability of the effectiveness of a medical treatment (e.g., overall or for a specific patient). Alternatively or additionally, the database 106 can provide data to a computing device 108 including one or more processors 104 to perform an estimation procedure.
[0019] NNM can utilize nonlinear activation functions (e.g., softmax activation function, rectified linear unit (ReLu) activation function) and loss functions (e.g., cross entropy loss, minimum square error (L 2 )).
[0020] The NNM used to calculate the probability may include one or more feed-forward NNMs. Figure 2 An example feed-forward NNM 200 that can be used for implementations of the present disclosure is shown. The NNM 200 includes multiple layers of computational units interconnected in a feed-forward manner. Each neuron in a layer has a directed connection to a neuron in a subsequent layer. There are many choices for activation functions (such as sigmoid, ReLU, etc.) that link neurons in adjacent layers. The parameters of the NNM can be calculated by a stochastic gradient descent algorithm. The NNM can be trained by searching for hyperparameters of the NNM (e.g., number of layers, number of neurons in each layer, etc.) based on the performance of the NNM on a validation set, for example, to avoid overfitting problems and / or to find the lowest validation error.
[0021] In some implementations, only one feedforward neural network is used. Such an implementation provides at least two advantages. First, the relationship between multiple outputs can be automatically handled through connections in the lower layers. Second, it is computationally simple and provides the ability to perform fast hyperparameter searches in a relatively small space. The NNM is represented here by f(X), where X represents the covariate vector.
[0022] Figure 3 and FIG. 4A to FIG. 4B Depicted are two example processes that can be performed according to implementations of the present disclosure to determine the effectiveness of a medical treatment. For ease of description, the processes described herein are categorized into two methods. Those skilled in the art will appreciate that a system can benefit from either or both of these two methods. First method:
[0023] The first method is Figure 3 and may be performed by one or more computing devices (e.g., Figure 1 The one or more computing devices receive a data set {(xi,yi), i=1,...,q} (302) from a database such as database 106, for example. The data set includes a set of covariate vectors X and a set of response indicators Y. Each covariate vector xi includes clinical covariates of the corresponding patient. Each response indicator (also referred to as "endpoint response" or "endpoint responder") yi is associated with a corresponding covariate vector xi.
[0024] The response indicator varies across the range. The range is divided into levels C 1 <C 2 <... <C k ∈supp(Y). Each C j In some implementations, the one or more computing devices receive a segmentation indicating a response category, such as from a database (304). In some implementations, a computing device receives the segmentation from an operator (eg, user 116).
[0025] To determine the efficacy of a medical treatment, the probability of a patient's response indicator, P(Y <C 1 )、P(Y <C 2 ),…,P(Y <C k). If a treatment can result in a high probability of one or more key response categories (e.g., two), the efficacy of a drug or treatment can be confirmed. For example, the percentage change (PCHG) of a key biomarker associated with asthma can be studied to determine the effectiveness of an asthma medication. Lower values of the biomarker indicate a healthier condition; therefore, a lower PCHG indicates that the medical treatment is more effective for asthma. In this example, PCHG is the response indicator. The range of change in the response indicator can be divided into response categories C using three splits of -50%, -25%, and 0. 1 <-50%, -50% <C 2 <-25%, ..., -25% <C 3 <0 and C4>0. In this example, PCHG falls into category C 1 or C 2 The higher the probability, the more effective the asthma medication is.
[0026] The one or more computing devices convert each response indicator received in 302 into a one-hot encoded vector based on the response category indicated by the segmentation (306). (A one-hot encoded vector is a vector in which only one bit is "on", e.g., has a value of 1 instead of 0.) In other words, the computing device generates a zi vector for each xi covariate vector (and its corresponding yi response indicator) by: z i =(I(y i <C 1 ),I(C 1 ≤y i <C 2 ),..,I(C k-1 ≤y i <C k ), I(y i ≥C k )) (1)
[0027] The one or more computing devices train the NNM (308) with a set of covariate vectors and their corresponding one-hot encoded vectors (i.e., {xi,zi}i=1,2,...,q) to obtain a trained NNM model
[0028] Applying the trained NNM to each covariate vector {xi}i=1,...,q, the trained NNM provides the probability of the response class associated with the covariate vector, namely:
[0029] The computing device estimates (310) the therapeutic effect of a medical treatment by applying the trained NNM to the data {xi,yi}i=1,2,...,q. To do so, the computing device determines (or receives from a database) a first subset of covariate vectors for patients receiving the medical treatment and a second subset of covariate vectors for patients not receiving the medical treatment. The computing device then compares (i) a first probability with (ii) a second probability for one or more response categories. The first probability is the probability that a covariate vector in the first subset of covariate vectors is associated with the one or more response categories, and the second probability is the probability that a covariate vector in the second subset of covariate vectors is associated with the one or more response categories. The computing device receives the first probability and the second probability from the trained NNM. The one or more response categories (for which the first probability is compared with the second probability) may be all response categories or specific response categories among the response categories (e.g., in the above example, response category C 1 <-50% and -50% <C 2 <-25%).
[0030] More specifically, consider C j As the key response category and T as the component of X representing treatment (eg, T=1 for treatment and T=0 for no treatment), the treatment effect was calculated by: in and
[0031] Such an aggregate estimator (see Eqs. (3) and (4)) improves the accuracy of estimating treatment effects with fairly large sample sizes. Since the covariates are correlated and there is no linearity assumption in (deep) NNMs, observations containing true inherent associations (e.g., components of a covariate vector, response indicators) can be applied to obtain accurate results even with complex relationships between covariates or between covariates and the output of the NNM.
[0032] The one or more computing devices provide the estimated treatment effect for presentation. For example, the computing device may provide the estimated treatment effect for presentation in a graphical user interface (e.g., Figure 1 The display is displayed on a graphical user interface of the computing device 102 (312).
[0033] As described above, the probability of a patient's response indicator is calculated for each response category, i.e., P(Y <C 1 )、P(Y <C 2 ),…,P(Y <Ck )。Using ordinal classification (or partitioning) for the response indicator range can impose hard constraints on the output of the NNM. To simplify the training of the NNM, partitioning can be performed as categorical classification rather than ordinal classification. This categorical classification can provide disjoint response categories, such as, (Y < C1), (C1 ≤ Y < C2), …, (Ck−1 ≤ Y < Ck), (Y ≥ Ck). The output of the NNM for these disjoint response categories will be in the following form: P(Y < C1), P(C1 ≤ Y < C2), …, P(Ck−1 ≤ Y < Ck), P(Y ≥ Ck), which results in k + 1 components in the output. Through the trained and a specific set of covariates, the computing device can obtain the estimated probability of the ordinal classification response category by summing up: Such a procedure offers the following advantages: (i) a one-to-one transformation with no information loss; and (ii) obtaining a standard classification problem after the transformation, which can be accurately and efficiently solved by a deep neural network.
[0034] The first method described above may include bootstrapping the data received at 302. Bootstrapping can be beneficial when there are not enough samples available for training the neural network or when more samples than the available samples are desired for training. Bootstrapping the received data includes bootstrapping (or resampling from) a set of covariate vectors and their corresponding response indicators. The original data and the bootstrapped data can be used for the training and / or validation of the NNM. Second method:
[0035] The second method is depicted in FIG. 4A to FIG. 4B and can be performed by one or more computing devices (e.g., Figure 1 the computing devices 102 or 108 in
[0036] ). The second method eliminates the conversion (or transformation) of the response indicator (yi) into a one-hot encoder vector (zi). Instead, the NNM is used to directly model the relationship between the covariate vector X and the response indicator Y. As an advantage, eliminating the transformation of the response indicator to the encoder vector can reduce power loss.
[0037] In the first level of bootstrapping, the second method bootstraps the received data n times to obtain n sets of bootstrap data. Each set of bootstrap data is used to train the NNM and obtain the corresponding trained NNM f B (X). The second method then collects the bi (X) For each covariate vector (Xi) the predicted response indicator {y^(b i ),b i =1,...,n}. The method estimates the efficacy of a medical treatment based on the probability that the predicted response indicator (for each covariate vector) is within one or more specific response categories. The following paragraphs provide a detailed description of the second method.
[0038] In a second method, the one or more computing devices (as described above) receive a first data set {(xi,yi), i=1,...,q} (402), for example, from a database such as database 106. The first data set includes a set of covariate vectors X and a set of response indicators Y. Each covariate vector xi includes clinical covariates for a corresponding patient. Each response indicator yi is associated with a corresponding covariate vector xi.
[0039] The response indicator varies across the range. The range is divided into levels C 1 <C 2 <... <C k ∈supp(Y). Each C j In some implementations, the one or more computing devices receive a segmentation indicating a response category, such as from a database (404). In some implementations, a computing device receives the segmentation from an operator (eg, user 116).
[0040] The one or more computing devices train a NNM with a first data set {(xi, yi), i=1, ..., q} to obtain a first NNM (406). The NNM may include a feed-forward neural network. The training may include training the NNM using a plurality of hyperparameters, and selecting a set of hyperparameters (from the plurality of hyperparameters) with a lowest validation error for the first NNM.
[0041] The computing device bootstraps the first data set n times to obtain n sets of second data sets {(xi,yi)(b p );i=1,...,q;p=1,...,n}(408). The computing device trains a NNM with each of the second data sets to obtain n sets of second NNMs(410). Training one or more second NNMs may be performed in a process similar to the process of training the first NNM described in the previous paragraph.
[0042] The one or more computing devices obtain n+1 predicted responses for each covariate vector of the first data set from the trained NNMs (i.e., from the first NNM and the n second NNMs) (412). Each predicted response is obtained by applying the first NNM and a corresponding one of the n second NNMs to the covariate vector. The obtained covariate vector x i The predicted response can be expressed as {y^ i (b p ),bp=1,...,n+1}.
[0043] The probability that the covariate vector is a responder to the medical treatment can indicate the likelihood that a patient with clinical covariates similar to the covariate vector will respond or has responded to the medical treatment. The computing device estimates the covariate vector x by estimating the associated probability of the covariate vector for one or more key response categories. i is the probability of being a responder to the medical treatment. For example, the computing device may estimate the covariate vector x for each response category identified at 404 i The associated probability (414).
[0044] The covariate vector x can be calculated by applying the indicator function to each of the n+1 predicted responses to obtain n+1 outputs of response categories, and by normalizing the aggregation of the obtained n+1 outputs. i In other words, all answer categories C j The covariate vector x in i The association probability can be calculated by the following formula:
[0045] The one or more computing devices estimate the therapeutic effect of the medical treatment based on the association probability of the covariate vector associated with the patients who have received the medical treatment and the association probability of the covariate vector associated with the patients who have not received the medical treatment. More specifically, the computing device determines (or receives from a database) a first subset of covariate vectors for patients who have received the medical treatment and a second subset of covariate vectors for patients who have not received the medical treatment (416). The computing device then compares (i) a first normalized aggregate of the association probabilities of the first subset with (ii) a second normalized aggregate of the association probabilities of the second subset for one or more key response categories (418). In other words, the computing device estimates the therapeutic effect by the following formula: where t i =1 indicates treatment, and t i=0 indicates no treatment. The one or more key response categories may include all response categories, or specific response categories.
[0046] The one or more computing devices provide the estimated treatment effect for presentation. For example, the computing device may provide the estimated treatment effect for presentation in a graphical user interface (e.g., Figure 1 The display (420) is displayed on a graphical user interface of the computing device 102 in FIG. 1 .
[0047] In addition to the above procedures, the one or more computing devices can estimate the uncertainty of the estimated treatment effect by using a second-level bootstrapping. In the second level, the uncertainty is estimated by bootstrapping the first data set m times to obtain m third data sets, calculating the treatment effect for each of the m third data sets to obtain m treatment effects, and calculating the uncertainty based on the distribution of the m treatment effects. Example uncertainties include, but are not limited to, confidence intervals and standard deviations for the m treatment effects.
[0048] Figure 5 A schematic diagram of an example computing system 500 is depicted. The system 500 can be used to perform operations described with respect to one or more implementations of any of the first method or the second method according to the present disclosure. For example, the system 500 can be included in any or all of the server components or other one or more computing devices discussed herein. The system 500 can include one or more processors 510, one or more memories 520, one or more storage devices 530, and one or more input / output (I / O) devices 540. The components 510, 520, 530, and 540 can be interconnected using a system bus 550.
[0049] The processor 510 may be configured to execute instructions within the system 500. The processor 510 may include a single-threaded processor or a multi-threaded processor. The processor 510 may be configured to execute or otherwise process instructions stored in one or both of the memory 520 or the storage device 530. The execution of one or more instructions may cause the graphical information to be displayed or otherwise presented via a user interface on the I / O device 540.
[0050] The memory 520 can store information within the system 500. In some implementations, the memory 520 is a computer-readable medium. In some implementations, the memory 520 can include one or more volatile memory units. In some implementations, the memory 520 can include one or more non-volatile memory units.
[0051] The storage device 530 may be configured to provide mass storage for the system 500. In some implementations, the storage device 530 is a computer-readable medium. The storage device 530 may include a floppy disk device, a hard disk device, an optical disk device, a magnetic tape device, or other types of storage devices. The I / O device 540 may provide I / O operations for the system 500. In some implementations, the I / O device 540 may include a keyboard, a pointing device, or other devices for data input. In some implementations, the I / O device 540 may include an output device, such as a display unit for displaying a graphical user interface or other types of user interfaces.
[0052] The described features may be implemented in digital electronic circuit systems, or in computer hardware, firmware, software, or a combination thereof. The apparatus may be implemented in a computer program product tangibly embodied in an information carrier (e.g., in a machine-readable storage device) for execution by a programmable processor, and the method steps may be performed by a programmable processor executing an instruction program to perform the functions of the described implementation by operating on input data and generating output. The described features may advantageously be implemented in one or more computer programs that may be executed on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that may be used directly or indirectly in a computer to perform an activity or cause a result. A computer program may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0053] Suitable processors for executing a program of instructions include, by way of example, general and special purpose microprocessors, and the sole processor or one of multiple processors of any type of computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer include a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer may also include one or more mass storage devices for storing data files, or be operatively coupled to communicate therewith; such devices include magnetic disks (such as internal hard disks and removable disks), magneto-optical disks, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks and removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into an application specific integrated circuit (ASIC).
[0054] To provide for interaction with a user, these features may be implemented on a computer having a display device (such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user, and a keyboard and pointing device (such as a mouse or trackball) through which the user can provide input to the computer.
[0055] These features can be implemented in a computer system including a back-end component (such as a data server), or a computer system including a middleware component (such as an application server or an Internet server), or a computer system including a front-end component (such as a client computer with a graphical user interface or an Internet browser), or any combination thereof. The components of the system can be connected by any digital data communication form or medium (such as a communication network). Examples of communication networks include, for example, a local area network (LAN), a wide area network (WAN), and the computers and networks that form the Internet.
[0056] A computer system may include clients and servers. The client and server are generally remote from each other and typically interact through a network (such as the described network). The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0057] In addition, the logic flows depicted in the figures do not require the particular order or sequential order shown to achieve the desired results. In addition, other steps can be provided, or steps can be eliminated from the described processes, and other components can be added to or removed from the described systems. Therefore, other implementations are within the scope of the appended claims.
[0058] Many implementations of the present disclosure have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the appended claims. In summary, the present invention includes but is not limited to the following items: 1. A computer-implemented method executed by one or more processors, the method comprising: A data set is received from a database, the data set comprising: a set of covariate vectors, each of which includes the clinical covariates of the corresponding patient, and a set of response indicators, each response indicator being associated with a corresponding covariate vector, wherein the response indicators in the set of response indicators vary in range; receiving from the database a plurality of partitions over the range of the response indicator to define a plurality of response categories for the response indicator; converting each answer indicator into a corresponding one-hot encoded vector based on the answer category indicated by the plurality of segmentations; training a neural network model based on each covariate vector and a corresponding one-hot encoded vector associated with the covariate vector; and The treatment effect of medical treatment is estimated by: determining a first subset of covariate vectors for patients who received the medical treatment and a second subset of covariate vectors for patients who did not receive the medical treatment, and for one or more response categories, comparing (i) a first probability that a covariate vector in a first subset of the covariate vectors is associated with the one or more response categories with (ii) a second probability that a covariate vector in a second subset of the covariate vectors is associated with the one or more response categories, wherein the first probability and the second probability are calculated by using a trained neural network model; and The estimated treatment effect is provided for display on a graphical user interface of the computing device. 2. A method according to item 1, wherein the neural network model utilizes non-linear activation functions and loss functions. 3. A method according to claim 1, wherein when estimating the treatment effect, the one or more response categories include all of the multiple response categories. 4. A method according to claim 1, wherein training the neural network model further comprises bootstrapping the set of covariate vectors and the set of response indicators. 5. A method according to clause 1, wherein at least two of the plurality of response categories are disjoint. 6. A method according to item 1, wherein the neural network model utilizes a softmax activation function and a cross entropy loss function. 7. A method according to item 1, wherein the neural network is a feed-forward neural network. 8. A computer-implemented method executed by one or more processors, the method comprising: A first data set is received from a database, the first data set comprising: a set of covariate vectors, each of which includes the clinical covariates of the corresponding patient, and a set of response indicators, each response indicator being associated with a corresponding covariate vector, wherein the response indicators in the set of response indicators vary in range; receiving from the database a plurality of partitions over the range of the response indicator to define a plurality of response categories for the response indicator; training a neural network model (NNM) using the first data set to obtain a first NNM; bootstrapping the first data set n times to obtain n second data sets; Training the NNM with each of the n second data sets to obtain n sets of second NNMs; For each covariate vector: obtaining n+1 predicted responses for the covariate vector from the first NNM and the n second NNMs, each predicted response being obtained by applying a corresponding one of the first NNM and the n second NNMs to the covariate vector, and For each response category: The association probability indicating that the covariate vector is associated with the response category is estimated by: For the response category, applying an indicator function to each of the n+1 predicted responses to obtain n+1 outputs for the response category, and calculating the association probability by normalizing an aggregation of the obtained n+1 outputs for the response category; The treatment effect of medical treatment is estimated by: determining a first subset of covariate vectors for patients who received the medical treatment and a second subset of covariate vectors for patients who did not receive the medical treatment, and For one or more response categories, comparing (i) a first normalized aggregate of the associated probabilities of the first subset with (ii) a second normalized aggregate of the associated probabilities of the second subset; and The estimated treatment effect is provided for display on a graphical user interface of the computing device. 9. A method according to item 8, wherein the one or more response categories include all of the multiple response categories. 10. The method of claim 8, further comprising estimating the uncertainty of the treatment effect by: bootstrapping the first data set m times to obtain m third data sets; Calculate the treatment effect for each of the m third data sets to obtain m treatment effects; The uncertainty is calculated based on the distribution of the m treatment effects. 11. A method according to item 10, wherein the uncertainty of the treatment effect includes at least one of the confidence interval and standard deviation of the m treatment effects. 12. The method of clause 8, wherein the NNM utilizes a non-linear activation function and a minimum variance loss function. 13. A method according to item 12, wherein the non-linear activation function is a rectified linear unit (ReLU) function. 14. The method according to clause 8, wherein said training said NNM with said first data set to obtain said first NNM comprises: training the NNM using a plurality of hyperparameters; and A hyperparameter set having a lowest validation error is selected for the first NNM, the hyperparameter set being selected from the plurality of hyperparameters. 15. A non-transitory computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising: A data set is received, the data set comprising: a set of covariate vectors, each of which includes the clinical covariates of the corresponding patient, and a set of response indicators, each response indicator being associated with a corresponding covariate vector, wherein the response indicators in the set of response indicators vary in range; receiving a plurality of partitions on the range of the response indicator to define a plurality of response categories for the response indicator; converting each answer indicator into a corresponding one-hot encoded vector based on the answer category indicated by the plurality of segmentations; training a neural network model based on each covariate vector and a corresponding one-hot encoded vector associated with the covariate vector, the neural network model utilizing a non-linear activation function and a loss function; and The treatment effect of medical treatment is estimated by: determining a first subset of covariate vectors for patients who received the medical treatment and a second subset of covariate vectors for patients who did not receive the medical treatment, and for one or more response categories, comparing (i) a first probability that a covariate vector in a first subset of the covariate vectors is associated with the one or more response categories with (ii) a second probability that a covariate vector in a second subset of the covariate vectors is associated with the one or more response categories, wherein the first probability and the second probability are calculated by using a trained neural network model; and The estimated treatment effect is provided for display on a graphical user interface of the computing device. 16. A non-transitory computer-readable medium according to item 15, wherein when estimating the treatment effect, the one or more response categories include all of the multiple response categories. 17. A non-transitory computer-readable medium according to item 15, wherein training the neural network model further comprises bootstrapping the set of covariate vectors and the set of response indicators. 18. The non-transitory computer-readable medium of clause 15, wherein at least two of the plurality of response categories are disjoint. 19. A non-transitory computer-readable medium according to item 15, wherein the neural network model utilizes a softmax activation function and a cross entropy loss function. 20. The non-transitory computer-readable medium of clause 15, wherein the neural network is a feed-forward neural network.
Claims
1. A computer-implemented method executed by one or more processors, the method include: A data set is received from a database, the data set comprising: a set of covariate vectors, each covariate vector including clinical covariates as clinical characteristics of a corresponding patient in a group of patients, and a set of response indicators for the set of patients, each response indicator being associated with a respective clinical covariate in the respective covariate vector and being within a particular range assigned to the respective clinical covariate, wherein each range includes one or more response categories for the respective clinical covariate, each response category being part of the respective range; as well as The treatment effect of medical treatment is estimated by: determining a first subset and a second subset of covariate vectors from the set of covariate vectors, the first subset comprising clinical covariates for patients who received the medical treatment and the second subset comprising clinical covariates for patients who did not receive the medical treatment, and The first normalized probability is obtained by: applying the trained neural network model to a first subset of the covariate vectors to obtain a first set of probabilities, and Calculating a normalized aggregation of the first set of probabilities to obtain the first normalized probability; The second normalized probability is obtained by: applying the trained neural network model to a second subset of the covariate vectors to obtain a second set of probabilities, and Calculating a normalized aggregation of the second set of probabilities to obtain the second normalized probabilities; For one or more response categories, comparing the first normalized probability to the second normalized probability; and The estimated treatment effect is provided for display on a computing device.
2. The method of claim 1, wherein the neural network model utilizes nonlinear activation functions and loss functions.
3. The method of claim 1 or 2, wherein in estimating the treatment effect, the one or more response categories include all response categories.
4. A method according to claim 1 or 2, wherein the estimated treatment effect is provided for display on a graphical user interface of a computing device.
5. The method of claim 1 or 2, wherein at least two of the plurality of response categories are disjoint.
6. The method of claim 1 or 2, wherein the neural network model utilizes a softmax activation function and a cross entropy loss function.
7. The method according to claim 1 or 2, wherein the neural network model is a feedforward neural network model.
8. A computer-implemented method executed by one or more processors, the method include: A first data set is received from a database, the first data set comprising: a set of covariate vectors, each covariate vector including clinical covariates as clinical characteristics of a corresponding patient in a group of patients, and a set of response indicators for the set of patients, each response indicator being associated with a respective clinical covariate in the respective covariate vector and being within a particular range associated with the respective clinical covariate, wherein each range includes a plurality of response categories for the respective clinical covariate, each response category being a portion of the respective range; For each covariate vector: obtaining n+1 predicted responses for the covariate vector, each predicted response being obtained by applying a corresponding one of n+1 neural network models (NNMs) to the covariate vector, and For each response category, estimating an association probability indicating that the covariate vector is associated with the response category by, wherein the association probability is estimated for the response category based on n+1 predicted responses, The treatment effect of medical treatment is estimated by: determining a first subset of covariate vectors and a second subset of covariate vectors from the set of covariate vectors, the first subset comprising clinical covariates for patients who received the medical treatment, the second subset of covariate vectors comprising clinical covariates for patients who did not receive the medical treatment, and For one or more response categories, comparing (i) a first association probability estimated for the first subset with (ii) a second association probability estimated for the second subset; and providing the estimated treatment effect for display on a computing device.
9. The method of claim 8, wherein the one or more response categories include all of the response categories.
10. The method of claim 8 or 9, further comprising estimating the uncertainty of the treatment effect by: bootstrapping the first data set m times to obtain m third data sets; Calculate the treatment effect for each of the m third data sets to obtain m treatment effects; The uncertainty is calculated based on the distribution of the m treatment effects. The method of claim 10 , wherein the uncertainty of the treatment effect comprises at least one of a confidence interval and a standard deviation of the m treatment effects.
12. The method according to any one of claims 8 to 9, wherein the NNM utilizes a non-linear activation function and a minimum variance loss function.
13. The method of claim 12, wherein the non-linear activation function is a rectified linear unit (ReLU) function.
14. The method of any one of claims 8 to 9, wherein the estimated treatment effect is provided for display on a graphical user interface of a computing device.
15. A non-transitory computer readable medium coupled to one or more processors and storing instructions executable by a computer system to perform the method of any one of claims 1 to 14.