Systems and methods for predicting and optimizing clinical trial outcomes

The AI-based trial outcome predictor with an explainability model optimizes clinical trial configurations by generating accurate, interpretable predictions, addressing the limitations of existing methods in predicting clinical trial success and resource efficiency.

JP2025539982APending Publication Date: 2025-12-11VALO HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522159
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-17
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing methods for predicting clinical trial outcomes lack accuracy and mechanistic insight, are inflexible, and do not allow for 'what if' scenario construction, making it difficult to optimize clinical trial designs efficiently.

Method used

A system and method using an AI-based trial outcome predictor trained on historical data to generate explainable predictions and optimize clinical trial configurations by identifying key contributors to outcomes, incorporating an explainability model to provide insight into the model's output and allowing for data normalization and PII removal.

Benefits of technology

Enhances prediction accuracy and optimization of clinical trials by providing clear insights into success factors, reducing resource waste, and improving the likelihood of successful trial outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539982000001_ABST
    Figure 2025539982000001_ABST
Patent Text Reader

Abstract

A vector associated with a clinical trial is obtained from one or more data sources. The vector includes one or more fixed elements and one or more optimizable elements. A trial outcome predictor trained on data related to multiple past clinical trials is used to determine a probabilistic model of the outcome of the clinical trial based on the vector. An explainability model is used to calculate a contribution score for the vector based on the probabilistic model. Each contribution score indicates the relative contribution of an associated element of the vector to the outcome of the clinical trial. An explainable prediction of the trial outcome of the clinical trial is generated based on the probabilistic model and the one or more contribution scores associated with the one or more optimizable elements. The explainable prediction is output for review by a user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 429,775, filed December 2, 2022, which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates to automatically modeling the outcomes of future clinical trials. In particular, but not by way of limitation, the present disclosure relates to predicting the probability that a future clinical trial will achieve a trial outcome. In particular, but not by way of limitation, the present disclosure relates to optimizing the configuration of a future clinical trial to improve the probability that a trial outcome will be achieved. [Background technology]

[0003] Approximately 1 in 10 drug candidates successfully progress through clinical trials and gain regulatory approval. Accurately predicting clinical trial outcomes therefore offers multiple opportunities in clinical development, including prioritizing drug development investments, modifying existing trial portfolios to maximize success, and discovering underappreciated molecules. Summary of the Invention

[0004] According to one aspect of the present disclosure, systems and methods, and computing instructions configured for execution on one or more processors, are provided for generating explainable predictions of trial outcomes of a clinical trial. A trial configuration vector associated with the clinical trial is obtained from one or more data sources. The trial configuration vector includes one or more fixed elements and one or more optimizable elements. A probabilistic model of the outcome of the clinical trial is determined using a trial outcome predictor based on the trial configuration vector. The trial outcome predictor is trained on data related to multiple past clinical trials. An explainability model based on the probabilistic model is used to calculate multiple contribution scores for the trial configuration vector. Each of the multiple contribution scores indicates the relative contribution of an associated element of the trial configuration vector to the outcome of the clinical trial. An explainable prediction of the trial outcome of the clinical trial is generated based on the probabilistic model and one or more of the multiple contribution scores. The one or more contribution scores are associated with one or more optimizable elements. The explainable prediction is output for review by a user.

[0005] According to further aspects of the present disclosure, systems and methods for optimizing parameters of a clinical trial, as well as computing instructions configured to execute on one or more processors, are provided. A first test configuration associated with the clinical trial is obtained from one or more data sources. The first test configuration includes values ​​associated with one or more fixed test parameters and at least one optimizable test parameter. An outcome predictor is obtained, and the outcome predictor estimates a relationship between the test configuration of the clinical trial and an outcome of the clinical trial. The first test configuration is optimized to improve the outcome of the clinical trial by calculating updated values ​​of the at least one optimizable test parameter using the outcome predictor and the first test configuration, such that the first estimated outcome of the clinical trial is greater than a second estimated outcome of the clinical trial, and creating an updated test configuration including the updated value of the at least one optimizable test parameter. The first estimated outcome is determined from the outcome predictor based on the updated test configuration, and the second estimated outcome is determined from the outcome predictor based on the first test configuration. The updated test configuration is output for user review.

[0006]

[0010] In accordance with the above and the disclosure herein, the present disclosure includes, at a minimum, improvements in computer functionality or other technology that disclose an artificial intelligence (AI)-based model, e.g., a trial outcome predictor, trained using data from multiple past clinical trials, which, when deployed on an underlying system, allows the disclosed systems and methods to run with fewer iterations and use fewer computing resources than related systems and methods of the prior art. That is, the present disclosure describes improvements in the functionality of the computer itself, or "any other technology or technical field," because the increased prediction improvement provided by the trial outcome predictor allows the underlying computer system to utilize fewer processing and memory resources compared to prior art systems and methods, because the trial outcome predictor can generate or determine a probability model, or other outcome, of a clinical trial with a higher likelihood of success using fewer computational cycles or other iterations, with less impact on the underlying computing device, compared to conventional prior art systems and methods. Stated another way, the systems and methods of the present disclosure are an improvement over the prior art because at least prior art systems and methods require empirical or trial-and-error approaches that may involve real-world testing and / or data inputs that may lead to and require large databases and memory and processor usage to arrive at analogous real-world or simulated test results with the same or similar high accuracy or predictive results. In contrast, the disclosed systems and methods describe the generation and / or use of test configuration vectors that define a streamlined set of elements (e.g., fixed elements and optimizable elements) using a more limited set of known data related to the elements, requiring less memory and / or processing utilization compared to conventional approaches in which large sets of unknown and potentially irrelevant data are used or required.

[0007] Additionally, the present disclosure relates to improvements in other technical or technological fields, at least because it discloses the generation and / or use of an explainable model. The explainable model improves upon conventional prior art AI-related models by providing technical clarity in the form of a visual or data view of the disclosed AI model's output or results (e.g., a test outcome predictor). Stated differently, conventional AI models and related algorithms generally do not provide clarity, insight, or other explanations regarding the model's generation as to how the output or results are achieved. Such prior art methods operate as black-box computational structures, providing little or no insight into the model or its training methods. Such prior art techniques may be disadvantageous in training or generating AI models because technical biases or errors may be implicitly built into such AI models, resulting in technical biases or errors in the model's output that cannot be discovered, improved upon, or otherwise determined. In contrast, the disclosed explainability model provides a view into a model (e.g., a test outcome predictor) and its associated output by providing a database and / or visual representation, or otherwise explanation, of how training data affects or otherwise determines the output of the disclosed AI model, e.g., a test outcome predictor. In other words, the explainability model visualizes how the AI ​​model is currently trained and how such training affects the output results. This allows the AI ​​model to be retrained or reconfigured using different training data, e.g., different test configuration vectors with different fixed and / or optimizable parameters, and / or using different data from additional data sources, to eliminate errors and / or biases in a second version of the AI ​​model, e.g., a test outcome predictor, and its associated output.

[0008] Additionally, the present disclosure relates to improvements over other technologies or technical fields because at least the disclosed systems and methods provide for normalization and / or data formatting of data received, ingested, and / or otherwise obtained from one or more data sources to create, generate, or otherwise obtain test configuration vectors used to train AI models, e.g., test outcome predictors as described herein. In particular, data received from various data sources may include data from different databases, data sinks, or other data locations, and such data may not be compatible (e.g., as raw data or in other as-received form). The systems and methods of the present disclosure may operate to normalize or format such data, e.g., to create a set of normalized data for use in training a test outcome predictor as described herein.

[0009] Additionally, the present disclosure relates to improvements in other technical fields, at least because the disclosed systems and methods can reduce data sets and improve security by removing personally identifiable information (PII) from data received by a data source that includes PII. PII may include sensitive data, such as personal health data. Such data reduction and / or normalization can improve the security of the systems or methods described herein by eliminating data stored in memory, while also reducing the risk of security breaches of sensitive data.

[0010] The present disclosure includes certain features other than what are well understood, routine, and conventional activities in the art and / or adds unconventional steps that otherwise limit the present disclosure to particular useful applications, e.g., systems and methods for generating interpretable predictions of clinical trial test outcomes and / or optimizing clinical trial parameters.

[0011] Further features and aspects of the present disclosure are set forth in the accompanying claims.

[0012] The present disclosure will now be described, by way of example only, with reference to the accompanying drawings in which: [Brief explanation of the drawings]

[0013] [Figure 1] 1 illustrates a system for generating explainable predictions of test outcomes for clinical trials according to an embodiment of the present disclosure. [Figure 2] 1 illustrates an example test configuration according to an embodiment of the present disclosure. [Figure 3A] 10 illustrates an example of a contribution score according to an embodiment of the present disclosure. [Figure 3B] 10 illustrates an example of a contribution score according to an embodiment of the present disclosure. [Figure 4] 10 illustrates a portion of an example report according to an embodiment of the present disclosure. [Figure 5] 1 illustrates a process for optimizing a test configuration according to an embodiment of the present disclosure. [Figure 6A] 1 shows the results of predicting the outcomes of multiple clinical trials according to embodiments of the present disclosure. [Figure 6B] 1 shows the results of predicting the outcomes of multiple clinical trials according to embodiments of the present disclosure. [Figure 7A] 1 shows predicted success rates for two clinical trials according to embodiments of the present disclosure. [Figure 7B] 1 shows predicted success rates for two clinical trials according to embodiments of the present disclosure. [Figure 8A] 1 illustrates a method for generating explainable predictions of test outcomes in clinical trials according to an embodiment of the present disclosure. [Figure 8B] 1 illustrates a method for generating explainable predictions of test outcomes in clinical trials according to an embodiment of the present disclosure. [Figure 9] 1 illustrates a method for optimizing parameters of a clinical trial according to an embodiment of the present disclosure. [Figure 10] 1 illustrates an exemplary computing system for performing the methods of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] The ability to predict likely clinical trial outcomes is a critical step in prioritizing drug development investments, modifying existing trial portfolios to maximize success, and discovering underappreciated molecules. Clinical trials can be proposed trials, i.e., trials that have not yet begun and may be in the design and development phase. Typically, the intervention has been identified, but other factors, such as the sponsor and the sponsor's trial site, have yet to be determined. Predicting the likely outcomes of such trials at this stage can help avoid conducting trials that are unlikely to be successful. This can help avoid unnecessary patient and public involvement in trials that are likely to have limited public benefit. Alternatively, a clinical trial may have already begun but could still be improved. Existing methods for predicting clinical trial outcomes generally cannot predict the likelihood of trial success with high accuracy and mechanistic insight. Such methods lack flexibility and involve only limited dimensionality of data. Furthermore, the basic algorithms utilized in these existing methods do not allow for the construction of various "what if" scenarios for a clinical trial. The level of insight gained from such methods is always limited, as users are unable to clearly understand the factors contributing to the likelihood of a clinical trial's success or failure. Therefore, automatically optimizing clinical trial designs using existing methods is a difficult and inefficient task.

[0015] The systems and methods of the present disclosure aid in identifying and quantifying key sources of risk for assets and clinical trials by providing explainable estimates of the likelihood of a clinical trial's likely outcomes. Additionally, the present disclosure provides systems and methods for optimizing the configuration of a clinical trial. Such optimization improves the likelihood of clinical trial success by identifying aspects of the clinical trial design that can be modified prior to conducting the clinical trial, thereby aiding in more efficient use of resources.

[0016] FIG. 1 illustrates a system 100 for generating explainable predictions of test outcomes for clinical trials according to an embodiment of the present disclosure.

[0017] The system 100 includes a trial outcome predictor 102, an explainability model 104, and an optimizer 106. In an embodiment, the system 100 further includes a training unit 108 that trains the trial outcome predictor using historical clinical trial data 110. FIG. 1 also shows one or more data sources 112 in communication with the trial outcome predictor 102, and a report 114 viewable by a user 116. FIG. 1 also shows a trial configuration vector 118 associated with the clinical trial, a probabilistic model 120 of the outcome of the clinical trial, a plurality of contribution scores 122, an explainable prediction 124 of the trial outcome, a first updated trial configuration vector 126, and a second updated trial configuration vector 128.

[0018] A clinical trial is described or represented by a trial configuration vector 118. The clinical trial may be a proposed clinical trial that has not yet begun or a clinical trial that is already underway. The system 100 predicts the trial outcome of a clinical trial based on feature values ​​or elements in the trial configuration vector 118. The trial configuration vector 118 is obtained from one or more data sources 112 and passed to the trial outcome predictor 102. The trial configuration vector 118 includes one or more fixed elements associated with one or more fixed trial features and one or more optimizable elements associated with one or more optimizable trial features. The trial outcome predictor 102, trained based on data related to multiple past clinical trials, determines a probabilistic model 120 of the clinical trial outcome based on the trial configuration vector 118. In one embodiment, the trial outcome predictor 102 is trained by the training unit 108 using the past clinical trial data 110. The explainability model 104 calculates multiple contribution scores 122 for the trial configuration vector 118 based on the probabilistic model 120. Each contribution score in the plurality of contribution scores 122 indicates the relative contribution of the associated element of the trial configuration vector 118 to the outcome of the clinical trial. An explainable prediction 124 of the trial outcome of the clinical trial is generated based on the probabilistic model 120 and one or more contribution scores in the plurality of contribution scores 122. The one or more contribution scores from which the explainable prediction 124 is generated are associated with one or more optimizable elements of the trial configuration vector 118. The explainable prediction 124 is output for review by the user 116. In one embodiment, the explainable prediction 124 is included in the report 114 that is output for review by the user 116.

[0019] The explainable predictions 124 provide insight into the factors that contribute to the outcome (e.g., success or failure) of a clinical trial. The insights provided by the explainable predictions 124 can help facilitate the creation of improved clinical trial configurations (plans), thereby increasing the likelihood of achieving the trial results. Thus, the system 100 enables the user 116 to explore various "what if" scenarios for a clinical trial in an efficient manner. The output of the system 100 can also help quantify various success and risk factors for a clinical trial, thereby helping to influence the decision-making process when deciding whether to conduct a clinical trial. Thus, the greater insight provided by the present disclosure can aid in the development of improved clinical trials with increased likelihood of success and reduced risk. Additionally, the present disclosure provides a higher level of explainability and understanding of the relationship between clinical trial configurations and predicted trial outcomes, which can in turn help improve the understanding and optimization of clinical trials.

[0020] The trial configuration vector 118, alternatively referred to as a trial configuration or trial vector, corresponds to the configuration of a clinical trial. That is, the trial configuration vector 118 encodes aspects related to the clinical trial's design and protocol. Each element or value of the trial configuration vector 118 is associated with a feature of the clinical trial and may be binary, integer, or real-valued. For example, an element associated with a categorical sponsor type feature may be coded using techniques such as one-hot encoding to represent different possible values ​​of this feature (e.g., "corporate" or "academic institution"), while an element associated with a feature corresponding to the number of investigators involved in the clinical trial may take on a non-zero integer or real-valued value. In embodiments, elements associated with non-binary numeric features may be transformed (e.g., normalized, log-transformed, etc.) before being passed to the trial outcome predictor 102.

[0021] FIG. 2 illustrates an example test configuration according to an embodiment of the present disclosure.

[0022] 2 illustrates a trial configuration vector 202 associated with a plurality of features 204 of a clinical trial. Trial configuration vector 202 includes elements 206-1, 206-2, 206-3, 208-1, 208-2, 210, and 212. A first group of elements 214 includes elements 206-1, 206-2, and 206-3 associated with feature "A" in the plurality of features 204. A second group of elements 216 includes element 210 and is associated with feature "C" in the plurality of features 204. First group of elements 214 are obtained from a first data source 218, and second group of elements 216 are obtained from a second data source 220. The remaining elements of trial configuration vector 202 may be obtained from first data source 218, second data source 220, or some other data source (not shown). Data received from various data sources may be obtained, received, or otherwise contained in different databases or data sinks and may not be compatible as raw data or in other received formats. The systems and methods of the present disclosure may operate to normalize or format such data, e.g., to create normalized data sets for use in populating elements of test configuration vectors and / or for use in training test outcome predictors. Additionally, for example, certain aspects may reduce data and improve system security by removing personally identifiable information (PII) from data received by data sources that contain PII. PII may include sensitive data, such as personal health data. Such data reduction and / or normalization may improve the security of the systems or methods described herein by eliminating data stored in memory, while also reducing the risk of security breaches of sensitive data.

[0023] In the example shown in FIG. 2, feature "A" corresponds to an encoded feature (e.g., one-hot encoded) used to represent categorical feature values. For example, feature "A" corresponds to a drug target type and can take one of three values: "enzyme," "receptor," or "ion channel." Elements 206-1, 206-2, and 206-3 in first group of elements 214 may be binary and indicate which of the three drug target types the clinical trial pertains to. For example, element [1,0,0] may indicate an enzyme drug target type, while element [0,0,1] may indicate an ion channel drug target type.

[0024] 2 corresponds to a numeric feature. For example, feature "C" may correspond to the number of successful trials a clinical trial sponsor has. Thus, the element 210 in the trial configuration vector 202 associated with feature "C" may take on a non-zero integer value.

[0025] As shown in FIG. 2 , a trial configuration vector represents multiple features related to a clinical trial. Thus, a trial configuration vector, such as trial configuration vector 202, includes a concatenation of elements associated with clinical trial features. The focus of elements used by a trial configuration vector allows for a streamlined set of elements (e.g., fixed and optimizable elements) that defines a more limited, known dataset related to the elements, resulting in less memory and / or processing utilization compared to traditional approaches that use or require large amounts of unknown, potentially unrelated datasets. As a result, the multiple features associated with a trial configuration vector (e.g., multiple features 204) may include a combination of features from different categories or groups of features, such as biological features, chemical features, geographic features, sponsor features, design and operational features, which may include investigator features, keyword features, and other features. It will be understood that each of these feature categories is associated with a corresponding element in a separate vector within trial configuration vector 202. Incorporating data from various sources in this manner allows for a wider variety of clinical trials to be represented. This increased representational power contributes to improved prediction accuracy.

[0026] The plurality of features associated with the trial configuration vector may include biological features related to the target associated with the clinical trial. The biological features may include mechanism-of-action features. The mechanism-of-action features attempt to quantify various mechanisms of action of the drugs targeted by the clinical trial (e.g., angiogenesis inhibitors, immunostimulants, tubulin inhibitors, etc.). The mechanism-of-action features may be represented as n-dimensional vectors, where n corresponds to the number of different mechanisms of action that can be represented in the system. As such, elements of the trial configuration vector associated with mechanism-of-action features may be coded using categorical encoding techniques such as one-hot encoding, dummy encoding, effect encoding, and hash encoding. In embodiments, the biological features may include hierarchical mechanism-of-action features. Such features extend the mechanism-of-action coding described above to include higher-level groupings. That is, rather than coding only the master name of the mechanism of action of the drug in the clinical trial, names from each hierarchical level of the mechanism of action are coded. For example, consider the antidopaminergic mechanism of action of a dopamine D2 receptor antagonist. Using the master name encoding technique (as described above), a dopamine D2 receptor antagonist can be represented by a single binary index value in the test configuration vector. Representing this mechanism of action hierarchically, the master name ("dopamine D2 receptor antagonist") would be split into "dopamine" at the first level, "D2 receptor" at the second level, and "antagonist" at the third level. Each level could then be represented by an index value in the test configuration vector. In such a hierarchical representation, a dopamine D3 receptor antagonist would be represented identically at the first and second levels ("dopamine" and "D3 receptor") but differently at the third level ("antagonist"). Thus, a hierarchical representation allows for more precise quantification of the similarities between different mechanisms of action. This facilitates identifying the contributions of different hierarchies of mechanisms of action, and thus, representing mechanisms of action at a finer granularity ultimately improves the interpretability of the model.

[0027] The chemical features relate to targets associated with clinical trials. The chemical features may include chemical structure data related to the target drug. For example, the chemical structure data may include a vectorized representation of the target's SMILES string. The chemical structure data may also include vectors representing molecular data related to the target, such as molecular weight and categorically coded molecule type (e.g., using one-hot encoding, dummy encoding, effect encoding, hash encoding, etc.). The chemical structure data may include an indicator variable indicating the presence or absence of chemical structure data for the target.

[0028] Design and operational characteristics are associated with aspects of a clinical trial, such as its design, protocol, and operation. Design and operational characteristics may include geographic characteristics, sponsor characteristics, and / or investigator characteristics. Geographic characteristics may be included in a vector representing the study country(ies) and study region(ies) in which the clinical trial will be conducted or has been conducted. For example, a clinical trial conducted in Germany, Canada, and the UK would have a categorically coded geographic feature vector representing the three countries and study regions in Europe and North America associated with the trial. Sponsor characteristics include characteristics related to the study's sponsor(s), such as the number of sponsors, the type of sponsor (e.g., government, pharmaceutical manufacturer, contract research organization, etc.), and sponsor experience. Sponsor experience includes the number of previous studies the sponsor has been involved in that have been completed, discontinued, had positive Phase I / II / III results, or had negative Phase I / II / III results. Investigator characteristics may include data about the investigators involved in the clinical trial, such as the number of investigators and the investigator experience.

[0029] Keyword features relate to study keywords associated with a clinical trial. Keyword features may include coded representations associated with study keywords such as "randomized," "open label," and "pharmacodynamics." Example coding techniques include one-hot encoding, dummy encoding, effect encoding, and hash encoding. Keyword features may also include coded representations associated with annotations associated with a clinical trial, such as "indication expansion," "expanded access," and "investigator initiated." Keyword features may also include coded representations of Medical Subject Heading (MeSH) terms.

[0030] Other features relate to various aspects of the clinical trial not included in the above feature groupings, such as route of administration (e.g., injectable, inhaled, topical, etc.), drug origin (e.g., chemical, biological, etc.), or therapeutic area (e.g., oncology, autoimmunology, etc.). In all such instances, categorical features may be coded using any suitable categorical coding technique, such as one-hot encoding, dummy encoding, effect encoding, hash encoding, etc.

[0031] The elements of the test configuration vector associated with each of the above features may be obtained from many different data sources. As shown in FIG. 2, elements 206-1, 206-2, and 206-3 of the test configuration vector 202 associated with feature "A" are obtained from a first data source 218, while element 210 of the test configuration vector 202 associated with feature "C" is obtained from a second data source 220. In this example, the first data source 218 may correspond to a pharmacological database or other source containing information related to the uses, effects, etc. of different drugs. The second data source 220 may correspond to a database or other source related to the past performance of clinical trial sponsors. Features such as biological features, chemical features, and sponsor features may be obtained from publicly available databases such as the US and EU clinical trial registry databases, ChemBL, etc. Some features, such as keyword features, may be extracted from metadata associated with records in such databases (e.g., from web pages associated with the trial).

[0032] Each element of a test configuration vector for a clinical trial can be fixed or optimizable. Fixed elements of a test configuration vector should be understood to be immutable, i.e., the fixed elements of the test configuration vector do not change during subsequent processing or optimization. Optimizable elements of a test configuration vector should be understood to be mutable, i.e., the optimizable elements of the test configuration can be changed or altered during subsequent processing or optimization. In embodiments, fixed or optimizable test configuration elements are predetermined. These elements may be identified by metadata associated with multiple features of the clinical trial. Whether an element is fixed or optimizable may depend on the feature to which it pertains. For example, features related to the pharmacology of a drug being tested in a clinical trial may be fixed, while certain features related to the design and conduct of the clinical trial may be optimizable. Advantageously, by dichotomizing the test configuration vector into fixed and optimizable elements, the configuration of the clinical trial can be manipulated and / or optimized while retaining meaningful results. Identifying fixed and optimizable elements in this manner helps ensure that the updated or optimized clinical trial configuration has achievable results, thereby enabling optimization of the clinical trial configuration.

[0033] Referring again to FIG. 1, the test configuration vector 118 is used by the test outcome predictor 102 to determine a probabilistic model 120 of the clinical outcome.

[0034] The test outcome predictor 102 includes a machine learning model, which may be an unsupervised model, a supervised model, or an ensemble model (i.e., an ensemble of unsupervised and / or supervised models). The machine learning model may be one of a k-nearest neighbor model, a random forest model, an elastic net model, or a support vector machine (SVM) model. Those skilled in the art will understand that the present disclosure is not intended to be limited to only such models, and that any suitable machine learning or predictive model (e.g., rule-based model, fuzzy model, probabilistic model, etc.) may be used.

[0035] In one embodiment, the test outcome predictor 102 includes an ensemble model that combines predictions from a set of unsupervised and / or supervised models by defining weighting coefficients for each model in the ensemble that minimize cross-validation risk (e.g., mean squared error). Each model in the ensemble model may include one or more hyperparameters, such as the neighborhood size parameter k in a k-nearest neighbor model or the minimum node size parameter in a random forest model. Each model may be associated with a set of possible hyperparameters. The ensemble model may then be trained by identifying the best-performing model (i.e., model + hyperparameter selection) from among these sets.

[0036] In one embodiment, the ensemble models include a k-nearest neighbor model, a random forest model, an elastic net model, and a support vector machine (SVM) model. The k-nearest neighbor model has k=[2,10] possible parameter sets. The random forest model has a minimum node size parameter taken from the set {1,2,3} and a minimum node size parameter taken from the set {1,2,3}.

number

[0037] In one embodiment, the trial outcome predictor 102 includes a causal model. For example, the trial outcome predictor 102 may include a Bayesian network or a deconfounder-based model. One example of such a model is a probabilistic principal component analysis (PPCA) model fitted using stochastic variational inference (SVI) and evidence lower bound (ELBO) optimization. The causal model thus estimates the causal relationship between the configuration and outcome of a clinical trial. The causal model can then be used for causal inference (i.e., determining the outcome of a clinical trial when elements in a clinical trial vector change).

[0038] As shown in FIG. 1 , the trial outcome predictor 102 may be trained on historical clinical trial data 110 using a training unit 108. Here, the training unit 108 can be understood as a computational unit or a unit that trains a machine learning model on the training data. Thus, the training unit 108 may be separate from other units of the system 100. For example, the training unit 108 may be part of an external system specifically configured to utilize specialized hardware and / or software to train the outcome predictor. The training unit 108 may utilize a suitable training algorithm, such as stochastic gradient descent, ADAM, etc., to generate the trained machine learning model. In the case of an ensemble model, the training unit 108 may simultaneously train (i.e., fit) individual models using a suitable training technique and determine the best-performing model and a weighted average of all models. Those skilled in the art will appreciate that any suitable training algorithm for the machine learning model used may be utilized by the training unit 108.

[0039] In one implementation, the historical clinical trial data includes trial configuration data corresponding to 9,297 historical clinical trials conducted before 2018. This data includes information related to 5,409 positive trials (i.e., clinical trials with successful outcomes) and 3,888 negative trials (i.e., clinical trials with unsuccessful outcomes). Each clinical trial in the data is represented by a trial configuration vector having 553 elements with features related to the biological features, chemical features, design and operational features, keyword features, and other features described above. The trial outcome predictor 102 may be trained to predict a single outcome for a given trial configuration vector. For example, a first trial outcome predictor may be used to predict the probability that a clinical trial will progress from Phase I to Phase II, and a second trial outcome predictor may be used to predict the probability that a serious adverse event will occur. Other possible trial outcomes are described in more detail below.

[0040] The test outcome predictor 102 receives as input a test configuration vector 118 and provides as output a probabilistic model 120 of the clinical trial outcome. The probabilistic model may include a probability score or value associated with the clinical trial outcome. The probability score or value represents the probability that the clinical trial outcome will be achieved. The probabilistic model may further include an uncertainty estimate. The uncertainty estimate may be associated with the probability score. In embodiments, the probabilistic model is a probability distribution, such as a probability density function, associated with the clinical trial outcome. The probability density function may be determined from the predictions obtained by the test outcome predictor 102 using parametric or non-parametric density estimation techniques.

[0041] The probabilistic model 120 represents the probability of achieving a clinical trial outcome. As such, different trial outcome predictors can be trained and used to provide different trial outcome predictions. In one embodiment, the trial outcome corresponds to the overall success of the clinical trial, whereby the trial outcome predictor is trained to predict a probability score including the probability of the clinical trial's success. In one embodiment, the trial outcome corresponds to the clinical trial transitioning from Phase 1 to Phase 2, whereby the trial outcome predictor is trained to predict a probability score including the probability of the clinical trial transitioning from Phase 1 to Phase 2. Alternatively, the trial outcome corresponds to the clinical trial transitioning from Phase 2 to Phase 3, whereby the trial outcome predictor is trained to predict a probability score including the probability of the clinical trial transitioning from Phase 2 to Phase 3. In a further embodiment, the trial outcome includes a serious adverse event occurring, whereby the trial outcome predictor is trained to predict a probability score including the probability of the serious adverse event occurring as part of the clinical trial. Examples of serious adverse events include interventions to prevent permanent impairment or damage, physical disability or permanent injury, hospitalization, and death. When training the different trial outcome predictors described above, outcomes (e.g., clinical trial success, occurrence of serious adverse events, etc.) are included as targets in the training data.

[0042] According to one aspect of the present disclosure, determining factors that contribute to predicted trial outcomes can help improve model interpretation and subsequent optimization of clinical trial configurations. These factors may be represented as contribution scores. In this manner, the explainability model 104 uses the probabilistic model 120 determined by the trial outcome predictor 102 to calculate multiple contribution scores 122 for the trial configuration vector 118. As will be described in further detail below, the multiple contribution scores 122 indicate the relative contribution of each element (feature value) in the trial configuration vector 118 to the clinical trial outcome.

[0043] The explainability model 104 includes an explainability algorithm that calculates a plurality of contribution scores 122. In an embodiment, the explainability algorithm utilizes both the probabilistic model 120 and the machine learning model of the test outcome predictor 102 to calculate the plurality of contribution scores 122.

[0044] Generally, the explainability algorithm used by the explainability model 104 determines the relative contribution, or influence, that each feature of the clinical trial vector 118 has on the outcome of the clinical trial. This relative contribution can be either positive or negative, such that a particular feature value or element of the clinical trial vector 118 can have a positive or negative impact on the outcome of the clinical trial. In one embodiment, the explainability algorithm uses a feature permutation technique to determine the relative contribution of the features. Baseline measurements s b (e.g., probabilities associated with clinical trial outcomes) are obtained from the trial outcome predictor 102 given the clinical trial vector 118. The elements i in the clinical trial vector 118 associated with the feature i are then permuted to generate a transformed clinical trial vector. Given this transformed clinical trial vector, the permuted measurements s i is obtained from the test outcome predictor 102. The difference between the baseline measurement and the permuted measurement, i.e., s b -s i is recorded, and the process of sorting the elements is repeated several times to obtain the average difference between the two measurements. This average represents the contribution of the element (i.e., feature) to the overall outcome of the clinical trial. A positive average value indicates improved performance (i.e., improved trial results) when this element is included in the clinical trial vector 118. A negative average value indicates decreased performance when this element is included in the clinical trial vector 118.

[0045] Alternatively, in a further embodiment, if the test outcome predictor 102 utilizes a random forest model, the explainability algorithm comprises a random forest feature importance algorithm based on the average reduction in impurity (e.g., mean squared error reduction, Gini, log loss, etc.), hi another embodiment, the explainability algorithm comprises a model-independent method such as breakDown, LIME, or SHAP.

[0046] The explainability algorithm may be applied to all features of the clinical trial vector 118 to obtain a contribution score for each element of the clinical trial vector 118. Alternatively, the explainability algorithm may be applied to a subset of the features of the clinical trial vector 118. For example, the explainability algorithm may be applied only to those elements of the clinical trial vector 118 that are optimizable. By focusing on the optimizable elements of the clinical trial vector 118, the plurality of contribution scores 122 provides a compact representation of the impact of modifiable clinical trial features and, therefore, provides insight into which features may be selected for further processing or optimization. In this manner, the disclosed explainability model provides a database representation and / or visual representation, or other explanation, of how training data impacts or otherwise determines the output of the disclosed AI model (e.g., a trial outcome predictor) and its associated output. In other words, the explainability model visualizes how the AI ​​model is currently being trained and how such training impacts the output results. This allows the AI ​​model to be retrained or reconfigured using different training data, e.g., different test configuration vectors with different fixed and / or optimizable parameters, and / or using different data from additional data sources, to eliminate errors and / or biases in a second version of the AI ​​model, e.g., a test outcome predictor and its associated output.

[0047] 3A and 3B show examples of contribution scores according to embodiments of the present disclosure.

[0048] FIG. 3A shows multiple contribution scores 302 (e.g., multiple contribution scores 122 shown in FIG. 1 ) for five different features “A” through “E.” Features “C,” “D,” and “E” are fixed features (i.e., these features have fixed components in the clinical trial vector), while features “A” and “B” are optimizable features (i.e., these features have adjustable components in the clinical trial vector), as indicated by the underlined lines. In the example shown in FIG. 3A , multiple contribution scores 302 are calculated using an explainability model for a random forest-based trial outcome predictor, whereby multiple contribution scores 302 indicate the average decrease in accuracy results for each feature. This metric can be understood as the loss of accuracy (i.e., in predicting the outcome of a clinical trial) that occurs when the corresponding feature is removed from the clinical trial vector. Thus, multiple contribution scores 302 encode the relative importance of each feature to the overall outcome of the clinical trial. In the example shown in FIG. 3A , the order is such that removing feature “E” results in the greatest decrease in outcome accuracy, thus indicating that feature “E” is the most important feature for the overall outcome of the clinical trial.

[0049] FIG. 3B shows multiple contribution scores 304 (e.g., multiple contribution scores 122 in FIG. 1 ) for five different features “F” through “J.” Features “H,” “I,” and “J” are fixed features, while features “F” and “G” are optimizable features. The multiple contribution scores 304 further include contribution scores for all other features in the clinical trial vector. FIG. 3B also shows the overall outcome determined from a probabilistic model of the clinical trial. The multiple contribution scores 304 shown in FIG. 3B were determined using breakDown, a model-independent explainability model, and correspond to the contribution that each of features “F” through “J” makes to the overall outcome when that feature assumes a particular value. That is, a feature’s contribution score corresponds to the contribution that feature makes to the overall outcome given that feature’s value or element in the clinical trial vector. In the example shown in FIG. 3B , feature “I” with element i1 in the clinical trial vector results in improved outcomes. This improvement is indicated by the left and right arrows that show the difference between the results without feature “I” (left side of the arrow) and with feature “I” (right side of the arrow). In contrast, if feature "F" has element f1 in the clinical trial vector, the results are degraded. The amount of this degradation is indicated by the double arrows showing the difference in results with feature "F" (to the left of the arrow) and without feature "F" (to the right of the arrow).

[0050] The contribution scores shown in Figures 3A and 3B provide insight into the impact of elements of the clinical trial vector on predicted trial outcomes. This insight can help drive improvements to the clinical trial, which may then improve the overall likelihood of achieving the trial outcome. Furthermore, distinguishing between the contributions provided by fixed and optimizable parameters can help facilitate the optimization process by identifying elements that, if optimized, have the greatest potential to improve trial outcomes.

[0051] 1 , the probabilistic model 120 and one or more of the plurality of contribution scores 122 are used to form an explainable prediction 124 of the trial outcome of the clinical trial. One or more of the plurality of contribution scores 122 used to generate the explainable prediction 124 correspond to contribution scores associated with optimizable elements of the clinical trial vector 118. As a result, the explainable prediction 124 indicates what improvements can be made to increase the likelihood that the trial outcome will be achieved.

[0052] The explainable predictions 124 are output for review by the user 116. The explainable predictions 124 may be output in a format similar to that described above in connection with Figures 3A and 3B. Alternatively, the explainable predictions 124 may be output in a structured format (e.g., a JSON file) for further processing or manipulation. In one embodiment, the explainable predictions 124 are included in the report 114 that is output for review by the user 116.

[0053] FIG. 4 illustrates a portion of an example report according to an embodiment of the present disclosure.

[0054] 4 illustrates a probabilistic model 402 of the outcome of a clinical trial and multiple contribution scores 404. An overall probability of success 406 is shown alongside the probabilistic model 402. The multiple contribution scores 404 include a first contribution score 408, a second contribution score 410, and a third contribution score 412. The third contribution score 412 is associated with an optimizable feature 414.

[0055] In the example of FIG. 4 , the clinical trial corresponds to a Phase II trial of two interventions in patients with advanced urothelial carcinoma. The outcome corresponds to the overall success of the clinical trial, whereby the probabilistic model 402 includes a posterior probability distribution of the probability of success of the clinical trial. The probabilistic model 402 may be determined using a trial outcome predictor, such as the trial outcome predictor 102 of the system 100 of FIG. 1 . The overall probability of success 406 is approximately 0.3, with uncertainty estimates (95% confidence interval) of 0.18 and 0.44. The plurality of contribution scores 404 may be calculated using an explainability model, such as the explainability model 104 of the system 100 of FIG. 1 . The plurality of contribution scores 404 are ordered according to the magnitude of their contribution to the overall probability of success 406 of the clinical trial. One skilled in the art will appreciate that the labeling of the features (e.g., “A,” “B,” etc.) in the plurality of contribution scores 404 is for illustrative purposes. A first contribution score 408 is associated with feature “A,” which corresponds to a design and operational feature of the clinical trial. Specifically, feature "A" corresponds to the number of discontinued trials associated with a sponsor of the clinical trial. The first contribution score 408 has an overall negative contribution to the predicted outcome of the clinical trial, thereby contributing to a lower overall probability of success. Thus, the first contribution score 408 indicates that the largest single factor contributing to the overall probability of success 406 is the number of discontinued trials associated with one of the sponsors of the clinical trial. The second contribution score 410 is associated with feature "B," which corresponds to another design and operational feature of the clinical trial. Specifically, feature "B" corresponds to the number of sponsors involved in the clinical trial. The second contribution score 410 has an overall positive contribution and therefore contributes to an improved overall probability of success. The third contribution score 412 is associated with feature "D," which corresponds to another design and operational feature of the clinical trial. Specifically, feature "D" corresponds to the number of principal investigators involved in the clinical trial. The third contribution score 412 has an overall negative contribution and therefore contributes to a lower overall probability of success. In this example, feature "D" is an optimizable feature, meaning that the elements in the clinical trial vector associated with feature "D" can be changed.This indicates that adjusting the number of investigators involved in a clinical trial may help improve the overall probability of success. Therefore, the multiple contribution scores 404 contained within the example report shown in Figure 4 can help identify potential improvements to a clinical trial. These improvements may optimize the probability of achieving the clinical trial's outcomes and may improve the overall design of the clinical trial.

[0056] In this manner, the information contained in the report may be used to update the configuration of the clinical trial (i.e., update one or more of the optimizable elements of the clinical trial vector). In one embodiment, the updated configuration of the clinical trial is obtained from an external source, such as a user or an external system. Alternatively, the updated configuration of the clinical trial is obtained by the optimization process.

[0057] Referring again to FIG. 1 , the first updated test configuration vector 126 may be obtained from an external source (e.g., user 116) to determine an updated probability model of the clinical trial outcome from the test outcome predictor 102 based on the first updated test configuration vector 126 (in a manner similar to that described above with respect to test configuration vector 118). The first updated test configuration vector 126 includes one or more elements optimized based on the explainable predictions 124, where the optimized elements in the first updated test configuration vector 126 correspond to updates, changes, or adjustments to the optimizable elements in the test configuration vector 118. In this manner, the first updated test configuration vector 126 corresponds to the test configuration vector 118 with one or more elements adjusted or optimized by an external source (e.g., user 116 or another system). The updated probability model may then be output for review by the user 116. In one embodiment, updated contribution scores for the updated test configuration are calculated using the explainability model 104 (in a manner similar to that described above with respect to contribution scores 122). The updated contribution scores may then be compared to the contribution scores 122 to determine what changes in the contribution scores result from updating one or more optimizable elements of the test configuration vector 118. The results of the comparison may be output for review by the user 116.

[0058] In an alternative embodiment, the test configuration vector 118 may be updated using an optimization process employed by the optimizer 106. An example of the optimization process is shown in FIG.

[0059] FIG. 5 illustrates an example optimization process 500 according to an embodiment of the present disclosure.

[0060] FIG. 5 illustrates a first test configuration 502 (test configuration vector), a first updated test configuration 504, and a second updated test configuration 506 associated with a clinical trial. The first test configuration 502 includes fixed test parameter values ​​508 and optimizable test parameter values ​​510-1. The first updated test configuration 504 includes fixed test parameter values ​​508 and first updated optimizable test parameter values ​​510-2. The second updated test configuration 506 includes fixed test parameter values ​​508 and second updated optimizable test parameter values ​​510-3. FIG. 5 also illustrates an optimizer 512 that can interface with a predictor 514 to determine the updated test configurations. In one embodiment, the optimizer 512 corresponds to the optimizer 106 of the system 100 of FIG. 1, and the predictor 514 corresponds to the test result predictor 102 of the system 100 of FIG. 1.

[0061] The example optimization process 500 includes steps i, i+1, ..., i+n. In a first step, i, a first clinical trial configuration 502 associated with a clinical trial is optimized to create a first updated trial configuration 504. The optimization performed in step i results in improved clinical trial outcomes (e.g., the optimization results in an increased probability of the clinical trial progressing from Phase I to Phase II). The first updated trial configuration 504 created in the first step, i, includes the same values ​​or factors as the first trial configuration 502 for fixed test parameter values ​​508, but has first updated optimal possible test parameter values ​​510-2 determined by an optimizer 512. In a next step, i+1, the process is repeated, but this time using the first updated trial configuration 504 to determine a new trial configuration that improves the clinical trial outcomes. This process is repeated until a final step, i+n, where the process terminates. Thus, the second updated test configuration 506 determined in step i+(n-1) is output from the optimization process 500 as the final, or optimized, test configuration.

[0062] In one embodiment, the optimizer 512 obtains or generates the updated test configuration using a greedy heuristic. At a general level, such an approach evaluates the performance of several candidate test configurations at each step and selects the best-performing candidate test configuration as the updated test configuration. Here, performance may be measured using a predictor 514 and thus corresponds to the estimated outcome (e.g., probability of success) of the clinical trial given the test configuration. The candidate test configuration may be determined by obtaining or generating configurations in the neighborhood of the current test configuration (e.g., by permuting the optimizable test parameter value(s)). The candidate test configuration that results in the greatest improvement in the estimated outcome is then selected as the best-performing candidate test configuration. In other embodiments, the optimizer 512 obtains or generates the updated test configuration using an optimization algorithm such as hill climbing, tabu search, or simulated annealing. Those skilled in the art will understand that the present disclosure is not intended to be limited to such optimization approaches and that any suitable algorithm or method may be used to obtain an optimized test configuration that results in improved outcomes of the clinical trial.

[0063] In one embodiment, the optimizeable test parameter values ​​510-1 to which the optimization process 500 is applied are selected based on the contribution scores associated with those values. For example, contribution scores for the optimizeable elements of the test configuration vector (e.g., elements in the plurality of contribution scores 122 shown in FIG. 1 ) can be obtained or generated. An optimizeable element is selected for optimization if it has a contribution score that meets predetermined criteria. Examples of predetermined criteria include a negative contribution, a contribution score that is below a predetermined threshold, and a contribution score that is associated with a particular characteristic. An optimization process (e.g., optimization process 500) is used to optimize the optimizeable elements, thereby improving the overall outcome of the clinical trial. In this way, the system can automatically identify aspects of a clinical trial that could be improved and optimize these elements to improve the likelihood that the clinical trial's trial outcomes will be met. This provides an efficient and effective mechanism for improving the design of a clinical trial and helps improve the likelihood that a clinical trial will achieve its trial outcomes before the trial begins.

[0064] Once the optimization technique used by optimizer 512 is complete, a final test configuration is obtained, i.e., second updated test configuration 506 in Figure 5. The optimization technique may terminate after a predetermined number of steps have been performed. Alternatively, the optimization technique may terminate after subsequent iterations have achieved less than a predetermined amount of improvement to the clinical trial results or have remained unchanged over a set number of iterations.

[0065] The systems described in connection with Figures 1 through 5 above can be used to provide efficient and accurate predictions of clinical trial outcomes. Providing explainable predictions of outcomes can identify the contributions of different clinical trial features, thereby enabling deeper insight into the predictions. Furthermore, explainable predictions can help facilitate clinical trial optimization by identifying features of a clinical trial that can be optimized to help improve the likelihood of achieving the trial outcome.

[0066] 6A and 6B show the results of applying the system of the present disclosure to predicting the outcomes of multiple clinical trials.

[0067] The results shown in Figures 6A and 6B correspond to the prediction of the outcome (positive or negative) of a Phase III cancer clinical trial using the methodology described in connection with Figures 1-5. The trial outcome predictor used an elastic net model and was trained on 1,982 trials completed before January 1, 2018. The results shown in Figures 6A and 6B were obtained from a retention test set of 168 trials completed after January 1, 2018. The training data included 779 successful trials and 1,203 failed trials. The test data included 66 successful trials and 102 failed trials. Characteristics of each trial included the sponsor type (e.g., government, pharmaceutical company, etc.), target type, mechanism of action, MESH terms associated with the trial, trial region, and trial country.

[0068] Figure 6A shows the receiver operating characteristic (ROC) curve for the true positive rate (sensitivity) and false positive rate (1-specificity) of the results obtained on the test set. Figure 6B shows the precision-recall graph for the results obtained on the test set. The system achieved an AUC of 0.773, an AUPR of 0.699, and a precision of approximately 85% and a recall of 10%.

[0069] 7A and 7B show predicted success probabilities for two clinical trials according to embodiments of the present disclosure.

[0070] FIG. 7A shows the success probability obtained by the system 100 of FIG. 1 for the first clinical trial of axitinib for renal cell carcinoma (RCC). FIG. 7B shows the success probability obtained by the system 100 of FIG. 1 for the second clinical trial of axitinib for RCC. As shown, the predicted success probability for the first clinical trial is 0.29 (±0.05), and the predicted success probability for the second clinical trial is 0.84 (±0.03). Notably, both the first and second clinical trials had the same sponsor, the same symptoms, and the same medication. However, by incorporating richer features from different categories (e.g., biological features, chemical features, design and operational features, etc.) into the clinical trial vector, the disclosed approach was able to correctly predict that the first clinical trial was likely to fail, while the second clinical trial was likely to succeed.

[0071] FIG. 8A illustrates a method 800 for generating explainable predictions of test outcomes for clinical trials according to an embodiment of the present disclosure.

[0072] The method 800 includes step 802 of obtaining a test configuration vector, step 804 of determining a probability model based on the test configuration vector, step 806 of calculating a contribution score for the test configuration vector based on the probability model, step 808 of generating an explainable prediction of the test outcome based on the contribution score and the probability model, and step 810 of outputting the explainable prediction.

[0073] In an obtaining step 802, a trial configuration vector associated with a clinical trial (e.g., trial configuration vector 118 of system 100 of FIG. 1 ) is obtained from one or more data sources (e.g., one or more data sources 112 of system 100 of FIG. 1 ). Additionally or alternatively, obtaining the trial configuration vector may include generating the trial configuration vector from one or more data sources. In such aspects, generating may include modifying elements (e.g., to be fixed and / or to be optimized) of the (otherwise updated) trial configuration to determine, select, or create the trial and / or trial configuration vector. The trial configuration vector encodes aspects related to the clinical trial design and protocol and includes one or more fixed elements and one or more optimizable elements. Each element or value of the trial configuration vector is associated with a characteristic of the clinical trial and may be binary, integer, or real-valued. The clinical trial to which the trial configuration vector relates may be a proposed clinical trial that has not yet begun or an ongoing clinical trial that has already begun.

[0074] The trial configuration vector may include one or more elements associated with one or more biological features, where the one or more biological features are related to a target associated with the clinical trial. The one or more biological features may include at least one hierarchical mechanism-of-action feature. The trial configuration vector may include one or more elements associated with one or more chemical features, where the one or more chemical features are related to a target associated with the clinical trial. The trial configuration vector may include one or more elements associated with one or more design and operational features of the clinical trial. The one or more design and operational features may include one or more geographic features related to an investigative site associated with the clinical trial. The one or more design and operational features may include one or more sponsor features related to a sponsor associated with the clinical trial. The one or more design and operational features may include one or more investigator features related to an investigator associated with the clinical trial. The trial configuration vector may include one or more elements associated with keywords associated with the clinical trial. The trial configuration vector may include other features related to various aspects of the clinical trial not covered by the feature groupings above.

[0075] In determining step 804, a probabilistic model of the outcome of the clinical trial (e.g., probabilistic model 120 of system 100 of FIG. 1 ) is determined using a trial outcome predictor (e.g., trial outcome predictor 102 of system 100 of FIG. 1 ) based on the trial configuration vector. In one embodiment, the outcome corresponds to the overall success of the clinical trial, whereby the trial outcome predictor predicts a probability score including the probability of the clinical trial's success. In one embodiment, the outcome corresponds to the clinical trial progressing from Phase 1 to Phase 2, whereby the trial outcome predictor predicts a probability score including the probability of the clinical trial transitioning from Phase 1 to Phase 2. Alternatively, the outcome corresponds to the clinical trial progressing from Phase 2 to Phase 3, whereby the trial outcome predictor predicts a probability score including the probability of the clinical trial transitioning from Phase 2 to Phase 3. In a further embodiment, the outcome includes the occurrence of a serious adverse event, whereby the trial outcome predictor predicts a probability score including the probability of the serious adverse event (e.g., intervention to prevent permanent impairment or damage, disability or permanent damage, hospitalization, and death) occurring as part of the clinical trial.

[0076] The trial outcome predictor includes a predictor trained on data related to multiple past clinical trials. Further details regarding training the trial outcome predictor are provided above in connection with the training unit 108 of the system 100 of FIG. 1. The trial outcome predictor includes a machine learning model, which may be an unsupervised model or a supervised model. Examples of such models include k-nearest neighbor models, random forest models, elastic net models, and support vector machines. Alternatively, the trial outcome predictor may include an ensemble model.

[0077] The probabilistic model of a clinical trial outcome includes a probability score associated with the outcome of the clinical trial. The probability score or value represents the probability that the outcome of the clinical trial will be achieved. The probability score may include the probability that the clinical trial will progress from Phase 1 to Phase 2. The probability score may include the probability that the clinical trial will progress from Phase 2 to Phase 3. The probability score may include the probability of a serious adverse event occurring as part of the clinical trial. The probabilistic model of a clinical trial outcome may further include an uncertainty estimate.

[0078] In a calculating step 806, a probabilistic model-based explainability model (e.g., explainability model 104 of system 100 of FIG. 1 ) is used to calculate a plurality of contribution scores (e.g., a plurality of contribution scores 122 of system 100 of FIG. 1 ) for the trial configuration vector. Each of these plurality of contribution scores indicates the relative contribution of an associated element of the trial configuration vector to the outcome of the clinical trial. Examples of contribution scores are shown and described in connection with FIGS. 3A and 3B above.

[0079] In a generating step 808, an explainable prediction of the test outcome of the clinical trial (e.g., explainable prediction 124 of FIG. 1 ) is generated based on the probabilistic model and one or more contribution scores of the plurality of contribution scores. These one or more contribution scores are associated with one or more optimizable factors. In some embodiments, the explainable prediction further includes one or more additional contribution scores associated with one or more fixed factors.

[0080] In an outputting step 810, the explainable predictions are output for review by a user (e.g., user 116 shown in FIG. 1). In some embodiments, the explainable predictions are included in a report (e.g., report 114 shown in FIG. 1), which is then output for review by a user. A portion of an example report is shown and described in connection with FIG. 4 above.

[0081] FIG. 8B illustrates a method 812 that includes additional steps that may be performed as part of the method 800 of FIG. 8A according to an embodiment of the present disclosure.

[0082] The steps of method 812 may be performed after completing the steps of method 800. In particular, method 812 may be performed after step 808 of generating explainable predictions or step 810 of outputting explainable predictions.

[0083] The method 812 includes step 814 of obtaining an updated test configuration vector, step 814 of determining an updated probability model based on the updated test configuration vector, step 818 of calculating an updated contribution score based on the updated probability model, step 820 of determining changes to the contribution score, and step 822 of outputting the changes to the contribution score.

[0084] In an obtaining step 814, an updated test configuration vector associated with the clinical trial (e.g., the first updated test configuration vector 126 or the second updated test configuration vector 126 shown in FIG. 1 ) is obtained. The updated test configuration vector includes one or more elements that have been optimized based on the explainable predictions. The updated test configuration vector may be obtained from a user or an external source, such as an external computer system. Additionally or alternatively, obtaining the first test configuration (updated or not) may include generating the first test configuration from one or more data sources. In such an aspect, generating may include modifying elements (e.g., to be fixed and / or to be optimized) of the test configuration (updated or not) to determine, select, or create the test and / or test configuration vector.

[0085] In a determining step 816, an updated probability model of the outcome of the clinical trial is determined based on the updated trial configuration vector using a trial outcome predictor.

[0086] Optionally, after an updated probabilistic model is determined, method 812 outputs the updated probabilistic model of the clinical trial results for a user to review. The updated probabilistic model provides feedback to the user regarding changes in the trial results that result from changes made to the trial configuration vector. This feedback may aid in the explainability and / or optimization of the clinical trial.

[0087] In a calculating step 818, updated contribution scores for the updated test configuration vector are calculated using an explainability model based on the updated probability model.

[0088] In a determining step 820, one or more changes to the plurality of contribution scores are determined based on a comparison of the plurality of contribution scores to the plurality of updated contribution scores.

[0089] In an outputting step 822, the one or more changes to the plurality of contribution scores are output for user review. The user may then review the changes to the study results and the contribution of each feature to the study results as a result of modifying the study configuration vector. This feedback may provide further insight into clinical trial design and operational features and optimizations, helping to improve clinical trial planning and execution.

[0090] FIG. 9 illustrates a method 900 for optimizing parameters of a clinical trial according to an embodiment of the present disclosure.

[0091] Method 900 includes obtaining 902 a first test configuration associated with a clinical trial, obtaining 904 an outcome predictor, optimizing 906 the first test configuration to improve the outcome of the clinical trial, and outputting 908 the updated test configuration. Optimizing 906 includes calculating 910 updated values ​​for optimizable test parameters of the first test configuration, and creating 912 an updated test configuration including the updated values ​​of the optimizable test parameters. In some embodiments, method 900 further includes generating 914 a report and transmitting 916 the report.

[0092] In an obtaining step 902, a first test configuration (e.g., test configuration vector 118 shown in FIG. 1) associated with a clinical trial is obtained from one or more data sources (e.g., one or more data sources 112 shown in FIG. 1). The first test configuration includes values ​​associated with one or more fixed test parameters and at least one optimizable test parameter.

[0093] An outcome predictor (e.g., trial outcome predictor 102 of system 100 of FIG. 1) is obtained in an obtaining step 904. The outcome predictor estimates the relationship between the trial configuration of a clinical trial and the outcome of the clinical trial.

[0094] As described above, the outcome predictor may include a supervised model, an unsupervised model, or an ensemble model. In one embodiment, the outcome predictor includes a causal model, such that the relationship determined by the outcome predictor between the test configuration of the clinical trial and the outcome of the clinical trial includes the causal relationship determined by the causal model.

[0095] In an optimization step 906, the first trial configuration is optimized (e.g., by the optimizer 106 of the system 100 shown in FIG. 1) to improve the outcome of the clinical trial. In this manner, the system can automatically identify aspects of the clinical trial that could be improved and optimize these elements to improve the likelihood that the clinical trial's trial outcomes will be met. This provides an efficient and effective mechanism for improving the design of a clinical trial and helps improve the likelihood that the clinical trial will achieve its trial outcomes before the clinical trial begins. In one embodiment, optimizable elements in the trial configuration are identified for optimization. The optimizable elements may be identified manually (e.g., by a user) or automatically based on a contribution score associated with the optimizable elements. For example, the identified optimizable elements may correspond to the optimizable elements in the clinical trial vector that provide the largest negative contribution to the overall outcome of the clinical trial. In such an embodiment, an explainability model (e.g., the explainability model 104 of the system 100 of FIG. 1) may be used to calculate contribution scores for the optimizable elements of the clinical trial vector prior to the optimization step 906.

[0096] In a calculating step 910, an updated value of at least one optimizable test parameter is calculated using the outcome predictor and the first test configuration such that a first estimated outcome of the clinical trial is greater than a second estimated outcome of the clinical trial. The first estimated outcome is determined from the outcome predictor based on the updated test configuration, and the second estimated outcome is determined from the outcome predictor based on the first test configuration. In one embodiment, the updated value of the at least one optimizable test parameter is calculated using a greedy heuristic, as described above. Alternatively, an optimization algorithm such as hill climbing, tabu search, or simulated annealing is used to obtain the updated value of the at least one test parameter.

[0097] In a creating step 912, an updated test configuration (eg, second updated test configuration vector 128 shown in FIG. 1) is created that includes an updated value of at least one optimizable test parameter.

[0098] In an outputting step 908, the updated test configuration is output for review by a user (e.g., user 116 shown in FIG. 1 ). Optionally, estimation results associated with the updated test configuration are also output for review by the user.

[0099] In an embodiment, method 900 further includes generating 914 a report that includes one or more of the updated test configuration values. A portion of an example report is shown and described above in connection with FIG. 4. Method 900 may further include transmitting 916 the report for display to a user. For example, the report may be generated on a first device or system and transmitted (e.g., via a local area network, a wide area network, the Internet, etc.) to a second device or system, where the report can be displayed to the user. In such a configuration, the systems and data used to predict clinical trial outcomes and generate the report may be kept separate and secure from the user, thereby reducing user access to potentially sensitive data used to generate the report.

[0100] In some embodiments, the method 900 further includes calculating a first plurality of contribution scores for the updated test configuration using an explainability model (e.g., the explainability model 104 of the system 100 shown in FIG. 1 ). Each contribution score of the first plurality of contribution scores indicates a relative contribution of an associated value of the updated test configuration to the first estimated result. In some embodiments, the report further includes one or more of the first plurality of contribution scores for the updated test configuration.

[0101] The method 900 may further include calculating a second plurality of contribution scores for the first test configuration using the explainability model. Each contribution score of the first plurality of contribution scores indicates a relative contribution of an associated value of the first test configuration to the second estimated outcome. In some embodiments, the report further includes one or more of the second plurality of contribution scores for the first test configuration. In further embodiments, the report further includes a comparison of the first plurality of contribution scores for the updated test configuration with the second plurality of contribution scores for the first test configuration.

[0102] The optimization process in Figure 9 provides an efficient and effective mechanism for improving clinical trial planning and understanding. This optimization process further helps improve the probability that a clinical trial will achieve its trial outcomes before the trial begins.

[0103] The systems and methods of the present disclosure (described in connection with FIGS. 1-9 above) may be implemented in hardware or a combination of hardware and software. For example, they may be implemented as dedicated hardware devices, software libraries, or network packages integrated into network applications. In one embodiment, the present disclosure is implemented in software, such as a program running on an operating system.

[0104] Figure 10 illustrates an exemplary computing system for performing methods of the present disclosure. Specifically, Figure 10 illustrates a block diagram of one embodiment of a computing system according to exemplary aspects and embodiments of the present disclosure.

[0105] The computing system 1000 can be configured to perform any of the operations disclosed herein, such as, for example, any of the operations described with reference to FIGS. 1-10 . The computing system includes one or more computing devices 1002. The one or more computing devices 1002 of the computing system 1000 include one or more processors 1004 and memory 1006. The one or more processors 1004 can be any general-purpose processor(s) configured to execute a set of instructions. For example, the one or more processors 1004 can be one or more general-purpose processors, one or more field programmable gate arrays (FPGAs), and / or one or more application-specific integrated circuits (ASICs). In one embodiment, the one or more processors 1004 include a single processor. Alternatively, the one or more processors 1004 include multiple processors operatively connected. The one or more processors 1004 are communicatively coupled to the memory 1006 via an address bus 1008, a control bus 1010, and a data bus 1012. The memory 1006 may be random access memory (RAM), read-only memory (ROM), persistent storage such as a hard drive, and / or erasable programmable read-only memory (EPROM), etc. The one or more computing devices 1002 further include an input / output (I / O) interface 1014 communicatively coupled to the address bus 1008, the control bus 1010, and the data bus 1012.

[0106] The memory 1006 may store information accessible by the one or more processors 1004. For example, the memory 1006 (e.g., one or more non-transitory computer-readable storage media, memory devices) may include computer-readable instructions (not shown) that may be executed by the one or more processors 1004. The computer-readable instructions may be software written in any suitable programming language and may be implemented in hardware. Additionally, or alternatively, the computer-readable instructions may execute in logically and / or virtually separate threads on the one or more processors 1004. For example, the memory 1006 may store instructions (not shown) that, when executed by the one or more processors 1004, cause the one or more processors 1004 to perform operations such as any of the operations and functions for which the computing system 1000 is configured, as described herein. Additionally, or alternatively, the memory 1006 may store data (not shown) that may be retrieved, received, accessed, written, manipulated, created, and / or stored. In some implementations, one or more computing device(s) 1002 may retrieve data from and / or store data in one or more memory device(s) located remotely from the computing system 1000.

[0107] The computing system 1000 further includes a storage device 1016, a network interface 1018, an input controller 1020, and an output controller 1022. The storage device 1016, the network interface 1018, the input controller 1020, and the output controller 1022 are communicatively coupled via the I / O interface 1014.

[0108] The storage device 1016 is a computer-readable medium, optionally a non-transitory computer-readable medium, that includes one or more programs that, when executed by the one or more processors 1004, include instructions that cause the computing system 1000 to perform the method steps of the present disclosure. Alternatively, the storage device 1016 is a transitory computer-readable medium. The storage device 1016 can be a persistent storage device, such as a hard drive, cloud storage, or any other suitable storage device.

[0109] The network interface 1018 may be a Wi-Fi module, a network interface card, a Bluetooth module, and / or any other suitable wired or wireless communication device. In one embodiment, the network interface 1018 is configured to connect to a network, such as a local area network (LAN) or a wide area network (WAN), the Internet, or an intranet.

[0110] FIG. 10 illustrates an example of a computing system 1000 that can be used to implement the present disclosure. Other computing systems can also be used. Computing tasks described herein as being performed by and / or performed by one or more functional unit(s) can instead be performed remotely from the respective system, or vice versa. Such configurations can be implemented without departing from the scope of the present disclosure. The use of computer-based systems allows for the organization, combination, and division of a wide variety of tasks and functions among components. Computer-implemented operations can be performed on a single component or across multiple components. Computer-implemented tasks and / or operations can be performed serially or in parallel. Data and instructions can be stored in a single storage device or across multiple storage devices.

Claims

1. 1. A computer-implemented method for generating explainable predictions about test outcomes of a clinical trial, the computer-implemented method comprising: obtaining, by one or more processors, a trial configuration vector associated with the clinical trial from one or more data sources in communication with the one or more processors, the trial configuration vector including one or more fixed elements and one or more optimizable elements; determining, by the one or more processors, a probabilistic model of the outcome of the clinical trial based on the trial configuration vector using a test outcome predictor, the test outcome predictor being trained with data related to a plurality of past clinical trials; calculating, by the one or more processors, a plurality of contribution scores for the trial configuration vector based on the probabilistic model using an explainability model, each contribution score of the plurality of contribution scores indicating a relative contribution of an associated element of the trial configuration vector to the outcome of the clinical trial; generating, by the one or more processors, an explainable prediction of the test outcome of the clinical trial based on the probabilistic model and one or more contribution scores of the plurality of contribution scores, the one or more contribution scores being associated with the one or more optimizable elements; and outputting, by the one or more processors, the explainable predictions for review by a user. Computer-implemented methods.

2. obtaining, by the one or more processors, an updated trial configuration vector associated with the clinical trial, the updated trial configuration vector including one or more optimized elements based on the explainable prediction; and 2. The computer-implemented method of claim 1, further comprising: determining, by the one or more processors, an updated probability model of the outcome of the clinical trial using the test outcome predictor based on the updated test configuration vector.

3. 3. The computer-implemented method of claim 2, further comprising: outputting, by the one or more processors, the updated probabilistic model regarding the outcome of the clinical trial for review by the user.

4. calculating a plurality of updated contribution scores for the updated test configuration vector based on the updated probability model using the explainability model; and 3. The computer-implemented method of claim 2, further comprising: determining, by the one or more processors, one or more changes to the plurality of contribution scores based on a comparison of the plurality of contribution scores to the plurality of updated contribution scores.

5. The computer-implemented method of claim 4 , further comprising outputting, by the one or more processors, the one or more changes to the plurality of contribution scores for review by a user.

6. 10. The computer-implemented method of claim 1, wherein the test outcome predictor comprises a model selected from a list comprising a k-nearest neighbor model, a random forest model, an elastic net model, and a support vector machine.

7. The computer-implemented method of claim 1 , wherein the test outcome predictor comprises an ensemble model.

8. The computer-implemented method of claim 1 , wherein the explainable prediction further comprises one or more additional contribution scores associated with the one or more fixed factors.

9. 10. The computer-implemented method of claim 1, wherein the trial configuration vector includes one or more elements associated with one or more biological features, the one or more biological features relating to a target associated with the clinical trial.

10. The computer-implemented method of claim 9 , wherein the one or more biological features include at least one hierarchical mechanistic feature.

11. 10. The computer-implemented method of claim 1, wherein the trial configuration vector includes one or more elements associated with one or more chemical features, the one or more chemical features relating to a target associated with the clinical trial.

12. The computer-implemented method of claim 1 , wherein the trial configuration vector includes one or more elements associated with one or more design and operational features of the clinical trial.

13. 13. The computer-implemented method of claim 12, wherein the one or more design and operational features include one or more geographic features associated with investigative sites associated with the clinical trial.

14. 13. The computer-implemented method of claim 12, wherein the one or more design and operational characteristics include one or more sponsor characteristics related to a sponsor associated with the clinical trial.

15. 13. The computer-implemented method of claim 12, wherein the one or more design and operational features include one or more investigator features related to an investigator associated with the clinical trial.

16. The computer-implemented method of claim 1 , wherein the trial configuration vector includes one or more elements associated with keywords associated with the clinical trial.

17. The computer-implemented method of claim 1 , wherein the explainable predictions are included in a report, whereby the report is output for review by a user.

18. The computer-implemented method of claim 1 , wherein the probabilistic model of the outcome of the clinical trial includes a probability score associated with the outcome of the clinical trial.

19. 20. The computer-implemented method of claim 18, wherein the probability score comprises a probability that the clinical trial will progress from Phase 1 to Phase 2.

20. 20. The computer-implemented method of claim 18, wherein the probability score comprises a probability that the clinical trial will progress from Phase 2 to Phase 3.

21. 20. The computer-implemented method of claim 18, wherein the probability score comprises a probability of a serious adverse event occurring as part of the clinical trial.

22. 20. The computer-implemented method of claim 18, wherein the probabilistic model of the outcome of the clinical trial further comprises an uncertainty estimate.

23. The computer-implemented method of claim 1 , wherein obtaining the test configuration vector from the one or more data sources comprises generating the test configuration vector from the one or more data sources.

24. 1. A non-transitory machine-readable medium storing instructions for generating explainable predictions of test outcomes of a clinical trial, the instructions, when executed by one or more processors, causing the one or more processors to: obtaining, from one or more data sources in communication with the one or more processors, a trial configuration vector associated with the clinical trial, the trial configuration vector including one or more fixed elements and one or more optimizable elements; determining a probabilistic model of the outcome of the clinical trial based on the trial configuration vector using a test outcome predictor, the test outcome predictor being trained with data related to a plurality of past clinical trials; calculating a plurality of contribution scores for the trial configuration vector based on the probabilistic model using an explainability model, each contribution score of the plurality of contribution scores indicating a relative contribution of an associated element of the trial configuration vector to the outcome of the clinical trial; generating an explainable prediction of the test outcome of the clinical trial based on the probabilistic model and one or more contribution scores of the plurality of contribution scores, the one or more contribution scores being associated with the one or more optimizable elements; and outputting the explainable predictions for user review.

25. 1. A system configured to generate explainable predictions of test outcomes of a clinical trial, the system comprising: one or more processors; a memory, when executed by the one or more processors, that causes the one or more processors to: obtaining, from one or more data sources in communication with the one or more processors, a trial configuration vector associated with the clinical trial, the trial configuration vector including one or more fixed elements and one or more optimizable elements; determining a probabilistic model of the outcome of the clinical trial based on the trial configuration vector using a test outcome predictor, the test outcome predictor being trained with data related to a plurality of past clinical trials; calculating a plurality of contribution scores for the trial configuration vector based on the probabilistic model using an explainability model, each contribution score of the plurality of contribution scores indicating a relative contribution of an associated element of the trial configuration vector to the outcome of the clinical trial; generating an explainable prediction of the test outcome of the clinical trial based on the probabilistic model and one or more contribution scores of the plurality of contribution scores, the one or more contribution scores being associated with the one or more optimizable elements; and outputting the explainable prediction for user review. Equipped with system.

26. 1. A computer-implemented method for optimizing parameters of a clinical trial, the method comprising: obtaining, by one or more processors, from one or more data sources in communication with the one or more processors, a first test configuration associated with the clinical trial, the first test configuration including values ​​associated with one or more fixed test parameters and at least one optimizable test parameter; obtaining, by the one or more processors, an outcome predictor, the outcome predictor estimating a relationship between a test configuration of the clinical trial and an outcome of the clinical trial; optimizing, by the one or more processors, the first test configuration to improve an outcome of the clinical trial, the optimizing step comprising: calculating, by the one or more processors, an updated value of the at least one optimizable test parameter using the outcome predictor and the first test configuration such that a first estimated outcome of the clinical trial is greater than a second estimated outcome of the clinical trial; generating, by the one or more processors, an updated test configuration that includes the updated value of the at least one optimizable test parameter; optimizing, wherein the first estimated result is determined from the result predictor based on the updated test configuration, and the second estimated result is determined from the result predictor based on the first test configuration; outputting, by the one or more processors, the updated trial configuration for user review; Including, Computer-implemented methods.

27. The outputting step includes: generating, by the one or more processors, a report including one or more values ​​of the updated test configuration; transmitting, by the one or more processors, the report for display to a user.

27. The computer-implemented method of claim 26.

28. 28. The computer-implemented method of claim 27, further comprising: calculating, using an explainability model, a first plurality of contribution scores for the updated test configuration, each contribution score of the first plurality of contribution scores indicating a relative contribution of an associated value of the updated test configuration to the first estimated outcome.

29. 30. The computer-implemented method of claim 28, wherein the report further includes one or more of the first plurality of contribution scores for the updated test configuration.

30. 29. The computer-implemented method of claim 28, comprising: using the explainability model to calculate a second plurality of contribution scores for the first test configuration, each contribution score of the first plurality of contribution scores indicating a relative contribution of an associated value of the first test configuration to the second inferred outcome.

31. 31. The computer-implemented method of claim 30, wherein the report further includes one or more of the second plurality of contribution scores for the first test configuration.

32. 31. The computer-implemented method of claim 30, wherein the report further includes a comparison of the first plurality of contribution scores for the updated test configuration and the second plurality of contribution scores for the first test configuration.

33. 27. The computer-implemented method of claim 26, wherein the outcome predictor comprises a causal model.

34. 34. The computer-implemented method of claim 33, wherein the relationship determined by the outcome predictor between the test configuration of the clinical trial and the outcome of the clinical trial comprises a causal relationship determined by the causal model.

35. 27. The computer-implemented method of claim 26, wherein obtaining the first test configuration from the one or more data sources comprises generating the first test configuration based on the one or more data sources.

36. 1. A non-transitory machine-readable medium having stored thereon instructions for optimizing parameters of a clinical trial, the instructions, when executed by a device including one or more processors, causing the one or more processors to: obtaining, from one or more data sources in communication with the device, a first test configuration associated with the clinical trial, the first test configuration including values ​​associated with one or more fixed test parameters and at least one optimizable test parameter; obtaining an outcome predictor, the outcome predictor estimating a relationship between a test configuration of a clinical trial and an outcome of the clinical trial; optimizing the first test configuration to improve an outcome of the clinical trial, the optimizing step comprising: calculating an updated value of the at least one optimizable test parameter using the outcome predictor and the first test configuration such that a first estimated outcome of the clinical trial is greater than a second estimated outcome of the clinical trial; creating an updated test configuration including an updated value of the at least one optimizable test parameter; optimizing, wherein the first estimated result is determined from the result predictor based on the updated test configuration, and the second estimated result is determined from the result predictor based on the first test configuration; outputting the updated test configuration for user confirmation; Non-transitory machine-readable media.

37. 1. A system configured to optimize parameters of a clinical trial, the system comprising: one or more processors; a memory, when executed by the one or more processors, that causes the one or more processors to: obtaining, from one or more data sources in communication with the one or more processors, a first test configuration associated with the clinical trial, the first test configuration including values ​​associated with one or more fixed test parameters and at least one optimizable test parameter; obtaining an outcome predictor, the outcome predictor estimating a relationship between a test configuration of a clinical trial and an outcome of the clinical trial; optimizing the first test configuration to improve an outcome of the clinical trial, the optimizing step comprising: calculating an updated value of the at least one optimizable test parameter using the outcome predictor and the first test configuration such that a first estimated outcome of the clinical trial is greater than a second estimated outcome of the clinical trial; creating an updated test configuration including an updated value of the at least one optimizable test parameter; optimizing, wherein the first estimated result is determined from the result predictor based on the updated test configuration, and the second estimated result is determined from the result predictor based on the first test configuration; outputting the updated test configuration for user confirmation; a memory storing instructions; system.