Method and system for machine learning-assisted simulation techniques in clinical data
A machine learning algorithm optimizes clinical trial designs by mapping multidimensional parameters, reducing simulation requirements and enhancing efficiency and adaptability in clinical trials.
Patent Information
- Application Number
- PCT/IL2025/050521
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-06-16
- Publication Date
- 2025-12-26
AI Technical Summary
Clinical trials face inefficiencies due to the high computational demand and complexity of simulating various configurations, leading to increased costs and reduced responsiveness to real-time data, especially in advanced and complex trials with multiple variables.
A machine learning algorithm is employed to map the multidimensional parameter space of clinical trial designs, reducing the need for extensive simulations by learning the relationship between design parameters and operating characteristics, thereby optimizing trial setups.
The method significantly enhances the speed and efficiency of clinical trial design by requiring at least 25 times fewer simulations while maintaining predictive accuracy, allowing for a broader exploration of design parameters and increased adaptability.
Smart Images

Figure IL2025050521_26122025_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM FOR MACHINE LEARNING-ASSISTED SIMULATION TECHNIQUES IN CLINICAL DATA
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to computer-implemented methods and systems for data processing and decision support in clinical studies, and more particularly to a technically implemented framework for identifying optimal clinical trial design parameters using machine learning trained on simulation outcomes.
[0004] The disclosure addresses technical challenges associated with identifying an optimal working point in a high-dimensionality space by implementing computational techniques.
[0005] BACKGROUND
[0006] Clinical trials, particularly advanced and complex ones, face significant challenges due to the absence of straightforward analytical methods to calculate their operating characteristics.
[0007] Designers often need to conduct extensive simulation, which is not only expensive, but also slow, owing to the high volume of simulations needed for accuracy. Furthermore, complex trials typically have multiple degrees of freedom (e.g., number of interim analyses, timing of interims, and the like), making manual optimization nearly impossible without relying on heuristic methods, based on prior trials and subjective intuition.
[0008] Brute force simulation involves running a large number of simulations to calculate the operating characteristics of a trial under various configurations. This process is computationally intensive as each simulation might require substantial time and processing power, particularly when the trials are complex with many variables. This computational demand translates into higher costs, both in terms of the time invested by research teams and the financial cost associated with high-performance computing resources.
[0009] Each simulation provides estimates that are inherently uncertain. Therefore, to achieve reliable results, many iterations are needed. However, the more complex the trial design (e.g., adaptive trials with multiple interim analyses and multiple adaptations), the greater the number of simulations required to meet statistical requirements, such as type 1 error and the like. This can lead to inefficiencies, as the process becomes not only slower, but also less responsive to real-time data and adjustments.
[0010] As the number of variables increases (such as different dosage levels, timing of doses, or number of interim analyses), the dimensionality of the simulation space explodes. The complexity of design process can make it impractical to explore all potential combinations of trial parameters thoroughly, thereby limiting the scope of trial designs that can be feasibly evaluated.
[0011] Thus, there is a need to provide a method and system for calculating and estimating the operating parameters of clinical trials, in an efficient manner saving cost, time and computational resources, without compromising accuracy.
[0012] SUMMARY
[0013] According to some embodiments, methods and systems for identifying optimal clinical trial design parameters using machine learning trained on simulation outcomes is presented herein.
[0014] According to some embodiments, the method disclosed herein employs a machine learning algorithm using results of a series of simulations to learn the relationship between a trial’s design parameters and its operating characteristics. By integrating advanced computational techniques, the machine learning model efficiently maps out the multidimensional parameters space and supports the required optimization process.
[0015] According to some embodiments, the method of the present disclosure, presents a robust model that uses machine learning to estimate and predict the operating characteristics of clinical trials, while taking into consideration various design adjustments.
[0016] According to some embodiments, the methods and systems presented herein advantageously reduces the need for extensive simulations dramatically, achieving the required predictive accuracy with at least 25 times fewer simulations compared to brute force methods.
[0017] It is understood that the reduced number of required simulations enable running simulations over various scenarios. For example, for a Phase I trial simulations may be run for various anticipated toxicity scenarios.
[0018] According to some embodiments, through rigorous training / testing methodologies, the machine learning algorithm model has been validated to ensure high accuracy and reliability in predicting trial outcomes.
[0019] According to some embodiments, the implementation of the machine learning model in clinical trial design advantageously enhances the speed and efficiency of trial setup significantly. By reducing the dependency on extensive simulations, trial designers can explore a broader array of design parameters more quickly and with greater precision. This not only saves time and resources but also potentially increases the efficacy and adaptability of clinical trials.
[0020] According to some embodiments, in one aspect, a method for identifying optimal clinical trial design parameters is presented herein. The method comprises: a. receiving / inputting a plurality of optional clinical trial design parameters collectively defining a space of working points, optionally the first subset of working points comprises 80-5000 of working points out of the totality of working points in the space of working point; b. selecting a first subset of working points from within the first space of working points, wherein each working point is defined by a different set of clinical trial design parameters, optionally wherein the first subset of working points comprises 80-5000 of working points out of the totality of working points in the space; c. running a small plurality of simulations for each of the selected working points to obtain their respective simulated outcome; d. training an ML model on said small plurality of simulations and their respective simulated outcomes to obtain an ML model configured to output predicted simulation outcomes for non-simulated working points within the space of working points, optionally wherein the plurality of simulations comprises between 50 and 5000, or between 50-1000 simulations; e. applying the trained ML model on non-simulated working points from within the space of working points, thereby mapping the space; f. defining an improved space of working points based on the mapping; g. selecting a second subset of working points from within the improved space of working points, optionally the second subset of working points comprises 80- 5000 of working points out of the totality of working points in the space of working point; h. running a second plurality of simulations on the second subset of working points; , optionally wherein the second small plurality of simulations comprises between 50 and 5000, or between 50-1000 simulations; i. updating the trained ML model, based on simulated treatment outcomes of the second small plurality of simulations to obtain an updated trained ML model; j. repeating steps f-i until obtaining an optimal working point comprising a defined set of clinical trial design; k. outputting for example on a user interface, a clinical trial design comprising the defined set of clinical trial design parameters.
[0021] Advantageously, the herein disclosed method and systems achieves a predetermined predictive accuracy with at least 25 times fewer simulations as compared to brute force methods.
[0022] According to some embodiments, for each iteration a “objective function” or “reward function “configured to evaluate the fitness of a simulated working point may be applied. The fitness i.e. ‘how good an evaluated working point it is’ is then utilized to guide the qualitative change in the working points utilized for a next iteration. That is, according to some embodiments, the step of selecting working points from within the space of working points may include applying the objective function.
[0023] According to some embodiments, the a small plurality of simulations may include about 2, 3, 4, 5 ,6, 10, 20, 50, or 100 simulations (or any range therebetween). Each possibility is a separate embodiment. According to some embodiments, each additional iteration (step g) may include the same or a larger number of simulations (but less than 5000, and preferably less than 1000).
[0024] According to some embodiments, the method further comprises running 100000 simulations after identifying / outputting the optimal working point with an optimized set of clinical trial design parameters (or on a chosen ‘close to optimal’ working point, according to regulatory requirements. According to some embodiments, identifying / outputting an optimal set of clinical trial design parameters comprises optimizing sample size, cost of the clinical trial, duration of the clinical trial, estimated treatment efficacy of the trial, number of interims, probability of success of the trial or any combination thereof.
[0025] According to some embodiments, the method may further include applying a Paretto Curve on the optimal working point identified in order to assess the trade-off between efficacy and toxicity and / or efficacy and cost associated with choosing a close to optimal working point instead of the computed optimal working point. For example, if the optimal working point suggests a sample size of 155, the applying of the Paretto curve may output how chances of success are affected if the sample size is reduced to 150 and how does it influence overall cost. According to some embodiments, the ‘close to optimal working points’ and their associated cost may also be displayed via the user interface.
[0026] According to some embodiments, the method further comprises outputting, for the optimal set of trial design parameters, one or more of: a probability of overall trial success, a probability of finding a best treatment as a function of the number of patients included in the trial, estimated distribution of cost and time of the trial overall, estimated distribution of cost and time until identification of failure, estimated distribution of cost and time until identification of success, distribution of estimated treatment effect, distribution of statistical measures or any combination thereof. Each possibility and combination of possibilities is a separate embodiment.
[0027] According to some embodiments, at least a portion of the clinical and / or statistical input parameters comprise value ranges.
[0028] According to some embodiments, the value ranges are predetermined.
[0029] According to some embodiments, the method may further include determining / computing suitable ranges for the portion of clinical and / or statistical input parameters.
[0030] According to some embodiments, in case of discrete parameters, such as but not limited to, number of patients, number of interims and the like, the method may include a step of adjustment of the optimal working point if a non-discrete value is obtained for one or more of the parameters (e.g. 179.5 patients or 3.2 interims). In short, at times the theoretical best value might lie between possible discrete choices. Therefore adjustments (typically in the form of applying searching techniques) need to be made to find the close as possible to optimal discrete value. That is, the aim is to find a configuration (combination of values limited to discrete values for discrete data) that approaches what would be the “optimum” if the problem had a continuous relaxation.
[0031] According to some embodiments, the method may be applied for a plurality of scenarios as described herein.
[0032] According to some embodiments, the scenarios may be user selected. Additionally or alternatively, the method may include automatically applying the method for at least two predefined scenarios. According to some embodiments, the at least two scenarios may be utilized as reference scenarios.
[0033] According to some embodiments, the selection of working points of step (b) is given and / or computed.
[0034] According to some embodiments, the number of simulations included in the plurality of simulations is predetermined.
[0035] According to some embodiments, the number of simulations included in the plurality of simulations is determined based on a number of simulations required to obtain an accuracy above a predetermined threshold.
[0036] According to some embodiments, the plurality of simulations comprises between 50 and 5000 simulations, between 100 and 1000 simulations or any other range within the range of 50-5000. Each possibility is a separate embodiment.
[0037] According to some embodiments, optional (or close to optimal) clinical trial design parameters comprise clinical and statistical input parameters.
[0038] According to some embodiments, the clinical parameters are selected from primary endpoint, delay, number of arms, futility threshold efficacy, efficacy threshold how good before deciding success assumed clinical efficacy, recruitment rate, primary endpoint metrics, secondary endpoints, number of interims and any combination thereof. Each possibility is a separate embodiment.
[0039] According to some embodiments, the statistical input parameters are selected from target power (chance of succeeding per number of patients), type I error, allocation logic, statistical test and threshold and any combination thereof. According to some embodiments, defining the improved space of working points comprises selecting clinical trial design parameters optimizing operating characteristics of the clinical trial design and / or clinical and / or statistical input parameters optimizing the power of the ML model.
[0040] According to some embodiments, the method further comprises conducting a large plurality of simulations for the identified optimal clinical trial design parameters.
[0041] According to some embodiments, in another aspect a method for identifying optimal clinical trial design parameters is presented herein. The method comprises:
[0042] (a) receiving a plurality of clinical trial simulations and their associated simulation outcomes for each of a plurality of working points, wherein each working point is selected from a space of working points defined by optional clinical trial design parameters, optionally the first subset of working points comprises 80-5000 of working points out of the totality of working points in the space of working point;
[0043] (b) training a machine learning (ML) model on the received number of simulations and their simulated outcomes,
[0044] (c) applying the trained ML model on additional, non-simulated working points from within the space of working points to obtain their respective predicted simulation outcomes, thereby mapping the space of working points;
[0045] (d) defining an improved space of working points based on the mapping;
[0046] (e) selecting a second subset of working points from within the improved space of working points, optionally the second subset of working points comprises 80-5000 of working points out of the totality of working points in the space of working point;
[0047] (f) running a second plurality of simulations on the second subset of working points; optionally wherein the second small plurality of simulations comprises between 50 and 5000, or between 50-1000 simulations;
[0048] (g) updating the trained ML model, based on predicted simulation outcomes of a second small plurality of simulation to obtain an updated ML model, (h) repeating steps c-g until obtaining an optimal working point comprising a defined set of clinical trial design parameters; and
[0049] (i) identifying / outputting a clinical trial design comprising the defined set of trial design parameters.
[0050] Advantageously, the hereindisclosed system provides improved processing capabilities to the processor thus allowing it to reliably and robustly identify optimal working points for clinical trial designs utilizing a fraction (e.g. about 10%, 20%, 25%, 50% or 75% - or any range therebetween) of the working points evaluated using brute force. Each possibility is a separate embodiment.
[0051] According to some embodiments, the method further includes transmitting the clinical study plan to a clinical decision support system.
[0052] According to some embodiments, the method further includes conducting a clinical study using the produced clinical study plan.
[0053] According to some embodiments, the outputting comprises displaying the clinical trial design on a user interface (UI). According to some embodiments, the UI is configured to allow a user to interact and affect changes to one or more subgroups defined by the features defining the multidimensional feature space. According to some embodiments, the changes made by the user are transmitted to or retrieved by the processor. According to some embodiments, the processor is further configured to recompute the optimal working point, based on the affected changes. According to some embodiments, the processor is further configured to regenerate the clinical study design based on the affected changes.
[0054] According to some embodiments, the UI may be utilized by a user to define which parameters are rigid (i.e. the user may decide a rigid cost range, a number of interims or the like) and which parameters should be optimized in view of the values of the rigid parameters. It is understood that both input and output parameters can be defined as rigid or optimizable. For example, an input parameter such as ‘number of interims” may be set as rigid, i.e. the user may request to only provide a working point(s) that includes 3 interims. Similarly, an output parameter, such as “chance of success’ may be set as rigid, i.e. the user may request to only provide a working point(s) that provide an overall chance of success to be at least 85%. According to some embodiments a parameter may be both rigid and optimizable. For example, a rigid range may be provided (e.g., the number of subjects included in the trial must be between 200-500), while the exact number within the range is optimized.
[0055] According to some embodiments, the choice of ML model may depend on the problem. According to some embodiments, the algorithms utilized can vary throughout the process. For example, at early stages a simple model e.g. logistic or linear models can be utilized to aid in guiding the following simulations into the correct part of the design space. As the process continues more sophisticated models (e.g. random forest, gaussian process regression, a boosted regression tree, a regularized GLM, and / or a neural network model), using the results of all prior simulations, are trained. The complex models may estimate the performance for non-monotonous parameters, without degrading the overall performance.
[0056] According to some embodiments, the method further comprising outputting, for the optimal set of trial design parameters, one or more of: a probability of getting overall trial success, a probability of finding a best treatment as a function of the number of patients included in the trial, estimated distribution of cost and time of the trial overall, estimated distribution of cost and time until identification of failure, estimated distribution of cost and time until identification of success, distribution of estimated treatment effect, distribution of statistical measures. Each possibility and combination of probabilities is a separate embodiment.
[0057] According to some embodiments, the number of simulations included in the small plurality of simulations is predetermined.
[0058] According to some embodiments, the number of simulations included in the small plurality of simulations is determined based on a number of simulations required to obtain an accuracy above a predetermined threshold.
[0059] According to some embodiments, the small plurality of simulations comprises between 50 and 5000 simulations or between 50 and 1000 simulations. According to some embodiments, the small plurality of simulations comprises less than 5000, lest than 2000, less than 1000, less than 750 or less than 500 simulations. Each possibility is a separate embodiment.
[0060] According to some embodiments, defining the improved space of working points comprises selecting clinical trial design parameters optimizing operating characteristics of the clinical trial and / or clinical and / or clinical trial design parameters optimizing the accuracy of the ML model(s).
[0061] According to some embodiments, the method further comprises conducting a large plurality of simulations for the identified optimal set of clinical trial design parameters.
[0062] According to some embodiments, the method may be applied for a plurality of scenarios as described herein.
[0063] According to some embodiments, there is provided a system including a memory and a processor coupled to the memory programmed with executable instructions, configuring the processor to: a. receive / input a small plurality of optional clinical trial design parameters collectively defining a space of working points, b. select a first subset of working points from within the first space of working points, wherein each working point is defined by a different set of clinical trial design parameters, optionally wherein the first subset of working points comprises 80-5000 of working points out of the totality of working points in the space; c. run a small plurality of simulations for each of the selected working points to obtain their respective simulated outcome; d. train an ML model on said small plurality of simulations and their respective simulated outcomes to obtain an ML model configured to output predicted simulation outcomes for non-simulated working points within the space of working points, optionally wherein the plurality of simulations comprises between 50 and 5000, or between 50-1000 simulations; e. apply the trained ML model on non-simulated working points from within the space of working points, thereby mapping the space; f. define an improved space of working points based on the mapping; g. select a second subset of working points from within the improved space of working points, optionally the second subset of working points comprises 80- 5000 of working points out of the totality of working points in the space of working point; h. run a second plurality of simulations on the second subset of working points, optionally wherein the second small plurality of simulations comprises between 50 and 5000, or between 50-1000 simulations; i. update the trained ML model, based on simulated treatment outcomes of the second small plurality of simulations to obtain an updated trained ML model; j. repeat steps f-i until obtaining an optimal working point comprising a defined set of clinical trial design; k. output, for example on a user interface, a clinical trial design comprising the defined set of clinical trial design parameters
[0064] According to some embodiments, the system further includes a memory (coupled to the processor) configured to store the selected working points, simulation outcomes, trained ML models and updates thereto. Each possibility and combination of possibilities is a separate embodiment.
[0065] According to some embodiments, the system further comprising the user interface (UI). According to some embodiments, the UI is configured to allow a user to interact and affect changes to one or more subgroups defined by the features of the multidimensional feature space. According to some embodiments, the changes made by the user are transmitted to or retrieved by the processor. According to some embodiments, the processor is further configured to recompute the optimal working point, based on the affected changes. According to some embodiments, the processor is further configured to regenerate and redisplay the clinical study design based on the affected changes on the UI.
[0066] According to some embodiments, the UI may be utilized by a user to define which parameters are rigid (i.e. the user may decide a rigid cost range, a number of interims or the like) and which parameters should be optimized in view of the values of the rigid parameters. It is understood that both input and output parameters can be defined as rigid or optimizable. For example, an input parameter such as ‘number of interims” may be set as rigid, i.e. the user may request to only provide a working point(s) that includes 3 interims. Similarly, an output parameter, such as “chance of success’ may be set as rigid, i.e. the user may request to only provide a working point(s) that provide an overall chance of success to be at least 85%. According to some embodiments a parameter may be both rigid and optimizable. For example, a rigid range may be provided (e.g., the number of subjects included in the trial must be between 200-500), while the exact number within the range is optimized.
[0067] According to some embodiments, the user can, via the UI set a number of scenarios for which optimal working points are each retrieved. For example, for a phase I trial, various toxicity scenarios can be evaluated
[0068] Certain embodiments of the present disclosure may include some, all, or none of the above advantages. One or more other technical advantages may be readily apparent to those skilled in the art from the figures, descriptions, and claims included herein. Moreover, while specific advantages have been enumerated above, various embodiments may include all, some, or none of the enumerated advantages.
[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In case of conflict, the patent specification, including definitions, governs. As used herein, the indefinite articles “a” and “an” mean “at least one” or “one or more” unless the context clearly dictates otherwise.
[0070] BRIEF DESCRIPTION OF THE FIGURES
[0071] Some embodiments of the disclosure are described herein with reference to the accompanying figures. The description, together with the figures, makes apparent to a person having ordinary skill in the art how some embodiments may be practiced. The figures are for the purpose of illustrative description and no attempt is made to show structural details of an embodiment in more detail than is necessary for a fundamental understanding of the disclosure. For the sake of clarity, some objects depicted in the figures are not drawn to scale. Moreover, two different objects in the same figure may be drawn to different scales. In particular, the scale of some objects may be greatly exaggerated as compared to other objects in the same figure.
[0072] In the figures: FIG. 1 schematically shows a flowchart of a method 100 for identifying optimal clinical trial design parameters, according to some embodiments;
[0073] FIG. 2 schematically shows a flowchart of a method 200 for identifying optimal clinical trial design parameters, according to some embodiments; and
[0074] FIG. 3 schematically shows table 1 which compares the designs of the method presented herein with an alternative method design and a fixed design.
[0075] DETAILED DESCRIPTION
[0076] The principles, uses, and implementations of the teachings herein may be better understood with reference to the accompanying description and figures. Upon perusal of the description and figures present herein, one skilled in the art will be able to implement the teachings herein without undue effort or experimentation. In the figures, same reference numerals refer to same parts throughout.
[0077] In the description and claims of the application, the words “include” and “have”, and forms thereof, are not limited to members in a list with which the words may be associated.
[0078] As used herein, the term “about” may be used to specify a value of a quantity or parameter (e.g. the length of an element) to within a continuous range of values in the vicinity of (and including) a given (stated) value. According to some embodiments, “about” may specify the value of a parameter to be between 80 % and 120 % of the given value. For example, the statement “the length of the element is equal to about 1 m” is equivalent to the statement “the length of the element is between 0.8 m and 1.2 m”. According to some embodiments, “about” may specify the value of a parameter to be between 90 % and 110% of the given value. According to some embodiments, “about” may specify the value of a parameter to be between 95 % and 105 % of the given value.
[0079] As used herein, according to some embodiments, the terms “substantially” and “about” may be interchangeable.
[0080] According to some embodiments, methods for identifying optimal clinical trial design parameters, are presented herein. As used herein, the term “clinical trial” refers to prospective biomedical (or behavioral) research studies on human participants designed to answer specific questions about new treatments. They generate data on dosage, safety and efficacy and typically include four phases. According to some embodiments, the clinical trial may be an exploratory phase II clinical trial.
[0081] As used herein, the term “trial simulation” refers to the study of the effects of a drug in virtual patient populations using computational mathematical models. That is, the term "simulation” refers to the use of computational models and statistical algorithms to mimic the behavior of clinical trials before they are actually conducted. The goal is to test, evaluate, and optimize different trial designs under various hypothetical scenarios to improve decisionmaking and increase trial efficiency. According to some embodiments, simulation in clinical trial design may include the process of generating synthetic data based on statistical models of disease progression, patient response, recruitment patterns, or other trial components, to evaluate how different trial setups (e.g., sample sizes, endpoints, randomization schemes) might perform in silico.
[0082] According to some embodiments, the term “arms”, refers to treatment groups of a clinical trial, and may refer to a clinical trial including a single treatment group, two treatment groups (e.g. first medicament and second medicament, first dose and second dose etc.), three treatment groups, four treatment groups, five treatment groups or more. Each possibility is a separate embodiment.
[0083] As used herein, according to some embodiments, the term “clinical trial design parameters” refers to the parameters which define the clinical trial, for example, the number of arms being evaluated, primary outcomes, minimal clinical value required for authorization, expected time to clinical results, historical data, etc.
[0084] According to some embodiments, the clinical trial design parameters may be statistical parameters and / or clinical parameters that influence the simulation outcome and / or the operating characteristics of the clinical trial.
[0085] As used herein, the term “operating characteristics” refers to information on a clinical trial design’s expected behavior under specific conditions. A few of the most common measures are the expected probability of success (i.e. power), the expected effect size of the treatment at the end of the trial, the expected duration of the trial, and the required sample size. As used herein, the term “expected” refers to the fact that these operating characteristics are not deterministic and can vary from trial to trial, even when the “truth” is constant.
[0086] According to some embodiments the clinical trial design parameters may be classified into three types of parameters:
[0087] According to some embodiments the clinical trial design parameters may be internal parameters. As used herein, the term “internal parameters” refers to parameters which have at least one freedom degree, which influence the simulation outcome but are not directly of interest to the client in and of themselves. A non-limiting example of internal parameters is the method used for statistical significance testing as long as it is accepted by the regulatory body.
[0088] According to some embodiments the clinical trial design parameters may be reality parameters. As used herein, the term “reality parameters” refers to parameters which cannot be directly controlled, however, affect directly the operating characteristics of the clinical trial. According to some embodiments, the client is required to assume a value of the parameters, but does not directly control the value that will manifest in practice.
[0089] A non-limiting example of internal parameters is the efficacy of the medicament tested. According to some embodiments, the reality parameters may be parameters provided by the client. As a non-limiting example, a reality parameter provided by the client may be the recruitment rate of patients.
[0090] According to some embodiments the clinical trial design parameters may be external parameters. As used herein, the term “external parameters” refers to parameters which have at least one degree of freedom, and which affect the operating characteristics of the clinical trial. A non-limiting example of an external parameter is the number of interim analyses included in the clinical trial, in that each interim analysis may affect the cost of the trial and / or the duration of the trial and therefore the number of interim analysis is limited by the rigidity of the operating characteristics of the clinical trial.
[0091] According to some embodiments, the term “optimal parameters” and “optimal design parameters” may be used interchangeably and refer to a set / combination of parameters that significantly (e.g. by at least 0.5% or at least 1%) improve one or more of the operating characteristics of the clinical trial. As used herein, according to some embodiments, the term “working point” refers to a point where all the clinical trial design parameters are defined. For example, a first working point may be defined as a trial comparing 3 treatment arms of different doses and a control arm, using response adaptive randomization (RAR) with 2 interims, after 40% of patients are recruited and after 70%, using Thompson sampling for re-allocation based on the probability of each arm being best, without futility or efficacy stopping, with control allocation matched to the leading arm, and with acceptance based on a proportion test with a threshold of 0.01 for each arm. The endpoint is binary (response) and the assumed effect size is from a binomial distribution with a probability of 0.1 for the placebo, 0.15 for arm 1, 0.2 for arm 2 and 0.3 for arm 3. The expected recruitment rate is 1 patient per site per month with an s profile (slow recruitment initially, fast in the mid stage and slow again towards the end of the trial), and with a dropout of 15%. The required power is 0.8 and the required family wise type I error rate is 0.025.
[0092] As used herein, according to some embodiments, the term “a space of working points” refers to the space generated by a large number of working points each defined by a different set / combination of clinical trial design parameters.
[0093] As used herein, according to some embodiments, the term “interim analysis” refers to an analysis of data from an ongoing trial before data collection has been completed. The number of interim as well as their timing is determined prior to commencing the trial. According to some embodiments, interim analysis results may cause modifications in the conduct of the trial. Depending on the results, an interim analysis may lead to changes, such as stopping one treatment arm or changing the number of participants in a group, or stopping the trial altogether.
[0094] As used herein, the term “minimal clinical value” refers to the minimal effect of the medicament that has clinical value. For example, in pancreatic cancer any drug prolonging life expectancy by a month may be considered to have clinical value due to the short overall life expectancy of pancreatic cancer patients.
[0095] As used herein, the term “assumed clinical value” refers to the effect that the client expects to observe in the trial, e.g., due to previous clinical data or pre-clinical data. For example, in pancreatic cancer while any drug prolonging life expectancy by a month may be considered to have clinical value, the client may assume the effect will be substantially larger based on previous clinical or pre-clinical data possessed and design the trial powered to detect this effect size.
[0096] According to some embodiments, the herein disclosed method may include the following steps:
[0097] Step 1 : Randomly sample a set of design configurations from a relevant range of parameters. These are sampled based on optimal experimental design concepts to remain within reasonable bounds, while still optimizing the expected information gain (i.e. more samples at the edge of the design space).
[0098] Step 2: Running a small number of simulations for each of the configurations (can be as few as a single simulation). These simulations output both the binary outcome - success or failure, but also a more detailed output, i.e the test statistic, the effect size to support better training of the ML model.
[0099] Step 3: Train a flexible machine learning model on the results of the simulations to estimate the relationship between the configuration and the operating characteristics. The choice of the specific model may depend on the problem. Moreover, different algorithms can be applied and can vary throughout the process. For example, at early stages a simple model (such as logistic or linear models) is utilized to aid in guiding the following simulations into the correct part of the design space. As the process continues, more sophisticated models, using the results of all prior simulations, can be trained. Advantageously, these models can estimate the performance for non-monotonous parameters, without degrading the overall performance. According to some embodiments, this step may further include searching techniques to optimize discrete values for a non-discrete theoretical value, as essentially described herein.
[0100] Step 4: Estimating the uncertainty of the model. According to some embodiments, the estimation of the uncertainty can be based on analytical results for some of the models, while a more complex Bayesian approach can be applied for more sophisticated models (utilizing the model uncertainties).
[0101] According to some embodiments, steps 1-4 are run iteratively. Each time the newly sampled points are nearer the optimal design and the ML model fits better to the optimal set of parameters and the expected uncertainty decreases, such that it converges to the required design (optimal design). Step 5: Confirming the statistical performance of the identified configuration by running a large number of simulations at the chosen point. According to some embodiments, the loses / gains associated with choosing a close to optimal working point may also be computed.
[0102] The process can terminate after a given number of simulations, a given number of iterations or after converging to the required accuracy.
[0103] According to some embodiments, the process can run while a wide range of assumptions are still examined, and these can be included into the ML model parameters. As a result, the optimal point may be based on a range of assumptions rather than a single point.
[0104] Advantageously, when comparing the number of iterations required for the herein disclosed model to that of brute force search, approximately 25 times fewer simulations are required.
[0105] Reference is now made to FIG. 1 which schematically shows a flowchart of a method 100 for identifying optimal clinical trial design parameters, according to some embodiments.
[0106] In step 110, a plurality of optional clinical trial design parameters collectively defining a space of working points is received and / or inputted. According to some embodiments, the optional clinical trial design parameters include clinical and statistical input parameters. According to some embodiments, a portion of the parameters may be rigid and the remaining optimizable. According to some embodiments, all of the parameters may be optimizable. According to some embodiments, the user may define (via a UI) which parameters are rigid or within a rigid range (semi rigid). According to some embodiments, some of the parameters may be continuous, while others may be discrete.
[0107] In step 120, according to some embodiments, a first subset of working points from within the first space of working points is selected, wherein each working point is defined by a different set of clinical trial design parameters. According to some embodiments, the selection of working points in step 120 is given, random and / or computed. According to some embodiments, the first subset of working points includes between 80-5,000, or between 80- 1,000 or between 80-500 working points Each possibility is a separate embodiment. According to some embodiments, in step 130, a small plurality of simulations for each of the selected working points is run to obtain their respective simulated outcome. According to some embodiments, the number of simulations included in the plurality of simulations is predetermined. According to some embodiments, the number of simulations included in the plurality of simulations is determined based on a number of simulations required to obtain an accuracy above a predetermined threshold. According to some embodiments, the plurality of simulations comprises between 50 and 5000 simulations, between 100 and 1000 simulations or between 100 and 500 simulations. Each possibility is a separate embodiment.
[0108] According to some embodiments, in step 140, a machine learning (ML) model is trained on said plurality of simulations and their respective simulated outcomes, to obtain an ML model configured to output predicted simulation outcomes for additional (simulated and / or nonsimulated) working points within the space of working points. According to some embodiments, the ML model is a logistic regression model, according to some embodiments the ML model is a random forest classifier, according to some embodiments the ML model is a gaussian process regression, according to some embodiments the ML model is a boosted regression tree, according to some embodiments the ML model is a regularized GLM, according to some embodiments the ML model is a neural network. According to some embodiments, various models are applied sequentially. According to some embodiments, initially a simple ML model (such as a linear model and / or logistic regression model) may be applied for a first one or more iteration whereafter a complex ML model (such as a random forest classifier, a boosted regression tree, a regularized GLM, a neural network or any combination thereof - each possibility and combination of possibilities is a separate embodiment) for a subsequent one or more iterations. As a non-limiting example a logistic model may initially be applied followed by utilization of more complex models, such as but not limited to random forest models.
[0109] In step 150, according to some embodiments, the trained ML model is applied on additional non-simulated working points (and optionally also on the simulated working points), to obtain their respective predicted simulation outcomes, thereby mapping the space of working points.
[0110] In step 160, according to some embodiments, an improved space of working points is defined based on the mapping, e.g. according to the contributive predictive power of regions of the space. For example, the predictive power of the ML model may not be expected to improve by including a 4thand 5thinterim analyses, and the second space of working point may therefore exclude working points that include 4thand 5thinterim analyses. According to some embodiments, an “objective function” or “reward function” configured to evaluate the fitness of a simulated working point may be applied to guide the selection of working points.
[0111] Then, a second plurality of simulations is, according to some embodiments, then conducted on working points from the second space of working points, and the trained ML model is updated, based on the simulated treatment outcomes (operating characteristics). As a non-limiting example a logistic model may initially be applied followed by utilization of more complex models, such as but not limited to random forest models.
[0112] According to some embodiments, step 170, further conducting iterations until an optimal working point is obtained. According to some embodiments, the optimality is predefined and combines different relevant parameters. In case an optimal working point is not obtained, steps 120-170 may be repeated until such is obtained.
[0113] When an optimal working point is obtained, according to some embodiments, at step 180, an optimal set of clinical trial design parameters is identified / outputted, based on the predicted simulation outcomes conducted on a plurality of working points from the optimal space of working points. According to some embodiments the optimal space of working points may be encompassed by the first space of working points.
[0114] According to some embodiments, defining the optimal set of clinical trial design parameters includes optimizing sample size, cost of the clinical trial, duration of the clinical trial, estimated treatment efficacy of the trial, probability of success of the trial or any other operating characteristic or combination thereof. Each possibility and combination of possibilities is a separate embodiment.
[0115] According to some embodiments, method 100 further includes a step of displaying the optimal working point and its parameter values to a user via a UI.
[0116] According to some embodiments, at least a portion of the clinical and / or statistical input parameters include value ranges. According to some embodiments, the value ranges may be predetermined. According to some embodiments, the value ranges may be determined during the simulation and may change due to simulation progress. Additionally or alternatively, method 100 further comprises determining / computing suitable ranges for the portion of clinical and / or statistical input parameters.
[0117] According to some embodiments, the clinical parameters are selected from: primary endpoint which is the main result at the end of the trial to see if a given treatment worked, delay (reality parameter), number of arms included (external parameter), futility threshold efficacy (i.e. how bad does the treatment output need to be to stop the trial (internal parameter)), efficacy threshold (how good does the treatment output need to be before deciding success (internal parameter)) assumed clinical efficacy, recruitment rate, primary endpoint metrics (the metrics for measuring success (external parameters)), secondary endpoints (which may provide supportive information about a treatment's effect on the primary endpoint or demonstrate additional effects on the disease or condition) and any combination thereof.
[0118] According to some embodiments, the statistical input parameters are selected from: target power (chance of succeeding per number of patients), allocation logic, statistical test and any combination thereof.
[0119] According to some embodiments, defining the improved space of working points comprises selecting clinical trial design parameters optimizing operating characteristics of the clinical trial design and / or clinical and / or statistical input parameters optimizing the power of the ML model.
[0120] According to some embodiments, the method may be run for a number of scenarios for which optimal working points are each retrieved. For example, for a phase I trial, various toxicity scenarios can be evaluated, such as:
[0121] 1. A ‘no Dose-Limiting Toxicity’ (No DLT) scenario in which all doses tested are well tolerated and only mild adverse events (AEs), such as headaches, nausea, or fatigue, are reported.
[0122] 2. A ‘Mild Toxicity Across All Doses’ scenario in which mild (Grade 1-2) adverse effects are observed at all dose levels (no DLTs).
[0123] 3. An ‘Acute Dose-Limiting Toxicity (DLT) at Low Dose’ scenario in which severe adverse effects (Grade 3 or higher) occur at early dose levels. 4. A ‘Delayed Toxicity’ scenario in which severe toxicity appears after a delay, often after repeated dosing.
[0124] 5. A ‘Cumulative Toxicity’ scenario in which toxic effects increase over time or with repeated dosing.
[0125] 6. An ‘Intermittent Toxicity’ (Reversible) scenario in which toxicity appears, resolves after cessation, and reappears on re-exposure.
[0126] 7. An ‘Idiosyncratic or Immunologic Toxicity’ scenario in which rare, unpredictable toxicities (e.g., drug-induced hepatitis, rash, anaphylaxis) not related to dose are observed.
[0127] 8. An ‘Organ-Specific Toxicity’ scenario in which adverse effect specific to certain organs (e.g. hepatotoxicity, nephrotoxicity or cardiotoxicity) is observed.
[0128] According to some embodiments, the scenarios may be user selected. Additionally or alternatively, the method may include automatically applying the method for at least two predefined scenarios. According to some embodiments, the at least two scenarios may be utilized as reference scenarios. Running such “reference scenarios” may advantageously ensure that skewing of the results is avoided.
[0129] It is understood that the type of scenario influences the optimal working point of the trial design.
[0130] According to some embodiments, the method further comprising conducting a large plurality of simulations for the identified optimal working point. According to some embodiments, the large plurality of simulations may include at least 50,000 simulations or at least 100,000 simulations per working points from the optimal space of working points.
[0131] According to some embodiments, the method further comprises running 100000 simulations after identifying / outputting the optimal set of clinical trial design parameters, as may be required by regulations, or any other number of simulations that may be required by regulations.
[0132] According to some embodiments, the ML model capabilities may be expanded to integrate real-time data from ongoing trials, thereby further optimize trial designs adaptively, making them more responsive to interim results and external factors. According to some embodiments, a method for identifying optimal clinical trial design parameters is presented.
[0133] Reference is now made to FIG. 2 which schematically shows a flowchart of a method 200 for identifying optimal clinical trial design parameters, according to some embodiments.
[0134] In step 210, according to some embodiments, a plurality of clinical trial simulations and their associated simulation outcomes for each of a plurality of working points are received, wherein each working point is selected from a space of working points defined by optional clinical trial design parameters.
[0135] In step 220 according to some embodiments, a machine learning (ML) model is trained on the received number of simulations and their simulated outcomes, to obtain an ML model configured to output predicted simulation outcomes for additional (simulated or non-simulated) working points within the space of working points. According to some embodiments, the ML model is a logistic regression, according to some embodiments the ML model is a random forest classifier, according to some embodiments the ML model is a gaussian process regression, according to some embodiments the ML model is a boosted regression tree, according to some embodiments the ML model is a regularized GLM, according to some embodiments the ML model is a neural network. According to some embodiments, various models are applied sequentially. As a non-limiting example, a logistic model may initially be applied followed by utilization of more complex models, such as but not limited to random forest models.
[0136] In step 230, according to some embodiments, the trained ML model is applied on additional, non-simulated working points from the space of working points to obtain their respective predicted simulation outcomes, thereby mapping the space of working points.
[0137] In step 240, according to some embodiments, an improved space of working points is defined based on the mapping, as essentially described herein above.
[0138] In step 250, according to some embodiments, the trained ML model is updated based on simulation outcomes computed for a plurality of working points within the improved space of working points. According to some embodiments, the updating further comprises exchanging one type of model with another (e.g. logistic model with random forest model) According to some embodiments, in step 260, iterations are conducted until an optimal of working point is identified. In case an optimal working point is not obtained, steps 240-250 may be repeated for the improved space of working points until such is obtained.
[0139] According to some embodiments, defining the optimal set of clinical trial design parameters includes optimizing sample size, cost of the clinical trial, duration of the clinical trial, estimated treatment efficacy of the trial, probability of success of the trial or any other operating characteristic or combination thereof. Each possibility is a separate embodiment.
[0140] According to some embodiments method 200 further comprising outputting, for the optimal set of trial design parameters, one or more of: a probability of getting overall trial success, a probability of finding a best treatment as a function of the number of patients included in the trial, estimated distribution of cost and time of the trial overall, estimated distribution of cost and time until identification of failure, estimated distribution of cost and time until identification of success, distribution of estimated treatment effect, distribution of statistical measures. Each possibility is a separate embodiment.
[0141] According to some embodiments, the number of simulations included in the plurality of simulations is predetermined. According to some embodiments, the number of simulations included in the plurality of simulations is determined based on a number of simulations required to obtain an accuracy above a predetermined threshold. According to some embodiments, the plurality of simulations comprises between 50 and 5000 simulations, between 100 and 1000 simulations or between 100 and 500 simulations. Each possibility is a separate embodiment.
[0142] According to some embodiments, defining the improved space of working points comprises selecting clinical trial design parameters optimizing operating characteristics of the clinical trial and / or clinical and / or clinical trial design parameters optimizing the power of the ML model.
[0143] According to some embodiments, the method further includes conducting a large plurality of simulations for the identified optimal set of clinical trial design parameters. According to some embodiments, the large plurality of simulations may include at least 50,000 simulations or at least 100,000 simulations per working points from the optimal space of working points. According to some embodiments, the ML model capabilities may be expanded to integrate real-time data from ongoing trials, thereby further optimize trial designs adaptively, making them more responsive to interim results and external factors.
[0144] According to some embodiments, a method for identifying optimal clinical trial design parameters is presented.
[0145] The following examples are presented in order to more fully illustrate some embodiments of the invention. They should in no way be construed, however, as limiting the broad scope of the invention. One skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the invention.
[0146] EXAMPLES
[0147] Example 1
[0148] An example of a use case is now presented herein, of a sponsor (also referred to herein as “client”) in the process of developing a novel peptide for a rheumatic disease. The sponsor is interested in examining the potential benefits of an adaptive design.
[0149] Clinical trial design parameters:
[0150] The trial was a phase 2 trial, with multiple doses and treatment regiments. The sponsor was also interested in exploring a potential combination with an existing drug.
[0151] A set of 4 treatment arms, as well as a control arm (that would allow the sponsor to obtain the information required for phase 3) were examined.
[0152] The Estimated effect size in the control arm was expected to be low (5-10%) based on previous data.
[0153] The assumed effect size for the treatment was 30% per arm, and the sample size was calculated to allow for 80% power and a type I error of 2.5%.
[0154] The estimated recruitment rate was high (30 patients per month). The expected time to measure the endpoint for supporting adaptation was 1 month.
[0155] Based on the above, the most relevant degrees of freedom in the trial were:
[0156] 1. Number of interim analysis (1-5)
[0157] 2. Timing of the interim analysis (any time after 2 months until 5 months)
[0158] 3. Maximal sample size (from 100 to 300)
[0159] 4. Test statistic threshold for efficacy at the end of the trial (up to 0.025).
[0160] 5. Type of adaptive design (RAR, GSD, SSR, combinations thereof)
[0161] 6. Aggressiveness (i.e. focus on the “best” arm or on “any promising” arm.)
[0162] For each of the input parameters mentioned above, there are many additional potential configurations (for example each interim can be at any time within the range mentioned, for simplicity it is assumed, equally spaced interims from the first interim to the end of the trial).
[0163] In addition, the sponsor was interested in assessing the performance of the design under higher efficacy assumptions (effect size of 0.45), and lower efficacy assumptions (0.2). The sponsor also presented a number of possible scenarios - a single effective treatment arm, a scenario with two effective arms, and a scenario where all the arms are at least partially effective.
[0164] In addition, the selected scenario was further assessed under a slower recruitment rate, which better aligned with historical precedent.
[0165] The sponsor also wanted to consider a design with a strict type I family-wise error control, as it would better serve as supporting evidence in a regulatory setting.
[0166] Accordingly, even at this high level of initial designing of the trial, the number of potential configurations is very large: assuming only 5 options for each of the input parameters that are examined the total number of combinations is:
[0167] 5 X 5 X 5 X 5 X 5 X5 = 15,625
[0168] All these need to be examined under 3X3X2X3 i.e. 54 assumptions. Thus, in the absence of the herein disclosed method and system almost a million optional combinations would need to undergo the 100K simulations required by regulation bodies. In addition, eventually, the sponsor decided to examine an additional treatment arm, which in the absence of the herein disclosed method and system would require that the whole process be rerun.
[0169] Results
[0170] The method presented herein was applied on the initial trial design, described above. Given the large number of degrees of freedom, the initial phase included running simulations for 2-3 configurations of each of the input parameters - thus decreasing the number of combinations by 3 orders of magnitude. Moreover, for each of these configurations only 10 simulations were run, thus decreasing the total number of simulations by 4 more orders of magnitude.
[0171] This initial step allowed getting a very rough estimate of the range of operating characteristics, using a simple logistic model. This model was only aimed at assessing the directional impact of each parameter to support more relevant sampling at the next iteration.
[0172] Then about 150 additional working points were simulated, with 50 simulations each, focusing on the most promising areas of the space of working points (2 -3 interims, max sample size of 230, test statistic in the range of 0.015 to 0.02, RAR design).
[0173] Advantageously, after completing 6 iterations, and gradually introducing more sophisticated ML models (e.g. starting with logistic and then random forest), an optimal working point and its set of operating parameters were outputted namely: an 33% saving, reduced sample size (175 as opposed to 240) utilizing response adaptive randomization, with 2 interims, at 35% and 68% of the trial, with a maximal sample size of 175 patients, using a less aggressive adaptation at the first interim, and more aggressive adaptation at the second interim.
[0174] Moreover, it was able to assess the performance under different scenarios and assumptions.
[0175] Finally, 100,000 simulations were run at the proposed design configuration, which simulation validated the result. Moreover, a set of simulations with 10,000 runs at 3 additional assumptions was also conducted and the robustness of the design was validated. 1 The proposed design almost quadrupled the expected savings the sponsor was suggested by an alternative method.
[0176] FIG. 3 schematically shows a table, which compares the designs of the method presented herein with an alternative method design and a fixed design (see Hartung J. Biom J. 2006 Aug;48(4):521-36.).
[0177] As can be seen from the table, while in the fixed design only the power and the required samples are defined and no adaptation is / can be applied, the alternative method applies a group sequential design (GSD) adaptation and requires 240 samples, 2 interims and a first interim at 50% of the trial, which leads to a saving of 8%. The method presented herein, applies a response adaptive randomization (RAR) and defines 2 interims with the first interim at 35% of the trial, only 175 required samples, leads to a saving of 33%., thus clearly indicating its superiority.
[0178] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the disclosure. No feature described in the context of an embodiment is to be considered an essential feature of that embodiment, unless explicitly specified as such.
[0179] Although stages of methods, according to some embodiments, may be described in a specific sequence, the methods of the disclosure may include some or all of the described stages carried out in a different order. In particular, it is to be understood that the order of stages and sub-stages of any of the described methods may be reordered unless the context clearly dictates otherwise, for example, when a later stage requires as input an output of a former stage or when a later stage requires a product of a former stage. A method of the disclosure may include a few of the stages described or all of the stages described. No particular stage in a disclosed method is to be considered an essential stage of that method, unless explicitly specified as such.
[0180] Although the disclosure is described in conjunction with specific embodiments thereof, it is evident that numerous alternatives, modifications, and variations that are apparent to those skilled in the art may exist. Accordingly, the disclosure embraces all such alternatives, modifications, and variations that fall within the scope of the appended claims. It is to be understood that the disclosure is not necessarily limited in its application to the details of construction and the arrangement of the components and / or methods set forth herein. Other embodiments may be practiced, and an embodiment may be carried out in various ways.
[0181] The phraseology and terminology employed herein are for descriptive purpose and should not be regarded as limiting. Section headings are used herein to ease understanding of the specification and should not be construed as necessarily limiting.
Claims
CLAIMS1. A method for identifying optimal clinical trial design parameters, the method comprising: a. inputting a plurality of optional clinical trial design parameters collectively defining a space of working points; b. selecting a first subset of working points from within the first space of working points, wherein each work point is defined by a different set of clinical trial design parameters, wherein each working point is defined by a different set of clinical trial design parameters; c. running a small plurality of simulations for each of the selected working points to obtain their respective simulated outcome, wherein the small plurality of simulations comprises between 50 and 5000 simulations; d. training a machine learning (ML) model on said plurality of simulations and their respective simulated outcomes to obtain an ML model configured to output predicted simulation outcomes for non-simulated working points within the space of working points; e. applying the trained ML model on non-simulated working points from within the space of working points, thereby mapping the space; f. defining an improved space of working points based on the mapping; g. selecting a second subset of working points from within the improved space of working points; h. running a second plurality of simulations on the second subset of working points; i. updating the trained ML model, based on simulated treatment outcomes of the second plurality simulations to obtain an updated trained ML model; j . repeating steps d-i for the improved space of working points until obtaining an optimal working points, comprising a defined set of clinical trial design parameters; k. outputting a clinical trial design comprising the defined set of clinical trial design parameter values.
2. The method of claim 1, further comprising running 100000 simulations on the optimal working point.
3. The method of claim 1 or 2, wherein outputting an optimal set of clinical trial design parameters comprises optimizing sample size, cost of the clinical trial, duration of the clinical trial, estimated treatment efficacy of the trial, probability of success of the trial or any combination thereof.
4. The method of any one of claims 1-3, further comprising outputting, for the optimal set of trial design parameters, one or more of: a probability of overall trial success, a probability of finding a best treatment as a function of the number of patients included in the trial, estimated distribution of cost and time of the trial overall, estimated distribution of cost and time until identification of failure, estimated distribution of cost and time until identification of success, distribution of estimated treatment effect, distribution of statistical measures.
5. The method of any one of claims 1-4, wherein at least a portion of the clinical and / or statistical input parameters comprise value ranges.
6. The method of claim 5, wherein the value ranges are predetermined.
7. The method of claim 5 or 6, wherein the method further comprises determining / computing suitable ranges for the portion of clinical and / or statistical input parameters.
8. The method of any one of claims 1-7, wherein the selection of working points of step (b) is given and / or computed.
9. The method of any one of claims 1-8, wherein the selecting of the second subset of working points comprises applying a reward function.
10. The method of any one of claims 1-9, wherein the plurality of simulations comprises between 50 and 1000 simulations.
11. The method of any one of claims 1-10, wherein optional clinical trial design parameters comprise clinical and statistical input parameters.
12. The method of claim 11, wherein the clinical parameters are selected from primary endpoint, delay, number of arms, futility threshold efficacy, efficacy threshold, assumed clinical efficacy, recruitment rate, primary endpoint metrics, secondary endpoints and any combination thereof.
13. The method of claim 11 or 12, wherein the statistical input parameters are selected from target power (chance of succeeding per number of patients), allocation logic, statistical test and any combination thereof.
14. The method of any one of claims 1-13, further comprising conducting a large plurality of simulations for the identified optimal working point.
15. The method of any one of claims 1-14, further comprising applying one or more search techniques to optimize discrete values for non-discrete theoretical values.
16. The method of any one of claims 1-15, wherein the ML model comprises a random forest model.
17. The method of any one of claims 1-16, wherein the ML model comprises a simple ML model in step e and a complex model in at least some of the repeating of step j .
18. The method of claim 17, wherein the simple model comprises a logistic model or a linear model and the complex model is random forest classifier, a gaussian process regression, a boosted regression tree, a regularized GLM, and / or a neural network model.
19. The method of claim 18, wherein the simple model comprises a logistic model and the complex model is random forest classifier.
20. The method of any one of claims 1-19, further comprising conducting a clinical trial utilizing the output clinical trial design.
21. The method of any one of claims 1-20, achieving a predetermined required predictive accuracy with at least 25 times fewer simulations as compared to brute force methods22. A system for identifying optimal clinical trial design parameters, the system comprising:(a) a memory for storing a dataset comprising:• plurality of optional clinical trial design parameters collectively defining a space of working points;• simulation outcomes; and• one or more trained ML models;(b) a processor coupled to the memory and configured to execute instructions that cause the system to: a. input a plurality of optional clinical trial design parameters collectively defining a space of working points; b. select a first subset of working points from within the first space of working points, wherein each work point is defined by a different set of clinical trial design parameters, wherein each working point is defined by a different set of clinical trial design parameters; c. run a small plurality of simulations for each of the selected working points to obtain their respective simulated outcome, wherein the small plurality of simulations comprises between 50 and 5000 simulations; d. train a machine learning (ML) model on said plurality of simulations and their respective simulated outcomes to obtain an ML model configured to output predicted simulation outcomes for non-simulated working points within the space of working points; e. apply the trained ML model on non-simulated working points from within the space of working points, thereby mapping the space; f. define an improved space of working points based on the mapping; g. select a second subset of working points from within the improved space of working points; h. run a second plurality of simulations on the second subset of working points; i. update the trained ML model, based on simulated treatment outcomes of the second plurality simulations to obtain an updated trained ML model;j . repeat steps d-i for the improved space of working points until obtaining an optimal working points, comprising a defined set of clinical trial design parameters; k. output a clinical trial design comprising the defined set of clinical trial design parameter values.
23. The system of claim 22, further comprising a user interface (UI) configured to present the output design parameter values to a user.
24. The system of claim 22 or 23, wherein the UI is configured to allow a user to interact and affect changes to one or more subgroups defined by the features defining the multidimensional feature space.
25. The system of claim 24, wherein the processor is further configured to reidentify a revised optimal working point, based on the affected changes.
Citation Information
Patent Citations
Procedure for determining a clinical trial design to predict added benefit (AMNOG prediction)
DE102021002951A1
Interactive trial design platform
US20210241866A1
Methods and system for reducing computational complexity of clinical trial design simulations
US20210319158A1