System and method for analysing success of clinical trials, interpreting risk factors and for managing pharmaceutical portfolio using artificial intelligence
The use of AI and machine learning in a hierarchical system addresses the challenges of predicting clinical trial success and managing pharmaceutical portfolios by processing risk factor data and historical trial information, resulting in accurate and timely predictive assessments.
Patent Information
- Application Number
- PCT/CA2024/051726
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-26
AI Technical Summary
The pharmaceutical industry faces significant challenges in predicting the success of clinical trials and managing pharmaceutical portfolios effectively, due to the complexity of risk factors involved and the limitations of traditional methods.
A system and method using artificial intelligence and machine learning to analyze the success of clinical trials, interpret risk factors, and manage pharmaceutical portfolios. This involves a hierarchical computerized system that processes data on risk factors and historical clinical trials to predict multiple success levels of clinical trials, including operational, scientific, phase transition, regulatory, and commercial success.
The system provides accurate and real-time predictive assessments of clinical trial success, enabling informed decision-making and optimizing resource allocation in drug development and portfolio management.
Smart Images

Figure CA2024051726_26062025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR ANALYSING SUCCESS OF CLINICAL TRIALS, INTERPRETING RISK FACTORS AND FOR MANAGING PHARMACEUTICAL PORTFOLIO USING ARTIFICIAL INTELLIGENCECross-Reference to Related Applications
[0001] The present patent application claims the benefits of priority of United States Provisional Patent Application No. 63 / 614,374, entitled “METHOD AND SYSTEM FOR PHARMACEUTICAL PORTFOLIO STRATEGIC MANAGEMENT DECISION SUPPORT BASED ON ARTIFICIAL INTELLIGENCE” and filed at the United States Patent and Trademark Office on December 22, 2023, the content of which is incorporated herein by reference.Field of the Invention
[0002] The present invention generally relates to systems, methods and platforms for analyzing success of clinical trials, interpreting risk factors and / or managing pharmaceutical portfolio using machine learning. More particularly, the present invention relates to systems and computer-implemented methods for predictive assessment of the likelihood of success of the drug development candidate at any phase of interventional clinical development and for any conditions.Background of the Invention
[0003] The discovery and clinical development of new drugs is the core business of the innovative pharmaceutical industry and clinical trials are the cornerstone of this innovative process, as a critical step towards bringing new medicines safely to market. However, this discovery and development process is known to be long, complex, very risky and increasingly costly. The statistics associated with the new drug discovery and development process would make this business activity unattractive a priori. Indeed, it's estimated that the entire process may cost more than US$2 billion and may require more than 10 years of development for a new molecular entity before reaching the market. In addition, it is estimated that only around 10% of potential drugs entering Phase I clinical trials will reach the market, and not all drugs that gain marketing authorization will be financially profitable taking individually. In fact, a study estimates that only about 20% of New Molecular Entities (NME) cover their average capitalized R&D expenses and not all marketed drugs will generate revenues that match or exceed R&D costs (Vernon et al., 2010). More importantly, it is in the late stages of clinical drug development that failures are costliest. However, even if clinical drug development is ahigh-risk endeavor, it's a high-return business, guaranteeing a huge return on investment in the event of success. But the continuously increasing R&D expenditures of drug development and the low rate of introduction of new products on the market over time result in a complex and critical business model issue for pharmaceutical innovation. Because of this, it is therefore vital for the pharmaceutical industry to continuously seek new strategies to better control risk of failure in drug development. Having the means to select the best candidates efficiently and accurately with a high probability of success, and focus resources on them, is the Holy Grail. Numerous approaches have been proposed over the last few decades to remedy this highly complex problem. Roughly speaking, on the one hand, there are upstream approaches that focus solely on optimizing drug discovery steps, on the premise that a better structure-activity relationship of the molecule is a guarantee of success. On the other hand, downstream approaches focus more on optimizing clinical research processes, on the hypothesis that operational issues such as patient recruitment, non-adherence, participants selection, site selection etc. are the key to success. Despite large investments and technological advancements, a significant number of drugs still fail at the clinical development stage.
[0004] The missing part on these approaches, is the emphasis on one silo of the pharmaceutical innovation and forgetting in the blind spot, an activity of great importance, which is the strategic pharmaceutical portfolio management.
[0005] In fact, strategic pharmaceutical portfolio management can be roughly defined as the process of maximizing the value of R&D pharmaceutical portfolio through the appropriate allocation of resources, requiring alignment of portfolio management activities with the company's strategic objectives. And this value maximization of R&D pharmaceutical portfolio concerns both therapeutic (scientific) value and commercial value of the product. Therefore, a better approach to risk management in drug development must therefore consider not only the optimization of the scientific value (i.e., pharmacological effects) of the molecule under development, but also better assess and optimize its commercial value. But we would also add the critical operational and strategic aspects, because if pharmacological effects (i.e., efficacy and safety) are well-known failure factors in drug development, the literature is replete with a multitude of other important failure risk factors of a strategic, operational, financial and commercial nature. The decisions associated with these risk factors are major contributors to failure in late stages and are difficult for humans to control in their entirety. For these reasons, it is imperative to implement means to better control the risks associated with clinical drug development, and to accelerate and optimize drug development by delivering actionableinsights that empower decision-makers to make informed, critical Go / No Go decisions with greater precision and confidence.
[0006] As previously mentioned, a multitude of different factors are associated with clinical drug development failures. Among these factors, some well-known are the focus of emerging approaches aimed at reducing risk of failures. However, those factors are only the tip of the iceberg. Beyond them lies a range of less understood variables, which represent other powerful silent drivers that subtly influence the success or failure of clinical drug development. It is widely acknowledged that the lack of efficacy is a major contributor of drug development failure, but there are many factors which can prevent potentially effective drugs from demonstrating efficacy such as underpowered clinical trial due to poor recruitment, high participants dropouts or even inappropriate study design, or unsuitable statistical endpoint measures (Fogel, 2018). Indeed, studies shown that to maintain the equivalent statistical power of clinical trials, it is necessary to increase patient recruitment by 50% to 200 % when patient non-adherence rates reach 30 % to 50 %, respectively (Alsumidaie, 2017). And recruiting additional patients can be very costly. For example, the average cost per patient in a Phase III trial is reported to be US$42,000 (Fogel, 2018). Further, authors have shown that factors such as trial participants' income levels, language and socio-cultural barriers can impact the recruitment rate and retention of participants in the trial (Huang et al., 2019). In fact, a survey has mentioned that difficulty to understand the informed consent form (ICF) is reported twice as often by patients who prematurely abandoned clinical trials than by those who completed them (Lopienski, 2015). Overly, the clinical trial designers must determine the appropriate patient’s eligibility criteria to ensure scientific validity of evidence in general patient population. However, restrictive inclusion criteria can significantly increase patient recruitment duration, leading to amend trial protocols. Indeed, a study have reported than 16 % of protocol amendments are due to changes in population criteria (K. A. Getz et al., 2011). (Consequently, protocol amendment can be costly, with the median direct cost amounting to $141,000 and $535,000 for a phase II and phase III protocol, respectively (K. A. Getz et al., 2016). Additionally, analyses have shown that the study research environment including the stringency of regulatory requirements and the level of oversight during study operations can present a hidden substantial challenge that contribute to recruitment delays (Briel et al., 2021). And recruitment delays represent a major operational challenge, with 80% of clinical trials unable to complete recruitment on time and 20% experiencing delays of at least six months (Nuttall, 2012). These delays can be extremely costly, as studies have shown that each day ofdelay in clinical development has a financial impact ranging from $600,000 to $8 million (Chaudhari et al., 2020) . Moreover, numerous studies have highlighted that, beyond the traditional model, collaborative and partnership-based approaches with academic and hospital research centers, along with the globalization of clinical trials to emerging regions, can enhance patient recruitment and reduce operational costs (Khanna, 2012). A recent estimates of clinical trials costs conducting in countries like China and India indicates that expenses can be up to 60% lower compared to other regions (Collier, 2009). In addition, traditional statistical approaches to estimating the probability of success or PTRS take into account macro factors such as clinical phase, condition, therapeutic area. However, several trends have been documented in the literature that support the relationship of therapeutic area with protocol complexity and increased investigative site work burden (K. A. Getz et al., 2008). In fact, in recent years, to better balance the risk-benefice ratio, scientists have added many procedures, aspects, and methods to aid in interpreting study outcomes, and guiding strategic decision making in trials. While these enhancements have been essential, in an environment of increasingly stringent regulatory requirements, and other third-party payers requests, they have also contributed to increased protocol complexity (K. Getz et al., 2023). Ultimately, the increasing complexity of protocols considerably reduces the productivity of clinical trials. According to an analysis of the clinical development pipeline, the success rate for oncology products was low, at just 8%, despite representing almost 30% of the products in pipeline. And the findings of the analysis indicated a strong correlation with protocol complexity (IQVIA, 2019). In other hand, the complexity of clinical trial protocols increases the financial burden significantly. A recent study reports that a 22% of study’s budget (approximately $1.3 million) is dedicated to procedures ensuring regulatory compliance, while 18% (approximately $1.1 million) is allocated to procedures for supplementary secondary, tertiary, and exploratory endpoints (ASPE, 2014)
[0007] Undoubtedly, it has become increasingly difficult for human to fully envision, control and anticipate this plethora of diverse and multidimensional risks. Moreover, these analyses highlight the complex interconnections and correlations between these factors, revealing the challenge of human control and the limitations of traditional methods in analyzing the macroscopic and microscopic dimensions of risk in clinical drug development. Artificial Intelligence offers capability to streamline drug development process by enabling intelligent control of the multitude of avoidable operational, scientific, strategic, financial and commercial risk factors. Specifically, machine learning offers the potential to overcome the majorchallenges facing the innovative pharmaceutical industry, and in this case, the plurality and multi-dimensionality of risk in clinical drug development and at the highest level, the optimization of the complex decision-making processes associated with strategic pharmaceutical portfolio management.
[0008] Artificial intelligence is having a profound impact in almost every field, but one of the key factors in its success is the availability of abundant, high-quality data to build machine learning models (Zha et al., 2023). This is particularly true and challenging in the pharmaceutical industry. Indeed, many authors have mentioned the difficulty of collecting and organizing historical data for developing predictive models to assess the probability of success in clinical trials. Some note that one of the biggest challenges in estimating the success rate of clinical trials is access to accurate data on clinical trial features and outcomes since collecting such data is high-priced, time-consuming, and susceptible to error (Wong et al., 2019). To build a high-quality and suitable database to properly train Machine learning models in this context, several challenges exist to be overcome.
[0009] One of these challenges is the requirement of a comprehensive list of risk factors impacting clinical trial success or failure. Ideally, this involves data collection and extraction through a meticulous, rigorous review of scientific literature and open-access databases, complemented by the inclusion of unpublished and private data to achieve a complete and holistic perspective. This list of factors will enable experts to annotate historical clinical trial data. Of course, this is a Herculean task, given the diversity and multidimensionality of the risks included, which are also reflected in the multidisciplinary nature of the experts involved (e.g., pharma portfolio managers, biostatisticians, clinical research design experts, market access experts, pharmacologists, regulatory affairs experts, clinical research operations, medical affairs, drug safety specialists, life sciences investments, value, and pricing, etc.).
[0010] Another major challenge lies in data acquisition, i.e. the ability to gather information from multiple data sources, structure it and normalize it in a single aggregated database. Indeed, public databases such as ClinicalTrials.gov allow users to consult the data for each clinical trial. However, using this single data source poses several problems. Although these data are historical and prospective, they provide information that is not linked between the different phases of clinical trial files. Further, the database includes both mandatory and optional aspects, and missing data make the quality of the raw information sometimes difficult to exploit as it stands. These limitations result in inconsistent data quality (Tse et al., 2018). More importantly, two keys’ aspects are essential for building a comprehensive clinical trial databaseto support predictive tools in pharmaceutical portfolio management. Firstly, collecting data from multiple sources to capture the full granularity of each clinical trial protocol and its results, and secondly, linking all historical clinical trials related to each drug candidate for a given indication. In fact, clinical trial data can be extracted separately for each study phase, which can be useful fortraining a predictive model for each trial phase outcome separately. However, to predict a phase transition, for example, it becomes essential to link clinical trial phases together, and thus to reconstruct the historical data pipeline for at least 10 years. To the best of our knowledge, there is currently no public database providing access to a large, high-quality historical data pipeline linking the phases of clinical trials together, enabling us to track the progress of a drug's development over time, from first-in-human studies to market launch. Building and accessing such a database opens the possibility of efficiently and accurately training predictive machine learning models to solve part of the complex problem of accurately selecting the best candidates with a high probability of success with confidence.
[0011] Another challenge is to develop not only a risk assessment and predictive analysis tool that will be integrated into such a unique database, but a predictive analytics tool with a prescriptive approach that overcome the “black box” nature of traditional machine learning approaches, providing an interpretation of risks and deploying them in an understandable way to guide strategic decision-making. Traditionally, pharmaceutical portfolio management has relied on the use of historical estimates of regulatory approval rates and human judgment to support drug development decisions. However, relying on human judgment and historical estimates results in obvious limitations and inaccurate results. These estimates are both quantitative and qualitative. For example, when they are quantitative, an obvious limitation is that the statistical estimates used to calculate the probability of technical and regulatory success (PTRS) consider only a few risk factors, such as therapeutic class, condition, stage of clinical development etc. As a result, this approach is unable to assess the effect of many other risk factors. It is also limited by its inability to provide decision-makers with a means of analyzing the impact of each major risk factor, predicting, anticipating, and controlling it, and thus dynamically optimizing the probability of success. When qualitative, estimates are based on the views of key opinion leaders. Of course, these opinions are subject to many obvious biases of human j udgment.
[0012] Recently, the literature revealed a trend towards developing approaches using machine learning to optimize the performance of clinical trials, and to predict important milestones in the clinical development of drugs, such as regulatory approval or clinical phase transitions.Indeed, a recent publication has developed machine learning algorithms to forecast five efficiency metrics including screen failure ratio, dropout ratio, pre-enrollment duration, enrollment duration and study duration using a total of 23 trial design features characterizing 2050 clinical trials conducting by pharmaceutical company Roche (Wu et al., 2022). Following a similar approach, another computing platform have been developed using at least one machine learning model trained on the different trial design parameters to generate and predict multiple scores of successes including overall success score, participant recruitment and retention rate, etc. (Vrouwenvelder et al., 2021). Focusing on clinical trial design, an inventor has developed an interactive platform to evaluate, compare and select an optimal trial design, using a simulation methods and machine learning to predict different trial performance aspects like; study duration, study budget, study complexity, time to market, probability of regulatory acceptance, etc. (BHATTACHARYA et al., 2021). According to widely recognized factors in the literature, drug efficacy and safety are key contributors to success or failure in clinical development. An applied data-driven approach has been explored to predict the likelihood of toxicity in clinical trials using 48 chemical properties and target-based features for a dataset that contains 828 drugs (Gay vert et al., 2016). Another inventor has advanced a virtual clinical trial platform to facilitate biomedical decision making and optimize treatment options in oncology based on patient-centric approach (SHRAGER et al., 2020). However, other published works have developed a machine learning model predicting a drug approval and clinical phase transition based on some drug and trial attributes (Lo et al., 2018). These pioneering publications demonstrate the feasibility of predicting the likelihood of success in clinical trials with the ability to consider many other risk factors other than pharmacological efficacy or safety as traditionally.
[0013] Overall, the scenarios proposed in these approaches do not consider a certain number of important elements such as the business perspectives of the various stakeholders (e.g., pharmaceutical companies, clinical research organization, Biotech, Investments funds, etc.) involved in the drug development process and the generality of applicability in various geographical contexts. In fact, consideration of the business perspectives, very often neglected in the approaches traditionally proposed, is crucial to proposing a workable solution. Numerous stakeholders are involved to varying degrees throughout the drug development process. Their involvement is very often a source of decisions. And so, a subject to be considered in the analysis of the internal and external decision-making environment. A practical solution needs to offer a pragmatic tool that doesn't just work in the laboratory, but also adapts perfectly to thebusiness realities of each stakeholder involved in the process. For example, a solution that works for a clinical research organization (CRO) is not sufficient for a pharmaceutical portfolio manager or regulatory affairs department. The reality of a small biotech company is not the same as that the large pharmaceutical company. While the term clinical trial success is used generically in the literature, it's important to note that the definition of success varies pragmatically, whether you're a CRO, a large pharmaceutical company, a life sciences investment fund, a biotech, a hedge fund, a patient, the public payers when relevant, etc. It's essential to take all these aspects into consideration. Moreover, from a broader perspective, the publications cited above demonstrate the possibility of making predictions of critical milestones in clinical drug development but applied to a subspace of potential risks. While the successful clinical development of a drug, from the laboratory through preclinical studies and human clinical trials to regulatory approval, remains a very important, if not the most critical, step, it is unfortunately not synonymous with commercial or financial success for the drug candidate. Consequently, defining and predicting the success of a clinical trial solely in terms of regulatory approval or phase transition is not sufficient to optimize strategic decisions in pharmaceutical portfolio management. This is why a more comprehensive consideration of the other dimensions of the Go / No-Go decision-making process is more reasonable. And the “Go / No-Go” decision-making process is at the heart of vital strategic pharmaceutical portfolio management activities. In the challenging dynamic environment, where pharmaceutical portfolio managers should take complex Go / No-Go decisions by considering multitude risk factors, it has therefore become essential for pharmaceutical portfolio managers to seek new and effective approaches to better decision-making.SUMMARY OF THE INVENTION
[0014] In the present invention, in some embodiments, methods, systems and platform are provided to end-users a priori predictive assessment of the likelihood of success of the drug development candidate at any phase of interventional clinical development and for any condition. For instance, the methods, systems, and platform can be configured to estimate a priori (before starting clinical trial) probability of success of a clinical trial candidate and a drug candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0015] Moreover, methods, systems and platform are provided to end-users a real-time predictive assessment of the likelihood of success of the drug development candidate at any phase of interventional clinical development and for any conditions. For instance, the methods, systems, and platform can be configured to estimate in real time (during execution of a clinicaltrial) probability of success of a clinical trial candidate and a drug candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0016] In one aspect of the present invention, a hierarchical computerized system for analyzing and predicting multiple success levels of a clinical trial candidate using machine learning models is provided. The system comprising a server comprising a data storage comprising data associated to risk factors impacting success of a clinical trial, historical clinical trials pipeline relating to a plurality of drugs for any conditions, clinical trials data related to a plurality of drugs candidates comprising a plurality of variables derived from the data associated to risk factors. The server further comprises a data processing engine configured to preprocess data from the data storage and a hierarchical machine learning engine comprising at least one first level model trained with the data of the data storage and configured to calculate an outcome of intermediate factors based on the said analyzed data and a second level model trained with the calculated predictions of the first level models, wherein the machine learning engine is configured to calculate predictions of multiple success levels of the one or more clinical trials.
[0017] The system may comprise more than two levels of successive models, each of the models using output of the previous level model as an input and may comprise six levels of trained models.
[0018] The first level model may estimate recruitment rate, protocol deviation, and patient dropout for a clinical trial candidate, the second level model estimating probability of operational success of a clinical trial candidate and interpretability of such prediction, the third level model estimating probability of scientific success of a clinical trial candidate and interpretability of such prediction, the fourth level model estimating probability of transition from one phase of clinical trial candidate to the next and interpretability of such prediction, the fifth level model estimating probability of regulatory success of a drug candidate and interpretability of such prediction, the sixth level model estimating a probability of financial or commercial success of a drug candidate.
[0019] The hierarchical machine learning engine may be further configured to calculate an outcome of intermediate factors. The intermediate factors may be any one of the followings protocol deviation, recruitment rate and patient dropout. The hierarchical machine learning engine may be further configured to compute explainability / interpretability of predictions of the multiple predicted success of a clinical trial candidate. The computing of the explamability / interpretability of predictions may further comprise computing globalexplainability and local explainability. The system may incorporate the global explainability tool to visualize the major contributions of the variables that the model has chosen to meet its target learning objective and local explainability tool to visualize the contributions of the system variables to a specific prediction. The system of claim 1 may comprise a decisionmaking system for the clinical drug development. The system may comprise a decision-making system for pharmaceutical portfolio management. The system may further comprise a decisionmaking system for early assessment of clinical drug development. The system may further comprise a decision-making system for early assessment of drug candidate in any clinical phase for any condition.
[0020] The first-level model may be configured to calculate a plurality of outcomes of intermediate factors or latent characteristics representing non-operational variables. The hierarchical machine learning engine may comprise a plurality of first level models, each of the first level models being configured to calculate an intermediate prediction of intermediate factor or latent characteristic. The plurality of the first level models may comprise at least one of the following models: a model to calculate prediction of recruitment rate in the clinical trial, a model to calculate a prediction of protocol deviation in the clinical trial and a model to calculate other intermediate factors impacting the success of the clinical trial overall. The intermediate predictions of intermediate factors of each of the first level models may be inputted in next level model to calculate the prediction of the success of the clinical trial.
[0021] The system may further comprise a module to interpret and explain the calculated prediction of multiple success levels of the clinical trial. The module to interpret and explain the calculated predictions of multiple success levels of the clinical trial may comprise generating logical rules used to calculate the prediction. The module to interpret and explain the calculated predictions of multiple success levels of the clinical trial may comprise any of the followings: global explainability, local explainability, comparative analysis with similar clinical trials and counterfactual scenario analysis.
[0022] The system may comprise a data acquisition module in data communication with a plurality of external data sources comprising data relating to clinical trials, regulatory approvals, research databases, economic and reimbursement information, pharmacological information, commercial information, corporate information, and expert knowledge. The data acquisition module being ma be configured to classify the data source in a plurality of repositories. The repositories may comprise any one of the following type of data: historical clinical trials protocols manually labeled with one or more risk factors of operational successor failure of clinical trials, list of risk factors of success or failure of regulatory approval in one or more jurisdictions, list of risk factors relates to commercial success, list of risk factors relates to a molecule in study / pipeline.
[0023] In another aspect of the invention, a computer-implemented method for predicting multiple success levels of a clinical trial candidate is provided. The method comprises storing data associated to risk factors impacting the success of a clinical trial, historical clinical trials pipeline relating to a plurality of drugs for any condition, and clinical trials data related to a plurality of drugs candidates comprising a plurality of variables derived from the data associated to risk factors, processing, transforming, and normalizing the acquired data using a natural language processor and executing at least one first level machine learning model trained with the processed, transformed, and normalized data to analyze data relating to the clinical trial and to calculate an outcome of intermediate factors based on the said analyzed data, executing at least one second level model trained with the calculated predictions of the first level models and using a machine learning engine to calculate predictions of the multiple success levels of the clinical trial.
[0024] The method may comprise training a plurality of first-level models, each of the first- level models calculating an intermediate prediction of an outcome of a specific latent characteristic of the clinical trial. Each of the plurality of first-level models may calculate one of the followings: a prediction of recruitment rate in the clinical trial, a prediction of protocol deviation in the clinical trial and other intermediate factors relating to the clinical trial.
[0025] The second-level model may use each of the intermediate predictions calculated by the first-level models to calculate the prediction of a probability of operational success of the clinical trial. The method may further comprise monitoring in real-time progress characteristics of the clinical study using such characteristics to calculate the prediction of the multiple success levels of the clinical trial. The characteristics may comprise recruitment rate, protocol deviation, and patient dropout. The execution of the machine learning model may further comprise calculating any one of the followings: probability of operational success of the clinical trial candidate, probability of scientific success of the clinical trial candidate, probability of transition from one phase of the clinical trial candidate to the next, probability of regulatory success of a drug of the clinical trial candidate, probability of financial or commercial success of a drug of the clinical trial candidate.
[0026] The execution of the machine learning model may further generate recommendations for optimizing study design and optimize study conduct to improve probabilities of the multiple success levels of the clinical trial. The method may further comprise developing a plurality of machine learning models for the clinical trial, training the developed models with acquired data and selecting one or more of the developed models based on performance metrics. The method may further comprise acquiring data from a plurality of external data sources comprising any one of the followings: data relating to clinical trials, regulatory approvals, research databases, economic and reimbursement information, pharmacological information, commercial information, corporate information and expert knowledge. A computer-readable medium storing instructions for executing the method as described above.
[0027] As noted above, a workable solution requires an appropriate and pragmatic definition of clinical drug development success that will consider the practical and technical perspectives of conducting clinical trials, the pharmaceutical portfolio management decision making processes, the business perspectives of stakeholders, as well as the challenges of developing predictive machine learning models. Because of this, in the context of the present invention, the success of the clinical trial will be understood as either an operational success of a clinical trial candidate, or a scientific success of a clinical trial candidate, or a success in moving from one phase of a clinical trial to the next, or a success in obtaining regulatory approval of a drug candidate, or, finally, a commercial or financial success of a drug candidate.
[0028] Thus, firstly, operational success of a clinical trial is understood as success in the execution of the clinical trial from study startup until last patient last visit milestone and closing the clinical trial database. Thus, operational success includes the successful execution of a clinical trial by avoiding operational risks such as low patient recruitment rates, delayed execution, inappropriate site selection, poor patient retention, poor site management, budget and contract negotiations, etc. For example, achieving operational success is a key business objective and priority for many contract research organizations, and even for the clinical research departments of large pharmaceutical companies. What's more, delays due to poorly controlled operational risks are catastrophic in terms of costs and time-to-market. In addition, recruitment and retention failures can considerably affect statistical power, and indirectly result further in an inability to prove efficacy. And so, for this reason, it would be valuable to provide a predictive tool to assess the risks of operational success in clinical trials, enabling these organizations to optimize their activities, better allocate resources, reduce costs, and increase productivity. For these reasons, the methods, systems and platform can be configured toestimate and provide to end-users such as Clinical Research Organizations or Clinical project manager a priori and a real time probability of operational success of a clinical trial candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0029] Secondly, scientific success is the achievement of a conclusive clinical trial result following analysis of the clinical trial data. This is the biostatistical and clinical aspects of clinical trials. However, scientific success is dependent on the operational success of the clinical trial, because as in the example above, recruitment and retention failures can considerably affect statistical power (b power). Statistical power is the probability that a trial's intervention effect will be detected if the effect is there. Thus, a drop in statistical power can result in the inability to detect an effect. Another example is a non-rigorous selection of participants who will be non-adherent in the clinical trial could result in the scientific failure of the clinical trial. As stated above, studies have shown that a 20% non-adherence rate among subjects requires a 50% increase in a sample size to maintain equivalent statistical power. For example, for a sponsor such as pharmaceutical company or biotech firm developing a molecule, achieving scientific success in way of gaining understand of the full scientific value of the drug candidate is crucial. For scientists in clinical research, it enables them to gather clinical evidence to support and convince regulatory authorities of a drug candidate's efficacy and safety. For decision-makers in pharmaceutical portfolio management or investment, to assess the scientific value of the pharmaceutical asset in the portfolio and analyze the investment potential. Then, for instance, methods, systems and a platform can be configured to estimate and provide to end-users such as teams of medical affairs or clinical research of a pharmaceutical industry or a biotech, a priori and a real time probability of scientific success of a clinical trial candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0030] Thirdly, once an operational and scientific success of a candidate clinical trial has been achieved, a decision must be made as to whether a drug candidate will proceed from one clinical phase to the next. For example, the transition from Phase II to Phase III. The decision to phase transition obviously depends as much on operational success as on scientific success of clinical trials, but not only that. Indeed, certain studies which complete an operational success, and a scientific success are often abandoned or (“deprioritized”), quite simply for strategic, commercial, or other considerations. This gives an idea of the major importance of integrating the above-mentioned aspects of pharmaceutical portfolio management into predictive models. Because of these reasons, considering all perspectives, methods, systems,and a platform can be configured to estimate and provide to end-users such as portfolio managers a priori and a real time probability of phase transition success of a drug candidate in a clinical trial, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0031] Fourthly, the clinical trial success most often cited in the literature, regulatory success, which of course is the achievement of regulatory approval by a marketing regulatory agency (e.g., FDA, Health Canada). Achieving regulatory success depends successively on operational and scientific success in all phases (I, II, III), and moreover on successful phase transition if this has not begun at phase III. In addition, most approaches in the literature rely on FDA approval to determine the regulatory success of a drug. Because of these reasons, considering all perspectives, a methods, systems, and a platform can be configured to estimate and provide to end-users such as regulatory affairs team, portfolio managers a priori and a real time probability of regulatory success of a drug candidate in a clinical trial, whether in Phase I, Phase II or Phase III, for any condition being studied and tailored to different regulatory agencies.
[0032] Finally, commercial, or financial success is the achievement of the expected return on investment. Marketing a drug is dependent on obtaining regulatory approval from the relevant regulatory agency (e.g., Health Canada in Canada), and therefore on regulatory success. However, regulatory success does not guarantee commercial or financial success since not all drugs that gain marketing authorization will be financially profitable taking individually. Thus, the development of a tool for portfolio manager's such as the probability of technical and regulatory success (PTRS), which estimates a chance of regulatory success, remains a limited tool as it does not consider crucial aspects of commercial success. Because of these reasons, considering all perspectives, methods, systems, and a platform can be configured to estimate and provide to end-users such as portfolio managers a priori and a real time probability of commercial or financial success of a drug candidate in a clinical trial, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0033] The present methods, systems, and a platform are developed for end-user’s organizations such as large pharmaceutical companies, biotech companies, clinical research organizations, life science investment funds, investment banks, hedge funds and consulting and pharmaceutical asset management firms. The methods, systems and a platform are applicable to the specific fields of clinical drug development planning, design, management and execution, and risk management, strategic pharmaceutical portfolio management, financialinvestments in life sciences, pharmaco-economics, market access and pricing strategies, stock market, due diligence assessment in the context of mergers and acquisitions (M&A).
[0034] As noted above, an effective and accurate solution requires access to a large and high- quality database that includes an exhaustive and a full granularity related to each clinical trial protocol and to each drug candidate from multiple data sources. Because of this, according to some embodiments, methods, systems, and platform are provided to process a unique aggregated, multimodal, structured database, methodically and manually built from a variety of heterogeneous, structured, and unstructured external data sources. This unique database contains information that comprehensively captures meticulously risk factors affecting the clinical drug development process, from early phase I studies through to market launch. Risk factors are linked, as indicated above, to our definition of clinical drug development success i.e., operational success of a clinical trial, scientific success of a clinical trial, success in moving from one phase of a clinical trial to the next, success in obtaining regulatory approval of a drug, and commercial success of a drug. Moreover, this unique database integrates clinical trials, whether in Phase I, Phase II or Phase III, and for any condition being studied. In some embodiments, a relational database system architecture is developed and contains more than 23,000 clinical trials protocol featured with among than 180 and more variables that are extracted from multisource and annotated manually by domain experts (e.g., drug development experts, pharmacologists, pharmaceutical portfolio manager, clinical research experts, medical affairs experts, regulatory affairs experts, pharmaco-economists, biostatisticians, etc.). These variables, which characterize the success or failure factors of clinical trials, are derived from an extensive review of the scientific literature (e.g., PubMed, Medline, Embase, Web of Science, etc.) over the last 20 years, an analysis of open databases and the collection of private data from companies such as contract research organizations, and finally a compilation of semistructured interviews conducted with pharmaceutical industry experts. The gathering, organization, integration, and preprocessing of these data is based on a proprietary standard operating procedure developed in the frame of a Data-Centric approach as illustrated in (FIG. 11). To the best of our knowledge, such an extensive review of risk factors of success or failure of clinical trial does not exist in the literature. Consequently, the more than 180 variables linked to each clinical trial protocol and each drug candidate are factors associated to each prediction objective, namely, in order, the operational success of a clinical trial, the scientific success of a clinical trial, the successful transition from one clinical trial phase to the next, the regulatory success of a drug candidate and, finally, the commercial or financial success of a drugcandidate. For instance, among others variables such as, therapeutic area, study endpoints, number of sites, study results, recruitment rate, statistical consideration, study design, type of drug in the study (biologic or small molecule), patent duration, primary purpose, number of patients enrolled, protocol deviation, biomarkers, dropouts, study allocation, masking, biological target, mechanism of action (MoA), outcomes measures, study duration, conditions, type of population in the study, sponsor type, metabolism pathway, inclusion and exclusion criteria, study environment, breakthrough therapy designation by FDA, fast track approval, orphan status, prevalence of disease etc. are featured in the relational database. We develop a comprehensive drug development risk network (DDRN) mapping which is in fact a data of relational map of potential risk factors in a clinical trial.
[0035] As mentioned above, developing an appropriate solution for strategic pharmaceutical portfolio management requires an extensive database that associates all clinical trials specific to each drug candidate for a given indication. In addition, the database which completely reconstructs the history of a drug candidate's journey from its phase 1 to market launch through regulatory approvals (e.g., FDA, Health Canada, etc.). Thus, database allow users to capture the evolution history of a drug candidate and its associated clinical trials as well as all successes and failures (i.e., operational success of a clinical trial, scientific success of a clinical trial, successful transition from one clinical trial phase to the next, the regulatory success of a drug candidate and the financial or commercial success of a drug candidate). In addition to the more than 180 variables annotated by experts and characterized each clinical trial separately, other relevant variables were added during the pipeline construction. These new variables represent a junction factor that allow to enrich the understanding of each drug candidate history throughout its clinical development process.
[0036] As mentioned above, an effective, accurate and workable Al-based solution relies on access to a large and high-quality database. Because of this, methods, systems, and platform adopt a data-centric Al approach (FIG. 11). By constructing a unique and a comprehensive database that is well curated by domain experts, diverse practical scenarios and all risk factors affecting clinical trials are taken into consideration. This approach ensures a representative dataset that not only enables effective generalization to unseen data but also the development of solutions tailored to pragmatic needs of diverse end-users. In addition, to address data inconsistency quality, a rigorous and validated annotation protocol is followed to collect a comprehensive set of multiple variables (e.g., 180 or more, etc.). For instance, this unique database integrates clinical trials dataset, whether in Phase I, Phase II or Phase III, and for anycondition being studied characterized by multiple variables related to, without limitation, operational clinical trials risk factors, regulatory approval risk factors, pharmacological risk factors, commercial success risk factors.
[0037] This new source of rich and unique data is the cornerstone for the training of an explainable machine learning models. For this reason, database is built considering that the prediction targets are the probability of operational success of a clinical trial candidate, or the probability of scientific success of a clinical trial candidate, or the probability of transition of one clinical trial candidate to the next, or the probability of regulatory success of drug candidate and or the probability of commercial or financial success of a drug candidate. For instance, in some embodiments, this new source of rich and unique database is used to train explainable machine learning model to generate for a clinical trial candidate whether in Phase I, Phase II or Phase III, and for any condition being studied an a priori and a real time probability of operational success.
[0038] However, as explained above, the prediction targets defined have a certain degree of dependency at varying levels. For example, while regulatory success of a drug (marketing approval) is no guarantee of its commercial success, commercial success is highly dependent on regulatory success (marketing approval). And scientific success of a clinical trial is also dependent on operational success of a clinical trial, and so on. Furthermore, at the level of granular considerations, certain risk factors affecting one prediction target may have causal relationships with other risk factors strongly affecting another prediction target. For example, as explained before, the risk of low participant retention or high dropout rates, or the risk of low recruitment rates strongly affect operational success of a clinical trial, but these factors are also known to affect the demonstration of scientific efficacy because of their impact on the B- power of the trial if left uncontrolled. Because of all these facts and considering the business feasibility for optimal use of the platform by the various stakeholders, it is appreciated to develop an explainable machine learning system based on a hierarchical model. Thus, in some embodiments, the unique relational database is used to train explainable machine learning system based on a hierarchical architecture.
[0039] Although end-to-end modeling is based on the objective of directly predicting a target such as the regulatory approval of a drug under development, as is frequently seen in approaches proposed in the literature, it nevertheless has limitations if applied to the scenarios proposed in the present invention. Indeed, in addition to the fact that the prediction targets (i.e. Operational success of a clinical trial candidate, scientific success of a clinical trial candidate,clinical phase transition success, regulatory success of a drug candidate, commercial or financial success of a drug candidate) are linked by dependencies at different levels, a fact of crucial importance for the development of an Al-based model is the existence of two types of variables among the risk factors of interest: non-operational variables and operational variables.
[0040] In fact, unlike operational variables, which are known at the beginning of the clinical trial, the values of non-operational variables can only be known at the end of the clinical trial. And this particularity is linked to the context in which a clinical trial is carried out. For example, the number of patients to be recruited (Target recruitment) is known at the activation of the trial, because it is provided in the trial protocol, whereas the actual total number of patients recruited (Actual recruitment) will be known at the end of the trial.
[0041] This problem poses a challenge for the end-to-end design of Al models that are workable in realistic contexts, even though they are theoretically valid. Indeed, an Al model for predicting regulatory approval of a drug that is trained with historical clinical trial data which includes variables among others such as Target recruitment and Actual recruitment in the model's input data will generate a predictive model that may theoretically perform well. However, if we wanted to make it workable in a real-world setting of a company, we'd run up against the problem of what value to enter for Actual recruitment as an input variable for a new clinical study seeking predicting regulatory approval. As a result, these models either become unusable in a real-life prospective context, or non-operational variables are omitted in model training, in which case the model is completely disconnected from industry reality. These scenarios are observed with approaches proposed in the literature, which although theoretically functional, are not truly practical in real-life context. For these reasons, a system based on a hierarchical Al model is needed, which on the one hand allows latent features (non-operational variables) to be incorporated into the modeling, and on the other hand provides end-users with successive predictions of targets by considering dependencies at different levels between them.
[0042] As cited above, many non-operational variables strongly affect the clinical drug development success and should be incorporated into the modeling process. Because of this, in the context of the present invention, each prediction target (i.e., Operational success of a clinical trial candidate, scientific success of a clinical trial candidate, clinical phase transition success, regulatory success of a drug candidate, commercial or financial success of a drug candidate) is broken down into sub-objectives, simplifying the tasks and enabling access to non-operational variables. As a result, in some embodiments, more than 180 operationalvariables included in the unique database are used to train independently a machine learning models to generate for a clinical trial candidate whether in Phase I, Phase II or Phase III, and for any condition being studied an a priori prediction of many non-operational (intermediate) variables at level 1 (e.g., participant recruitment, protocol deviation, participant dropout, etc.). For instance, these intermediate predictions will be aggregated as input to train a machine learning hierarchical model to generate for clinical trial candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied an a priori prediction of operational success at level 2.
[0043] In some embodiments, based on the new unique database including more than 180 operational variables and on the intermediate (non-operational) variables predicted at level 1, a machine learning model configurated on hierarchical architecture is trained to estimate and provide to end-users successively an a priori probability of operational success of a clinical trial candidate at level 2, followed by a probability of scientific success of a clinical trial candidate at level 3, followed by a probability of phase transition at level 4, followed by a probability of regulatory approval of a drug candidate at level 5, and finally at the level 6, the probability of commercial or financial success for a drug candidate, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0044] As mentioned above, to guide a strategic decision-making, not only a predictive analysis tools are requiring but also a prescriptive one that overcome the ‘’black box” nature of many machines learning approaches. In fact, interpretability and explainability are mentioned as crucial in the literature to enhance users’ confidence and demonstrate the relevance of the solution to be deployed, but the techniques and methods to meet these needs are not well defined. In addition, the trade-off between developing a high performing ML model and well interpreted predictions is known as a major challenge. In fact, the use of "black box" algorithms is gaining in popularity in many fields, including decision support. In the healthcare industry, and more specifically in the pharmaceutical industry, these algorithms can be a valuable aid to clinical trial managers responsible for managing clinical trials and performing early clinical drug development assessment. Nevertheless, in this context, the difficulty of understanding the algorithm's predictions greatly hinders the appropriation of the tool by decision-makers. Indeed, interpretability is mentioned as important in all the articles, but the techniques and methods for meeting this need are not presented. More precisely, the interpretability of a model refers to its ability to be understood and explained transparently, both in terms of the characteristics of the data considered, and the internal mechanisms used tomake predictions. Beter interpretability helps to reduce the "black box" character of models, and promotes informed decision-making, which is essential when it comes to predicting the success of a pharmaceutical product in the portfolio. The main challenge for the successful deployment of a decision support tool for clinical trial management lies in users' confidence in the tool. Because of these reasons, in the present invention, an approach based on ‘’Glass box” machine learning models are developed. As a result, the methods, systems, and platform are configurated to provide to end-users a priori and a real time explainable predictions of a probability of operational success of clinical trial candidate, a probability of scientific success of clinical trial candidate, a probability of phase transition, a probability of regulatory approval of a drug candidate, and a probability of commercial or financial success of a drug candidate whether in Phase I, Phase II or Phase III, and for any condition being studied. Moreover, such explainable machine learning approach offers both global and local explainability.
[0045] Firstly, the global explainability focuses on providing an overall understanding of the varying degree of influence that input variables can have on the ML model’s behavior. For instance, based on the entire unique database, among more than 180 variables on which the hierarchical machine learning models are trained, the methods, systems, and platform are adapted to provide a global explainability that maps the specific features (e.g., biomarkers, study design, statistical consideration etc.) significantly contribute on the model’s decision making process on a broader scale to generate as either a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate , a probability of phase transition, a probability of regulatory approval of a drug candidate, and a probability of commercial or financial success of a drug candidate whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0046] Secondly, the local explainability focuses on explaining why the hierarchical machine learning models made a specific decision for a particular clinical trial or for a specific drug. For instance, among more than 180 variables characterizing a specific clinical trial, the methods, systems, and platform are designed to provide a local explainability that maps the specific features (e.g., drug ’s mechanism of action (MoA), active comparator, interim analysis, multicentric study etc.) significantly contribute on the model’s decision-making process for individual cases to generate as either a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate , a probability of phase transition success, a probability of regulatory approval success of a drug candidate, and a probability of commercial or financial success of a drug candidate whether in Phase I, Phase IIor Phase III, and for any condition being studied. Then, the methods, systems, and platform are configured to aggregate the major features generated by local explainability into predefined groups of features (e.g., investigated drug characteristics, clinical trial plan information, clinical trial location characteristics, etc.) to make representation of results more understandable for users. More precisely, the impact of each group of features is calculated by aggregating the impact of major individual predictive variables also impacting the predicted success targets. For instance, the methods, systems, and platform are adapted to display a local explainability for a particular clinical trial that can map the investigated drug characteristics as a major group of features impacting the model’s decision-making process for a clinical trial candidate to predict a probability of scientific success. The impact of that group of features aggregates the impact of many individual predictive variables also impacting the model’s decision-making process including, without limitation, drugs mechanism of action, route of a drug administration.
[0047] Further, unlike global explainability, local explainability provides actionable insights tailored to a particular clinical trial and a specific drug. It identifies the variables that predominantly influence predictions and clarifies their impact whether positive or negative on the predicted outcomes, enabling more informed strategic decision-making. For instance, among more than 180 variables characterizing a specific clinical trial and a specific drug, the methods, systems, and platform are adapted to provide a local explainability that maps the specific variables configured to be aggregated into groups of features and that significantly contribute to the model’s decision-making process for individual cases and their positive or negative impact. Then, the methods, systems, and platform are adapted to allow end-users to open a particular group of features and identify the individual variables included. This enable users to determine how optimizing as either a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate, a probability of phase transition success, a probability of regulatory approval of a drug candidate, and a probability of commercial or financial success of a drug candidate whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0048] In some embodiments, in addition of local explainability, the methods, systems, and platform are configured to provide to end-users with a comparative analysis based on probability distribution of each predicted success target (i.e., a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate , a probability of phase transition, a probability of regulatory approval of a drugcandidate, and a probability of commercial or financial success of a drug candidate) compared against all the similar studies and drugs in the same therapeutic area available in the training database, guiding end-users toward maximum achievable value of probability for each target predictions of a clinical trial candidate and a drug candidate.
[0049] In some embodiments, in addition of prescriptive approaches described above, methods, systems, and platform are configurated enabling end-users to simulate counterfactual scenarios by adjusting individual features values and observing how these changes influence the target predictions. For instance, based on the knowledge of the major contributors, as well as their positive or negative impact on each target predictions, the end-users can adjust one or more of these variables to optimize and maximize the probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate , a probability of phase transition, a probability of regulatory approval of a drug candidate, and a probability of commercial or financial success of a drug candidate whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0050] Based on the business perspectives of each end-user, a large pharmaceutical compagnies have active portfolio management with multiple competing pharmaceutical products in different stage of drug development. These companies employ techniques like Discounted Cash Flow to manage their pharmaceutical portfolio. However, this valuation approach falls short in providing sufficient quantitative insights into the risks associated with a drug candidate. Because of this, in some embodiments, methods, systems, and platform is applied to address a strategic pharmaceutical portfolio management challenge. Indeed, this approach provides a comprehensive view within a single project included in a portfolio by providing insights into multiple levels in a hierarchical framework. At each level, the probability of success and the most impactful factors are provided. For instance, at level 2, the probability of operational success of a clinical trial candidate and the most impactful factors are represented. At level 3, the probability of scientific success of a clinical trial candidate and the most impactful factors are represented, and so on. In broadest terms, this approach allows end-users to select high-potential drug candidate, prioritize and optimize the portfolio, and perform risk assessment for a drug candidate through the clinical development process, whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0051] In some embodiments, for a clinical trial candidate whether in Phase I, Phase II or Phase III, and for any condition being studied, a methods, systems, and platform are applied to address a specific clinical trial design challenge with particular emphasis on identifying and mitigatingpotential risks and optimizing clinical trial protocol prior to trial initiation. According to some embodiments and based on a comprehensive operational parameters of study protocol, the explainable machine learning system is configurated on hierarchical models to provide an a priori non-operational variables at level 1, followed by the probability of operational success of a clinical trial candidate at level 2 and so on until the probability of commercial or financial success of a drug candidate in a clinical trial at level 6. Further, the system provides a prescriptive analysis of the major actionable features that influence these probabilities at each level. And the system generates also a comparative analysis with similar studies (e.g., studies in the same therapeutic area) using a visualization tool to illustrate the probabilities distribution across studies within the same therapeutic area. These two options serve as recommendations to guide end-users on the maximum achievable value for each target predictions and to identify the predictive drivers that need to be adjusted to optimize the clinical protocol design.
[0052] In some embodiments, the system is configurated to allow end-users to simulate counterfactual scenarios by modifying some variables and observing how each target predictions varies until the recommended maximum value for each target prediction is achieved. From a pragmatic perspective, each end-user has access only to those variables that can be modified in his practice. For instance, a clinical research organization (CRO) can adjust a limited set of features to optimize clinical trial operations. In contrast, sponsors, i.e., pharmaceutical companies, have broader flexibility to modify several aspects of clinical trial protocols (e.g., participants eligibility criteria, study endpoints, study location).
[0053] It is widely acknowledged that portfolio management should be a dynamic re- evaluation process. During the clinical development, several changes in information flow exist in addition of significant and constant uncertainty. Because of this, in some embodiments, within the scope of a single project in strategic pharmaceutical portfolio management, a methods, systems, and platform are applied to address a specific clinical trial oversight challenge with particular emphasis on identifying and mitigating potential risks and optimizing clinical trial productivity throughout its execution. Indeed, as a clinical trial progress, all intermediate variables (e.g., participant recruitment, protocol deviation, participant dropout etc.) can become known and their values evolve continuously throughout the execution of a clinical trial. According to some embodiments and based on a comprehensive set of study protocol variables and trial parameters updated in real-time, the explainable machine learning system is configurated on hierarchical models to directly provide a real-time probability of operational success of a clinical trial candidate at level 1, followed by a probability of scientificsuccess of a clinical trial candidate at level 2 and so on until the probability of a financial or commercial success of a drug candidate at level 5. Further, as described above, based on the knowledge of the key predictive drivers and the distribution of probabilities across similar studies provided by the system, the end-users can make well-informed decisions. Indeed, the insights provided by the system permit not only evaluating value of drug candidate or clinical trial candidate effectively but also assessing risk structure of ongoing clinical trials. For instance, a CRO can optimize clinical trial execution, for example by implementing strategies to enhance patient recruitment, improve patient retention, etc. while a sponsors, i.e., pharmaceutical companies, can use these actionable insights to maintain a balanced pharmaceutical portfolio in terms of high-risk and low-risk projects. Further, it serves to make well-informed decisions about the progression of clinical trials. This may involve accelerating promising studies, deprioritizing or even terminating the unpromising ones.
[0054] A variety of methods can be used to manage the pharmaceutical portfolio and to optimize investment decision. The Net Present Value (NPV) is the simple approach calculated by each team in the industry using historical data. It is used to prioritize and rank all active projects to make Go / No Go decisions, thereby maximizing the portfolio’s value. However, this method ignores risks and probabilities. That why other methods incorporating risks and probabilities are often used by Go / No Go decisions makers. For instance, the risk-adjusted Net Present Value (rNPV) accounts for the probabilities of phase transition success at each stage of clinical drug development, as well as regulatory approval success. Additionally, the probability of commercial success is also considered to calculate the Expected Commercial Value of projects (ECV). However, these methods are dependent on historical quantitative data for success probabilities, which may not accurately represent the current drug characteristics in development. Additionally, they can lack standardization across all industries and might fail to identify multiple interconnected risk factors associated with each clinical drug development at a micro level, such as trial design, trial locations. This can limit decision-makers’ ability to evaluate pharmaceutical asset not only based on their expected value but also on whether the associated risk factors are controllable or not. Because of these reasons, in the present invention, the methods, systems, and platform are configured to generate to end-users the probabilities of phase transition success, the probability of regulatory approval success and the probability of financial or commercial success outputted by explainable hierarchical machine learning models. These models generate a specific prediction based on input over 180 variables that characterize the clinical development candidate. In addition, these models are trained onstructured database comprising a plurality of clinical trials from various pharmaceutical industries, enabling them to generate generalized predictions for new data related to a clinical development process. Further, in some embodiments, according to the local explainability provided by the present invention and explained above, the methods, systems, and platform provide to end-users the major risk factors impacting the model’s decision making for each probability may be predicted for a clinical trial candidate and a drug candidate (i.e., probability of operational success of a clinical trial candidate, probability of scientific success of a clinical trial candidate, probability of clinical phase transition success, probability of regulatory approval of a drug candidate, and probability of commercial / financial success of a drug candidate). This enables end-users to structure risk factors, determine whether they are controllable, thereby maximizing value and balancing the pharmaceutical portfolio. Further, it facilitates the optimization the investment decisions.
[0055] Other and further objects and advantages of the present invention will be obvious upon an understanding of the illustrative embodiments about to be described or will be indicated in the appended claims, and various advantages not referred to herein will occur to one skilled in the art upon employment of the invention in practice.Brief Description of the Drawings
[0056] The above and other aspects, features and advantages of the invention will become more readily apparent from the following description, reference being made to the accompanying drawings in which:
[0057] FIG. 1 is a diagram of an embodiment of communication systems between database systems, artificial intelligence systems and interface system of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0058] FIG. 2A-2B-2C is a block diagram of the high-level architecture of an embodiment of a system to generate predictive and prescriptive information to support strategic decision making a priori and during the clinical drug development phases in accordance with the principles of the present invention.
[0059] FIG. 3A-3B is a workflow block diagram of the data sources of the system of FIG. 2A- 2B-2C
[0060] FIG. 4A-4B is an architecture block diagram of an embodiment of a data system and processing engine of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0061] FIG. 5A-5B is an architecture block diagram of the execution and deployment module of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0062] FIG. 6 is a is workflow block diagram of hierarchical machine learning model architecture of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0063] FIG.7 is a high-level illustration of an end-to-end machine learning model of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0064] FIG. 8 is a high-level illustration of a hierarchical machine learning model of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0065] FIG. 9A-9B is a block diagram of workflow linkage process of building a clinical trial pipeline database of a system for managing pharmaceutical portfolio based on artificial intelligence.
[0066] FIG. 10 is a high-level illustration of an embodiment of an in-depth drug development risk network (DDRN) understanding and mapping of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0067] FIG. 11 is a diagram of an embodiment of an iterative data centric approach (Data centric Al) of a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.
[0068] FIG 12. is a diagram of a trial success definition considering needs, expectations, and business perspectives of stakeholders using a system for managing pharmaceutical portfolio based on artificial intelligence in accordance with the principles of the present invention.Detailed Description of the Preferred Embodiment
[0069] A method, system, and platform method for analyzing the success of clinical trials, interpreting risk factors and for managing strategic pharmaceutical portfolio managementbased using explainable artificial intelligence will be described hereinafter. Although the invention is described in terms of specific illustrative embodiments, it is to be understood that the embodiments described herein are by way of example only and that the scope of the invention is not intended to be limited thereby.
[0070] The present invention generally relates to systems, methods and platforms for a priori predictive assessment, and continuous analysis of the evolution of clinical trial success, counterfactual analysis, and recommendations, as well as interpretability of risk factors associated with clinical drug development. The said systems, methods, and platforms are generally applicable to the specific areas of strategic pharmaceutical portfolio management, the planning, the design and the conduct or execution of clinical trials, and even the optimization of investment. Understandably, the described systems, methods and platforms of the present invention may be applied to other areas, especially in the field of pharmaceutical portfolio and clinical trials.
[0071] An embodiment of a method for analyzing success of clinical trials, interpreting risk factors, and managing pharmaceutical portfolio based on explainable machine learning is provided. The method comprises collecting data relating to risk factors impacting operational success of a clinical trial candidate, training a first set of machine learning models having hierarchical architecture using a training dataset of several clinical trials data related to a plurality of drugs candidate at different phases of approval for any condition being studied to take as input all available variables before starting the clinical trials, predicting a plurality of non-operational variable values using the first set of trained machine learning models, generating a prediction at level 1 of a multiple non-operational variable values not available before starting clinical trials in a real-word practice context, training a second set of machine learning models having hierarchical architecture using a training dataset based on the generated prediction at level 1, using the trained second set of machine learning models to successively estimate a plurality of probabilities. The probabilities may comprise probability of operational success at level 2, of scientific success at level 3, of phase transition success at level 4, of regulatory approval success at level 5 and of commercial or financial success at level 6. Understandably, the present estimation is not limited to the above-mentioned estimates.
[0072] The collection of data relating to risk factors may comprise collecting a comprehensive catalog of risk factors impacting operational success of a clinical trial candidate or scientific success of a clinical trial candidate or phase transition from one clinical trial phase to the next or regulatory success of a drug candidate or financial / commercial success of a drug candidate.The sources of catalog of risk factors may comprise but are not limited to an extensive review of the scientific literature, a private data from companies including, without limitation, contract research organizations; and a semi-structured interviews conducted with pharmaceutical industry experts including, without limitation, pharma portfolio managers, biostatisticians, clinical research design experts, market access experts, pharmacologists, regulatory affairs experts, clinical research operations experts, medical affairs experts, drug safety specialists, life sciences investments experts,
[0073] The method may further comprise storing data of a historical clinical trials pipeline of a plurality of drugs from one or more phases of approval. The phases may comprise phase I to the final regulatory approval agencies decision. The clinical trials may be aggregated based on similar characteristics, such as but not limited to the same drug under investigation, the same medical condition, the same population criteria and / or other granularities including, without limitation, route of drug administration. The data may be stored in any data source, such as but not limited to a relational database.
[0074] The method may further comprise storing a plurality of clinical trials data related to a plurality of drugs candidates. The data may be related to Phase I, Phase II or Phase III and may be related to any condition being studied. Each clinical trial may be featured by among 180 and more variables derived from the collected risk factors catalog, extracted from external multisource and / or annotated manually by domain experts including, without limitation, pharma portfolio managers, biostatisticians, clinical research design experts, market access experts, pharmacologists, regulatory affairs experts, clinical research operations experts, medical affairs experts, drug safety specialists, life sciences investments experts. The variables characterizing each clinical trial in the database, whether in Phase I, Phase II or Phase III, and for any condition being studied may comprise, but are not limited to at least one of a plurality of variables characterizing investigated drug, a plurality of variables characterizing medical condition, a plurality of variables characterizing clinical trial plan, a plurality of variables characterizing clinical trial participants, a plurality of variables characterizing clinical trial location, an outcome of a clinical trial (operational success, scientific success, clinical phase transition success) and / or an outcome of an investigated drug (regulatory approval success, financial or commercial success).
[0075] The training of the first set of machine learning models may further comprise using a training dataset of several clinical trials data related to a plurality of drugs candidate, whether in Phase I, Phase II or Phase III for any condition being studied. One or more of the machinelearning models may be trained to take as input all available variables before starting clinical trials in the real-word practice context (operational variables) and may be configured to predict at a first iteration (or level 1) a plurality of non-operational variable values not available before starting clinical trials in the real-word practice context (e.g., participant recruitment, protocol deviation, participant dropout, etc.).
[0076] The method may further comprise validating each of the trained machine learning models of the first set on validating dataset of several clinical trials data related to a plurality of drugs candidates. The drug candidates may be in Phase I, Phase II or Phase III for any condition being studied. The validated data set may be used to generate the prediction at level 1 of a multiple non-operational variable values not available before starting clinical trials in the real-word practice context (e.g., participant recruitment, protocol deviation, participant dropout, etc.).
[0077] The training of the second set of machine learning models may further use a training dataset derived from validating dataset at level 1 of several clinical trials data related to a plurality of drugs candidate. The drug candidate may be in Phase I, Phase II or Phase III for any condition being studied. Each of the machine learning model of the second set may be trained to take as input all available variables before starting clinical trials in the real-word practice context (operational variables) and the predicted values at level 1 of a multiple non- operational variable (e.g., participant recruitment, protocol deviation, participant dropout, etc.) to estimate and provide to end-users successively a probability of operational success of a clinical trial candidate at level 2, followed by a probability of scientific success of a clinical trial candidate at level 3, followed by a probability of phase transition success at level 4, followed by a probability of regulatory approval success of a drug candidate at level 5, and finally at the level 6, the probability of commercial or financial success of a drug candidate.
[0078] The method may further comprise validating each trained machine learning models of the second set on validating dataset derived from validating dataset specific to level 1 of several clinical trials data related to a plurality of drugs candidate. The drug candidate may be in Phase I, Phase II or Phase III for any condition being studied. The validated data may be used to provide to end-users successively a probability at other levels than level 1. In some embodiments, the level 2 is the probability of operational success followed by a probability of scientific success at level 3, followed by a probability of phase transition success at level 4, followed by a probability of regulatory approval success at level 5, and finally at the level 6, the probability of commercial or financial success.
[0079] The method may further comprise testing the second set of trained machine learning models on inference testing dataset of several new clinical trials data related to a plurality of drugs candidate. The tested models may be used to estimate and / or provide to end-users successively a multiple non-operational variable value. The variable values may comprise but are not limited to participant recruitment, protocol deviation, participant dropout, etc. In one embodiment, the estimation may be provided at level 1, followed by a probability of other levels. The estimation may comprise probability of operational success at level 2, followed by a probability of scientific success at level 3, followed by a probability of phase transition success at level 4, followed by a probability of regulatory approval success at level 5, and finally at the level 6, the probability of commercial or financial success.
[0080] The method may further comprise evaluating the level of performance of each trained machine learning models and selecting a trained machine learning model based on the evaluated performance level. The evaluation of the level of performance may be based on the prediction performance of each trained machine learning, one type of trained machine learning model for each target prediction. The target predictions may comprise but are not limited to participant recruitment, protocol deviation, participant dropout, probability of operational success, a probability of scientific success, a probability of phase transition success, a probability of regulatory approval success, and finally the probability of commercial or financial success. The models’ selection is based on comparing the performance metrics for each trained machine learning model specific to each target prediction comprising a Fl score, Area Under the Curve (AUC), accuracy, recall, precision metrics;
[0081] The method may further comprise generating to an end-user and via an interface, a global explainability that determines specific variables significantly contribute on the model’s decision-making process across entire dataset to generate as either a probability of operational success, a probability of scientific success, a probability of phase transition success, a probability of regulatory approval success, and a probability of commercial or financial success of a drug candidate in clinical trial, whether in Phase I, Phase II or Phase III, and for any condition being studied;
[0082] The method may further comprise identifying, via an interface, to an end-user, the major predictive variables influencing the model’s decision-making across entire dataset and for each target prediction. The method may further comprise receiving, via an interface, under than 40 variables represented as questions to be answered by the end-user.
[0083] The method may further comprise generating, through a predefined transformation rule, expanded derived variables representing a more than 180 variables characterizing a particular clinical trial candidate and a particular drug candidate. The generation of expanded derived variables may use variables as inputs to execute the trained machine learning models configured on hierarchical architecture and receiving as output a probability of operational success of a clinical trial candidate, a probability scientific success of a clinical trial candidate, a probability of successful transition from one clinical trial phase to the next, a probability of regulatory success of a drug candidate and a probability of financial or commercial success of a drug candidate to be launched on the market.
[0084] The method may further comprise identifying, such as via an interface, to an end-user, a predicted probability for each of a plurality of success targets including but not limited to a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate, a probability of successful transition from one clinical trial phase to the next, a probability of regulatory success of a drug candidate, and a probability of financial or commercial success of a drug candidate to be launched.
[0085] The method may further comprise identifying, such as via an interface, to an end-user, a comparative analysis comprising a probability distribution of each target prediction for a clinical trial candidate, compared against similar clinical trials stored in the database; and determining the maximum achievable value of probability for each target predictions of a clinical trial candidate.
[0086] The method may further comprise generating to an end-user and via an interface, a local explainability mapping that identifies a plurality of individual features that significantly contributes on the model’s decision making process to generate for a particular clinical trial candidate as either a probability of operational success, a probability of scientific success, a probability of successful transition from one clinical trial phase to the next; and for a particular drug candidate a probability of regulatory success and a probability of financial or commercial success.
[0087] The method may further comprise identifying, such as via an interface, to an end-user, the most important or major individual features generated by a local explainability and configured to be aggregated into groups of features. The identification of the said features may further comprise receiving local explainability data comprising individual major variables and their impact whether positive or negative on each target prediction, clustering the variables intopredefined groups and calculating the aggregated score of impact for each group and / or generating a grouped output, wherein each group represents features with related impact to the model’s decision-making process for each target prediction.
[0088] The method may further comprise receiving an interface action configured to receive one or more end-user selection of one or more groups of features to open in response to a local explainability and a comparative analysis not satisfying predicted probability of one or more target prediction.
[0089] The method may further comprise identifying, such as via an interface, to an end-user, the individual variables aggregated in selected groups of features associated with their impact to the model’s decision-making process.
[0090] The method may further comprise receiving an interface action configured to receive one or more end-user modification of one or more individual predictive variables to execute the trained machine learning models configurated on hierarchical architecture for the clinical trial candidate and for the drug candidate with the modified one or more variables, a predicted probability for each of a plurality of success targets will result.
[0091] The method may further comprise identifying, such as via an interface, to an end-user, a new predicted probability for each of a plurality of success targets of a drug candidate to be launched. The success targets may comprise a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate, a probability of successful transition from one clinical trial phase to the next, a probability of regulatory success of a drug candidate and a probability of financial or commercial success of the drug candidate to be launched.
[0092] The machine learning models may be trained to predict for any condition being studied, such as but not limited to a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate, a probability of clinical phase transition success, a probability of regulatory approval success of a drug candidate and / or a probability of commercial / financial success of a drug candidate.
[0093] The method may be based on explainable machine learning models configured within a hierarchical architecture to optimize the design of a clinical trial candidate before its initiation, enhance oversight of a clinical trial candidate during its execution, improve the clinical trial risk management and / or improve the strategic management of a pharmaceutical portfolio; and optimize investment decisions.
[0094] The method comprising estimating a priori explainable prediction of probabilities for a plurality of success targets related to a clinical trial candidate and to a drug candidate permitting an end-user to optimize the design of a clinical trial candidate before its initiation and improve its risk management.
[0095] The method may be performed in real to enhance oversight of a clinical trial candidate and improve its risk management. The method may comprise providing an explainable machine learning models configured within a hierarchical architecture, capable of any of generating probabilities of clinical phase transition success across all stages of development, from early phase I to regulatory submission, generating a probability of regulatory approval success, generating a probability of financial or commercial success. The probabilities may be generated from inputting specific variables related to a clinical drug development candidate, permitting an end-user preform at least one of estimating the risk adjusted Net Present Value (rNPV), estimating the Expected Commercial Value of project (ECV) and interpreting all impactful risk factors associated with each predicted probability of a plurality of success target. The plurality of success targe may comprise a probability of operational success of a clinical trial candidate, a probability of scientific success of a clinical trial candidate, a probability of clinical phase transition success, a probability of regulatory approval success of a drug candidate, and / or a probability of commercial / financial success of a drug candidate), permitting end-users to make well-informed decision in risk management of a pharmaceutical asset to improve the strategic pharmaceutical portfolio management and optimize investment decisions.
[0096] The method may further comprise inputting, into an explainable machine learning model configured within a hierarchical architecture, all predictive variables known before starting a clinical trial candidate to predict, a priori and at level 1, values of latent variables not known before starting a clinical trial candidate, further based on all known variables and predicted values at level 1. The explainable machine learning model may be configured within hierarchical architectures and trained to predict, in a priori, probabilities for a plurality of success targets including but not limited to a probability of operational success of a clinical trial candidate at level 2, a probability of scientific success of a clinical trial candidate at level 3, a probability of clinical phase transition success at level 4, a probability of regulatory approval success of a drug candidate at level 5 and a probability of commercial / financial success of a drug candidate at level 6. The prediction level 1 may comprise using any one of patient recruitment, protocol deviation, patient dropout, other latent variables.
[0097] The method may comprise inputting into the explainable machine learning model, all predictive variables known before starting a clinical trial candidate and on variables whose values become known as the trial progresses, to predict in a real-time the probabilities for a plurality of success targets, including, but not limited to a probability of operational success of a clinical trial candidate at level 1, a probability of scientific success of a clinical trial candidate at level 2, a probability of clinical phase transition success at level 3, a probability of regulatory approval success of a drug candidate at level 4 and a probability of commercial / fmancial success of a drug candidate at level 5.
[0098] The method may further comprise a user visualizing a representation of the prediction. The representation may comprise a global explainability generating ranked variables related to clinical trial candidate and to drug candidate impacting the model’s decision-making process to predict each of a plurality of success targets for all clinical trials and for all drugs stored in the database based on explainable machine learning hierarchical models.
[0099] The method may further comprise visualizing a global explainability generating ranked variables related to clinical trial candidate and to drug candidate impacting the model’s decision-making process including, without limitation, biomarkers, study design, statistical consideration, participant recruitment.
[0100] The representation may also comprise a comparative analysis generating a distribution of probabilities of a plurality of success targets for a clinical trial candidate and for a drug candidate and the maximum achievable value of each probability compared against all other similar clinical trials and drugs stored in the database. The method may further comprise providing a comparative analysis generating a probability distribution for clinical trials in the same therapeutic area.
[0101] The method may further comprise providing a local explainability generating major or most important features impacting the model’s decision-making process to predict each of a plurality of success targets for a clinical trial candidate and for a drug candidate based on explainable machine learning hierarchical models.
[0102] The representation may comprise ranked groups of features impacting the model’s decision-making process including, without limitation, investigated drug characteristics, medical condition characteristics, clinical trial plan information, clinical trial participants characteristics, clinical trial location characteristics.
[0103] The method may further comprise generating ranked groups of features impacting the model’s decision-making process. The impact of each group may be calculated by aggregating the impacts of individual features related to clinical trial candidate and to drug candidate, including, without limitation, the number of sites where the clinical trial will be conducted, the number of cities involved, and the clinical investigative site experience, all belonging to clinical trial location characteristics group.
[0104] The method may comprise permitting the end-user to select one or more groups of features and modify the values of one or more individual features related to clinical trial candidate and to drug candidate aggregated within the selected group.
[0105] The method may further comprise generating a counterfactual scenario analysis providing a new probability for a plurality of success targets related to a clinical trial candidate and to a drug candidate in response of modified values of one or more individual features related to clinical trial candidate and to drug candidate based on the explainable machine learning model configured within a hierarchical architecture.
[0001] Referring now to FIG. 1, a high-level representation of an embodiment of the communication systems 100 is shown. The communication systems 100 generally comprises database systems 102, e artificial intelligence systems 105 and a user interfacing system 110, such as a web-based platform. The database system 102 may comprise a knowledge management system 103 and a user clinical trial data catalog management system 104. The knowledge management system 103 is typically configured to store clinical trial data in a data source, such as but not limited to a relational database. The clinical trial data is fed an artificial intelligence system. The clinical trial data may originate from structured and / or unstructured (e.g. CSV, XML etc.) from data sources 101, typically external data source. The data may be collected, annotated by experts and stored in the database system 102. The database system may be any type of database such as PostgreSQL. The database system 102 generally comprises data visualization tools, such as but not limited to pgAdmin. The user clinical trial data catalog management system 104 further comprises a database system, such as a structured database (e.g. PostgreSQL) which is configured to store user clinical trial catalog. The catalog management system 104 may further comprise a module for creating annotation protocol logic, a network drive which help to backup automatically the clinical trial catalog database and incorporating data visualization tools. The user clinical trial catalog management system 104 is typically in data communication with the interfacing system 110, such as the web-based platform.
[0002] The artificial intelligence systems 105 is configured to process relational and structured database (i.e., register knowledge) to train machine learning models based on a hierarchical approach to generate predictive insights. The artificial intelligence systems 105 illustrate two levels of prediction of the hierarchical architecture. At a first level (also referred as level 1), the prediction of the model 106 is used to train multiple machine learning models which will predict the outcomes of latent characteristics (i.e., non-operational parameters). The models’ prediction a level 1 106 shows the intermediates models for predicting participant recruitment, protocol deviation etc. The second level of models 107 are trained using the output of the prediction of the level 1 106 as input and to predict operational success of a clinical trial candidate and generate explainability of prediction. In the present embodiment, at level 2, the system is configured to generate a probability of success of a clinical trial candidate, global and local explainability and allows what-if (contrafactual) analysis which could help to modify parameters of a clinical trial design to simulate scenarios while designing or to optimize probability of success or by taking informed decision.
[0003] The user-interfacing system 110 comprises interfaces allowing end-users 120 to interact with a platform or an application, such as through a network connection, such as the Internet or other LAN or WAN network. The end-users 120 may comprise individuals working in a specific field, such as, but not limited to, clinical trial design, pharmaceutical portfolio management, regulatory affairs, medical affairs, clinical trial operations management, life sciences venture capital, clinical trial etc. The user-interfacing system 110 may be configured to present data input form as questions about a clinical trial candidate and a drug candidate. Based on the inputted answers, the probability of a success of a clinical trial candidate and a drug candidate may be performed and presented to the user.
[0004] Referring now to FIG. 2A-2B-2C, a block diagram illustrates a high-level architectural overview of an embodiment of a decision support system model for analyzing the success of clinical trials, interpreting risk factors and for strategic pharmaceutical portfolio management based on explainable machine learning 200. The system 200 is generally implemented as an intelligent platform allowing users to perform a priori predictive assessment and continuous analysis of the evolution of clinical trial success, for counterfactual analysis and recommendations, as well as interpretability of risk factors associated with clinical drug development. The system 200 may be configured to generate both predictive and prescriptive information to support strategic decision-making during the clinical drug development phases.
[0005] The system 200 typically comprises a data acquisition module 210, a data system and processing module 230 and one or more applicative programs 250. The data acquisition module 210 generally comprises programs or application to fetch or obtain data records from one or more external data sources 211. The programs may use application programming interfaces (API), network protocols, data fdes to retrieve the data records or any known method to fetch data from an external source using a computer network. The data acquisition module 210 may be configured to connect with public domain databases 212, industrial partners or private databases 216, information from the Internet 220, government, or official databases 225 and / or knowledge from experts 228.
[0006] The data sources 211 may comprise open and non-open databases, structured and unstructured information. In some embodiments, public domain databases 212 may comprise clinical trials databases 213, such as but not limited to ClinicalTrials.gov, EudraCT, Health Canada’s Clinical Trials Database etc. As an example, ClinicalTrials.gov is an online database belonging to the United States government containing more than 400,000 clinical trials conducted worldwide. The EudraCT database (European Union Drug Regulating Authorities Clinical Trials Database) contains information on interventional clinical trials on medicines conducted in the European Union (EU), or the European Economic Area (EEA) which started after 1 May 2004. And Health Canada’s Clinical Trials Database is populated with information about each clinical trial after Health Canada issues the NOL (No objection letter).
[0007] The public domain databases 212 may also include databases from regulatory agencies 214, such as, but not limited to, the US Food and drug Administration (FDA), Health Canada’s Drug products, the European Medicines Agency (EMA), national registries of authorized drugs in the EU, etc. These databases generally contain information on medicines that have received marketing authorization from regulatory agency in their jurisdiction. In addition, other valuable information can be found, such as submission dossiers, summary reviews, monographs, approval letters, updates on post-approval follow-ups, etc.
[0008] The industrial partners private data sources 216 may comprise private databases 217, such as but not limited to clinical research organization (CRO) databases, pharmaceutical or biotech companies etc. The information from the Internet 220 may comprise research databases 221, such as MEDLINE™, PubMed™ etc., company drug pipeline 222, such as pharmaceutical company pipeline. The government or official databases 225 may comprise public reimbursement databases 226 such as but not limited to INESSS (Institut National d'Excellence en Sante et en Services Sociaux) and CADTH (Canadian Agency for Drugs &Technologies in Health) databases in Canada, and commercial databases 227 such as market research databases. The CADTH database is downloadable in CSV format and contains more than 1000 entries. The INESSS is the Health Technology Assessment agency responsible for recommending reimbursement decisions to the Ministry of Health in Quebec and contains more than 5000 entries, at the time of fding. The knowledge from experts’ source 228 may comprise, but is not limited to, Clinical Research Organization (CRO), biotech and pharmaceutical executives, pharmaceutical portfolio managers, biostatisticians, clinical research designers, market access experts, pharmacologists, regulatory affairs specialists, clinical research operations experts, medical affairs experts, drug safety specialists, life sciences investments partners, global access strategy, and pricing, etc., and others Key Opinion Leader and Go / No Go decision makers 229.
[0009] Still referring to FIG. 2A-2B-2C, the system 200 is generally configured to organize and process the data acquired from different external structured and unstructured data sources. The artificial intelligence program 250 is configured to perform a priori predictive assessment, and continuous analysis of the evolution of clinical trial success, counterfactual analysis, and recommendations, as well as interpretability of risk factors associated with clinical drug development. Hence, the system is performed for generating predictive but also prescriptive information to support strategic decision making during the clinical drug development phases. The system 200 is mainly based on explainable machine learning.
[0010] Still referring to FIG. 2A-2B-2C, the processing module 230 is configured to classify the data into a plurality of repositories 231, such as but not limited to clinical trials data 232, regulatory approval data 234, pharmacological data 236, commercial data 238. The module 230 may be configured to store risk factors associated with different factors. The risk factors may be stored as catalogs. The factors may comprise but are not limited to clinical trial operational risk factors 233, regulatory approval risk factors 235, pharmacological risk factors 237 and / or commercial success risk factors 239. The module 230 may be further configured to curate data. The curation process may be performed experts in the field, according to their fields of expertise. The curation process may combine data from several sources. The combined data may be annotated with identified risk factors, organized and stored in a data storage, such as a database. The system 200 may comprise a data processing system configured to process the data and structures the data in a data storage, such as relational database 240 used to train explainable machine learning models.
[0011] The system may comprise an internal data source, such as a database, comprising the structured annotated internal data 240. The internal data storage 240 may be a structured relational database comprising the data from collected and annotated databases and may comprise relationship allowing for the linking of scattered data collected from various structured and unstructured external databases. As an example, the structured relational database 240 was developed and contains more than 23,000 clinical trials protocol featured with among the 180 and more variables that are extracted from multisource with constitute the catalog of risk factors. The risk factor may comprise but are not limited to clinical trial operational risk factors 233, regulatory approval risk factors 235, pharmacological risk factors 237, commercial success risk factors 239. In such exemplary database, the 23,000 clinical trial protocols were manually annotated by domain experts, such as but not limited to drug development experts, pharmacologists, pharmaceutical portfolio manager, clinical research experts, medical affairs experts, regulatory affairs experts, pharmaco-economists, biostatisticians. The clinical trial protocols were further processed by the data processing system to generate the structured annotated internal data 240.
[0012] The catalog of risk factors generally characterizes the success or failure factors of clinical trials. In some embodiments, the catalogs or risk are derived from an extensive review of the scientific literature (e.g., PubMed, Embase, Web of Science, etc.) over the last 20 years, an analysis of open databases and the collection of private data from companies such as contract research organizations, and finally a compilation of semi-structured interviews conducted with pharmaceutical industry experts. The gathering, organization, integration, and preprocessing of these data is based on a proprietary standard operating procedure developed in the frame of a data-centric approach as illustrated in FIG. 11. Such an extensive review of risk factors of success or failure of clinical trial is believed not to exist in the literature. Consequently, the more than 180 variables of the present example are each linked to clinical trial protocol and each drug candidate are factors associated to each prediction objective, namely, preferably in order, the operational success of a clinical trial, the scientific success of a clinical trial, the successful transition from one clinical trial phase to the next, the regulatory success of a drug candidate and, finally, the commercial or financial success of a drug candidate. For instance, among others variables such as but not limited to, therapeutic area, study endpoints, number of sites, study results, recruitment rate, statistical consideration, study design, type of drug in the study (biologic or small molecule), patent duration, primary purpose, number of patients enrolled, protocol deviation, orphan status, biomarkers, dropouts, study allocation, masking,biological target, mechanism of action (MoA), first-in-class, outcomes measures, study duration, conditions, fast track approval type of population in the study, sponsor type, metabolism pathway, inclusion and exclusion criteria, study environment, breakthrough therapy designation by FDA, prevalence of disease etc. are featured in the relational database240.
[0013] Referring now to FIG. 2A-2B-2C, in some embodiments, domain experts manually annotate raw clinical research data 232 by labeling data with catalog operational clinical success risk factors 233 based on a predetermined annotation protocol which is believed to be rigorous. The domain experts may further manually annotate regulatory approval raw data 234 by labeling data with catalog regulatory risk factors catalog 235. The domain experts may further manually annotate pharmacological raw data 236 by labeling data with pharmacological risk factors 237 and commercial data 238 by labeling data with catalog of commercial success risk factors 239. Therefore, the internal data storage or database 240 may be a structured relational database comprising the data from collected and annotated by domain experts and may comprise relationship allowing for the linking of scattered data collected from various databases.
[0014] The processing module 230 may further comprise a data governance and integration subsystem 244 comprising modules managing data access 245, data management 246, security of data 247 and operations on data 248. The structured data is feed to a machine learning unit241. The structured data might also feed to a deep learning unit 243 and / or the natural language processing unit 242. A machine learning unit 241, a natural language processing unit 242, and / or a deep learning unit 243 are configured to be trained using the structured database 240. In some embodiments, the system 200 is configured to organize and structure at least some of the unstructured external data sources, such as using Natural Language Processing (NLP) techniques. The system further comprises an application module 250. The application module 250 is configured to execute instruction implementing the algorithms and to produce various analyses. In such an embodiment, a user of the system 200 will be able to obtain the following knowledge prediction of operational success of a clinical trial candidate, scientific success of a clinical trial candidate, transition success from a clinical trial candidate to the next, regulatory success of a drug candidate and / or commercial or financial success of a drug whether in Phase I, Phase II or Phase III, and for any condition being studied.
[0015] In such an embodiment, a user of the system 200 will be able to analyze the estimated predictions to optimize the probability of success. For instance, based on the knowledge of themajor contributors generated by system 200, as well as their positive or negative impact on each target predictions, the end-users may adjust one or more of these variables to optimize and maximize the probability of operational success.
[0016] In yet other embodiments, the system 200 may be configured to calculate and display to the users’ recommendations for the optimization of each probability of success. The system 200 may further be configured to calculate and display classifications of risk factors, recommendations for study design and / or estimation of recruitment rates. The calculated recommendation generally allows improved control of risk factors in the planning and execution of clinical trials through a priori and real-time monitoring of a clinical trial. The calculated recommendations may further provide efficient resources allocations. The resources allocations may comprise proactive classification of risk factors for each project by Al would allow for effective prioritization of clinical trials. The calculated recommendations may further provide proactive management of clinical trial conduct, such as but not limited to anticipation of recruitment needs, anticipation of impacting events such as adherence or drop-out. The calculated recommendations may further provide diligent analysis to evaluate investment opportunities, such as but not limited to Merger & Acquisition and Venture Capital Investments.
[0017] Still referring to FIG. 2A-2B-2C, the applications 250 are configured to use the results of the processed data to experiment, test, and fine tune the said data. The output of the said experimentations, testing and tuning are deployed to predict, recommend, optimize, analyze, discover and / or report based on the processed data and the artificial intelligence systems.
[0018] Referring to FIGS. 3 and 4, the system 230 may also be divided into a data processing engine 260 and a model engineering 270. The information collected in the data acquisition module 210 is sent to the data processing engine 260 for classifying, processing and / or standardizing the data. The data processing engine 260 may comprise a transformation module 261 configured to transform the data into an acceptable format, a normalization / standardization module 262, an imputation module 263 and an encoding module 264. The imputation module 263 may allow the treatment of missing data and incomplete data sets and the reduction of bias due to missing data. The encoding module 264 may allow the conversion of labelled data points into numerical variables for model training. Once the data is processed through one or more of the data processing modules 261, 262, 263 and / or 264, the said data is sent to data preprocessing pipeline 265. The data preprocessing pipeline 265 sends the processed data to a selection sampler 266. The selection sampler 266 is configured to dividethe data into a testing 267 and training 268 sets. The testing set 267 is configured to test different models and the training set 268 is configured to train machine learning models.
[0019] The model engineering 270 is configured to receive the training set 268 and / or the testing set 267 as a basis for developing a plurality of models 271 and to select and validate the developed models 275. The model development module 272 is configured to develop machine learning models such as, but not limited to, Random Forests (RF), Support Vector Machines (SVM), Gradient Boosted Trees (GBT), Logistic Regression (LR), and Neural Networks (NN), etc. The model development module 272 may further be configured to generate explainability and interpretability the developed models. Explainability enhances end-user confidence by offering insights that clarify model behavior, helping humans comprehend its internal logic. Unlike the classic black box model, where humans cannot understand the machine's decisions, explainable models (i. e. GlassBox™) offer the possibility for humans to understand the internal processes involved during model training or decision-making, enabling a clear understand of how an Al-based system has produced a particular prediction. The more explainable a model is, the more the decisions it makes are understood by humans.
[0020] The model selection and validation module 274 may be used to select the best model developed in the model development module 272. The model selection and validation module 274 may thus comprise selecting the best model 275 based on performance metrics, such as the accuracy, Fl -Score, precision, Area Under the Curve (AUC) and the recall performance metrics. The model selection and validation module 274 may further comprise generating explanations and interpretation 276 of the decision used in the said model. The model selection and validation module 274 may further comprise validating the generated explanations and interpretations 277 with domain experts. For instance, the validation may comprise examining sensitive parameters that affect the predictions significantly. The model selection and validation module 274 may also comprise finding counterfactual scenarios 278. In some embodiments, the finding of counterfactual scenarios may comprise Targeted Maximum Likelihood Estimation models (TMLE). The selected models are then executed and deployed in production.
[0021] Referring now to FIGS. 4 and 5, once the model is validated by the model selection and validation module 274, the validated model is sent to the application module 250. The application module may comprise execution 251 and deployment 255. The execution 251 comprises an experimentation component 252, a testing component 253 and a tuning component 254. The execution module 251 thus allows to experiment, test, and fine-tune themodel. The deployment module 255 may use the chosen trained and tested model to make decisions based on new data. The deployment module 255 may thus comprise a prediction component 256 which may predict operational success of a clinical trial candidate, scientific success of a clinical trial candidate, transition success from a clinical trial to the next, regulatory success of a drug candidate and commercial or financial success of a drug candidate.
[0022] The deployment module 255 may further comprise an analysis component 257, allowing the user to assess the impacting features, vary these features and tune them and assess the counterfactual success predictions. The deployment module 255 may also comprise a recommendation component 258, which may generate recommendations on which modifiable factor may be changed to influence the predictions generate in the prediction component 256. The deployment module 255 may further comprise an optimization component 259, allowing the users to have access to a set of diverse trained models for different phases of clinical trials while providing the users with explanations on models’ decisions. The deployment module 255 may further comprise a discover component 291, allowing the user to discover the importance of certain features and reveal the impact of features on the final decisions. The deployment module 255 may further comprise a report component 292 configured to display on a user interface a summary of the predictions, the predictive factors, the models and the counterfactual recommendations and explanations.
[0023] In another embodiment of the invention, the execution and deployment module are implemented as analytical, predictive, and prescriptive applications to support strategic decision making a priori and during the clinical drug development phases.
[0024] Referring to FIG. 6, a flowchart of an embodiment of an architecture of the hierarchical system 300 showing the detailed architecture of the various levels that make up the explainable machine learning system is presented. The hierarchical system 300 comprises a data processing pipeline 302, a system of model prediction at a first level 305 and a system of model prediction at a second level 318. The data preprocessing pipeline 302 is configured to process relational and structured data storage 301 to train machine learning models at the first and second levels 305 and 318 based on a hierarchical approach to generate predictive insights and interpretability of those predictions at the second level. Once data is processed, data processing pipeline 302 sends the processed data to a selection sampler. The selection sampler divides the data into training dataset 303, validation dataset 310 and testing dataset 304. The testing dataset 304 is configured to test different models developed at the first level 305 and at the second level 318. The validation dataset 310 is used to validate the trained models for latent risk factorsdeveloped at the first level 305. The training dataset 303 is used to train at the first level 305 multiple machine learning models to predict the outcomes of latent risk factors, such as non- operational parameters.
[0025] As illustrated, the system 200 is configured to develop, optimize and / or select best configuration 324 of prediction. The system of model predictions at the first level 305 may be further configured to generate intermediate models. The intermediate models may be used to predict patient recruitment 306, prediction of protocol deviation 307 and / or prediction of participant dropout 309. The predictions of the intermediate models are typically performed until all the other intermediate latent risk factors required have been trained into models 308.
[0026] The resulting intermediate predictions of the latent characteristics are inputted at into the machine learning second level models of the hierarchical system 200. The two datasets resulting of the splitting of the validation dataset 310 used at level 1 are inputted in the second level model 318. As such, the second level model 318 receives at least one training dataset 322, validation dataset 323 and testing dataset 304. The testing dataset 304 is configured to test different models developed at the first level models 305 and at the second level models 318. The validation dataset 323 is used to validate the trained models for the operation success of a clinical trial candidate 319 at the second level models 318. The second level models 318 are configured to generate explainability of prediction 321.
[0027] As illustrated, the system 200 is configured to develop, optimize and / or select best configuration for predicting operational success of clinical trial candidate. The system 200 is configured for the second level model 318 to generate probability of success of at least clinical trial candidate, global and local interpretability. The second level model 318 may further perform what-if (contrafactual) analysis. Such analysis maybe used to modify and or optimize parameters of a clinical trial design to simulate scenarios and increase probability of operational success or by taking informed Go / No-Go decision.
[0028] The system 200 may further be configured to display predictions of operational success of a clinical trial candidate and global and local explanations. The system 200 may be embodied as a user interface through web-based platform (SaaS).
[0029] The use of a hierarchical system 300 generally aims at breaking down complex decision-making process into a plurality of levels (1, 2, 3, 4 or more.). Furthermore, the use of a hierarchical model 300 discloses the specific impact of each of the level and the associated variables over the resulting prediction at each level (e.g., probability of operational success ofa clinical trial candidate at second level models 318) of the clinical trial. In such embodiment, each level of models represents a specific aspect or consideration in the decision-making process and one or more predictions are calculated for each of the level models. The hierarchical system approach generally aims at providing a systematic and comprehensive analysis. The predictions of each level are typically used as an input source for the next level. As such, the next level uses the result from the previous level to calculate another prediction, which is enhanced in relation to the intermediate predictions. As such, the system may break down the complex problem of predicting operational success of a clinical trial candidate for instance, into plurality of calculation of predictions for smaller problematics. The use of hierarchical model further allows improving overall performance of the engine.
[0030] In embodiments using a hierarchical model generally aims at unlocking or calculating predictions for specific situations or aspects of the global prediction by providing a more targeted and modular approach to problem-solving. By breaking down a complex problem into simpler steps, each intermediate model can focus on a specific aspect and solve sub-problems. Interpreting the results and understanding the information provided by these models or levels is crucial to comprehend which issue is being resolved and utilize the information provided by the models effectively. For example, hierarchical models may be used in predicting the likelihood of success in achieving the target number of participants for a phase II clinical trial candidate or assessing the likelihood of a phase III clinical trial candidate belonging to a high- risk for protocol deviation.
[0031] Each level of a hierarchical model generally requires a separate training, evaluation and testing process, as well as the management of intermediate inputs and outputs. For example, the hierarchical model may comprise a first level linked to the estimation of recruitment rate, the estimation of protocol deviation, and the estimation of a patient dropout for a clinical trial candidate, a second level linked to the estimation of probability of operational success of a clinical trial candidate and interpretability of such prediction, a third level linked to the estimation of probability of scientific success of a clinical trial candidate and interpretability of such prediction, a fourth level linked to the estimation of probability of transition from one phase of clinical trial candidate to the next and interpretability of such prediction, a fifth level linked to the estimation of probability of regulatory success of a drug candidate and interpretability of such prediction, a sixth level linked to the estimation of probability of financial or commercial success of a drug candidate. Understandably, any other number of levels may be contemplated within the scope of the present invention. Finally, in the illustratedworkflow, in the hierarchical system 300, the model is expected to evolve with the prediction of subsequent targets until financial or commercial success of drug candidate is attained.
[0032] Referring now to FIG. 7, a diagram of a high-level illustration of an embodiment of an end-to-end machine learning model in accordance with the principles of the present invention is presented. Referring now to FIG. 8, a diagram of a high-level illustration of an embodiment of a hierarchical machine learning model is presented.
[0033] Referring now to FIG. 9A-9B, a block diagram of an embodiment of a workflow linkage process for construction a clinical trial pipeline database is illustrated. The workflow linkage process 400 generally comprises a subsystem configured to perform clinical trial data analysis 410, a module for aggregation per phase based on a start date 420 and a historical clinical trials pipeline construction module 430, . The data analysis subsystem 410 is configured to analyze the characteristics of each clinical trial stored in the data storage. The analysis may be performed using a plurality of variables, such as but not limited to trial drug, medical condition, trial participants, phase of clinical development, and / or other granularities that may be used to aggregate the same clinical trials by phase. The data analysis subsystem 410 may further be configured to construct a sequential clinical trials aggregation per phase based on a start date 420. The historical clinical trials pipeline construction module 430 is generally configured to track the progress of drug development for any condition over time. As an example, the module 430 may track progress from phase I through to the decision of each regulatory agencies regarding new submissions.
[0034] Referring now to FIG. 10, a high-level illustration of an embodiment of an in-depth drug development risk network (DDRN) understanding and mapping is illustrated. It has always been understood in the literature review that the clinical drug development is associated with the plurality and multi-dimensionality of risk of various natures. Indeed, while pharmacological efficacy and innocuity are well-known cited failure factors of drug development, the scientific literature abounds with a multitude of other important failure risk factors of a strategic, operational, financial and commercial nature. The decisions associated with these risk factors are important contributors to failure in late stages and are difficult for humans to control. The comprehensive drug development risk network (DDRN) mapping is in fact a data of relational map of potential risk factors in a clinical trial. As an example, for just a few of the more than 180 variables used to annotate data, it shows how the said variables are linked by association or causality relationships, etc. The DDRN is useful for understanding therelationships between these variables, and supports, among other things, the construction of explicability approaches.
[0035] Referring now to FIG. 11, is a diagram of an embodiment of an iterative data centric approach (Data centric Al) is presented. The availability of abundant, high-quality data for building machine learning models is an essential factor in the successful development of artificial intelligence. This is particularly true and challenging in the pharmaceutical industry. Collecting and organizing data for the development of predictive models for assessing probability of success of clinical trials is time consuming, highly costly and a herculean task. Recently, the role of data in Al has grown considerably, giving rise to the emerging concept of data-centric Al. The focus of researchers has gradually shifted from the model design to improving the quality and quantity of data. In fact, in the conventional model-centric Al lifecycle, researchers focus primarily on identifying more efficient models to improve Al performance while keeping data largely unchanged (Zha et al., 2023). However, this modelcentric paradigm overlooks potential quality issues and undesirable data defects, such as missing values, incorrect labels and anomalies. Complementing existing efforts in model improvement, data-centric Al emphasizes the systematic engineering of data to build Al systems, shifting our attention from the model to the data (Zha et al., 2023).
[0036] The present invention relates generally to systems, method and platforms based on a data centric Al approach. In some embodiments the system requires a wide range of data, from clinical research, regulatory approvals, pharmacology, and commercial, as well as the knowhow of experts in the pharmaceutical industry. The present invention uses an approach to database collection and development that enables generation of a very large, high quality and unique database designed for the task of prediction of operational success of a clinical trial candidate, prediction of scientific success of a clinical trial candidate, prediction of the likelihood of transition from one phase of clinical trial to the next, prediction of regulatory success of a drug candidate and finally the prediction of financial or commercial success of a drug candidate.
[0037] Referring now to FIG 12, a diagram illustrating a pragmatical and functional clinical trial success definition considering needs, expectations, and business perspectives of stakeholders is presented. The operational success of a clinical trial is understood as success in the execution of the clinical trial from beginning of a study until last patient last visit milestone and closing the clinical trial database. The scientific success of a clinical trial is the achievement of a conclusive clinical trial result following analysis of the clinical trial data. The phasetransition success involves deciding on whether a drug candidate will proceed from one clinical phase to the next, taking into account strategic, commercial, or other relevant considerations. The regulatory success of a drug candidate is the achievement of approvals and authorization from one or more regulatory agencies (e.g., FDA, EMA, etc.). The financial or commercial success of a drug candidate is the achievement of the expected return on investment.
[0038] While illustrative and presently preferred embodiments of the invention have been described in detail hereinabove, it is to be understood that the inventive concepts may be otherwise variously embodied and employed and that the appended claims are intended to be construed to include such variations except insofar as limited by the prior art.
[0039] REFERENCES
[0040] Alsumidaie, M. (2017, April 24). Non-Adherence: A Direct Influence on Clinical Trial Duration and Cost. Applied Clinical Trials. https: / / www.appliedclinicaltrialsonline.com / view / non-adherence-direct-influence-chnical- trial-duration-and-cost
[0041] ASPE. (2014, July 24). Examination of Clinical Trial Costs and Barriers for Drug Development. ASPE. https: / / aspe.hhs.gov / reports / examination-clinical-trial-costs-barriers- drug-development-0
[0042] BHATTACHARYA, J., Bolognese, J., Buer, A., Edwards, E., Huang, S. Y., Jemiai, Y., Mehta, C., Patel, N., Pelz, A., Sathe, A. P., Schultz, J. A., & Senchaudhuri, P. (2021). Trial design platform (United States Patent US20210241859A1). https: / / patents.google.com / patent / US20210241859Al / en?oq=US2021241859Al
[0043] Briel, M., Eiger, B. S., McLennan, S., Schandelmaier, S., von Elm, E., & Satalkar, P. (2021). Exploring reasons for recruitment failure in clinical trials: A qualitative study with clinical trial stakeholders in Switzerland, Germany, and Canada. Trials, 22(1), 844. https: / / doi.org / 10.1186 / sl3063-021-05818-0
[0044] Chaudhari, N., Ravi, R., Gogtay, N. J., & Thatte, U. M. (2020). Recruitment and retention of the participants in clinical trials: Challenges and solutions. Perspectives in Clinical Research, 11(2), 64-69. https: / / doi.org / 10.4103 / picr.PICR_206_19
[0045] Collier, R. (2009). Rapidly rising clinical trial costs worry researchers. CMAJ:Canadian Medical Association Journal = Journal de I Association Medicale Canadienne, 180(3), 277-278. https: / / doi.org / 10.1503 / cmaj.082041
[0046] Fogel, D. B. (2018). Factors associated with clinical trials that fail and opportunities for improving the likelihood of success: A review. Contemporary Clinical Trials Communications, 11, 156-164. https: / / doi.Org / 10.1016 / j.conctc.2018.08.001
[0047] Gayvert, K. M., Madhukar, N. S., & Elemento, O. (2016). A Data-Driven Approach to Predicting Successes and Failures of Clinical Trials. Cell Chemical Biology, 23(10), 1294- 1301. https: / / doi.Org / 10.1016 / j.chembiol.2016.07.023
[0048] Getz, K. A., Stergiopoulos, S., Short, M., Surgeon, L., Krauss, R., Pretorius, S., Desmond, J., & Dunn, D. (2016). The Impact of Protocol Amendments on Clinical Trial Performance and Cost. Therapeutic Innovation & Regulatory Science, 50(A), 436-441. https: / / doi.org / 10.1177 / 2168479016632271
[0049] Getz, K. A., Wenger, J., Campo, R. A., Seguine, E. S., & Kaitin, K. I. (2008). Assessing the impact of protocol design changes on clinical trial performance. American Journal of Therapeutics , 75(5), 450-457. https: / / doi.org / 10.1097 / MJT.0b013e31816b9027
[0050] Getz, K. A., Zuckerman, R., Cropp, A. B., Hindle, A. L., Krauss, R., & Kaitin, K. I. (2011). Measuring the Incidence, Causes, and Repercussions of Protocol Amendments. Drug Information Journal, 45(3), 265-275. https: / / doi.org / 10.1177 / 009286151104500307
[0051] Getz, K., Smith, Z., & Kravet, M. (2023). Protocol Design and Performance Benchmarks by Phase and by Oncology and Rare Disease Subgroups. Therapeutic Innovation & Regulatory Science, 57(1), 49-56. https: / / doi.org / 10.1007 / s43441-022-00438-5
[0052] Huang, B., De Vore, D., Chirinos, C., Wolf, J., Low, D., Willard-Grace, R., Tsao, S., Garvey, C., Donesky, D., Su, G., & Thom, D. H. (2019). Strategies for recruitment and retention of underrepresented populations with chronic obstructive pulmonary disease for a clinical trial. BMC Medical Research Methodology, 19(f), 39. https: / / doi.org / 10.1186 / sl2874-019-0679-y
[0053] IQVIA. (2019). Global Oncology Trends 2019. https: / / www.iqvia.com / insights / the- iqvia-institute / reports-and-publications / reports / global-oncology-trends-2019
[0054] Khanna, I. (2012). Drug discovery in pharmaceutical industry: Productivity challenges and trends. Drug Discovery Today, 77(19-20), 1088-1102. https: / / doi.Org / 10.1016 / j.drudis.2012.05.007
[0055] Lo, A. W., Siah, K. W., & Wong, C. H. (2018). Machine Learningwith Statistical Imputation for Predicting Drug Approvals (SSRN Scholarly Paper 2973611). https: / / doi.org / 10.2139 / ssm.2973611
[0056] Lopienski, K. (2015). Retention In Clinical Trials - Keeping Patients On Protocols. https: / / www.meddeviceonline.com / doc / retention-clinical-trials-keeping-patients-protocols- 0001
[0057] Nuttall, A. (2012). Considerations For Improving Patient Recruitment Into Clinical Trials, https: / / www.clinicalleader.com / doc / considerations-for-improving-patient-0001
[0058] SHRAGER, J. C., Tenenbaum, J. M., PORTER, C. K., Hoos, W. A., & SHAPIRO, M. A. (2020). Platforms for conducting virtual trials (United States Patent US20200411199A1). https: / / patents.google.com / patent / US20200411199Al / en?oq=US2020411199A1
[0059] Tse, T., Fain, K. M., & Zarin, D. A. (2018). How to avoid common problems when using ClinicalTrials.gov in research: 10 issues to consider. The BMJ, 361, k!452. https: / / doi.org / 10.1136 / bmj.kl452
[0060] Vernon, J. A., Golec, J. H., & Dimasi, J. A. (2010). Drug development costs when financial risk is measured using the Fama-French three-factor model. Health Economics, 19(f), 1002-1005. https: / / doi.org / 10.1002 / hec.1538
[0061] Vrouwenvelder, A., Carraway, S. A., Hefferman, J., & Kenna, K. D. (2021). Using machine learning to facilitate design and implementation of a clinical trial with a high likelihood of success (United States Patent US20210357769A1). https: / / patents.google.com / patent / US20210357769Al / en?oq=US2021357769Al
[0062] Wong, C. H., Siah, K. W., & Lo, A. W. (2019). Estimation of clinical trial success rates and related parameters. Biostatistics, 20(2), 273-286. https : / / doi.org / 10.1093 / biostatistics / kxx069
[0063] Wu, K., Wu, E., D Andrea, M., Chitale, N., Lim, M., Dabrowski, M., Kantor, K., Rangi, H., Liu, R., Garmhausen, M., Pal, N., Harbron, C., Rizzo, S., Copping, R., & Zou, J. (2022). Machine Learning Prediction of Clinical Trial Operational Efficiency. The AAPS Journal, 24(3), 57. https: / / doi.org / 10.1208 / sl2248-022-00703-3
[0064] Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2023).Data-centric Artificial Intelligence: A Survey (arXiv:2303.10158). arXiv. https: / / doi.org / 10.48550 / arXiv.2303.10158
Claims
Claims1) A hierarchical computerized system for analyzing and predicting multiple success levels of a clinical trial candidate using machine learning models comprising: a server comprising: a data storage comprising: data associated to risk factors impacting success of a clinical trial; historical clinical trials pipeline relating to a plurality of drugs for any condition; clinical trials data related to a plurality of drugs candidates comprising a plurality of variables derived from the data associated to risk factors; a data processing engine configured to preprocess data from the data storage; and a hierarchical machine learning engine comprising: at least one first level model trained with the data of the data storage and configured to calculate an outcome of intermediate factors based on the said analyzed data; a second level model trained with the calculated predictions of the first level models; wherein the machine learning engine is configured to calculate predictions of multiple success levels of the one or more clinical trials.2) The system of claim 1 comprising more than two levels of successive models, each of the models using output of the previous level model as an input.3) The system of claim 2 comprising six levels of trained models.4) The system of claim 3, the first level model estimating recruitment rate, protocol deviation, and patient dropout for a clinical trial candidate, the second level model estimating probability of operational success of a clinical trial candidate and interpretability of such prediction, the third level model estimating probability of scientific success of a clinical trial candidate and interpretability of such prediction, the fourth level model estimatingprobability of transition from one phase of clinical trial candidate to the next and interpretability of such prediction, the fifth level model estimating probability of regulatory success of a drug candidate and interpretability of such prediction, the sixth level model estimating a probability of financial or commercial success of a drug candidate.5) The system of claim 1, the hierarchical machine learning engine being further configured to calculate an outcome of intermediate factors.6) The system of claim 5, the intermediate factors being any one of the followings protocol deviation, recruitment rate and patient dropout.7) The system of claim 1, the hierarchical machine learning engine being further configured to compute explainability / interpretability of predictions of the multiple predicted success of a clinical trial candidate.8) The system of claim 7, the computing of the explainability / interpretability of predictions further comprising computing global explainability and local explainability.9) The system of claim 8, incorporates the global explainability tool to visualize the major contributions of the variables that the model has chosen to meet its target learning objective and local explainability tool to visualize the contributions of the system variables to a specific prediction.10) The system of claim 1 further comprising a decision-making system for the clinical drug development.11) The system of claim 1 further comprising a decision-making system for pharmaceutical portfolio management.12) The system of claim 1 further comprising a decision-making system for early assessment of clinical drug development.13) The system of claim 1 further comprising a decision-making system for early assessment of drug candidate in any clinical phase for any condition.14) The system of claim 1, the first-level model being configured to calculate a plurality of outcomes of intermediate factors or latent characteristics representing non-operational variables.15) The system of claim 1, the hierarchical machine learning engine comprising a plurality of first level models, each of the first level models being configured to calculate anintermediate prediction of intermediate factor or latent characteristic.16) The system of claim 15, the plurality of the first level models comprising at least one of the following models: a model to calculate prediction of recruitment rate in the clinical trial; a model to calculate a prediction of protocol deviation in the clinical trial; and a model to calculate other intermediate factors impacting the success of the clinical trial overall.17) The system of claim 15, the intermediate predictions of intermediate factors of each of the first level models being inputted in next level model to calculate the prediction of the success of the clinical trial.18) The system of claim 1 further comprising a module to interpret and explain the calculated prediction of multiple success levels of the clinical trial.19) The system of claim 18, the module to interpret and explain the calculated predictions of multiple success levels of the clinical trial comprising generating logical rules used to calculate the prediction.20) The system of claim 19, the module to interpret and explain the calculated predictions of multiple success levels of the clinical trial comprising any of the followings: global explainability; local explainability; comparative analysis with similar clinical trials; and counterfactual scenario analysis.21) The system of claim 1 comprising a data acquisition module in data communication with a plurality of external data sources comprising data relating to clinical trials, regulatory approvals, research databases, economic and reimbursement information, pharmacological information, commercial information, corporate information, and expert knowledge.22) The system of claim 21, the data acquisition module being further configured to classify the data source in a plurality of repositories.23) The system of claim 22, the repositories comprising any one of the following type of data: historical clinical trials protocols manually labeled with one or more risk factors ofoperational success or failure of clinical trials, list of risk factors of success or failure of regulatory approval in one or more jurisdictions, list of risk factors relates to commercial success, list of risk factors relates to a molecule in study / pipeline.24) A computer-implemented method for predicting multiple success levels of a clinical trial candidate comprising: storing data associated to risk factors impacting the success of a clinical trial, historical clinical trials pipeline relating to a plurality of drugs for any condition, and clinical trials data related to a plurality of drugs candidates comprising a plurality of variables derived from the data associated to risk factors; processing, transforming, and normalizing the acquired data using a natural language processor; and executing at least one first level machine learning model trained with the processed, transformed, and normalized data to analyze data relating to the clinical trial and to calculate an outcome of intermediate factors based on the said analyzed data; executing at least one second level model trained with the calculated predictions of the first level models; using a machine learning engine to calculate predictions of the multiple success levels of the clinical trial.25) The method of claim 24 further comprising training a plurality of first-level models, each of the first-level models calculating an intermediate prediction of an outcome of a specific latent characteristic of the clinical trial.26) The method of claim 25, each of the plurality of first-level models calculating one of the followings: a prediction of recruitment rate in the clinical trial; a prediction of protocol deviation in the clinical trial; and other intermediate factors relating to the clinical trial.27) The method of claim 24, the second-level model using each of the intermediate predictions calculated by the first-level models to calculate the prediction of a probability of operational success of the clinical trial.28) The method of claim 24 further comprising monitoring in real-time progress characteristics of the clinical study using such characteristics to calculate the prediction of the multiple success levels of the clinical trial.29) The method of claim 28, the characteristics comprising recruitment rate, protocol deviation, and patient dropout.30) The method of claim 24, the execution of the machine learning model further calculating any one of the followings: probability of operational success of the clinical trial candidate, probability of scientific success of the clinical trial candidate, probability of transition from one phase of the clinical trial candidate to the next, probability of regulatory success of a drug of the clinical trial candidate, probability of financial or commercial success of a drug of the clinical trial candidate.31) The method of claim 30, the execution of the machine learning model further generating recommendations for optimizing study design and optimizing study conduct to improve probabilities of the multiple success levels of the clinical trial.32) The method of claim 24 further comprising developing a plurality of machine learning models for the clinical trial, training the developed models with acquired data and selecting one or more of the developed models based on performance metrics.33) The method of claim 24 further comprising acquiring data from a plurality of external data sources comprising any one of the followings: data relating to clinical trials, regulatory approvals, research databases, economic and reimbursement information, pharmacological information, commercial information, corporate information and expert knowledge.34) A computer-readable medium storing instructions for executing the method of claim 24.
Citation Information
Patent Citations
Systems and methods for designing clinical trials
US20200105380A1
System and interfaces for processing and interacting with clinical data
US20200321083A1
Using machine learning to facilitate design and implementation of a clinical trial with a high likelihood of success
US20210357769A1
Method and system for pharmaceutical portfolio strategic management decision support based on artificial intelligence
WO2023245301A1
Cited By
Systems and methods for interim clinical trial analysis
WO2026009178A1