System and method for quantitative clinical trial database modeling to provide information for project decisions
By extracting and processing information from multiple databases, generating structured datasets, and applying natural language processing and machine learning, the problem of unstructured data across database systems is solved, achieving the unification of clinical trial knowledge and accurate prediction of drug trial success probabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2026-04-10
AI Technical Summary
Clinical trial knowledge stored across multiple clinical trial database systems is often unstructured, and the data structuring methods differ between different systems, making it difficult to unify and effectively utilize the information.
The system extracts information from multiple databases using at least one processor, generates structured datasets, applies natural language processing techniques to process unstructured text, generates additional data, provides a graphical user interface to receive user analysis input, and executes machine learning modules to determine the probability of success and approval for drug trials.
It achieves standardization and unification of clinical trial knowledge across database systems, supports accurate prediction of the success probability and approval likelihood of drug trials, and improves data utilization efficiency and analysis reliability.
Smart Images

Figure CN121844387A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to European Patent Application No. 23306518.4, filed on 14 September 2023, and European Patent Application No. 24305220.6, filed on 9 February 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] Some of the embodiments disclosed herein relate to the field of medical informatics, and more specifically, to systems and methods that use medical informatics primarily for planning, implementing, and analyzing clinical trials and their outcomes. Background Technology
[0003] Clinical trial knowledge stored in or accessed by various conventional clinical trial database systems may differ in its structure and management. For example, clinical trial knowledge may include information about clinical trial outcomes, including safety and efficacy data based on results published in different journals, information about drugs evaluated during clinical trials, and information about drug outcomes evaluated in clinical trials. Clinical trial knowledge stored in or accessed by various conventional clinical trial database systems may also include information about comparisons of different clinical trials based on safety and efficacy data. However, clinical trial knowledge stored across multiple clinical trial database systems can often be unstructured or structured according to specific formats associated with the independent clinical trial database systems within those systems. Therefore, data and information representing clinical trial knowledge may be structured differently across different clinical trial database systems. Summary of the Invention
[0004] Embodiments of this disclosure include a method comprising: extracting or accessing information from one or more databases by at least one processor, the information including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; generating, by the at least one processor, multiple structured datasets comprising multiple data types based on the extracted information, the multiple structured datasets including trial summary datasets, trial group datasets, baseline datasets, outcome datasets, and outcome comparison datasets; applying natural language processing to at least a portion of the extracted or accessed information, including unstructured text, by the at least one processor; generating, by the at least one processor, additional data corresponding to at least one of the multiple data types based on the natural language processing for use in at least some of the multiple structured datasets; and providing a graphical user interface for receiving input from a user regarding an analysis to be performed using the multiple structured datasets.
[0005] The method may include a method in which the plurality of structured datasets further include a regulatory event dataset. The method may include a method in which at least a portion of the extracted information, including unstructured data, includes recent information not yet obtained from the one or more databases as structured clinical trial information. The method may include a method further comprising: extracting or accessing second information from at least one of the one or more databases by at least one processor, the second information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generating second data corresponding to at least one of the plurality of data types by the at least one processor based on the second information, to add to at least one dataset of the plurality of datasets or update at least some data in at least one dataset of the plurality of datasets; and applying natural language processing to at least a portion of the second information, including unstructured text, by the at least one processor, and generating second data or revised data corresponding to at least one of the plurality of data types based on the natural language processing for use in at least some of the plurality of structured datasets.
[0006] The method according to some embodiments may further include: periodically extracting or accessing supplementary information from at least one of the one or more databases by at least one processor, the supplementary information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generating supplementary data corresponding to at least one of the plurality of data types by the at least one processor based on the supplementary information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; and applying natural language processing to at least a portion of the supplementary information, including unstructured text, by the at least one processor, and generating supplementary data or revised data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets.
[0007] The method according to some embodiments may further include: the at least one processor executing a machine learning module to determine the probability of success in advancing a clinical drug trial to the next phase of clinical trials based on the plurality of structured datasets. The method may further include: the at least one processor executing a machine learning module to determine the probability of success in obtaining regulatory approval for the drug in the current clinical trial based on the plurality of structured datasets. The method may further include: the processor executing a machine learning module to determine the probability of success in obtaining regulatory approval for the drug in the current clinical trial, for each of a plurality of regulatory pathways, based on the plurality of structured datasets. The method may further include: the at least one processor identifying one or more scientific, design, regulatory, or operational factors that contribute to the historical approval probability of at least some of the plurality of clinical drug trials.
[0008] In some embodiments, the method may include a method wherein the one or more scientific factors include a mechanism of action associated with each of the one or more clinical drug trials, the one or more design factors include an endpoint associated with each of the one or more clinical trials, and the one or more regulatory factors include designating the one or more clinical drug trials as a breakthrough therapy. The method may further include: assigning a trial status category to each clinical trial in the plurality of structured datasets, wherein these trial status categories include one or more of the following: completed with a final result, terminated or ongoing but without a final result, and no result. The method may further include: having the at least one processor simulate the outcomes of one or more phase III clinical trials based on one or more proof-of-concept study outcomes. The method may include a method wherein the GUI includes: a first field for receiving a first input including disease-related information, the first field specifying variables, combinations of variables of interest, or both; and a second field for identifying the output to be generated. The method may include a method wherein the first input includes or identifies a disease-specific database or a subset of disease-specific databases, the disease-specific database being at least partially based on standardized clinical trials associated with the disease. The method may include a method in which the identified output to be generated includes one or more of the following: a description of one or more first clinical trials associated with the disease; a characteristic or outcome associated with the one or more first clinical trials; or a prediction of the outcome of the one or more first clinical trials for a drug used to treat the disease.
[0009] In some embodiments, the method may include a method wherein the trial dataset includes data types for associating trial outcomes with treatment information. The method may further include generating a predictive model or computer trial simulation for predicting trial outcomes. The method may include a method wherein the extracted or accessed trial outcome information includes endpoints associated with each of these clinical trials. The method may further include generating a predictive model or computer trial simulation for predicting these endpoints, wherein the predictive model or computer trial simulation is at least partially based on a machine learning model.
[0010] Embodiments of this disclosure include a system comprising: a storage device configured to store a unified clinical trial database; a memory storing one or more instructions; and at least one processor operatively coupled to the storage device, wherein the at least one processor is configured or programmed to read the one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: extracting or accessing information from one or more databases, including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; and generating, based on at least a portion of the extracted information, multiple structures comprising multiple data types. The database comprises multiple structured datasets, including trial summary datasets, trial group datasets, baseline datasets, outcome datasets, and outcome comparison datasets; at least one processor applies natural language processing to at least a portion of the extracted or accessed information, including unstructured text; the at least one processor generates additional data corresponding to at least one of the multiple data types based on the natural language processing for use in at least some of the multiple structured datasets; the multiple structured datasets are stored in the storage device, wherein the unified clinical trial database includes the multiple structured datasets; and a graphical user interface is provided to receive input from the user regarding the analysis to be performed using the unified clinical trial database.
[0011] Implementations of the system may include a system in which the plurality of structured datasets further include a regulatory event dataset. The system may further be a system in which at least a portion of the extracted information, including unstructured data, includes recent information not yet obtained from the unified clinical trial database as structured clinical trial information. The system may further be a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: extracting or accessing second information from at least one database in the unified clinical trial database, the second information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generating second data corresponding to at least one of the plurality of data types based on the second information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; and applying natural language processing to at least a portion of the second information, including unstructured text, and generating second data or revised data corresponding to at least one of the plurality of data types based on the natural language processing for use in at least some of the plurality of structured datasets.
[0012] The system may further include a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: periodically extracting or accessing supplementary information from at least one database in the clinical trial database, the supplementary information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generating supplementary data corresponding to at least one of the plurality of data types based on the supplementary information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; and applying natural language processing to at least a portion of the supplementary information, including unstructured text, and generating supplementary data or revised data corresponding to at least one of the plurality of data types based on the natural language processing for use in at least some of the plurality of structured datasets.
[0013] The system may further include a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: execute a machine learning module to determine the probability of success in advancing a clinical drug trial to the next phase of clinical trials based on the plurality of structured datasets. The system may further include a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: execute a machine learning module to determine the probability of success in obtaining regulatory approval for a drug in the current clinical trial based on the plurality of structured datasets, for each of a plurality of regulatory pathways.
[0014] The system may further be a system in which the at least one processor is configured or programmed to read the one or more instructions stored in the memory, thereby causing the at least one processor to: identify by the at least one processor one or more scientific factors, design factors, regulatory factors, or operational factors that contribute to the historical approval probability of at least some of the plurality of clinical drug trials. The system may further be a system in which the one or more scientific factors include a mechanism of action associated with each of the one or more clinical drug trials, the one or more design factors include an endpoint associated with each of the one or more clinical trials, and the one or more regulatory factors include designating the one or more clinical drug trials as a breakthrough therapy. The system may further include a system in which the at least one processor is configured or programmed to read the one or more instructions stored in the memory, thereby causing the at least one processor to: assign a trial status category to each clinical trial in the plurality of structured datasets, wherein these trial status categories include one or more of the following: completed with final results, terminated or ongoing but without final results, and no results.
[0015] Embodiments of the system may further include a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operation: simulating one or more phase III clinical trial outcomes based on one or more proof-of-concept study outcomes. The system may further be a system in which the GUI includes: a first field for receiving a first input including disease-related information, the first field specifying variables, combinations of variables of interest, or both; and a second field for identifying the output to be generated. The system may further be a system in which the first input includes or identifies a disease-specific database or a subset of disease-specific databases. The system may further be a system in which the identified output to be generated includes one or more of the following: a description of one or more first clinical trials associated with the disease; a characteristic or outcome associated with the one or more first clinical trials; or a prediction of the outcome of the one or more first clinical trials for a drug used to treat the disease. The system may further be a system in which the trial dataset includes data types that associate trial outcomes with treatment information. The system may further be a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operation: generating a predictive model or computer simulation for predicting the outcome of an experiment.
[0016] An embodiment of the system may further be a system in which the extracted or accessed trial outcome information includes an endpoint associated with each of these clinical trials. The system may include a system in which the at least one processor is configured or programmed to read one or more instructions stored in the memory, causing the at least one processor to: generate a predictive model or computer trial simulation for predicting progression-free survival (PFS) based on the overall response rate (ORR)—also referred to herein as the objective response rate (ORR). The system may further include a system in which the information also includes publications, conference proceedings, journal articles, and regulatory information associated with one or more clinical trials.
[0017] Embodiments of this disclosure include a non-transitory computer-readable medium comprising instructions that, when executed by a processing device, cause the processing device to: extract or access information from one or more databases by at least one processor, the information including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; generate, by the at least one processor, multiple structured datasets comprising multiple data types based on at least a portion of the extracted information, the multiple structured datasets including trial summary datasets, trial group datasets, baseline datasets, outcome datasets, and outcome comparison datasets; apply natural language processing to at least a portion of the extracted or accessed information, including unstructured text, by the at least one processor; generate, by the at least one processor, additional data corresponding to at least one of the multiple data types based on the natural language processing for use in at least some of the multiple structured datasets; store the multiple structured datasets in the storage device, wherein the unified clinical trial database includes the multiple structured datasets; and provide a graphical user interface for receiving input from a user regarding an analysis to be performed using the unified clinical trial database.
[0018] Embodiments of the non-transitory computer-readable medium include a medium in which the plurality of structured datasets further include a regulatory event dataset. The non-transitory computer-readable medium also includes a medium in which at least a portion of the extracted information, comprising unstructured data, includes recent information not yet obtained from the unified clinical trial database as structured clinical trial information. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to perform the following operations: extract or access second information from at least one database in the unified clinical trial database by at least one processor, the second information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generate second data corresponding to at least one of the plurality of data types by the at least one processor based on the second information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; and apply natural language processing to at least a portion of the second information, including unstructured text, by the at least one processor, and generate second data or revised data corresponding to at least one of the plurality of data types based on the natural language processing for use in at least some of the plurality of structured datasets.
[0019] Embodiments of the non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: periodically extract or access supplementary information from at least one database in the clinical trial database, the supplementary information including one or more of drug study design information, biomarker information, drug treatment information, and efficacy information for at least one clinical trial; and one or both of the following: generating supplementary data corresponding to at least one of the plurality of data types based on the supplementary information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; and applying natural language processing to at least a portion of the supplementary information, including unstructured text, and generating supplementary data or revised data corresponding to at least one of the plurality of data types based on the natural language processing for use in at least some of the plurality of structured datasets. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: execute a machine learning module to determine the probability of success of advancing a clinical drug trial to the next stage of clinical trials based on the plurality of structured datasets.
[0020] Embodiments of the non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: execute a machine learning module to determine, based on the plurality of structured datasets, the probability of obtaining regulatory approval for the drug in the current clinical trial. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: execute a machine learning module to, for each of the plurality of regulatory pathways, determine, based on the plurality of structured datasets, the probability of obtaining regulatory approval for the drug in the current clinical trial. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: identify, by the at least one processor, one or more scientific, design, regulatory, or operational factors contributing to the historical approval probability of at least some of the plurality of clinical drug trials.
[0021] Embodiments of the non-transitory computer-readable medium may include a medium wherein the one or more scientific factors include a mechanism of action associated with each of the one or more clinical drug trials, the one or more design factors include an endpoint associated with each of the one or more clinical trials, and the one or more regulatory factors include designating the one or more clinical drug trials as a breakthrough therapy. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: assign a trial status category to each clinical trial in the plurality of structured datasets, wherein these trial status categories include one or more of the following: completed with a final result, terminated or ongoing but without a final result, and no result. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: simulate the outcomes of one or more phase III clinical trials by the at least one processor based on one or more proof-of-concept study outcomes.
[0022] The non-transitory computer-readable medium may include a GUI comprising: a first field for receiving a first input including disease-related information, the first field specifying variables, combinations of variables of interest, or both; and a second field for identifying the output to be generated. The non-transitory computer-readable medium may include a medium in which the first input includes or identifies a disease-specific database or a subset of disease-specific databases. The non-transitory computer-readable medium may include a medium in which the identified output to be generated includes one or more of the following: a description of one or more first clinical trials associated with the disease; a characteristic or outcome associated with the one or more first clinical trials; or a prediction of the outcome of the one or more first clinical trials for a drug used to treat the disease. The non-transitory computer-readable medium may include a medium in which the trial dataset includes a data type that associates trial outcomes with treatment information. Embodiments of the non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to: generate a predictive model or computer trial simulation for predicting trial outcomes.
[0023] Embodiments of this non-transitory computer-readable medium may include a medium in which the extracted or accessed trial outcome information includes overall response rate (ORR) data and progression-free survival (PFS) data. The non-transitory computer-readable medium may further include instructions that, when executed by the processing device, cause the processing device to perform the following operations: generate a predictive model or computer trial simulation for predicting PFS based on ORR using a tree-based regression machine learning model or an associated computer trial simulation. The non-transitory computer-readable medium may also include a medium in which the information further includes publications, conference proceedings, journal articles, and regulatory information associated with one or more clinical trials.
[0024] Exemplary embodiments may include another tree-based regression model method for predicting median progression-free survival (mPFS) or median overall survival (mOS) of test subjects in response to drug treatment using overall response rate (ORR) data. Specifically, the tree-based regression model is a weighted regression using the sample size of the experimental group as weights. In some embodiments, the tree-based regression model depends on some or all of the following: treatment category data, treatment line count data, biomarker status information, and the subject's current cancer stage.
[0025] Some embodiments provide a method for determining a predictive model for median progression-free survival (mPFS) or median objective survival (mOS) in late-stage clinical trials, the predictive model being based on the overall response rate (ORR) in early-stage clinical trials. The method includes: extracting or accessing clinical trial information, including ORR information, mPFS information, mOS information, treatment category information, indication information, line-of-treatment information, disease stage information, and biomarker status information for multiple clinical trials obtained from one or more databases. The method further includes: defining variable categories that include at least some or all of the variables based on ORR information, treatment category information, indication information, line-of-treatment information, disease stage information, and biomarker status information. The method further includes: determining a tree-based regression machine learning model for ORR-based mPFS or ORR-based mOS, the determination including: performing a forward partitioning of the predictor variable space based on these variable categories, dividing it into multiple subspaces until no heterogeneity exists for ORR-based mPFS or until no heterogeneity exists for ORR-based mOS; and determining a regression model for each subspace.
[0026] In some embodiments, one or more of these treatment categories include anti-PD1 and / or anti-PLD1 therapy (PD1 / PDL1 therapy); and the biomarker status information includes PD1 and / or PDL1 status information.
[0027] In some embodiments, performing a forward partitioning of the predictor variable space into multiple subspaces until there is no heterogeneity in ORR-based mPFS or until there is no heterogeneity in ORR-based mOS includes: determining a slope for each variable class and grouping variable classes with homogeneous slopes to form grouped variable classes; and for each grouped variable class, splitting the grouped variable class if there is a significant heterogeneous slope among the variable classes in that grouped variable class.
[0028] In some embodiments, determining the slope for each variable category and grouping variable categories with homogeneous slopes to form grouped variable categories includes: determining a first variable category with the largest sample size compared to the sample sizes of other variable categories that are one or more second variable categories; determining the slope of each variable category by comparing: a first regression of the mPFS or mOS on the product of the ORR and the value of the category variable, and a second regression of the mPFS information or mOS information on the ORR information plus the value of the category variable. Determining the slope for each variable category and grouping variable categories with homogeneous slopes to form grouped variable categories also includes, for each of the one or more second variable categories: determining the slope difference between the slope of the second variable category and the slope of the first variable category; and grouping the variable categories if the p-value of the slope difference between the variable categories is higher than a specified grouping threshold or the sample size of the second variable category is lower than a specified sample size threshold. In some embodiments, the above steps are performed for each remaining ungrouped second variable category.
[0029] In some embodiments, splitting a grouping variable category for each grouping variable category, where there is a significant heterogeneous slope within that category, includes: for each variable with more than one grouped category, comparing the following: a first regression of the mPFS or mOS on the product of the ORR and the value of the grouping category variable, and a second regression of the mPFS or mOS information on the ORR information plus the value of the grouping category variable; and selecting the grouping variable category to be split if the ANOVA p-value is less than a specified splitting threshold and the ANOVA p-value is most significant for that variable compared to all other variables.
[0030] In some embodiments, the method further includes, for each split grouping variable category: determining a first variable category having the largest sample size in the split grouping variable category compared to the sample sizes of other variable categories in the split grouping variable category, the other variable categories being one or more second variable categories; and determining the slope of the variable category for each variable category in the split grouping variable category by comparing: a first regression of the mPFS or mOS on the product of the ORR and the value of the categorical variable, and a second regression of the mPFS information or mOS information on the ORR information plus the value of the categorical variable; and for each of the one or more second variable categories: determining the slope difference between the slope of the second variable category and the slope of the first variable category; and grouping the variable categories if the p-value of the slope difference between the variable categories is higher than a specified grouping p-threshold or the sample size of the second variable category is lower than the specified sample size threshold.
[0031] In some embodiments, determining a regression model for each subspace includes: for each variable, comparing: a first regression of the mPFS information or mOS information to the ORR information plus the value of the categorical variable, and a second regression of the mPFS information or mOS information to the ORR; and selecting variables whose p-value is less than a specified p-threshold.
[0032] In some embodiments, for at least some subspaces, the prediction model of mPFS for ORR is a log regression of the log of mPFS on the log of ORR; or for at least some subspaces, the prediction model of mOS for ORR is a log regression of the log of mOS on the log of ORR.
[0033] In some embodiments, the prediction model for at least some subspaces is further based on one or more of indication information and disease staging information.
[0034] In some embodiments, treatment category information includes chemotherapy, a combination of chemotherapy and PD1 / PDL1, or a combination of two or more of chemotherapy, PD1 / PDL1, or targeted therapy.
[0035] In some embodiments, the indication information includes one or both of non-small cell lung cancer data and melanoma data.
[0036] In some embodiments, the treatment line number information includes one or both of first-line (1L) treatment and above-first-line (1L+) treatment.
[0037] In some embodiments, the methods described herein are performed by a computing system, which includes at least one processor or one or more processors.
[0038] In some embodiments, the extracted or accessed clinical trial information is obtained from a quantitative clinical trial database. In some embodiments, this quantitative clinical trial database is generated by a method comprising the following steps: extracting or accessing information from one or more databases by at least one processor, the information including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; generating multiple structured datasets comprising multiple data types based on the extracted information by the at least one processor, the multiple structured datasets including trial summary datasets, trial group datasets, baseline datasets, outcome datasets, and outcome comparison datasets; applying natural language processing to at least a portion of the extracted or accessed information, including unstructured text, by the at least one processor; and generating additional data corresponding to at least one of the multiple data types based on the natural language processing, for use in at least some of the multiple structured datasets. Attached Figure Description
[0039] Figure 1 This is a flowchart of a process that can be performed according to some embodiments of this disclosure.
[0040] Figure 2 The logical information flow between various exemplary databases, modules, and computing environments according to some embodiments of this disclosure is schematically depicted.
[0041] Figure 3 It includes charts of different datasets that can be incorporated into a unified clinical trial database according to some embodiments disclosed herein.
[0042] Figure 4 The present disclosure is a logic diagram illustrating the process for generating a unified clinical trial database dataset based on some embodiments of this disclosure.
[0043] Figure 5 The present disclosure is a logic diagram illustrating the process for generating a unified clinical trial database using clinical trial group datasets.
[0044] Figure 6 The present disclosure is a logic diagram illustrating the process for generating a baseline dataset of clinical trials for a unified clinical trial database.
[0045] Figure 7 The present disclosure is based on a logic diagram of some embodiments, depicting the process for generating a dataset of clinical trial outcomes for a unified clinical trial database.
[0046] Figure 8The present disclosure is based on a logic diagram of some embodiments, depicting the process for generating a clinical trial outcome comparison dataset for a unified clinical trial database.
[0047] Figure 9 The present disclosure is a logic diagram illustrating the process of applying natural language processing to unstructured text stored in a database to extract clinical trial outcomes and clinical trial outcome comparison data, based on some embodiments of this disclosure.
[0048] Figure 10 An example computer system that can be used to perform some embodiments of this disclosure is schematically depicted.
[0049] Figure 11 The present disclosure is a logic diagram illustrating a process of accessing clinical trial data on a host device from a client device using an application programming interface, based on some embodiments thereof.
[0050] Figure 12 These are examples of clinical trial statuses at different organizational levels, based on some embodiments of this disclosure.
[0051] Figure 13 The present disclosure is a logic diagram illustrating the process of extracting clinical trial data from one or more external databases for use in a unified clinical trial database and using the extracted clinical trial data as input to a trial simulator.
[0052] Figure 14 The clinical trial simulator graphical user interface, based on some embodiments of this disclosure, depicts the regression of progression-free survival on overall response rate.
[0053] Figure 15 The present disclosure describes a clinical trial simulator graphical user interface that depicts the probability of a drug treatment used in a clinical trial progressing from one development stage to the next or to approval for the total number of progression-free survival events, and the probability of a drug treatment used in a clinical trial progressing from one development stage to the next or to approval for a given event size.
[0054] Figure 16 The flowcharts, based on some embodiments of this disclosure, illustrate a tree-based regression process for progression-free survival associated with patients in a clinical trial, based on the overall response rate of patients to drug treatment for a specific disease in the clinical trial, including graphs showing the analysis results.
[0055] Figure 17 It is a table that includes the distribution of predictor variables for experimental groups included in examples of some embodiments according to this disclosure.
[0056] Figure 18 It is a tree-based machine learning model that regresses the median progression-free survival (mPFS) on the total remission rate (ORR) in some examples of embodiments of this disclosure.
[0057] Figures 19A to 19E Examples of ORR regression to mPFS, including graphs with heterogeneity within tree-based machine learning models, are included in some embodiments of this disclosure.
[0058] Figure 19A This is a graph showing the heterogeneity of ORR regression on mPFS across treatment categories (all trial groups) in the example.
[0059] Figure 19B This is a graph showing the heterogeneity of ORR regression on mPFS in the tree-based machine learning model in terms of the number of inside lines (treatment category group 1) in the example.
[0060] Figure 19C This is a graph showing the heterogeneity of ORR regression on mPFS in the tree-based machine learning model (treatment category group 2).
[0061] Figure 19D This is a graph showing the heterogeneity of ORR regression on mPFS within the tree-based machine learning model for indications (treatment category grouping 1 1L / 1L+).
[0062] Figure 19E This is a graph showing the heterogeneity of ORR regression on mPFS within the tree-based machine learning model for indications (treatment category group 2 remaining lines).
[0063] Figure 20 This is a tree-based machine learning model that regresses ORR on median objective survival (mOS) according to some embodiments of this disclosure.
[0064] Figures 21A to 21C Examples of ORR regression to mOS, including graphs with heterogeneity within tree-based machine learning models, are included in some embodiments of this disclosure.
[0065] Figure 21A This is a graph showing the heterogeneity of ORR regression on mOS across all treatment categories in the tree-based machine learning model within the example.
[0066] Figure 21B This is a graph showing the heterogeneity of ORR regression on mOS within the treatment category group 2 indications in the example tree-based machine learning model. Figure 21CThis is a graph showing the heterogeneity of ORR regression on mOS in the treatment category grouping 2 NSCLC / melanoma within a tree-based machine learning model.
[0067] Figure 22 This includes a comparison of the mean squared error (MSE) of different regression models, such as ORR regressing mPFS and ORR regressing mOS, in the examples.
[0068] Figures 23A to 23D This includes survival probability estimates for phase II and III studies that used progression-free survival (PFS) and objective survival (OS) as endpoints, as shown in the examples. Figure 23A This is the success probability diagram for scenario 1. Figure 23B This is the success probability diagram for scenario 2. Figure 23C It is the success probability graph for scenario 3, and Figure 23D This is the success probability diagram for scenario 4. Detailed Implementation
[0069] Some embodiments of this disclosure include computer-implemented methods, systems, and nontransitory computer-readable storage media that provide standardized, quantified clinical trial data received or extracted from multiple database systems, associated with one or more drugs for treating diseases at various approval stages. The disclosure of some embodiments describes a process for standardizing and normalizing clinical trial data obtained from various database systems that store similar clinical trial data using different data types and formats. Some clinical trial data stored in certain conventional clinical trial database systems may be unstructured text, including relevant outcome results similar to structured clinical trial data stored in other database systems; however, due to the unstructured nature of the text, any meaningful information associated with this unstructured text cannot be easily standardized based on clinical trial data stored in other databases. Some embodiments disclosed herein provide a process that obtains clinical trial data associated with certain measurements or outcome quantities (also referred to as variables) from multiple different database systems, generates standardized and normalized data based on the obtained clinical trial data, and stores the standardized and normalized clinical trial data in a disease (indication)-specific standardized database.
[0070] Some embodiments disclosed herein provide methods for generating a unified clinical trial database, as well as systems that include or employ a unified clinical trial database. Some embodiments disclosed herein also provide methods for employing a unified clinical trial database.
[0071] Some of the embodiments disclosed herein provide users (e.g., researchers, clinical trial sponsors, physicians, investors, etc.) with the probability or likelihood that a particular drug used in a clinical trial will advance through various stages of the clinical trial, and to identify which factors (if any) significantly contribute to the likelihood that the particular drug will advance through certain stages of the clinical trial until approval.
[0072] Some conventional independent database systems can provide predictive analyses that generate probabilities or likelihoods of a drug's success in advancing to each phase of clinical trials. However, simply summing or combining these probabilities or likelihoods from multiple different independent conventional database systems only provides the sum or combination of independent probabilities or likelihoods that a drug will advance through the different phases of clinical trials and ultimately to approval. Furthermore, given that the metrics associated with the outcomes of clinical trials corresponding to a drug differ between different databases storing clinical trial data, the methods used to calculate probabilities and / or likelihoods may differ for different independent conventional database systems. As a result, any combination of independent probabilities or likelihoods may not provide a more accurate representation of the probability or likelihood that a drug will advance through the various phases of clinical trials and to approval than the independent probabilities or likelihoods themselves.
[0073] Some of the methods, systems, and nontransitory computer-readable media described herein are based on normalized, or standardized and structured clinical trial datasets, accurately providing the probability and likelihood of a drug advancing through each phase of a clinical trial and ultimately to approval. In addition, some of the described methods, systems, and nontransitory computer-readable media identify factors that significantly contribute to the probability and likelihood of a drug associated with a clinical trial advancing through each phase and to approval. In some embodiments, the described methods, systems, and nontransitory computer-readable media may identify different lines of treatment that contribute to the probability and likelihood of a drug associated with a clinical trial advancing through each phase and to approval.
[0074] In some embodiments, safety and efficacy information stored across a clinical trial database system can be used to determine or model the likelihood that a drug will advance through different phases of clinical trials and ultimately to approval. In some embodiments, this safety and efficacy information can also be used to identify certain factors that contribute to a drug advancing through different phases of clinical trials and to approval. This information is associated with each independent clinical trial, providing insights into the likelihood that a phenotyped drug will advance through different phases of clinical trials and, potentially, to approval for each clinical trial.
[0075] Therefore, standardizing safety and efficacy data obtained from different clinical trial databases increases the likelihood of identifying or modeling phenotypes of drugs that will advance through different stages of clinical trials and ultimately to approval, as well as the ability to identify certain factors that contribute to advancing through different stages of clinical trials and ultimately to approval.
[0076] In some embodiments, the described methods, systems, and nontransitory computer-readable media apply natural language processing (NLP) to unstructured text associated with clinical trials obtained from a database system to identify relevant outcome events associated with the outcome of a clinical trial corresponding to a drug administered during the clinical trial. In some embodiments, the methods, systems, and nontransitory computer-readable media detect portions of the unstructured text, such as clinical trial design information and outcomes associated with the clinical trial. For example, in some embodiments, the methods, systems, and nontransitory computer-readable media apply NLP to the unstructured text to detect information associated with a description of how the study will be conducted, including control group design and strategies for blinding and allocating subjects. In response to applying NLP to the unstructured text, in some embodiments, these methods, systems, and nontransitory computer-readable media automatically annotate sentences associated with clinical trial design information and corresponding outcomes associated with the clinical trial outcome, apply tags to sentences and phrases in the unstructured text, and apply pattern matching techniques to identify and extract data associated with the clinical trial outcome (such as treatment group, biomarker type, efficacy estimate, confidence interval, and p-value). In some embodiments, the extracted data are used by these methods, systems and nontransitory computer-readable media to determine the probability or likelihood of a drug associated with a clinical trial advancing through various approval stages, and any factors that significantly contribute to that probability or likelihood of success.
[0077] In some embodiments, the described methods, systems, and non-transitory computer-readable media provide a graphical user interface (GUI) including an application programming interface (API) that enables a user to access and interact with structured clinical trial data from one or more indication-specific unified clinical trial databases. In some embodiments, the GUI also provides the user with the ability to locate clinical trial data based on certain research area and project-specific needs. For example, the GUI may have an interface optimally designed for oncology research and another interface optimally designed for another research area. In some embodiments, the GUI receives input from the user, and in response, the system simulates the outcomes of a clinical trial based on clinical trial data from one or more unified clinical trial databases.
[0078] Figure 1An example method 100, which can be performed according to some embodiments of this disclosure, is described. Method 100 provides a user with access to the analysis of structured clinical trial data from one or more databases (e.g., multiple clinical trial databases) via a graphical user interface. One or more steps of the method may be performed by one or more computing devices or computing systems (e.g., Figure 2 The computing device 205 or Figure 10 The method is executed by computing system 1000. At 101, the method retrieves or accesses clinical trial information associated with one or more clinical trials from one or more databases (e.g., multiple clinical trial databases). For example, an account hosted on computing device 205 can request an access token by submitting a client ID and password to an identity service, and can receive an access token from the identity service confirming the account's identity. The computing device can then use the issued access token to retrieve or access information associated with the clinical trial stored or hosted on clinical trial databases 201 and 202. The following is combined with... Figure 11 It further describes gaining access to one or more clinical trial databases.
[0079] The obtained clinical trial information may include drug study design information, biomarker information, drug treatment information, and efficacy information associated with one or more clinical trials from one or more databases (e.g., clinical trial databases 201 and 202). The drug study design information may include, but is not limited to, a description of how studies have been, are being, or will be conducted, including one or more of the control group design and strategies for blinding and allocating subjects. The biomarker information may include, but is not limited to, measurable biological parameters that can be assessed throughout the clinical trial as indicators of normal biological, pathological, and / or pharmacological responses to certain drugs. Efficacy information may include, but is not limited to, efficacy endpoints. For some disease areas, such as cancer, in some embodiments, the efficacy endpoint may include overall response rate (ORR), progression-free survival (PFS), and overall survival (OS).
[0080] After extracting or accessing relevant information about clinical trials from clinical trial databases 201 and 202, the method can then generate datasets associated with clinical trial summaries, clinical trial groups, and at least some or all of the baselines corresponding to the clinical trials, clinical trial outcomes, and outcome comparisons between clinical trials. For example, according to some embodiments, computing device 205 can generate multiple structured datasets, including trial summary dataset 203a, trial group dataset 203b, etc. Figure 2 ) and baseline dataset 203c ( Figure 2 ), Ending Dataset 203d ( Figure 2 Comparison of datasets 203e and 203e (and their outcomes)Figure 2 At least some or all of the following. Some clinical trials may not have baseline information available for generating baseline dataset 203c. Some clinical trials may not have outcome information available for generating outcome dataset 203d. Some clinical trials may not have outcome comparison information available for generating outcome comparison dataset 203e. Where regulatory information is available, according to some embodiments, computing device 205 may also generate regulatory information dataset 203f (where regulatory information is available). Figure 2 ).
[0081] The trial summary dataset 203a may include study-level information extracted from clinical trial databases 201 and / or 202. In some embodiments, this information may include the original dataset associated with the study itself. For example, in some embodiments, the original dataset may include one or more of the following: a New Clinical Trial Identifier (NCT ID) associated with the drug in the clinical trial, the industry associated with drug development, the number of groups associated with the clinical trial, the start year of the clinical trial, and the completion year of the clinical trial if the clinical trial is completed. In some embodiments, the original dataset may also include information associated with the primary completion year, which is the date on which the last subject was examined or intervened for the final collection of data for the primary outcome, regardless of whether the clinical study ended or was terminated according to a pre-specified protocol. In some embodiments, the original dataset may also include the sponsor name and / or sponsor type (e.g., academic institution, industry, government agency). In some embodiments, the trial summary dataset 203 may also include information associated with the drug administered throughout the clinical trial or associated with the clinical trial.
[0082] The trial dataset 203b may include treatment information associated with the clinical trial. The baseline dataset 203c may include any or all of the original datasets from the clinical trial databases 201 and / or 202, which include results of a subject group in the clinical trial, baseline measurements associated with subjects in each group of the clinical trial, and baseline counts associated with the clinical trial.
[0083] The outcome dataset 203d may include measurements associated with clinical trial outcomes, counts associated with clinical trial outcomes, information about the various groups of participants in the clinical trial, and any or all of the drug information associated with the clinical trial outcomes. This information may include any or all of the drug event information, drug phase, and drug approval information.
[0084] Outcome comparison dataset 203e may include analytical data comparing outcomes between clinical trials.
[0085] In some embodiments, the plurality of structured datasets also includes a safety dataset. This outcome comparison dataset may include inter-treatment comparisons within the study regarding efficacy or safety endpoints. Because pairwise comparisons can be made between different treatments, for different clinical endpoints, and for different analysis populations, the relevant comparisons need to be standardized for potential pooled analyses. Clinical safety measures include, but are not limited to, adverse events and safety laboratory values, which can be pooled together for analysis.
[0086] Method 100 then proceeds to step 102 and determines whether the extracted or accessed clinical trial information includes unstructured text. If the method determines that the extracted or accessed clinical trial information does indeed include unstructured text (yes), then method 100 proceeds to step 104 and applies natural language processing to the unstructured text.
[0087] At step 104, method 100 applies a natural language processing module, such as natural language processing module 204, to unstructured text associated with the clinical trial information stored or hosted in clinical trial databases 201 and 202. According to some embodiments, the output of natural language processing module 204 may include structured data based on the unstructured text extracted or accessed from clinical trial databases 201 and 202. Method 100 proceeds to step 105, in which structured data corresponding to at least one of a plurality of data types is generated based on the output of the natural language processing of the unstructured text for at least some of the plurality of structured datasets. The following section discusses… Figure 9 This includes further descriptions of examples of the natural language processing module.
[0088] Generate structured data corresponding to at least one of the plurality of data types, and method 100 provides a graphical user interface (GUI) to a user, which receives input from the user regarding the analysis to be performed using the plurality of structured datasets (step 110).
[0089] Returning to step 102, if method 100 determines that the extracted or accessed clinical trial information does not include unstructured text (no), then at step 103, method 100 generates structured data corresponding to at least one of the plurality of data types based on the extracted or obtained clinical trial information, for at least some of the plurality of structured datasets.
[0090] In some embodiments, at 106, method 100 generates one or more variables based on at least some data types of some structured datasets. For example, the method may combine information from clinical trials associated with trial summaries stored in clinical trial databases 201 and 202, and then generate variables capable of screening the outcomes of all clinical trials to determine which outcomes are associated with clinical trial data stored in clinical trial database 201 or clinical trial database 202. In some embodiments, the variables generated based on these data types are stored in a unified clinical trial database.
[0091] At 110, a graphical user interface (GUI) 205 is provided to the user for performing or having performed analyses on one or more of the structured datasets 203. In some embodiments, the GUI may receive information or criteria from the user regarding data to be copied or accessed from different structured datasets 205. As an example, the GUI may receive a request from the user via GUI 205 to determine the probability of a drug successfully advancing through Phase III clinical trials. In some embodiments, the GUI may receive from the user information identifying the drug, and / or information about a set of parameters to be included in the analysis.
[0092] Figure 2 These are schematic diagrams based on some embodiments of this disclosure, depicting the logical information flow between exemplary databases, modules, and computing devices to generate and provide a unified clinical trial database and perform analyses (e.g., simulate clinical trial outcomes) on datasets stored in that database. Flow 200 illustrates the relationship between clinical trial data stored in clinical trial databases 201 and 202 and one or more datasets (trial summary dataset 203a, trial group dataset 203b, baseline dataset 203c, outcome dataset 203d, outcome comparison dataset 203e, and regulatory information dataset 203f) generated from clinical trial information obtained from clinical trial databases 201 and 202 and stored in the unified clinical trial database.
[0093] Examples of clinical trial databases include, but are not limited to: the Clinical Trials Summary Content (AACT) database, a publicly available relational database containing all information (e.g., protocol and outcome data elements) for each study registered on clinicaltrials.gov; and Informa Pharma Intelligence Trialtrove (Trialtrove) and Biomedtracker, databases that provide clinical trial data including datasets, clinical research information, and data from FDA clinical trials. In some embodiments, one or more additional clinical trial databases may be used. In some embodiments, clinical trial databases other than AACT or Trialtrove may be used. In some embodiments, information about clinical trials may also be obtained from publicly accessible sources (e.g., news articles) of one or more non-clinical trial databases. In some embodiments, information about clinical trials may also be obtained from one or more internal organizational sources (e.g., internal reports, internal databases, etc.).
[0094] According to some embodiments, the one or more datasets may be stored in one or more uniform standardized clinical trial databases 203g, which may be stored on server 203 or multiple servers. Clinical trial databases 201 and 202 include clinical trial data that can be used to create any of the one or more structured datasets or to operate in response to applying NLP module 204 to clinical trial data stored in clinical trial databases 201 and 202. For example, some clinical trial data stored in clinical trial databases 201 or 202 may be unstructured text data from which NLP module 204 can utilize meaningful structured data, which can then be added to one or more datasets. In some embodiments, the structured data generated by NLP module 204 may be data associated with variables already existing in one or more datasets. For example, NLP module 204 may generate structured data associated with outcome dataset 203d and / or outcome comparison dataset 203f based on unstructured data stored in one of clinical trial databases 201 or 202. The unstructured data may be data associated with the event type and date of one of the clinical trials, and according to some embodiments, NLP module 204 may determine the outcome or outcome comparison based on the event type and date. According to some embodiments, the outcome and outcome comparison associated with the event type and the date the event type occurred may be additional data added to the outcome dataset 203d and the outcome comparison dataset 203f. (See below for more details.) Figure 9 Examples of natural language processing according to some embodiments are described in more detail.
[0095] Users can use the GUI 205a on computing device 205 to retrieve, interact with, and / or analyze data in any dataset. In some embodiments, users can create datasets based on data requested by the user, which may include additional data not included in any of the independent datasets discussed herein, or may exclude data included in any of the independent datasets discussed herein. Users can enter commands via GUI 205a to generate analyses of the one or more datasets. In some embodiments, the analysis includes determining the probability that a drug will successfully complete all phases of a clinical trial.
[0096] Figure 3 This includes a GUI 205a (see [link to GUI 205a], which can be included in the unified clinical trial database 203g for users to access via computing device 205) Figure 2 The graphs use different datasets. In some embodiments, datasets obtained from a unified clinical trial database are used to simulate the outcomes of clinical trials. Trial summary dataset 303 describes general study attributes. These attributes may include clinical trial identifiers (e.g., NCT_ID), the name of the trial or study (e.g., Trial_title), the stage of the trial or study (e.g., stage), the number of groups included in the trial or study, planned or actual enrollment information provided in the protocol section of the trial or study record (e.g., timing variables), the design of the trial or study, and subgroup information that can be compiled from patient segmentation information in the trial or study record. Subgroup attributes may be variables generated based on combinations of information associated with patients.
[0097] According to some embodiments, server 203 creates trial summary dataset 303 by first unifying or standardizing certain clinical trial data from one or more clinical trial databases (e.g., clinical trial database 201 and clinical trial database 202), as described below. Figure 4 The explanation given.
[0098] According to some embodiments, the trial summary 303 may also include data on the primary investigational drug used in the clinical trial, as well as data on other drugs used in the clinical trial, in combination with the primary investigational drug, or as a control drug used as the primary investigational drug; these drugs are also referred to herein as other related drugs. According to some embodiments, for the primary investigational drug (e.g., drugclass_primary, drugsubclass_primary) and other (multiple) related drugs (e.g., (multiple) combination drug components or (multiple) control drugs) (e.g., drugclass_other, drugsubclass_other), the category data and subclass data are encoded as variables. According to some embodiments, the server can generate data such as... based on information obtained from a clinical trial database. Figure 3 The experiment summary dataset 303 shown includes one or more variables. Each of these variables in the experiment summary dataset will be described below. Figure 4 The example embodiments are discussed in detail in the description of the examples.
[0099] According to some embodiments, the experimental group dataset 306 can be generated based on AACT data stored in the clinical trial database 201. According to some embodiments, the experimental group dataset 306 can be used to organize data related to all treatment groups and treatment subgroups. It contains treatment groups for each study focusing on one or more drug components.
[0100] The baseline dataset 304 may also include clinical trial patient segmentation variables, such as symptoms or disease stages, symptoms or disease types, such as cancer type, number of lines of treatment, or gene expression. The baseline dataset 304 may also include variables associated with the treatments offered in each group of the clinical trial. In some embodiments, these variables may include data on drug combinations used to treat symptoms or diseases, classifications of drug combinations used to treat symptoms or diseases, or subcategories of drug combinations used to treat symptoms or diseases. The combination of structured datasets standardizes the variables used in the clinical trial database, specified by data type. Therefore, information related to treatments in each group of the clinical trial can be used to organize treatment groups.
[0101] In some embodiments, additional subgroup and dosage data for each group in the study can be compiled and added as supplementary data to further enrich the dataset. In some embodiments, drug coding can be based on the World Health Organization Drug Dictionary (WHODD) standard coding. In some embodiments, if the exact code for a drug is not included in the WHODD, a drug code can be generated to better classify the drug. According to some embodiments, the server can generate codes such as those obtained from a clinical trial database. Figure 3The experimental group dataset 306 shown includes one or more variables. Each of these variables in experimental group dataset 306 will be discussed below regarding... Figure 5 The example embodiments are discussed in detail in the description of the examples.
[0102] Baseline dataset 304 may be created at least in part based on information obtained from a clinical trial database (e.g., AACT data). In some embodiments, baseline dataset 304 may include demographic data for the entire clinical trial participant population, as well as baseline metrics and data for each clinical trial group or control group. Some information included in the baseline data may include a description of each baseline or demographic characteristic measured in the clinical trial. For example, baseline data may include age, sex / gender, race, ethnicity (if collected according to the protocol), and any other measurements(s) assessed at baseline and used in the analysis of primary outcome measurements(s). According to some embodiments, the server may generate data such as... Figure 3 The baseline dataset 304 shown includes one or more variables. Each of these variables in the baseline will be discussed below. Figure 6 The example embodiments are discussed in detail in the description of the examples.
[0103] To standardize baseline variables across different clinical trial reports, baseline category (bl_cat), baseline subcategory (bl_subcat), baseline subcategory level (bl_subcat_level), and baseline subcategory sublevel (bl_subcat_sublevel) variables can be derived from information obtained from clinical trial data in clinical trial databases 201 and 202, so that any baseline variable can be encoded using a hierarchical structure. In some embodiments, baseline data may be obtained from only a single clinical trial database (e.g., only from AACT).
[0104] Server 230 creates outcome dataset 305 by first unifying or standardizing certain data associated with clinical trial outcomes reported for each group of clinical trials, obtained from clinical trial databases 201 and 202, as described below. Figure 7 As explained. According to some embodiments, some data associated with clinical trial outcomes may include efficacy data, such as the overall response rate (ORR), progression-free survival (PFS), or overall survival (OS) endpoint variables observed in patients during the clinical trial.
[0105] ORR, PFS, or OS variables are exemplary efficacy endpoints used in oncology-related studies. For other disease domains, variables corresponding to different efficacy endpoints can be used. Relevant endpoints can be extracted, and these relevant endpoints can be based on information obtained from one or more clinical trial databases (e.g., AACT clinical trial data or Biomedtracker clinical trial data). Other data that may be included in the Outcome 305 dataset may include dictionaries for defining clinical trial outcomes. According to some embodiments, the Outcome 305 dataset may also include patient segmentation information for the trial group, which may be based at least in part on information about group variables from clinical trial databases (e.g., AACT clinical trial data and Trialtrove clinical trial data). In some embodiments, these variables may include data about drugs or drug combinations used to treat a condition or disease, classifications of drug combinations used to treat a condition or disease, or subcategories of drug combinations used to treat a condition or disease. According to some embodiments, the server can generate data such as information obtained from clinical trial databases. Figure 3 The outcome dataset 305 shown includes one or more variables. Each of these variables in the outcome dataset will be discussed below regarding... Figure 7 The example embodiments are discussed in detail in the description of the examples.
[0106] Server 230 creates outcome comparison dataset 307 by first unifying or standardizing certain data from information obtained from clinical trial databases 201 and 202 that are associated with outcome comparison results between treatment groups in clinical trials, as described below. Figure 8 This is as explained. Some of this data may include variables associated with one or more of the following: parameter value type (e.g., treatment difference, treatment ratio, odds ratio, etc.), p-value associated with the clinical trial, confidence interval associated with the clinical trial, and statistical test method used in the clinical trial. According to some embodiments, the server can generate data such as... based on information obtained from the clinical trial database. Figure 3 The outcome comparison dataset 307 shown includes one or more variables. Each of these variables in outcome comparison dataset 307 will be discussed below. Figure 8 The example embodiments are discussed in detail in the description of the examples.
[0107] As described above, server 203 creates a trial summary dataset 303 by first unifying or standardizing clinical trial information related to the drugs used during the clinical trial, obtained from one or more clinical trial databases (e.g., both clinical trial database 201 and clinical trial database 202). Server 203 can begin the process of standardizing the clinical trial information used to summarize the study by extracting data associated with certain variables that capture specific attributes of the study and stored in clinical trial databases 201 and 202. Figure 4 In the illustrated example embodiment, server 203 standardizes or unifies information related to variables associated with AACT data 301, Trialtrove data 302a, and BMT data 302b to generate a trial summary dataset (Trial_Summary 402).
[0108] In some embodiments, the variables in the trial summary dataset may be a default set of variables based on information that can be accessed or extracted from clinical trial databases 201 and 202. These variables may be presented to the user via a GUI. In some embodiments, the GUI may receive input from the user instructing server 203 to extract data from clinical trial databases 201 and 202 associated with other variables not included in the variable list associated with AACT 301 data, Trialtrove 302a data, and BMT 302b data. In other embodiments, the GUI may receive input from the user instructing server 203 to remove variables associated with AACT 301 data, Trialtrove 302a data, and BMT 302b data from the trial summary.
[0109] Figures 4 to 8 This description is based on data from the AACT, Trialtrove, and Biomedtracker (BMT) databases and is for illustrative purposes. Those skilled in the art will understand from this disclosure that embodiments using other clinical trial databases with different data formats or data types also fall within the scope of this invention.
[0110] Figure 4The following is a logic diagram 400, based on some embodiments of this disclosure, depicting a process for generating a dataset of clinical trial summary 303. This process may begin with server 203 connecting to clinical trial databases 201 and 202 to extract raw summary data associated with AACT clinical trial data 301. Data associated with AACT clinical trial data 301 may include variables such as: the NCT ID associated with the clinical trial, the number of groups associated with the clinical trial (num_of_arms), the year the clinical trial started (start_year), the year the clinical trial completed (completion_year), the primary year of completion (primary_completion_year), the country, state, or province where the clinical trial was conducted (region), the names of the investigators(s) who conducted the clinical trial (investigator), the trial title (trial_title), and the overall status of the clinical trial (overall_status).
[0111] Data associated with AACT clinical trial data 301 may also include sponsor(s) names and sponsor(sponsor_type). Sponsor(sponsor) types can be industry sponsors, academic sponsors, or government sponsors. AACT clinical trial data 301 may further include the enrollment completion time (enr_comp_time). Data associated with AACT clinical trial data 301 may also include the estimated total number of subjects to be enrolled (target number) or the total number of subjects actually enrolled in the clinical study (enrl_num).
[0112] AACT clinical trial data 301 may also include an outcome type variable, which can be used to screen a range of outcome event types associated with the clinical trial. Event types may include data associated with clinical trial outcomes, data associated with reported events in the clinical trial, and data associated with the clinical trial's participant flow. Data associated with reported events may include summary information about reported adverse events (any unexpected or adverse medical events to a participant, including abnormal physical examinations, laboratory findings, symptoms, or illnesses), including serious adverse events, other adverse events, and death. Data associated with the clinical trial's participant flow may include enrollment data related to the enrollment process and pre-assignment details (i.e., significant events in the study that occur after participant enrollment but before participant assignment). Data associated with the participant flow includes data on the participant flow applicable to all milestones of the clinical trial.
[0113] Server 203 can set the outcome type variable to "outcome" to extract outcome data related to clinical trials associated with the same NCT ID from clinical trial databases 201 and 202.
[0114] Server 203 can also extract raw data associated with Pharma Intelligence clinical trial data from clinical trial databases 201 or 202. This raw data may include Trialtrove data 302a and BMT data 302b. The raw data may include variables associated with: the clinical trial identifier (trialId), the clinical trial title (trialTitle), and the clinical trial status (trialStatus). The raw data may also include the clinical trial start date (trialStartDate), the clinical trial design (trialStudyDesign), and the clinical trial last modified date (trialLastModifiedDate).
[0115] The raw data may further include data related to the outcome of the clinical trial (trialOutcomeDetails), data related to the results of the clinical trial (trialResults), and any notes about the clinical trial (trialNotes).
[0116] The raw data may also include the primary endpoint of the clinical trial and any other endpoints associated with the clinical trial. Furthermore, the clinical trial data may include a list of trial sponsors and certain therapeutic areas (e.g., oncology). In some embodiments, the clinical trial data may include a first group of primary drugs tested during the clinical trial. In some embodiments, the clinical trial data may also include one or more of the following: a second group of primary drugs tested during the clinical trial, a first group of alternative drugs tested during the clinical trial that are different from the first group of primary drugs, and a second group of alternative drugs tested during the clinical trial that are different from the second group of primary drugs.
[0117] Pharma Intelligence clinical trial data may also include a unique identifier (Drug ID) for identifying the specific drug included in the Pharma Intelligence clinical trial data, as well as any brand name associated with that Drug ID, and indication group (drug indication + pivotal study (yes / no), e.g.: Breastcancer_Pivotal: Yes).
[0118] Server 203 can generate disease subtype data (e.g., data associated with melanoma, non-small cell lung cancer such as squamous cell carcinoma, and adenocarcinoma) based at least in part on data related to the trial study design and patient segmentation data associated with the trial. This data is used to determine the number of treatment lines for diseases corresponding to the different disease subtypes identified in the disease subtype data. In addition to determining the disease subtype data and the data related to the trial study design, server 203 can also generate gene expression data associated with the diseases corresponding to the different disease subtypes identified in the disease subtype data.
[0119] After server 203 extracts AACT clinical trial data and PharmaIntelligence clinical trial data from clinical trial databases 201 and 202, server 203 can apply a text mining keyword matching algorithm to extract data associated with cancer subtype, gene expression, stage, and treatment line number variables based on the patientSegment variable. For example, if a column of patientSegment includes information related to different subtypes of the disease, such as "Adenocarcinoma," "Squamous Cell," and "Large Cell," i.e., different subtypes of the disease, then server 203 can group the different subtypes of the disease into a single value, which can be represented as Adenocarcinoma|Squamous Cell|Large Cell. Similarly, if a column of patientSegment includes information related to different stages of the disease, such as "Stage I" and "Stage II" for cancer, server 203 can group the different stages of the disease into a single value, which can be represented as Stage I|Stage II. Server 203 can also apply text mining keyword matching algorithms to process gene expression and treatment line count variables to generate single values for both. In response to applying text mining keyword matching algorithms to combined AACT clinical trial data and Pharma Intelligence clinical trial data, server 203 can generate indicators that identify the data source from which the data associated with cancer subtype, gene expression, stage, and treatment line count variables originated. This allows tracking the sources of the data types associated with patientSegment (multiple sources).
[0120] Due to the complexity of drug nomenclature and mechanism-of-action classification used by different entities conducting clinical trials, there is a lack of consistency in drug information between AACT clinical trial data and Pharma Intelligence clinical trial data. Therefore, according to some embodiments, a drug dictionary is created as a standard reference to encode drug names and classify mechanisms of action.
[0121] Server 203 generates, at least in part, a set of variables that can be used to uniquely describe the drug being administered in a clinical trial, based on the variables trialPrimaryDrugsTested, trialOtherDrugsTested, bmtprimarydrugstested, and bmtotherdrugtested.
[0122] Server 203 can expand `trialPrimaryDrugsTested` and extract variables such as `drugId`, `drugPrimaryName`, `drugName`, `mechanismOfAction`, `mechanismSynonyms`, and `directMechanism` associated with the primary drug administered during the clinical trial corresponding to the `Trialtrove` clinical trial data. Server 203 can also expand `trialSecondaryDrugsTested` and extract variables associated with other drugs in the `Trialtrove` data. Server 203 can expand the `bmtprimarydrugstested` variable and extract variables such as `drugID`, `brandname`, and `Target` associated with the primary drug administered during the clinical trial corresponding to the `BMT` clinical trial data. Server 203 can also expand the `bmtotherdrugtested` variable and extract `drugID` associated with other drugs administered during the clinical trial, also corresponding to the `BMT` clinical trial data.
[0123] Server 203 can merge all four expanded datasets via drugid. The result of a full connection between the Trialtrove dataset and the BMT dataset is a complete list of all drugs designed to treat (these) diseases (e.g., variables All.Names), which are included in the first subset of the Pharma Intelligence clinical trial data and the second subset of the PharmaIntelligence clinical trial data.
[0124] Variables such as Target, directMechanism, and mechanismSynonyms can be used to classify drugs according to drug dictionaries (such as the World Health Organization (WHO) Drug Dictionary).
[0125] After a complete drug list is created and a predetermined list of variables is selected by server 203, the drug list needs to be coded according to a drug dictionary. In some embodiments, the drug dictionary may be stored locally on server 203. The drug dictionary includes data associated with Trialtrove clinical trial data and data associated with BMT clinical trial data. The drugs are then coded according to WHODD or specified disease-specific rules, as drug categories and drug subcategories.
[0126] According to some embodiments, an established drug dictionary can be used to classify each drug into certain categories and subcategories with similar or identical names. Server 203 can extract trial group data associated with each drug being investigated in each clinical trial, and if more than one drug is used in a group, generate combination drugs used in each trial. Using the merged drug dictionary, each therapeutic drug and drug combination is encoded and classified into drug combination categories and drug combination subcategories. Figure 5 Logic diagram 500 illustrates the process by which server 203 creates a clinical trial group dataset to add classifications and subclassifications to the drugs used as treatments in the experimental group.
[0127] According to some embodiments, server 203 may initiate the process by loading trial group data from clinical trial database 201. This trial group data includes information about different design groups, interventions, terms, phrases, or names synonymous with a specific intervention, and the interventions applied to different design groups. Server 203 may extract design group data associated with the trial group from clinical trial database 201. This data may include information about subjects in a protocol-specified group, subgroup, or cohort who are assigned to receive (multiple) specific interventions or observations in a clinical trial according to a protocol used to administer certain drugs to subjects in that specified group, subgroup, or cohort.
[0128] Server 203 can extract intervention data associated with trial groups from clinical trial database 201. This data may include specific medical interventions or exposures, including but not limited to drugs, medical devices, procedures, vaccines, and other products of interest in the clinical trial or associated with the trial group or subgroups within the trial group. Server 203 can extract terms, phrases, or names from clinical trial database 201 that are synonymous with interventions applied to different design subgroups. Each term, phrase, or name is associated with an intervention related to the clinical trial. For example, server 203 may extract three different names or phrases, each associated with a specific intervention for that trial group.
[0129] According to some embodiments, server 203 can extract design group-specific intervention data from clinical trial database 201, which is associated with specific interventions applied to specific groups in a clinical trial. The design group-specific intervention data serves as a cross-reference between groups and corresponding interventions applied to those groups. For example, according to some embodiments, if a clinical trial has multiple groups and multiple interventions, the design group-specific intervention data can specify which interventions are associated with which groups.
[0130] Server 203 can also be programmed to extract specific variables associated with intervention data. For example, server 203 can identify and extract data associated with the NCT ID of a clinical trial, data describing the type of intervention being applied in the trial group, which may include specific procedures as well as drugs, biological interventions, radiation therapy, etc. Additional intervention data that server 203 can extract includes the names and descriptions of one or more drugs used during the clinical trial.
[0131] Server 203 can also be programmed to extract design group-specific variables, including NCT ID, group title, and description of the group. The description of the group may include information identifying the effect of the intervention received by the subject, and / or may include different types of groups, including but not limited to experimental group, positive control group, placebo control group, sham treatment control group, and no intervention group.
[0132] Then, according to some embodiments, server 203 combines data associated with information corresponding to different design groups, interventions, terms, phrases, and / or names used to refer to specific interventions, and interventions applied to different design groups to create a table including all the extracted variables described above. Server 203 can combine this data by using common key variables to combine data associated with information corresponding to different design groups, interventions, terms, phrases, or names used to refer to specific interventions, and interventions applied to different design groups. These common key variables may include any one of the following: NCT ID, an identifier associated with an intervention, an identifier associated with an intervention, or an identifier associated with a design group. The resulting table (drug table) is then combined with a drug dictionary, as explained below, resulting in the assignment of drug categories and subcategories to drugs administered in the experimental group.
[0133] According to some embodiments, server 203 can select from the drug table variables associated with the NCT ID, names associated with the drug administered in the experimental group, identifiers associated with each design group, drug descriptions, titles associated with the experimental group, descriptions of the experimental group, and the group types included in the experimental group, and can group the data associated with the table generated by merging the drug table and the drug dictionary according to the selected variables. The group types included in the experimental group can be experimental groups or control groups. Figure 5As shown, according to some embodiments, AACT 301 may include the variable NCT ID, group type, design group ID associated with each design group, name, arm title associated with the trial group, intervention type, drug intervention associated with the intervention type, and drug description.
[0134] The drug dictionary 401 may contain a list of all different names associated with the drugs administered in the test group (All_Names). The drug dictionary 401 may also include classification information (drugclass) associated with each drug included in the drug dictionary and subclass information (drugsubclass) associated with each drug included in the drug dictionary.
[0135] Then, according to some embodiments, server 203 can perform word segmentation on the name variables in the drug table, and can merge the drug table and drug dictionary based on the segmented name variables and the reconstructed name variables in the drug dictionary. Then, the drugs in the drug table are encoded into drug categories and drug subcategories.
[0136] After server 203 merges the modified drug table and the modified drug dictionary, the data in the table resulting from the merger of the modified drug table and the modified drug dictionary is grouped together according to the variables associated with the NCT ID from the drug table, the names associated with the drugs administered in the experimental groups, the identifiers associated with each design group, the descriptions of the drugs, the titles associated with the experimental groups, the descriptions of the experimental groups, and the types of different experimental groups.
[0137] According to some embodiments, after server 203 groups the data in the resulting table according to the aforementioned variables, server 203 can generate a variable associated with the most commonly used drug name of the drug in that group. This variable can be represented as a drugname_combo as shown in Trial_Arm 501 by combining each drug component of Trial_Arm. This drugname_combo variable can be encoded as drugclass_combo and drugsubclass_combo as shown in Trial_Arm 501. Furthermore, Trial_Arm 501 may include an NCT ID (NCT_ID), a trial arm title (Ta_arm_title), and a group type (group_type).
[0138] According to some embodiments, each clinical trial may have corresponding baseline data or outcome data. Baseline data may include data collected for all subjects and each experimental group at the start of the clinical trial. This data may include, but is not limited to, demographic data such as age, sex, race, and ethnicity, as well as study-specific indicators (e.g., systolic blood pressure, prior antidepressant treatment, etc.). Server 203 can, according to... Figure 6 The logic diagram 600 extracts baseline data associated with clinical trials.
[0139] Server 203 can begin the process of extracting baseline data associated with a clinical trial by first determining whether baseline data exists for that clinical trial. For some clinical trials, baseline data may not exist, for example, because the clinical trial has been abandoned or has not yet started. According to some embodiments, server 203 can determine whether any baseline data exists in clinical trial database 201 by attempting to access the baseline data stored in clinical trial database 201.
[0140] If server 203 determines that baseline data does exist for the clinical trial, server 203 can retrieve the baseline data from clinical trial database 201. Baseline data associated with the clinical trial may include outcome data associated with one or more subject groups. For example, according to some embodiments, outcome data associated with the one or more groups may include an integrated or summary list of group titles and descriptions that can be used to report summary outcome information.
[0141] Outcome data associated with the one or more subject groups may include NCT IDs and outcome type variables that indicate clinical trial outcomes. For example, according to some embodiments, the value of the outcome type variable may be “Baseline,” “Outcome,” “Reported Event,” or “Participant Flow.” In some embodiments, server 203 may be programmed to create filters that, when retrieving data associated with outcome type variables from clinical trial database 201, will only extract certain outcome type data, such as baseline outcome type data. According to some embodiments, outcome data associated with the one or more subject groups may also include a title associated with the baseline of the trial group in which the one or more subject groups are included, and a description of the baseline of that trial group. The description of the baseline of the trial group may include data collected for all subjects and each trial group or control group at the start of the clinical trial. This data may include demographic data, such as age, sex, race, and ethnicity, and study-specific indicators (e.g., systolic blood pressure, prior antidepressant treatment, etc.).
[0142] As described above, according to some embodiments, the baseline data may further include a set of baseline measurements, which may include an NCT ID, a title associated with the baseline of the test group, a description of the baseline of the test group, a classification, a set of units, a parameter type, a value associated with the parameter type, a quantity associated with the parameter value, a dispersion type, a dispersion value associated with a particular type of dispersion, a lower limit or limit of dispersion, or an upper limit or limit of dispersion.
[0143] As described above, according to some embodiments, baseline data may also include a set of baseline count information. In some embodiments, the baseline can be generated based on data obtained from a single clinical trial database (e.g., AACT). In some embodiments, baseline measurements and baseline counts are used to generate baseline data. The corresponding baseline data is represented in... Figure 6 The variable list is in AACT 301.
[0144] In some embodiments, server 203 may merge variable list AACT 301 with trial summary 402 to establish relationships between baseline data and disease subtype classification, disease stage in subjects, number of treatment lines used to treat subjects, and expression of genes associated with a subject's disease, according to some embodiments. For example, the disease may be a type of cancer, and staging may include the stage of that cancer type for all subjects; the number of treatment lines may include a specific sequence of administration of different therapies to subjects through different stages as the disease progresses; and gene expression may include expressed or unexpressed genes associated with the disease. After server 203 merges variable list AACT 301 with trial summary 402, according to some embodiments, server 203 may further merge trial group 501 into the merged variable list AACT 301 and trial summary 402 using NCT ID. Server 203 merges trial group 501 with the merged variable list AACT 301 and trial summary 402 to determine if there is any similarity between the title associated with trial group 501 and the title associated with the baseline of trial group 501. Server 203 can determine whether the two titles are exactly matched, closely matched, or similar by segmenting the titles associated with test group 501 and the baseline of test group 501, and then comparing the number of segmented words in the titles. According to some embodiments, server 203 compares the number of segmented words included in the baseline group title with the number of the same segmented words included in the test group 501 title, and determines which of the test groups associated with 501 has the maximum number of the same segmented words as the baseline group title. Server 203 can then add information from the test group with the maximum number of segmented words common to the baseline group title to the baseline dataset. For example, in some embodiments, patient segmentation information available in the summary data is subsequently added to the baseline dataset for future patient subgroup identification.
[0145] According to some embodiments, server 203 may also determine whether the title of the experimental group includes phrases such as "all subjects" or "total," which indicate whether all subjects are included in the experimental group. Determining that the title includes these phrases or phrases synonymous with them indicates whether the baseline count includes all subjects. If the title does include a phrase indicating that all subjects are included in the experimental group, server 203 may create a flag indicating that the experimental group includes all subjects. If the title does not include a phrase indicating that all subjects are included in the experimental group, according to some embodiments, server 203 may set the flag to a value indicating that the experimental group does not include all subjects. This flag may be referred to as the overall population flag and may be represented by (tot_pop_flg) in the baseline 602 dataset.
[0146] To standardize the naming of baseline variables from different sources, a hierarchical structure is used to uniformly encode them, according to some embodiments. Three variables, bl_variable, bl_description, and classification, can be derived from the AACT table to derive bl_cat1, bl_cat2, bl_level, and bl_class_std. This dataset can be named baseline_code_dictionary_v4. Furthermore, the derived variables baseline_cat1 to bl_cat, bl_cat2 to bl_subcat, and bl_level can be renamed to bl_subcat_level and bl_class_std to bl_subcat_sublevel.
[0147] After server 203 completes the above-mentioned merging processes, server 203 can generate, as follows: Figure 6 The baseline 602 dataset is shown. The server can receive requests for data included in the baseline 602 dataset stored on the server 203 from the computing device 205 via the GUI 205a, and the server 203 can send the requested data to the computing device 205.
[0148] According to some embodiments, when server 203 sets the value of the result type variable to "Outcome", it is in accordance with... Figure 7 Consistent with logic diagram 700, server 203 can be programmed to extract only outcome type data when retrieving data associated with the outcome type variable from clinical trial database 201. Outcome data associated with a clinical trial may include outcome data associated with one or more subject groups. For example, according to some embodiments, outcome data associated with the one or more groups may include an integrated or summary list of group titles and descriptions that can be used to report summary outcome information. This summary outcome information may include a summary of the outcomes associated with the clinical trial. Outcome type data may include an NCT ID associated with the clinical trial, an identifier (result_group_id) associated with the outcome of a group of one or more subjects being observed in the clinical trial, a title (oc_arm_title) associated with the trial group corresponding to the outcome data, and a description (oc_arm_description) associated with the trial group corresponding to the outcome data.
[0149] According to some embodiments, the outcome data may further include a set of outcome measures, which may include summary data associated with primary and secondary outcome measures for each subgroup included in the clinical trial. This summary data may include parameter estimates (e.g., overall response rate, median progression-free survival, etc., for some embodiments) and measures of dispersion or precision. The outcome data may further include outcome count information, which may include the sample size or count included in the analysis of each outcome for each subgroup in the clinical trial. This count or sample size may represent the number of subjects, but may also represent other units of measurement, such as “lesions” of the eye or other characteristics associated with the subject’s body. The outcome data may further include the outcome population, i.e., the patient population associated with the outcome measure. For example, in some embodiments, this would include patients with high PD1 expression. The aforementioned outcome data may be extracted by server 203 from clinical trial database 201.
[0150] Data associated with a set of outcome measures can be represented by several variables. This set of outcome measures can be characterized by a title (oc_title), a description of the clinical trial outcome (oc_description), and the units used for capture during the clinical trial (e.g., in some embodiments, this would include median progression-free survival in weeks, months, or years). According to some embodiments, the set of outcome measures can be further characterized by parameter type, values associated with that parameter type, the number associated with that parameter value, dispersion type, dispersion values associated with a particular dispersion type, a lower limit or bound of dispersion, or an upper limit or bound of dispersion.
[0151] According to some embodiments, after server 203 has extracted different outcome data from clinical trial database 201, it can combine the different outcome data using NCT ID, an identifier associated with the outcome of grouping one or more subjects being observed in the clinical trial, and / or an identifier associated with the outcome.
[0152] Server 203 can extract additional outcome data from clinical trial database 202, which may include information associated with one or more drug efficacy endpoints corresponding to the clinical trial. According to some embodiments, an event data table, denoted as the variable `eventDataTable`, may include a description of the experimental group or information about the treatment administered to a subject group during the experimental group period. This information may be represented as the variable `oc_arm_title`. The reason this information can be represented as the variable `oc_arm_title` is that the description of the experimental group is typically included in the experimental group's title. The event data table may also include information about the number of subjects who received the treatment. This information may be represented as the variable `oc_sample size`, also known as a count of the number of subjects. A results data table, denoted as the variable `dataTableResults`, may include information associated with a description of the efficacy endpoint of the experimental group. This description may be represented as the variable `oc_title`.
[0153] Server 203 can combine the data extracted from clinical trial databases 201 and 202 by simply stacking the data extracted from them one on top of the other. The resulting dataset can be represented as the variable result_sub.
[0154] The `eventType` variable can include different types of outcomes associated with a clinical trial, including top-line results (regardless of the trial's phase or duration) indicating statistical significance, final results, published results, and updated results. Final results can be those associated with the completion of a clinical trial. Published results can include those associated with the completion of a clinical trial that has also been published. Updated results can be those that may or may not have been published, but include new outcome data that did not exist at the time point prior to the creation of the updated result. Server 203 can extract one or more drug efficacy endpoints from any of the above results. For example, server 203 can extract overall response rate (ORR), progression-free survival (PFS), and / or overall survival (OS) efficacy endpoints from one or more of the above results.
[0155] Server 203 can determine which clinical trials include these efficacy endpoints by determining whether the outcome title variable (oc_title) of each clinical trial included in result_sub contains any keywords, such as those associated with ORR, PFS, or OS. Server 203 can compare the words contained in the outcome variable corresponding to each clinical trial title with keywords (such as the phrase "overall survival" and other relevant keywords) to identify and unify the outcome. Then, in some embodiments, server 203 can determine that the clinical trial does indeed include, for example, overall survival data associated with that clinical trial. Similar algorithms are applicable to other trial outcomes, such as overall response rate, partial response, duration of response, etc.
[0156] According to some embodiments, after performing the keyword search described above, server 203 can encode efficacy endpoints across different clinical trials. The encoding of efficacy endpoints corresponds to categories and subcategories. The category can be represented as an outcome category variable (oc_class) and an outcome subcategory variable (oc_subclass). Server 203 generates an outcome dictionary that includes both the outcome category variable and the outcome subcategory variable.
[0157] According to some embodiments, the outcome dictionary can be generated by filtering out data associated with the outcome time variable `oc_title` extracted from clinical trial databases 201 and 202, and filtering out data associated with the outcome classification variable `oc_classification` extracted from clinical trial database 201. Outcome categories and subcategories are defined by different disease domain requirements and analytical needs, and can be updated when new endpoints of interest need to be added. In addition to comparing the keywords associated with the efficacy endpoints listed above with words in the outcome title `oc_title`, server 203 can also compare the same keywords with words included in the outcome description variable `oc_description`. Each efficacy endpoint corresponds to a category in the outcome dictionary, and the category in the outcome dictionary can be represented as `oc_class`. The outcome dictionary also includes corresponding outcome subcategories, and efficacy endpoints also correspond to those subcategories.
[0158] In addition to the `oc_title` variable, the ending dictionary can include the ending description variable `oc_description` and the ending classification variable `oc_classification`. After server 203 generates the ending dictionary, server 203 can merge the ending dataset variable `result_sub` with the ending dictionary on the `oc_title`, `oc_description`, and `oc_classification` variables. Figure 7In this context, the merge can be presented as a concatenation of the AACT 301 data and the BMT 302b data with the outcome_variable_dic 702, so that the oc_class and oc_subclass variables are created or added to the outcome 701 dataset produced by the merge, and the outcome variables are classified through this dictionary.
[0159] like Figure 7 As shown, according to some embodiments, the list of variables included in AACT 301 data may include any or all of the following variables: oc_classification, oc_description, oc_dispersion_type, oc_dispersion_value_num, oc_dispersion_lower_limit, oc_dispersion_upper_limit, time_frame, and population (which may also be represented as oc_population). According to some embodiments, AACT 301 data may further include NCT ID, sample_size, oc_units, oc_title, oc_param_value_num, oc_type, oc_arm_title, and oc_arm_desc. AACT 301 data is extracted from the Clinical Trials Database 201. BMT 302b data may include eventType, eventDate, eventPhase, drugPhase, likelihoodofApproval, averageApprocal, changetoLikelihoodApproval, numberOfEvaluablePatients, and Endpoint_description and treatment description (…). Figure 7 Any or all of the treatment_descr variables in the dataset.
[0160] `numberOfEvaluablePatients` corresponds to the sample size or count of subjects in a clinical trial. Since this value is the same as the count or sample size in AACT 301 data, server 203 can store the values associated with the `sample size`, `count`, and `numberOfEvaluablePatients` variables in the `sample_size` variable. The `Endpoint_description` variable can be associated with the `oc_title` variable because, as explained above, the outcome title can contain words related to one or more efficacy endpoints. Therefore, server 203 can store the description provided in the `Endpoint_description` variable in the `oc_title` variable. The `treatment_descr` variable can include a description of the treatment group, which is similar to the treatment group description associated with the `oc_arm_desc` variable; therefore, server 203 can store the description stored in the `treatment_descr` variable in the `oc_arm_desc` variable.
[0161] According to some embodiments, after server 203 merges the outcome dictionary with the result_sub variable, the resulting dataset can be merged with test group 501 using the NCT ID to generate, for example... Figure 7 The dataset shown is the outcome 701 dataset. After merging the experimental group 501 with the resulting dataset, the server 203 can perform word segmentation on the outcome experimental group title variable (oc_arm_title) and the treatment or experimental group title (arm_title), and merge the outcome experimental group title and the experimental group title based on the similarity of the title names.
[0162] To extract genesub and cancersub, three functions (e.g., R functions) such as match_gene, match_sign, and match_cancer can be generated programmatically from oc_arm_title, trial_title, oc_description, oc_arm description, and oc_classification.
[0163] According to some embodiments, PopulationSub can be extracted using keyword search algorithms on variables such as group title, group description, ending title, title, category, and description (in order).
[0164] In response to a user request for outcome data associated with one or more relevant clinical trials, server 203 can display the outcome 701 dataset on the display of computing device 205 via GUI 205a.
[0165] Figure 8 The following is a logic diagram, based on some embodiments of this disclosure, depicting a process 800 for generating a clinical trial outcome comparison dataset. The outcome comparison dataset may include statistical analysis results, such as odds ratios, hazard ratios, differences in response rates, differences in risk, relative risk ratios, or other estimated parameters between different treatments. Server 203 may extract comparison results from clinical trial database 201 and clinical trial database 202, and stack them vertically.
[0166] Server 203 may send requests to clinical trial database 201 to access and retrieve data associated with statistical outcome analyses of primary and secondary clinical trial outcomes. According to some embodiments, this data may include estimates of treatment effects, confidence intervals, measures of other dispersion, and p-values. This data may be referred to as outcome analysis data and may be represented as the variable `outcome_analysis`. Server 203 may also send requests to clinical trial database 201 to access and retrieve data identifying comparison groups associated with each statistical outcome analysis. This data may be referred to as outcome analysis group data and may be represented as the variable `outcome_analysis_group`. Server 203 may also send requests to clinical trial database 201 to access and retrieve data associated with an integrated or summary list of grouped clinical trial titles and descriptions used to report summary outcome information. Server 203 may also send requests to clinical trial database 202 to access and retrieve data associated with the outcomes of clinical trial data stored in clinical trial database 202.
[0167] Outcome analysis grouping data may include NCT IDs associated with different clinical trials, identifiers associated with outcome analyses of different clinical trials, and identifiers associated with different subject groups participating in outcome comparisons of different subject groups.
[0168] Outcome analysis data may include the NCT ID, an identifier associated with the outcome analysis (which may be represented as the variable `outcome_analysis_id`), and the type of statistical test applied to the clinical trial data for analysis. According to some embodiments, this statistical test may be any or all of a superiority test, a non-inferiority test, or an equivalence test. Outcome analysis data may also include a description of the statistical test.
[0169] According to some embodiments, outcome analysis data may also include a dispersion-type variable that measures the dispersion of clinical trial group data in a specific group relative to another group. An example could be the standard error of the mean. The dispersion-type variable may be represented as `dispersion_type`. According to some embodiments, outcome analysis data may also include corresponding values for the dispersion-type variable. According to some embodiments, outcome analysis data may also include modifiers, symbols, or notations indicating whether the p-value associated with the statistical test is less than or greater than a certain value. This modifier may be represented as the variable `p_value_modifier`. Outcome analysis data may also include a p-value variable associated with the statistical test, which may be represented as `p_value`.
[0170] According to some embodiments, outcome analysis data may also include information associated with confidence intervals for the statistical test. For example, outcome analysis data may include a variable associated with the number of sides (one-sided or two-sided) of the confidence interval, and may be represented as ci_n_sided. Additional information associated with the confidence interval may include the level associated with the confidence interval, which may be represented as a percentage and captured by the variable ci_percent. According to some embodiments, the lower limit (which may be represented as the variable ci_lower_limit) and the upper limit (which may be represented as the variable ci_upper_limit) associated with the confidence interval are also included in the outcome analysis data.
[0171] According to some embodiments, outcome analysis data may also include statistical methods for analyzing clinical trial group data. For example, this method may be a statistical test for calculating a p-value. It can be expressed as a variable method and may include any of the following tests: analysis of covariance (ANCOVA), analysis of variance (ANOVA), chi-square test, corrected chi-square test, Cochran-Mantel-Haenszel test, Fisher's exact test, Kruskal-Wallis test, log-rank test, Mantel-Haenszel test, McNemar test, mixed model analysis, regression analysis, Cox regression, linear regression, logistic regression, sign test, one-sided t-test (1-Sided), two-sided t-test (2-Sided), Wilcoxon (Mann-Whitney) test, or any other statistical test.
[0172] According to some embodiments, outcome analysis data may also include information about the procedures used to estimate the intervention effect, and may be, for example, statistical tests of hypotheses. Information about the procedures used to estimate the intervention effect may be represented as the variable `estimate_description`. Outcome analysis data may also include descriptions of groupings, and may be represented as the variable `groups_description`, such as, in some embodiments, the intention-to-treat population, the safety population, or the high PD1 expression population.
[0173] Server 203 can also select the variables to be included in outcome data 70: nct_id, oc_arm_title, oc_geneSub, oc_geneSign, oc_cancerSub, populationSub, drug_name_combo, drug_intervention, drugclass_combo, and drugsubclass_combo. According to some embodiments, server 203 can also extract outcome analysis-related data from clinical trial database 202 and extract NCT_ID, Title, result_type, group_type, arm_title, and arm_description from clinical trial database 201.
[0174] For BMT outcome analysis, server 203 can consider the raw result_sub data and filter it for the variables result_type="BMT_outcome" and group_type="difference". In some embodiments, the outcome dictionary generation method summarized above is used to generate separate BMT and AACT comparison dictionaries. The BMT comparison data is concatenated with the BMT comparison dictionary on outcome_title. Then, the BMT comparison dataset is concatenated with the outcome table on nct_id to retrieve the experimental group title and population variables. For AACT, analysis_group.rdata and analysis r data are merged using NCT_ID, ctgov_group_code, and result_group_id, and then merged with result_sub on NCT_ID to retrieve the oc_arm title. In some embodiments, outcome_bmt comparison data is stacked with AACT comparison data.
[0175] The server merges the outcome comparison dataset with the final outcome dataset on oc_title and NCT_DI to obtain arm title and arm title2 as key variables for obtaining drug_level information.
[0176] The above describes the specific variable names. Figures 4 to 8 This is for illustrative purposes only. Those skilled in the art will understand from this disclosure that any other variable names may be used.
[0177] Figure 9 This is a logic diagram based on some embodiments of the present disclosure, depicting a process of applying natural language processing to unstructured text stored in a database to extract clinical trial outcomes and comparative data of clinical trial outcomes. The natural language processing may include one or more computer-readable instructions that, once executed by server 203, cause the server to send a prompt to the clinical trial database or databases 201 and / or 202, requesting clinical trial data associated with the outcomes or results corresponding to a clinical trial. The requested clinical trial data may include unstructured text interpreting the outcomes or results of a clinical trial. When server 203 executes the computer-readable instructions for natural language processing, it can detect portions within the unstructured text corresponding to design and outcome information associated with the clinical trial. Server 203 may enter credentials, such as a user's username and password, into a webpage and request the raw data or pre-configured datasets disclosed herein from a database hosting raw data or pre-configured datasets (e.g., Trialtrove 901), which are stored in file 902.
[0178] The natural language processing (NLP) can extract the requested data (906) from database (901) by detecting portions of the file corresponding to the requested data. For example, when the server executes instructions associated with NLP, it can detect certain portions of the raw data, such as the results portion or the portion associated with the clinical trial design. The server can also clean or normalize the data to adapt it to the data format that will be used to generate one or more pre-configured datasets discussed herein. The NLP instructions can further cause server (203) to generate file (910) containing cleaned text, portions containing data related to the requested data, and specific sentences associated with the requested data. When executing the NLP instructions, server (202) can use file (910) to automatically annotate (907) sentences with phrases and / or tags related to the requested data.
[0179] Configuration function 904, when executed by server 203, causes server 203 to use one or more named entity models 905 when executing instructions associated with natural language processing to detect specific data in cleaned text, such as disease, cell type, dose, dose intensity, endpoint, drug name, gene expression data, etc.
[0180] For example, the natural language processing may include instructions that cause server 203 to annotate unstructured text with tags that identify certain named entities, such as any or all diseases, cell types, dosages and intensities of drugs administered to subjects during clinical trials, efficacy endpoints, types of drugs used, statistics collected throughout the clinical trial, and the number or count of patients at the clinical trial outcome or result. Server 203 may annotate the unstructured text by adding tags to sentences, phrases, or words that identify any of the aforementioned entities or any other relevant information that may be used to identify data corresponding to each dataset stored in server 203.
[0181] After server 203 completes automatic annotation 907, the server can then perform entity pattern matching 908 between the entities identified above and entities identified in the unstructured text. For example, the aforementioned list of entities can be input into a natural language processing unit, which can then enable server 203 to perform entity association, in which server 203 can identify information such as biomarker levels, treatment type, tumor type, treatment dose, treatment group or trial group, and subgroup, all of which are associated with clinical trials within the unstructured text. After server 203 completes pattern matching, server 203 can generate a database 909 and a mapping between the identified information and the automatically annotated sentences or phrases. The extracted information can then be used by server 203 to create one or more datasets stored on server 203. After generating the database (e.g., a combined clinical database), server 203 can also output a file containing one or more structured datasets based on data requested by the user.
[0182] In some embodiments, information or data is periodically (e.g., daily, weekly, etc.) obtained from clinical trial databases, and new or changed data is processed and integrated into a combined standardized clinical trial database. In some embodiments, this combined standardized clinical trial database utilizes up-to-date, unbiased data to support and enable rapid, comprehensive meta-analysis compared to conventional systems used for analyzing clinical trial data. In some embodiments, information processing and integration may be tailored and guided by therapeutic domain knowledge and drug development experience. In some embodiments, this combined standardized clinical trial database can better support robust decision-making in clinical development and planning.
[0183] Figure 10 An exemplary computing device or system 1000 is schematically depicted (e.g., Figure 2 The system 1000 can be used to perform operations described in one or more embodiments of any method disclosed herein. For example, the system 1000 can be included in any or all server components or other computing devices(s) discussed herein. The system 1000 may include one or more processors 1010, one or more memories 1020, one or more storage devices 1030, and one or more input / output (I / O) devices 1040. Components 1010, 1020, 1030, and 1040 may be interconnected using a system bus 1050.
[0184] Processor 1010 may be configured to execute instructions within system 1000. Processor 1010 may include a single-threaded processor or a multi-threaded processor. In some embodiments, the one or more processors 1000 may include one or more graphics processing units. Processor 1010 may be configured to execute or otherwise process instructions stored in one or both of memory 1020 or storage device 1030. Executing the instructions(s) may cause graphical information to be displayed or otherwise presented via a user interface on I / O device 1040. In some embodiments, the one or more processors may include one or more graphics processing units.
[0185] The memory 1020 may store information within the system 1000. In some embodiments, the memory 1020 is a computer-readable medium. In some embodiments, the memory 1020 may include one or more volatile memory cells. In some embodiments, the memory 1020 may include one or more non-volatile memory cells.
[0186] Storage device 1030 can be configured to provide mass storage for system 1000. In some embodiments, storage device 1030 is a computer-readable medium. Storage device 1030 may include floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, or other types of storage devices. I / O device 1040 can provide I / O operations for system 1000. In some embodiments, I / O device 1040 may include a keyboard, pointing device, or other devices for data input. In some embodiments, I / O device 1040 may include output devices, such as a display unit for displaying a graphical user interface or other types of user interfaces.
[0187] Figure 11This is a logic diagram based on some embodiments of the present disclosure, depicting a process 1100 of accessing clinical trial data on a host device from a client device 1101 using an application programming interface (API) 1104. This process can be used to obtain information and data from a clinical trial database to generate a combined standardized clinical trial database. Process 1100 illustrates the exchange of credentials between a client 1101 (such as, for example, computing device 205) and the API 1104 hosted on a server 203. The client 1101 can request an access token from an identity service 1102 that stores access tokens by executing instructions according to request access token 1111. The client 1101 can send credentials to the identity service 1102, which include a client ID, password, authorization type, and scope. Upon approval of the credentials submitted by the client 1101, the execution of instructions returns an access token 1112, and the identity service 1102 returns the access token as a response. The identity service 1102 may be hosted on a server provided by a third-party provider. The client 1101 can use the access token in an API request 1113 when attempting to request data from the API 1104. API 1104 can return the requested data 1114. Process 1100 is executed by computing device 205 in response to user input of credentials used to request data from server 203.
[0188] Figure 12 This is a table of clinical trial status organized according to different tiers, based on some embodiments of this disclosure. Table 1200 illustrates how different clinical trials are organized according to their status. Table 1200 includes multiple tiers, including a first tier (tier 1) for clinical trials that have been completed and include final results (such as outcome data included in outcome 203d). Table 1200 also includes a second tier (tier 2) for clinical trials that have been terminated or are ongoing but have no final results. Table 1200 also includes a third tier (tier 3) for clinical trials that have no results at all. And Table 1200 may include a fourth tier (tier 4), which may be an overarching category for clinical trials whose status is unknown. The endpoint associated with each tier may be the efficacy or safety of a drug administered throughout the clinical trial.
[0189] Because Tier 1 includes the end results associated with the clinical trial, the data associated with the end results can be used to construct the endpoints of the clinical trial. Users can use combinations of different tiers to design future clinical trials based on data recorded across different tiers. Table 1200 also includes the percentage of all clinical trials categorized into one of the different tiers, as well as the percentage of clinical trials sponsored or conducted by a certain number of pharmaceutical companies.
[0190] Figure 13This is a logic diagram based on some embodiments of the present disclosure, depicting a process of extracting clinical trial data from one or more external databases to generate a combined standardized clinical trial database and using clinical trial data from this combined standardized clinical trial database as input to a trial simulator. Process 1300 may begin with server 203 extracting or accessing data stored in clinical trial databases 201 and 202, logically represented as external information 1215. The data stored in external information 1215 is data collected by a third party and related to one or more clinical trials. External information 1215 may be standardized, combined, or integrated and stored in the combined standardized clinical trial database 1211 in the form of one or more of the aforementioned different datasets (e.g., trial summary 203a, trial group 203b, baseline 203c, outcome 203d, outcome comparison 203e, and / or regulatory information 203f). This combined standardized clinical trial database may also be referred to herein as a quantitative knowledge database.
[0191] In some embodiments, internal information 1216 (e.g., non-disclosure information or information generated within the organization) may be data associated with one or more clinical trials that a user requesting access to the combined standardized clinical trial database 1211 has access to but is not included in external information 1215. Server 203 may extract or access internal information 1216 in response to a request from computing device 205 to the server storing internal information 1216 for access to server 203. Once external information 1215 and internal information 1216 are combined and integrated, clinical trial database 1211 applies study population specification 1201 to the data stored in the combined clinical trial database and stores it in study-specific database 1210. This study-specific database and control group history database may be project-specific databases compiled from the original combined standardized clinical trial database. For example, an oncology chemotherapy dataset may be further extracted from the combined standardized clinical trial database and compiled into trials focusing on chemotherapy treatment. The established chemotherapy dataset can be used for clinical trial design and decision-making, with the chemotherapy dataset serving as a control group. An example of a study population may be patients with metastatic non-small cell lung cancer receiving first-line treatment including chemotherapy. For example, data stored in a research-specific database 1210 (which may be stored on server 203 or another server or local storage) can be used to study a treatment ORR->PFS prediction model 1202 in order to generate a meta-analysis model 1213. As an example, it is possible to... Figure 16The process 1600 generates a tree-based meta-analysis model 1213. Database 1210 stores treatments, treatment outcomes, and study design and baseline factors specific to the disease and population of interest, enabling the development of predictive models of interest. One or more control group hypotheses, such as chemotherapy in first-line metastatic non-small cell lung cancer patients, can be applied to data stored in the study-specific database 1210, and the resulting data can be stored in the control group historical study database 1212. For example, the control group hypothesis could be that chemotherapy will be used in first-line metastatic non-small cell lung cancer patients. Data stored in the control group historical study database 1212 and the meta-analysis linear model 1213 can be used to simulate the outcomes and success probabilities of ongoing clinical trials or user-designed de novo clinical trials. Database 1212 may also be derived from 1211, but with a focus on the control group of interest, namely chemotherapy in first-line metastatic non-small cell lung cancer patients.
[0192] In some embodiments, a customized user interface (e.g., a GUI) can be implemented to access and analyze data from a combined standardized clinical trial database. In some embodiments, the GUI is customized based on the therapeutic area of interest and treatment needs. For example, an R-shiny-based interface for oncology was implemented, and... Figure 14 As shown in the figure. In some embodiments, this GUI enables non-modeling personnel to interactively explore the database and test different modeling and simulation scenarios through interface design. Figure 14 These are screenshots of a clinical trial simulator graphical user interface according to some embodiments of this disclosure, depicting the regression of progression-free survival on overall response rate. Various trial simulators can be created based on user needs; this example simulation is used to predict later trial outcomes by inputting early trial outcomes (i.e., proof-of-concept (POC) studies) using a meta-analysis model established after POC study results are available (post-POC). Additional predictive modeling can be performed before the POC study (pre-POC). Additional modeling (sensitivity analysis) can also be performed at each stage to ensure the robustness of analyses based on relevant data in the database. A trial simulation GUI interface with a GUI for accessing data can be implemented. In some embodiments, this GUI can be designed based on user needs or expectations.
[0193] Figure 15Screenshot 1500 of a graphical user interface for a clinical trial simulator according to some embodiments of this disclosure depicts the probability of a drug treatment used in a clinical trial progressing from one development phase to the next or to approval for the total number of progression-free survival events, and the probability of a drug treatment used in a clinical trial progressing from one development phase to the next or to approval for a given event size. Users can use the simulator to predict the success of the next trial and optimize the next trial design.
[0194] Some embodiments provide estimates of drug approval probabilities based at least in part on structured datasets in a unified clinical trial database. In some embodiments, machine learning methods can be used to determine the estimates of drug approval probabilities. For example, the method described by Lo et al. in "Machine Learning with Statistical Imputation for Predicting Drug Approvals," published on October 4, 2019, in Harvard Data Science Review, Vol. 1.1, pp. 1–42, is analyzed and is incorporated herein by reference in its entirety.
[0195] Lo's publications describe predictive models for assessing the probability of approval for drug candidates in two scenarios: after Phase II trials and after Phase III trials. Lo's publications describe the use of statistical imputation to address missing data issues, such as k-nearest neighbor (kNN) (e.g., 5NN) imputation. Lo describes the use of machine learning techniques to generate predictions, including cross-validation for training and a retained test set for performance evaluation, and the use of the standard Area Under the Receiver Operating Characteristic (AUC) metric to measure model performance. AUC is an estimated probability that a classifier will rank a positive outcome above a negative outcome. Lo's publications describe the use of a random forest (RF) classifier model to predict approval.
[0196] Some embodiments include methods and applications for analyzing structured datasets based on a unified clinical trial database. Figure 16 According to some embodiments of this disclosure, an overview of a tree-based regression analysis model for predicting progression-free survival (PFS) associated with treatment of a specific disease based on the overall response rate (ORR) of patients in clinical trials is illustrated. Additional descriptions of exemplary tree-based regression analyses are provided below. PFS can be determined based on computer trial simulations.
[0197] A unified clinical trial database for cancer was constructed based on clinical trial information obtained from the AACT, Trialtrove, and BMT databases. A graphical user interface was generated for analysis based on structured datasets from the unified clinical trial database (see...). Figure 14 (GUI in the text).
[0198] Analysis was performed using structured datasets to predict progression-free survival (PFS) based on proof-of-concept (POC) objective response rates (ORR) of PD-1 checkpoint response inhibitors. Only intervention trials initiated by industry pursuants after 2010 were included. Late-stage PD-1-related clinical trials enrolling advanced patients were included. Trials were categorized into the following treatment classes: PD-1 monotherapy, chemotherapy monotherapy, other targeted therapy monotherapy, PD-1 in combination with chemotherapy, PD-1 in combination with other targeted therapies, chemotherapy in combination with other targeted therapies, and PD-1 in combination with chemotherapy and other targeted therapies. Only trials with paired ORR / PFS were included. Only ORR / PFS in the primary analysis population (PAS) were included. This analysis included approximately 200 groups with paired ORR / PFS. Key variables included: ORR; line number: 1L / 1L+ and 2L / 2L+; stage: III / IV and IV; PAS PDL1 status: positive and unspecified; cancer type: melanoma, NSCLC, and others; and treatment class: PD-1 monotherapy, chemotherapy monotherapy, other targeted therapy monotherapy, PD-1 + chemotherapy, PD-1 + other targeted therapy, chemotherapy + other targeted therapy, and PD-1 + chemotherapy + other targeted therapy. It was assumed that the baseline population / ORR was the same for both Phase II and Phase III treatments.
[0199] Figure 16A tree-based model of ORR-based PFS is schematically depicted according to some embodiments of this disclosure, including a graph of Log(ORR) versus Log(PFS) for each split. The first split in the tree-based regression analysis is performed by line number, where the ORR / PFS slope exhibits the greatest heterogeneity across line number compared to stage, cancer category, and PAS PDL1 status. The second split in the tree-based regression is performed by treatment category, where within “1L / 1L+”, the ORR / PFS slope exhibits the greatest heterogeneity across treatment category (chemotherapy-related and non-chemotherapy-related) compared to stage, cancer category, and PAS PDL1 status. The third split in the tree-based regression analysis is performed by stage, where within “1L / 1L+ and non-chemotherapy-related”, the ORR / PFS slope exhibits the greatest heterogeneity across stage compared to cancer category and PAS PDL1 status. Splitting is stopped if no clinically significant ORR / slope heterogeneity is found, or if the sample size of the subgroups after splitting is less than or equal to 5. Each subsequent prediction model after the split is Log(PFS) ~ Log(ORR) + U, where U is a variable with a significant intercept. The final tree-based regression model can be compared with other regression-based models to evaluate and compare model performance (e.g., the mean squared error (MSE) of model predictions in the training and testing split settings).
[0200] The features described herein can be implemented in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The apparatus can be implemented in a computer program product tangibly embodied in an information carrier (e.g., in a machine-readable storage device) for execution by a programmable processor; and the method steps can be executed by the programmable processor executing an instruction program to perform the functions of the described embodiments by manipulating input data and generating output. The described features can advantageously be implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive and transmit data and instructions from a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform an activity or produce a result. Computer programs can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0201] Suitable processors for executing instruction programs include, by way of example, general-purpose and special-purpose microprocessors, as well as a single processor or one of multiple processors in any kind of computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The components of a computer may include a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer may also include one or more mass storage devices for storing data files, or operatively coupled to and communicating with them; such devices include disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Suitable storage devices for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into application-specific integrated circuits (ASICs).
[0202] To provide interaction with the user, these features can be implemented on a computer with a display device for displaying information to the user (such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor), and a keyboard and pointing device (such as a mouse or trackball) through which the user can provide input to the computer.
[0203] These features can be implemented in computer systems that include back-end components (such as data servers), or middleware components (such as application servers or internet servers), or front-end components (such as client computers with graphical user interfaces or internet browsers), or any combination thereof. The components of the system can be connected via digital data communication of any form or medium, such as communication networks. Examples of communication networks include, for example, local area networks (LANs), wide area networks (WANs), and the computers and networks that form the Internet.
[0204] Computer systems can include clients and servers. Clients and servers are typically geographically separated and usually interact through networks such as the network described. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other.
[0205] Furthermore, the logical flow depicted in the accompanying drawings does not require the specific order or sequential arrangement shown to achieve the desired result. Additionally, other steps can be provided from the described flow, or steps can be eliminated, and other components can be added to or removed from the described system. Example: Using objective response rate to develop PD1 / PDL1 combination therapies to predict progression-free survival and overall survival.
[0206] In the development of PD1 / PDL1 combination therapies for solid tumors, objective response rate (ORR) is a commonly used clinical endpoint in early-stage studies, while progression-free survival (PFS) and overall survival (OS) are widely used in later-stage studies. Developing predictive models for median PFS (mPFS) and median OS (mOS) using early clinical outcomes such as ORR can inform the optimization of later-stage trial design and the assessment of probability of success (POS). Existing literature includes studies on the association and alternatives between ORR / mPFS / mOS, but the number of included clinical trials is limited. In this example, instead of establishing alternatives, we predict mPFS and mOS based on early efficacy using ORR and optimize later-stage trial designs for PD1 / PDL1 combination therapy development. To include a sufficient number of qualified clinical trials, a comprehensive quantitative clinical trial panorama database (QLD), also known as a unified clinical trial database, was constructed by combining information from various sources (such as clinicaltrial.gov, publications, and company press releases) regarding relevant indications and therapies. A generalizable algorithm was developed to systematically extract structured data, which was then manually processed for scientific accuracy and completeness. Ultimately, over 150 late-stage clinical trials were identified for ORR / mPFS (ORR / mOS) prediction model development, compared to a maximum of 40 in existing literature. A tree-based machine learning regression model was derived to simultaneously leverage and explain the heterogeneity of the ORR / mPFS (ORR / mOS) relationship, ensuring the prediction model is robust and has a clear clinical interpretive structure. Through over 1000 cross-validations, the proposed model's mean squared error is comparable to that of random forests and limiting gradient boosting, and through cross-validation, it significantly outperforms commonly used additive or interactive regression models. An example is provided illustrating the application of the proposed ORR / mPFS (ORR / mOS) prediction model in the POS assessment of late-stage trials of PD1 / PDL1 combination therapy. 1. Introduction
[0207] Anti-PD1 / anti-PDL1 therapies, such as pembrolizumab (FDA approved in 2014), nivolumab (FDA approved in 2014), and atezolizumab (FDA approved in 2016), have been developed over the past decade for the treatment of various cancers, such as NSCLC and melanoma. In recent years, PD1 / PDL1 combination therapies, such as PD1 / PDL1 combined with other targeted therapies, have become increasingly popular due to better clinical efficacy. For clinical trials of PD1 / PDL1 combination therapies in solid tumors, ORR is a commonly used clinical endpoint in early-stage studies, while PFS / OS is widely used in later-stage studies. ORR is an attractive early clinical endpoint because it allows for shorter trial durations, smaller patient cohorts, and even single-arm trial designs (Clarke JM, Wang X., and Ready NE (2015). Surrogate clinical endpoints to predict overall survival in non-small cell lung cancer trials – are we in a new era? [Translational Lung Cancer Research], 4, 804–808.). The relevance or surrogate nature of ORR to PFS / OS remains a core consideration in determining its effectiveness; therefore, various studies have investigated the association between ORR / mPFS / mOS. For example, Zhu et al. (Zhu Andrew X, Lin Yong, Ferry David Raymond, Widau Ryan C, and Saha Abhijoy (2022). Meta-analysis of surrogate endpoints for survival in patients with unresectable hepatocellular carcinoma treated with immune checkpoint inhibitor-based regimens).The American Society of Clinical Oncology (ACS) has investigated the association between ORR / mPFS / mOS of immune checkpoint inhibitors, as well as the work of Ye et al. (Ye Jiabu, Ji Xiang, Dennis Phillip A, Abdullah Hesham, and Mukhopadhyay Pralay (2020). Relationship between progression-free survival, objective response rate, and overall survival in clinical trials of PD-1 / PD-L1 immune checkpoint blockade: a meta-analysis [PD-1 / PD-L1 immune checkpoint blockade clinical trials: a meta-analysis]. Clinical Pharmacology and Therapeutics, 108(6), 1274–1288), and Nie et al. (Nie Runcong, Chen Foping, Yuan Shuqiang, Luo Yingshan, Chen Shi, Chen Yongming, Chen Xiaojiang, Chen Yingbo, Li Yuanfang, and Zhou Zhiwei (2019)). Evaluation of objective response, disease control, and progression-free survival as surrogate endpoints for overall survival in anti-programmed death-1 and anti-programmed death-ligand 1 trials.European Journal of Cancer, 106, 1–11) and Ito et al. (Ito Kentaro, Miura Satoru, Sakaguchi Tadashi, Murotani Kenta, Horita Nobuyuki, Akamatsu Hiroaki, Uemura Kohei, Morita Satoshi and Yamamoto Nobuyuki (2019). The impact of high PD-L1 expression on the surrogate endpoints and clinical outcomes of anti-PD-1 / PD-L1 antibodies in non-small cell lung cancer. Lung Cancer, 128, 113–119) investigated the ORR / mPFS / mOS relationship of PD1 / PDL1-related therapies. Blumenthal et al. (Blumenthal Gideon M, Karuri Stella W, Zhang Hui, Zhang Lijun, Khozin Sean, Kazandjian Dickran, Tang Shenghui, Sridhara Rajeshwari, Keegan Patricia, and Pazdur Richard (2015). Overall response rate, progression-free survival, and overall survival with targeted and standard therapies in advanced non-small-cell lung cancer: US Food and Drug Administration trial-level and patient-level analyses.Journal of Clinical Oncology, 33(9), 1008; Goring et al. (Goring Sarah, Varol Nebibe, Waser Nathalie, Popoff Evan, Lozano-Ortega Greta, Lee Adam, Yuan Yong, Eccles Laura, TranPhuong and Penrod John R. (2022). Correlations between objective response rate and survival-based endpoints in first-line advanced non-small cell lung cancer: a systematic review and meta-analysis [Lung Cancer, 170, 122–132] and Hua et al. (Hua Tiantian, Gao Yuan, Zhang Ruyang, Wei Yongyue and Chen Feng (2022). Validating ORR and PFS as surrogate endpoints in phase II and III clinical trials for NSCLC patients: difference exists in the strength of surrogacy in various trials [Settings [Validating ORR and PFS as surrogate endpoints in phase II and III clinical trials for NSCLC patients: surrogate strength varies across trial settings].BMC Cancer, 22(1), 1–13, investigated the ORR / mPFS / mOS relationship in NSCLC. However, all existing literature focusing on the ORR / mPFS / mOS association contains conflicting results. Some tumor types or therapies show ORR as a substitute for PFS or OS, but some studies have failed to establish this relationship. One reason may be the inclusion of only a limited number of clinical studies and the heterogeneity of relationships. Therefore, predictive models have not yet been clearly developed for all tumor types or all therapies. In this example, instead of establishing endpoint substitution, a sufficient number of clinical trials across various tumor types, treatment categories, etc., were integrated to develop a comprehensive predictive model for PD1 / PDL1 combination therapy. Therefore, the efficacy observed in earlier studies can be used to better predict the success of later studies or to better plan later studies.
[0208] To include a sufficient number of clinical trials to capture ORR / mPFS (ORR / mOS) heterogeneity in predictive model development, a comprehensive quantitative panorama database (QLD) was derived by including all PD1 / PDL1, NSCLC, and melanoma-related clinical trials from clinicaltrial.gov and Informa. All PD1 / PDL1-related clinical trials were included because the inventors aimed to focus on building predictive models for PD1 / PDL1 combination therapy development. All NSCLC and melanoma-related clinical trials were included to enrich the group of treatment categories less common in PD1 / PDL1-related trials (chemotherapy, other targeted therapies, and chemotherapy + other targeted therapies, which are systematically defined in Section 2.2), thereby facilitating predictive model development. NSCLC and melanoma were chosen as the two most common solid tumors because they have more clinical trials than other tumor types. All PD1 / PDL1, NSCLC, and melanoma-related trials available on clinicaltrial.gov and Informa up to July, September, and December 2022 were included to generate the QLD. Specifically, the inventors first utilized clinicaltrial.gov and its associated clinical studies and sponsor-uploaded study efficacy outcomes. Furthermore, Informa was subsequently integrated, providing supplementary data from publications, conference abstracts and presentations, and company press releases.
[0209] Next, all eligible clinical trials from QLD were selected, and a tree-based regression prediction model for the development of PD1 / PDL1 combination therapies was constructed. A key idea of tree-based regression machine learning models is to forward partition the predictor variable space into multiple simple subspaces until ORR / mPFS (ORR / mOS) heterogeneity is eliminated. Compared to existing simple prediction models (such as linear regression models), the proposed method is more flexible and capable of explaining ORR / mPFS (ORR / mOS) heterogeneity. Compared to existing machine learning prediction models (such as random forests and extreme gradient boosting models), the proposed method has a well-defined structure for clinical interpretation. The predictive performance of the proposed method was compared with variable-selective multiple linear regression models with / without interaction, random forests, and extreme gradient boosting. An example of the application of the developed tree-based machine learning prediction model in late-stage trial POS assessment is provided. Furthermore, the main application in this example is predicting late-stage mPFS / mOS based on early ORR to provide information for late-stage trial POS. Late-stage studies with both ORR and mPFS / mOS were selected within the same study. 2. Methods 2.1QLD
[0210] One purpose of creating the QLD (also known as the Unified Clinical Trials Database) is to organize information related to historical clinical studies into an indication- or mechanism-of-action (MOA)-specific database. The Unified Clinical Trials Database (QLD) is used to generate insights to prioritize development plans, optimize and mitigate project and study risks, and ultimately facilitate pivotal exploration and timely decision-making. A key focus is the analysis of efficacy and safety characteristics and their correlation with success probability modeling and influencing factor identification. Current input data sources are clinicaltrial.gov and Informa. Specifically, in this example, all PD1 / PDL1, NSCLC, and melanoma-related clinical trials with structured information from clinicaltrial.gov and Informa were included to generate the QLD. The QLD contains at least three types of information, including trial summaries, outcomes, and outcome comparisons.
[0211] Trial Summary: This dataset includes general research-level attributes such as study ID, study name, period, number of groups, study start / end date, study design, cancer type, number of lines of treatment, cancer stage, primary study drug / drug class / drug subclass, and other related drugs / drug classes / drug subclasses (either combination drug components or control drugs).
[0212] Outcomes:This dataset contains reported clinical outcomes for each study group. For each outcome record, the following information is provided: study ID, study name, group title, drug / drug combination, outcome title and description, outcome value, and number of patients. The corresponding drug / drug combination categories / subcategories and endpoint categories (e.g., ORR, PFS, OS, etc.) are compiled using a combination of algorithms and manual processing.
[0213] Outcome Comparison: This dataset contains outcome comparisons between treatment groups. For each outcome comparison record, the study ID, study name, outcome comparison p-value, confidence interval, statistical test method, and primary and comparative drugs / drug combinations are provided. The corresponding primary and comparative drug / drug combination categories / subcategories, endpoint categories (e.g., ORR, PFS, OS, etc.), and comparison categories (e.g., difference, odds ratio, hazard ratio, etc.) are compiled by a combination of algorithms and manual analysis.
[0214] In this example, information from the trial summary and outcomes is used to develop a predictive model. 2.2 Data Selection and Predictor Variables
[0215] The following inclusion and exclusion criteria were used to screen eligible trial groups from the QLD in order to develop an ORR / mPFS (ORR / mOS) predictive model for PD1 / PDL1 combination therapy: (1) All clinical trials initiated after 2010 were included due to more recent and prevalent mechanisms of action (MOA). (2) Unknown or early-stage (Phase I, I / II, Phase II) clinical trials were excluded due to potentially immature PFS and OS. (3) Trials involving patients with non-advanced cancer (Phase I / II) were excluded. (4) Trials without matched ORR / mPFS (ORR / mOS) were excluded.
[0216] For each included trial group, only the primary analysis population with paired ORR / mPFS and ORR / mOS was included in the predictive model development. The primary analysis population was chosen because it is the most important indicator affecting trial POS and study design. In addition to ORR, mPFS, and mOS, treatment category, indication, line of treatment, cancer stage, and PD1 / PDL1 status were extracted for each included trial group in the primary analysis population. Note that all information is either at the study level or at the treatment group level; no patient-level data were available. Treatment category was defined as one of the following seven categories: chemotherapy, PD1 / PDL1 therapy, other targeted therapies, PD1 / PDL1 + chemotherapy combination therapy, PD1 / PDL1 + other targeted therapy combination therapy, other targeted therapy + chemotherapy combination therapy, and PD1 / PDL1 + other targeted therapy + chemotherapy combination therapy, where other targeted therapies refer to all targeted therapies other than PD1 / PDL1 therapy. All targeted therapies other than PD1 / PDL1 therapies were grouped into Other Targeted Therapies because the primary objective in this example was the development of PD1 / PDL1 combination therapies; therefore, the number of included groups for each targeted therapy category other than PD1 / PDL1 was relatively small. Therefore, it was realistic to merge them and assume they shared the same ORR / mPFS (ORR / mOS) relationship. Indications were grouped into three categories: NSCLC, melanoma, and others (Hua et al. 2022; Nie et al. 2019). Treatment lines were grouped into three categories: 1L, 1L+, and remaining. Cancer stages were grouped into two categories: Stage IV only and Stage III / IV mixed. PD1 / PDL1 status was grouped into two categories: positive and unspecified. Remaining biomarker (excluding PD1 / PDL1) statuses from the primary analysis population were not extracted for analysis. This is because their associated therapies were not abundant enough, resulting in a relatively low prevalence of biomarker positivity in the analysis dataset. Furthermore, since the primary objective is the development of PD1 / PDL1-related therapies, the heterogeneity of ORR / mPFS (ORR / mOS) relationships across biomarker (excluding PD1 / PDL1) states is not addressed in this example. Therefore, it is assumed that ORR / mPFS (ORR / mOS) relationships across biomarker (excluding PD1 / PDL1) states are homogeneous. 2.3 Tree-based Machine Learning Regression Model
[0217] In this section, a tree-based machine learning regression model is developed to predict ORR to mPFS / mOS. In addition to ORR, treatment category, indication, number of lines of treatment, PD1 / PDL1 status, and cancer stage are also included as predictors to better explain the homogeneity / heterogeneity of ORR / mPFS (ORR / mOS). Compared to traditional linear prediction models (such as additive models), the proposed model explains the heterogeneous relationship of ORR / mPFS (ORR / mOS). Compared to modern black-box machine learning models (such as random forests and extreme gradient boosting), the proposed model has a clear structure for clinical interpretation. A key idea of the proposed model is to forward partition the predictor variable space into multiple simple regions until there is no significant heterogeneity in the ORR / mPFS (ORR / mOS) relationship. The proposed tree-based machine learning regression model is summarized in the following steps.
[0218] Step 1 (grouping of variable categories with homogenous slopes): Each variable The category is represented as The corresponding sample size decreases. By comparison and ,Evaluate Each of them The slope difference. If P-value > 0.25 or Then each and Merge. Repeat the above process for the remaining categories until there are no remaining categories. For each variable The categories after grouping are , denoted as categorical grouping variable .
[0219] Step 2 (splitting of grouped variable categories in Step 1 in the presence of significant heterogenous slopes): For satisfying Each variable By comparison and To assess the slope difference. Select according to The conditions for splitting into branches are: (1) ANOVA P value < 0.01, (2) ANOVA P value is the most significant among all candidate variables.
[0220] Step 3:
[0221] Scenario 1: If a split exists in step 2, then for The residual variable Θ on each branch Repeat steps 1 and 2.
[0222] Case 2 (Potential Heterogeneous Intercept Adjustment): If there is no split in step 2, then for those satisfying... Each variable By comparison and To assess the intercept effect, select all p-values < 0.01. The prediction model on the given branch is: .
[0223] Specifically, in steps 1 through 3, all regression models are weighted regressions, with the sample size of each experimental group used as the weight. In the development of the ORR / mPFS (ORR / mOS) prediction model, Θ includes treatment category, indication, number of lines of treatment, PD1 / PDL1 status, and cancer stage. 3. Results
[0224] For the ORR-to-mPFS prediction model, 154 trials (321 trial groups) were ultimately included. For the ORR-to-mOS prediction model, 123 trials (249 trial groups) were ultimately included. The number of late-stage clinical trials included is significantly greater than that in existing literature (maximum of 40 included) (Ye et al., 2020). This demonstrates the comprehensiveness of QLD and ensures the generalizability and reproducibility of the developed prediction models. The distribution of each predictor variable is summarized in Table 1. For continuous variables (such as ORR), the mean and standard deviation are summarized. For categorical variables (such as treatment category, indication, number of lines of treatment, PD1 / PDL1 status, and cancer stage), the number and proportion of each category are provided. Table 1. Distribution of predictor variables in the included experimental groups Dataset 1 is used to analyze ORR / mPFS, and Dataset 2 is used to analyze ORR / mOS.
[0225] The tree-based machine learning prediction model for ORR / mPFS is summarized in Figure 18 Specifically, for treatment categories, chemotherapy, other targeted therapies plus chemotherapy, and PD1 / PDL1 plus other targeted therapies plus chemotherapy are grouped together (treatment category group 1) because the ORR / mPFS relationship is homogeneous within these three treatment categories. Similarly, PD1 / PDL1 therapy and PD1 / PDL1 plus other targeted therapies are grouped together (treatment category group 2). Figure 19A ).(Note: Figure 19AIn the diagram, C represents chemotherapy, P represents PD1 / PDL1 therapy, and T represents other targeted therapies. The ORR / mPFS relationships among treatment category group 1, treatment category group 2, other targeted therapies, and PD1 / PDL1 + chemotherapy combination are heterogeneous and are preserved as branches of a tree. This indicates that the same ORR may lead to different mPFS in different treatment categories. Furthermore, within treatment category group 1 and treatment category group 2, different treatment lines lead to heterogeneous ORR / mPFS relationships. Figure 19B and Figure 19C Furthermore, within the 1L / 1L+ range of treatment category group 1 and the remaining treatment lines of treatment category group 2, different indications resulted in heterogeneous ORR / mPFS relationships. Figure 19D and Figure 19E Tree-based regression methods can leverage the ORR / mPFS (ORR / mOS) relationship when homogeneity is identified among influencing factors such as treatment categories, tumor types, and number of treatment lines.
[0226] A summary of tree-based machine learning prediction models for ORR / mOS is provided below. Figure 20 Specifically, due to the homogeneity of ORR / mOS relationships, chemotherapy, chemotherapy + other targeted therapies, and PD1 / PDL1 + other targeted therapies + chemotherapy are grouped together (treatment category group 1). Similarly, other targeted therapies, PD1 / PDL1 + chemotherapy, and PD1 / PDL1 + other targeted therapies are grouped together (treatment category group 2). Figure 21A ).(Note: Figure 21A In this context, C represents chemotherapy, P represents PD1 / PDL1 therapy, and T represents other targeted therapies. The ORR / mOS relationships among treatment category group 1, treatment category group 2, and PD1 / PDL1 therapy are heterogeneous. Furthermore, within treatment category group 2, different indications lead to heterogeneous ORR / mOS relationships. Figure 21B Within treatment category 2 NSCLC and melanoma, different lines of treatment resulted in heterogeneous ORR / mOS relationships ( Figure 21C ).
[0227] Next, the predictive performance of the proposed tree-based machine learning regression model was evaluated against other commonly used weighted regression models and black-box machine learning models, including M1: mPFS(mOS) ~ ORR, M2: mPFS(mOS) ~ ORR + treatment category + indication + number of treatment lines + PD1 / PDL1 status + cancer stage, M3: variables selected from M2, M4: mPFS(mOS) ~ ORR × (treatment category + indication + number of treatment lines + PD1 / PDL1 status + cancer stage), M5: variables selected from M4, M6: random forest, M7: extreme gradient boosting, and M8: our proposed tree-based machine learning regression model. Predictive performance was evaluated using 80%-20% cross-validation. The entire dataset was divided into 80% and 20% portions for 1000 iterations. In each iteration, 80% of the data was used to train the model, and 20% was used to test the mean squared error (MSE) of predictions. Specifically, for both Random Forest and Limiting Gradient Boosting, various combinations of hyperparameter tuning parameters were searched, and combinations with the minimum median mean squared error (MSE) in 1000 iterations were reported. For Random Forest, the following hyperparameter tuning parameters were searched: number of trees in the forest ∈ {500, 1000}, maximum number of nodes per tree ∈ {10, 20, 30}, and minimum number of observations for terminal nodes per tree ∈ {10, 20, 30}. For Limiting Gradient Boosting, the following hyperparameter tuning parameters were searched: maximum number of iterations ∈ {25, 50, 100}, maximum depth per tree ∈ {1, 2, 3, 4, 5}, and minimum number of observations for terminal nodes per tree ∈ {10, 20, 30}.
[0228] exist Figure 22 Clearly, regardless of whether ORR is used to predict mPFS or mOS, the predicted MSE of the tree-based machine learning model is comparable to that of M6 and M7 (black-box machine learning models) and significantly smaller than that of M1, M2, M3, M4, and M5 (regression models). This demonstrates the robustness of the model and structure developed through tree-based machine learning regression. Specifically, M1, M2, and M3 fail to account for ORR / mPFS (ORR / mOS) heterogeneity, M4 and M5 fail to adequately leverage it, and M6 and M7 lack clear clinical interpretation. Overall, compared to other competitors, the proposed tree-based machine learning regression model is robust, leverages and detects ORR / mPFS (ORR / mOS) heterogeneity, and possesses a clear structure for clinical interpretation. 4. Application
[0229] This section describes how to apply a developed tree-based machine learning prediction model for late-stage trial POS assessment. Two scenarios are considered: Scenario 1: A 1:1 randomized phase III clinical trial, with the treatment group receiving PD1 / PDL1 plus targeted therapy (excluding PD1 / PDL1), and the control group receiving Keytruda. The target population is previously untreated patients with PD1 / PDL1-positive advanced / metastatic NSCLC. Scenario 2: A 1:1 randomized phase III clinical trial, with the treatment group receiving PD1 / PDL1 plus chemotherapy, and the control group receiving chemotherapy. The target population is previously untreated patients with advanced / metastatic NSCLC. The sample size for each group is N. For simplicity, a censoring rate of 0% is assumed. For the treatment group, the phase II study is a single-arm study with a sample size of n. For Scenario 1, the median progression-free survival (mPFS) and median overall survival (mOS) of Keytruda are assumed to be 5.40 months (standard deviation 0.48 months) and 16.7 months (standard deviation 1.48 months), respectively (Mok et al., 2019). For scenario 2, we assume that the median progression-free survival (mPFS) and median overall survival (mOS) for chemotherapy are 5.00 months (standard deviation 0.48 months) and 13.0 months (standard deviation 1.07 months), respectively (Gogishvili et al., 2022). Furthermore, we assume that the survival curves for both the treatment and control groups follow an exponential distribution. The position of survival (POS) in phase III trials is assessed using PFS and OS as clinical endpoints, as shown below. Step 1: ORR was sampled from the treatment group in a phase II study. Step 2: A tree-based machine learning prediction model for ORR / mPFS (ORR / mOS) was established by sampling 100% of patients with replacement, and the mPFS (mOS) of the treatment group was predicted based on the ORR sampled in step 1. Step 3: The mPFS (mOS) of the control group was sampled. Step 4: The survival curve for each group is derived from the mPFS (mOS) sampled in steps 2 and 3. Step 5: Simulate each group of N patients from the survival curves in step 4, and derive the hazard ratio P-value using a proportional hazards model. Step 6: Repeat steps 1 through 5 a total of 5000 times to assess the proportion of simulation trials with a p-value < 0.05.
[0230] POS in Figures 23A to 23D The study summarized the sample sizes n and N for different phase II / III studies. For scenario 1, where OS was the endpoint, the POS was relatively small when the ORR was 40%. Figure 23AThis is because the predicted mean mOS is 18 months, similar to that of competitor Keytruda. In this case, the likelihood of proceeding to a Phase III study is low. For the remaining cases, as the sample sizes n and N of the Phase II / III studies increase, the POS increases due to the decrease in variance between the predicted mPFS (mOS) and the estimated hazard ratio in the simulated trial. Furthermore, the increase in POS decreases with increasing n and N. For example, for a Phase III study sample size N, the POS increases more when N increases from 100 to 300 compared to N increasing from 300 to 500. For a Phase II study sample size n, the POS increases more when n increases from 10 to 30 compared to n increasing from 30 to 50. When the sample size of the Phase II / III studies is sufficient, the POS cannot be significantly improved, which may help justify the expected sample size in the Phase II / III study design. 5. Discussion
[0231] In this example, a comprehensive QLD and tree-based machine learning model are built to enhance the ability to predict mPFS / mOS for PD1 / PDL1 combination therapy development via ORR. The QLD integrates information from various sources, such as clinicaltrial.gov and Informa, to include a sufficient number of historical clinical trials. A generalizable algorithm is proposed to systematically extract trial-level summaries and clinical outcomes from structured data from clinicaltrial.gov and Informa, supplemented by manual tidying to ensure scientific accuracy and completeness. Next, a tree-based machine learning regression model is developed to simultaneously leverage and explain the heterogeneity of the ORR / mPFS (ORR / mOS) relationship while possessing a well-defined structure for clinical interpretation. As determined through cross-validation, the model and structure developed by the proposed method are more robust than popular existing machine learning prediction models, such as random forests and extreme gradient boosting. This prediction model allows for better prediction of the success of later studies or better planning of later studies by leveraging ORR observed in earlier studies.
[0232] According to some embodiments, a potential extension of this example is to include historical clinical trials from other tumor types besides NSCLC and melanoma for ORR / mPFS (ORR / mOS) predictive model development. This can enrich each targeted therapy class and enable a systematic exploration of potential ORR / mPFS (ORR / mOS) heterogeneity between each targeted therapy class and the target biomarker status of the primary analysis population. Therefore, in some embodiments, predictive models derived from a wider range of tumors can support various late-stage study designs related to targeted therapies, and the embodiments are not limited to use in PD1 / PDL1 combination therapy development. Furthermore, depending on the availability of data sources (e.g., clinicaltrial.gov and Informa), some embodiments may in the future include other potential predictive variables besides treatment class, tumor type, number of lines of treatment, cancer stage, and biomarker status of the primary analysis population.
[0233] Several embodiments of this disclosure have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other embodiments are within the scope of the claims.
Claims
1. A method comprising: Information is extracted or accessed from one or more databases by at least one processor, including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; The processor generates multiple structured datasets of various data types based on the extracted information. These multiple structured datasets include a trial summary dataset, a trial group dataset, a baseline dataset, an outcome dataset, and an outcome comparison dataset. The at least one processor applies natural language processing to at least a portion of the extracted or accessed information, including unstructured text. The at least one processor generates additional data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets; as well as Provide a graphical user interface (GUI) to receive input from users about the analysis to be performed using the multiple structured datasets.
2. The method as described in claim 1, wherein, The multiple structured datasets also include a regulatory event dataset.
3. The method as claimed in claim 1 or claim 2, wherein, The extracted information includes at least a portion of unstructured data, including recent information that has not yet been obtained from the one or more databases as structured clinical trial information.
4. The method according to any one of claims 1 to 3, further comprising: The second information is extracted or accessed by at least one processor from at least one of the one or more databases, the second information including one or more of drug study design information, biomarker information, drug treatment information and efficacy information for at least one clinical trial; as well as One or both of the following: The at least one processor generates second data corresponding to at least one of the plurality of data types based on the second information, so as to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; as well as The at least one processor applies natural language processing to at least a portion of the unstructured text in the second information, and the at least one processor generates second data or revised data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets.
5. The method according to any one of claims 1 to 3, further comprising: At least one processor periodically retrieves or accesses supplementary information from at least one of the one or more databases, including one or more of the following: drug study design information for at least one clinical trial, biomarker information, drug treatment information, and efficacy information; and One or both of the following: The at least one processor generates additional data corresponding to at least one of the plurality of data types based on the additional information, to add to at least one of the plurality of datasets or update at least some data in at least one of the plurality of datasets; as well as The at least one processor applies natural language processing to at least a portion of the additional information, including unstructured text, and the at least one processor generates additional or revised data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets.
6. The method according to any one of claims 1 to 5, further comprising: The at least one processor executes a machine learning module to determine the probability of success in advancing a clinical drug trial to the next phase of clinical trials based on the multiple structured datasets.
7. The method of any one of claims 1 to 6, further comprising: The at least one processor executes a machine learning module to determine the probability of obtaining regulatory approval for the drug in the current clinical trial based on the multiple structured datasets.
8. The method of any one of claims 1 to 7, further comprising: The processor executes a machine learning module to determine the probability of obtaining regulatory approval for a drug in the current clinical trial, based on the multiple structured datasets, for each of the multiple regulatory pathways.
9. The method of any one of claims 1 to 8, further comprising: The at least one processor identifies one or more scientific, design, regulatory, or operational factors that contribute to the historical approval likelihood of at least some of the plurality of clinical drug trials.
10. The method of claim 9, wherein, The one or more scientific factors include the mechanism of action associated with each of the one or more clinical drug trials, the one or more design factors include the endpoints associated with each of the one or more clinical trials, and the one or more regulatory factors include designating the one or more clinical drug trials as breakthrough therapies.
11. The method of any one of claims 1 to 10, further comprising: Assign a trial status category to each clinical trial in the multiple structured datasets, where these trial status categories include one or more of the following: completed with final results, terminated or in progress but without final results, and no results.
12. The method of any one of claims 1 to 11, further comprising: The processor simulates one or more phase III clinical trial outcomes based on one or more proof-of-concept study outcomes.
13. The method according to any one of claims 1 to 12, wherein, The GUI includes: The first field is used to receive a first input including disease-related information, which specifies variables, combinations of variables of interest, or both; and The second field is used to identify the output to be generated.
14. The method of claim 13, wherein, The first input includes or identifies a disease-specific database or a subset thereof, which is at least partially based on standardized clinical trials associated with the disease.
15. The method of claim 13 or claim 14, wherein, The identified outputs to be generated include one or more of the following: a description of one or more first clinical trials associated with the disease; a characteristic or outcome associated with the one or more first clinical trials; or a prediction of the outcome of the one or more first clinical trials for a drug used to treat the disease.
16. The method according to any one of claims 1 to 15, wherein, The trial dataset includes data types used to correlate trial outcomes with treatment information.
17. The method of any one of claims 1 to 16, further comprising: Generate predictive models or computer simulations for predicting trial outcomes.
18. The method according to any one of claims 1 to 16, wherein, The extracted or accessed trial outcome information includes the endpoints associated with each of these clinical trials.
19. The method of claim 18, further comprising generating a predictive model or computer simulation for predicting these endpoints, wherein, The predictive model or computer simulation is at least partially based on a machine learning model.
20. A method for determining a predictive model for median progression-free survival (mPFS) or median objective survival (mOS) in a later-stage clinical trial, the predictive model being based on the objective response rate (ORR) in an earlier-stage clinical trial, the method comprising: Extract or access clinical trial information, including ORR, mPFS, mOS, treatment category, indication, line of treatment, disease stage, and biomarker status information for multiple clinical trials obtained from one or more databases; Define variable categories that include at least some or all of the variables based on ORR information, treatment category information, indication information, line of treatment information, disease stage information, and biomarker status information; as well as Determine the tree-based regression machine learning model for ORR-based mPFS or ORR-based mOS, this determination includes: Based on these variable categories, the predictor variable space is forward-segmented into multiple subspaces until there is no heterogeneity of ORR-based mPFS or until there is no heterogeneity of ORR-based mOS. as well as Determine the regression model for each subspace.
21. The method of claim 20, wherein, One or more of these treatment categories include anti-PD1 and / or anti-PLD1 therapy (PD1 / PDL1 therapy); and The biomarker status information includes PD1 and / or PDL1 status information.
22. The method of claim 20 or 21, wherein, Perform a forward partitioning of the predictor space into multiple subspaces until no ORR-based mPFS heterogeneity or no ORR-based mOS heterogeneity exists, including: Determine the slope for each variable category and group variable categories with homogeneous slopes to form grouped variable categories; and For each grouping variable category, split the grouping variable category if there is a significant heterogeneous slope within the variable categories in that grouping variable category.
23. The method of claim 22, wherein, Determine the slope for each variable category and group variable categories with homogeneous slopes to form grouped variable categories. These variable categories include: Determine the first variable category with the largest sample size compared to the sample sizes of other variable categories that are one or more second variable categories; The slope of each variable category is determined by comparing the following items: The first regression of the mPFS or mOS on the product of the ORR and the value of the categorical variable, and The mPFS or mOS information is used to perform a second regression on the ORR information plus the value of the categorical variable. For each of the one or more second variable categories: Determine the slope difference between the slope of the second variable category and the slope of the first variable category; and If the p-value of the slope difference between the variable categories is higher than a specified grouping threshold or the sample size of the second variable category is lower than a specified sample size threshold, then these variable categories are grouped.
24. The method of claim 23, further comprising performing the steps of claim 23 for each remaining ungrouped second variable category.
25. The method of claim 23 or 24, wherein, For each grouping variable category, splitting that grouping variable category involves the following if there is a significant heterogeneous slope within that category: For each variable with more than one grouped category, compare the following: The first regression of the mPFS or mOS on the product of the ORR and the value of the grouping categorical variable, and The mPFS or mOS information is used in a second regression of the ORR information plus the value of the grouping categorical variable; and Choose the grouping variable category to split if the ANOVA p-value is less than the specified splitting threshold and the ANOVA p-value is the most significant for that variable compared to all other variables.
26. The method of claim 25, further comprising, for each split grouping variable category, performing the following steps: Identify the first variable category with the largest sample size among the split grouping variable categories compared to the sample sizes of other variable categories in the split grouping variable category, wherein the other variable categories are one or more second variable categories; The slope of each variable category in the split grouping variable categories is determined by comparing the following: The first regression of the mPFS or mOS on the product of the ORR and the value of the categorical variable, and The mPFS or mOS information is used to perform a second regression on the ORR information plus the value of the categorical variable. as well as For each of the one or more second variable categories: Determine the slope difference between the slope of the second variable category and the slope of the first variable category; as well as If the p-value of the slope difference between the variable categories is higher than the specified grouping p-threshold or the sample size of the second variable category is lower than the specified sample size threshold, then these variable categories are grouped.
27. The method according to any one of claims 20 to 26, wherein, The extracted or accessed trial outcome information includes the endpoints associated with each of these clinical trials.
28. The method according to any one of claims 20 to 27, wherein, Treatment category information includes chemotherapy, a combination of chemotherapy and PD1 / PDL1, or a combination of two or more of chemotherapy, PD1 / PDL1, or targeted therapy.
29. A system comprising: A storage device configured to store a unified clinical trial database; A memory that stores one or more instructions; as well as At least one processor, operatively coupled to the storage device, wherein the at least one processor is configured or programmed to read one or more instructions stored in the memory, thereby causing the at least one processor to perform the following operations: Information is extracted or accessed from one or more databases by at least one processor, including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; The at least one processor generates multiple structured datasets comprising multiple data types based on at least a portion of the extracted information. These multiple structured datasets include a trial summary dataset, a trial group dataset, a baseline dataset, an outcome dataset, and an outcome comparison dataset. The at least one processor applies natural language processing to at least a portion of the extracted or accessed information, including unstructured text. The at least one processor generates additional data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets; The multiple structured datasets are stored in the storage device, wherein the unified clinical trial database includes the multiple structured datasets; and Provides a graphical user interface to receive input from users regarding the analysis to be performed using the unified clinical trial database.
30. A non-transitory computer-readable medium comprising instructions that, when executed by a processing device, cause the processing device to perform the following operations: Information is extracted or accessed from one or more databases by at least one processor, including drug study design information, biomarker information, drug treatment information, and efficacy information for multiple clinical trials; The at least one processor generates multiple structured datasets comprising multiple data types based on at least a portion of the extracted information. These multiple structured datasets include a trial summary dataset, a trial group dataset, a baseline dataset, an outcome dataset, and an outcome comparison dataset. The at least one processor applies natural language processing to at least a portion of the extracted or accessed information, including unstructured text. The at least one processor generates additional data corresponding to at least one of the plurality of data types based on the natural language processing, for use in at least some of the plurality of structured datasets; The multiple structured datasets are stored in the storage device, wherein the unified clinical trial database includes the multiple structured datasets; as well as Provides a graphical user interface to receive input from users regarding the analysis to be performed using the unified clinical trial database.