Artificial intelligence for identifying one or more predictive biomarkers
By combining AI-trained predictive models and xenograft mouse models with independent analysis of multiple data sources, the problem of identifying and validating predictive biomarkers for cancer treatment response was solved, improving prediction accuracy and reducing the probability of false drug target detection.
Patent Information
- Application Number
- CN202480026611.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-23
- Filing Date
- 2024-03-25
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to effectively identify and validate predictive biomarkers for cancer treatment response, especially given the systematic biases and errors among datasets from multiple sources, leading to erroneous drug target discovery.
By using artificial intelligence to train a predictive model, combined with xenograft mouse models and independent analysis of multiple data sources, cancer biomarkers are screened and validated. Learning algorithms such as random forest regressors and feature selection algorithms, as well as dimensionality reduction techniques such as UMAP and PCA, are used to reduce systematic bias and improve prediction accuracy.
It improves the accuracy of predicting cancer treatment response, reduces the probability of false drug target discovery, and provides more reliable biomarkers for personalized treatment decisions.
Smart Images

Figure CN121002198A_ABST
Abstract
Description
Background Technology
[0001] This disclosure relates to systems, software, and methods for identifying predictive biomarkers, particularly predictive biomarkers for cancer treatment response. Summary of the Invention
[0002] According to one aspect of this disclosure, a method is provided for using at least one hardware processor to train artificial intelligence to identify one or more predictive biomarkers for cancer treatment.
[0003] In some implementations, the method includes: training at least one learning algorithm in a prediction model using at least one cancer drug discovery dataset, wherein the trained prediction model is capable of defining the multiple drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and validating the prediction model using a xenograft mouse model, wherein human cancer tissue exhibiting a specific set of cancer biomarkers is implanted in parallel into multiple immunodeficient mice, and the multiple immunodeficient mice are subsequently treated in parallel with one of the multiple drugs or drug combinations, and the biological response to the treatment is fed back into the prediction model to further train the learning algorithm.
[0004] In some embodiments of the disclosed method, human cancer tissue exhibiting a specific set of cancer biomarkers is in situ implanted into multiple immunodeficient mice.
[0005] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0006] In some embodiments of the disclosed method, for each cancer category, at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
[0007] In some embodiments of the disclosed method, for each drug or drug combination under a cancer category, at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
[0008] In some embodiments of the disclosed method, cancer classification data includes information from cancer cell lines screened in vitro, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from groups comprised of cancer cell line identification, TCGA classification, tissue type, tissue subtype, or a combination thereof.
[0009] In some implementations of the disclosed method, patient information includes information selected from cancer type, patient tumor biomarkers, patient age, patient gender, patient weight, patient family history, patient past health status, and combinations thereof.
[0010] In some embodiments of the disclosed methods, cancer biomarker data include gene expression data. In some embodiments, the gene expression data is normalized using transcripts per million (TPM). In some embodiments, the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0011] In some embodiments of the disclosed methods, the cancer biomarker data includes gene methylation data. In some embodiments, the cancer biomarker data includes protein biomarker data.
[0012] In some embodiments of the disclosed method, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes substructural features of the drug. In some embodiments, the substructural features of the drug include descriptors from the SMILES specification of the drug.
[0013] In some embodiments of the disclosed methods, biological response data include in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof. In some embodiments, IC50, AUC, or a combination thereof are converted to a normal distribution, and the normalized score is reported as a standardized z-score or normalized and then scaled from 0 to 1 or from 0 to 100%.
[0014] In some embodiments of the disclosed methods, the biological response data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergistic scores obtained using a highest single-agent (HSA) model, a Loewe additivity model, a Bliss independent model, or a zero-interaction-efficacy (ZIP) model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe, HSA, Bliss, or ZIP synergistic scores. In some embodiments, the in vitro drug combination Loewe, HSA, Bliss, or ZIP synergistic scores are reported as standardized z-scores or standardized and then scaled from 0 to 1 or from 0% to 100%.
[0015] In some embodiments of the disclosed methods, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or combinations thereof. In some embodiments, efficacy assessments include tumor size measurements, tumor burden assessments, efficacy biomarker assessments, response assessments, survival assessments, or combinations thereof. In some embodiments, response assessments, survival assessments, or combinations thereof are converted and reported as scaled efficacy scores from 0 to 1 or from 0% to 100%.
[0016] In some embodiments of the disclosed methods, biological response data include xenograft drug monotherapy or drug combination pharmacology results. In some embodiments, xenograft pharmacology results include tumor size measurements, tumor growth inhibition (TGI), AUC, or combinations thereof. In some embodiments, the TGI or AUC is converted to a normal distribution and reported as a standardized z-score or scaled score from 0 to 1 or from 0% to 100%.
[0017] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from the Cancer Drug Sensitivity Genomics (GDSC).
[0018] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0019] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from NCI-ALMANAC.
[0020] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas project.
[0021] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0022] In some embodiments of the disclosed method, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0023] In some implementations of the disclosed method, at least one learning algorithm includes a regression algorithm, preferably a random forest regressor.
[0024] In some implementations of the disclosed method, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0025] In some embodiments of the disclosed method, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0026] In some implementations, the disclosed method also includes applying a dimensionality reduction algorithm, preferably UMAP, PCA, or both, to the prediction model.
[0027] In some implementations, the disclosed method also includes independently resampling data elements in each dataset.
[0028] According to another aspect of this disclosure, a system is provided, the system comprising: an input unit capable of downloading at least one cancer drug discovery dataset; and a training unit capable of training at least one learning algorithm in a predictive model using the at least one cancer drug discovery dataset, wherein the trained predictive model is capable of defining the plurality of drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and wherein the training unit is capable of further training the learning algorithm using data obtained from a xenograft mouse model, wherein in the xenograft mouse model, human cancer tissue exhibiting a specific set of cancer biomarkers is implanted in parallel into a plurality of immunodeficient mice, each of the plurality of immunodeficient mice being treated in parallel with one of the plurality of drugs or drug combinations, and a measurement of the biological response to the treatment is used as new training data for the predictive model.
[0029] In some implementations of the disclosed system, human cancer tissue exhibiting a specific set of cancer biomarkers is in situ implanted into multiple immunodeficient mice.
[0030] In some implementations of the disclosed system, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0031] In some implementations of the disclosed system, for each cancer category, at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
[0032] In some implementations of the disclosed system, for each drug or drug combination under a cancer category, at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
[0033] In some implementations of the disclosed system, cancer classification data includes information from cancer cell lines screened in vitro, patient information from clinical studies, or a combination thereof. In some implementations, cancer cell line information includes information selected from groups comprised of cancer cell line identification, TCGA classification, tissue type, tissue subtype, or a combination thereof.
[0034] In some implementations of the disclosed system, patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health status, and combinations thereof.
[0035] In some embodiments of the disclosed system, cancer biomarker data includes gene expression data. In some embodiments, the gene expression data is normalized using transcripts per million (TPM). In some embodiments, the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0036] In some embodiments of the disclosed system, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0037] In some embodiments of the disclosed system, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes substructural features of the drug. In some embodiments, the substructural features of the drug include descriptors from the SMILES specification of the drug.
[0038] In some embodiments of the disclosed system, biological response data include in vitro drug screening results. In some embodiments, in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof. In some embodiments, IC50, AUC, or a combination thereof are converted to a normal distribution, and the normalized score is reported as a standardized z-score or normalized, and then scaled from 0 to 1 or from 0 to 100%.
[0039] In some embodiments of the disclosed system, the bioresponse data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are presented as synergistic scores obtained using a highest single-agent (HSA) model, a Loewe additivity model, a Bliss independent model, or a zero-interaction-efficacy (ZIP) model. In some embodiments, the in vitro drug combination screening results are presented as Loewe synergistic scores, HSA scores, Bliss scores, or ZIP scores. In some embodiments, the in vitro drug combination Loewe, HSA, Bliss, or ZIP synergistic scores are reported as standardized z-scores or standardized and then scaled from 0 to 1 or from 0% to 100%.
[0040] In some implementations of the disclosed system, biological response data include clinical study results. In some implementations, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or combinations thereof. In some implementations, efficacy assessments include tumor size measurements, tumor burden assessments, efficacy biomarker assessments, response assessments, survival assessments, or combinations thereof. In some implementations, response assessments, survival assessments, or combinations thereof are converted and reported as scaled efficacy scores from 0 to 1 or from 0% to 100%.
[0041] In some embodiments of the disclosed system, biological response data include xenograft drug monotherapy or drug combination pharmacology results. In some embodiments, xenograft pharmacology results include tumor size measurements, tumor growth inhibition (TGI), AUC, or combinations thereof. In some embodiments, the TGI or AUC is converted to a normal distribution and reported as a standardized z-score or scaled score from 0 to 1 or from 0% to 100%.
[0042] In some implementations of the disclosed system, at least one cancer drug discovery dataset includes data from the Cancer Drug Sensitivity Genomics (GDSC).
[0043] In some implementations of the disclosed system, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0044] In some implementations of the disclosed system, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas (TCGA) project.
[0045] In some implementations of the disclosed system, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0046] In some embodiments of the disclosed system, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0047] In some implementations of the disclosed system, at least one learning algorithm includes a regression algorithm, preferably a random forest regressor.
[0048] In some implementations of the disclosed system, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0049] In some implementations of the disclosed system, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0050] In some implementations, the disclosed system also includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0051] In some implementations, the disclosed system also includes independently resampling data elements in each dataset.
[0052] According to another aspect of this disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a processor, the instructions cause the processor to: train at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the trained predictive model is capable of defining the plurality of drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and to train the predictive model using a xenograft mouse model, wherein human cancer tissue exhibiting a specific set of cancer biomarkers is implanted in parallel into a plurality of immunodeficient mice, and the plurality of immunodeficient mice subsequently each receive in parallel one of the plurality of drugs or drug combinations, and the biological response to the treatment is fed back to the predictive model to further train the learning algorithm.
[0053] In some embodiments of the disclosed computer-readable medium, human cancer tissue exhibiting a specific set of cancer biomarkers is in situ implanted into multiple immunodeficient mice.
[0054] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0055] In some embodiments of the disclosed computer-readable medium, for each cancer category, at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
[0056] In some embodiments of the disclosed computer-readable medium, for each drug or drug combination under a cancer category, at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
[0057] In some embodiments of the disclosed computer-readable medium, cancer classification data includes information from cancer cell lines screened in vitro, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from groups comprised of cancer cell line identification, TCGA classification, tissue type, tissue subtype, or combinations thereof.
[0058] In some embodiments of the disclosed computer-readable medium, patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health conditions, and combinations thereof.
[0059] In some embodiments of the disclosed computer-readable medium, cancer biomarker data includes gene expression data. In some embodiments, the gene expression data is normalized using transcripts per million (TPM). In some embodiments, the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0060] In some embodiments of the disclosed computer-readable medium, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0061] In some embodiments of the disclosed computer-readable medium, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes substructural features of the drug. In some embodiments, the substructural features of the drug include descriptors from the SMILES specification of the drug.
[0062] In some embodiments of the disclosed computer-readable medium, the biological response data includes in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of the drug response curve), or a combination thereof. In some embodiments, the IC50, AUC, or combination thereof are converted to a normal distribution, and the normalized score is reported as a standardized z-score or normalized and then scaled from 0 to 1 or from 0 to 100%.
[0063] In some embodiments of the disclosed computer-readable medium, the biological response data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergistic scores obtained using a highest single-agent (HSA) model, a Loewe additivity model, a Bliss independent model, or a zero-interaction-efficacy (ZIP) model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe, HSA, Bliss, or ZIP synergistic scores. In some embodiments, the in vitro drug combination Loewe, HSA, Bliss, or ZIP synergistic scores are reported as standardized z-scores or standardized and then scaled from 0 to 1 or from 0% to 100%.
[0064] In some embodiments of the disclosed computer-readable medium, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or combinations thereof. In some embodiments, efficacy assessments include tumor size measurements, tumor burden assessments, efficacy biomarker assessments, or combinations thereof. In some embodiments, response assessments, survival assessments, or combinations thereof are converted and reported as scaled efficacy scores from 0 to 1 or from 0% to 100%.
[0065] In some embodiments of the disclosed computer-readable medium, the biological response data includes xenograft drug monotherapy or drug combination pharmacology results. In some embodiments, xenograft pharmacology results include tumor size measurements, tumor growth inhibition (TGI), AUC, or combinations thereof. In some embodiments, the TGI or AUC is converted to a normal distribution and reported as a standardized z-score or scaled score from 0 to 1 or from 0% to 100%.
[0066] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from the Cancer Drug Sensitivity Genomics (GDSC) dataset.
[0067] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0068] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from NCI-ALMANAC.
[0069] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas Project (TCGA).
[0070] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0071] In some embodiments of the disclosed computer-readable medium, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0072] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0073] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a regression algorithm, preferably a random forest regressor.
[0074] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0075] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0076] In some implementations, the disclosed computer-readable medium also includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0077] In some implementations, the disclosed computer-readable medium also includes independently resampling data elements in each dataset.
[0078] In some implementations of the disclosed methods, systems, and non-transitory computer-readable media, the predictive model does not provide a cancer diagnosis.
[0079] In some implementations of the disclosed methods, systems, and nontransitory computer-readable media, the predictive model does not provide cancer prognosis.
[0080] In some implementations of the disclosed methods, systems, and nontransitory computer-readable media, the predictive model does not provide cancer diagnosis or prognosis. Attached Figure Description
[0081] The purpose and features of the present invention can be better understood by referring to the following detailed description and accompanying drawings.
[0082] Figure 1 This is a schematic diagram of a method, system, and non-transitory computer-readable medium for developing software to identify and validate predictive chemical molecular features and genetic biomarkers for cancer treatment, according to the present disclosure.
[0083] Figure 2 This is a schematic diagram detailing the algorithms used to create molecular chemical predictors and predictive genetic biomarkers for identifying cancer treatments. Training data for each cancer type (X) is processed separately to build a machine learning model to predict treatment methods for each cancer type.
[0084] Figure 3 The results of feature optimization are detailed, showing that only 5-10 genes are needed to achieve peak optimal prediction accuracy in both the A) colorectal cancer treatment prediction model and the B) lung cancer treatment prediction model.
[0085] Figure 4The accuracy results of our in vitro therapy prediction models are detailed, namely A) a prediction model for monotherapy and B) a prediction model for a combination of monotherapy and combination therapy.
[0086] Figure 5 This demonstrates the results of our machine learning model on the clinical treatment response of cancer patients.
[0087] Figure 6 This is an overview of the pharmacological experiments validating the PDX model, in which we demonstrate the drug response of the orthotopic PDX model for triple-negative breast cancer to four predicted AI therapies. Detailed Implementation
[0088] This disclosure provides a method, system, and software for identifying and validating predictive biomarkers that can predict a person's response to cancer treatment.
[0089] Genomic and proteomic analyses provide a wealth of information about the quantity and form of proteins expressed in cells and offer the possibility of identifying the expressed protein profile of a specific cellular state for each cell. In some cases, this cellular state may characterize abnormal physiological responses associated with disease. Therefore, identifying and comparing the cellular states of patients with disease with corresponding cellular states of healthy patients can provide opportunities for diagnosing and controlling disease treatments.
[0090] Recent advances in transcriptomics and proteomics analysis techniques have enabled the application of computational methods to detect changes in expression patterns and their association with disease conditions, thereby helping to identify biomarkers that may contribute to combinations of multiple biomarkers with highly accurate diagnostic performance.
[0091] While high-throughput screening methods provide abundant datasets of gene expression information, the challenge in bioinformatics remains developing robust methods to organize data into patterns that can reproducibly diagnose diverse individual populations. A commonly accepted approach is to pool data from multiple sources to form a combined dataset, then divide the dataset into discovery / training and test / validation sets. However, both transcriptional and protein expression profiling data are typically characterized by a large number of variables relative to the number of available samples.
[0092] Differences in observed expression profiles between patient and control groups are often masked by: (1) biological variations or unknown subphenotypes in the disease or control population; (2) site-specific bias due to differences in study protocols, sample handling, etc.; (3) bias due to differences in instrument conditions (e.g., chip batches); and / or (4) variations due to measurement errors. False discovery of drug targets remains a serious problem, especially considering the cost and effort typically required for post-discovery work such as protein / gene identification and further validation of potential biomarkers.
[0093] This disclosure recognizes that sometimes systematic bias due to site-specific factors can only be detected through careful analysis and comparison of data from multiple sources. This disclosure provides systems, software, and methods for analyzing expression profiling data from multiple sources (e.g., clinical trial sites) to overcome potential systematic biases in expression data that are typically generated in such analyses, thereby reducing the probability of false drug target discovery. In a preferred aspect, the invention will use bioinformatics and a combination of expression profiling from samples from multiple sources to screen, identify, and validate biomarkers for specific biological states or diseases of interest. Measurements of these biomarkers in patient samples can provide information about the presence, absence, or severity of a disease or trait in the patient, such as in humans. In one aspect, a disease or trait is the presence, predisposition, or risk of recurrence of a disease.
[0094] In some embodiments, this disclosure provides bioinformatics tools to analyze expression profiling data from samples from two or more independent sources, thereby reducing sources of variability and bias that could lead to the identification of false targets in drug discovery. In some embodiments of this disclosure, data from multiple sources are not pooled into a combined dataset and are then split into discovery / training sets and test / validation sets. In some embodiments, data from multiple sources (e.g., multiple different clinical trial sites) are analyzed separately and independently.
[0095] For each source, adequate sample size and statistical resampling methods (e.g., guided analysis) help identify biomarkers that perform well in representative populations and are consistent across different randomly selected subgroups. The use of resampling procedures reduces the combined effects of biological variability and the large number of variables in gene expression profiling data.
[0096] In some embodiments, this disclosure relates to developing at least two distinct learning sets (discovery datasets) developed independently of each other. Each learning set includes subject data (data points) from multiple subjects. The subject data from each subject indicates the phenotype (a category of biological state or a form of pathological state) to which the subject belongs, and each subject is classified into one of multiple different pathological categories. Different phenotypes are typically associated with pathology, such as diseased versus normal, different stages of disease, etc. However, they can include any measurable biological characteristic. Each learning set has subject data from at least two subjects belonging to each phenotype. The subject data from each subject includes measurements of multiple data elements from each subject's sample.
[0097] In some implementations, the results of individual and independent analyses are cross-compared to identify a subset of potential biomarkers that have comparable performance levels on data from each individual source and exhibit the same up / down regulation pattern across different sample groups across multiple data sources.
[0098] The biomarkers selected from the cross-comparisons are then used to develop a multivariate classification model to classify the sample (e.g., cancerous tissue from a patient) into one of the biological state categories or diseases. Preferably, another independent validation dataset is used to further validate this subset of potential biomarkers. Furthermore, additional samples and other methods (e.g., including but not limited to immunoassays) are preferably used to identify these potential biomarkers and validate their performance.
[0099] In a preferred aspect, the expression profiling data evaluated is proteomic profiling data (i.e., data relating to the expression of proteins and their modified and processed forms). For example, the method is particularly suitable for mass spectrometry-based proteomic analysis. Therefore, in one aspect, the AI method disclosed herein is used to screen, identify, and validate predictive biomarkers for cancer treatment. Data from independent datasets (e.g., the types of expressed biomarkers, the expression levels of each biomarker) are cross-compared to identify those biomarkers that represent one or more features of the predictive dataset. Such features may include the presence of a condition common to members of the dataset, such as the presence of a disease.
[0100] The expression profiles (e.g., presence, absence, quantity) of biomarkers in a sample (e.g., cancerous tissue from a patient) can be used to identify the state of cells, tissues, organs, and / or the patient. In some respects, the expression profile of a single biomarker indicates the state. In others, the expression profiles of multiple biomarkers indicate the state.
[0101] definition As used herein, unless otherwise specified, the following terms have the meanings assigned to them.
[0102] Unless the context clearly specifies otherwise, as used in this specification and claims, the singular forms “a,” “an,” and “the” include a plural of indicators. For example, the term “a cell” includes a plurality of cells, including mixtures thereof. The term “a protein” includes a plurality of proteins.
[0103] Furthermore, as used herein and in the following claims, “in” means both “in” and “on”, unless the context clearly specifies otherwise.
[0104] The combinations described herein, such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof", include any combination of A, B, and / or C, and may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" may be only A, only B, only C, A and B, A and C, B and C, or A and B and C, and any such combination may contain one or more members of its constituent elements A, B, and / or C. For example, a combination of A and B may include one A and multiple B, multiple A and one B, or multiple A and multiple B.
[0105] In the context of this invention, "biomarker" refers to a biomolecule, such as a protein or its modified, cleaved, or fragmented form, nucleic acid, carbohydrate, metabolite, intermediate, etc., that varies in a sample and whose presence, absence, or quantity indicates the state of the sample source (e.g., cells, tissue, patient). The terms "biomarker" and "marker" are used interchangeably.
[0106] A “dataset” is a collection of data points as its elements.
[0107] A “data point” refers to an element of a dataset, such as a subject sample, which is identified by, for example, a label or patient number that identifies the source of the sample.
[0108] A “biological state category” refers to the biological characteristic to which a data point can be classified. Each dataset contains data points 1 to i, and will have at least two data points representing at least two forms of a biological state category, present in the sample source providing the data point (+1 category) or absent in the sample source providing the data point (-1 category). In one respect, a -1 category data point represents a control (e.g., negative for the disease), but this is not always the case. For example, in some respects, a +1 category sample represents one stage of the disease (e.g., malignant cancer), while a -1 category represents another stage of the disease (e.g., benign cells). What the state category represents will depend on the nature of the diagnostic test to which the biomarker is selected. Examples of biological state categories are pathology (pathological vs. non-pathological (e.g., cancer vs. non-cancer)), drug response (responsive vs. non-responsive), toxicity (toxic vs. non-toxic), prognosis (progressing vs. non-progressing), and the most common phenotype (presence vs. absence of phenotypic conditions).
[0109] A "data element" refers to a characteristic of a data point, representing its features. For example, in one respect, a data element represents the expression values of multiple different genes in a sample. In another respect, a data element represents a peak detected by mass spectrometry. Furthermore, a data element can represent various phenotypic characteristics, such as the level of any biologically important analyte (e.g., in a clinical chemistry or hematology laboratory group), responses to questions in an evaluation test, elements of medical history, etc.
[0110] "Data element value" refers to the value assigned to a data element. The value can be qualitative or quantitative, such as "present or absent", "high, medium or low", or a measured numerical value.
[0111] A "qualified" data element refers to assigning a value to a data element that can be selected based on applicable criteria.
[0112] "Selection criteria" refers to one or more standards established by a user-implemented method applied to qualifiers to select data elements into an initial subset. Selection criteria may be a cutoff value for numerical qualifiers or a category for qualitative qualifiers. Examples of cutoff criteria are "data elements ranking in the top 10 percent in discriminative power" or "data elements providing at least 80% specificity and at least approximately 70% sensitivity." Examples of category criteria are "good" or "bad" data elements based on qualifiers; this will depend to some extent on the nature of the biological state category of interest, as for diseases with fewer diagnostic markers, data elements with lower specificity or sensitivity may be selected, using lower numerical or qualitative qualifiers. The initial selection criteria may be that the data element consistently outperforms other data elements among multiple data points in the dataset in identifying the biological state category.
[0113] As used herein, the term "corresponding" refers to the ability of a system or system component to receive input data from another system or system component and to provide an output response in response to the input data. "Output" can be in the form of data or in the form of an action taken by the system or system component.
[0114] As used herein, “gene expression level” or “gene expression level” refers to the properties and quantity of a molecule (e.g., RNA or polypeptide) encoded by a gene. The expression level of an mRNA molecule is intended to include both the quantity of mRNA (determined by the transcriptional activity of the gene encoding the mRNA) and the stability of the mRNA (determined by the half-life of the mRNA). Gene expression level is also intended to include the quantity of a polypeptide corresponding to a given amino acid sequence encoded by a gene. Therefore, the expression level of a gene may correspond to the quantity of mRNA transcribed from the gene, the quantity of a polypeptide encoded by the gene, or both. The expression level of a gene product can be further classified according to the expression level of different forms of the gene product. For example, RNA molecules encoded by a gene may include differentially expressed splice variants, transcripts with different initiation or termination sites, and / or other differentially processed forms. Gene-encoded polypeptides may encompass cleaved and / or modified forms of the polypeptide. Furthermore, polypeptides with a given type of modification may exist in multiple forms. For example, a polypeptide may be phosphorylated at multiple sites and express different levels of differentially phosphorylated proteins.
[0115] As used herein, “gene expression profile” refers to a characteristic representation of gene expression levels in a sample (such as cells or tissues). The determination of a gene expression profile in an individual sample represents the gene expression status of that individual. A gene expression profile reflects the expression of messenger RNA, or polypeptides, or their forms encoded by one or more genes in a cell or tissue. More generally, “expression profile” refers to the spectrum of biomolecules (nucleic acids, proteins, carbohydrates) that exhibit different expression patterns in different cells or tissues. The term “expression profile” encompasses the term “gene expression profile”.
[0116] As used herein, a “computer program product” means an organized set of instructions expressed in the form of statements in a natural language or programming language, contained in a physical medium of any nature (e.g., written, electronic, magnetic, optical, or otherwise), and usable with a computer or other automated data processing system of any nature (but preferably based on digital technology). When a computer or data processing system executes such programming language statements, it causes the computer or data processing system to perform according to the specific content of the statements. Computer program products include, but are not limited to, programs in source code and object code and / or tests or databases embedded in computer-readable media. Furthermore, computer program products that enable a computer system or data processing device to operate in a pre-selected manner may be provided in various forms, including but not limited to original source code, assembly code, object code, machine language, encrypted or compressed versions of the foregoing, and any and all equivalents.
[0117] Cancer discovery dataset This disclosure provides a data element selection method that reduces the chance of selecting a classifier whose discriminative power is biased towards sampling differences rather than differences in biological state categories. Specifically, the classifier can be a biomarker, such as a biomolecule exhibiting variability in expression profiles (transcriptional profiles, proteomic profiles, etc.) and clinical sampling. In a preferred aspect of the invention, the biomarker is obtained by proteomic analysis of patient samples. However, the classifier can also be any other phenotypic feature.
[0118] Datasets may contain biases or pre-analytical variables that produce “incorrect” classifiers / biomarkers—that is, biomarkers that differentiate groups based on specific biases rather than the underlying biological state under study. For example, if a dataset is gender-biased in terms of the presence / absence of a disease, some highly discriminatory classifiers / biomarkers might differentiate data points based on gender rather than disease. Similarly, if diseased and normal samples in a dataset are treated differently, classifiers / biomarkers might differentiate data points based on the difference in treatment rather than disease.
[0119] In independent datasets, the likelihood of the same bias is reduced. Therefore, classifiers / biomarkers common to all independent datasets are more likely to differentiate based on the biological state of interest rather than some experimental bias. Thus, two datasets are independent if they are collected in a way that significantly reduces the likelihood of being affected by the same bias; that is, the datasets are independent if the populations used to obtain these datasets show statistically significant differences in at least one pre-analytical variable. The best way to reduce bias between datasets is to collect data points from different locations in different geographic areas. This makes it more likely that bias factors will be randomized across different datasets, thus eliminating them in the possible crossover subsets of classifiers / biomarkers.
[0120] Additional or alternative methods to reduce bias include collecting data points from different times and / or groups that differ in one or more of these non-restrictive pre-analytical variables, such as: sex, age, race, sample collection parameters, sample processing parameters, weight, diet, medication status, medical condition, physical activity level, pregnancy and menstruation, presence and / or level of circulating antibodies, and clinical characteristics (e.g., PSA level, cholesterol level, family history, etc.). Preferably, the groups differ in many pre-analytical variables.
[0121] When selecting certain types of biomarkers (such as those associated with a specific disease), it may be particularly important to provide a population that differs from certain pre-analytical variables. For example, when identifying biomarkers for reduced protein C levels, it may be necessary to provide a population that differs from other thrombosis risk factors.
[0122] This disclosure recognizes that, in some embodiments, identifying characterization profiles (such as expression profiles of cells having a given cell state) leads to the discovery of classifiers, such as biomarkers, which can be used to identify said cell state with high probability (e.g., having at least about 80% specificity and at least about 70% sensitivity in diagnostic tests). Expression profiles can be derived from the expression of nucleic acids (e.g., RNA transcripts, including their differentially spliced or processed forms), proteins (including their modified and / or processed forms), carbohydrates (e.g., lectins), etc. In one aspect, the cell state reflects the state of the patient from whom said cells are derived and can diagnose the physiological processes the patient is experiencing (e.g., pathological responses experienced when a patient has or is developing a disease or recovering from a disease).
[0123] The first step is to acquire multiple independent datasets. Each dataset includes data points, such as labels indicating sample or patient numbers, representing multiple samples from multiple sample sources. Each dataset includes multiple forms of at least one biological state category, with multiple data points (samples) belonging to each form of that category. For example, biological state categories may include, but are not limited to: the presence / absence of disease in the sample source (i.e., the patient from whom the sample was obtained); disease stage; disease risk; likelihood of disease recurrence; shared genotypes of one or more gene loci (e.g., common HLA haplotypes; gene mutations; gene modifications, such as methylation); exposures to agents (e.g., toxic or potentially toxic substances, environmental pollutants, candidate drugs, etc.) or conditions (temperature, pH, etc.); demographic characteristics (age, sex, weight; family history; past health history, etc.); drug resistance; drug sensitivity (e.g., responsiveness to drugs), etc.
[0124] Datasets are independent of each other to reduce collection bias in the final classifier selection. For example, they may be collected from multiple sources and collected at different times and locations using different exclusion or inclusion criteria; that is, datasets may be relatively heterogeneous when considering features beyond those defining biological state categories. Factors contributing to heterogeneity include, but are not limited to, biological variations due to sex, age, and race; individual variations due to diet, exercise, and sleep behavior; and sample processing variations due to clinical protocols for blood handling. However, biological state categories may contain one or more common features (e.g., sample sources may represent individuals with the same disease and the same sex or one or more other common demographic features).
[0125] In one respect, datasets from multiple sources are generated by collecting samples from the same patient population at different times and / or under different conditions. However, datasets from multiple sources do not constitute a subset of a larger dataset, i.e., datasets from multiple sources are collected independently (e.g., from different locations and / or at different times, and / or under different collection conditions).
[0126] In a preferred aspect, multiple datasets are obtained from multiple different clinical trial sites, and each dataset includes multiple patient samples obtained at each individual trial site. Sample types include, but are not limited to, blood, serum, plasma, nipple aspirate, urine, tears, saliva, cerebrospinal fluid, lymph, cell and / or tissue lysates, laser-microdissected tissue or cell samples, embedded cells or tissue (e.g., paraffin blocks or frozen); fresh or archival samples (e.g., from autopsies). For example, samples may be derived from in vitro cell or tissue cultures. Alternatively, samples may be derived from an organism or a population of organisms, such as single-celled organisms. Thus, for example, in a method for discovering biomarkers for a specific cancer, blood samples may be collected from subjects selected by independent groups at two different testing sites, thereby providing samples from which independent datasets will be developed.
[0127] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0128] In some embodiments of the disclosed method, for each cancer category, at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
[0129] In some embodiments of the disclosed method, for each drug or drug combination under a cancer category, at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
[0130] In some embodiments of the disclosed method, cancer classification data includes information from cancer cell lines screened in vitro, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from groups comprised of cancer cell line identification, TCGA classification, tissue type, tissue subtype, or a combination thereof.
[0131] In some implementations of the disclosed method, patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health conditions, and combinations thereof.
[0132] In some embodiments of the disclosed methods, cancer biomarker data include gene expression data. In some embodiments, the gene expression data is normalized using transcripts per million (TPM). In some embodiments, the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0133] In some embodiments of the disclosed methods, the cancer biomarker data includes gene methylation data. In some embodiments, the cancer biomarker data includes protein biomarker data.
[0134] In some embodiments of the disclosed method, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes substructural features of the drug. In some embodiments, the substructural features of the drug include descriptors from the SMILES specification of the drug.
[0135] In some embodiments of the disclosed method, biological response data include in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve), or a combination thereof.
[0136] In some embodiments of the disclosed methods, the biological response data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergistic scores obtained using a highest single-agent (HSA) model, a Loewe additivity model, or a Bliss independent model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe synergistic scores.
[0137] In some embodiments of the disclosed methods, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or combinations thereof. In some embodiments, efficacy assessments include tumor size measurements, tumor burden assessments, efficacy biomarker assessments, or combinations thereof.
[0138] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes cancer drug sensitivity genomics (GDSC).
[0139] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0140] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas project.
[0141] In some implementations of the disclosed method, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0142] In vitro cell line drug screening dataset In some embodiments, this disclosure uses drug compound, drug response, tissue type, and genetic molecular data from large cell line sensitivity screenings such as CCLE, CTRP, and GDSC. These datasets include data points with cell line molecular data, cell line drug response data, and cell line tissue type, among others.
[0143] Molecular data from the CCLE, CTRP, and GDSC datasets were obtained directly from the source datasets. The molecular data includes gene expression, copy number, and mutation information. Expression data were obtained from next-generation RNA sequencing (RNA-seq), and the values are continuous. Copy number information was obtained from whole-exome sequencing and is continuous. Mutation information was obtained from whole-exome sequencing and is expressed as mutation frequencies of gene alleles.
[0144] Drug information for the CCLE, CTRP, and GDSC datasets was obtained from the PharmacoGx package or directly from the source database. The chemical structures of the drugs were obtained using PubChem's Simplified Molecular Input Lines (SMILES) system and converted into drug molecule descriptors using RDKit.
[0145] Drug response data include three metrics: IC50, AUC, and activity at 1 μM. IC50 (half-maximal inhibitory concentration) and AUC (area under the curve) information were obtained from the PharmacoGx package or directly from source databases (GDSC, CCLE, CTRP).
[0146] In vitro cell line drug combination dataset In some embodiments, this disclosure uses molecular information, drug response, tissue type, and genetic molecular data from large drug combination cell line sensitivity screening datasets such as NCI-ALMANAC and Drugcombo.org. These datasets include data points containing cell line molecular data, cell line drug response data, and cell line tissue type, among other things.
[0147] Molecular data from the NCI-ALMANAC and Drugcombo.org datasets were obtained directly from the source datasets or CCLE and GDSC. The molecular data includes gene expression, copy number, and mutation information. Expression data were obtained from next-generation RNA sequencing (RNA-seq), and the values are continuous. Copy number information was obtained from whole-exome sequencing and is continuous. Mutation information was obtained from whole-exome sequencing and is expressed as mutation frequencies of gene alleles.
[0148] Drug information from NCI-ALMANAC and Drugcombo.org was obtained from the source dataset. The chemical structures of the drugs were obtained through PubChem's Simplified Molecular Input Lines (SMILES) system and converted into drug molecule descriptors using RDKit.
[0149] Combination drug response values from the NCI-ALMANAC and Drugcombo.org datasets are obtained directly from the source datasets and include the following metrics: IC50, AUC, and synergy score. The synergy score is obtained from the highest single-dose (HSA) model, the Loewe additivity model, the Bliss independent model, or the zero-interaction power (ZIP) model.
[0150] Clinical research dataset In some implementations, this disclosure uses molecular information, drug response, tissue type, and genetic molecular data from clinical research data from the Cancer Genome Atlas Project (TCGA). These datasets include data points containing patient molecular data, patient drug response data, patient survival data, and patient cancer type.
[0151] The molecular data in the TCGA dataset were obtained from the Genome Data Sharing (GDC) platform. The molecular data includes gene expression, copy number, and mutation information. Expression data is from next-generation RNA sequencing (RNA-seq), and the values are continuous. Copy number information is from whole-exome sequencing and is also continuous. Mutation information is from whole-exome sequencing and is expressed as mutation frequencies of gene alleles.
[0152] Drug information from TCGA is obtained from the GDC data platform. The chemical structure of the drug is obtained through PubChem's Simplified Molecular Input Line (SMILES) system and converted into a drug molecule descriptor using RDKit.
[0153] Patient drug response in TCGA was defined according to the Evaluation of Clinical Efficacy in Solid Tumors (RECIST) and the Cheson criteria for hematologic malignancies. The RECIST criteria included the following categories: complete response (CR), partial response (PR), stable disease (SD), progressive disease (PD), and non-evaluable response (NE). The Cheson criteria included the following categories: complete response (CR), unconfirmed complete response (CRu), partial response (PR), stable disease (SD), relapsed disease (RD), and progressive disease (PD). Patient survival endpoints in TCGA were measured as continuous values in days.
[0154] Learning Algorithms This disclosure uses a cancer discovery dataset to train at least one learning algorithm for a predictive model, wherein the trained predictive model is able to constrain a predictive biological response to human cancer tissue exhibiting a specific set of cancer biomarkers based on a variety of drugs or drug combinations.
[0155] In some embodiments of the disclosed method, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0156] In some implementations of the disclosed method, at least one learning algorithm includes a multivariate regression algorithm, including random forests, support vector machines, and artificial neural networks.
[0157] In some implementations of the disclosed method, at least one learning algorithm includes a classification algorithm, including random forests, support vector machines, and artificial neural networks.
[0158] In some embodiments of the disclosed method, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0159] In some implementations, the disclosed method also includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0160] In some implementations, the disclosed method also includes independently resampling data elements in each dataset.
[0161] Select the initial subset Now, a subset of data elements, such as genes or proteins, is selected from each dataset based on selection criteria. Typically, the genes or proteins that are the “best” predictors in each dataset will be selected. For example, selection criteria might be the “top 10%” or “genes or proteins that provide a specified level of model prediction accuracy.” All data elements in each dataset that meet the selection criteria will be selected as the initial subset. For example, if one hundred genes or proteins are sorted in each dataset, the top 10% or the top 10 genes or proteins could be selected as the initial dataset.
[0162] Select the cross subset Most typically, these initial subsets are not identical in terms of the data elements they contain. However, if they contain common data elements, these data elements can be selected into the cross subset. Thus, for example, the initial subset of dataset 1 might contain genes or proteins 1, 3, 5, 7, and 9. The initial subset of dataset 2 might contain genes or proteins 1, 2, 3, 4, and 5. The cross subset might contain any or all of genes or proteins 1, 3, and 5 as common data elements between the two initial subsets.
[0163] More specifically, results from multiple datasets are cross-compared to identify a final set of common data elements with consistent expression patterns as a group of potential biomarkers. Therefore, data elements selected or limited to having good "values" or "weights" using the aforementioned learning algorithm in independently discovered datasets are compared to select a cross subset of data elements. The data elements in the cross subset are those that have good values across multiple datasets; that is, the data elements are consistently good biomarkers. In some cases, a "good value" refers to a data element that has a specificity greater than at least 80% and a sensitivity greater than at least about 70% in tests that detect or diagnose a category of biological state.
[0164] Collect data points and generate data elements Collect data points representing individual samples within the dataset. Each data point contains a data element. Multiple data points in the dataset share characteristics belonging to the same type of biological state category. For example, each data point belonging to the same biological state category could represent a sample from a patient identified as having a disease of interest for the biomarker being identified.
[0165] Data elements are features of data points, representing the characteristics of those data points. For example, in one aspect, a data element represents the expression value of multiple different genes in a patient sample suffering from a disease common to patients contributing samples to the dataset. Any expression profiling method known in the art can be used to obtain expression values and is covered within the scope of this invention.
[0166] Data elements (e.g., gene expression values) can be obtained through transcriptional analysis and / or proteomic analysis. Transcriptional analysis techniques include, but are not limited to: sequencing-by-synthesis (SBS), Northern blot, differential display methods based on qPCR and RT-PCR, nuclease protection, representational differential analysis (RDA), suppression subtractive hybridization (SSH) and enzymatic degradation subtraction (EDS), gene array analysis, cDNA fingerprinting, subtractive hybridization, gene expression serial analysis, or SAGE, etc. Proteomic analysis techniques include, but are not limited to: two-hybrid analysis, fluorescence resonance energy transfer (MET), two-dimensional gel electrophoresis, mass spectrometry (e.g., laser desorption / ionization mass spectrometry), fluorescence (e.g., sandwich immunoassay), surface plasmon resonance, ellipsometry, and atomic force microscopy.
[0167] Other types of biomolecules with differential expression can be analyzed to provide data elements. For example, carbohydrates such as lectins (e.g., glycans) have diverse expression patterns and can provide data values for data elements that constitute data points.
[0168] The preferred method for expression profiling analysis is high-throughput and obtains data elements from a dataset of more than about ten, more than about 50, more than about 100, more than about 200, or more than about 500 samples.
[0169] Preferably, algorithms such as UMAP or PCA are used to transform data elements to reduce the number of feature dimensions. For example, 20,000 genes can be reduced to only a few hundred dimensions.
[0170] Preferably, the data element is represented as a numerical vector, including a value representing the level of the sample component represented by the data element and at least one other characteristic of the sample component / data element, such as its name or descriptor.
[0171] Limited data elements In the next step, any type of multivariate analysis is used to qualify the data elements obtained from the expression profiling method. In one approach, qualifying involves using pattern recognition processes, such as classification models.
[0172] Classification models can be trained on pre-classified "known data elements" (e.g., cancer or non-cancer, cancer type). The data elements used to form the classification model can be called a "training dataset" or a "discovery dataset." Once trained, the classification model can identify patterns in the data elements from unknown samples. The classification model can then be used to classify unknown samples. For example, this can be used to predict whether a particular biological sample is associated with a certain biological disease (e.g., having or not having the disease, and the type of disease).
[0173] Any suitable statistical classification (or “learning”) method can be used to form a classification model that attempts to categorize data volumes into different classes based on objective parameters present in the data. Classification methods can be supervised or unsupervised. Supervised and unsupervised classification processes are known in the art and have been commented on, for example, in Jain, IEEE Transactions on Pattern Analysis and Machine Intelligence 22 (1): 4-37, 2000. When choosing a classification method, a balance must be struck between reducing the number of data elements to simplify analysis and minimizing the risk of losing useful information.
[0174] Unsupervised classification attempts to learn classification based on similarities found in the training dataset, without pre-classifying the data elements derived from the training dataset (e.g., representation data).
[0175] Unsupervised learning methods include cluster analysis. Cluster analysis attempts to divide data into "clusters" or groups, where ideally, members of the clusters or groups should be very similar to each other and very dissimilar to members of other clusters.
[0176] In supervised classification, training data containing instances of known categories is presented to a learning mechanism that uses a learning algorithm to learn one or more sets of relations defining each known category. New data can then be applied to the learning mechanism, and the learned relations are used to classify the new data. Differentially expressed sample components (i.e., data elements defining data points) can be identified using a set of data elements (whose values represent the expression of sample components) as training data, where the identity of each data point (i.e., the label corresponding to the sample number / patient number) is known in advance. Supervised learning techniques derive a classification model (classifier) that assigns data elements obtained from multiple data points to a predetermined number of known categories with minimal error. The contribution of each variable to the classification model is then analyzed as a measure of the data element values, i.e., which data elements are likely to serve as biomarkers with good discriminative power (i.e., the ability of biomarkers to distinguish between data points with and without biological states). For each dataset, each common data element in each data point is constrained based on the data element's ability to classify the data point into a biological state category, as a function of the data element value.
[0177] There are different methods for deriving classification models, and the types of classification methods commonly used are not a limiting feature of this invention.
[0178] In some implementations, this disclosure integrates resampling procedures into the evaluation of expressive data to reduce the impact of differences between samples within a dataset (e.g., samples from patients at clinical trial sites) and between different datasets (e.g., samples from patients at different clinical trial sites using different exclusion and inclusion criteria, and samples from sampling groups with different demographic characteristics). Resampling methods are preferably applied in a supervised learning environment, such as using the learning algorithms disclosed herein, including guided learning, bagging, boosting, Monte Carlo simulations, etc.
[0179] Therefore, in one aspect, multiple datasets are independently and repeatedly divided into subsets containing test data points (+1 class data points) and compared with reference or control data points (-1 class data points).
[0180] In each resampling run, data elements that make a significant and consistent contribution to separating data points with at least one common feature from those without at least one common feature are selected; that is, data points that can diagnose a biomarker with at least one common feature are identified. Parameters such as mean, variance, and confidence intervals (e.g., confidence scores for expressed data) of the sampled data elements are measured to determine the distribution of these parameters and identify outlier scores, thus forming a short list of candidate biomarkers represented by the data elements. For example, expressed values with high mean rank and small standard deviation (such as sequence read counts) can be selected for this list. By performing such analyses independently on each of the multiple datasets, the likelihood of data elements being selected due to bias or artifacts in the data is reduced, thereby reducing the possibility of false biomarker discovery.
[0181] This method identifies data elements with high confidence values (selected differences from a zero (random) distribution are accepted as statistically significant and with the fewest false findings, e.g., FDR ≤ 0.05) and expresses them qualitatively in the same way (overexpressed or underexpressed in both datasets). High-confidence outliers are ranked according to their expression difference between the test and reference data points (i.e., the most diagnostically significant outliers) to the smallest difference (i.e., the least diagnostically significant outliers).
[0182] Therefore, for example, gene or protein expression data from a sample set might yield expression data for over 20,000 genes or proteins: each is a data element, and its measured expression level is a data element value. After performing a selected form of analysis on the dataset, based on the expression level of each gene or protein, a particular sample (data point) can be classified as cancerous or non-cancer, and if it is cancerous, the specific cancer type category (in the form of a biological state category) can be determined or "qualified." Each gene or protein can then be ordered from most discriminative to least discriminative.
[0183] Using learning algorithms in predictive models Data elements in the cross subset can be used in a multivariate model to generate a multivariate regression algorithm. To build a multivariate prediction model, data from multiple datasets are combined and randomly divided into training and test sets.
[0184] The performance of the potential biomarker set identified through resampling and cross-comparison on the test set, and the resulting predictive model, is evaluated to identify those biomarkers that survive and still possess high diagnostic power for at least one common feature. The predictive model is validated on independent data elements from one or more new datasets, which share at least one common feature and do not involve the biomarker discovery and model building process. Validation can be performed independently on datasets containing large populations of data points, or analyzed using different methods (e.g., expression analysis techniques different from those used to initially acquire the data elements, such as by immunoassay) to obtain a validation training set that can be used to identify the most discriminative biomarkers among test subjects. Statistical methods for evaluating the validation dataset include sensitivity and specificity estimation and receiver operating characteristic (ROC) curve analysis.
[0185] The resulting multiple regression algorithm can then be tested against a separate "validation" dataset to determine its final effectiveness. This validation dataset should be independent of all discovery datasets used to discover the biomarkers that generate the regression algorithm.
[0186] Biomarkers can be evaluated after resampling, but more preferably, after cross-comparison, to identify additional features that can be used to characterize the validation dataset. For example, sequence information for peptide or nucleic acid biomarkers can be determined. These additional features can be used to generate probes to test for the presence of the biomarker in test samples (new data points) within the dataset used to validate the biomarker. Additional features may include sequence data about larger sequences of which the biomarker sequence is a subsequence (e.g., sequence data of genes or proteins derived from nucleic acids or peptides). Such data can be obtained by querying databases (such as gene sequence, protein sequence, or glycomic databases) using the biomarker sequences. Using this method, if sequences of other biomarkers are known in the database, those sequences can be identified.
[0187] Preferably, a data element is identified as a biomarker when it can predict the presence or absence of a feature of a dataset member with an accuracy greater than 70%, preferably greater than 80%, and still more preferably greater than 90%. In some aspects, a combination of multiple data elements can provide the desired predictive value. In some aspects, combinations with high predictive values may include data elements with lower confidence and may be more predictive than a single data element with a higher confidence value. For example, suitable combinations of data elements for use as biomarkers can be identified through pairing using ordered or random methods.
[0188] Xenotransplantation mouse model validation The predictive model utilizing at least one learning algorithm according to this disclosure is further validated using a xenograft mouse model, wherein, in the xenograft mouse model, human cancer tissue exhibiting a specific set of cancer biomarkers is implanted in parallel into multiple immunodeficient mice, each immunodeficient mouse is subsequently treated in parallel with one of a plurality of drugs or drug combinations, and the biological response to the treatment is fed back into the predictive model to further train the learning algorithm. In some embodiments of the disclosed method, human cancer tissue exhibiting a specific set of cancer biomarkers is orally implanted into multiple immunodeficient mice.
[0189] Patient-derived xenografts (PDXs) are cancer models in which tissue or cells from a human patient's tumor are implanted into an immunodeficient or humanized mouse. PDX models are used to create an environment that allows cancer to grow naturally, monitors cancer, and evaluates treatments in the original patient and patients with similar cancer characteristics.
[0190] Xenograft of tumors Several types of immunodeficient mice can be used to establish PDX models: athymic nude mice, severely compromised immunodeficient (SCID) mice, NOD-SCID mice, and recombinant activating gene 2 (Rag2) knockout mice. [2] The mice used must be immunodeficient to prevent transplant rejection. NOD-SCID mice are considered to be more immunodeficient than nude mice and are therefore more commonly used in PDX models because NOD-SCID mice do not produce natural killer cells.
[0191] When a human tumor is removed, necrotic tissue is eliminated, and the tumor can be mechanically fragmented into smaller pieces, chemically digested, or physically processed into a single-cell suspension. Utilizing discrete tumor fragments or single-cell suspensions each has its advantages and disadvantages. Tumor fragments retain cell-cell interactions and some of the original tumor's tissue structure, thus mimicking the tumor microenvironment. Alternatively, single-cell suspensions allow scientists to collect unbiased samples of the entire tumor, eliminating unintentionally selected spatially isolated subclones during analysis or tumor passage. However, single-cell suspensions expose surviving cells to intense chemical or mechanical forces, which may susceptible them to apoptosis, affecting cell viability and implantation success rates.
[0192] Ectopic and in situ implantation Unlike creating xenograft mouse models using existing cancer cell lines, no intermediate in vitro processing steps are required before tumor fragments are implanted into mouse hosts to create PDXs. Tumor fragments are implanted ectopically or orally into immunodeficient mice. Ectopic implantation is the implantation of tissue or cells into a region of the mouse that is not related to the original tumor site, usually subcutaneously or subcapsularly. The advantage of this method is that it can be implanted directly and tumor growth can be easily monitored. With orthotopic implantation, scientists transplant tumor tissue or cells from patients into the corresponding anatomical location in mice. Subcutaneous PDX models are unlikely to metastasize in mice and do not simulate the initial tumor microenvironment, with an implantation rate of 40-60%. Subcapsular PDXs retain the original tumor matrix and equivalent host matrix, with an implantation rate of 95%. Ultimately, tumor implantation takes about 2 to 4 months, depending on the tumor type, implantation site and immunodeficient mouse strain used; implantation failure can not be declared until at least 6 months later. [2] Researchers can use ectopic implantation to first implant tumors from patients into mice and then use orthotopic implantation to implant tumors that have grown in mice into offspring mice.
[0193] Implants of different generations The first generation of mice that receive tumor fragments from patients are usually designated F0. When the tumors in the F0 mice become large enough, researchers transplant the tumors into the next generation of mice. Each subsequent generation is designated F1, F2, F3...Fn. For drug development studies, mouse expansion between the F3 and F10 generations is often used to ensure that the PDX is genetically or histologically indistinguishable from the patient's tumor. [8] The predictive model was validated better than cancer cell lines. Without being bound by any particular theory, this disclosure recognizes that a combination of one or more factors makes the PDX mouse model a better validation tool than cancer cell line screening or cancer cell line-derived xenograft models (CDX). First, cancer cell lines are originally derived from patient tumors but acquire the ability to proliferate in vitro through cell culture. Due to in vitro manipulation, cell lines traditionally used for cancer research undergo genetic alterations that are irreversible when the cells are grown in vivo. Because the cell culture process includes an enzymatic environment and centrifugation, cells better suited to survival in culture can be selected, tumor-resident cells and proteins that interact with cancer cells can be eliminated, and the culture becomes phenotypically homogeneous.
[0194] Second, when implanted into immunodeficient mice, the cell lines are less likely to develop tumors, and any tumors that do successfully grow are genetically divergent, unlike tumors in heterogeneous patients. Researchers have begun to attribute the lack of tumor heterogeneity and the absence of the human stromal microenvironment to the fact that only 5% of anticancer drugs receive FDA approval after preclinical testing. Specifically, cell line xenografts often fail to predict drug responses in primary tumors because the cell lines do not follow drug resistance pathways or the influence of the microenvironment on drug responses found in human primary tumors.
[0195] Third, numerous PDX models have been successfully established for the treatment of breast cancer, prostate cancer, colorectal cancer, lung cancer, and many other cancers because using PDX offers unique advantages over cell lines for drug safety and efficacy studies and for predicting patient tumor responses to certain anticancer agents. Since PDX can be passaged without in vitro processing steps, PDX models allow patient tumors to proliferate and expand without significant genetic changes in tumor cells occurring over multiple generations in mice. In PDX models, patient tumor samples grow in a physiologically relevant tumor microenvironment that mimics the oxygen, nutrient, and hormone levels at the patient's primary tumor site. Furthermore, the implanted tumor tissue retains the genetic and epigenetic abnormalities found in the patient and xenograft tissue, including the surrounding human matrix, can be excised from the patient. Therefore, numerous studies have found that PDX models exhibit responses to anticancer agents similar to those of actual patients who provided the tumor samples.
[0196] Computer programs and systems The training dataset and classification model according to embodiments of the present invention can be embodied in computer code executed or used by a digital computer. The computer code can be stored on any suitable computer-readable medium, including optical discs or disks, sticks, magnetic tapes, transmission type media (such as digital and analog), etc., and can be written in any suitable computer programming language, including C, C++, Java, Python, etc.
[0197] The output data generated during training can be displayed on any graphical interface on a user device that can be connected to a digital computer or a server connected to such a computer (e.g., via the Internet). Suitable digital computers include miniature, small, or mainframe computers using any standard or proprietary operating system, such as Unix-based, Windows™, or Linux™ operating systems. The digital computer used may be physically separate from the instrument used to acquire the values of data elements in the analytical experiments. For example, the computer may be located remotely from the mass spectrometer used to create the spectrum of interest, or it may be coupled to the mass spectrometer. The graphical interface may also be located remotely from the computer, for example, as part of a wireless device that can be connected to a network.
[0198] This disclosure also includes a computer system having a database containing characteristics of data elements / biomarkers specific to different cancers. In one aspect, cell state includes one or more stages of differentiation; phenotypic expression; cell cycle proliferation or stage; response to stimuli, diseases, agents (e.g., toxins or potential toxic agents, known or candidate drugs; antibiotics; infectious or pathological organisms; environmental pollutants, etc.), conditions, etc.; or conditions (temperature, pH, etc.); and so on. In another aspect, cell state reflects the state of the cell's origin. For example, cell state may reflect disease or other physiological responses or disorders experienced by the patient of cell origin (e.g., such as old age; mental state; addiction; allergic reactions, etc.).
[0199] In one implementation, the database includes ordered or clustered biomarkers (i.e., biomarkers are divided into subsets based on their discriminative power). Biomarkers can be ordered or clustered based on their associations with various parameters. Such parameters include responses to toxins, diseases, contaminants, conditions, stressors, developmental stages, drugs, therapeutics, antibiotics, etc. The biomarkers contained in the database exhibit a relatively narrow range of variation in a population under a given cellular state, but possess high discriminative power across cellular states. For example, the biomarkers have reproducible associations with parameters (specificity of at least 80% and sensitivity of at least about 70% in tests for detecting or diagnosing parameters) and high discriminative power.
[0200] However, it should be noted that discriminative power is not a limiting characteristic of biomarkers. For example, for some diseases where there are few or no satisfactory diagnostic tests, biomarkers with lower specificity and / or sensitivity may still be valuable.
[0201] The system also includes a database management system. User requests or queries are formatted into an appropriate language that the database management system understands, and the database management system processes the queries to extract relevant information from the training set database.
[0202] The system may also include records from an external database, or be able to communicate with such an external database. Preferably, the system can be connected to a network server and a network to which one or more clients are connected. The network can be a local area network (LAN) or a wide area network (WAN) as known in the art. Preferably, the server includes the hardware required to run computer program products (e.g., software) to access database data and process user requests. For example, one type of user request might be for the system to identify biomarkers associated with a selected cell state. Such a request may provide optional data options, such as probe sources (e.g., links to sites that provide binding mates for the biomarkers, such as antibodies) that can be used to detect one or more biomarkers.
[0203] The system also includes an operating system (e.g., UNIX or Linux) for executing instructions from the database management system. In one aspect, the operating system also runs World Wide Web applications and World Wide Web servers, thereby connecting the servers to the network.
[0204] Preferably, the system includes one or more user devices comprising a graphical user interface (GUI) containing interface elements such as buttons, drop-down menus, scroll bars, and text input fields, which are common in graphical user interfaces known in the art. Requests entered on the GUI are transmitted to an application within the system (such as a web application) for formatting to search for relevant information in one or more system databases. User-input requests or queries can be constructed using any suitable database language (e.g., Sybase or Oracle SQL). In one embodiment, users of the user devices within the system can directly access data using an HTML interface provided by the system's web browser and web server.
[0205] The graphical user interface (GUI) can be generated from GUI code that is part of the operating system and can be used to input and / or display input data. The processed data results can be displayed on the interface, printed on a printer communicating with the system, stored in a storage device, and / or transmitted over a network, or provided in the form of a computer-readable medium.
[0206] Non-restrictive implementation scheme The following provides some non-limiting embodiments of different aspects of this disclosure.
[0207] 1. A method for training artificial intelligence to identify one or more predictive biomarkers for cancer treatment using at least one hardware processor, the method comprising: training at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the trained predictive model is capable of defining the plurality of drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and validating the predictive model using a xenograft mouse model, wherein in the xenograft mouse model, human cancer tissue exhibiting the specific set of cancer biomarkers is implanted in parallel into a plurality of immunodeficient mice, the plurality of immunodeficient mice are subsequently treated in parallel with one of the plurality of drugs or drug combinations, and the biological response to the treatment is fed back to the predictive model to further train the learning algorithm.
[0208] 2. The method as described in embodiment 1, wherein the human cancer tissue exhibiting the specific group of cancer biomarkers is in situ implanted into the plurality of immunodeficient mice.
[0209] 3. The method as described in embodiment 1 or 2, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0210] 4. The method as described in embodiment 3, wherein for each cancer category, the at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
[0211] 5. The method as described in embodiment 4, wherein for each drug or drug combination under the cancer category, the at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
[0212] 6. The method of any one of embodiments 3-5, wherein the cancer classification data includes information from in vitro screened cancer cell lines, patient information from clinical studies, or a combination thereof.
[0213] 7. The method of embodiment 6, wherein the cancer cell line information includes information selected from cancer cell line identification, TCGA classification, tissue type, tissue subtype or a combination thereof.
[0214] 8. The method of embodiment 6, wherein the patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health status and combinations thereof.
[0215] 9. The method as described in any one of embodiments 3-8, wherein the cancer biomarker data includes gene expression data.
[0216] 10. The method as described in embodiment 9, wherein the gene expression data are normalized using per million transcripts (TPM).
[0217] 11. The method of any one of embodiments 9-10, wherein the gene expression data is normalized to the geometric mean of at least four housekeeping genes with low variance.
[0218] 12. The method as described in any one of embodiments 3-11, wherein the cancer biomarker data includes gene methylation data.
[0219] 13. The method of any one of embodiments 3-12, wherein the cancer biomarker data includes protein biomarker data.
[0220] 14. The method of any one of embodiments 3-13, wherein the drug or drug combination data includes the chemical structure of the drug.
[0221] 15. The method of any one of embodiments 3-14, wherein the drug or drug combination data includes substructural features of the drug.
[0222] 16. The method of embodiment 15, wherein the substructural features of the drug include a descriptor from the SMILES specification of the drug.
[0223] 17. The method of any one of embodiments 3-16, wherein the biological response data includes in vitro drug screening results.
[0224] 18. The method of embodiment 17, wherein the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof.
[0225] 19. The method of any one of embodiments 3-18, wherein the biological response data includes in vitro drug combination screening results.
[0226] 20. The method of embodiment 19, wherein the in vitro drug combination screening results are in the form of synergistic scores obtained by the highest single-agent (HSA) model, the Loewe additivity model, the zero interaction power (ZIP) model, or the Bliss independent model.
[0227] 21. The method of embodiment 19, wherein the in vitro drug combination screening results are in the form of AUC scores.
[0228] 22. The method as described in any one of embodiments 3-21, wherein the biological response data includes clinical study results.
[0229] 23. The method as described in implementation scheme 22, wherein the clinical study results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
[0230] 24. The method of embodiment 23, wherein the efficacy evaluation includes tumor size measurement, tumor burden evaluation, efficacy biomarker evaluation, or a combination thereof.
[0231] 25. The method of any one of embodiments 1-24, wherein the at least one cancer drug discovery dataset includes cancer drug sensitivity genomics (GDSC).
[0232] 26. The method of any one of embodiments 1-25, wherein the at least one cancer drug discovery dataset comprises data from NCI Almanac or Drugcomb.org.
[0233] 27. The method of any one of embodiments 1-26, wherein the at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas project.
[0234] 28. The method of any one of embodiments 1-27, wherein the at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0235] 29. The method of any one of embodiments 1-28, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0236] 30. The method of any one of embodiments 1-28, wherein the at least one learning algorithm includes an unsupervised learning algorithm.
[0237] 31. The method according to any one of embodiments 1-30, wherein the at least one learning algorithm includes a regression algorithm, preferably a random forest regressor.
[0238] 32. The method of any one of embodiments 1-31, wherein the at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0239] 33. The method as described in any one of embodiments 1-32, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0240] 34. The method as described in any one of embodiments 1-33 further includes applying a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0241] 35. The method as described in any one of embodiments 1-34, further comprising independently resampling data elements in each dataset.
[0242] 36. A system comprising: an input unit capable of downloading at least one cancer drug discovery dataset; and a training unit capable of using the at least one cancer drug discovery dataset to train at least one learning algorithm in a prediction model, wherein the trained prediction model is capable of defining the plurality of drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and wherein the training unit is capable of further training the learning algorithm using a xenograft mouse model, wherein in the xenograft mouse model, human cancer tissue exhibiting the specific set of cancer biomarkers is implanted in parallel into a plurality of immunodeficient mice, and the plurality of immunodeficient mice subsequently receive, in parallel, one of the plurality of drugs or drug combinations, and the biological response to the treatment is fed back to the prediction model.
[0243] 37. The system as described in embodiment 36, wherein the human cancer tissue exhibiting the specific set of cancer biomarkers is in situ implanted into the plurality of immunodeficient mice.
[0244] 38. The system of embodiment 37, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0245] 39. The system of embodiment 38, wherein for each cancer type, the at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination to the cancer type.
[0246] 40. The system of embodiment 39, wherein for each drug or drug combination under the cancer type, the at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination to the cancer type.
[0247] 41. The system of any one of embodiments 38-40, wherein the cancer classification data includes information from cancer cell lines screened in vitro, patient information from clinical studies, or a combination thereof.
[0248] 42. The system of embodiment 41, wherein the cancer cell line information includes information selected from cancer cell line identification, TCGA classification, tissue type, tissue subtype or a combination thereof.
[0249] 43. The system as described in embodiment 41, wherein the patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health status and combinations thereof.
[0250] 44. The system of any one of embodiments 38-43, wherein the cancer biomarker data includes gene expression data.
[0251] 45. The system as described in embodiment 44, wherein the gene expression data is normalized using transcripts per million (TPM).
[0252] 46. The system of any one of embodiments 44-45, wherein the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0253] 47. The system of any one of embodiments 38-46, wherein the cancer biomarker data includes gene methylation data.
[0254] 48. The system of any one of embodiments 38-47, wherein the cancer biomarker data includes proteomic biomarker data.
[0255] 49. The system of any one of embodiments 38-48, wherein the drug or drug combination data includes the chemical structure of the drug.
[0256] 50. The system of any one of embodiments 38-49, wherein the drug or drug combination data includes substructural features of the drug.
[0257] 51. The system of embodiment 50, wherein the substructural features of the drug include a descriptor from the SMILES specification of the drug.
[0258] 52. The system as described in any one of embodiments 38-51, wherein the biological response data includes in vitro drug screening results.
[0259] 53. The system as described in embodiment 52, wherein the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof.
[0260] 54. The system as described in any one of embodiments 38-53, wherein the biological response data includes in vitro drug combination screening results.
[0261] 55. The system as described in embodiment 54, wherein the in vitro drug combination screening results are in the form of synergistic scores obtained by the highest single-agent (HSA) model, the Loewe additivity model, the zero interaction power (ZIP) model, or the Bliss independent model.
[0262] 56. The system as described in embodiment 54, wherein the in vitro drug combination screening results are in the form of AUC scores.
[0263] 57. The system as described in any one of embodiments 38-56, wherein the biological response data includes clinical study results.
[0264] 58. The system as described in embodiment 57, wherein the clinical study results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
[0265] 59. The system of embodiment 58, wherein the efficacy evaluation includes tumor size measurement, tumor burden evaluation, efficacy biomarker evaluation, or a combination thereof.
[0266] 60. The system of any one of embodiments 36-59, wherein the at least one cancer drug discovery dataset includes cancer drug sensitivity genomics (GDSC).
[0267] 61. The system of any one of embodiments 36-60, wherein the at least one cancer drug discovery dataset comprises data from NCI Almanac or Drugcomb.org.
[0268] 62. The system of any one of embodiments 36-61, wherein the at least one cancer drug discovery dataset comprises data from the Cancer Genome Atlas project.
[0269] 63. The system of any one of embodiments 36-62, wherein the at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0270] 64. The system of any one of embodiments 36-63, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0271] 65. The system of any one of embodiments 36-63, wherein the at least one learning algorithm includes an unsupervised learning algorithm.
[0272] 66. The system of any one of embodiments 36-65, wherein the at least one learning algorithm includes a regression algorithm, preferably a random forest regressor.
[0273] 67. The system of any one of embodiments 36-66, wherein the at least one learning algorithm comprises a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0274] 68. The system of any one of embodiments 36-67, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0275] 69. The system as described in any one of embodiments 36-68 further includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0276] 70. The system of any one of embodiments 36-69, wherein the system is capable of independently resampling data elements in each dataset.
[0277] 71. A non-transitory computer-readable medium storing instructions, wherein, when executed by a processor, the instructions cause the processor to: train at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the trained predictive model is capable of defining the plurality of drugs or drug combinations based on a predicted biological response to human cancer tissue exhibiting a specific set of cancer biomarkers; and to train the predictive model using a xenograft mouse model, wherein in the xenograft mouse model, human cancer tissue exhibiting the specific set of cancer biomarkers is implanted in parallel into a plurality of immunodeficient mice, the plurality of immunodeficient mice subsequently receiving, in parallel, one of the plurality of drugs or drug combinations, and feeding back the biological response to the treatment to the predictive model to further train the learning algorithm.
[0278] 72. The non-transient computer-readable medium as described in embodiment 71, wherein the human cancer tissue exhibiting the specific set of cancer biomarkers is in situ implanted into the plurality of immunodeficient mice.
[0279] 73. The non-transient computer-readable medium as described in embodiment 72, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0280] 74. The non-transient computer-readable medium of embodiment 73, wherein for each cancer type, the at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination to the cancer type.
[0281] 75. The non-transient computer-readable medium of embodiment 74, wherein for each drug or drug combination under the cancer type, the at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination to the cancer type.
[0282] 76. A non-transient computer-readable medium as described in any one of embodiments 73-75, wherein the cancer classification data includes information from in vitro screened cancer cell lines, patient information from clinical studies, or a combination thereof.
[0283] 77. The non-transient computer-readable medium of embodiment 76, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, tissue type, tissue subtype or combinations thereof.
[0284] 78. The non-transient computer-readable medium of embodiment 76, wherein the patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health conditions and combinations thereof.
[0285] 79. The non-transient computer-readable medium as described in any one of embodiments 73-78, wherein the cancer biomarker data includes gene expression data.
[0286] 80. The non-transient computer-readable medium as described in embodiment 79, wherein the gene expression data is normalized using per million transcripts (TPM).
[0287] 81. A non-transient computer-readable medium as described in any one of embodiments 79-80, wherein the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0288] 82. The non-transient computer-readable medium as described in any one of embodiments 73-81, wherein the cancer biomarker data includes gene methylation data.
[0289] 83. The non-transient computer-readable medium as described in any one of embodiments 73-82, wherein the cancer biomarker data includes protein biomarker data.
[0290] 84. The non-transient computer-readable medium as described in any one of embodiments 73-83, wherein the drug or drug combination data includes the chemical structure of the drug.
[0291] 85. The non-transient computer-readable medium as described in any one of embodiments 73-84, wherein the drug or drug combination data includes substructural features of the drug.
[0292] 86. The non-transient computer-readable medium of embodiment 85, wherein the substructural features of the drug include descriptors from the SMILES specification of the drug.
[0293] 87. A non-transient computer-readable medium as described in any one of embodiments 73-86, wherein the biological response data includes in vitro drug screening results.
[0294] 88. The non-transient computer-readable medium as described in embodiment 87, wherein the in vitro drug screening results are in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof.
[0295] 89. A non-transient computer-readable medium as described in any one of embodiments 73-88, wherein the biological response data includes in vitro drug combination screening results.
[0296] 90. The non-transient computer-readable medium as described in embodiment 89, wherein the results of in vitro drug combination screening are in the form of synergistic scores obtained by the highest single-agent (HSA) model, the Loewe additivity model, the zero interaction power (ZIP) model, or the Bliss independent model.
[0297] 91. The non-transient computer-readable medium as described in embodiment 89, wherein the results of in vitro drug combination screening are in the form of AUC scores.
[0298] 92. A non-transient computer-readable medium as described in any one of embodiments 73-91, wherein the biological response data includes clinical study results.
[0299] 93. The non-transient computer-readable medium as described in embodiment 92, wherein the clinical study results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
[0300] 94. The non-transient computer-readable medium as described in embodiment 93, wherein the efficacy evaluation includes tumor size measurement, tumor burden evaluation, efficacy biomarker evaluation, or a combination thereof.
[0301] 95. A non-transient computer-readable medium as described in any one of embodiments 71-94, wherein the at least one cancer drug discovery dataset includes cancer drug sensitivity genomics (GDSC).
[0302] 96. A non-transient computer-readable medium as described in any one of embodiments 71-95, wherein the at least one cancer drug discovery dataset comprises data from NCI Almanac or Drugcomb.org.
[0303] 97. A non-transient computer-readable medium as described in any one of embodiments 71-96, wherein the at least one cancer drug discovery dataset comprises data from the Cancer Genome Atlas project.
[0304] 98. A non-transient computer-readable medium as described in any one of embodiments 71-97, wherein the at least one cancer drug discovery dataset comprises data from the Genotype Tissue Expression Project (GTEx).
[0305] 99. A non-transient computer-readable medium as described in any one of embodiments 71-98, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0306] 100. A non-transient computer-readable medium as described in any one of embodiments 71-98, wherein the at least one learning algorithm comprises an unsupervised learning algorithm.
[0307] 101. A non-transient computer-readable medium as described in any one of embodiments 71-100, wherein the at least one learning algorithm comprises a regression algorithm, preferably a random forest regressor.
[0308] 102. The non-transient computer-readable medium as described in any one of embodiments 71-101 further includes a dimensionality reduction algorithm for data elements in each dataset, preferably UMAP, PCA, or both.
[0309] 103. The non-transient computer-readable medium as described in any one of embodiments 71-102, wherein the at least one learning algorithm comprises a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0310] 104. The non-transient computer-readable medium as described in any one of embodiments 71-103, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0311] 105. A non-transient computer-readable medium as described in any one of embodiments 71-104, wherein the non-transient computer-readable medium is capable of independently resampling data elements in each dataset.
[0312] 106. The method as described in any one of embodiments 1-34, the system as described in any one of embodiments 35-70, and the non-transient computer-readable medium as described in any one of embodiments 71-104, wherein the prediction model does not provide a cancer diagnosis.
[0313] 107. The method as described in any one of embodiments 1-34, the system as described in any one of embodiments 35-70, and the non-transient computer-readable medium as described in any one of embodiments 71-104, wherein the prediction model does not provide a cancer prognosis.
[0314] 108. The method as described in any one of embodiments 1-34, the system as described in any one of embodiments 35-70, and the non-transient computer-readable medium as described in any one of embodiments 71-104, wherein the prediction model does not provide cancer diagnosis or prognosis.
[0315] Non-limiting embodiments The following provides certain non-limiting embodiments based on different aspects of this disclosure. To examine how different aspects of this disclosure affect the resulting prediction accuracy, a series of models were established for single-therapy prediction (Figure XA) and combination therapy prediction including both single-therapy and combination therapy prediction (Figure XB).
[0316] Cancer drug sensitivity genomics (GDSC).
[0317] Drugcomb.org.
[0318] National Cancer Institute ALMANAC.
[0319] Cancer Genome Atlas Project.
[0320] Genotype Tissue Expression Project (GTEx).
[0321] Extensive machine learning methods have been applied to drug response prediction problems: regularized regression methods (e.g., lasso, elastic networks, ridge regression), partial least squares (PLS) regression, support vector machines (SVM), random forests (RF), neural networks and deep learning, logistic models, or kernel Bayesian matrix factorization (KBMF). However, there are currently no reports of systematic exploration of model training strategies based on multiple large-scale cell line screening data. Furthermore, cell line-based models have not been compared with xenograft-based models. This research bridges these gaps, with the ultimate goal of improving the accuracy of drug response prediction in both cell lines and xenografts.
[0322] First, the impact of different modeling parameters (e.g., response indices, number of features, feature types) on the predictive performance of drug sensitivity testing was systematically investigated. In this non-limiting example, four drugs (doxorubicin, gemcitabine, olaparib, and palbociclib) were used, with drug sensitivity data for all four drugs available in each dataset. Predictive performance was evaluated using a xenograft mouse model. The optimal parameter set and strategy were selected from this analysis, and it was examined whether models could be trained using cell line-based and clinical-based data to predict drug sensitivity in animal xenograft systems. Because each drug is tested on samples from specific tissue types in xenograft experiments, the cancer detection datasets are correspondingly limited; for example, for drug response prediction, all cell line-based and clinical-based research data used for this modeling task came from the same cancer classification as the xenograft samples. The process of this non-limiting example is as follows: Figure 1 As shown.
[0323] Predictive Model Categories of supervised learning Drug response prediction and mitotic rate / growth curve slope prediction are regression tasks. Figure 2 The task of predicting tissue type for a specific cancer AI model is a classification task. Additionally, some tasks predicting early drug response on the TCGA clinical dataset are classification tasks.
[0324] Modeling methods and hyperparameters For regression tasks, our best-performing modeling method is random forest. For classification tasks, our best-performing modeling method is also random forest.
[0325] Each modeling method has its own set of hyperparameters. For random forest regression, the most important hyperparameters include the maximum depth of a single tree (max_depth), the number of trees in the forest (n_estimators), and the criterion for splitting.
[0326] Feature selection To select a subset of all available features for modeling, feature selection is performed using only data from the training set. For the drug response prediction task, a multi-step feature selection process is used to fine-tune our model. Figure 2 First, the top 1000 drug molecule features that predict drug response for each cancer type are selected. In a random forest, "IncNodePurity," or variable importance value, is sorted, and the top 1000 are initially selected. Next, the Boruta algorithm is used to further refine the drug molecule feature selection, making it a smaller subset to achieve optimal model accuracy. Approximately 20,000 gene expression features (HK ratios) are added to the refined drug molecule features (about 30-40), and the model is trained on drug response to find the top 1000 gene features. Again, Boruta is used to reduce the number of genes to only a few important genes (30-40). Gene biomarker optimization is performed to further reduce the total number of genes to 5-10 genes. Figure 3 Now, molecular features and optimized genetic biomarkers are used to predict treatment options for a subset of cancer types used in the initial training. This task is performed for each subset of cancer types with available training data.
[0327] result Performance evaluation of prediction models Multiple datasets were created using bagging or guided aggregation, and the model's performance was evaluated on each dataset to validate the performance of the random forest regression model. Results of random forest regression predictions on a single-therapeutic drug in vitro dataset using growth curve endpoints (AUC) showed that the model performed well, with an average R-value of [missing value]. 2 The accuracy was 0.651, the RMSE was 0.094, and the accuracy rate was 90%. Figure 4 A). Then, the monotherapy and combination therapy were evaluated together by converting the AUC score of the monotherapy and the synergistic score of the combination therapy into normalized Z-scores. Combined, the average R-squared of the model... 2 The accuracy was 0.569, the RMSE was 0.686, and the accuracy rate was 83%. Figure 4 B).
[0328] For the TCGA clinical response evaluation, random forest classification was used to build a model of patient treatment response. Due to the small dataset size (2212 patients in total), the entire TCGA clinical response dataset was combined for this evaluation. RECIST standard categories were used to construct four response variables, classifying patients according to various treatment methods: complete response (CR), partial response (PR), stable disease (SD), and progressive disease (PD). The area under the receiver operating characteristic (ROC) curve (AUC) showed good model accuracy in predicting the correct response category, with a mean ROC AUC of 0.93. Figure 5 ).
[0329] Xenograft verification Gene expression was measured using sequential TPM (transcriptions per million mapped reads) values from RNA-seq. The TPM values were normalized to a set of housekeeping gene ratios (HK ratios). The HK ratios of each gene biomarker in the PDX were fed into appropriate cancer type prediction models to determine a ranking list of treatment options. Figure 2 Top-ranked therapies were tested in a PDX to validate the accuracy of treatment predictions. In one such instance, four AI-predicted therapies were tested on an orthotopic PDX model of triple-negative breast cancer. Results showed that the single therapy with the highest AI prediction completely suppressed tumor growth. Figure 6 ).
Claims
1. A method for training at least one learning algorithm using at least one hardware processor, wherein the at least one learning algorithm is in a predictive model for cancer treatment, the method comprising: The prediction model is trained using at least one cancer drug discovery dataset, wherein the prediction model is capable of defining the multiple drugs or drug combinations based on the predicted biological response of human cancer tissue exhibiting a specific set of cancer biomarkers; and the prediction model is validated using a xenograft mouse model, wherein human cancer tissue exhibiting the specific set of cancer biomarkers is implanted in parallel into multiple immunodeficient mice, and the multiple immunodeficient mice are subsequently treated in parallel with one of the multiple drugs or drug combinations, and the biological response of the treatment is fed back to the prediction model to further train the at least one learning algorithm.
2. The method of claim 1, wherein the human cancer tissue exhibiting the specific group of cancer biomarkers is in situ implanted into the plurality of immunodeficient mice.
3. The method of claim 1, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
4. The method of claim 3, wherein for each cancer category, the at least one learning algorithm defines the drug or drug combination based on the biological response of the drug or drug combination.
5. The method of claim 4, wherein for each drug or drug combination under the cancer category, the at least one learning algorithm defines the cancer biomarker based on the correspondence between the cancer biomarker and the biological response of the drug or drug combination.
6. The method of claim 3, wherein the cancer classification data includes information from in vitro screened cancer cell lines, patient information from clinical studies, or a combination thereof.
7. The method of claim 6, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, tissue type, tissue subtype, or combinations thereof.
8. The method of claim 6, wherein the patient information includes information selected from cancer type, patient age, patient gender, patient weight, patient family history, patient past health status, and combinations thereof.
9. The method of claim 3, wherein the cancer biomarker data includes gene expression data.
10. The method of claim 9, wherein the gene expression data is normalized using transcripts per million (TPM).
11. The method of claim 9, wherein the gene expression data is normalized to the geometric mean of at least four housekeeping genes with low variance.
12. The method of claim 3, wherein the cancer biomarker data includes gene methylation data.
13. The method of claim 3, wherein the cancer biomarker data includes protein biomarker data.
14. The method of claim 3, wherein the drug or drug combination data includes the chemical structure of the drug.
15. The method of claim 14, wherein the drug or drug combination data includes substructural features of the drug.
16. The method of claim 15, wherein the substructural features of the drug include a descriptor from the SMILES specification of the drug.
17. The method of claim 3, wherein the biological response data includes in vitro drug screening results.
18. The method of claim 17, wherein the in vitro drug screening result is in the form of IC50 (half-maximum inhibitory concentration), AUC (area under the curve of drug response), or a combination thereof.
19. The method of claim 3, wherein the biological response data includes in vitro drug combination screening results.
20. The method of claim 19, wherein the in vitro drug combination screening results are in the form of synergistic scores obtained by the highest single-agent (HSA) model, the Loewe additivity model, the zero interaction power (ZIP) model, or the Bliss independent model.
21. The method of claim 19, wherein the in vitro drug combination screening result is in the form of an AUC score.
22. The method of claim 3, wherein the biological response data includes clinical study results.
23. The method of claim 22, wherein the clinical study results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
24. The method of claim 23, wherein the efficacy evaluation includes tumor size measurement, tumor burden evaluation, efficacy biomarker evaluation, or a combination thereof.
25. The method of claim 1, wherein the at least one cancer drug discovery dataset comprises Cancer Drug Sensitivity Genomics (GDSC), data from NCI Almanac or Drugcomb.org, data from the Cancer Genome Atlas project, data from the Genotype Tissue Expression Project (GTEx), or a combination thereof.
26. The method of claim 1, wherein the at least one learning algorithm comprises a regression algorithm, and optionally wherein the regression algorithm comprises a random forest regressor.
27. The method of claim 1, wherein the at least one learning algorithm comprises a classification algorithm, and optionally wherein the classification algorithm comprises a random forest classifier, an SVM classifier, or both.
28. The method of claim 1, wherein the at least one learning algorithm comprises a feature selection algorithm, and optionally wherein the feature selection algorithm comprises the Boruta algorithm.
29. The method of claim 1, further comprising applying a dimensionality reduction algorithm, and optionally wherein the dimensionality reduction algorithm includes UMAP, PCA, or both.
30. The method of claim 1, wherein the prediction model does not provide a cancer diagnosis, and optionally wherein the prediction model does not provide a cancer prognosis.
Citation Information
Patent Citations
Biomarkers and methods to predict response to inhibitors and uses thereof
US20150240315A1
Breast and ovarian cancer methylation markers and uses thereof
US20190360052A1
Biological status determination using cell-free nucleic acids
US20200115762A1
Visible neural network framework
WO2022087540A1