Artificial intelligence for identifying one or more predictive biomarkers
By analyzing expression profiling data from multiple independent sources and employing resampling techniques, the method addresses biases in existing biomarker identification, improving the accuracy and reliability of predictive biomarkers for cancer treatment responses.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CERTIS ONCOLOGY SOLUTIONS INC
- Filing Date
- 2024-03-25
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for identifying predictive biomarkers for cancer treatment response are hindered by systematic biases and high false discovery rates in expression profiling data, particularly due to site-specific factors and variability across different research protocols and populations.
A method involving the use of bioinformatics to analyze expression profiling data from multiple independent sources separately, reducing variability and bias by developing independent training sets and employing resampling techniques, followed by cross-comparison and validation of biomarkers using a xenograft mouse model.
This approach reduces the likelihood of false discoveries and enhances the accuracy of identifying predictive biomarkers for cancer treatment responses, ensuring robust performance across diverse populations.
Smart Images

Figure 2026511166000001_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to predictive biomarkers, particularly to systems, software, and methods for identifying predictive biomarkers related to cancer treatment response.
Summary of the Invention
[0002] According to one aspect of the present disclosure, a method is provided for training artificial intelligence to identify one or more predictive biomarkers for cancer treatment using at least one hardware processor.
[0003] In some embodiments, the method comprises training at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the trained predictive model is capable of evaluating multiple drugs or drug combinations from the perspective of the predicted biological response to human cancer tissues presenting a specific set of cancer biomarkers, training, and validating the predictive model by using a xenograft mouse model, wherein in the xenograft mouse model, human cancer tissues presenting a specific set of cancer biomarkers are transplanted in parallel into multiple immunodeficient mice, and then each is treated in parallel with one of multiple drugs or drug combinations, and the biological response to the treatment is fed back to the predictive model for further training the learning algorithm, validating.
[0004] In some embodiments of the disclosed method, human cancer tissues presenting a specific set of cancer biomarkers are orthotopically transplanted into multiple immunodeficient mice.
[0005] In some embodiments of the disclosed method, the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0006] In some embodiments of the disclosed method, for each class of cancer, at least one learning algorithm evaluates a drug or drug combination in relation to their biological response.
[0007] In some embodiments of the disclosed method, for each drug or drug combination under a class of cancer, at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response.
[0008] In some embodiments of the disclosed method, cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0009] In some embodiments of the methods of this disclosure, patient information includes information selected from the type of cancer, the patient's tumor biomarkers, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0010] In some embodiments of the disclosed method, cancer biomarker data includes gene expression data. In some embodiments, gene expression data is normalized using TPM (transcripts per million). In some embodiments, gene expression data is normalized to the geometric mean of four housekeeping genes with small variance.
[0011] In some embodiments of the disclosed method, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0012] In some embodiments of the disclosed method, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes feature quantities of a substructure of the drug. In some embodiments, the feature quantities of a substructure of the drug include descriptors of the drug from the SMILES standard.
[0013] In some embodiments of the disclosed method, the biological response data includes in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof. In some embodiments, the IC50, AUC, or a combination thereof are converted to a normal distribution, and the normalized scores are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0–100%.
[0014] In some embodiments of the disclosed methods, biological response data include in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergy scores obtained through the best single agent (HSA) model, the Loewe additive model, the Bliss independence model, or the zero interaction potency (ZIP) model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe, HSA, Bliss, or ZIP synergy scores. In some embodiments, the Loewe, HSA, Bliss, or ZIP synergy scores of the in vitro drug combinations are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0%–100%.
[0015] In some embodiments of the disclosed methods, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or a combination thereof. In some embodiments, efficacy assessments include tumor sizing, tumor volume assessments, efficacy biomarker assessments, response assessments, survival assessments, or a combination thereof. In some embodiments, response assessments, survival assessments, or a combination thereof are converted and reported as efficacy scores scaled to 0-1 or 0%-100%.
[0016] In some embodiments of the disclosed methods, biological response data include pharmacological outcomes of xenotransplantation with monotherapy or combination therapy. In some embodiments, xenotransplantation pharmacological outcomes include tumor sizing, tumor growth inhibition (TGI), AUC, or a combination thereof. In some embodiments, TGI or AUC are converted to a normal distribution and reported as standardized z-scores, or scaled to 0–1 or 0–100%.
[0017] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from cancer genomic drug sensitivity (GDSC).
[0018] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0019] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from NCI-ALMANAC.
[0020] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from a cancer genome atlas program.
[0021] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0022] In some embodiments of the disclosed method, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0023] In some embodiments of the disclosed method, at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0024] In some embodiments of the disclosed method, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0025] In some embodiments of the disclosed method, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0026] In some embodiments, the disclosed method further includes applying a dimensionality reduction algorithm, preferably UMAP, PCA, or both, to a predictive model.
[0027] In some embodiments, the disclosed method further includes independently resampling data elements within each dataset.
[0028] According to another aspect of the present disclosure, there is provided a system including an input unit capable of downloading at least one drug discovery dataset, and a training unit capable of training at least one learning algorithm in a prediction model using at least one cancer drug discovery dataset. The trained prediction model can evaluate a plurality of drugs or drug combinations from the perspective of the predicted biological response to human cancer tissues presenting a specific set of cancer biomarkers. The training unit can further train the learning algorithm by using data obtained from a xenograft mouse model. In the xenograft mouse model, human cancer tissues presenting a specific set of cancer biomarkers are transplanted into a plurality of immunodeficient mice in parallel, and then each is treated in parallel with one of a plurality of drugs or drug combinations, and the measurement results of the biological response to the treatment are used as new training data for the prediction model.
[0029] In some embodiments of the disclosed system, human cancer tissues presenting a specific set of cancer biomarkers are orthotopically transplanted into a plurality of immunodeficient mice.
[0030] In some embodiments of the disclosed system, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0031] In some embodiments of the disclosed system, for each cancer class, at least one learning algorithm evaluates drugs or drug combinations with respect to their biological responses.
[0032] In some embodiments of the disclosed system, for drugs or drug combinations under each cancer class, at least one learning algorithm evaluates cancer biomarkers with respect to the correspondence of the drugs or drug combinations to biological responses.
[0033] In some embodiments of the disclosed system, cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from a group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0034] In some embodiments of the system of this disclosure, patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0035] In some embodiments of the disclosed system, cancer biomarker data includes gene expression data. In some embodiments, gene expression data is normalized using TPM (transcripts per million). In some embodiments, gene expression data is normalized to the geometric mean of four housekeeping genes with small variance.
[0036] In some of the disclosed systems, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0037] In some embodiments of the disclosed system, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes feature quantities of a substructure of the drug. In some embodiments, the feature quantities of a substructure of the drug include descriptors of the drug from the SMILES standard.
[0038] In some embodiments of the disclosed system, biological response data include in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof. In some embodiments, the IC50, AUC, or a combination thereof are converted to a normal distribution, and the normalized scores are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0–100%.
[0039] In some embodiments of the disclosed system, biological response data include in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergy scores obtained through the best single agent (HSA) model, the Loewe additive model, the Bliss independence model, or the zero interaction potency (ZIP) model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe synergy scores, HSA scores, Bliss scores, or ZIP scores. In some embodiments, the Loewe, HSA, Bliss, or ZIP synergy scores for in vitro drug combinations are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0%–100%.
[0040] In some embodiments of the disclosed system, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or a combination thereof. In some embodiments, efficacy assessments include tumor sizing, tumor volume assessments, efficacy biomarker assessments, response assessments, survival assessments, or a combination thereof. In some embodiments, response assessments, survival assessments, or a combination thereof are converted and reported as efficacy scores scaled to 0-1 or 0%-100%.
[0041] In some embodiments of the disclosed system, biological response data include pharmacological outcomes of xenotransplant drug monotherapy or drug combination therapy. In some embodiments, xenotransplant pharmacological outcomes include tumor sizing, tumor growth inhibition (TGI), AUC, or a combination thereof. In some embodiments, TGI or AUC are converted to a normal distribution and reported as standardized z-scores, or scaled to 0–1 or 0–100%.
[0042] In some embodiments of the disclosed system, at least one cancer drug discovery dataset includes data from cancer genomic drug sensitivity (GDSC).
[0043] In some embodiments of the disclosed system, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0044] In some embodiments of the disclosed system, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas Program (TCGA).
[0045] In some embodiments of the disclosed system, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0046] In some embodiments of the disclosed system, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0047] In some embodiments of the disclosed system, at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0048] In some embodiments of the disclosed system, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0049] In some embodiments of the disclosed system, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0050] In some embodiments, the disclosed system further includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0051] In some embodiments, the disclosed system further includes independently resampling data elements within each dataset.
[0052] In another aspect of the present disclosure, a non-temporary computer-readable medium is provided on which instructions are stored. When executed by a processor, the instructions cause the processor to train at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, the trained predictive model being capable of evaluating multiple drugs or drug combinations in terms of predicted biological responses to human cancer tissue exhibiting a particular set of cancer biomarkers, and to validate the predictive model by using a xenograft mouse model, in which human cancer tissue exhibiting a particular set of cancer biomarkers is transplanted in parallel into multiple immunodeficient mice, each of which is then treated in parallel with one of the multiple drugs or drug combinations, and the biological response to the treatment is fed back into the predictive model to further train the learning algorithm.
[0053] In some embodiments of the computer-readable media disclosed, human cancer tissue exhibiting a specific set of cancer biomarkers is orthotopically transplanted into multiple immunodeficient mice.
[0054] In some embodiments of the disclosed computer-readable media, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0055] In some embodiments of the disclosed computer-readable medium, for each class of cancer, at least one learning algorithm evaluates a drug or drug combination in relation to their biological response.
[0056] In some embodiments of the disclosed computer-readable medium, for each drug or drug combination under a cancer class, at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response.
[0057] In some embodiments of the disclosed computer-readable media, cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0058] In some embodiments of the computer-readable media of this disclosure, patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0059] In some embodiments of the disclosed computer-readable media, cancer biomarker data includes gene expression data. In some embodiments, gene expression data is normalized using TPM (transcripts per million). In some embodiments, gene expression data is normalized to the geometric mean of four housekeeping genes with small variance.
[0060] In some embodiments of the disclosed computer-readable media, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0061] In some embodiments of the disclosed computer-readable media, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes feature quantities of a substructure of the drug. In some embodiments, the feature quantities of a substructure of the drug include descriptors of the drug from the SMILES standard.
[0062] In some embodiments of the disclosed computer-readable media, the biological response data includes in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof. In some embodiments, the IC50, AUC, or a combination thereof are converted to a normal distribution, and the normalized scores are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0–100%.
[0063] In some embodiments of the disclosed computer-readable media, the biological response data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of synergy scores obtained through the best single agent (HSA) model, the Loewe additive model, the Bliss independence model, or the zero interaction potency (ZIP) model. In some embodiments, the in vitro drug combination screening results are in the form of Loewe, HSA, Bliss, or ZIP synergy scores. In some embodiments, the Loewe, HSA, Bliss, or ZIP synergy scores for in vitro drug combinations are reported as standardized z-scores, or as standardized and scaled to 0–1 or 0%–100%.
[0064] In some embodiments of the disclosed computer-readable media, biological response data include clinical study results. In some embodiments, clinical study results include safety assessments, efficacy assessments, dose-response assessments, pharmacodynamic assessments, pharmacokinetic assessments, progression-free survival assessments, or a combination thereof. In some embodiments, efficacy assessments include tumor sizing, tumor volume assessments, efficacy biomarker assessments, or a combination thereof. In some embodiments, response assessments, survival assessments, or a combination thereof are converted and reported as efficacy scores scaled to 0-1 or 0%-100%.
[0065] In some embodiments of the disclosed computer-readable media, biological response data include pharmacological results of xenotransplant drug monotherapy or drug combination therapy. In some embodiments, xenotransplant pharmacological results include tumor sizing, tumor growth inhibition (TGI), AUC, or a combination thereof. In some embodiments, TGI or AUC are converted to a normal distribution and reported as standardized z-scores or scaled to 0–1 or 0–100%.
[0066] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from cancer genomic drug sensitivity (GDSC).
[0067] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0068] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from NCI-ALMANAC.
[0069] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas Program (TCGA).
[0070] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0071] In some embodiments of the computer-readable media disclosed, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0072] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0073] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0074] In some embodiments of the disclosed computer-readable medium, at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0075] In some embodiments of the disclosed computer-readable media, at least one learning algorithm includes a feature selection algorithm, preferably a Boruta algorithm.
[0076] In some embodiments, the disclosed computer-readable medium further includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0077] In some embodiments, the disclosed computer-readable medium further includes independently resampling the data elements within each dataset.
[0078] In some embodiments of the disclosed methods, systems, and non-temporary computer-readable media, the predictive models do not provide cancer diagnosis.
[0079] In some embodiments of the disclosed methods, systems, and non-temporary computer-readable media, the predictive models do not provide cancer prognoses.
[0080] In some embodiments of the disclosed methods, systems, and non-temporary computer-readable media, the predictive models do not provide cancer diagnosis or prognosis.
[0081] The object and features of the present invention can be better understood by referring to the following detailed description and accompanying drawings. [Brief explanation of the drawing]
[0082] [Figure 1] This disclosure provides a schematic diagram of the method, system, and non-temporary computer-readable medium for constructing software to identify and validate predictive chemical molecular features and genetic biomarkers for cancer treatment. [Figure 2] This is a schematic diagram detailing the algorithms used to create software for identifying molecular chemical predictors and predictive genetic biomarkers for cancer treatment. Training data from each cancer type (X) is processed separately to build machine learning models for predicting treatments for each cancer type. [Figure 3] In A) the colon cancer treatment prediction model and B) the lung cancer treatment prediction model, we present in detail the results of feature optimization that shows only 5 to 10 genes that are shown to be necessary to achieve peak optimal prediction accuracy. [Figure 4]The present inventors' results regarding the accuracy of their in-vitro therapy prediction models, specifically A) a prediction model for monotherapy and B) a prediction model combining both monotherapy and combination therapy, are presented in detail. [Figure 5] The results of our machine learning model regarding clinical treatment response in cancer patients are shown. [Figure 6] This is an overview of a PDX model validation pharmacology experiment demonstrating the drug response of an orthotopic PDX model in triple-negative breast cancer to four predicted AI therapies. [Modes for carrying out the invention]
[0083] This disclosure provides methods, systems, and software for identifying and validating predictive biomarkers that predict human responses to cancer treatment.
[0084] Genomic and proteomic analyses provide rich information about the number and morphology of proteins expressed within cells, offering the possibility of identifying characteristic protein expression profiles for specific cellular states within each cell. In some cases, this cellular state may be a feature of abnormal physiological responses associated with disease. As a result, identifying cellular states from patients with disease and comparing them to corresponding cellular states from healthy patients can provide an opportunity to diagnose disease and control treatment.
[0085] Recent advances in transcription and proteomics profiling techniques have made it possible to apply computational methods to detect changes in expression patterns and their correlation with disease states, thereby facilitating the identification of markers that can contribute to multi-marker combinations with high-precision diagnostic capabilities.
[0086] While high-throughput screening methods provide large datasets of gene expression information, the challenge in bioinformatics remains developing robust methods for organizing data into reproducible and diagnostically viable patterns across diverse populations. A commonly accepted approach is to pool data from multiple sources to form a combined dataset, and then split that dataset into discovery / training and test / validation sets. However, both transcriptional profiling and protein expression profiling data are often characterized by a large number of variables relative to the number of available samples.
[0087] Observed differences between expression profiles of patient- or control-derived samples are typically masked by (1) biological variability or unknown subphenotypes within the disease or control population, (2) site-specific bias due to differences in research protocols, sample handling, etc., (3) bias due to differences in instrumentation (e.g., tip batches), and / or (4) variability due to measurement errors. False discovery of drug targets remains a serious problem, especially considering the costs and effort typically required for "post-discovery" work such as protein / gene identification and further validation of potential biomarkers.
[0088] This disclosure recognizes that systematic bias due to site-specific factors may only be detectable by carefully analyzing and comparing data from multiple sources. This disclosure provides systems, software, and methods for analyzing expression profiling data from multiple sources (e.g., clinical trial sites) to overcome potential systematic bias in expression data typically generated by such analyses, thereby reducing the probability of false discovery of drug targets. In a preferred embodiment, the invention combines the use of bioinformatics with expression profiling of samples from multiple sources to screen, identify, and validate biomarkers relating to a specific biological state or condition of interest. Measurement of these markers in patient samples can provide information that may be the presence or severity of a patient's condition or characteristic, such as a human. In one embodiment, the condition or characteristic may be the presence of a disease, a predisposition to disease, or a risk of recurrence.
[0089] In some embodiments, the Disclosure provides bioinformatics tools for analyzing expression profiling data from samples from two or more independent sources in a manner that reduces the variability and bias that can lead to the identification of false targets during the drug discovery process. In some embodiments of the Disclosure, data from multiple sources are not pooled together into a combined dataset and then split into discovery / training sets and test / validation sets. In some embodiments, data from multiple sources (e.g., multiple different clinical trial sites) are analyzed separately and independently of each other.
[0090] For each source, sufficient sample size and statistical resampling methods (e.g., bootstrap analysis) help discover biomarkers that function well in representative populations and consistently function well in different randomly selected subpopulations. The use of resampling procedures reduces biological variability and the combined effects of multiple variables in gene expression profiling data.
[0091] In some embodiments, the disclosure includes developing at least two different training sets (discovery datasets) developed independently of each other. Each training set includes subject data (data points) from multiple subjects. The subject data from each subject indicates the phenotype (class of biological conditions or pathology) to which the subject belongs, and each subject is classified into one of several different classes of pathologies. These different phenotypes are generally pathology-related, such as disease vs. normal, different disease stages, etc. However, they may also include any measurable biological characteristics. Each training set includes subject data from at least two subjects belonging to each of the phenotypes. The subject data from each subject includes measurements of multiple data elements from each subject sample.
[0092] In some embodiments, the results from analyses performed individually and independently are then cross-compared to achieve comparable performance for data from each individual source and to share the same upper / lower control pattern across different groups of samples from multiple sources.
[0093] Next, a multivariate classification model is developed that uses biomarkers selected from cross-comparison to classify a sample (e.g., patient cancer tissue) into one of a class or condition of biological status. This subset of potential biomarkers is preferably further validated using a separate, independent validation dataset. Furthermore, preferably, the identity of these potential biomarkers is identified and their performance is validated using additional samples and additional methods (including, but not limited to, immunoassays).
[0094] In one preferred embodiment, the expression profiling data to be evaluated is proteome profiling data (i.e., data relating to protein expression and their modifications and processed forms). For example, this method is particularly suitable for use in mass spectrometry-based analysis of the proteome. Thus, in one embodiment, the AI method disclosed herein is used for screening, identifying, and validating predictive biomarkers for cancer treatment. Data from independent datasets (e.g., the types of biomarkers expressed, the expression levels of each biomarker) are cross-compared to identify markers that predict one or more features of the dataset. Such features may include the presence of conditions shared by members of the dataset, such as the presence of disease.
[0095] The expression profiles (e.g., presence, quantity) of biomarkers in a sample (e.g., patient's cancer tissue) can be used to identify cells, tissues, organs, and / or the patient's condition. In certain embodiments, the expression profile of a single biomarker indicates the condition. In other embodiments, the expression profiles of multiple biomarkers indicate the condition.
[0096] definition As used herein, the following terms have the meanings set forth below unless otherwise specified.
[0097] As used herein and in the claims, the singular forms "a," "an," and "the" refer to multiple subjects unless the context clearly indicates otherwise. For example, the term "cell" refers to multiple cells, including mixtures thereof. The term "protein" refers to multiple proteins.
[0098] In this specification and as used throughout the following claims, the word “in” means both “inside” and “on the surface,” unless the context is clearly different.
[0099] The combinations described herein, for example, “at least one of A, B, or C,” “at least one of A, B, and C,” “at least one of A, B, and C,” and “A, B, C, or any combination thereof,” include any combination of A, B, and / or C, which may include multiple A's, multiple B's, or multiple C's. Specifically, “at least one of A, B, or C,” “at least one of A, B, or C,” “at least one of A, B, and C,” “at least one of A, B, and C,” and “A, B, C, or any combination thereof” may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, and any such combination may include one or more members of its constituent A, B, and / or C. For example, the combination of A and B may include one A and multiple B's, multiple A's and one B's, or multiple A's and multiple B's.
[0100] In the context of this invention, "biomarker" refers to a biomolecule, such as a protein, or its modified, cleaved, or fragmented form, nucleic acid, carbohydrate, metabolite, or intermediate, which is present differently in a sample and whose presence or amount indicates the state of the sample's source (e.g., cells, tissue, or patient). The term "biomarker" is used synonymously with the term "marker."
[0101] A "dataset" refers to a set of data where each element is a data point.
[0102] A "data point" refers to an element of a dataset, such as a subject sample, which is identified by a label or patient number that identifies the source of the sample.
[0103] A "biological state class" refers to a biological characteristic that allows data points to be classified. Each dataset containing data points 1-i has at least two data points that represent at least two forms of a biological state class: either present in the sample source providing the data points (class +1) or not present in the sample source providing the data points (class -1). In one embodiment, a class -1 data point represents a control (e.g., negative for the disease), but this is not necessarily the case. For example, in a particular embodiment, a class +1 sample represents a specific stage of the disease (e.g., malignant cancer), and a class -1 sample represents another stage of the disease (e.g., benign cells). What a state class represents is determined by the nature of the diagnostic test from which the biomarker is selected. Examples of classes of biological states include pathology (pathological vs. non-pathological (e.g., cancer vs. non-cancer)), drug response (drug responders vs. drug non-responders), toxicity response (toxic response vs. non-toxic response), prognosis (progression to disease state vs. non-progression to disease state), and most commonly, phenotype (presence of phenotypic state vs. absence of phenotypic state).
[0104] A "data element" refers to a feature of a data point that represents the characteristics of that data point. For example, in one aspect, a data element represents the expression levels of several different genes in a sample. In another aspect, a data element represents a peak detected by mass spectrometry. In yet another aspect, a data element represents various phenotypic characteristics, such as the level of any biologically significant analyte (e.g., a clinical chemistry or hematology test panel), responses to questions in an assessment test, or elements of medical history.
[0105] A "data element value" refers to a value assigned to a data element. The value can be qualitative or quantitative, and may be, for example, "presence or absence," "high, medium, or low," or a measured numerical value.
[0106] "Evaluating" a data element means assigning a value to a data element to which a selection criterion may be applicable.
[0107] A “selection criterion” refers to one or more criteria established by the user implementing the method and applied to the evaluation values to select data elements into an initial subset. The selection criterion can be a cutoff for numerical evaluation values or a class of qualitative evaluation values. Examples of cutoff criteria include “data elements within the top 10 percent of discriminative power” or “data elements that provide at least 80% specificity and at least about 70% sensitivity.” Examples of class criteria are “good” or “poor” data elements based on evaluation values, which depend to some extent on the nature of the class of biological condition in question, for diseases with few diagnostic markers. Data elements with low specificity or sensitivity may be selected by lower numerical or qualitative evaluation values. Initially, the selection criterion may be that, across multiple data points in the dataset, the data element is consistently better at identifying the class of biological condition than other data elements.
[0108] As used herein, the term “correspond / correspondence to” refers to the ability of a system or system component to receive input data from another system or system component and to provide an output response in response to the input data. “Output” may be in the form of data, or in the form of an action taken by the system or system component.
[0109] As used herein, “gene expression level” or “gene expression level” refers to the properties and quantity of the molecule encoded by the gene, e.g., RNA or polypeptide. The expression level of an mRNA molecule is intended to include the quantity of mRNA determined by the transcriptional activity of the gene encoding the mRNA, and the stability of the mRNA determined by its half-life. The gene expression level is also intended to include the quantity of polypeptide corresponding to a given amino acid sequence encoded by the gene. Thus, the gene expression level may correspond to the quantity of mRNA transcribed from that gene, the quantity of polypeptide encoded by that gene, or both. The expression levels of gene products can be further classified by the expression levels of different forms of gene products. For example, RNA molecules encoded by a gene may include differentially expressed splice variants, transcripts with different start or stop sites, and / or other differentially processed forms. Polypeptides encoded by a gene may include cleaved and / or modified forms of the polypeptide. Furthermore, there may be multiple forms of a polypeptide having a given type of modification. For example, a polypeptide can be phosphorylated at multiple sites, resulting in the expression of a protein with differentially phosphorylated proteins at different levels.
[0110] As used herein, “gene expression profile” refers to a characteristic representation of gene expression levels in a specimen, such as a cell or tissue. Determining a gene expression profile in a specimen derived from an individual represents the gene expression state of that individual. A gene expression profile reflects the expression of messenger RNA or polypeptides, or their forms encoded by one or more genes, in a cell or tissue. More generally, “expression profile” refers to a profile of biomolecules (nucleic acids, proteins, carbohydrates) that exhibit different expression patterns across different cells or tissues. The term “expression profile” encompasses the term “gene expression profile.”
[0111] As used herein, “computer program product” means a representation of a set of instructions, written in natural language or a programming language, stored on a physical medium of any nature (e.g., paper, electronic, magnetic, optical, etc.), and used with a computer or other automated data processing system (preferably based on digital technology). When such instructions in a programming language are executed by a computer or data processing system, they cause the computer or data processing system to operate according to the specific content of the instructions. Computer program products include, but are not limited to, programs in source and object code, and / or test or data libraries embedded in computer-readable media. Furthermore, computer program products that enable a computer system or data processing device to operate in a pre-selected manner are provided in a variety of forms, including, but not limited to, original source code, assembly code, object code, machine code, the aforementioned encrypted or compressed versions, and any equivalents.
[0112] Cancer detection dataset This disclosure provides a data element selection method that reduces the likelihood of selecting a classifier with discriminative power biased towards sampling differences rather than morphological differences in a class of biological states. In particular, the classifier may be a biomarker such as biomolecules exhibiting variability in expression profiling (transcriptional profiling, proteome profiling, etc.) and clinical sampling. In a preferred embodiment of the present invention, the biomarker is obtained from proteome analysis of a patient sample. However, the classifier may also be any other phenotypic trait.
[0113] Datasets may contain "false" classifiers / biomarkers, i.e., biases or pre-analysis variables that generate biomarkers that distinguish groups based on specific biases rather than on the underlying biological status of the subjects being studied. For example, if a dataset is gender-biased regarding the presence or absence of disease, certain highly discriminative classifiers / biomarkers may distinguish data points based on gender rather than disease. Similarly, if diseased samples and normal samples in a dataset are treated differently, classifiers / biomarkers may distinguish data points based on differences in handling rather than disease.
[0114] Independent datasets are less likely to have the same bias. Therefore, classifiers / biomarkers common to all independent datasets are more likely to identify based on the biological state of interest rather than some experimental bias. Thus, two datasets are independent if they were collected in a way that significantly reduces the likelihood of them being subject to the same bias. That is, datasets are independent if the populations used to obtain these datasets show a statistically significant difference with respect to at least one pre-analysis variable. The best way to reduce bias between datasets is to collect data points from different sites in different geographical locations. In this way, bias factors are randomized across different datasets and are therefore more likely to be excluded in crossover subsets of likely classifiers / biomarkers.
[0115] Additional or alternative methods to reduce bias include collecting data points at different time points and / or from different populations with respect to one or more non-limiting pre-analysis variables: e.g., sex, age, ethnicity, sample collection parameters, sample processing parameters, weight, diet, medication status, medical condition, exercise level, pregnancy and menstruation, presence and / or levels of circulating antibodies, clinical characteristics (e.g., PSA levels, cholesterol levels, family history, etc.). Preferably, the populations differ from many of the pre-analysis variables.
[0116] In selecting certain types of biomarkers (e.g., biomarkers associated with specific diseases), it may be particularly important to provide different populations with respect to specific pre-analysis variables. For example, when identifying biomarkers related to decreased protein C levels, it may be desirable to provide different populations with respect to other thrombotic risk factors.
[0117] This disclosure recognizes that, in some embodiments, identifying a profile exhibiting features such as the expression profile of a cell having a given cellular state can lead to the discovery of classifiers, such as biomarkers, that can be used to identify that cellular state with high probability (e.g., having a specificity of at least about 80% and a sensitivity of at least about 70% in a diagnostic test). Expression profiles may originate from the expression of nucleic acids (e.g., RNA transcripts, including differentially spliced or processed forms), proteins (including their modified and / or processed forms), carbohydrates (e.g., lectins), etc. In one embodiment, the cellular state reflects the patient's state from which the cell originates and serves as a diagnostic indicator of the physiological processes the patient is experiencing (e.g., pathological responses experienced when the patient has, is developing, or is recovering from a disease).
[0118] As a first step, multiple independent datasets are acquired. A dataset includes, for example, data points representing multiple samples from multiple sample sources, such as labels referring to sample numbers or patient numbers. Each dataset contains multiple forms of at least one class of biological conditions, and the multiple data points (samples) belong to each of the forms of that class. For example, a class of biological conditions may include, but is not limited to, the presence or absence of disease in the sample source (i.e., the patient from whom the sample was taken), the stage of the disease, the risk of the disease, the likelihood of disease recurrence, common genotypes at one or more loci (e.g., common HLA haplotypes, gene mutations, genetic modifications such as methylation), exposure to drugs (e.g., toxic or potentially toxic substances, environmental pollutants, candidate drugs, etc.) or conditions (e.g., temperature, pH, etc.), demographic characteristics (age, sex, weight, family history, medical history, etc.), drug tolerance, drug sensitivity (e.g., drug responsiveness), etc.
[0119] The datasets are independent of each other to reduce collection bias in the selection of the final classifier. For example, they may be collected from multiple sources, and may be collected at different times and in different locations using different exclusion or inclusion criteria. That is, the datasets may be relatively heterogeneous when considering features other than those that define the class of biological status. Factors contributing to heterogeneity include, but are not limited to, biological variability due to sex, age, and ethnicity; individual variability due to diet, exercise, and sleep behavior; and sample processing variability due to clinical protocols for blood processing. However, a class of biological status may contain one or more common features (for example, the sample source may represent individuals with the disease and the same sex, or one or more other common demographic characteristics).
[0120] In one embodiment, a dataset from multiple sources is generated by collecting samples from the same patient population at different times and / or under different conditions. However, a dataset from multiple sources does not constitute a subset of a larger dataset. That is, datasets from multiple sources are collected independently (e.g., from different sites and / or at different times and / or under different collection conditions).
[0121] In a preferred embodiment, multiple datasets are obtained from multiple different clinical trial sites, with each dataset containing multiple patient samples obtained from each individual site. Sample types include, but are not limited to, blood, serum, plasma, papillary aspirate, urine, tears, saliva, cerebrospinal fluid, lymph, cell and / or tissue lysates, laser-dissected tissue or cell samples, embedded cells or tissues (e.g., in paraffin blocks or frozen), and fresh or old samples (e.g., from dissection). Samples can be obtained, for example, from in vitro cell or tissue cultures. Alternatively, samples may originate from living organisms or populations of organisms, e.g., single-celled organisms. Therefore, for example, in a method for discovering biomarkers for a particular cancer, blood samples can be collected from subjects selected by independent groups at two different test sites, thereby providing samples for creating independent datasets.
[0122] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0123] In some embodiments of the disclosed method, for each class of cancer, at least one learning algorithm evaluates a drug or drug combination in relation to their biological response.
[0124] In some embodiments of the disclosed method, for each drug or drug combination under a class of cancer, at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response.
[0125] In some embodiments of the disclosed method, cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof. In some embodiments, cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0126] In some embodiments of the system of this disclosure, patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0127] In some embodiments of the disclosed method, cancer biomarker data includes gene expression data. In some embodiments, gene expression data is normalized using TPM (transcripts per million). In some embodiments, gene expression data is normalized to the geometric mean of four housekeeping genes with small variance.
[0128] In some embodiments of the disclosed method, cancer biomarker data includes gene methylation data. In some embodiments, cancer biomarker data includes protein biomarker data.
[0129] In some embodiments of the disclosed method, the drug or drug combination data includes the chemical structure of the drug. In some embodiments, the drug or drug combination data includes feature quantities of a substructure of the drug. In some embodiments, the feature quantities of a substructure of the drug include descriptors of the drug from the SMILES standard.
[0130] In some embodiments of the disclosed method, the biological response data includes in vitro drug screening results. In some embodiments, the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof.
[0131] In some embodiments of the disclosed method, the biological response data includes in vitro drug combination screening results. In some embodiments, the in vitro drug combination screening results are in the form of a best single agent (HSA) model, a Loewe additive model, or a Bliss independence model. In some embodiments, the in vitro drug combination screening results are in the form of a Loewe synergistic effect score.
[0132] In some embodiments of the disclosed methods, biological response data include clinical study results. In some embodiments, clinical study results include safety evaluations, efficacy evaluations, dose-response evaluations, pharmacodynamic evaluations, pharmacokinetic evaluations, progression-free survival evaluations, or a combination thereof. In some embodiments, efficacy evaluations include tumor sizing, tumor volume evaluation, efficacy biomarker evaluation, or a combination thereof.
[0133] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes genomic drug sensitivity (GDSC) data for cancer.
[0134] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from Drugcomb.org.
[0135] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from a cancer genome atlas program.
[0136] In some embodiments of the disclosed method, at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0137] In vitro cell line drug screening dataset In some embodiments, the disclosure utilizes drug compound, drug response, tissue type, and genetic molecular data from large-scale cell line sensitivity screenings such as CCLE, CTRP, and GDSC. These datasets include, among other things, data points having cell line molecular data, cell line drug response data, and cell line tissue type.
[0138] Molecular data from the CCLE, CTRP, and GDSC datasets were obtained directly from the source datasets. The molecular data included gene expression, copy number information, and mutation information. Expression data was obtained from next-generation RNA sequencing (RNAseq), and the values are continuous. Copy number information was obtained from whole-exome sequencing, and is continuous. Mutation information was obtained from whole-exome sequencing and is expressed as mutation frequency for each gene allele.
[0139] Drug information from the CCLE, CTRP, and GDSC datasets was obtained from the PharmacoGx package or directly from the source databases. Drug chemical structures were obtained from PubChem as a simplified molecular input line entry system (SMILES) and converted to drug molecular descriptors using RDKit.
[0140] Drug response data included three indicators: IC50, AUC, and survival rate at 1 μM. IC50 (median inhibitory concentration) and AUC (area under the drug response curve) information were obtained from the PharmacoGx package or directly from source databases (GDSC, CCLE, CTRP).
[0141] In vitro cell line drug combination dataset In some embodiments, the disclosure utilizes molecular information, drug response, tissue type, and genetic molecular data from large-scale cell line susceptibility screenings such as NCI-ALMANAC and Drugcombo.org. These datasets include, among other things, data points containing cell line molecular data, cell line drug response data, and cell line tissue type.
[0142] Molecular data from the NCI-ALMANAC and Drugcombo.org datasets were obtained directly from the source datasets or from CCLE and GDSC. Molecular data included gene expression, copy number information, and mutation information. Expression data was obtained from next-generation RNA sequencing (RNAseq), and the values are continuous. Copy number information was obtained from whole-exome sequencing, and is continuous. Mutation information was obtained from whole-exome sequencing and is expressed as mutation frequency for gene alleles.
[0143] Drug information was obtained from the source datasets NCI-ALMANAC and Drugcombo.org. Drug chemical structures were obtained from PubChem as a simplified molecular input line entry system (SMILES) and converted to drug molecular descriptors using RDKit.
[0144] Combination drug response values from the NCI-ALMANAC and Drugcombo.org datasets are obtained directly from the source datasets and include the following metrics: IC50, AUC, and synergy score. The synergy score is obtained from the best single agent (HSA) model, Loewe additive model, Bliss independence model, or zero interaction potency (ZIP) model.
[0145] Clinical research datasets In some embodiments, this disclosure uses molecular information, drug response, histological type, and genetic molecular data from clinical research data of the Cancer Genome Atlas Project (TCGA). These datasets include, among other things, data points including patient molecular data, patient drug response data, patient survival data, and patient cancer type.
[0146] The molecular data for the TCGA dataset was obtained from the Genomic Data Commons (GDC) data portal. The molecular data included gene expression, copy number information, and mutation information. Expression data was obtained from next-generation RNA sequencing (RNAseq), and the values are continuous. Copy number information was obtained from whole-exome sequencing, and is continuous. Mutation information was obtained from whole-exome sequencing and is expressed as mutation frequency for gene alleles.
[0147] Drug information for TCGA was obtained from the GDC Data Portal. Drug chemical structures were obtained from PubChem as a simplified molecular input line entry system (SMILES) and converted to drug molecular descriptors using RDKit.
[0148] Patient drug response to TCGA is defined under the criteria for evaluating treatment response in solid tumors (RECIST) and the Cheson criteria for hematological malignancies. The RECIST criteria include the following classifications: complete response (CR), partial response (PR), stable disease (SD), progression (PD), and not evaluable (NE). The Cheson criteria include the following classifications: complete response (CR), complete response not confirmed (CRu), partial response (PR), stable disease (SD), relapsed disease (RD), and progression (PD). Patient survival endpoints in TCGA are measured as continuous values on a daily basis.
[0149] Learning algorithms This disclosure describes how to train at least one learning algorithm used in a predictive model using a cancer detection dataset, and how the trained predictive model can evaluate multiple drugs or drug combinations in relation to their predicted biological responses to human cancer tissues exhibiting a specific set of cancer biomarkers.
[0150] In some of the disclosed methods, at least one learning algorithm includes a supervised learning algorithm. In some embodiments, at least one learning algorithm includes an unsupervised learning algorithm.
[0151] In some embodiments of the disclosed method, at least one learning algorithm includes a multivariate regression algorithm, which includes a random forest, a support vector machine, and an artificial neural network.
[0152] In some embodiments of the disclosed method, at least one learning algorithm includes a classification algorithm comprising a random forest, a support vector machine, and an artificial neural network.
[0153] In some embodiments of the disclosed method, at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0154] In some embodiments, the disclosed method further includes a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0155] In some embodiments, the disclosed method further includes independently resampling data elements within each dataset.
[0156] Selection of an initial subset Here, a subset of data elements, such as genes or proteins, is selected from each dataset based on selection criteria. Generally, the genes or proteins that are the "best" predictors are selected from each dataset. For example, the selection criteria might be the "top 10 percent" or "genes or proteins that provide a specified level of model predictive accuracy." All data elements from each dataset that meet the selection criteria are selected for the initial subset. For example, if each dataset has 100 ranked genes or proteins, the top 10 percent or the top 10 genes or proteins might each be selected for the initial dataset.
[0157] Selection of an Intersecting Subset In most cases, these initial subsets are not identical in terms of the data elements they load. However, if they contain common data elements, these elements can be selected within the crossover subset. For example, the initial subset from dataset 1 may contain genes or proteins 1, 3, 5, 7, and 9. The initial subset from dataset 2 may contain genes or proteins 1, 2, 3, 4, and 5. The crossover subset may contain any or all of genes or proteins 1, 3, and 5 as data elements common to both initial subsets.
[0158] More specifically, results from multiple datasets are cross-compared to determine a final set of common data elements with consistent expression patterns, forming a panel of potential biomarkers. Thus, data elements selected or determined to have good "values" or "weights" using the learning algorithm described above in independent discovery datasets are compared to select a cross-subset of data elements, where the data elements within the cross-subset have good values across multiple datasets; i.e., the data elements are consistently good biomarkers. In some cases, "good values" refer to data elements that have a specificity of at least 80% and a sensitivity of at least approximately 70% in a test for detecting or diagnosing a class of biological conditions.
[0159] Data point collection and data element generation Data points are collected to represent individual samples within a dataset. Each data point contains data elements. Multiple data points within a dataset are characterized by belonging to the same class of biological conditions. For example, each data point belonging to the same class of biological conditions may represent a sample from a patient identified as having the disease of interest for which a biomarker has been identified.
[0160] A data element is a feature of a data point that represents the characteristics of that data point. For example, in one embodiment, a data element represents the expression levels of several different genes in a sample from a patient having a disease that is commonly shared among patients contributing samples to the dataset. Expression levels can be obtained using any method for expression profiling known in the art and are included within the scope of this invention.
[0161] Data elements (e.g., gene expression values) can be obtained by transcriptional profiling and / or proteomic profiling. Transcriptional profiling techniques include, but are not limited to, synthetic sequencing (SBS), Northern blotting, qPCR and RT-PCR-based differential display methods, nuclease protection, differential retrieval analysis (RDA), suppression subtractive hybridization (SSH), and enzymatic degradation subtraction (EDS), gene array profiling, cDNA fingerprinting, subtractive hybridization, sequential analysis of gene expression, or SAGE. Proteomic profiling techniques include, but are not limited to, bihybrid analysis, fluorescence resonance energy transfer (MET), two-dimensional gel electrophoresis, mass spectrometry (e.g., laser desorption / ionization mass spectrometry), fluorescence (e.g., sandwich immunoassay), surface plasmon resonance, polarization analysis, and atomic force microscopy.
[0162] Other types of biomolecules expressed differentially may be profiled to provide data elements. For example, carbohydrates such as lectins (e.g., glycans) have diverse expression patterns that can provide data values for data elements containing data points.
[0163] A preferred method for expression profiling is high-throughput, acquiring data elements from approximately 10+, 50+, 100+, 200+, or 500+ samples within the dataset.
[0164] Preferably, the data elements are transformed to reduce the dimensionality of the features using algorithms such as UMAP or PCA. For example, 20,000 genes can be reduced to just a few hundred dimensions.
[0165] Preferably, the data element is represented as a vector of numbers containing values representing the level of the sample component represented by the data element and at least one other characteristic of the sample component / data element, such as its name or descriptor.
[0166] Evaluation of data elements In the next step, the data elements obtained from expression profiling are evaluated using any type of multivariate analysis. One method involves using pattern recognition processes, such as classification models, for eligibility.
[0167] A classification model can be trained on pre-classified "known data elements" (e.g., cancerous or non-cancerous, type of cancer). The data elements used to form the classification model may be referred to as the "training dataset" or "discovery dataset." Once trained, the classification model can recognize patterns in the data derived from data elements from unknown samples. The classification model can then be used to classify unknown samples into classes. This can be useful, for example, in predicting whether a particular biological sample is associated with a particular biological condition (e.g., whether it has a disease or not, and the type of disease).
[0168] A classification model can be formed using any appropriate statistical classification (or "learning") method that attempts to separate the body of data into classes based on objective parameters present in the data. The classification method can be either supervised or unsupervised. Supervised and unsupervised classification processes are known in the art and are outlined, for example, in Jain, IEEE Transactions on Pattern Analysis and Machine Technology 22(1):4-37, 2000. When selecting a classification method, a balance must be struck between minimizing the risk of losing useful information and reducing the number of data elements to simplify the analysis.
[0169] Unsupervised classification attempts to learn classifications based on the similarity of discovery / training datasets, without pre-classifying the data elements (e.g., expression data) from which the training dataset is derived.
[0170] Unsupervised learning methods include cluster analysis. Cluster analysis attempts to divide data into “clusters” or groups that are, ideally, very similar to each other and very different from the members of other clusters.
[0171] In supervised classification, training data containing examples of known categories is presented to a learning mechanism, which uses a learning algorithm to learn one or more sets of relationships that define each of the known classes. New data may then be applied to the learning mechanism, which then uses the learned relationships to classify the new data. Differentially expressed sample components (i.e., defining the data elements of data points) and identities (i.e., labels corresponding to sample number / patient count) can be identified by using a set of data elements whose values represent the expression of sample components, as pre-known training data. Supervised learning techniques derive a classification model (classifier) that assigns data elements taken from multiple data points to a predefined number of known classes with the smallest possible error. The contribution of individual variables to the classification model is then analyzed as a measure of the values of the data elements, i.e., which data elements are likely to function as biomarkers with good discriminative power (i.e., the ability of a biomarker to distinguish between data points with a biological state and data points without a biological state). Each common data element within each data point is evaluated independently of each dataset, as a function of the data element values, based on the data element's ability to classify the data point into a class of biological states.
[0172] There are various approaches to deriving classification models, and generally, the type of classification approach used is not a feature that limits the present invention.
[0173] In some embodiments, the present disclosure integrates resampling procedures into the evaluation of expression data to mitigate the effects of variability between samples within a dataset (e.g., patient-derived samples from clinical trial sites) and between different datasets (e.g., samples from patients at various clinical trial sites using different exclusion and selection criteria, and sampling of populations with different demographic characteristics). Resampling methods such as bootstrapping, bagging, boosting, and Monte Carlo simulation are preferably applied in the context of supervised learning, for example, using the learning algorithms disclosed herein.
[0174] Therefore, in one embodiment, multiple datasets are independently and repeatedly divided into subsets containing test data points (class + 1 data points) and compared with reference or control data points (class - 1 data points).
[0175] In each resampling run, data elements(s) that significantly and consistently contribute to separating data points with at least one common feature from data points that do not share this feature are selected. That is, to identify biomarkers diagnosed as having at least one common feature. Parameters such as the mean, variance, and confidence interval of the sampled data elements (e.g., confidence score of expression data) are measured to determine the distribution of parameters, identify outlier scores, and form a short list of candidate biomarkers represented by the data elements. For example, expression values (such as sequence read counts) with high mean ranks and small standard deviations may be selected for this list. Performing such analyses independently for each of multiple datasets reduces the likelihood of selecting data elements as a result of data bias or artifacts, thereby reducing the possibility of misidentifying biomarkers.
[0176] This method identifies data elements with high confidence (selected differences from a randomized distribution are accepted as statistically significant with minimal false discoveries, e.g., FDR ≤ 0.05) and qualitatively expresses them in the same way (overexpression or underexpression in both datasets). Highly reliable outliers are ranked from those showing the largest difference in expression between test data points and reference data points (i.e., most diagnostic) to those showing the smallest difference (i.e., least diagnostic).
[0177] Therefore, for example, gene or protein expression data from a collection of samples may yield expression data for more than 20,000 genes or proteins, each being a data element, with its measured expression level being the data element value. After applying the dataset to the selected analysis format, specific samples (data points) are classified as cancerous or non-cancerous based on their expression levels, and if cancerous, their ability to belong to a particular cancer class (a form of biological condition) is determined or "evaluated." Each gene or protein may then be ranked from most discriminative to least discriminative.
[0178] Use of learning algorithms in predictive models Data elements within a crossover subset can be used in a multivariate model to generate a multivariate regression algorithm. To build a multivariate predictive model, data from multiple datasets is combined and randomly split into discovery / training and test sets.
[0179] The performance of a panel of potential biomarkers identified through resampling and cross-comparison, and the derived predictive models, is evaluated on their test set to identify biomarkers that still retain high diagnostic capability for at least one common feature. The predictive models are validated against independent data elements from one or more novel datasets that share at least one common feature and were not involved in the biomarker discovery and model building process. Independent validation may be performed against the datasets being analyzed, either by including data points from a larger population or by using different methods (e.g., using expression profiling techniques different from those initially used to acquire data elements, such as immunoassays), thereby obtaining a validation training set that can be used to identify the most highly discriminative biomarkers among those being tested. Statistical methods for evaluating the validation datasets include sensitivity and specificity estimation, as well as receiver operating characteristic (ROC) curve analysis.
[0180] The multivariate regression algorithm thus generated can be tested against a separate, independent "validation" dataset to determine the algorithm's final power. The validation dataset must be independent of all the discovery datasets used to discover the biomarkers from which the regression algorithm was generated.
[0181] Biomarkers can be evaluated after resampling, but more preferably after cross-comparison, to identify additional features of the biomarkers that can be used to characterize the validation dataset. For example, sequence information of peptide or nucleic acid biomarkers may be determined. Additional features may be used to generate probes to test for the presence of the biomarker in test samples (novel data points) within the dataset used to validate the biomarkers. Additional features may include sequence data for larger sequences (e.g., sequence data of the gene or protein from which the nucleic acid or peptide originates) in which the biomarker sequence is a subsequence. Such data may be obtained by querying databases, e.g., gene sequence databases, protein sequence databases, or carbohydrate databases, using the biomarker sequence. Using this method, sequences of other markers can be identified if they are known in the database.
[0182] Preferably, a data element is identified as a biomarker if it can predict the presence or absence of a characteristic of a component of the dataset with an accuracy of more than 70%, preferably more than 80%, and more preferably more than 90%. In certain embodiments, combining multiple data elements can provide a desired predictive value. In certain embodiments, a combination with a high predictive value may include a data element with low confidence, and may be more predictive than a single data element with a higher confidence value. Combinations of data elements suitable for use as biomarkers may be identified, for example, by pairing them using an ordered or random approach.
[0183] Verification using xenograft mouse models The predictive model utilizing at least one learning algorithm provided in this disclosure is further validated by using a xenograft mouse model in which human cancer tissue exhibiting a specific set of cancer biomarkers is transplanted in parallel into multiple immunodeficient mice, each of which is then treated in parallel with one of multiple drugs or drug combinations, and the biological response to the treatment is fed back into the predictive model to further train the learning algorithm. In some embodiments of the disclosed method, human cancer tissue exhibiting a specific set of cancer biomarkers is orthotopically transplanted into multiple immunodeficient mice.
[0184] Patient-derived xenografts (PDX) are cancer models in which tumor-derived tissue or cells from a human patient are transplanted into immunodeficient or humanized mice. PDX models are used to create an environment that allows for the natural growth of cancer, its monitoring, and corresponding treatment evaluation in the original patient and in patients with similar cancer profiles.
[0185] Tumor xenotransplantation To establish a PDX model, several types of immunodeficient mice can be used: thymus-deficient nude mice, severely immunodeficient (SCID) mice, NOD-SCID mice, and recombinant activating gene 2 (Rag2) knockout mice. The mice used must be immunodeficient to prevent graft rejection. NOD-SCID mice are considered to be more immunodeficient than nude mice, and therefore, because they do not produce natural killer cells, they are more commonly used in PDX models.
[0186] When a human tumor is excised, necrotic tissue is removed, and the tumor can be mechanically sectioned into smaller fragments, chemically digested, or physically manipulated to form a single-cell suspension. There are advantages and disadvantages to using either individual tumor fragments or single-cell suspensions. Tumor fragments mimic the tumor microenvironment, as they retain intercellular interactions and some of the tissue structures of the original tumor. Alternatively, single-cell suspensions allow researchers to collect unbiased sampling of the entire tumor, thereby eliminating spatially isolated subclones that might otherwise be accidentally selected during analysis or tumor passage. However, single-cell suspensions expose viable cells to strong chemical or mechanical forces, which can increase their sensitivity to anokis and significantly impact cell viability and engraftment success.
[0187] Ectopic and orthotopic transplants Unlike the creation of xenograft mouse models using existing cancer cell lines, there are no intermediate in vitro processing steps before transplanting tumor fragments into a mouse host to create PDX. Tumor fragments are transplanted ectopically or orthotopically into immunodeficient mice. In ectopic transplantation, tissue or cells are transplanted into a region of the mouse unrelated to the original tumor site, generally subcutaneously or subrenal. The advantages of this method are direct access for transplantation and ease of monitoring tumor growth. In orthotopic transplantation, researchers transplant the patient's tumor tissue or cells into the corresponding anatomical location in the mouse. Subcutaneous PDX models do not induce metastasis in mice, do not mimic the initial tumor microenvironment, and have a graft survival rate of 40-60%. Subrenal PDX maintains the original tumor matrix and equivalent host matrix, with a graft survival rate of 95%. Ultimately, it takes approximately 2-4 months for the tumor to engraft, varying depending on the tumor type, transplantation site, and the strain of immunodeficient mice used. Graft failure should not be judged until at least 6 months. Researchers can use ectopic transplantation for the initial engraftment from the patient to the mouse, and then use orthotopic transplantation to transplant the tumor grown in the mouse to subsequent generations of mice.
[0188] The generation that has taken root The first generation of mice that receive tumor fragments from a patient are generally designated as F0. Once the tumor in the F0 mice reaches a sufficient size, researchers pass the tumor to the next generation of mice. Each subsequent generation is designated as F1, F2, F3…Fn. In drug development research, it is common to grow mice from the F3 to F10 generations to confirm that the PDX is not genetically or histologically deviant from the patient's tumor.
[0189] Superiority over cancer cell lines in predictive model validation While not intended to be bound by any particular theory, this disclosure recognizes one or a combination of several factors that make the PDX mouse model a better validation tool than cancer cell line screening or cancer cell line-derived xenograft (CDX) models. Firstly, cancer cell lines, originally derived from a patient's tumor, acquire the ability to grow in in vitro cell culture. As a result of in vitro manipulation, cell lines conventionally used in cancer research undergo genetic transformations that are not reversible when the cells are grown in vivo. Due to the cell culture process, including the enzymatic environment and centrifugation, cells better adapted to survive in culture are selected, tumor resident cells and proteins that interact with cancer cells are eliminated, and the culture becomes phenotypically homogeneous.
[0190] Secondly, when transplanted into immunodeficient mice, the cell lines do not readily develop tumors, and the resulting tumors, unlike heterogeneous patient tumors, are genetically branched. Researchers are beginning to attribute the lack of tumor heterogeneity and the absence of the human stromal microenvironment to why only 5% of anticancer drugs are approved by the U.S. Food and Drug Administration (FDA) after preclinical trials. Specifically, because the cell lines do not follow the pathways of microenvironmental influence on drug resistance or drug response seen in primary human tumors, cell line-xenografts often cannot predict the drug response in primary tumors.
[0191] Thirdly, many PDX models have been successfully established for breast cancer, prostate cancer, colorectal cancer, lung cancer, and many other cancers because there are significant advantages to using PDX over cell lines for drug safety and efficacy testing, and for predicting patient tumor responses to specific antitumor drugs. Since PDX can be passaged without in vitro processing steps, PDX models allow for the propagation and growth of patient tumors across multiple mouse generations without significant genetic transformation of tumor cells. Within PDX models, patient tumor samples grow in a physiologically relevant tumor microenvironment that mimics the levels of oxygen, nutrients, and hormones found in the patient's primary tumor site. Furthermore, transplanted tumor tissue maintains the genetic and epigenetic abnormalities found in the patient, and xenograft tissue can be excised from the patient to include the surrounding human stroma. As a result, many studies have found that PDX models exhibit responses to anticancer drugs similar to those seen in actual patients who provided tumor samples.
[0192] Receiving program and system The training datasets and classification models according to embodiments of the present invention can be embodied by computer code executed or used by a digital computer. The computer code can be stored on any suitable computer-readable medium, including optical or magnetic disks, sticks, tapes, and other transmission media such as digital and analog, and can be written in any suitable computer programming language, including C, C++, Java, and Python.
[0193] The output data obtained from the training can be displayed on a digital computer or on any graphical display interface on a user device that can connect to a server to which such a computer is connected (e.g., via the Internet). Suitable digital computers include microcomputers, minicomputers, or large computers using any standard or proprietary operating system, such as Unix®, Windows®, or Linux®-based operating systems. The digital computer used may be physically separated from the equipment used to acquire values for data elements in the profiling experiment. For example, the computer may be separated from the mass spectrometer used to create the spectrum of interest, or it may be coupled to the mass spectrometer. The graphical interface may also be remote from the computer and may be, for example, part of a wireless device that can connect to a network.
[0194] This disclosure also includes a computer system having a database containing data elements / biomarker features characteristic of various cancers. In one embodiment, the cellular state includes one or more of the following: stage of differentiation: phenotypic expression, cell cycle proliferation or stage, response to stimuli, disease, drugs (e.g., toxins or potentially toxic drugs, known or candidate drugs, antibiotics, infectious or pathological organisms, environmental pollutants, etc.), or conditions (e.g., temperature, pH, etc.). In another embodiment, the cellular state reflects the state of the cell's source. For example, the cellular state may reflect the disease or other physiological response(s) or condition(s) (e.g., old age, mental illness, addiction, allergic reaction, etc.) experienced by the patient from whom the cell originates.
[0195] In one embodiment, the database includes ranked or clustered biomarkers (i.e., biomarkers divided into subsets based on their discriminative power). Biomarkers may be ranked or clustered according to their association with various parameters. Such parameters may include responses to toxins, diseases, contaminants, conditions, stressors, developmental stages, drugs, therapeutic agents, antibiotics, etc. The database includes biomarkers that exhibit a relatively narrow range of variability in the population for a given cellular state, but have a high degree of discriminative power between cellular states. For example, the biomarkers are reproducibly associated with parameters (with a specificity of at least over 80% and a sensitivity of at least about over 70% in tests for detecting or diagnosing the parameters) and have high discriminative power.
[0196] However, it should be noted that discriminative power is not a limiting characteristic of biomarkers. For example, in the case of certain diseases for which there are few or no satisfactory diagnostic tests, biomarkers with lower specificity and / or sensitivity may still be valuable.
[0197] This system also includes a database management system. User requests or queries are formatted in an appropriate language that is understood by the database management system, which processes the queries and extracts relevant information from the training set database.
[0198] The system may further include records from an external database, or may communicate with such an external database. Preferably, the system is connectable to a network to which a network server and one or more clients are connected. The network may be a local area network (LAN) or a wide area network (WAN), as is well known in the art. Preferably, the server includes hardware necessary to run computer program products (e.g., software) to access database data for processing user requests. For example, one type of user request may be one to which the system identifies biomarkers associated with a selected cell state. Such a request may provide optional data options, such as sources of probes that can be used to detect one or more biomarkers (e.g., links to sites that provide binding partners for biomarkers such as antibodies).
[0199] The system also includes an operating system (e.g., UNIX or Linux) to execute instructions from the database management system. In one embodiment, the operating system also runs a World Wide Web application and a World Wide Web server, thereby connecting the server to the network.
[0200] Preferably, the system includes one or more user devices, each containing a graphical display interface, which includes interface elements such as buttons, pull-down menus, scroll bars, and fields for entering text, commonly found in graphical user interfaces known in the art. Requests entered on the user interface are sent to an application program within the system (such as a web application) and formatted to retrieve relevant information from one or more of the system databases. The requests or queries entered by the user may be constructed in any suitable database language (e.g., Sybase or Oracle SQL). In one embodiment, a user of a user device within the system can directly access the data using an HTML interface provided by the system's web browser and web server.
[0201] A graphical user interface may be generated by graphical user interface code as part of the operating system and may be used to input data and / or display the input data. The results of the processed data may be displayed on the interface, printed by a printer communicating with the system, stored in a memory device, and / or transmitted over a network, or provided in the form of computer-readable media.
[0202] Non-limiting embodiments Provided below are specific, non-limiting embodiments of different aspects of this disclosure.
[0203] 1. A method for training artificial intelligence using at least one hardware processor to identify one or more predictive biomarkers for cancer treatment, the method comprising: training at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, the trained predictive model being capable of evaluating multiple drugs or drug combinations in terms of predicted biological responses to human cancer tissue exhibiting a particular set of cancer biomarkers; and validating the predictive model by using a xenograft mouse model, in which the human cancer tissue exhibiting a particular set of cancer biomarkers is transplanted in parallel into multiple immunodeficient mice, each of which is subsequently treated in parallel with one of the multiple drugs or drug combinations, and the biological response to the treatment is fed back into the predictive model to further train the learning algorithm.
[0204] 2. The method according to Embodiment 1, wherein the human cancer tissue exhibiting a specific set of cancer biomarkers is orthotopically transplanted into the plurality of immunodeficient mice.
[0205] 3. The method according to Embodiment 1 or 2, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0206] 4. The method according to Embodiment 3, wherein for each class of cancer, the at least one learning algorithm evaluates a drug or drug combination in relation to its biological response.
[0207] 5. The method according to Embodiment 4, wherein, for each drug or drug combination under the class of cancer, the at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response.
[0208] 6. The method according to any one of Embodiments 3 to 5, wherein the cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof.
[0209] 7. The method according to Embodiment 6, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0210] 8. The method according to Embodiment 6, wherein the patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0211] 9. The method according to any one of Embodiments 3 to 8, wherein the cancer biomarker data includes gene expression data.
[0212] 10. The method according to Embodiment 9, wherein the gene expression data is normalized using TPM (transcripts per million).
[0213] 11. The method according to any one of claims 9 to 10, wherein the gene expression data is normalized to the geometric mean of at least four housekeeping genes with small variance.
[0214] 12. The method according to any one of Embodiments 3 to 11, wherein the cancer biomarker data includes gene methylation data.
[0215] 13. The method according to any one of Embodiments 3 to 12, wherein the cancer biomarker data includes protein biomarker data.
[0216] 14. The method according to any one of Embodiments 3 to 13, wherein the drug or drug combination data includes the chemical structure of the drug.
[0217] 15. The method according to any one of Embodiments 3 to 14, wherein the drug or drug combination data includes feature quantities of a substructure of the drug.
[0218] 16. The method according to Embodiment 15, wherein the feature quantities of the substructure of the drug include descriptors of the drug from the SMILES standard.
[0219] 17. The method according to any one of Embodiments 3 to 16, wherein the biological response data includes in vitro drug screening results.
[0220] 18. The method according to Embodiment 17, wherein the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof.
[0221] 19. The method according to any one of Embodiments 3 to 18, wherein the biological response data includes in vitro drug combination screening results.
[0222] 20. The method according to Embodiment 19, wherein the in vitro drug combination screening results are in the form of a synergistic effect score obtained through the best single agent (HSA) model, the Loewe additive model, the zero interaction potency (ZIP) model, or the Bliss independence model.
[0223] 21. The method according to Embodiment 19, wherein the in vitro drug combination screening results are in the form of an AUC score.
[0224] 22. The method according to any one of Embodiments 3 to 21, wherein the biological response data includes clinical research results.
[0225] 23. The method according to Embodiment 22, wherein the clinical research results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
[0226] 24. The method according to Embodiment 23, wherein the efficacy evaluation includes tumor size measurement, tumor volume assessment, efficacy biomarker evaluation, or a combination thereof.
[0227] 25. The method according to Embodiments 1 to 24, wherein the at least one cancer drug discovery dataset includes cancer genomic drug sensitivity (GDSC).
[0228] 26. The method according to Embodiments 1 to 25, wherein the at least one cancer drug discovery dataset includes data from NCI Almanac or Drugcomb.org.
[0229] 27. The method according to any one of Embodiments 1 to 26, wherein the at least one cancer drug discovery dataset includes data from the Cancer Genome Atlas Program.
[0230] 28. The method according to any one of Embodiments 1 to 27, wherein the at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0231] 29. The method according to any one of Embodiments 1 to 28, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0232] 30. The method according to any one of Embodiments 1 to 28, wherein the at least one learning algorithm includes an unsupervised learning algorithm.
[0233] 31. The method according to any one of Embodiments 1 to 30, wherein the at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0234] 32. The method according to any one of embodiments 1 to 31, wherein the at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0235] 33. The method according to any one of Embodiments 1 to 32, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0236] 34. A method according to any one of embodiments 1 to 33, further comprising a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0237] 35. The method according to any one of Embodiments 1 to 34, further comprising independently resampling data elements within each dataset.
[0238] 36. A system comprising: an input unit capable of downloading at least one cancer drug discovery dataset; and a training unit capable of using the at least one drug discovery dataset to train at least one learning algorithm in a predictive model, wherein the trained predictive model is capable of evaluating multiple drugs or drug combinations in terms of predicted biological responses to human cancer tissue exhibiting a particular set of cancer biomarkers; the training unit is capable of further training the learning algorithm by using a xenograft mouse model, wherein in the xenograft mouse model, the human cancer tissue exhibiting a particular set of cancer biomarkers is transplanted in parallel into multiple immunodeficient mice, each of which is then treated in parallel with one of the multiple drugs or drug combinations, and the biological response to the treatment is fed back to the predictive model.
[0239] 37. The system according to Embodiment 36, wherein the human cancer tissue exhibiting a specific set of cancer biomarkers is orthotopically transplanted into the plurality of immunodeficient mice.
[0240] 38. The system according to Embodiment 37, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0241] 39. The system according to Embodiment 38, wherein for each cancer type, the at least one learning algorithm evaluates a drug or drug combination in relation to the biological response to that cancer type.
[0242] 40. The system according to Embodiment 39, wherein for each drug or drug combination under each cancer type, the at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response of the cancer type.
[0243] 41. The system according to any one of Embodiments 38 to 40, wherein the cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof.
[0244] 42. The system according to Embodiment 41, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0245] 43. The system according to Embodiment 41, wherein the patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0246] 44. The system according to any one of embodiments 38 to 43, wherein the cancer biomarker data includes gene expression data.
[0247] 45. The system according to Embodiment 44, wherein the gene expression data is normalized using TPM (transcripts per million).
[0248] 46. The system according to any one of claims 44 to 45, wherein the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0249] 47. The system according to any one of embodiments 38 to 46, wherein the cancer biomarker data includes gene methylation data.
[0250] 48. The system according to any one of embodiments 38 to 47, wherein the cancer biomarker data includes protein biomarker data.
[0251] 49. The system according to any one of Embodiments 38 to 48, wherein the drug or drug combination data includes the chemical structure of the drug.
[0252] 50. The system according to any one of embodiments 38 to 49, wherein the drug or drug combination data includes feature quantities of the substructure of the drug.
[0253] 51. The system according to Embodiment 50, wherein the feature quantities of the substructure of the drug include descriptors of the drug from the SMILES standard.
[0254] 52. The system according to any one of embodiments 38 to 51, wherein the biological response data includes in vitro drug screening results.
[0255] 53. The system according to Embodiment 52, wherein the in vitro drug screening results are in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof.
[0256] 54. The system according to any one of embodiments 38 to 53, wherein the biological response data includes in vitro drug combination screening results.
[0257] 55. The system according to Embodiment 54, wherein the in vitro drug combination screening results are in the form of a synergistic effect score obtained through the best single agent (HSA) model, the Loewe additive model, the zero interaction potency (ZIP) model, or the Bliss independence model.
[0258] 56. The method according to Embodiment 54, wherein the in vitro drug combination screening results are in the form of an AUC score.
[0259] 57. The system according to any one of embodiments 38 to 56, wherein the biological response data includes clinical research results.
[0260] 58. The system according to Embodiment 57, wherein the clinical research results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
[0261] 59. The system according to Embodiment 58, wherein the effectiveness evaluation includes tumor size measurement, tumor volume assessment, effectiveness biomarker evaluation, or a combination thereof.
[0262] 60. The system according to Embodiments 36 to 59, wherein the at least one cancer drug discovery dataset includes cancer genomic drug sensitivity (GDSC).
[0263] 61. The system according to embodiments 36-60, wherein the at least one cancer drug discovery dataset includes data from NCI Almanac or Drugcomb.org.
[0264] 62. The system according to any one of embodiments 36 to 61, wherein the at least one cancer drug discovery dataset includes data from a cancer genome atlas program.
[0265] 63. The system according to any one of embodiments 36 to 62, wherein the at least one cancer drug discovery dataset includes data from the Genotype Tissue Expression Project (GTEx).
[0266] 64. The system according to any one of embodiments 36 to 63, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0267] 65. The system according to any one of embodiments 36 to 63, wherein the at least one learning algorithm includes an unsupervised learning algorithm.
[0268] 66. The system according to any one of embodiments 36 to 65, wherein the at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0269] 67. The system according to any one of embodiments 36 to 66, wherein the at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0270] 68. The system according to any one of embodiments 36 to 67, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0271] 69. A system according to any one of embodiments 36 to 68, further comprising a dimensionality reduction algorithm, preferably UMAP, PCA, or both.
[0272] 70. The system according to any one of embodiments 36 to 69, wherein the system is capable of independently resampling data elements within each dataset.
[0273] 71. A non-temporary computer-readable medium storing instructions, wherein, when executed by a processor, the instructions cause the processor to: train at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the trained predictive model is capable of evaluating multiple drugs or drug combinations in terms of predicted biological responses to human cancer tissue exhibiting a particular set of cancer biomarkers; and validate the predictive model by using a xenograft mouse model, wherein in the xenograft mouse model, the human cancer tissue exhibiting the particular set of cancer biomarkers is transplanted in parallel into a plurality of immunodeficient mice, each of which is then treated in parallel with one of the plurality of drugs or drug combinations, and the biological response to the treatment is fed back to the predictive model; the non-temporary computer-readable medium.
[0274] 72. The non-transient computer-readable medium according to Embodiment 71, wherein the human cancer tissue exhibiting a specific set of cancer biomarkers is orthotopically transplanted into the plurality of immunodeficient mice.
[0275] 73. The non-temporary computer-readable medium according to Embodiment 72, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
[0276] 74. A non-temporal computer-readable medium according to Embodiment 73, wherein for each cancer type, the at least one learning algorithm evaluates a drug or drug combination in relation to their biological response to the cancer type.
[0277] 75. A non-temporary computer-readable medium according to Embodiment 74, wherein, for each drug or drug combination under a cancer type, the at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response of the cancer type.
[0278] 76. The cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof, in a non-temporary computer-readable medium according to any one of embodiments 73 to 75.
[0279] 77. A non-temporary computer-readable medium according to Embodiment 76, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
[0280] 78. A non-temporary computer-readable medium according to Embodiment 76, wherein the patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
[0281] 79. The cancer biomarker data includes gene expression data in a non-temporary computer-readable medium according to any one of embodiments 73 to 78.
[0282] 80. The gene expression data is normalized using TPM (transcripts per million) in the non-temporary computer-readable medium according to Embodiment 79.
[0283] 81. The non-temporal computer-readable medium according to any one of claims 79 to 80, wherein the gene expression data is normalized to the geometric mean of four housekeeping genes with low variance.
[0284] 82. The cancer biomarker data includes gene methylation data in a non-temporary computer-readable medium according to any one of embodiments 73 to 81.
[0285] 83. The cancer biomarker data is the non - transient computer - readable medium according to any one of embodiments 73 to 82, including protein biomarker data.
[0286] 84. The drug or drug combination data is the non - transient computer - readable medium according to any one of embodiments 73 to 83, including the chemical structure of the drug.
[0287] 85. The drug or drug combination data is the non - transient computer - readable medium according to any one of embodiments 73 to 84, including feature quantities of the partial structure of the drug.
[0288] 86. The feature quantities of the partial structure of the drug are the non - transient computer - readable medium according to embodiment 85, including descriptors from the SMILES specification of the drug.
[0289] 87. The biological response data is the non - transient computer - readable medium according to any one of embodiments 73 to 86, including in vitro drug screening results.
[0290] 88. The in vitro drug screening results are in the form of IC50 (half - maximal inhibitory concentration), AUC (area under the drug - response curve), or a combination thereof, which is the non - transient computer - readable medium according to embodiment 87.
[0291] 89. The biological response data is the non - transient computer - readable medium according to any one of embodiments 73 to 88, including in vitro drug combination screening results.
[0292] 90. The in vitro drug combination screening results are in the form of synergy scores obtained through the highest single agent (HSA) model, Loewe additivity model, zero interaction potency (ZIP) model, or Bliss independence model, which is the non - transient computer - readable medium according to embodiment 89.
[0293] 91. The in vitro drug combination screening results are in the form of an AUC score in the non-temporary computer-readable medium according to Embodiment 89.
[0294] 92. The biological response data, including clinical research results, is provided in a non-temporary computer-readable medium as described in any one of embodiments 73 to 91.
[0295] 93. The clinical research results, including safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof, are presented in a non-temporary computer-readable medium as described in Embodiment 92.
[0296] 94. The efficacy evaluation includes tumor size measurement, tumor volume assessment, efficacy biomarker assessment, or a combination thereof, in a non-temporary computer-readable medium as described in Embodiment 93.
[0297] 95. The at least one cancer drug discovery dataset is a non-temporary computer-readable medium according to Embodiments 71-94, comprising cancer genomic drug sensitivity (GDSC).
[0298] 96. The at least one cancer drug discovery dataset is a non-temporary computer-readable medium according to Embodiments 71-95, comprising data from NCI Almanac or Drugcomb.org.
[0299] 97. The at least one cancer drug discovery dataset is a non-temporary computer-readable medium according to any one of embodiments 71 to 96, comprising data from a cancer genome atlas program.
[0300] 98. The at least one cancer drug discovery dataset is a non-temporary computer-readable medium according to any one of embodiments 71 to 97, comprising data from the Genotype Tissue Expression Project (GTEx).
[0301] 99. The non-temporary computer-readable medium according to any one of embodiments 71 to 98, wherein the at least one learning algorithm includes a supervised learning algorithm.
[0302] 100. The non-temporary computer-readable medium according to any one of embodiments 71 to 98, wherein the at least one learning algorithm includes an unsupervised learning algorithm.
[0303] 101. The non-temporal computer-readable medium according to any one of embodiments 71 to 100, wherein the at least one learning algorithm includes a regression algorithm, preferably a random forest regression.
[0304] 102. A non-temporal computer-readable medium according to any one of embodiments 71 to 101, further comprising a dimensionality reduction algorithm for data elements within each dataset, preferably UMAP, PCA, or both.
[0305] 103. The non-temporal computer-readable medium according to any one of embodiments 71 to 102, wherein the at least one learning algorithm includes a classification algorithm, preferably a random forest classifier, an SVM classifier, or both.
[0306] 104. The non-temporal computer-readable medium according to any one of embodiments 71 to 103, wherein the at least one learning algorithm includes a feature selection algorithm, preferably the Boruta algorithm.
[0307] 105. The non-temporary computer-readable medium according to any one of embodiments 71 to 104, wherein the non-temporary computer-readable medium is capable of independently resampling data elements within each dataset.
[0308] 106. The prediction model is the method according to any one of Embodiments 1 to 34, the system according to any one of Embodiments 35 to 70, and the non - transient computer - readable medium according to any one of Embodiments 71 to 104, which do not provide cancer diagnosis.
[0309] 107. The prediction model is the method according to any one of Embodiments 1 to 34, the system according to any one of Embodiments 35 to 70, and the non - transient computer - readable medium according to any one of Embodiments 71 to 104, which do not provide cancer prognosis.
[0310] 108. The prediction model is the method according to any one of Embodiments 1 to 34, the system according to any one of Embodiments 35 to 70, and the non - transient computer - readable medium according to any one of Embodiments 71 to 104, which do not provide cancer diagnosis or cancer prognosis.
[0311] Non - limiting examples The following are specific non - limiting examples of different aspects of the present disclosure. To examine how different aspects of the present disclosure affect the resulting prediction accuracy, a range of models were constructed for single - agent therapy prediction (Figure XA), and combination therapy (Figure XB) including both single - agent therapy prediction and combination therapy prediction.
[0312] Genomic Drug Sensitivity in Cancer (GDSC).
[0313] Drugcomb.org.
[0314] National Cancer Institute ALMANAC.
[0315] Cancer Genome Atlas Program.
[0316] Genotype Tissue Expression Project (GTEx).
[0317] A wide range of machine learning methods have been applied to the drug response prediction problem: regularized regression (e.g., Lasso, ElasticNet, Ridge regression), partial least squares (PLS) regression, support vector machines (VM), random forests (RF), neural networks and deep learning, logical models, or kernelized Bayesian matrix factorization (KBMF). However, a systematic exploration of model training strategies based on data from screening multiple large cell lines has not been reported. Furthermore, cell line-based models have not yet been compared to xenograft-based models. The ultimate goal of this study is to bridge these gaps by improving the accuracy of drug response prediction in cell lines and xenografts.
[0318] Initially, the effects of different modeling parameters (e.g., response metrics, number of features, feature types) on predictive performance in drug sensitivity testing were systematically studied. In this non-limiting example, four drugs (doxorubicin, gemcitabine, olaparib, and palbociclib) were used, and all drug sensitivities for the four drugs were available in each dataset. Predictive performance was evaluated by validation of xenograft mouse models. From this analysis, the optimal parameter set and strategy were selected to confirm whether drug sensitivity in animal xenograft systems could be predicted by training models from cell line-based and clinical-based data. In the xenograft experiments, each drug was tested on samples from specific tissue types, so the cancer detection datasets were accordingly limited; for example, in drug response prediction, all cell line-based and clinical-based data used for this modeling task were from the same cancer classification as the xenograft samples. The process of this non-limiting example is shown in Figure 1.
[0319] Predictive model Supervised learning classes The drug response prediction task and the slope prediction of the mitotic rate / proliferation curve were regression tasks (Figure 2). The histological type prediction task, used to evaluate which cancer AI model to use, was a classification task. Furthermore, some of the initial drug response prediction tasks for the TCGA clinical dataset were classification tasks.
[0320] Modeling methods and hyperparameters For regression tasks, the best-performing modeling method was random forest. For classification tasks, the best-performing modeling method was also random forest. Each modeling method has its own set of hyperparameters. For random forest regression, the most important hyperparameters include the maximum depth of individual trees (max_depth), the number of trees in the forest (n_estimators), and the criterion for splitting.
[0321] Function Selection Feature selection was performed using only data from the training set to select a subset of all available features for modeling. For the drug response prediction task, a multi-step feature selection process is used to fine-tune our model (Figure 2). First, we selected the top 1000 drug molecule features that predict drug response in each cancer type. In Random Forest, the "IncNodePurity" or variable importance value is sorted, and the top 1000 are selected first. Next, the Boruta algorithm is used to further refine the selection of drug molecule features into a smaller subset that produces optimal accuracy for the model. Approximately 20,000 gene expression features (HK ratio) are added to the refined drug molecule features (about 30-40), and trained against drug response to find the top 1000 gene features. Here again, Boruta is used to reduce the genes to just a handful of important genes (30-40). Gene biomarker optimization is performed to further reduce the total number of genes to 5-10 (Figure 3). Here, molecular features and optimized gene biomarkers are used to predict treatments for the cancer type subsets used in initial training. This task is performed for all cancer type subsets for which training data is available.
[0322] result Performance evaluation of predictive models Using bagging or bootstrap agglutination, multiple datasets were created, and the performance of each model was evaluated to validate the performance of the random forest regression model. The results of random forest regression prediction on an in-vitro dataset of monotherapy drugs, using the growth curve endpoint (AUC), indicate that this model performs well, on average R 2 The model performs well, achieving a score of 0.651, an RMSE of 0.094, and a precision of 90% (Figure 4A). Next, both monotherapy and combination therapy were evaluated together by converting the AUC score from monotherapy and the synergistic effect score from combination therapy into normalized Z-scores. Combined, the mean R of the model... 2The ratio was 0.569, the RMSE was 0.686, and the accuracy was 83% (Figure 4B).
[0323] In the TCGA clinical response evaluation, random forest classification was performed to model patient treatment responses. Due to the small size of the dataset (2212 in total), the entire clinical response dataset from TCGA was joined for this evaluation. Using RECIST criteria classes, four response variables were constructed to classify patients after various therapies into the following categories: complete response (CR), partial response (PR), stable disease (SD), and progression (PD). The area under the receiver operating characteristic (ROC) curve (AUC) indicated good model accuracy for predicting the correct response class, with a mean ROC AUC of 0.93 (Figure 5).
[0324] Verification of xenografts Gene expression was measured as consecutive TPM (transcripts per million mapped reads) values from RNA sequencing (RNA-seq). TPM values were normalized for a set of housekeeping genes (HK ratio). The HK ratios of each gene biomarker from the PDX were fed into an appropriate cancer type prediction model to determine a ranked list of treatments (Figure 2). The top-ranked treatments were tested in the PDX to validate the accuracy of the treatment predictions. In one such example, four AI-predicted therapies were tested against an orthotopic PDX model of triple-negative breast cancer. The results showed complete tumor growth suppression in the top-ranked monotherapies according to AI predictions (Figure 6).
Claims
1. A method for training at least one learning algorithm using at least one hardware processor, wherein the at least one learning algorithm is within a predictive model for cancer treatment, and the method is Training the at least one learning algorithm in a predictive model using at least one cancer drug discovery dataset, wherein the predictive model is capable of evaluating multiple drugs or drug combinations in terms of predicted biological responses to human cancer tissue exhibiting a particular set of cancer biomarkers, Validating a predictive model by using a xenograft mouse model, wherein in the xenograft mouse model, the human cancer tissue exhibiting a specific set of cancer biomarkers is transplanted in parallel into a plurality of immunodeficient mice, each of which is subsequently treated in parallel with one of the plurality of drugs or drug combinations, and the biological response to the treatment is fed back into the predictive model to further train the at least one learning algorithm. The method, including the method described above.
2. The method according to claim 1, wherein the human cancer tissue exhibiting a specific set of the cancer biomarkers is orthotopically transplanted into the plurality of immunodeficient mice.
3. The method according to claim 1, wherein the at least one cancer drug discovery dataset includes cancer classification data, cancer biomarker data, drug or drug combination data, and biological response data.
4. The method according to claim 3, wherein for each class of cancer, the at least one learning algorithm evaluates a drug or drug combination in relation to its biological response.
5. The method according to claim 4, wherein, for each drug or drug combination under the class of cancer, the at least one learning algorithm evaluates cancer biomarkers in relation to the correspondence of the drug or drug combination to the biological response.
6. The method according to claim 3, wherein the cancer classification data includes cancer cell line information from in vitro screening, patient information from clinical studies, or a combination thereof.
7. The method according to claim 6, wherein the cancer cell line information includes information selected from the group consisting of cancer cell line identification, TCGA classification, histological type, histological subtype, or a combination thereof.
8. The method according to claim 6, wherein the patient information includes information selected from the type of cancer, the patient's age, the patient's sex, the patient's weight, the patient's family history, the patient's medical history, and combinations thereof.
9. The method according to claim 3, wherein the cancer biomarker data includes gene expression data.
10. The method according to claim 9, wherein the gene expression data is normalized using TPM (transcripts per million).
11. The method according to claim 9, wherein the gene expression data is normalized to the geometric mean of at least four housekeeping genes with low variance.
12. The method according to claim 3, wherein the cancer biomarker data includes gene methylation data.
13. The method according to claim 3, wherein the cancer biomarker data includes protein biomarker data.
14. The method according to claim 3, wherein the drug or drug combination data includes the chemical structure of the drug.
15. The method according to claim 14, wherein the drug or drug combination data includes feature quantities of the substructure of the drug.
16. The method according to claim 15, wherein the characteristic features of the substructure of the drug include descriptors from the SMILES standard for the drug.
17. The method according to claim 3, wherein the biological response data includes in vitro drug screening results.
18. The method according to claim 17, wherein the in vitro drug screening result is in the form of IC50 (median inhibitory concentration), AUC (area under the drug response curve), or a combination thereof.
19. The method according to claim 3, wherein the biological response data includes in vitro drug combination screening results.
20. The method according to claim 19, wherein the in vitro drug combination screening results are in the form of a synergistic effect score obtained through a best single agent (HSA) model, a Loewe additive model, a zero interaction potency (ZIP) model, or a Bliss independence model.
21. The method according to claim 19, wherein the in vitro drug combination screening result is in the form of an AUC score.
22. The method according to claim 3, wherein the biological response data includes clinical research results.
23. The method according to claim 22, wherein the clinical research results include safety evaluation, efficacy evaluation, dose-response evaluation, pharmacodynamic evaluation, pharmacokinetic evaluation, progression-free survival evaluation, or a combination thereof.
24. The method according to claim 23, wherein the efficacy evaluation includes tumor size measurement, tumor volume assessment, efficacy biomarker evaluation, or a combination thereof.
25. The method according to claim 1, wherein the at least one cancer drug discovery dataset includes data from cancer genomic drug sensitivity (GDSC), NCI Almanac or Drugcom.org, data from the Cancer Genome Atlas Program, data from the Genotype Tissue Expression Project (GTEx), or a combination thereof.
26. The method according to claim 1, wherein the at least one learning algorithm includes a regression algorithm, and optionally the regression algorithm includes random forest regression.
27. The method according to claim 1, wherein the at least one learning algorithm includes a classification algorithm, optionally the classification algorithm includes a random forest classifier, an SVM classifier, or both.
28. The method according to claim 1, wherein the at least one learning algorithm includes a feature selection algorithm, and optionally the feature selection algorithm includes a Boruta algorithm.
29. The method according to claim 1, further comprising applying a dimensionality reduction algorithm, wherein the dimensionality reduction algorithm optionally includes UMAP, PCA, or both.
30. The method according to claim 1, wherein the predictive model does not provide cancer diagnosis, and optionally, the predictive model does not provide cancer prognosis.