Personalized ranking of cancer drugs
By processing mRNA expression levels through normalization and ranking algorithms, the method addresses the challenge of individualized cancer treatment, enhancing treatment efficacy and reducing toxicity by identifying the most effective drugs or combinations for each patient.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE CLEVELAND CLINIC FOUND
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Current cancer treatment approaches, particularly chemotherapy, lack individualized effectiveness due to the variability in drug response among patients, leading to unnecessary toxicity and limited clinical benefit, as high-throughput drug screening is time-consuming and costly, making it difficult to tailor treatments to individual patients.
A method involving the processing of raw mRNA expression levels using normalization, median finding, and ranking algorithms to rank cancer drugs or drug combinations based on gene signatures, enabling personalized treatment recommendations.
This approach allows for rapid identification of the most effective cancer drugs or combinations for individual patients, reducing unnecessary toxicity and improving treatment efficacy by ensuring drugs with proven efficacy are administered.
Smart Images

Figure IMGF000006_0001 
Figure IMGF000039_0001 
Figure IMGF000043_0001
Abstract
Description
[0001] L
[0002] PERSONALIZED RANKING OF CANCER DRUGS
[0003] The present application claims priority to U. S. Provisional application serial number 63 / 715,015 filed November 1, 2024, and is herein incorporated by reference in its entirety.
[0004] This invention was made with government support under 5R37CA244613 and 1F30CA257076 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0005] FIELD OF THE INVENTION
[0006] Provided herein are compositions, systems, and methods for ranking cancer drugs for treating a subject's cancer cells, where a plurality of gene signatures (each with a plurality of gene signature genes) with associated cancer drugs are processed with raw mRNA expression levels for genes in the sample. The processing (e.g.. by computer) can comprise: i) applying a normalization algorithm to generate normalized mRNA expression values for signature genes, ii) applying a median finding algorithm to the normalized mRNA expression values in each of the plurality of drug gene signatures to generate a plurality of median values, and iii) applying a ranking algorithm such that the median values are ranked from highest value to lowest value (or vice versa), with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drugs.
[0007] BACKGROUND
[0008] Tumors evolve in response to ever-changing selection pressures, making individualized treatment a significant challenge. Most standard chemotherapy regimens adopt a one-size-fits-all approach, where patients will receive drugs that are superior for a clinical trial cohort. Although such treatments may benefit a large portion of cancer patients, they are not guaranteed to be effective from individual to individual. Precision medicine seeks to address the variability in drug response among individual patients by tailoring treatment on an individual basis. Patients with targetable mutations can receive targeted therapy, but most cancer patients lack such mutations and cannot benefit from this form of precision medicine. In fact. 14% of all cancer patients are eligible for targeted therapy, and a mere 7% of all cancer patients respond to treatment. Thus, a majority of L
[0009] individuals with cancer will not receive precision medicine and instead are likely to be treated with a one-size-fits-all chemotherapy regimen.
[0010] Across cancer types, clinicians usually treat patients with combination therapy because these regimens generally yield more favorable responses compared to single agents. Several hypotheses for this benefit exist. Traditionally, drug combinations have been assembled based on the idea that the probability of resistance arising against multiple drugs with distinct mechanisms at once is less probable than resistance arising against a single agent. Interactions among drugs in a combination can also exist, and there are multiple models for these interactions. One is the Bliss independence model, where drugs in a combination act independently of each other, and the combined effect of the combination is equal to the sum of its parts. Similarly, the Loewe model of drug additivity assumes that the effect of a drug combination is equal to the sum of its components, but this model does not require the assumption that each drug acts independently of one another. From these models, synergistic and antagonistic interactions can also be defined, where the effect of one drug enhances or impedes the effect of another drug, respectively. Synergistic combinations produce a therapeutic effect greater than the sum of its parts, while antagonistic combinations produce an effect less than the sum of its parts. Retrospective studies on clinical trial data have revealed that the benefit of using combination chemotherapy is conferred primarily by the independent drug action model, where one treatment does not change the activity of another. Thus, treating a tumor with multiple drugs increases the probability that a patient will respond to at least one of the drugs being used. This means that if a patient is resistant to one or more dings in the combination, this approach risks increased toxicity with limited clinical benefit. With this understanding, patients should ideally receive drugs with known efficacy against their tumor, minimizing unnecessary toxicity from agents that are ineffective against their tumor. However, drug sensitivity is typically not measured before treatment for individual patients. As a result, it is currently unclear whether or not a given patient will truly benefit from each of the drags they receive. Ideally, tumor biopsies would be screened for drug response prior to treatment so that only drags with proven efficacy for the individual are administered. However, high-throughput drug screening requires substantial tissue, costly materials, and weeks to months of time depending on the facility. By the time the results are generated and returned to the clinician, the patient may have evolved a different sensitivity profile. For this reason, it is important that pre-treatment analyses of drag response are performed quickly to allow timely implementation of treatment regimens. L
[0011] SUMMARY
[0012] Provided herein are compositions, systems, and methods for ranking cancer drugs for treating a subject's cancer cells, where a plurality of gene signatures (each with a plurality of gene signature genes) with associated cancer drugs are processed with raw mRNA expression levels for genes in the sample. The processing (e.g., by computer) can comprise: i) applying a normalization algorithm to generate normalized mRNA expression values for signature genes, ii) applying a median finding algorithm to the normalized mRNA expression values in each of the plurality of drug gene signatures to generate a plurality of median values, and iii) applying a ranking algorithm such that the median values are ranked from highest value to lowest value (or vice versa), with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drugs.
[0013] In some embodiments, provided herein are methods of ranking cancer drugs effectiveness for treating a subject's cancer cells comprising: a) obtaining, from a biological sample from a subject. mRNA expression levels for: i) at least 5 signature genes (e.g.. 5, 6. 7, 8, 9. 10. 11. 12, 13, 14, 15, 16, 17, or more) in each of at least two drug gene signatures, wherein each of the drug gene signature comprises the at least 5 signature genes and is associated with one cancer drug, or combination of two cancer drugs, selected from a plurality of different cancer drugs, and ii) at least one non-signature gene (e.g., a house-keeping gene) not present in either of the at least two drug gene signatures; b) processing the mRNA expression levels with a processing system comprising: i) a computer processor, and ii) non-transitory computer memory comprising one or more computer programs and a database, wherein the one or more computer programs comprise: a signature gene mRNA expression normalization algorithm, a median finding algorithm, and a ranking algorithm, wherein the database comprises the at least two drug gene signatures with associated different cancer drugs, wherein the one or more computer programs, in conjunction with the computer processor and database, is / are configured to: A) apply the normalization algorithm to the mRNA expression level for each of the signature genes in view of the at least one non-signature gene mRNA expression level to generate a normalized mRNA expression level for each of the signature genes, B) apply the median finding algorithm to the normalized mRNA expression levels in each of the at least two drug gene signatures to generate a median value for each of the at least two drug gene signatures, and C) apply the ranking algorithm to the median values for each of the at least two drug gene signatures such that the median values are ranked from highest value to lowest value (or vice versa), with the highest value being associated with the most effective cancer drug, or most L
[0014] effective combination of two cancer drugs, of the plurality of different cancer drugs. In certain embodiments, the methods further comprise: treating the subject with the most effective cancer drug or most effective combination of two cancer drugs, and optionally not treating the subject with any other of the difference cancer drugs.
[0015] In further embodiments, the methods further comprise: generating a written and / or electronic report that provides a ranked list of the cancer drugs, or combinations of two cancer drugs. In other embodiments, the written and / or electronic report is provided to the subject and / or medical personnel treating the subject. In some embodiments, step a) is conducted at a first time point, and the method further comprises repeating step a) at a second time point (and optionally a third time point). In certain embodiments, the obtaining mRNA expression levels comprises sequencing the signature genes and the at least one non-signature gene. In other embodiments, the method is conducted without ever contacting the subject's cancer cells with any of the plurality of different cancer drugs in vitro. In further embodiments, the mRNA expression levels were determined by a method selected from: quantitative sequencing, Northern blotting, Nuclease Protection Assays, In Situ hybridization, and RT-PCR. In further embodiments, the obtaining mRNA expression levels comprises receiving a written or electronic report from a laboratory.
[0016] In particular embodiments, provided herein are systems for ranking cancer drugs effectiveness for treating a subject's cancer cells comprising: a) a computer processor, and b) non-transitory computer memory comprising: i) a data receiving component configured to receive, from a biological sample from a subject, mRNA expression level data for: A) at least 5 (e.g., at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more) signature genes in each of at least two drug gene signatures, and B) at least one non-signature gene not present in either of the at least two drug gene signatures; ii) a database comprises the at least two drug gene signatures each: i) comprising the at least 5 signature genes, and ii) having one associated cancer drug, or combination of two cancer drugs, selected from a plurality of difference cancer drugs, iii) one or more computer programs comprise: a signature gene mRNA expression normalization algorithm, a median finding algorithm, and a ranking algorithm, wherein the one or more computer programs, in conjunction with the computer processor, data receiving component, and database, is / are configured to: A) apply the normalization algorithm to the mRNA expression level for each of the signature genes in view of the at least one non-signature gene mRNA expression level to generate a normalized mRNA expression level for each of the signature genes, B) apply the median finding algorithm to the normalized mRNA expression levels in each of the at least two drug gene signatures to generate a L
[0017] median value for each of the at least two drug gene signatures, and C) apply the ranking algorithm to the median values for each of the at least two drug gene signatures such that the median values are ranked from highest value to lowest value (or vice versa), with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drugs, of the plurality of different cancer drugs. In further embodiments, the systems further comprise: a printing or electronic communication component configured to generate a written and / or electronic report that provides a ranked list of the cancer drugs, or combinations of two cancer drugs, from the most effective cancer drug, or combination, to least effective cancer drugs, or combination, against the tumor of the subject. In additional embodiments, the systems further comprise the biological sample from a subject.
[0018] In some embodiments, the biological sample comprises tumor cells or circulating mRNA, or any other sample comprising mRNA from the subject. In particular embodiments, the at least two drug gene signatures comprises at least three drug gene signatures. In additional embodiments, the at least two drug gene signatures comprises at least four drug gene signatures, at least five drug gene signatures, or at least ten gene signatures. In further embodiments, the at least one nonsignature gene comprises: i) at least one house keeping gene; or ii) at least two house-keeping genes; or iii) at least 200 non-signature genes, or iv) at least 10,000 non-signature genes. In certain embodiments, the normalization algorithm divides each of the mRNA expression levels of the signature genes by the mRNA expression level of: i) the at least one non-signature gene, or ii) the mean expression level of at least 20 (e.g., 20... 50... 100... 1,000,.... 10,000) non-signature genes.
[0019] In additional embodiments, the at least one non-signature gene is at least 20 or 200 nonsignature genes, and wherein the one or more computer programs further comprise: a mean expression level algorithm, a standard deviation algorithm, and wherein the normalization algorithm is a Z-score algorithm, and wherein the one or more computer programs, in conjunction with the computer processor and database, is / are configured to: A) apply the mean expression level algorithm to the mRNA expression levels for the at least 20 or 200 non-signature genes to generate a sample mean expression level value, B) apply the standard deviation algorithm to the mRNA expression levels for the at least 20 or 200 non- signature genes to generate a sample standard deviation value, C) apply the Z-score algorithm which, for each signature gene, subtracts the sample mean expression level value from a particular signature gene's expression level and divides by the sample standard deviation value, to generate a signature gene Z-score for each of the signature L
[0020] genes, and D) apply the median algorithm to the signature gene Z-scores in each of the plurality of drug gene signatures to generate the plurality of median values. In particular embodiments, the standard deviation algorithm applies the formula:
[0021]
[0022] wherein G is the population standard deviation, wherein N is the number of the at least 20 or 200 non-signature genes, wherein yi is each expression level for the at least 20 or 200 non-signature genes, and wherein p is the sample mean expression level value.
[0023] In certain embodiments, the median finding algorithm: i) when there are an odd number of normalized mRNA expression levels in a given genetic profile, identifies the middle normalized mRNA expression level when such normalized mRNA expression levels are numerically ordered highest to lowest, and ii) when there is an even umber of normalized mRNA expression levels in a given genetic profile, identifies the two middle normalized mRNA expression levels when such normalized expression levels are numerically ordered from highest to lowest and adds these two middle normalized mRNA expression levels and divides by two. In some embodiments, wherein the median Z-score algorithm: i) when there are an odd number of signature gene Z-scores in a given genetic profile, identifies the middle signature gene Z-score when such signature gene Z-scores are numerically ordered highest to lowest, and ii) when there is an even umber of signature gene Z-scores in a given genetic profile, identifies the two middle signature gene Z-scores when such signature gene Z-scores are numerically ordered from highest to lowest and adds these two middle signature gene scores and divides by two.
[0024] In particular embodiments, the at least two different cancer drugs is selected from:
[0025] Gemcitabine, Topotecan, Irinotecan, Cisplatin, Cyclophosphamide, Vinblastine, Cytarabine, Vorinostat, Luminespib. 5 -Fluorouracil, Vincristine, Docetaxel, Paclitaxel, I-BRD9, GSK591, EPZ004777, Nilotinib, Alpelisib. and YK-4-279. In other embodiments, the subject lacks known targetable mutations associated with the plurality of different cancer drugs. In additional embodiments, the subject is treated with the most effective cancer drag and no other cancer drug. In some embodiments, the subject is a human. In other embodiments, wherein: i) the at least one non-signature gene comprises at least at least 20 or 25 non-signature genes, or at least 100 nonsignature genes, or at least 200 non-signature genes, or at least 5,000 non-signature genes, or at least L
[0026] 10,000 non-signature genes, and / or ii) the at least 5 signature genes is at least 6, 7, 8, 9, 10, or 11 signature genes.
[0027] BRIEF DESCIPTION OF THE DRAWINGS
[0028] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.
[0029] Figure 1. A. The performance of each signature was measured by calculating a hazard ratio using the same method described by Scarborough et al. First, signature scores (median expression z-score of the signature’s genes) were calculated for all GDSC cell lines where response to the drug of interest existed. Cell lines were then grouped by score quintiles, and survival between the top and bottom quintiles (cell lines predicted to be the most and least sensitive) was compared. From this analysis, a hazard ratio was calculated and is visualized in the figures as a vertical blue line. The same process was repeated using randomly generated signatures of the same length, and the gray histogram represents the hazard ratios calculated from this null bootstrap. The dotted red lines denote the 95% confidence intervals for each null distribution, and the solid red line indicates the mean of the null signatures’ hazard ratios. B. Signature scores for all cell lines are calculated for all 10 gene signatures. Each column depicts the signature scores generated from one signature, and each row represents a single cell line. The higher the score, the greater the predicted sensitivity. C. The fraction of surviving cells is measured at standardized doses against each of the 10 drugs. Each column represents survival against one drug, and each row represents a single cell line in the same order as shown in panel B. A higher survival indicates greater resistance against a drug, and a lower survival indicates greater sensitivity against a drug.
[0030] Figure 2. Gene signatures predict the rank order of sensitivity for individual cell lines. Signature scores for each drug are calculated to predict sensitivity in each cell line. Relative sensitivity rankings are determined by ordering the signature scores from highest to lowest value, where the highest value indicates the greatest predicted sensitivity.
[0031] Figure 3. Quantifying relative predictive accuracy of response signatures. To determine the true sensitivity ranking, observed drug response is measured by calculating the surviving fraction of cells at a standardized dose for each cancer type. Standard doses are calculated by finding the median EC50 of all cell lines in a cancer subtype. L
[0032] Figure 4. Quantifying relative predictive accuracy of response signatures. Accuracy is scored by comparing the predicted rank order of sensitivity to the observed rank order of sensitivity.
[0033] Figure 5. Relative response predicted by the chemogram is more accurate than a null distribution. Cancer subtypes are abbreviated using TCGA study abbreviations. The boxed numbers above the x-axis panels c & d indicate the number of cell lines included in each cancer type. The lines in the center of the boxplots represent the mean accuracy, rather than the median. In panel a& c, predictive accuracy scores for each cell line are shown, and each point represents the prediction accuracy in a single cell line. Each point in panels b & d reflects the mean accuracy score for all cell lines within a given cancer type for one iteration of the null bootstrap. Since the null bootstrap was iterated 1,000 times, each blue violin plot shows the distribution of 1,000 average scores. The red points correspond to the mean accuracy scores of the chemogram using our evolutionary signatures. The points for the 3-signature chemogram (a & b) are more clearly discretized than the 10-signature chemogram (c & d) because there are far more possible scores with 10 signatures than 3 (6 possible scores vs 3,628,800).
[0034] Figure 6. Gene expression normalization. A. Among all cell lines in GDSC, GAPDH consistently exhibits high expression levels. As such, gene expression (for use within the chemogram) was normalized to GAPDH expression within each cell line. B. Signature scores derived for all drugs and cell lines are compared when calculated using GAPDH-normalized expression or expression z-scores. Signature scores derived from either normalization method are highly correlated for all signatures. Spearman’s rho is indicated at the top of each subpanel.
[0035] Figure 7. Relative response predicted by the chemogram across all cancer types. This is the same figure as Fig.5, but includes both epithelial and non-epithelial cancers. The results are very similar, generally good overall but a few cancer types average below 0.5. Results are also similar between the 3-drug and 10-drug results.
[0036] Figure 8. Predicted efficacy is consistent with observed survival. Distributions of observed survival were stratified by drug and predicted sensitivity ranking, where a ranking of 1 indicates the highest predicted sensitivity and 3 indicates the lowest predicted sensitivity. Each point corresponds with an individual cell line where the indicated drug was ranked either 1, 2, or 3 for sensitivity. The numbers below the boxplots and above the x-axis denote the number of cell lines included in each individual boxplot. Results are shown for both the 3 drug chemogram (a, c) and the 10 drug chemogram (b, d). Panels a and b show these results in cell lines of epithelial origin, and panels c and d show these results in all GDSC cell lines. For the 3 drug chemograms, it is seen that groups L
[0037] with higher sensitivity rankings have lower survival, and vice versa. A similar trend is observed for the 10-drug chemograms, although the trend is not as consistent as the 3-drug chemograms.
[0038] Figure 9. Frequency of rankings per drag. The frequency of each drag appearing in each sensitivity ranking is shown for the 3 drug chemogram (a, c) and 10 drag chemogram (b, d). Panels a and b show these results in cell lines of epithelial origin, and panels c and d show these results in all GDSC cell lines.
[0039] DETAILED DESCRIPTION
[0040] Provided herein are compositions, systems, and methods for ranking cancer drags for treating a subject's cancer cells, where a plurality of gene signatures (each with a plurality of gene signature genes) with associated cancer drags are processed with raw mRNA expression levels for genes in the sample. The processing (e.g., by computer) can comprise: i) applying a normalization algorithm to generate normalized mRNA expression values for signature genes, ii) applying a median finding algorithm to the normalized mRNA expression values in each of the plurality of drag gene signatures to generate a plurality of median values, and iii) applying a ranking algorithm such that the median values are ranked from highest value to lowest value (or vice versa), with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drags.
[0041] The general work flow for embodiments with Z-score normalization, including how the math works (e.g., for the 3 and 10 gene signatures in Example 1), is provided below for an example with three gene signatures each associated with one drug. Initially, for a given sample from a subject, raw bulk RNA-sequencing data is obtained (e.g., from cancer cells from Jane Doe). The genes in the gene signatures ("gem" for gemcitabine; "topo" for topotecan, and "iri" for irinotecan) for these three drags (used for this example) are shown in Table 1 below. Raw expression of the 10 genes in the gemcitabine signature, 11 genes in the topotecan signature, and 14 genes in the irinotecan signature are provided in Table 1. L
[0042] Table 1 - Three Gene Signatures with Associated Drug
[0043]
[0044] Across all genes in the individual sample (such as about 17,000 genes), expression values are normalized by calculating a z-score for each gene. This step happens once for each sample, regardless of how many signatures are being used. The formula employed is as follows:
[0045] Z-score per gene = (expression value of the particular gene signature gene - mean expression of all genes in the individual sample) / standard deviation of all genes in the individual sample
[0046] For Jane Doe's tumor cell sample, the average expression of all 17,737 genes is 4.84. The standard deviation is 2.08. Translating this from words into variables and replace the numbers we know, we have z=(x-4.84) / 2.08. Looking at the first gene in the first gene signature (ARNTL2) as an example, to calculate the z-score for the gene ARNTL2, we replace x with 9.3. which gives z=(9.3-4.84) / 2.08 -> z=2.14. For Jane Doe's sample, we can look at the z-scores for each gene in the same three signatures as above (shown in table 2 below). Again, in this Example, z-scores are calculated using data from all genes per one sample, not just for one signature at a time. L
[0047] Table 2 - Z scores for Each Gene in a Gene Signature with Associated Drug
[0048]
[0049] To calculate the signature score, find the median of the values listed above Table 2. In this case, the gemcitabine signature score for Jane Doe's sample is 0.99. The signature score for the topotecan signature is 1.20, and the signature score for irinotecan is 0.91. If there are more gene signatures used, the steps are repeated again for each additional predictive gene signature, starting from the same, normalized bulk-RNAseq data. All signature scores are then ranked from highest to lowest (or vice versa) to indicate predicted sensitivity. Higher expression of a signature is associated with higher sensitivity to the drug, so the higher the signature score is, the higher the predicted sensitivity. For Jane Doe's sample and the three drugs shown above, it is predicted that the sample will be most sensitive to topotecan (score of 1.20), second most sensitive to gemcitabine (score of 0.99), and the least sensitive to irinotecan (score of 0.91). These rankings are relative only to the drugs included here and do not necessarily mean, for instance, that irinotecan will not be effective. Instead, it is just that it is predicted that irinotecan will be relatively less effective than gemcitabine and irinotecan. The process outlined above can be repeated with the 10 drug signature shown in Table 3 below, instead of the 3 drug gene signature above. TABLE 3 - Ten Gene Signatures with associated cancer drug cisplatin cytarabine 5-fu gemcitabine irinotecan luminespib paclitaxel topotecan vinblastine vorinostat ADAT2 ACN9 ADAT2 ARNTL2 AIM2 ADAMTS6 ADAT2 AIM2 ACN9 ARID3B ATP1B3 ADAT2 ATP5D CRLF3 BCL2A1 ARHGAP22 AMTN FOXL2 ADAT2 C9orfl52 C15orf41 ASNS C1Q. BP CXCL1 BNC1 BNC1 ARTN IL1B ASNS CECR5 C1 BP C12orf57 CHCHD10 ELK3 FOXL2 CDH13 Clorf74 KRT5 C15orf41 CLYBL CDC7 CCNB1IP1 CLYBL GLIPR1 HMGA1 CSGALNACT2 C1QBP LY6K C1QBP DCXR CDCA7 CLYBL COQ.3 MLKL ITPRIP CTHRC1 CDCA7 MMP10 C6orfl70 DQX1 FKBP14 DLEU1 CSTA POLR3G LY6K DSE COQ.3 RAB38 CCNB1IP1 EFNA3 KRT5 DPH5 DSG3 PROCR MMP10 DYNC2H1 CSTA SERPINB4 COQ3 FHIT LRRC8C F12 F12 RELB PGM2 DZIP1 CYB5R4 SLFN11 CSTA FKBP4 LY6K FAR1 FAM83F SLFN11 SERPINB4 FAM101B DSG3 SOX7 ECSIT FRAT2 MMP10 FASTKD1 FASTKD1 SLC6A15 IL27RA FAT2 WDFY2 FGF11 IL17RB NPM3 GNPNAT1 FRAT2 SLFN11 ITPRIP FRAT2 FSD1 JHDM1D PSAT1 MMP10 GMDS SOX7 JARID2 GBP6 LRRC49 MYB RIOK1 MTHFD2 MRPL2 TRAF3 KRT14 GNA15 MMP10 NOC2L SLFN11 MYC MUC13 LRRC8C GPC2 NOC3L NRARP STOML2 NOB1 MYB MFAP2 HOXDIO NPM3 PDCD2L USP31 NPM3 NPM3 PDLIM4 JARID2 RIOK1 PEX7 WDR3 POLR1D NRARP POLR3G KRT5 RPF2 PHGR1 ZNF75O PSAT1 PDSS1 POPDC3 KRT6A SEH1L RIOK1 SFXN4 PIP5K1B RFTN1 KRT6B STOML2 SDHAF1 SIGMAR1 POLR1D SH3PXD2B MARK1 TAF4B TIMM8B SLC27A5 PPP1R1B SLC31A2 MMP10 TAF5 TMEM168 SNRPA1 REG4 SLC4A7 MMP13 TBPL1 TMEM183A TAF4B RPL22L1 SPHK1 MYC TMEM206 TOX3 TUBE1 SFXN4 TM4SF19 NPM3 UQCRH TRAP1 SLC27A5 TWIST1 NRARP
[0050] SPINK4 VEGFC PKP1
[0051] UQCRH WDFY2 PREP
[0052] VSNL1 PVRL1
[0053] ZNF511 REL
[0054] RIOK1
[0055] S100A7
[0056] SH3BP1
[0057] SLC27A5
[0058] SOX7
[0059] TAF4B
[0060] TAF5
[0061] TMEM2O6
[0062] TP63
[0063] TRERF1
[0064] USP31
[0065] WDFY2
[0066] ZNF75O L
[0067] The general work flow for embodiments 1 house-keeping gene (or minimal genes), including how the math works (e.g., for the 3 and 10 gene signatures), is provided below for an example with three gene signatures each associated with one drug. Initially, for a given sample from a subject, raw bulk RN A- sequencing data is obtained (e.g., from cancer cells from Jane Doe). The genes in the gene signatures ("gem" for gemcitabine; "topo" for topotecan, and "iri" for irinotecan) for these three drugs (used for this example) are shown in Table 4 below. Raw mRNA expression of the 6 genes in the gemcitabine signature, 11 genes in topotecan signature, and 14 genes in irinotecan signature are as follows for this cell line as shown in Table 4.
[0068] TABLE 4
[0069]
[0070] These raw expression values are normalized by dividing the expression value of a given gene to the GAPDH (house-keeping gene) expression value for the same sample. The normalized value for these three gene signatures are shown in Table 5 below. L
[0071] TABLE 5
[0072]
[0073] To calculate the signature score, find the median of the values listed in Table 5 above. In this case, the gemcitabine signature score for this sample is 0.54. The score for topotecan is 0.58 and the score for irinotecan is 0.53. If there are more signatures used, these steps are repeated again for each additional predictive signature, always starting from the same, normalized bulk- RNAseq data. All signature scores are then ranked from highest to lowest (or vice versa) to indicate predicted sensitivity. Higher expression of a signature is associated with higher sensitivity to the drag, so the higher the signature score is, the higher the predicted sensitivity. In this example, it is predicted that the sample will be most sensitive to topotecan, second most sensitive to gemcitabine, and the least sensitive to irinotecan. These rankings are relative only to the drugs included here and do not necessarily mean, for instance, that irinotecan will not be effective, just that it is predicted it will be relatively less effective than gemcitabine and irinotecan.
[0074] Additional gene signatures (with associated drug) that one could employ are shown in Table 6 below. L
[0075] Table 6 - Additional Drug Gene Signatures
[0076] Gemcitabine CRLF3, MLKL, POLR3G, PROCR, RELB, SLFN11
[0077] Topotecan AIM2 FOXL2 IL1B KRT5 LY6K MMP10 RAB38 SERPINB4 SLFN11 SOX7 WDFY2
[0078] Irinotecan AIM2 BCL2A1 BNC1 FOXL2 HMGA1 ITPRIP LY6K MMP10 PGM2 SERPINB4 SLC6A15 SLFN11 SOX7 TRAF3
[0079] Cisplatin LRRC8C, LY6K, MMP10, SLFN11, STOML2, USP31, WDR3, ZNF75O
[0080] Cyclophosphamide AKR1A1 C6orfl36 C9orfl23 CECR5 COQ3 DYM FRAT2 GCAT HACE1 HMG20B MARK1 MRPL2 MTHFD2 NT5DC2 PCCB RIOK1 SDF2L1 SLC27A5 TAF4B TDP1 UQCRH
[0081] Vinblastine ACN9 ADAT2 ASNS C15orf41 C1QBP C6orfl70 CCNB1IP1 COQ3 CSTA ECSIT FGF11 FSD1 LRRC49 MMP10 N0C3L NPM3 RI0K1 RPF2 SEH1L STOML2 TAF4B TAF5 TBPL1 TMEM206 UQCRH
[0082] Cytarabine ACN9 ADAT2 ASNS C12orf57 CCNB1IP1 CLYBL DLEU1 DPH5 F12 FAR1 FASTKD1 GNPNAT1 MMP10 MTHFD2 MYC NOB1 NPM3 P0LR1D PSAT1 SFXN4 SIGMAR1 SLC27A5 SNRPA1 TAF4B TUBE1
[0083] Vorinostat ARID3B C9orfl52 CECR5 CLYBL DCXR DQX1 EFNA3 FHIT FKBP4 FRAT2 IL17RB JHDM1D MYB NOC2L NRARP PDCD2L PEX7 PHGR1 RIOK1 SDHAF1 TIMM8B TMEM168 TMEM183A TOX3 TRAP1 TTC39A
[0084] Lurninespib ADAMTS6 ARHGAP22 BNC1 CDH13 CSGALNACT2 CTHRC1 DSE DYNC2H1 DZIP1 FAM101B IL27RA ITPRIP JARID2 KRT14 LRRC8C MFAP2 PDLIM4 POLR3G POPDC3 RFTN1 SH3PXD2B SLC31A2 SLC4A7 SPHK1 TM4SF19 TWIST1 VEGFC WDFY2
[0085] 5-Fluorouracil ADAT2 ATP5D C1QBP CHCHD10 CLYBL COQ3 CSTA DSG3 F12 FAM83F FASTKD1 FRAT2 GMDS MRPL2 MUC13 MYB NPM3 NRARP PDSS1 PIP5K1B POLR1D PPP1R1B REG4 RPL22L1 SFXN4 SLC27A5 SPINK4 UQCRH VSNL1 ZNF511
[0086] Vincristine ADAT2 ADSL C18orf21 C1QBP C6orfl70 GART GEMIN6 GLMN GNL3 HACE1 HNRNPC LTV1 METAP1 METTL8 MIB1 MRPS9 PIAS2 PMS1 POLR3G PSMG1 RNF138 RPP4O RPRD1A STOML2 TAF4B TAF5 TBPL1 TDP1 TMEM206 UQCRH USP14 WDR3 L
[0087] Docetaxel ACN9 ARTN ASNS SLC27A5
[0088] Paclitaxel ADAT2 AMTN ARTN Clorf74 C1QBP CDCA7 COQ3 CSTA CYB5R4 DSG3 FAT2 FRAT2 GBP6 GNA15 GPC2 HOXDIO JARID2 KRT5 KRT6A KRT6B MARK1 MMP10 MMP13 MYC NPM3 NRARP PKP1 PREP PVRL1 REL RI0K1 S100A7 SH3BP1 SLC27A5 S0X7 TAF4B TAF5 TMEM206 TP63 TRERF1 USP31 WDFY2 ZNF750
[0089] I-BRD9 ADAT2 NT5DC2 PDSS1 SIGMAR1 SLC27A5 SLC6A15
[0090] GSK591 ADAT2 CSTA I LIB KRT5 MARK1 NUDT8 SKP2 S0X7 TP63 WDFY2
[0091] EPZ004777 ADAT2 C0Q3 DLX6 FAT2 FRAT2 MARK1 NT5DC2 RI0K1 S1PR5 SLC27A5 TAF5
[0092] Nilotinib ANKRD30A C17orf58 CREB3L4 DHTKD1 EFHD1 FRAT2 GCAT HK1 IDH2 KIAA1324 LRIG1 MGP MUCL1 MYB PCCB PPM1J PREXI SLC24A3 SLC27A5 SPDEF ZNF77
[0093] Alpelisib ABCA12 ABCC5 AIM1 ANK3 ANKRD22 ARTN Clorfl72 Clorf74 CDH1 CDH3 C0R02A CREB3L4 CSTA DDR1 DSG3 ELM03 EPPK1 ERBB2 ESRP1 ESRP2 FUR FAT2 FGFR3 F0XA1 FRAT2 FXYD3 GALNT3 GBP6 GGCT GPX2 GRB7 GRHL1 H00K2 IRF6 IRX2 JAG2 JARID2 JUP KBTBD7 KRT14 KRT15 KRT5 KRT6A LNX2 MAP7 MARK1 MARVELD3 MB MREG NRARP PIP PKP3 PLEKHF2 PROM2 PRSS8 PTGFRN PVRL1 RAB25 S100A14 SCGB1A1 SLC27A5 SPDEF SPINT1 ST6GALNAC2 SYTL1 TC2N TP63 TPRG1 TSPAN13 ZNF750
[0094] YK-4-279 ADAT2 CCDC138 CCND2 CDCA7 CENPH COQ3 CSNK1E FASTKD1 LRP4 NRARP NT5DC2 PSAT1 SLC27A5 TAF5 SOX7 SPINK5 SPINT1 SPINT2 ST6GALNAC2 SYK SYNE2 SYTL1 TMEM40 TNFAIP8 T FSF1O TP63 TRIM29 WDR66 WNT10A WNT4 WNT7A XDH ZNF165 ZNF204P ZNF750
[0095] In addition to the gene signature, with associated drugs or drug combinations, herein, one can extract other predictive signatures that may be used with the systems and methods herein. For example, one can extract predictive gene signatures using a previously established method in Scarborough et al, NPJ Precis. Oncol. 7, 38 (2023). herein incorporated by reference in its entirety, particularly for methods for determining gene signatures. For example, one can use the Genomics of Drug Sensitivity in Cancer (GDSC) database and The Cancer Genome Atlas (TCGA) to extract one predictive gene signature per drug of interest. For example, to begin extracting a signature for cisplatin, one first partitions all GDSC cell lines of epithelial origin into 5 folds, where each fold contains a random 80% of all the cell lines. For one fold at a time, one identifies which cell lines L
[0096] comprise the top 20% that are most resistant against cisplatin (or other drug under consideration) and which cell lines comprise the top 20% that are most sensitive to cisplatin. Once these cell lines are identified, one can use three differential expression methods (limma, multtest, sam) to identify the genes that are differentially upregulated in sensitive cell lines. For the next step, one can retain only the genes identified by all three differential expression methods and refer to them as seed genes. Using the seed genes, one then generates a coexpression network by calculating the pairwise Spearman correlation between the expression of each seed gene and every other gene from TCGA normal and tumor tissue expression data. The seed genes found in the top 20% of highly coexpressed genes are maintained. This process is repeated for all five folds, and the genes extracted in at least three of the five folds form the final gene signature.
[0097] The present disclosure is not limited to particular methods of detecting the level of the recited genes. Markers may be detected as DNA (e.g., cDNA) or RNA (e.g., mRNA). Exemplary methods for detecting the presence or absence of mRNA levels include, but are not limited to, polymerase chain reaction (PCR)-based technologies including, for example, reverse transcription PCR (RT-PCR) and quantitative or real-time RT-PCR (RT-qPCR). Other methods include microarray analysis, RNA sequencing (e.g., next- generation sequencing (NGS)), in situ hybridization, and Northern blot. For example, the presence or amount of biomarker nucleic acid (e.g., mRNA) in a sample is determined (e.g., to determine the presence or level of biomarker expression). Biomarker nucleic acid (e.g.. RNA, amplified cDNA, etc.) may be detected / quantified using a variety of nucleic acid techniques known to those of ordinary skill in the art, including but not limited to nucleic acid sequencing, nucleic acid hybridization, nucleic acid amplification (e.g., by PCR, RT-PCR. qPCR. etc.), micorarray, Southern and Northern blotting, sequencing, etc. Nonamplified or amplified nucleic acids can be detected by any conventional means. For example, in some embodiments, nucleic acids are detected by hybridization with a detectably labeled probe and measurement of the resulting hybrids. Nucleic acid detection reagents may be labeled (e.g., fluorescently) or unlabeled, and may by free in solution or immobilized (e.g., on a bead, well, surface, chip, etc.).
[0098] In some embodiments, nucleic acid sequencing methods are utilized for detection mRNA expression levels, hi some embodiments, the technology provided herein finds use in a Second Generation (a.k.a. Next Generation or Next-Gen), Third Generation (a.k.a. Next-Next-Gen), or Fourth Generation (a.k.a. N3-Gen) sequencing technology including, but not limited to, pyrosequencing, sequencing-by-ligation, single molecule sequencing, sequence-by- synthesis (SBS), L
[0099] semiconductor sequencing, massive parallel clonal, massive parallel single molecule SBS, massive parallel single molecule real-time, massive parallel single molecule real-time nanopore technology, etc. Morozova and Marra provide a review of some such technologies in Genomics, 92: 255 (2008), herein incorporated by reference in its entirety. Those of ordinary skill in the art will recognize that because RNA is less stable in the cell and more prone to nuclease attack experimentally RNA is usually reverse transcribed to cDNA before sequencing.
[0100] In certain embodiments, next generation sequencing is employed to detect and quantify mRNA levels. Such NGS methods share the common feature of massively parallel, high-throughput strategies, with the goal of lower costs in comparison to older sequencing methods (see, e.g., Voelkerding et al., Clinical Chem., 55: 641-658, 2009; MacLean et al., Nature Rev.
[0101] Microbiol., 7: 287-296; each herein incorporated by reference in their entirety). NGS methods can be broadly divided into those that typically use template amplification and those that do not.
[0102] Amplification-requiring methods include pyrosequencing commercialized by Roche as the 454 technology platforms (e.g., GS 20 and GS FLX), Life Technologies / Ion Torrent, the Solexa platform commercialized by Illumina, GnuBio, and the Supported Oligonucleotide Ligation and Detection (SOLiD) platform commercialized by Applied Biosystems. Non-amplification approaches, also known as single-molecule sequencing, are exemplified by the HeliScope platform commercialized by Helicos BioSciences, and emerging platforms commercialized by VisiGen, Oxford Nanopore Technologies Ltd., and Pacific Biosciences, respectively.
[0103] In some embodiments, hybridization methods are utilized for detecting and quantitating mRNA levels. Illustrative non-limiting examples of nucleic acid hybridization techniques include, but are not limited to, in situ hybridization (ISH), microarray, and Southern or Northern blot.
[0104] In some embodiments, a computer-based analysis program is used to translate the raw data generated by the detection assay e.g., mRNA levels) into data of predictive value for a clinician. The clinician can access the predictive data using any suitable means. Thus, in some preferred embodiments, the present invention provides the further benefit that the clinician, who is not likely to be trained in genetics or molecular biology, need not understand the raw data. The data is presented directly to the clinician in its most useful form. The clinician is then able to immediately utilize the information in order to optimize the care of the subject (e.g., if the subject should be treated with a particular drug or not).
[0105] The present disclosure contemplates any method capable of receiving, processing, and transmitting the information to and from laboratories conducting the assays, information provides, L
[0106] medical personal, and subjects. For example, in some embodiments of the present invention, a sample (e.g., a biopsy) is obtained from a subject and submitted to a profiling service e.g., clinical lab at a medical facility, genomic profiling business, etc.), located in any part of the world (e.g., in a country different than the country where the subject resides or where the information is ultimately used) to generate raw expression level data. Where the sample comprises a tissue or other biological sample, the subject may visit a medical center to have the sample obtained and sent to the profiling center, or subjects may collect the sample themselves and directly send it to a profiling center. Where the sample comprises previously determined biological information, the information may be directly sent to the profiling service by the subject (e.g., an information card containing the information may be scanned by a computer and the data transmitted to a computer of the profiling center using an electronic communication system). Once received by the profiling service, the sample is processed and a profile is produced which ranks drugs for the patient.
[0107] In some embodiments, the information is first analyzed at the point of care or at a regional facility. The raw data is then sent to a central processing facility for further analysis and / or to convert the raw data to information useful for a clinician or patient (e.g., ranked drugs for use to treat cancer). The central processing facility provides the advantage of privacy (all data is stored in a central facility with uniform security protocols), speed, and uniformity of data analysis. The central processing facility can then control the fate of the data following treatment of the subject. For example, using an electronic communication system, the central facility can provide data to the clinician, the subject, or researchers. In some embodiments, the subject or medical care provider is able to directly access the data using the electronic communication system. The subject may chose further intervention or counseling based on the results.
[0108] EXAMPLES EXAMPLE 1
[0109] Personalizing chemotherapy drug selection using a transcriptomic chemogram Gene expression signatures predictive of chemotherapeutic response have the potential to greatly extend the reach of precision medicine by allowing medical providers to plan treatment regimens on an individual basis. Most published gene signatures are only capable of predicting response for individual drugs, but currently, a majority of chemotherapy regimens utilize combinations of different agents. We propose a unified framework, called the chemogram, that uses predictive gene signatures to rank the predicted sensitivity of different drugs in any given L
[0110] individual. Using this approach, providers could efficiently screen against many therapeutics to identify the drugs that would fit best into a patient’s treatment plan at any given time. This can be easily reassessed at any point in time if treatment efficacy begins to decline due to therapeutic resistance.
[0111] To demonstrate the utility of the chemogram, we first extract predictive gene signatures using a previously established method for extracting pan-cancer signatures, inspired by convergent evolution. Across cancer cell lines of epithelial origin from the Genomics of Drug Sensitivity in Cancer (GDSC) database and The Cancer Genome Atlas, we derived 3 signatures for 3 commonly used cytotoxic drugs (cisplatin, gemcitabine, and 5 -fluorouracil), which are shown in Table 1 below.
[0112] We then used these signatures in the chemogram to predict and rank sensitivity among the drugs. To assess the accuracy of our method, we compared the rank order of predicted response to the rank order of observed response (fraction of surviving cells) against each of the 3 chemotherapies. Across all cancer types, our signatures were able to correctly predict the rank order of drag response with 71 % accuracy on average and were consistently more accurate than randomized prediction rankings, as well as prediction rankings made by randomly generated gene signatures. In addition to the chemograms ability to accurately rank sensitivity, this framework is easily scalable for any number of drags. We repeated the process described above for 10 drugs (see Table 2) and found that the accuracy of the predicted sensitivity rankings was maintained as the number of drags in the chemograms screen increased. Our framework demonstrates the ability of transcriptomic signatures to not only predict chemotherapeutic response but correctly assign rankings of drug sensitivity on an individual basis.
[0113] Introduction
[0114] Tumors evolve in response to ever-changing selection pressures, making individualized treatment a significant challenge. Most standard chemotherapy regimens adopt a one-size-fits-all approach, where patients will receive drags that are superior for a clinical trial cohort. Although such treatments may benefit a large portion of cancer patients, they are not guaranteed to be effective from individual to individual. Precision medicine seeks to address the variability in drug response among individual patients by tailoring treatment on an individual basis. Patients with targetable mutations can receive targeted therapy, but most cancer patients lack such mutations and L
[0115] cannot benefit from this form of precision medicine.1In fact, 14% of all cancer patients are eligible for targeted therapy, and a mere 7% of all cancer patients respond to treatment.2Thus, a majority of individuals with cancer will not receive precision medicine and instead are likely to be treated with a one-size-fits-all chemotherapy regimen.
[0116] Across cancer types, clinicians usually treat patients with combination therapy because these regimens generally yield more favorable responses compared to single agents.3Several hypotheses for this benefit exist. Traditionally, drug combinations have been assembled based on the idea that the probability of resistance arising against multiple drugs with distinct mechanisms at once is less probable than resistance arising against a single agent.4Interactions among drags in a combination can also exist, and there are multiple models for these interactions. One is the Bliss independence model, where drugs in a combination act independently of each other, and the combined effect of the combination is equal to the sum of its parts.5Similarly, the Loewe model of drag additivity assumes that the effect of a drag combination is equal to the sum of its components, but this model does not require the assumption that each drag acts independently of one another.6From these models, synergistic and antagonistic interactions can also be defined, where the effect of one drag enhances or impedes the effect of another drag, respectively. Synergistic combinations produce a therapeutic effect greater than the sum of its parts, while antagonistic combinations produce an effect less than the sum of its parts.
[0117] Retrospective studies on clinical trial data have revealed that the benefit of using combination chemotherapy is conferred primarily by the independent drag action model, where one treatment does not change the activity of another.7, 8Thus, treating a tumor with multiple drugs increases the probability that a patient will respond to at least one of the drugs being used.9This means that if a patient is resistant to one or more drugs in the combination, this approach risks increased toxicity with limited clinical benefit. With this understanding, patients should ideally receive drugs with known efficacy against their tumor, minimizing unnecessary toxicity from agents that are ineffective against their tumor. However, drag sensitivity is typically not measured before treatment for individual patients. As a result, it is currently unclear whether or not a given patient will truly benefit from each of the drugs they receive. Ideally, tumor biopsies would be screened for drag response prior to treatment so that only drags with proven efficacy for the individual are administered. However, high-throughput drug screening requires substantial tissue, costly materials, and weeks to months of time depending on the facility. By the time the results are generated and L
[0118] returned to the clinician, the patient may have evolved a different sensitivity profile. For this reason, it is critical that pre-treatment analyses of drug response are performed quickly to allow timely implementation of treatment regimens.
[0119] An established clinical workflow is already in place to personalize drug regimens for patients who need antibiotics. These patients will often have an antibiogram performed on a sample of their infection, where a drag screen is performed on the sample to select antibiotics that effectively target bacteria without having an unnecessarily broad-spectrum treatment regimen.10This pipeline is effective in a clinical setting because bacteria grow far more rapidly than human tissue and antibiotics are less expensive than chemotherapeutics. As aforementioned, drug screens for human tissue samples are time-consuming and costly, making them impractical for clinical purposes.11To perform a drug screen, the tissue needs to be expanded in vitro so there are enough cells to test on. This further complicates matters because growing human tissue in a dish requires the cells to adapt to an incredibly different environment. If the tissue survives this selection process, the resulting cells are an imperfect model of the parent tumor. Due to these significant hurdles, there are many ongoing efforts to predict drag response without requiring a drug screen.
[0120] Most published drug response prediction models utilize gene expression, whether it be exclusively or in conjunction with other data such as the chemical structures of drags, genomics, or proteomics. As the availability of multi-omics data increase exponentially, more models are including such information to predict drug response.12However, clinical implementation of approaches that utilize multiple types of data may be slowed due to the cost and speed of obtaining all the information needed to employ the model. Model complexity can vary from relatively simple, such as gene signatures, to complex, such as neural networks and other machine learning methods. Gene signatures are a straightforward tool that can be easily translated into a clinical setting, and several have already been incorporated into decision-making algorithms (e.g. Mammaprint13’14, OncotypeDx14, 15, Prolaris16, 17). Previously, we established a method of extracting gene signatures predictive of drag response based on exploiting convergent evolution.18We derived a pan-cancer sensitivity signature for cisplatin across epithelial- origin cancer cell lines in The Genomics of Drug Sensitivity in Cancer (GDSC). We first identified similarities in differential gene expression across all epithelial-origin cancer cell lines that exhibited sensitivity to cisplatin, then compared their transcriptomes against cell lines that showed resistance to cisplatin. From this subset, the genes found to be the most highly co-expressed in tumor samples from The Cancer Genome Atlas L
[0121] (TCGA) were isolated to form the final cisplatin signature. We demonstrated that this signature can predict cisplatin response within cell lines from GDSC, and expression levels of the signature align with clinical trends observed in tumor samples from TCGA and the Total Cancer Care databases. As a preliminary validation, we conducted a case study in a novel muscle-invasive bladder cancer dataset to assess the signature’s ability to estimate risk level and found the signature to be predictive among patients who have received cisplatin-containing treatment.
[0122] Here, we utilize our evolution-inspired signature model to generate gene expression signatures predictive of several common chemotherapeutics and present a scalable framework for predicting and ranking drug sensitivity among those therapies. Our proposed workflow, which we refer to as a transcriptomic chemogram, would serve a similar purpose as an antibiogram. Just as an antibiogram provides clinical decision support for antibiotic selection, a chemogram would inform oncologists which chemotherapies would be optimal for an individual cancer patient by using predictive gene signatures to assess drug sensitivity across multiple drugs. With this method, the sensitivity profile of a tumor can be extensively characterized in an efficient manner that forgoes the need for a drug screen and can be performed at scale. Furthermore, a tumor’s sensitivity profile could easily be reassessed with biopsies of recurrent tumors or metastatic lesions. A chemogram could also be used to determine appropriate treatments for patients who have cancer types that lack well-defined options for second-line chemotherapy. This is often the case for individuals with rare cancers that lack extensive research.19, 20Using a chemogram to personalize chemotherapy and adjust treatment as the tumor evolves, patients with any type of cancer would be less likely to receive drugs that lack efficacy against their tumor at any time throughout treatment. We introduce this and demonstrate the application of a transcriptomic chemogram in vitro using several predictive pan-cancer signatures. L
[0123] RESULTS
[0124] Gene signatures can predict and rank drug sensitivity of individual samples
[0125] Using the evolution-inspired extraction pipeline defined by Scarborough et al. (ref. 18, herein incorporated by reference), we extracted 70 predictive gene signatures using The Cancer Genome Atlas (TCGA) and the Genomics of Drug Sensitivity in Cancer (GDSC) public datasets.18For each drug we derive a signature for, this extraction method identifies differentially upregulated genes between the most sensitive and most resistant GDSC cell lines across cancer types of epithelial origin. Three differential expression analysis methods are used, and only the genes found by all three methods are maintained for the next step. The resulting subset of genes is then filtered to maintain the most highly coexpressed genes among patient samples in TCGA. This method is repeated 5 times with a randomized 80% of cell lines, and the genes found in at least 3 of the 5 runs are included in the final signature. For simplicity, we first demonstrated the chemogram with cisplatin, 5 -fluorouracil, and gemcitabine, which are commonly used in various clinical settings. Later, we exhibit the scalability of the chemogram using the top 10 signatures by individual accuracy (Fig.l A, see Methods). We use the method described by Scarborough et al to assess drug sensitivity by generating a signature score (the median normalized expression among the genes in a given signature). For each tumor sample, one signature score can be generated for each predictive signature. The signature scores for each drug are then compared to each other to determine the rank order of predicted sensitivity, as depicted by Fig.2.
[0126] Using the gene expression data associated with each cell line provided by GDSC, we applied the chemogram to 616 untreated cell lines, 394 of which are of epithelial origin. First, raw gene expression values were normalized to GAPDH expression within each cell line (each gene’s expression value was divided by GAPDH expression for that cell line; Fig. 6). Using these normalized values, we then calculated the signature scores for each cell line-gene signature pair. The predicted sensitivity for the three drugs in each cell line was then ranked by descending signature score. The signature that yielded the highest score identified the drug predicted to cause the most cell death relative to the other two drugs. Likewise, the signature yielding the lowest score identified the drug predicted to cause the least cell death compared to the other drugs. Because each cell line’s gene expression can vary, even within the same cancer type, each cell line can have a L
[0127] different prediction ranking among the three drugs. The summary scores for these cell lines are depicted in Fig. IB.
[0128] Prediction rankings are assessed for accuracy by comparing to observed survival against each drug
[0129] To assess the accuracy of the chemogram’s predictions, we compared the predicted rank order of sensitivity (derived in Fig. 2) to the rank order of observed survival against each of the three drugs (calculated in Fig. 3). To measure observed survival in each cell line, we calculated survival at a standard dose for each cancer type. We defined the standard dose per drug as the median EC50 of all cell lines within a cancer type. For instance, to determine a standard dose for gemcitabine in all lung cancer cell lines, we find the median EC50 against gemcitabine across all lung cancer cell lines. This is repeated for cisplatin and 5 -fluorouracil. At these median EC50s. we then calculated the fraction of surviving cells for each individual cell line using the raw doseresponse data from GDSC fitted with the gdscIC50 package.21
[0130] Because cell lines of the same disease type can differ in their drug response, the measured survival can differ on an individual basis, similar to how different patients with the same type of cancer can have variable drug responses. This process was applied to all GDSC cell lines that had drug response data for all three drugs of interest, and the rank order of observed sensitivity was determined from this. The fraction of surviving cells against the standardized doses for the three drugs across all of these cell lines is shown in Fig. 1C.
[0131] Using multiple signatures allows for the prediction of sensitivity rank order
[0132] To quantify the accuracy of the chemogram’s predictions, the rank order of predicted sensitivity was compared with the rank order of observed survival per cell line. The accuracy of the predictions, relative to the observed survival, was scored from 0 to 1. With this scoring metric, 1 represented a perfectly accurate prediction of the survival rank order and 0 represented a prediction that was the opposite of the survival rank order. The remaining accuracy scores will be determined based on how close the drugs with greater predicted efficacy are to truly being more effective. The closer the drugs with greater predicted sensitivity are to truly showing more efficacy, the closer to 1 the accuracy score will be. Each increment of the accuracy score increases linearly, and the number L
[0133] of possible scores increases with the number of drugs used. An example of the scoring metric and its usage is shown in Fig. 4.
[0134] We first demonstrated the chemogram using only three predictive signatures for cisplatin, gemcitabine, and 5-fluorouracil, as shown in Table 1 below.
[0135] TABLE 1
[0136]
[0137] Using these signatures, we performed the prediction, validation, and accuracy scoring as outlined in Fig. 2, 3, and 4. Among the majority of cancer cell lines in the GDSC dataset, the chemogram predicted sensitivity rank order with over 50% accuracy. Comparing these results to a null bootstrap, we found that the accuracy of the chemogram-generated predictions were consistently higher than the null distribution. To perform the null bootstrap, three null gene signatures (each the same length as one of the three evolutionary gene signatures) were randomly generated and used to predict and rank chemosensitivity in the same manner as before, as depicted by Fig.2. The accuracy of the random signatures’ prediction would then be scored using the same method as described in Fig. 3 and 4. This process was repeated for every cell line 1000 times, and the accuracy scores for each cell line in every iteration are represented by the blue boxplots in Fig.
[0138] 5a. The mean accuracy scores from each iteration of the null bootstrap, as well as the average scores from the chemogram, are shown in Fig. 5b. Across all cancer types, the random signatures were only accurate about 50% of the time on average, and the chemogram consistently outperformed the null distribution. Signature-based chemograms are scalable for any number of drugs
[0139] To demonstrate the scalability of the chemogram, we repeated our validation with 10 signatures, all extracted using the same method as before.18The signatures used were associated with sensitivity to paclitaxel, vinblastine, vorinostat, 5 -fluorouracil, topotecan, gemcitabine, luminespib, cytarabine, cisplatin, and irinotecan, as shown in Table 2 below.
[0140] Table 3 shows an example of 10 gene signature with associated single drugs.
[0141] cisplatin cytarabine5.fugemcitabine irinotecan luminespib paclitaxel topotecan vinblastine vorinostat ADAT2 ACN9 ADAT2 ARNTL2 AIM2 ADAMTS6 ADAT2 AIM2 ACN9 ARID3B ATP1B3 ADAT2 ATP5D CRLF3 BCL2A1 ARHGAP22 AMTN F0XL2 ADAT2 C9orfl52 C15orf41 ASNS C1Q. BP CXCL1 BNC1 BNC1 ARTN IL1B ASNS CECR5 C1QBP C12orf57 CHCHD1O ELK3 F0XL2 CDH13 Clorf74 KRT5 C15orf41 CLYBL CDC7 CCNB1IP1 CLYBL GLIPR1 HMGA1 CSGALNACT2 C1QBP LY6K C1Q. BP DCXR CDCA7 CLYBL COQ3 MLKL ITPRIP CTHRC1 CDCA7 MMP10 C6orfl70 DQX1 FKBP14 DLEU1 CSTA POLR3G LY6K DSE COQ3 RAB38 CCNB1IP1 EFNA3 KRT5 DPH5 DSG3 PROCR MMP10 DYNC2H1 CSTA SERPINB4 COQ.3 FHIT LRRC8C F12 F12 RELB PGM2 DZIP1 CYB5R4 SLFN11 CSTA FKBP4 LY6K FAR1 FAM83F SLFN11 SERPINB4 FAM101B DSG3 SOX7 ECSIT FRAT2 MMP10 FASTKD1 FASTKD1 SLC6A15 IL27RA FAT2 WDFY2 FGF11 IL17RB NPM3 GNPNAT1 FRAT2 SLFN11 ITPRIP FRAT2 FSD1 JHDM1D PSAT1 MMP10 GMDS SOX7 JARID2 GBP6 LRRC49 MYB RIOK1 MTHFD2 MRPL2 TRAF3 KRT14 GNA15 MMP10 NOC2L SLFN11 MYC MUC13 LRRC8C GPC2 NOC3L NRARP STOML2 NOB1 MYB MFAP2 HOXDIO NPM3 PDCD2L USP31 NPM3 NPM3 PDLIM4 JARID2 RIOK1 PEX7 WDR3 POLR1D NRARP POLR3G KRT5 RPF2 PHGR1 ZNF750 PSAT1 PDSS1 POPDC3 KRT6A SEH1L RIOK1 SFXN4 PIP5K1B RFTN1 KRT6B STOML2 SDHAF1 SIGMAR1 POLR1D SH3PXD2B MARK1 TAF4B TIMM8B SLC27A5 PPP1R1B SLC31A2 MMP10 TAF5 TMEM168 SNRPA1 REG4 SLC4A7 MMP13 TBPL1 TMEM183A TAF4B RPL22L1 SPHK1 MYC TMEM206 TOX3 TUBE1 SFXN4 TM4SF19 NPM3 UQCRH TRAP1 SLC27A5 TWIST1 NRARP TTC39A SPINK4 VEGFC PKP1
[0142] UQCRH WDFY2 PREP
[0143] VSNL1 PVRL1
[0144] ZNF511 REL
[0145] RI0K1
[0146] S100A7 L
[0147] SH3BP1
[0148] SLC27A5
[0149] SOX7
[0150] TAF4B
[0151] TAF5
[0152] TMEM2O6
[0153] TP63
[0154] TRERF1
[0155] USP31
[0156] WDFY2
[0157] ZNF750
[0158] The performance of each signature is represented in Fig. 1A. The process of prediction, validation, and scoring for accuracy is the same as depicted in Fig. 2, 3. and 4, but with a greater number of drugs and associated signatures. Fewer cell lines were included in this analysis because several cell lines did not have associated drug response data for all 10 drugs. This portion of the analysis included 539 cell lines, whereas the 3-signature chemogram included 616. Of the 539 cell lines used in the 10-drug chemogram, 351 of these cell lines were of epithelial origin. From the 616 cell lines used in the 3-drug chemogram, 394 were of epithelial origin.
[0159] The accuracy of the chemogram was maintained, and even slightly improved after it was scaled to use 10 drugs instead of 3. The similarity of these results can be observed by comparing between Fig. 5a & b (3 drugs) and Fig. 5c & d (10 drugs). In most cancer types, the 10-signature chemogram consistently produced accuracy scores above 0.5 and the scores were distributed higher than that of the random predictions (Fig. 5c & d). The null distributions were generated in the same way as described previously, but the bootstrap was performed with 10 random signatures. The null distributions for the 3-drug and 10-drug chemograms are both uniformly distributed. The points in Fig. 5c appear more scattered than those in Fig. 5a because there is a vast difference in the number of possible accuracy scores between these chemograms. With 3 signatures, there are only 6 possible accuracy scores. With 10 signatures, there are 3,628,800 possible accuracy scores.
[0160] Standard chemotherapy generally takes a one-size-fits-all approach, where patients without targetable mutations will receive some combination of drugs that work for most patients, or improve outcomes to some degree compared to previous regimens, based on clinical trials. However, every cancer patient has a unique chemosensitivity profile, and what may work for most people is not guaranteed to work for an individual. Further, increased benefits from combination chemotherapy L
[0161] may largely be attributed to at least one, but not necessarily all, of the drugs being effective.3, 7 9In this scenario, patients are receiving more than what is necessary for effective treatment and experiencing excessive toxicity as a result. Precision medicine, particularly targeted therapy, improves on the issues imposed by the one-size-fits-all approach to chemotherapy and interpatient tumor heterogeneity by targeting driver mutations present in a tumor, and such therapies are only available to patients who harbor such mutations. However, few patients are eligible for targeted therapy and efficacy can be limited even for those who harbor actionable mutations.2Even patients who respond to treatment initially are not immune to disease recurrence due to the evolution of resistance. Personalizing chemotherapy using predictive gene signatures would make precision medicine much more accessible to patients, regardless of mutation status, while also aiding the treatment of patients with chemoresistant tumors.
[0162] The chemogram can be easily adapted for any number of drugs, as long as a predictive signature exists for the treatment in question, and does not require a plethora of time-consuming assays to be performed. This framework is similar to that of an antibiogram, which is used to determine what antibiotics a patient is sensitive to so that proper treatment can be administered.10To use a chemogram, gene expression of a tumor biopsy would be measured, predictive signatures would be applied to the gene expression profile, sensitivity scores would be calculated, and the best drugs to give a patient can be determined promptly. Since the chemogram is based on gene expression, it is applicable to every cancer patient, regardless of subtype, stage, or mutation status. The results that are returned to clinicians would be a ranked list of drugs, avoiding any confusion and leaving little room for subjective interpretation.
[0163] The chemogram would be an ideal tool for guiding evolution-informed therapy. Evolution-informed therapy is an approach to treatment where a drug regimen is continually adjusted according to the phenotypic changes in the disease. This method has been explored in theoretical models and bacteria, but has yet to be tested in the context of cancer.22 25Considering that tumors are constantly evolving and adapting to selective pressures, this dynamic approach to treatment would be highly relevant for use with chemotherapeutic s. The simplicity and flexibility of the chemogram makes it a promising device to guide an evolution-informed chemotherapy regimen. Depending on the physical accessibility of the tumor, regular biopsies could be taken and assessed using the chemogram to adjust treatment as a tumor continues to evolve so that the patient is continuously treated as effectively as possible. L
[0164] The exemplary chemogram in this example, was limited to monotherapy because the predictive signatures were only extracted for one drug at a time. However, most patients needing chemotherapy receive combinations of drugs. To better synergize the chemogram with standard practices, signatures that are predictive of drug combinations could be included in the screen. For instance, a chemogram could include a signature for drug A, drug B, and drugs A and B together. This would help clinicians determine if all drugs in a standard combination are appropriate for use on a case-by-case basis, thereby preventing excessive toxicity. Alternatively, potential drug interactions among different combinations could be assessed after the ranked list is generated. This might be accounted for by calculating an interaction score of some kind among the drugs predicted to be most effective. For example, if drugs A, B, and C are the top 3 drugs, one could determine whether the combination is appropriate by calculating a score for the combinations AB, AC, BC, and ABC, then treating the patient with the combination that yields the highest synergy score. Several ongoing efforts are being made to create a method for predicting and quantifying drug interactions, but many of them are based on different definitions of synergy and do not agree with each other as a result.26 28However, the potential problem of adverse drug interactions in patients may not be a serious concern based on recent findings. A vast majority of approved drug combinations only seem to exhibit additivity - true synergy appears to be rare, as is antagonism.3, 8However, rarity does not eliminate all potential, and drug interactions should be kept in mind.
[0165] The ranked list of drugs generated by the chemogram is relative to each individual patient. Therefore, if a patient is resistant to all of the drugs screened in the chemogram, their caregiver would still receive a ranked list even if the top-ranked drugs are unlikely to be effective. Defining a threshold for individual sensitivity scores would allow clinicians to make clear distinctions about what drugs would be effective for treatment. If many drugs exceed the threshold, the chemogram would aid in determining the best options among those choices.
[0166] METHODS
[0167] Data Collection & Pre-Processing
[0168] All data cleaning, analysis, and plotting were performed using R (Version 4.3.0) with RStudio.34 L
[0169] GDSC Gene Expression Data
[0170] Microarray mRNA expression data for 1018 cell lines was downloaded from GDSC, which can be accessed from www.cancerrxgene.org / .35The expression data was collected using the Human Genome U219 96-Array Plate with the Gene Titan MC instrument (Affymetrix). This data was normalized using the robust multi-array analysis (RMA) algorithm.36’37For signature extraction, expression data was then converted to z-scores per gene within samples. The chemogram was used with both z-score normalized expression as well as GAPDH-normalized expression, and the results for each were nearly identical (Fig. 6B). GAPDH was selected as the housekeeping gene to normalize the data to because of its consistently high expression level among all cell lines (Fig.
[0171] 6A). Cancer cell lines of epithelial origin were defined based on these GDSC tissue descriptors: head_and_neck, oesophagus, breast, biliary _tract. large_intestine, liver, adrenal_gland, stomach, kidney, lung_NSCLC_adenocarcinoma, lung_NSCLC_squamous-_cell_carcinoma, mesothelioma, pancreas, skin_other, thyroid, Bladder, cervix, endometrium, ovary, prostate, testis, urogenital_system_other, uterus.
[0172] GDSC Drug Response Data
[0173] Both raw drag response data and reported IC50 values for 809 cell lines were downloaded from GDSC. Reported IC50s were used for signature extraction and raw drag response data was used for chemogram validation presented by Fig. 3. The raw dose-response data was cleaned and normalized using the gdscIC50 package in R.21Normalization involved converting raw fluorescence values to percent viability, relative to the negative and positive controls. Doseresponse curves for each cell line-drug pair were fitted using a non-linear mixed effects model with the fitModelNlmeData function defined in the gdscIC50 package. The fitted equations generated from this were used to calculate EC50s and observed survivals per cell line.
[0174] Not all cell lines had associated drug responses for all of the drags used in our demonstration, which reduced the number of cell lines we could include in our analysis. Excluding cell lines with an unclassified cancer subtype, we were left with 763 cell lines for the 3-signature chemogram and 546 cell lines for the 10-signature chemogram. From these respective cohorts, there were 398 epithelial-origin cell lines for the 3-signature chemogram and 356 epithelial-origin cell lines for the 10-signature chemogram. L
[0175] TCGA Gene Expression Data
[0176] The Cancer Genome Atlas (TCGA) gene expression data was used for signature extraction. The RNA-Seq by Expectation Maximization (RSEM) normalized gene expression was downloaded using the RTCGAToolbox package (version 2.30.0).38Expression values were measured via Illumina HiSeq RNAseq V2 and were log2 transformed.
[0177] Sensitivity Signature Extraction
[0178] Predictive signatures were extracted using the pipeline described by Scarborough et al.18This work made use of the High Performance Computing Resource in the Core Facility for Advanced Research Computing at Case Western Reserve University. The GDSC data was partitioned into 5 folds, where each fold contained a randomized 80% of all cell lines in the dataset. For each signature, GDSC cell lines (within each fold) were ordered from most resistant to most sensitive. Cell lines in the top 20-30% (20% for all signatures except 5 -fluorouracil, which used 30%) of each extreme were identified, and differential expression was measured between the groups using limma, sam, and multtest.39-41The genes found to be differentially expressed by all 3 methods were used as seed genes for a coexpression network that was built using TCGA expression data. Seed genes found to be highly coexpressed in at least 3 of the 5 folds were included in the final signature.
[0179] Using this method, we extracted signatures for 70 chemotherapy agents and assessed the accuracy of each using the hazardratio metric described by Scarborough et al. The performance of each signature was assessed by calculating the signature score (median z-score expression of the signature genes) for all GDSC cell lines, then comparing the distribution of IC50 between cell lines in the highest and lowest quintiles of the signature scores. From this, we then calculated a hazard ratio, and this metric was used to rank all 70 signatures from best to worst accuracy. For the 10 signatures used in this study, the hazard ratio of the extracted signatures compared to a null distribution is shown in Fig. 1A. The null distribution was generated by using 1000 randomly generated gene signatures to stratify cell lines and calculate hazard ratios (per drug, all 1000 random signatures were the same length as the original signature). Based on the hazard ratio, we selected L
[0180] the top 10 signatures to demonstrate the chemogram. All but the irinotecan signature outperformed the cisplatin signature, which is the same as CisSig derived by Scarborough et al. All the signatures were extracted using the same parameters as the cisplatin signature, with the exception of the irinotecan, topotecan, and vorinostat signatures, which used a differential expression cutoff of 0.15 instead of 0.2.
[0181] REFERENCES
[0182] 1. Fountzilas et al., Clinical trial design in the era of precision medicine. Genome medicine 14, 101 (2022).
[0183] 2. Haslam et al., Updated estimates of eligibility for and response to genome-targeted oncology drugs among us cancer patients, 2006-2020. Annals Oncol. 32, 926-932 (2021).
[0184] 3. Pomeroy, et al., Drug independence and the curability of cancer by combination chemotherapy. Trends Cancer (2022).
[0185] 4. Pritchard et al., Understanding resistance to combination chemotherapy. Drug Resist. Updat. 15, 249-257 (2012).
[0186] 5. Bliss, C. I. The toxicity of poisons applied jointly 1. Annals applied biology 26, 585-615 (1939).
[0187] 6. Loewe, S. The problem of synergism and antagonism of combined drugs. Arzneimittel-forschung 3, 285-290 (1953).
[0188] 7. Palmer et al., Predictable clinical benefits without evidence of synergy in trials of combination therapies with immune-checkpoint inhibitors. Clin. Cancer Res. 28, 368-377 (2022).
[0189] 8. Chen et al. Independent drug action and its statistical implications for development of combination therapies. Contemp. Clin. Trials 98, 106126 (2020).
[0190] 9. Palmer, A. C. & Sorger, P. K. Combination cancer therapy can confer benefit via patient-to-patient variability without drug additivity or synergy. Cell 171, 1678-1691 (2017).
[0191] 10. Joshi, S. Hospital antibiogram: a necessity. Indian journal medical microbiology 28, 277-280 (2010). L
[0192] 11. Vargas, R. et al. Case study: patient-derived clear cell adenocarcinoma xenograft model longitudinally predicts treatment response. NPJ Precis. Oncol. 2, 14 (2018).
[0193] 12. Lin-Rahardja, K., Weaver, D. T., Scarborough, J. A. & Scott, J. G. Evolution-informed strategies for combating drug resistance in cancer. Int. J. Mol. Sci. 24, 6738 (2023).
[0194] 13. Audeh et al., Prospective validation of a genomic assay in breast cancer: The 70-gene mammaprint assay and the mindact trial. Acta medica academica 48 (2019).
[0195] 14. Xin et al., The era of multigene panels comes? the clinical utility of oncotype dx and mammaprint. World journal oncology 8, 34 (2017).
[0196] 15. Siow, Z. R., De Boer, R. H., Lindeman, G. J. & Mann, G. B. Spotlight on the utility of the oncotype dxR breast cancer assay. Int. journal women’s health 10, 89 (2018).
[0197] 16. Brawer, M. K. et al. Prolaris: A novel genetic test for prostate cancer prognosis. (2013).
[0198] 17. Ebell, M. H. Prolaris test for prostate cancer risk assessment. Am. Fam. Physician 100, 311-312 (2019).
[0199] 18. Scarborough, J. A., Eschrich, S. A., Torres-Roca, J., Dhawan, A. & Scott, J. G. Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature. NPJ Precis. Oncol. 7, 38 (2023).
[0200] 19. Huang, M. & Lucas, K. Current therapeutic approaches in metastatic and recurrent ewing sarcoma. Sarcoma 2011 (2010).
[0201] 20. Chen, C., Dorado Garcia, H., Scheer, M. & Henssen, A. G. Current and future treatment strategies for rhabdomyosarcoma. Front, oncology 9, 1458 (2019).
[0202] 21. Lightfoot, H., van der Meer, D. & Vis, D. J. gdscIC50: Pipeline for GDSC Curve Fitting (2022). R package version 0.99.4.
[0203] 22. Yoon, N., Vander Velde, R., Marusyk, A. & Scott, J. G. Optimal therapy scheduling based on a pair of collaterally sensitive drugs. Bull, mathematical biology 80, 1776-1809 (2018).
[0204] 23. Yoon, N., Krishnan, N. & Scott, J. Theoretical modeling of collaterally sensitive drug cycles: shaping heterogeneity to allow adaptive therapy. J. Math. Biol. 83, 1-29 (2021). L
[0205] 24. Maltas, J. & Wood, K. B. Pervasive and diverse collateral sensitivity profiles inform optimal strategies to limit antibiotic resistance. PLoS biology 17, e3000515 (2019).
[0206] 25. Weaver et al., Reinforcement learning informs optimal treatment strategies to limit antibiotic resistance. bioRxiv 2023-01 (2023).
[0207] 26. Yang et al. Stratification and prediction of drug synergy based on target functional similarity, npj Syst. Biol. Appl. 6, 16 (2020).
[0208] 27. Sheng et al., Advances in computational approaches in identifying synergistic drug combinations. Briefings Bioinforma. 19, 1172-1182 (2018).
[0209] 28. Pemovska et al., Recent advances in combinatorial drug screening and synergy scoring. Curr. opinion pharmacology 42, 102-110 (2018).
[0210] 29. Guerrero, et al. Analysis of racial / ethnic representation in select basic and applied cancer research studies. Sci. reports 8, 13978 (2018).
[0211] 30. Ma et al., Population-based differences in treatment outcome following anticancer drag therapies. The lancet oncology 11, 75-84 (2010).
[0212] 31. Saijo, The role of pharmacoethnicity in the development of cytotoxic and molecular targeted drags in oncology. Yonsei medical journal 54, 1-14 (2013).
[0213] 32. Patel, Cancer pharmacogenomics: implications on ethnic diversity and drag response.
[0214] Pharmacogenetics Genomics 25, 223-230 (2015).
[0215] 33. O’Donnell & Dolan, Cancer pharmacoethnicity: ethnic differences in susceptibility to the effects of chemotherapy. Clin. Cancer Res. 15, 4806-4814 (2009).
[0216] 34. R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria (2023).
[0217] 35. Yang et al. Genomics of drug sensitivity in cancer (gdsc): a resource for therapeutic biomarker discovery in cancer cells. Nucleic acids research 41, D955-D961 (2012).
[0218] 36. Irizarry et al. Exploration, normalization, and summaries of high density oligonucleotide array probe level data. Biostatistics 4, 249-264 (2003).
[0219] 37. Iorio et al. A landscape of pharmacogenomic interactions in cancer. Cell 166, 740-754 (2016). 38. Samur, Rtcgatoolbox: a new tool for exporting tcga firehose data. PloS one 9, e106397 (2014).
[0220] 39. Ritchie et al. limma powers differential expression analyses for ma-sequencing and microarray studies. Nucleic acids research 43, e47-e47 (2015).
[0221] 40. Tusher et al., Significance analysis of microarrays applied to. (2001).
[0222] 41. Pollard et al., Multiple testing procedures: the multtest package and applications to genomics. In Bioinformatics and computational biology solutions using R and bioconductor, 249-271 (Springer, 2005).
[0223] All publications and patents mentioned in the specification and / or listed below are herein incorporated by reference. Various modifications and variations of the described method and system of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the relevant fields are intended to be within the scope described herein.
Claims
CLAIMSWe Claim:
1. A method of ranking cancer drugs effectiveness for treating a subject's cancer cells comprising:a) obtaining, from a biological sample from a subject, mRNA expression levels for:i) at least 5 signature genes in each of at least two drug gene signatures, wherein each of said drug gene signatures comprises said at least 5 signature genes and is associated with one cancer drug, or combination of two cancer drugs, selected from a plurality of different cancer drugs, andii) at least one non-signature gene not present in either of said at least two drug gene signatures;b) processing said mRNA expression levels with a processing system comprising:i) a computer processor, andii) non-transitory computer memory comprising one or more computer programs and a database, wherein said one or more computer programs comprise: a signature gene mRNA expression normalization algorithm, a median finding algorithm, and a ranking algorithm, wherein said database comprises said at least two ding gene signatures with associated different cancer drugs,wherein said one or more computer programs, in conjunction with said computer processor and database, is / are configured to:A) apply said normalization algorithm to said mRNA expression level for each of said signature genes in view of said at least one non-signature gene mRNA expression level to generate a normalized mRNA expression level for each of said signature genes,B) apply said median finding algorithm to said normalized mRNA expression levels in each of said at least two drag gene signatures to generate a median value for each of said at least two drug gene signatures, andC) apply said ranking algorithm to said median values for each of said at least two drug gene signatures such that said median values are ranked from highest value to lowest value, with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drugs, of said plurality of different cancer drugs.
2. The method of claim 1, further comprising: treating said subject with said most effective cancer drug or most effective combination of two cancer drugs, and optionally not treating said subject with any other of said difference cancer drugs.
3. The method of claim 1, further comprising: generating a written and / or electronic report that provides a ranked list of said cancer drugs, or combinations of two cancer drugs.
4. The method of claim 3, wherein said written and / or electronic report is provided to said subject and / or medical personnel treating said subject.
5. The method of claim 1, wherein said biological sample comprises tumor cells or circulating mRNA.
6. The method of claim 1, wherein step a) is conducted at a first time point, and wherein the method further comprises repeating step a) at a second time point.
7. The method of claim 1, wherein said obtaining mRNA expression levels comprises sequencing said signature genes and said at least one non-signature gene.
8. The method of claim 1, wherein said median finding algorithm: i) when there are an odd number of normalized mRNA expression levels in a given genetic profile, identifies the middle normalized mRNA expression level when such normalized mRNA expression levels are numerically ordered highest to lowest, and ii) when there is an even umber of normalized mRNA expression levels in a given genetic profile, identifies the two middle normalized mRNA expression levels when such normalized expression levels are numerically ordered from highest to lowest and adds these two middle normalized mRNA expression levels and divides by two.
9. The method of claim 1, wherein said at least two drug gene signatures comprises at least three drug gene signatures, at least five drug gene signatures, or at least ten gene signatures.
10. The method of claim 1, wherein method is conducted without ever contacting said subject's cancer cells with any of said plurality of different cancer drugs in vitro.
11. The method of claim 1, wherein said at least one non-signature gene comprises: i) at least one house keeping gene; or ii) at least two house keeping genes; or iii) at least 200 non-signature genes, or iv) at least 10,000 non-signature genes.
12. The method of claim 1, wherein said normalization algorithm divides each of said mRNA expression levels of said signature genes by the mRNA expression level of: i) said at least one nonsignature gene, or ii) the mean expression level of at least 200 non-signature genes.
13. The method of claim 1, wherein said at least one non-signature gene is at least 200 non-signature genes, and wherein said one or more computer programs further comprise: a mean expression level algorithm, a standard deviation algorithm, and wherein said normalization algorithm is a Z-score algorithm, andwherein said one or more computer programs, in conjunction with said computer processor and database, is / are configured to:A) apply said mean expression level algorithm to said mRNA expression levels for said at least 200 non-signature genes to generate a sample mean expression level value,B) apply said standard deviation algorithm to said mRNA expression levels for said at least 200 non- signature genes to generate a sample standard deviation value,C) apply said Z-score algorithm which, for each signature gene, subtracts said sample mean expression level value from a particular signature gene's expression level and divides by said sample standard deviation value, to generate a signature gene Z-score for each of said signature genes, and D) apply said median algorithm to said signature gene Z-scores in each of said plurality of drug gene signatures to generate said plurality of median values.
14. The method of claim 13, wherein said standard deviation algorithm applies the formula:wherein G is the population standard deviation,wherein N is the number of said at least 200 non-signature genes,wherein %j is each expression level for said at least 200 non-signature genes, and wherein p is said sample mean expression level value.
15. The method of claim 13, wherein said median Z-score algorithm: i) when there are an odd number of signature gene Z- scores in a given genetic profile, identifies the middle signature gene Z-score when such signature gene Z-scores are numerically ordered highest to lowest, and ii) when there is an even umber of signature gene Z-scores in a given genetic profile, identifies the two middle signature gene Z-scores when such signature gene Z-scores are numerically ordered from highest to lowest and adds these two middle signature gene scores and divides by two.
16. The method of claim 1, Gemcitabine, Topotecan, Irinotecan, Cisplatin, Cyclophosphamide. Vinblastine, Cytarabine, Vorinostat, Luminespib, 5 -Fluorouracil, Vincristine, Docetaxel, Paclitaxel, I-BRD9, GSK591, EPZ004777, Nilotinib, Alpelisib, and YK-4-279.
17. The method of claim 1, wherein said subject lacks known targetable mutations associated with said plurality of different cancer drugs.
18. The method of claim 1, wherein said subject is treated with said most effective cancer drug and no other cancer drug.
19. The method of claim 1, wherein said subject is a human.
20. The method of claim 1, wherein said mRNA expression levels were determined by a method selected from: quantitative sequencing, Northern blotting, Nuclease Protection Assays, In Situ hybridization, and RT-PCR.
21. The method of claim 1, wherein said obtaining mRNA expression levels comprises receiving a written or electronic report from a laboratory.
22. The method of claim 1, wherein:i) said at least one non-signature gene comprises at least at least 25 non-signature genes, or at least 100 non-signature genes, or at least 200 non-signature genes, or at least 5,000 non-signature genes, or at least 10.000 non-signature genes, and / orii) said at least 5 signature genes is at least 6, 7, 8, 9, 10, or 11 signature genes.
23. A system for ranking cancer drugs effectiveness for treating a subject's cancer cells comprising:a) a computer processor, andb) non-transitory computer memory comprising:i) a data receiving component configured to receive, from a biological sample from a subject, mRNA expression level data for:A) at least 5 signature genes in each of at least two drug gene signatures, and B) at least one non-signature gene not present in either of said at least two drug gene signatures;ii) a database comprises said at least two drug gene signatures each: i) comprising said at least 5 signature genes, and ii) having one associated cancer drug, or combination of two cancer drugs, selected from a plurality of difference cancer drugs,iii) one or more computer programs comprise: a signature gene mRNA expression normalization algorithm, a median finding algorithm, and a ranking algorithm,wherein said one or more computer programs, in conjunction with said computer processor, data receiving component, and database, is / are configured to:A) apply said normalization algorithm to said mRNA expression level for each of said signature genes in view of said at least one non-signature gene mRNA expression level to generate a normalized mRNA expression level for each of said signature genes,B) apply said median finding algorithm to said normalized mRNA expression levels in each of said at least two drug gene signatures to generate a median value for each of said at least two drug gene signatures, andC) apply said ranking algorithm to said median values for each of said at least two drug gene signatures such that said median values are ranked from highest value to lowest value, with the highest value being associated with the most effective cancer drug, or most effective combination of two cancer drugs, of said plurality of different cancer drugs.
24. The system of claim 23, further comprising: a printing or electronic communication component configured to generate a written and / or electronic report that provides a ranked list of said cancer drugs, or combinations of two cancer drugs, from said most effective cancer drug, or combination, to least effective cancer drugs, or combination, against said tumor of said subject.
25. The system of claim 23, wherein said biological sample comprises tumor cells.
26. The system of claim 23. wherein said biological sample comprises circulating mRNA.
27. The system of claim 23, wherein said mRNA expression levels are determined by sequencing.
28. The system of claim 23, wherein said at least two drug gene signatures comprises at least three drug gene signatures.
29. The system of claim 23, wherein said plurality of drug gene signatures comprises at least five drug gene signatures, or at least ten gene signatures.
30. The system of claim 23, wherein:i) said at least one non-signature gene comprises at least at least 25 non-signature genes, or at least 100 non-signature genes, or at least 200 non-signature genes, or at least 5,000 non-signature genes, or at least 10,000 non-signature genes, and / orii) said at least 5 signature genes is at least 6. 7, 8, 9, 10, or 11 signature genes.
31. The system of claim 23, wherein said at least one non-signature gene is at least 200 nonsignature genes, and wherein said one or more computer programs further comprise: a mean expression level algorithm, a standard deviation algorithm, and wherein said normalization algorithm is a Z-score algorithm, andwherein said one or more computer programs, in conjunction with said computer processor and database, is / are configured to:A) apply said mean expression level algorithm to said mRNA expression levels for said at least 200 non-signature genes to generate a sample mean expression level value,B) apply said standard deviation algorithm to said mRNA expression levels for said at least 200 non- signature genes to generate a sample standard deviation value,C) apply said Z-score algorithm which, for each signature gene, subtracts said sample mean expression level value from a particular signature gene's expression level and divides by said sample standard deviation value, to generate a signature gene Z-score for each of said signature genes, and D) apply said median algorithm to said signature gene Z-scores in each of said plurality of drug gene signatures to generate said plurality of median values.
32. The system of claim 31, wherein said standard deviation algorithm applies the formula:wherein G is the population standard deviation,wherein N is the number of said at least 200 genes,wherein % is each expression level for said at least 200 genes, andwherein p is said sample mean expression level value.
33. The system of claim 31, wherein said median Z-score algorithm: i) when there are an odd number of signature gene Z-scores in a given genetic profile, identifies the middle signature gene Z-score when such signature gene Z-scores are numerically ordered from lowest to highest, or highest to lowest, and ii) when there is an even umber of signature gene Z-scores in a given genetic profile, identifies the two middle signature gene Z-scores when such signature gene Z-scores are numerically ordered from lowest to highest, or highest to lowest, and adds these two middle signature gene scores and divides by two.
34. The system of claim 23. further comprising said biological sample from a subject.
35. The system of claim 23. wherein said mRNA expression levels were determined by a method selected from: quantitative sequencing, Northern blotting, Nuclease Protection Assays, In Situ hybridization, and RT-PCR.
36. The system of claim 23, wherein said plurality of different cancer drugs is selected from:Gemcitabine, Topotecan, Irinotecan, Cisplatin, Cyclophosphamide, Vinblastine, Cytarabine, Vorinostat, Luminespib, 5-Fluorouracil, Vincristine. Docetaxel, Paclitaxel, I-BRD9, GSK591. EPZ004777, Nilotinib, Alpelisib, and YK-4-279.
37. The system of claim 23, said subject is a human.
38. The system of claim 23. wherein said median finding algorithm: i) when there are an odd number of normalized mRNA expression levels in a given genetic profile, identifies the middle normalized mRNA expression level when such normalized mRNA expression levels are numerically ordered highest to lowest, and ii) when there is an even umber of normalized mRNA expression levels in a given genetic profile, identifies the two middle normalized mRNA expression levels when such normalized expression levels are numerically ordered from highest to lowest and adds these two middle normalized mRNA expression levels and divides by two.