Apparatus, method, and computer-readable storage medium for predicting patient response to PARP inhibitor treatment
A machine learning model using genetic markers predicts PARP inhibitor sensitivity in tumors, overcoming limitations of BRCA1/2 and HRD+ biomarkers, enhancing treatment efficacy for patients who have failed multiple therapies.
Patent Information
- Application Number
- PCT/IB2025/055622
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-31
- Publication Date
- 2025-12-04
AI Technical Summary
Current methods for identifying patients likely to benefit from PARP inhibitor therapy are insufficient, particularly for those with alternative HR deficiencies or mutations in DNA damage response genes, and existing biomarkers like BRCA1/2 mutations or HRD+ phenotype fail to predict responsiveness in patients who have failed multiple therapies.
A machine learning model is applied to clinical and molecular data, including genetic information from patients with tumors, to predict sensitivity to PARP inhibitors, using features such as TNFAIP3 mutation and other genetic markers, bypassing conventional biomarkers like BRCA1/2 mutations or HRD+ status.
The model effectively identifies patients likely to respond to PARP inhibitors, improving treatment outcomes by predicting sensitivity beyond conventional biomarkers, leading to prolonged overall survival and progression-free survival.
Smart Images

Figure IB2025055622_04122025_PF_FP_ABST
Abstract
Description
APPARATUS, METHOD, AND COMPUTER-READABLE STORAGE MEDIUMFOR PREDICTING PATIENT RESPONSE TO PARP INHIBITOR TREATMENTCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application No 63 / 654,731 filed on May 31, 2024, the content of which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Poly (ADP-ribose) polymerases (PARP) inhibitors (PARPi) are a class of anti-cancer drugs which compete with nicotinamide (NAD+) for the catalytically active site of PARP molecules (e.g., PARPI and / or PARP2, or other PARP proteins). PARPi target PARP enzymes (mainly PARPI and PARP2), which are DNA damage sensors that catalyze the formation of negatively charged poly(ADP -ribose) (PAR) chains to regulate protein assemblies and tune chromatin dynamics in response to genotoxic stress (see, e.g., Bai, P. Molecular Cell 58, June 18, (2015)). PARPi have been shown to be effective against homologous recombination repair deficient tumors in a synthetically lethal interaction. The synthetic lethality between PARP inhibition and BRCA1 / 2 mutation or depletion was first observed in 2005, where it was originally hypothesized that inhibition of PARPI activity would lead to replication fork collapse and the subsequent homologous recombination-dependent repair of these forks. An FDA-approved test may be used to identify patients with a homologous recombination deficient (HRD+) phenotype.
[0003] In 2014, olaparib (LYNPARZA®, AstraZeneca) was approved by the European Medicines Agency (EMA) and the US Food and Drug Administration (FDA) for the treatment of BRCA1 / 2 mutant ovarian cancers (see, e.g., Kraus, Mol. Cell 2015 v58(6): 902-910 (2015)). Several additional PARPi, including talazoparib, niraparib, rucaparib and veliparib have also recently been approved. PARPi target PARP enzymes (mainly PARPI and PARP2), which are DNA damage sensors that catalyze the formation of negatively charged poly(ADP-ribose) (PAR) chains to regulate protein assemblies and tune chromatin dynamics in response to genotoxic stress. Due to the relatively low frequency of BRCA1 / 2 mutations and the limited utility of FDA-approved tests for HRD+ phenotype, this limits the applicability of PARP inhibitors to the treatment of 10-15% of breast and ovarian tumors, 4-7% of pancreatic tumors and 1.5% of prostate carcinoma (see, Bryant et al., 2005; Iqbal et al., 2012; Oh et al., 2019).However, PARP inhibitors may have much wider applications, including the treatment of tumors with alternative HR deficiencies or mutations in other DNA damage response genes. Tumors with high levels of oxidative and replicative stress may also be sensitive to PARP inhibitors, irrespective of identification of HRD+ phenotype by currently approved methods (see, e.g., Majuelos-Melguizo et al., Oncotarget, 6(7) (2015); Michelena et al., Nat. Comm. 9(2678) (2018)). Moreover, HRD+ status as a predictive biomarker for PARPi response has been found to be highly insufficient for the majority of cancer patients who have failed multiple lines of therapy, leading to voluntary withdrawal for several PARPi in these contexts (see, Lee et al., J Gynecol Oncol. 2023 Mar;34(2):e51).
[0004] Thus, there is an unmet need to identify patients who can respond to and benefit from PARPi therapy but are not identifiable by BRCA1 / 2 mutations or the approved tests for HRD+ phenotype, particularly in the context of patients who have failed multiple previous therapies.SUMMARY
[0005] According to embodiments, the present disclosure further relates to apparatus and methods, and computer readable medium comprising instructions for determining whether a subject suffering from a cancer (e.g., breast cancer, ovarian cancer, pancreatic cancer, prostate cancer, and others) will benefit from treatment with an inhibitor of Poly (ADP-ribose) polymerases (PARPi).
[0006] In one aspect, a method disclosed herein is provided for predicting an outcome of treatment with a PARP inhibitor for a subj ect having a tumor, the method comprising: obtaining clinical and molecular data from the subject; applying a machine learning model to the clinical and molecular data to calculate a sensitivity metric corresponding to a likelihood that the subject will respond to the PARP inhibitor; and predicting the treatment outcome in the subject by evaluating the sensitivity metric.
[0007] In one embodiment, the machine learning model is based on one or more of a radiusneighbor, neural network, a random forest classifier, elastic net regularization model, gradient- boosted tress-based method, and a random forest regressor. In another embodiment,, when the machine learning model is a random forest classifier, the obtained clinical and molecular data is based on feature importance. In another embodiment, the feature importance is determined by SHapley Additive exPlanations (SHAP) analysis, Gini Importance, Mean Decrease inImpurity (MDI), Permutation Importance (Mean Decrease in Accuracy), or Local Interpretable Model-agnostic Explanations) (LIME).
[0008] In one embodiment, the obtained complex clinical and molecular data identified as an important feature include TNFAIP3 mutation, CTCF deletion, BCOR deletion, NFKBIA amplification, FGF3 amplification, BCL2 deletion, EGFR mutation, ERBB2 amplification, KRAS mutation and amplification, MCL1 amplification, NRAS mutation, ARID1 A mutation, GNAS mutation, AURKB deletion, SDHA amplification, FLCN deletion, GATA3 deletion, BRAF mutation, CTNNB1 mutation, VHL mutation, MYC amplification, RICTOR amplification, PIK3CA mutation, NF1 mutation, CDKN2A mutation and deletion, RBI mutation, PTEN mutation, STK11 mutation, CDKN2B deletion, TP53 mutation.
[0009] In one embodiment, the machine learning model comprises clinical and molecular information derived from subjects having the type of tumor and having been treated with the PARP inhibitor. In another embodiment, the clinical and molecular data comprises genetic information derived from a pre-treatment subject tumor or human tumor model system. In one embodiment, the clinical and molecular data comprising genetic information is derived from a tumor or human tumor model system that has failed one or more lines of therapy. In some embodiments, the method may not comprise use of conventional predictive biomarkers for PARP inhibitors as determined by methods to identify HRD status or BRCA mutations.
[0010] In one embodiment, the PARP inhibitor is selected from the group comprising olaparib, niraparib, rucaparib, talazoparib, saruparib, stenoparib, fuzuloparib, iniparib, pamiparib, mefuparib, saruparib, and venadaparib. In another embodiment, the PARP inhibitor is olaparib. In one embodiment, the PARP inhibitor is a next generation PARP inhibitor. In another embodiment, the PARP inhibitor is talazoparib or saruparib.
[0011] In one embodiment, the tumor is an ovarian cancer tumor, a breast cancer tumor, a pancreatic cancer tumor, a non-small cell lung cancer tumor, a prostate cancer tumor, a fallopian tube tumor, or a primary peritoneal tumor. In one embodiment, the tumor is a solid breast cancer tumor. In another embodiment, the tumor is a solid ovarian cancer tumor.
[0012] In one embodiment, the method further comprises administering one or more doses of the PARP inhibitor to the subject when the subject is determined to be sensitive to the PARP inhibitor. In some embodiments, a sensitivity metric is provided. In one embodiment, the sensitivity metric is a score and evaluating the sensitivity metric comprises comparing the score to a median threshold. In another embodiment, the sensitivity metric is a score and evaluatingthe sensitivity metric comprises comparing the score to an area under the dose response curve. In some embodiments, the sensitivity metric is an area under the curve.
[0013] Another aspect of the invention provides a method for predicting, for a subject having a tumor that tests positive for a BRCA 1 or BRCA2 mutation or a defect in the DNA homologous recombination repair (HRR) pathway, whether the subject will not respond to a PARP inhibitor, the method comprising: obtaining complex clinical and molecular data pertaining to the subject, applying a machine learning model to the clinical and molecular data to generate a composite biomarker for the subject, the composite biomarker comprising genes based on the type of tumor and a PARP inhibitor; calculating a probability that the subject will respond to the PARP inhibitor; and predicting the treatment outcome in the subject by comparing the calculated probability to a threshold probability for treatment responsiveness.
[0014] In one embodiment, the machine learning model comprises complex clinical and molecular information derived from subjects having the same type of tumor and having been treated with the same PARP inhibitor as is being contemplated for the subject. In another embodiment, the molecular data comprises genetic information that is derived from a pretreatment subject tumor or human tumor model system. In one embodiment, the genetic information is derived from a tumor or human tumor model system that has failed one or more lines of therapy.
[0015] In some embodiments, the PARP inhibitor is selected from the group comprising olaparib, niraparib, rucaparib, talazoparib, saruparib, stenoparib, fuzuloparib, iniparib, pamiparib, mefuparib, saruparib, and venadaparib. In one embodiment, the PARP inhibitor is olaparib. In another embodiment, the PARP inhibitor is a secnext generation PARP inhibitor. In another embodiment, the PARP inhibitor is talazoparib.
[0016] In one embodiment, the tumor is an ovarian cancer tumor, a breast cancer tumor, a pancreatic cancer tumor, a non-small cell lung cancer tumor, a prostate cancer tumor, a fallopian tube tumor, or a primary peritoneal tumor. In one embodiment, the tumor is a solid breast cancer tumor. In another embodiment, the tumor is a solid ovarian cancer tumor.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] FIG. 1 depicts a flow diagram of aspects of an exemplary method of the present disclosure.
[0018] FIG. 2 is a visualization of training and validation of an exemplary machine learning model of the method of the present disclosure. In particular, FIG. 2 is a visualization of a t- distributed Stochastic Neighbor Embedding (tSNE) projection, based on the tSNE used in Example 1 to predict patient sensitivity to a treatment. This visualization demonstrates that the exemplary machine learning model has effectively learned the relationship between genetic expression and sensitivity.
[0019] FIG. 3A depicts an exemplary Kaplan-Meier plot of outcome measures comparing ovarian cancer patients predicted, by the methods described herein, to be sensitive (light blue) or less-sensitive (grey) to a PARPi (e.g., olaparib, niraparib, rucaparib). Overall survival (OS) was significantly increased (>36 months, p-value < 0.05) in predicted PARPi sensitive patients compared to predicted PARPi less-sensitive patients. FIG. 3B is a cox analysis evaluating potential confounding factors, such as PARPi used, sample source (metastatic or primary tumor), diagnosis age, treatment age, cancer sub type, and model inputs. The analysis shows that the results of the machine learning model are based on the prediction of the sensitivity metric and not confounded by other clinical features.
[0020] FIG. 4A is a visualization of a gene set enrichment analysis (GSEA) performed to assess differences across the validation cohort of ovarian cancer patients considered to be sensitive to PARPi using the method described in Example 1. FIG. 4B is a graphical representation of sensitivity predictions generated by Example 2.
[0021] FIG. 5 depicts an illustration of a flow diagram of aspects of the method of the present disclosure. The flow diagram of FIG. 5 deploys the exemplary machine learning model used in Example 2 predict patient sensitivity to a treatment. FIG. 5 demonstrates how the exemplary machine learning model generates a composite biomarker in cancer patients using complex clinically relevant data informed by related but distinct data sets. These distinct data sets can include unique features of specific drugs, mechanism of action and target features associated with a drugs activity, and a broad corpus of information related to relevant clinical features. Other context can be provided by functional data (genome wide perturbation screens), proteinprotein interaction networks and protein-DNA interaction networks. The composite biomarker can include biologically interpretable drug response predictions (e.g., drug sensitivity) and vulnerability networks associated with the drug response.
[0022] FIG. 6 depicts a series of graphical representations characterizing the olaparib-treated ovarian cancer validation cohort discussed in Example 2. The graphical representations relateto, from left to right, type of ovarian tumor, age at cancer diagnosis (Dx), sample type (primary or metastatic), and the number of therapeutics (Tx) administered before olaparib treatment. Input data required for drug response predictions were obtained for 46 ovarian cancer patients. Treatment histories for these patients were complex, with diverse drug combinations and a variable number of prior lines of treatment prior to receiving olaparib. Samples were restricted to NGS data from biopsies collected no more than two years prior to the initiation of olaparib, with the aim of minimizing the impact of clonal evolution, genetic drift, and selective pressures from intervening treatments on the relevance of the Zephyr model drug response predictions for olaparib.
[0023] FIGS. 7A-7C depict a series of exemplary Kaplan-Meier plots of outcome measures evaluating olaparib response predictions using real-world ovarian cancer patient outcomes. Comparisons were performed across the conventional HRD methodology (FIG. 7A) with methods set forth in Example 2 (FIG. 7B) for outcomes related to real world (rw)-overall survival (OS) and rw-progression free survival (PFS). Additionally, similar to FIG. 3B, a sensitivity analysis was conducted to assess the presence of features confounding the sensitivity determination alongside of the known BRCA status (FIG. 7C). Conventional HRD' scores were derived from Whole Exome Sequencing (WES) data, using methods that assess Allelic Imbalance (Al), somatic and germline hits to HRR genes, and an HRD threshold of >42. Stratification of ovarian cancer patients treated with olaparib by conventional HRD scores does not result in significant differences in rw-overall (rw-OS) (FIG. 7A, top) or progression free (rw-PFS) (FIG. 7BA, bottom) survival. In contrast, stratification of the same patients by Zephyr’s method results in statistically and clinically significant prolonged rw-OS (FIG. 7B, top) and rw-PFS (FIG. 7B, bottom). Response predictions were not confounded by BRCA1 / 2 somatic or germline mutations, HRD status or assessed potential clinical confounders (FIG. 7C).
[0024] FIG. 8 is a visualization of a gene set enrichment analysis (GSEA) performed to assess differences across the validation cohort of ovarian cancer patients considered to be sensitive to olaparib using the method described in Example 2. This analysis was conducted to evaluate the differences across signaling networks in patients predicted to be sensitive vs resistant to a PARP inhibitor, e.g., olaparib.
[0025] FIG. 9 depicts a side-by-side comparison of a graphical representation of vulnerability networks generated by the exemplary machine learning model of Example 2, wherein the leftside displays a vulnerability network enriched in olaparib-sensitive ovarian tumors and the right side displays a vulnerability network enriched in olaparib-insensitive ovarian tumors.
[0026] FIG. 10 is a graphical representation showing rank order of feature importance assessed using Gini Importance, as deployed in the exemplary machine learning model of Example 3. These data were used to define a minimal subset of genes based on their variation across a population to optimize impact on model performance.
[0027] FIGS. 11A and 11B are a set of exemplary Kaplan-Meier plot of outcome measures, where patients predicted, using the machine learning model of Example 3, to be sensitive appear as light blue and patients predicted to be less-sensitive appear as grey. FIGS. 11 A and 1 IB each show (top) evaluation of olaparib response predictions using real-world breast and ovarian cancer patient outcomes to define model differences across real world overall survival and (bottom) a sensitivity analysis conducted to assess the presence of features that could confound the interpretation of the model. The sensitivity analysis was performed similarly to that of FIG. 3B and FIG. 7C. Outputs of the exemplary machine learning model classified 22 of 24 breast cancer patients as less-sensitive to the PARPi and 85 of 109 ovarian cancer patients as less-sensitive to the PARPi.
[0028] FIG. 12 is a visualization of a gene set enrichment analysis (GSEA) performed to assess differences across olaparib predicted insensitive vs. sensitive tumors using expression data using the method described in Example 3. This analysis was conducted to evaluate the differences across signaling networks as set forth in Example 3.
[0029] FIG. 13 is a graphical presentation of a correlation between score and area under the curve (AUC) values predicted by the exemplary machine learning model (e.g., random forest regressor) used in Example 4. The score (x-axis) predicted AUC values y-axis were plotted for test set samples (n = 739, Pearson’s correlation = -0.72).
[0030] FIG. 14 depicts an exemplary Kaplan -Mei er plot of outcome measures comparing ovarian cancer patients predicted, using the exemplary machine learning model of Example 4, to be sensitive (light blue) or less-sensitive (grey) to olaparib.
[0031] FIG. 15 is a visualization of a gene set enrichment analysis (GSEA) performed to assess differences across patients predicted, by the exemplary machine learning model of Example 4, to have tumors that insensitive or sensitive to treatment with olapirib. This analysis can be conducted to illustrate the differences across signaling networks using Example 4. Normalized enrichment scores (NES) (x-axis) indicate enrichment of corresponding gene sets (y-axis) inpredicted olaparib-sensitive tumors (NES >= 1) or predict olaparib less-sensitive tumors (<= - 1). Circle size indicates number of genes per set and color denotes a log-transformed p-value.DETAILED DESCRIPTION
[0032] The term “a” or “an” refers to one or more of that entity, i.e., can refer to plural referents. As such, the terms “a,” “an,” “one or more,” and “at least one” are used interchangeably herein. In addition, reference to “an element” by the indefinite article “a” or “an” does not exclude the possibility that more than one of the elements is present, unless the context clearly requires that there is one and only one of the elements.
[0033] Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the device or the method being employed to determine the value, or the variation that exists among the samples being measured. Unless otherwise stated or otherwise evident from the context, the term “about” means within 10% above or below the reported numerical value (except where such number would exceed 100% of a possible value or go below 0%). When used in conjunction with a range or series of values, the term “about” applies to the endpoints of the range or each of the values enumerated in the series, unless otherwise indicated. As used in this application, the terms “about” and “approximately” are used as equivalents.
[0034] Tumors with homologous recombination repair deficiencies (HRD+) are susceptible to PARP inhibitors (PARPi), which exploit a synthetic lethality by targeting an essential base excision repair (BER) pathway. Four drugs are approved for use with this biomarker (i.e., HRD+). However, current patient selection biomarkers designed to enrich for PARPi response (HDR+, BRCAl / 2mut) in late-line cancer patients are insufficient. This has led to withdrawals of these drugs in certain indications, including late-line ovarian cancer.
[0035] Disclosed herein are alternative, machine learning based methods for identifying PARPi (e.g., olaparib) sensitive cancer patients.
[0036] The indications for which PARP inhibitors (PARPi) have been approved for use in the United States are summarized in Table 1. In 2014, olaparib (LYNPARZA®) was the first PARPi approved by the Food and Drug Agency (FDA) and European Medicines Agency (EMA) as a monotherapy for the treatment of advanced, germline BRCA mutated ovarian cancer. In 2017, this was extended to include maintenance therapy of recurring ovarian,fallopian, and primary peritoneal tumors, regardless of BRCA mutational status. Olaparib has also been approved for the treatment of germline BRCA1 / 2 mutated HER2 -negative breast and metastatic pancreatic cancer in 2018 and 2019, respectively. Most recently, olaparib was approved for the treatment of HRD-positive metastatic castration-resistant prostate cancer.
[0037] Several other PARP inhibitors, including rucaparib (RUBRACA®), niraparib (ZEJULA), and talazoparib (TALZENNA®) have also been approved for use in various clinical settings. In 2016, rucaparib was granted an accelerated approval for the treatment of germline or somatic BRCAl / 2-mutated advanced ovarian carcinomas, following multiple chemotherapy treatments. Subsequently, rucaparib maintenance therapy was approved in 2018 for recurring ovarian, fallopian and primary peritoneal, regardless of BRCA mutational status. In May 2020, rucaparib gained FDA approval for the treatment BRCA1 / 2 mutated metastatic castration-resistant prostate cancer.Table 1 - Summary of US approved indications for PARP inhibitors
[0038] Targeting DNA damage repair with PARP inhibitors has been approved for the treatment of certain solid tumors including those of ovarian, breast, pancreatic, and prostate cancer. Presently, two types of tests are clinically used to determine PARPi sensitivity, including testing for DNA homologous repair deficiency (HRD+) and testing for mutation in BRCA1 or BRCA2 genes. HRD detection methods can use mutational profiles (e.g., HRDetect) which require whole-exome sequencing or whole-genome sequencing data to achieve accurate mutation profiles, the gene expression signature, or CNV features, including the genomic scar score (GSS), Amoy, Myriad HRD, and Foundation HRD methods. Unfortunately, these methods do not identify all patients that can benefit from PARPi treatment; for example, the HRD and BRCA tests only identify about -25% of ovarian cancer patients as being candidates for treatment with PARP inhibitors. Moreover, within the patients harboring such positive predictive biomarkers, only -60% respond in first line scenarios. HRD and BRCA1 / 2 status are even worse predictors of PARPi sensitivity in later lines of therapy (e.g., after a patient has failed first and second lines of therapy). Disclosed herein are methods for applying a novel PARPi sensitivity signature for patients with solid tumors (e.g., breast, ovarian, prostate, or pancreatic tumors), not defined by traditional HRD or BRCA1 / 2 genomic biomarkers and signatures, to identify patients that are likely sensitive to PARPi therapy.
[0039] The present disclosure describes methods for determining whether a subject (e.g., a patient) suffering from a cancer (e.g., ovarian cancer or another solid tumor such as a breast cancer, ovarian cancer, pancreatic cancer, and prostate cancer) is sensitive to treatment with a PARPi, such as olaparib or another PARPi, including, for example, niraparib, rucaparib, talazoparib, saruparib, stenoparib, fuzuloparib, iniparib, pamiparib, mefuparib, and venadaparib. FDA approved PARP inhibitors are listed in Table 1. In some embodiments, the PARP inhibitor is a second-generation PARP inhibitor (e.g., talazoparib) having the ability totrap PARP1 and / or PARP2 to the sites of DNA damage. The PARP enzyme-inhibitor complex “locks” onto damaged DNA and prevents DNA repair.
[0040] In some embodiments, the PARP inhibitor is a third-generation PARP inhibitor (e.g., saruparib) having higher selectivity toward PARP1 and lower or no selectivity toward PARP2 in order to reduce off-target effects and improve the safety profile compared to first generation PARP inhibitors such as niraparib. Next generation PARP inhibitors can potentially be administered at higher doses than the first-generation PARP inhibitors.
[0041] The method comprises applying a machine learning model to clinical data, genetic expression data, and molecular data from a patient diagnosed with cancer to predict sensitivity of the patient to one or more PARPi. The machine learning model may be one of the exemplary machine learning models deployed herein, including in Examples 1-4. The molecular data, which can include mutational and copy number status, includes, but is not limited to, detected point mutations, additions, deletions, frame shifting mutations, translocations and copy number gains, amplifications, losses, homozygous deletions and the like. The clinical data may be real- world data, which may include genomic sequencing data, gene expression data, RNA data, protein expression levels, age, gender, treatment regimen, geometric location, patient outcome, and the like, allowing the methods described herein to be applied without limits. In some embodiments, the plurality of genes may be related (e.g., in a common cellular pathway) or may be unrelated (e.g., in different cellular pathways). These relationships may be determined. The real -world data may include, for at least training purposes, sequencing data from subjects that have been clinically diagnosed with a cancer (e.g., ovarian cancer or another solid tumor such as a breast cancer, ovarian cancer, pancreatic cancer, and / or prostate cancer) and have received a PARPi as part of the regimen for treating that disease. Treatment outcomes for these patients are also used in training. Such treatment may include on-label FDA approved therapeutics, treatment through a clinical trial, or off-label use. The methods disclosed herein deploy a machine learning model trained on real-world data to predict responsiveness of a particular patient to a therapy. Such prediction may be considered a partition, whereby subjects are partitioned into two or more groups, where one group may derive benefit to or respond to treatment with a PARPi (“sensitive”) and one group may not derive benefit from or not respond to treatment with a PARPi (“non-sensitive”).
[0042] Generally, the exemplary machine learning models herein may be deployed in a method of the present disclosure to identify a patient as a responder to a therapy based on an analysisof clinical and molecular data. Each exemplary machine learning model has been previously trained on a patient population diagnosed with a tumor (e.g., an ovarian cancer solid tumor) and treated with a PARPi (e.g., olaparib). The method generally includes the steps of acquiring patient data, the patient data including but not limited to clinical data and molecular data, including genomic sequencing data of the patient, in addition to a corpus of data including unique features of specific drugs, mechanism of action and target features associated with a drugs activity, and a broad corpus of information related to relevant clinical features, and applying an exemplary machine learning model to the acquired patient data to generate a prediction regarding whether the patient would derive benefit from treatment with a PARPi (e.g., is sensitive to a PARP inhibitor, e.g., exhibits at least one improved clinical outcome, such as increase in progression free survival or overall survival). Other context can be provided by functional data (genome wide perturbation screens), protein-protein interaction networks and protein-DNA interaction networks. In some embodiments, the exemplary machine learning model may be trained on a corpus of reference data including the patient population and having at least corresponding clinical data and molecular data by methods known in the art, including supervised learning methods and unsupervised learning methods. For instance, the molecular data may include, for a given patient population, sequencing or copy number data and may include molecular information from publicly available patient data sets, electronic medical record systems, medical insurance companies, commercial sequencing companies, health networks, or the like. In some embodiments, the sequencing or copy number data may be obtained from a population of subjects diagnosed with ovarian cancer and treated with olaparib. In some embodiments, the sequencing and copy number data may be obtained from the patient samples by laboratory diagnostics such as e.g., karyotyping, fluorescence in situ hybridization (FISH), comparative genomic hybridization, polymerase chain reaction (PCR), DNA microarray, DNA sequencing, multiplex ligation-dependent probe amplification, single strand conformation polymorphism, denaturing gradient gel electrophoresis, heteroduplex analysis, restriction fragment length polymorphism, Whole Genome Sequencing (WGS), Whole Exome Sequencing (WES), Targeted DNA Sequencing, RNA Sequencing (RNA-Seq), Single-Cell Sequencing, Long-Read Sequencing, Short-Read Sequencing, Chromatin Immunoprecipitation Sequencing, Methylation Sequencing and High Throughput Sequencing.
[0043] The exemplary machine learning models described herein may benefit subjects or patients not yet approved for treatment with a PARPi (e.g., olaparib) but whose clinical data,molecular data demonstrate a high probability of sensitivity to treatment with a PARPi or have previously demonstrated sensitivity to treatment using a PARPi. Such similarity may include similar mutational statuses to subjects that demonstrated a high probability of sensitivity to treatment with a PARP inhibitor or have previously demonstrated sensitivity to treatment using a PARP inhibitor. In some embodiments, the similarity is defined by a comparison of a probability metric to a probability threshold, as will be described herein.
[0044] In determining therapeutic sensitivity of a new patient to a particular drug (e.g., PARPi), the methods disclosed herein include obtaining clinical and molecular data for the new patient. The molecular data, which can include mutational and copy number status, includes, but is not limited to, detected point mutations, additions, deletions, frame shifting mutations, translocations and copy number gains, amplifications, losses, homozygous deletions and the like. The clinical data may be real-world data, which may include genomic sequencing data, gene expression data, RNA data, protein expression levels, age, gender, treatment regimen, geometric location, patient outcome, and the like, allowing the methods described herein to be applied without limits. The clinical data may also include cancer subtype, patient sex, and tumor source (e.g., primary or metastatic)). An exemplary machine learning model may then be applied to the obtained clinical and molecular data. The output of the exemplary machine learning model may be a sensitivity metric including a probability that the new patient is sensitive or insensitive to treatment with the particular drug, a score that can be compared with a threshold, a calculation of an area under the curve, and / or a classification of the new patient as sensitive or insensitive to treatment. In certain embodiments, a patient predicted as a responder to treatment with a PARPi is one with a predicted improved outcome as measured by an increase in progression free survival or overall survival, an improved Response Evaluation Criteria in Solid Tumours (RECIST) criteria, or desired changes in other outcomes associated with therapeutic response.
[0045] In some embodiments, analysis of the genomic sequencing data comprises use of a variant annotation and effect prediction tool that can annotate and predict the effects of genetic variations such as amino acid changes. Exemplary such tools for use in annotating genomic molecular data include: snpEff; SIFT (see https: / / sift.bii.a-star.edu.sg / www / code.html); and Polyphen2 (see https: / / sift.bii.a-star.edu.sg / www / code.html). Additional methods of identifying commonly described hotspot mutations can be found, e.g., in Chang et al., Nat Biotechnol. 2016 Feb; 34(2): 155—163; Travino, V. , Database (Oxford). 2020; 2020: baaa025;or using Cancer Dependency Map (DepMap; see https: / / bioconductor.org / packages / release / data / experiment / html / depmap.html).
[0046] In some embodiments, if the new patient is determined to respond to treatment with a PARPi, the method further comprises administering the treatment to the new patient or generating instructions for administering the treatment to the new patient. Administering the treatment may include administering one or more doses of a PARP inhibitor to the subject for an appropriate amount of time or otherwise administering a PARP inhibitor as prescribed on the label.
[0047] In some embodiments, the method includes a step wherein an experimental sample from the patient is acquired. The experimental sample may comprise any biological sample including but not limited to blood, serum, plasma, urine, bile, sputum, tumor samples, stool, pleural fluid, synovial fluid, CSF fluid, any tissues, organs, saliva, DNA / RNA, hair, nail clippings, or any other cells or fluids provided from a human body. In some embodiments, the experimental sample may be at least one selected from the group consisting of: blood, serum, plasma, urine, bile, sputum, tissue, cerebrospinal fluid, bone marrow aspirate, breast milk, saliva, synovial fluid, swabs, stool, and bronchial fluid. In some embodiments, the experimental sample may be from a patient that may or may not be presently diagnosed with a cancer (e.g., breast cancer, ovarian cancer, pancreatic cancer, or prostate cancer) and may or may not have mutations in BRCA1 / 2 or be recognized as having an HRD+ phenotype.
[0048] In some embodiments, a sensitivity of the patient to a PARP inhibitor can be determined based on the output of the model. In some embodiments, a method is provided for applying a machine learning model capable of identifying patients with tumors as being sensitive to treatment with a PARPi.
[0049] FIG. 1 depicts a flow chart of aspects of an exemplary method deploying a machine learning model to predict whether cancer patients, e.g., ovarian cancer patients, are sensitive or insensitive to treatment with a PARPi, e.g., olaparib.
[0050] At step 102 of method 100, clinical and molecular data, including genomic sequencing data, can be obtained from a patient. Mutation and copy number status for each gene is available from most commercial multi-gene NGS panels (e.g., Foundation, MSK-IMPACT, DFCI- ONCOPANEL, etc.). The patient may be a subject diagnosed with, for example, ovarian cancer and may be being considered for treatment with, for example, olaparib. The obtained molecular data may include sequencing or copy number data and may include sequencing data from tumorcells, non-tumor cells, somatic cells, metastatic cells, or the like. The clinical data may include one or more of electronic health records, cancer-gene dependency data, disease, sex, and stage (or tumor source as a surrogate). Stage may include primary or metastatic. The genomic sequencing data may include molecular data collected in clinical settings.
[0051] In embodiments, the obtained genomic sequencing data includes an expression profile of the genetic makeup of the patient. The expression profile may include a complete set of genes or may include a subset of genes. A complete set of genes may one that is provided by next generation sequencing and / or by a commercially available genetic panel. When only a subset of genes is provided and the corresponding machine learning model requires expression of additional genes, expression of the additional genes can be estimated. The estimation can be performed using, in one instance, Mut2Ex, a machine learning-based approached developed by Ramchandran and Baron, (see Ramchandran and Baron, Reconstructing gene expression and knockout effect scores from DNA mutation (Mut2Ex): methodology and application to cancer prediction problems (2022)). Such an estimation method can be employed in the methods herein to estimate unknown expression within tumor gene expression profiles. The estimation method can be based on 1) genetic information readily available in real -world clinical settings, such as those data derived from commercial next generation sequencing panels and 2) a limited subset of clinical information (including, but not limited to an oncologic code, patient sex, and cancer stage of the tumor). When Mut2Ex is deployed, the input thereto may be structured as follows: genetic mutations may be encoded as either 0, indicating no hotspot mutation was detected in that gene or 1, indicating a hotspot mutation was detected in that gene; SnpEff, Polyphen2 and SIFT effect prediction tools may be used to annotate the mutations to identify potential deleterious variants as the “hotspot” mutations; and mutations that fit the following criteria are selected as hotspot mutations: mutations annotated as HIGH impact by snpEff, which include effects such as frame shift vari ant, stop gained, start lost, stop lost, etc.; mutations identified as loss of function variants by snpEff; mutations identified as nonsense- mediated mRNA decay variants by snpEff; mutations annotated as probably damaging or possibly damaging by Polyphen2; or annotated as deleterious by SIFT. The copy number variation from the genes may then be encoded in an expanded manner. Further, each gene may be expanded to two values in the input of the data - gene amplification and gene deletion. In some embodiments, gene amplification comprises at least four copies of the gene. In some embodiments, gene amplification comprises four copies, five copies, six copies, seven copies,eight copies, or more. Only such major copy number events may be considered and not intermediated loss or gains. In some embodiments, gene deletion comprises a homozygous deletion where there is evidence of loss of two copies or evidence of a functional loss of both copies, as with loss of heterozygosity of a gene. The clinical features of each sample can then be fed into a pretrained BioBERT see, e.g., Lee et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics; 2020 Feb 15 ;36(4): 1234-1240). The term “BioBERT” refers to a specialized version of BERT (Bidirectional Encoder Representations from Transformers) which is a pre-trained natural language processing model. The BioBERT model includes additional pre-training on biomedical literature, enabling it to understand and process biomedical text more effectively than the general BERT model on natural language model tasks.
[0052] Regardless of how the clinical and molecular data is obtained, a machine learning model may be applied to the data to generate at least a sensitivity metric at step 104 of the method 100. The sensitivity metric may be a probability that the new patient is sensitive or insensitive to treatment with the particular drug, a score that can be compared with a threshold, a calculation of an area under the curve, and / or a classification of the new patient as sensitive or insensitive to treatment, among others.
[0053] In embodiments, the machine learning model comprises a regression model. In various embodiments, the regression model is a linear regression model, a multiple linear regression model, a polynomial regression model, a logistic regression model, a ridge and lasso regression model, or an assumption of linear regression model. In embodiments, the machine learning model is one of a radius-neighbor, a neural network, an elastic net regularization model, a gradient-boosted trees-based method (e.g., XGBoost, LightGBM), a pre-trained BioBERT, a regression model, random forest classifier, and random forest regressor. As would be understood by one of ordinary skill in the art, a random forest classifier or regressor is an ensemble learning method used for classification tasks in machine learning, and it combines the predictions of multiple decision trees to make a final prediction. Advantages of random forest classifier or regressor include prevention of overfitting, handling of missing values, estimating feature importance.
[0054] In step 108 of the method 100, the calculated sensitivity metric is evaluated to predict whether patient is a responder to the PARPi.
[0055] Embodiments of the subject matter and the functional operations described in this disclosure, such as the training, applying, and predicting using machine learning models as described above with reference to FIG. 1 through FIG. 15, can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this disclosure and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this disclosure, such as the training, applying, and predicting using machine learning models, can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machinegenerated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0056] In embodiments, the methods described herein can be performed on data processing hardware encompassing all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. Such apparatuses can also include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit) and can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0057] A computer program implanting the methods describe herein, which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more Scripts stored in a markup languagedocument, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0058] The processes and logic flows described in this disclosure can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0059] Computers suitable for the execution of a computer program include, by way of example, general or special purpose microprocessors or both or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random-access memory or both. Elements of a computer include a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM, EEPROM) and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be Supplemented by, or incorporated in, special purpose logic circuitry.
[0060] To provide for interaction with a user, embodiments of the subject matter described in this disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LED (light-emitting diode), or LCD (liquid crystal display)-based monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse, bywhich the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser.
[0061] Embodiments of the subject matter described in this disclosure can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this disclosure, or any combination of one or more Such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network.
[0062] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the user device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received from the user device at the server.EXAMPLES
[0063] The disclosure will now be illustrated with working examples, which is intended to illustrate the working of disclosure and not intended to take restrictively to imply any limitations on the scope of the present disclosure. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice of the disclosed methods and compositions, the exemplary methods, devices, and materials are described herein. It is to beunderstood that this disclosure is not limited to particular methods, and experimental conditions described, as such methods and conditions may apply.Example 1. A Regression Model for Cancer Patient Sensitivity to a PARPi
[0064] Sensitivity of ovarian cancer patients to treatment with PARPi (e.g., olaparib, niraparib, rucaparib camsylate) was evaluated. To this end, obtained clinical and molecular data included patient-specific features and clinical genomic information from 1879 genes. The obtained clinical and molecular data were fed to a machine learning model (e.g., a regressor) to generate a score indicative of whether the patient is sensitive to the PARPi. The method is agnostic to the choice of regressor, which may be a radius-neighbor, random forest, or gradient-boosted method. The machine learning model was trained on potency data from cell lines treated with at least one of a plurality of PARPi (e.g., olaparib, niraparib, rucaparib) and patient-specific data that includes predicted treatment outcomes with the at least one PARPi. Training data was based on approximately 6000 ovarian cancer cases identified to support training of the machine learning model. 80% of the approximately 6000 ovarian cancer cases were used as a training set for the machine learning model for determining the relationship between gene expression (of the 1879 genes) and response to the at least one PARPi. The remaining 20% of the approximately 6000 ovarian cancer cases were used to assess the accuracy of the machine learning model and predictions generated by the machine learning model were used for tuning. The output of the machine learning model, or sensitivity metric, was a score that was compared to a median threshold or an area under the dose response curve to identify the patient as a responder or not.
[0065] The machine learning model was further validated using input data required for drug response predictions were obtained for 135 ovarian cancer patients with confirmed treatment with a PARPi, patient outcomes and next generation sequencing (NGS) data less than 5 years prior to PARPi treatment.
[0066] FIG. 2 depicts a visualization of clustered gene expression in a t-distributed stochastic neighbor embedding (t-SNE). t-SNE projection in both the training and validation data sets are shown. t-SNE projection of patient gene expression is colored by sensitivity (training: green = sensitive, n = 2365, gray = insensitive, n = 3716 and validation: red = sensitive, n = 94, blue = insensitive, n = 151). The visualization verifies that the regressor has effectively learned therelationship between expression and sensitivity (evaluation samples tend to share sensitivity status with their associated clusters).
[0067] FIG. 3 A depicts an exemplary Kaplan-Meier plot of outcome measures comparing 135 ovarian cancer patients predicted, by the methods described herein, to be sensitive (light blue) or less-sensitive (grey) to a PARPi (e.g., olaparib, niraparib, rucaparib). Each patient was treated with PARPi no more than 5 years out from initial biopsy. Overall survival (OS) was significantly increased (>36 months, p-value < 0.05) in predicted PARPi sensitive patients compared to predicted PARPi less-sensitive patients. FIG. 3B is a cox analysis evaluating potential confounding factors, such as PARPi used, sample source (metastatic or primary tumor), diagnosis age, treatment age, cancer sub type, and model inputs. The analysis shows that the results of the machine learning model are based on the prediction of the sensitivity metric and not confounded by other clinical features.Example 2. Stratification of Late-line Cancer Patients Sensitivity to PARP Inhibitors
[0068] Clinical, translational, and chemical data were transformed into rich ‘machine learningready’ representations using custom encoders. These data representations were supplied to a machine learning model, including a neural network, as shown in FIG. 5. The machine learning generates a composite biomarker including a sensitivity metric and a prediction of a vulnerability network, an example of which is shown in FIG. 9. In this way, explainable drug response predictions for individual cancer patients can be generated based on patient-specific input data, such as clinico-genomic data and specific features of the drug of interest (see FIG. 5, left panel).
[0069] A composite biomarker, as used herein, refers to a novel multimodal polygenic biomarker indicative of, e.g., drug sensitivity in a patient. More specifically, generating a composite biomarker comprises a series of steps which use molecular data, genetic data, complex clinical data, and drug sensitivity data from human tumor cells to predict patient sensitivity to a treatment and classify human tumor cells into groups called "Vulnerability Networks" herein. A vulnerability network comprises a gene or set of genes that, when perturbed, result(s) in the reduction of viability. A genetic vulnerability within a cell indicates a gene that is required for the viability of that specific cell. In relation to a tumor cell, a genetic vulnerability may also be classified as a genetic perturbation as related to a gene which has aberrantly high RNA expression levels, a gene duplication, or an oncogenic mutation whichenhances the proliferative ability and viability of a specific tumor cell. The composite biomarker comprises characterization of the vulnerability network and its network of perturbation sensitivities, including performing post hoc molecular evaluation, differentially enriched mutations and copy number alterations, and gene set enrichment analysis (GSEA) on differentially expressed genes, and incorporating patient outcomes in order to make a drug response prediction.
[0070] As introduced above, for each prediction, two outputs are generated see FIG. 5, right panel). These two outputs include a drug sensitivity prediction for a specific drug (i.e., sensitivity metric) and a Vulnerability Network™ (i.e., vulnerability network) representing predicted perturbation sensitivities in a tumor. Sensitivity thresholds are set based on predictions from large real-world patient cohorts.
[0071] For each prediction, two model outputs are generated (right panel): a drug sensitivity prediction for a specific drug and an operational network representing predicted perturbation sensitivities in a tumor. Sensitivity thresholds are set based on predictions from large real -world patient cohorts. For drug combination predictions, each drug in the combination is run through the model individually. A predicted response to the combination per patient, is made based on concordant predictions for each drug in the combination (e.g., predicted sensitive to all drugs in a combination is classified as 'sensitive' for that combination).
[0072] In generating the machine learning model, the machine learning model was trained on potency data from cell lines treated with olaparib and other PARP inhibitors, patient-specific data that includes predicted treatment outcomes with olaparib, and gene expression data. Additional inputs in the training data included unique features of specific drugs, mechanism of action and target features associated with a drugs activity, and a broad corpus of information related to relevant clinical features. Other context was provided by functional data (genome wide perturbation screens), protein-protein interaction networks and protein-DNA interaction networks.
[0073] To validate the machine learning model, input data required for drug response predictions were obtained for 48 ovarian cancer patients with confirmed treatment with the PARP inhibitor olaparib, patient outcomes and next generation sequencing (NGS) data less than 2 years prior to PARPi treatment. FIG. 6 shows the baseline characteristics of real-world olaparib-treated ovarian cancer patients for response prediction validation, including type ofovarian cancer, age at cancer diagnosis, number of patients (in this case, 46), and the number of therapeutic interventions of a patient prior to olaparib.
[0074] Results from conventional HRD+ stratification methods (BRCA1 / 2 mutation status, HRD+) for determining olaparib sensitivity were contrasted with the machine learning model predictions. Patient sensitivity outcomes were represented as real world-overall survival (rw- OS) or rw-progression-free survival (rw-PFS). Conventional HRD+ stratification did not result in statistically significant improvement in either of the rw outcomes, whereas, for the same cohort, patient stratification via the Composite Biomarker significantly improved outcomes for both real world end points. (Compare FIG. 7A to FIG. 7B). These results were unaffected by BRCA1 / 2 mutations, HRD+ status, or other clinical confounders (FIG. 7C), underscoring that a certain portion of the olaparib-sensitive patients identified by the methods disclosed herein cannot be detected through conventional HRD+ stratification methods.
[0075] Further analysis of the identified olaparib-sensitive patients showed evidence of intact DNA damage response pathways in these patients. Gene expression profiles for ovarian patient tumors were used as input for gene set enrichment analysis (GSEA) comparing the predicted olaparib sensitive vs. insensitive tumors identified by the methods disclosed herein.
[0076] GSEA conducted across olaparib predicted insensitive vs. sensitive ovarian cancer tumors were positively enriched for multiple DNA damage response pathways, suggesting that despite presence of BRCA1 / 2 mutations and HRD+ status these pathways remain functional in these tumors, potentially bypassing the synthetic lethality mechanism exploited by PARP inhibitors (FIG. 8). Predicted-sensitive tumors were positively enriched for proliferation signatures and negatively enriched for metastatic, stem-like, and EMT signatures.
[0077] Vulnerability networks represent predicted perturbation sensitivities in tumors. The olaparib-sensitive tumors identified by the methods disclosed herein are characterized by vulnerability networks enriched for HR machinery, suggesting a reliance on this pathway in these tumors that may explain PARP inhibitor sensitivity (FIG. 9A). Conversely, as supported by GSEA (FIG. 8), networks enriched in insensitive tumors show no predicted susceptibility to HR pathway perturbation. Instead, these tumors are characterized by a predicted dependence on cell motility and stress response mechanisms, which could potentially bypass PARPinhibitor effects and contribute to the poor olaparib response observed in these patients (FIG. 9B).Example 3: Feature Importance as a Guide to Reinforce Model Performance
[0078] In this Example, a machine learning model was trained based on inputs similar to those reported above in Example 7, and included genetic data (e.g. DNA) and clinical embeddings.
[0079] Feature importance was determined using a method such as SHAP (SHapley Additive exPlanations) analysis, Gini Importance, Mean Decrease in Impurity (MDI), Permutation Importance (Mean Decrease in Accuracy), or LIME (Local Interpretable Model-agnostic Explanations)). Accordingly, a minimal subset of genes based on their variation across a population was defined to optimize impact on model performance (FIG. 10).
[0080] Refinement of features used to impact model performance was determined using the top 33 most impactful genomic features along with clinical details were selected for model usage, as follows: TNFAIP3 mutation, CTCF deletion, BCOR deletion, NFKBIA amplification, FGF3 amplification, BCL2 deletion, EGFR mutation, ERBB2 amplification, KRAS mutation and amplification, MCL1 amplification, NRAS mutation, ARID 1 A mutation, GNAS mutation, AURKB deletion, SDHA amplification, FLCN deletion, GATA3 deletion, BRAF mutation, CTNNB1 mutation, VHL mutation, MYC amplification, RICTOR amplification, PIK3CA mutation, NF1 mutation, CDKN2A mutation and deletion, RBI mutation, PTEN mutation, STK11 mutation, CDKN2B deletion, TP53 mutation (see FIG. 10).
[0081] Training data consistent with content listed in Example 1 was used in this Example and training and validation were performed as described. The 33 features used to impact model performance alongside a subset of specific clinical features captured in an embedding were selected from the combined training data to train a machine learning model to predict that a patient belongs to a class: “sensitive” or ”less-sensitive”. In other words, the output the machine learning model, which may be a random forest classifier, is a classification.
[0082] Validation of the machine learning model using feature importance was consistent as described above with the addition of 24 breast cancer patients. The outcome of the feature importance on rw-OS and rw-PFS are depicted in FIG. 11A and FIG. 11B along with thesensitivity analysis that demonstrates that these results are not confounded by the variables assessed. A GSEA was also performed, the results of which are shown in FIG. 12.Example 4: Feature Score and Random Forest Regressor
[0083] In this Example, a machine learning model was trained based on inputs similar to those reported above in Example 1 and based on predicted PARPi responses from Example 2. Prior to training, the top 10% of responding patients were identified. This population of patients was used to determine differential expression and isolate features used in the training. This resulted in two sets of features 1) combined score of 632 genes (Table 2) upregulated from each patient and 2) combined score of 269 genes (Table 3) downregulated from each patient. The two sets of feature scores were then divided resulting in a single combined score that was log2 transformed and then used to train a regressor (e.g., random forest regressor or ranger forest regressor) as described in Example 1. Testing and unique cohort validation of the model were conducted in a similar method as described above in Example 1. FIG. 13 depicts the distribution of the predicted AUC for 739 samples used in the test data set.
[0084] The data depicted in FIG. 14 demonstrate the olaparib rw-OS response predictions and the data in FIG. 15 demonstrate the gene set enrichment analysis (GSEA) that is associated with the sensitive (top of graph) and insensitive cohorts (bottom of graph) used for validation. Table 2. Upregulated Gene SetTable 3. Downregulated Gene Set
[0085] While this disclosure contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments.
[0086] Certain features that are described in this disclosure in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a sub-combination.
[0087] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0088] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.INCORPORATION BY REFERENCE
[0089] All references, articles, publications, patents, patent publications, and patent applications cited herein are incorporated by reference in their entireties for all purposes. However, mention of any reference, article, publication, patent, patent publication, and patent application cited herein is not, and should not be taken as an acknowledgment or any form of suggestion that they constitute valid prior art or form part of the common general knowledge in any country in the world.
Claims
CLAIMS1. A method for predicting an outcome of treatment with a PARP inhibitor for a subject having a tumor, the method comprising: obtaining clinical and molecular data from the subject; applying a machine learning model to the clinical and molecular data to calculate a sensitivity metric corresponding to a likelihood that the subject will respond to the PARP inhibitor; and predicting the treatment outcome in the subject by evaluating the sensitivity metric.
2. The method of claim 1, wherein the machine learning model is based on one or more of a radius-neighbor, neural network, a random forest classifier, elastic net regularization model, gradient-boosted tress-based method, and a random forest regressor.
3. The method of claim 2, wherein, when the machine learning model is a random forest classifier, the obtained clinical and molecular data is based on feature importance.
4. The method of claim 3, wherein the feature importance is determined by SHapley Additive exPlanations (SHAP) analysis, Gini Importance, Mean Decrease in Impurity (MDI), Permutation Importance (Mean Decrease in Accuracy), or Local Interpretable Model-agnostic Explanations) (LIME).
5. The method of claim 3, wherein the obtained clinical and molecular data identified as an important feature include TNFAIP3 mutation, CTCF deletion, BCOR deletion, NFKBIA amplification, FGF3 amplification, BCL2 deletion, EGFR mutation, ERBB2 amplification, KRAS mutation and amplification, MCL1 amplification, NRAS mutation, ARID 1 A mutation, GNAS mutation, AURKB deletion, SDHA amplification, FLCN deletion, GATA3 deletion, BRAF mutation, CTNNB1 mutation, VHL mutation, MYC amplification, RICTOR amplification, PIK3CA mutation, NF1 mutation, CDKN2A mutation and deletion, RBI mutation, PTEN mutation, STK11 mutation, CDKN2B deletion, TP53 mutation.
6. The method of claim 1, wherein the machine learning model comprises clinical and molecular information derived from subjects having the type of tumor and having been treated with the PARP inhibitor.
7. The method of claim 1, wherein the clinical and molecular data comprises genetic information derived from a pre-treatment subject tumor or human tumor model system.
8. The method of claim 1, wherein the clinical and molecular data comprises genetic information is derived from a tumor or human tumor model system that has failed one or more lines of therapy.
9. The method of claim 1, wherein the method may not comprise use of conventional predictive biomarkers for PARP inhibitors as determined by methods to identify HRD status or BRCA mutations.
10. The method of claim 1, wherein the PARP inhibitor is olaparib.
11. The method of claim 1, wherein the tumor is a solid breast cancer tumor.
12. The method of claim 1, wherein the tumor is a solid ovarian cancer tumor.
13. The method of claim 1, further comprising administering one or more doses of the PARP inhibitor to the subject when the subject is determined to be sensitive to the PARP inhibitor.
14. The method of claim 1, wherein the sensitivity metric is a score and evaluating the sensitivity metric comprises comparing the score to a median threshold.
15. The method of claim 1, wherein the sensitivity metric is a score and evaluating the sensitivity metric comprises comparing the score to an area under the dose response curve.
16. The method of claim 1, wherein the sensitivity metric is an area under the curve.
17. A method for predicting, for a subject having a tumor that tests positive for a BRCA 1 or BRC A2 mutation or a defect in the DNA homologous recombination repair (HRR) pathway, whether the subject will not respond to a PARP inhibitor, the method comprising: obtaining complex clinical and molecular data pertaining to the subject; applying a machine learning model to the clinical and molecular data to generate a composite biomarker for the subject, the composite biomarker comprising genes based on the type of tumor and a PARP inhibitor; calculating a probability that the subject will respond to the PARP inhibitor; and predicting the treatment outcome in the subject by comparing the calculated probability to a threshold probability for treatment responsiveness.
18. The method of claim 17, wherein the machine learning model comprises complex clinical and molecular information derived from subjects having the same type of tumor and having been treated with the same PARP inhibitor as is being contemplated for the subject.
19. The method of claim 17, wherein the molecular data comprises genetic information that is derived from a pre-treatment subject tumor or human tumor model system.
20. The method of claim 18, wherein the genetic information is derived from a tumor or human tumor model system that has failed one or more lines of therapy.
21. The method of claim 17, wherein the PARP inhibitor is selected from the group comprising olaparib, niraparib, rucaparib, talazoparib, saruparib, stenoparib, fuzuloparib, iniparib, pamiparib, mefuparib, saruparib, and venadaparib.
22. The method of claim 20, wherein the PARP inhibitor is olaparib.
23. The method of claim 20, wherein the PARP inhibitor is a next generation PARP inhibitor.
24. The method of claim 17, wherein the tumor is an ovarian cancer tumor, a breast cancer tumor, a non-small cell lung cancer tumor, a prostate cancer tumor, a fallopian tube tumor, or a primary peritoneal tumor.
25. The method of claim 24, wherein the tumor is an ovarian cancer tumor.
26. The method of claim 24, wherein the tumor is a breast cancer tumor.
Citation Information
Patent Citations
Method for determining sensitivity to PARP inhibitor or DNA damaging agent using non-functional transcriptome
US20230383363A1
Diagnostic test for predicting responsiveness to treatment with poly(ADP-ribose) polymerase (PARP) inhibitor
WO2011058367A2