Systems and methods for cancer treatment monitoring
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PREDICINE INC
- Filing Date
- 2023-05-18
- Publication Date
- 2026-05-26
AI Technical Summary
There is a lack of clinical biomarkers to identify patients with hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) metastatic breast cancer who may not respond to CDK4/6 inhibition (CDK4/6i) in combination with endocrine therapy (ET).
Perform genome-wide circulating tumor DNA (ctDNA) analysis to determine tumor mutational burden (bTMB) and copy number burden (bCNB) using comprehensive next-generation sequencing (NGS) to identify features associated with resistance to ET and CDK4/6i, including whole exome sequencing (WES) and low-pass whole genome sequencing (LP-WGS).
High bTMB and bCNB levels are associated with poor patient outcomes, predicting lack of clinical benefit and shorter progression-free survival, enabling early identification of patients requiring alternative treatment strategies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 343,749, filed May 19, 2022, which is hereby incorporated by reference in its entirety.
Background Art
[0002] Cancer is a leading cause of death worldwide. Detection of cancer in an individual can be important to provide treatment and improve the patient's outcome. Cancer can be caused by genetic aberrations that can lead to unregulated growth of cells. Detection of genetic aberrations can be important for cancer detection. Sequencing of nucleic acids in patient-derived samples can be used to detect genetic aberrations.
Summary of the Invention
[0003] CDK4 / 6 inhibition (CDK4 / 6i) in combination with endocrine therapy (ET) improves survival in patients with hormone receptor-positive (HR+) / human epidermal growth factor receptor 2-negative (HER2-) metastatic breast cancer (MBC). However, there is a lack of clinical biomarkers to identify patients who may not respond. The inventors performed genome-wide circulating tumor DNA (ctDNA) analysis to identify features associated with resistance to ET and CDK4 / 6i.
[0004] In one aspect, the present disclosure provides a method comprising: (a) obtaining or deriving a biological sample from a subject, wherein the subject has cancer, had cancer previously, or is suspected of having cancer; (b) assaying cell-free deoxyribonucleic acid (cfDNA) molecules obtained or derived from the biological sample, the assaying comprising sequencing at least a portion of the cfDNA molecules or derivatives thereof to generate a set of sequencing reads, the sequencing including at least one of whole exome sequencing (WES) and whole genome sequencing (WGS); and (c) determining at least one of a tumor mutational burden and a copy number burden of the subject, based at least in part on processing the set of sequencing reads.
[0005] In some embodiments, the biological sample is selected from the group consisting of a plasma sample, a serum sample, a buffy coat sample, a urine sample, a saliva sample, a tissue biopsy sample, a pleural fluid sample, a peritoneal fluid sample, an amniotic fluid sample, a cerebrospinal fluid sample, a lymph fluid sample, a sweat sample, a tear sample, a semen sample, derivatives thereof, and combinations thereof. In some embodiments, the biological sample includes a plasma sample. In some embodiments, the biological sample includes a urine sample.
[0006] In some embodiments, the biological sample is a single biological sample of the subject. In some embodiments, the biological sample is a plurality of biological samples of the subject.
[0007] In some embodiments, the biological sample is obtained or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, another blood collection tube, or a circulating tumor cell (CTC) collection tube.
[0008] In some embodiments, the method further comprises subjecting a biological sample to conditions sufficient to isolate, enrich, or extract cfDNA molecules.
[0009] In some embodiments, the method further comprises fractionating a whole blood sample of the subject to obtain cfDNA molecules.
[0010] In some embodiments, the sequencing further comprises WES. In some embodiments, the sequencing further comprises WGS. In some embodiments, WGS further comprises low-pass WGS. In some embodiments, the sequencing in (b) further comprises next-generation sequencing, low-pass sequencing, targeted sequencing, methylation-aware sequencing, bisulfite sequencing, or a combination thereof. In some embodiments, the sequencing further comprises methylation-aware sequencing or bisulfite sequencing.
[0011] In some embodiments, (b) further comprises amplifying at least a portion of the cfDNA molecules or derivatives thereof. In some embodiments, amplifying further comprises polymerase chain reaction (PCR). In some embodiments, amplifying further comprises isothermal amplification. In some embodiments, (b) further comprises using a microarray.
[0012] In some embodiments, the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, liver cancer, and combinations thereof.
[0013] In some embodiments, the cancer comprises breast cancer. In some embodiments, the breast cancer is metastatic breast cancer. In some embodiments, the breast cancer is hormone receptor positive (HR+) breast cancer. In some embodiments, the breast cancer is HER2-negative (HER2-) breast cancer. In some embodiments, the breast cancer is HR+ and HER2- breast cancer.
[0014] In some embodiments, the subject is asymptomatic with respect to cancer.
[0015] In some embodiments, processing in (c) further includes using a trained machine learning algorithm.
[0016] In some embodiments, the trained machine learning algorithm is trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 independent training samples.
[0017] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples related to the presence of cancer and a second set of independent training samples related to the absence of cancer.
[0018] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples related to the presence of cancer relapse or recurrence and a second set of independent training samples related to the absence of cancer relapse or recurrence.
[0019] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples related to the presence of cancer drug treatment or resistance to drug treatment, and a second set of independent training samples related to the absence of cancer drug treatment or resistance to drug treatment. In some embodiments, the trained machine learning algorithm further comprises an unsupervised machine learning algorithm. In some embodiments, the trained machine learning algorithm further comprises a supervised machine learning algorithm. In some embodiments, the supervised machine learning algorithm further comprises a deep learning algorithm, a neural network, or a random forest.
[0020] In some embodiments, (c) further comprises using the trained machine learning algorithm or another trained machine learning algorithm to process a set of clinical health data of the subject.
[0021] In some embodiments, the method further comprises determining a recurrence or relapse of the subject's cancer, at least in part, based on at least one of the subject's tumor mutation load and copy number load.
[0022] In some embodiments, the method further comprises at least partially determining a recurrence or relapse of the subject's cancer, at least in part, based on at least one of the subject's tumor mutation load and copy number load being at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%.
[0023] In some embodiments, the method further comprises determining, at least in part, a recurrence or relapse of the subject's cancer based at least in part on both the tumor mutation burden and the copy number burden of the subject being at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%.
[0024] In some embodiments, the method further comprises determining a recurrence or relapse with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% accuracy.
[0025] In some embodiments, the method further comprises determining a recurrence or relapse with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% sensitivity.
[0026] In some embodiments, the method further comprises determining a recurrence or relapse with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% specificity.
[0027] In some embodiments, the method further comprises determining a recurrence or relapse with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% positive predictive value.
[0028] In some embodiments, the method further includes determining relapse or recurrence at a negative concordance rate of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0029] In some embodiments, the method further includes determining cancer resistance to a drug treatment based at least in part on at least one of the subject's tumor mutation load and copy number load.
[0030] In some embodiments, the method further includes at least partially determining cancer resistance to a drug treatment based at least in part on at least one of the subject's tumor mutation load and copy number load being at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%.
[0031] In some embodiments, the method further includes at least partially determining cancer resistance to a drug treatment based at least in part on both the subject's tumor mutation load and copy number load being at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%.
[0032] In some embodiments, the method further includes determining resistance with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0033] In some embodiments, the method further includes determining resistance with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0034] In some embodiments, the method further includes determining resistance with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0035] In some embodiments, the method further includes determining resistance with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0036] In some embodiments, the method further includes determining resistance with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0037] In some embodiments, the method further includes determining the prognosis of the subject's cancer based at least in part on at least one of the subject's tumor mutation burden and copy number burden. In some embodiments, the method further includes determining the prognosis of the subject's cancer based at least in part on one of the subject's tumor mutation burden and copy number burden. In some embodiments, the prognosis includes the likelihood of progression-free survival, the length of progression-free survival, the likelihood of overall survival, the length of overall survival, or combinations thereof.
[0038] In some embodiments, (a) further comprises the step of obtaining or deriving a biological sample from a subject (i) before the subject undergoes a cancer clinical intervention, (ii) while the subject is undergoing a cancer clinical intervention, (iii) after the subject has undergone a cancer clinical intervention, or combinations thereof. In some embodiments, the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen blockade therapy, and combinations thereof.
[0039] In some embodiments, the method further comprises determining a clinical intervention for the subject based at least in part on at least one of the subject's tumor mutation burden and copy number burden. In some embodiments, the method further comprises determining a clinical intervention for the subject based at least in part on both the subject's tumor mutation burden and copy number burden. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions.
[0040] In some embodiments, the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, endocrine therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the clinical intervention comprises a CDK4 / 6 inhibitor. In some embodiments, the CDK4 / 6 inhibitor comprises palbociclib. In some embodiments, the clinical intervention comprises endocrine therapy. In some embodiments, the endocrine therapy comprises letrozole or fulvestrant. In some embodiments, the clinical intervention comprises endocrine therapy and a CDK4 / 6 inhibitor.
[0041] In some embodiments, the method further comprises administering a clinical intervention to the subject.
[0042] In some embodiments, the set of sequencing reads includes a quantitative measure of a set of cancer-associated genomic loci. In some embodiments, the set of cancer-associated genomic loci includes one or more members selected from the group consisting of the genes listed in Table 3, the genes listed in Table 4, the genes listed in Table 6, and the genes listed in Table 7. In some embodiments, the set of cancer-associated genomic loci includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 3, the genes listed in Table 4, the genes listed in Table 6, and the genes listed in Table 7. In some embodiments, the set of cancer-associated genomic loci includes one or more members selected from the group consisting of the genes listed in Table 3. In some embodiments, the set of cancer-associated genomic loci includes one or more members selected from the group consisting of the genes listed in Table 4. In some embodiments, the set of cancer-associated genomic loci includes one or more members selected from the group consisting of the genes listed in Table 6. In some embodiments, the set of cancer-associated genomic loci includes one or more members selected from the group consisting of the genes listed in Table 7.
[0043] In some embodiments, (b) further comprises using a nucleic acid primer or probe configured to selectively enrich a biological sample for DNA molecules corresponding to a set of genomic loci. In some embodiments, the nucleic acid primer or probe has sequence complementarity with at least a portion of the nucleic acid sequences of the set of genomic loci. In some embodiments, the nucleic acid primer or probe comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different nucleic acid primers or probes.
[0044] In some embodiments, the method further comprises monitoring at least one of the tumor mutation burden and the copy number burden of the subject, the monitoring comprising evaluating at least one of the tumor mutation burden and the copy number burden of the subject at each of a plurality of time points. In some embodiments, a difference in the evaluation of at least one of the tumor mutation burden and the copy number burden of the subject at the plurality of time points indicates one or more clinical indications selected from the group consisting of (i) diagnosis of cancer, (ii) prognosis of cancer, and (iii) effectiveness or ineffectiveness of a clinical intervention for treating the subject's cancer.
[0045] In some embodiments, the processing in (c) further comprises detecting tumor-related changes selected from the group consisting of copy number alterations (CNA), copy number losses (CNL), single nucleotide variants (SNV), insertions or deletions (indels), and rearrangements.
[0046] In some embodiments, the method further comprises filtering at least a subset of the set of sequencing reads based on a quality score.
[0047] In some embodiments, the method further comprises performing error correction on the set of sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules.
[0048] In some embodiments, the method further comprises performing at least one of single-stranded consensus calling and double-stranded consensus calling on the set of sequencing reads, thereby suppressing sequencing and PCR errors in the set of sequencing reads.
[0049] In some embodiments, the method further comprises determining the mutant allele frequency of the set of somatic mutations.
[0050] In another aspect, the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine-executable code that, when executed by the one or more computer processors, implements a method comprising: (a) obtaining or deriving a biological sample from a subject, where the subject has cancer, had cancer previously, or is suspected of having cancer; (b) assaying cell-free deoxyribonucleic acid (cfDNA) molecules obtained from or derived from the biological sample, the assaying comprising sequencing at least a portion of the cfDNA molecules or derivatives thereof to generate a set of sequencing reads, the sequencing comprising at least one of whole exome sequencing (WES) and whole genome sequencing (WGS); and (c) determining at least one of a tumor mutation burden and a copy number burden of the subject, based at least in part on processing the set of sequencing reads.
[0051] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method comprising: (a) obtaining or deriving a biological sample from a subject, where the subject has cancer, had cancer previously, or is suspected of having cancer; (b) assaying cell-free deoxyribonucleic acid (cfDNA) molecules obtained from or derived from the biological sample, the assaying comprising sequencing at least a portion of the cfDNA molecules or derivatives thereof to generate a set of sequencing reads, the sequencing comprising at least one of whole exome sequencing (WES) and whole genome sequencing (WGS); and (c) determining at least one of a tumor mutation burden and a copy number burden of the subject, based at least in part on processing the set of sequencing reads.
[0052] Also provided herein are systems and methods for detecting the presence or absence of cancer in a subject. The systems and methods provided herein include assaying polynucleotides to identify biomarkers of cancer in a subject. Detection of the type of cancer or specific biomarkers for a given cancer can enable an effective treatment to be provided to an individual and can result in an improved outcome. For multiple types of cancer, specific biomarkers indicative of a specific cancer type (or subtype) can be used to identify the prognosis of an individual suffering from the cancer. To provide accurate detection and prognosis of cancer, multiple analytes can be examined. By analyzing more analytes (and sets of biomarkers from the analytes), the detection of cancer (or cancer parameters) can be improved, enabling a recommendation of an effective treatment and making the prognosis more accurate.
[0053] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, the method comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules and cell-free ribonucleic acid (cfRNA) molecules from a biological sample obtained from or derived from the subject to detect a first set of biomarkers from the cfDNA molecules and a second set of biomarkers from the cfRNA molecules; and (b) computationally processing the first set of biomarkers and the second set of biomarkers to detect the presence or absence of cancer in the subject.
[0054] In some embodiments, the biological sample is selected from the group consisting of a cell-free deoxyribonucleic acid (cfDNA) sample, a cell-free ribonucleic acid (cfRNA) sample, a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, an erythrocyte sample, a urine sample, a saliva sample, a tissue biopsy, a pleural effusion sample, an ascites sample, an amniotic fluid sample, a cerebrospinal fluid sample, a lymph fluid sample, a sweat sample, a tear sample, a semen sample, or any derivative thereof, and any combination thereof. In some embodiments, the biological sample comprises a plasma sample. In some embodiments, the biological sample comprises a urine sample.
[0055] In some embodiments, the cfDNA molecules and cfRNA molecules are obtained from or derived from a single biological sample of the subject. In some embodiments, the cfDNA molecules and cfRNA molecules are obtained from or derived from different biological samples of the subject.
[0056] In some embodiments, the biological materials are obtained from or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tubes, and a CTC collection tube.
[0057] In some embodiments, (a) comprises subjecting the biological sample to conditions sufficient to isolate, concentrate, or extract a set of cfDNA molecules and cfRNA molecules.
[0058] In some embodiments, the method further comprises fractionating a whole blood sample of the subject to obtain cfDNA molecules and cfRNA molecules.
[0059] In some embodiments, at least one of the cfDNA molecules and cfRNA molecules is assayed using nucleic acid sequencing to generate nucleic acid sequencing reads. In some embodiments, the cfDNA molecules are assayed using DNA sequencing. In some embodiments, the DNA sequencing is selected from the group consisting of next-generation sequencing, whole-genome sequencing, low-pass sequencing, targeted sequencing, methylation-aware sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing, and combinations thereof. In some embodiments, the DNA sequencing comprises low-pass whole-genome sequencing. In some embodiments, the DNA sequencing comprises whole-exome sequencing. In some embodiments, the DNA sequencing comprises methylation-aware sequencing, enzymatic methylation sequencing, or bisulfite methylation sequencing.
[0060] In some embodiments, cfRNA molecules are assayed using RNA sequencing. In some embodiments, the RNA sequencing is selected from the group consisting of next-generation sequencing, transcriptome sequencing, mRNA-seq, totalRNA-seq, smallRNA-seq, exosome sequencing, and combinations thereof. In some embodiments, the RNA sequencing includes reverse transcribing the cfRNA molecules into complementary DNA (cDNA) molecules and performing DNA sequencing on the cDNA molecules.
[0061] In some embodiments, nucleic acid sequencing includes nucleic acid amplification. In some embodiments, it includes polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, nucleic acid sequencing includes the use of substantially simultaneous reverse transcription (RT) and polymerase chain reaction (PCR).
[0062] In some embodiments, at least one of the cfDNA molecules and cfRNA molecules is assayed using a polymerase chain reaction (PCR) assay, a microarray, or isothermal amplification.
[0063] In some embodiments, the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. In some embodiments, the cancer includes prostate cancer. In some embodiments, the prostate cancer is selected from the group consisting of hormone sensitive prostate cancer (HSPC), castrate-resistant prostate cancer (CRPC), metastatic prostate cancer, and combinations thereof. In some embodiments, the subject is asymptomatic for the cancer. In some embodiments, the cancer includes breast cancer. In some embodiments, the cancer includes bladder cancer.
[0064] In some embodiments, (b) includes processing a first set of biomarkers and a second set of biomarkers using a trained algorithm. In some embodiments, the trained algorithm is trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 independent training samples related to the presence or absence of cancer. In some embodiments, the trained algorithm is trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 independent training samples related to cancer recurrence. In some embodiments, the trained algorithm is trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 independent training samples related to drug treatment or resistance to drug treatment.
[0065] In some embodiments, the trained algorithm is trained using a first set of independent training samples associated with the presence of cancer and a second set of independent training samples associated with the absence of cancer. In some embodiments, the trained algorithm is trained using a first set of independent training samples associated with the presence of cancer and a second set of independent training samples associated with cancer recurrence. In some embodiments, the trained algorithm is trained using a first set of independent training samples associated with the presence of cancer and a second set of independent training samples associated with drug treatment resistance to drug treatment.
[0066] In some embodiments, the method further includes processing a set of clinical health data of a subject using the trained algorithm or another trained algorithm to determine the presence or absence of cancer. In some embodiments, the method further includes processing a set of clinical health data of a subject using the trained algorithm or another trained algorithm to determine cancer recurrence. In some embodiments, the method further includes processing a set of clinical health data of a subject using the trained algorithm or another trained algorithm to determine drug treatment or resistance to drug treatment.
[0067] In some embodiments, the trained algorithm includes an unsupervised machine learning algorithm. In some embodiments, the trained machine learning algorithm includes a supervised algorithm. In some embodiments, the supervised machine learning algorithm includes a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.
[0068] In some embodiments, (b) includes detecting the presence or absence of cancer in a subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0069] In some embodiments, (b) comprises detecting the presence or absence of cancer in a subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0070] In some embodiments, (b) comprises detecting the presence or absence of cancer in a subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0071] In some embodiments, (b) comprises detecting the presence or absence of cancer in a subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0072] In some embodiments, (b) comprises detecting the presence or absence of cancer in a subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0073] In some embodiments, the biological sample is obtained from or derived from the subject before the subject undergoes cancer treatment. In some embodiments, the biological sample is obtained from or derived from the subject during cancer treatment. In some embodiments, the biological sample is obtained from or derived from the subject after the subject has undergone cancer treatment.
[0074] In some embodiments, the treatment is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.
[0075] In some embodiments, the method further comprises the step of identifying a clinical intervention for the subject, at least in part based on the presence or absence of the detected cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises the step of administering the clinical intervention to the subject.
[0076] In some embodiments, the first set of biomarkers comprises a quantitative measure of the first set of cancer-related genomic loci. In some embodiments, the first set of cancer-related genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 1. In some embodiments, the first set of cancer-related genomic loci comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 1. In some embodiments, the first set of cancer-related genomic loci comprises PTEN, TP53, or RB1. In some embodiments, the first set of cancer-related genomic loci comprises PTEN, TP53, and RB1. In some embodiments, the first set of cancer-related genomic loci comprises PTEN. In some embodiments, the first set of cancer-related genomic loci comprises FGFR3 or ERBB2.
[0077] In some embodiments, the second set of biomarkers includes a quantitative measure of the second set of cancer-related genomic loci. In some embodiments, the second set of cancer-related genomic loci includes two or more members selected from the group consisting of the genes listed in Table 2. In some embodiments, the second set of cancer-related genomic loci includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.
[0078] In some embodiments, the method further includes using a probe configured to selectively enrich a biological sample for nucleic acid molecules corresponding to a set of genomic loci. In some embodiments, the probe is a nucleic acid primer. In some embodiments, the probe has sequence complementarity with at least a portion of the nucleic acid sequence of the set of genomic loci. In some embodiments, the probe includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.
[0079] In some embodiments, the method further includes determining the likelihood of determining the presence or absence of cancer in a subject.
[0080] In some embodiments, the method further includes monitoring the presence or absence of cancer in a subject, the monitoring including evaluating the presence or absence of cancer in the subject at each of a plurality of time points.
[0081] In some embodiments, the differences in the assessment of the presence or absence of cancer in a subject between multiple time points indicate one or more clinical indications selected from the group consisting of (i) cancer diagnosis, (ii) cancer prognosis, and (iii) the effectiveness or ineffectiveness of the treatment process for treating the subject's cancer. In some embodiments, the prognosis includes the expected progression-free survival (PFS) or overall survival (OS).
[0082] In some embodiments, the method further comprises assaying germline DNA (gDNA) molecules obtained from or derived from the subject to detect a third set of biomarkers, and computer-processing the third set of biomarkers to detect the presence or absence of cancer in the subject.
[0083] In some embodiments, the first set of biomarkers from cfDNA molecules includes tumor-related changes selected from the group consisting of copy number alterations (CNA), copy number losses (CNL), loss of heterozygosity (LOH), single nucleotide variants (SNV), insertions or deletions (indels), rearrangements, and epigenetic changes such as methylation. In some embodiments, the first set of biomarkers from cfDNA molecules includes copy number variations. In some embodiments, the first set of biomarkers from cfDNA molecules includes copy number losses. In some embodiments, the first set of biomarkers from cfDNA molecules includes single nucleotide variants.
[0084] In some embodiments, the second set of biomarkers from cfRNA molecules includes tumor-related changes selected from the group consisting of alternative splicing variants, fusions, single nucleotide variants (SNV), and insertions or deletions (indels).
[0085] In some embodiments, the method further includes filtering at least a subset of the nucleic acids of the sequencing reads based on a quality score.
[0086] In some embodiments, the method further includes performing error correction on the nucleic acids of the sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules and the cfRNA molecules.
[0087] In some embodiments, the method further includes performing at least one of single-stranded consensus scoring and double-stranded consensus scoring on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.
[0088] In some embodiments, the method further includes determining the mutant allele frequency of a set of somatic mutations among a first set of biomarkers. In some embodiments, the method further includes determining a blood copy number burden based on copy number changes or copy number losses of the first set of biomarkers.
[0089] In some embodiments, the method further includes determining a circulating tumor DNA (ctDNA) fraction of the cancer of the subject, at least partially based on a set of mutant allele frequencies.
[0090] In some embodiments, the method further includes determining a plasma tumor mutational burden (pTMB) of the cancer of the subject, at least partially based on a set of mutant allele frequencies.
[0091] In some embodiments, the method further includes determining a plasma tumor mutational burden (pTMB) of the cancer of the subject, at least partially based on a set of mutant allele frequencies that includes microsatellites.
[0092] In some embodiments, the method further includes determining an abnormality score of the subject's cancer, at least in part based on a set of mutant allele frequencies.
[0093] In some embodiments, the method further includes determining a methylation related score of the subject's cancer, at least in part based on a set of mutant allele frequencies.
[0094] In one aspect, the present disclosure provides a method for detecting the presence or absence of prostate cancer in a subject, the method comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules and germline DNA (gDNA) molecules obtained from or derived from a biological sample from the subject to detect a first set of biomarkers from cfDNA molecules and a second set of biomarkers from gRNA molecules, wherein at least one of the first set of biomarkers and the second set of biomarkers comprises an androgen receptor (AR) alteration; and (b) computationally processing the first set of biomarkers and the second set of biomarkers to detect the presence or absence of prostate cancer in the subject.
[0095] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the above methods or the methods described elsewhere herein.
[0096] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto. The computer memory comprises machine-executable code that, when executed by the one or more computer processors, implements any of the above methods or the methods described elsewhere herein.
[0097] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which only shows and describes exemplary embodiments of the present disclosure. As will be understood, the present disclosure is capable of other embodiments and different embodiments, and various details thereof can be modified in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.
[0098] Incorporation by reference All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference into this specification to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the present disclosure contained in this specification, this specification is intended to supersede and / or prevail over any such conflicting material.
Brief Description of the Drawings
[0099] The novel features of the present invention are particularly described in the appended claims. A better understanding of the features and advantages of the present invention can be obtained by referring to the following detailed description that describes exemplary embodiments, and the principles of the present invention and the accompanying drawings (further. The terms "Figure" and "FIG" are used herein).
Fig. 1A
Fig. 1B
Fig. 1C
Fig. 2A
Fig. 2B
Fig. 2C
Fig. 2D
Fig. 2E
Fig. 2F
Fig. 2G
Fig. 2H
Fig. 3
Fig. 4A
Fig. 4B
Fig. 4C
Fig. 5A
Fig. 5B
Fig. 6A
Fig. 6B
Fig. 6C
Fig. 7A
Fig. 7B
Fig. 7C
Fig. 8
Fig. 9
Fig. 10A
Fig. 10B
Fig. 11
Fig. 12A
Fig. 12B
Fig. 13
Fig. 14A
Fig. 14B
Fig. 15
Fig. 16A
Fig. 16B
Fig. 16C
Fig. 16D
Fig. 17
[0100] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.
[0101] Also provided herein are systems and methods for detecting the presence or absence of cancer in a subject. The systems and methods provided herein include assaying polynucleotides to identify cancer biomarkers in a subject. The biomarkers can be processed to identify the presence or absence of cancer. The methods described herein can process multiple types of analytes to determine the presence or absence of cancer. The multiple types of analytes can include DNA or RNA, such as cfDNA or cfRNA. The multiple analytes can be cfDNA, germline DNA, and cfRNA. By analyzing multiple different analytes, the method can enable improved detection or determination of prognosis compared to methods performed on fewer analytes or only one of many different analytes.
[0102] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, the method comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules and cell-free ribonucleic acid (cfRNA) molecules obtained from or derived from a subject in a biological sample to detect a first set of biomarkers from the cfDNA molecules and a second set of biomarkers from the cfRNA molecules; and (b) computer-processing the first set of biomarkers and the second set of biomarkers to detect the presence or absence of cancer in the subject.
[0103] The subject can be a subject suspected of having cancer. The cancer can be specific to or derived from an organ or other region of the subject. For example, the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and combinations thereof. The cancer can be hormone-sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), metastatic prostate cancer, and combinations thereof. The cancer can include biomarkers specific to a particular cancer. The specific biomarker can indicate the presence of a particular cancer. For example, the biomarker can indicate the presence of castration-resistant prostate cancer. Identification of the presence of a cancer type can enable the determination of treatment options or recommendations.
[0104] In some cases, the subject can be asymptomatic with respect to the cancer. For example, the cancer may show no symptoms and the subject may be unaware of the presence of the cancer. The methods described herein can enable identification at an earlier stage than otherwise. Identification of the presence of cancer at an early stage can enable treatment options or recommendations to be determined at an early stage and can enable the subject to have an improved prognosis.
[0105] A biological sample can contain nucleic acids. The biological sample is a cell-free deoxyribonucleic acid (cfDNA) sample or a cell-free ribonucleic acid (cfRNA) sample. The biological sample can contain genomic DNA or germline DNA (gDNA). The nucleic acid can be DNA (e.g., double-stranded DNA, single-stranded DNA, single-stranded DNA hairpin, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), RNA (e.g., cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA, tiRNA, snRNA), or a DNA / RNA hybrid. The biological sample can be derived from or contain a biological fluid. For example, the biological sample can be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample, or other body fluid samples. The biological sample can include or be a pleural effusion sample, an ascites sample, an amniotic fluid sample, a cerebrospinal fluid sample, a lymph fluid sample, a sweat sample, a tear sample, a semen sample, or any combination of biological fluids. In some cases, the sample can contain RNA and DNA. For example, the sample can contain cfDNA and cfRNA, and the cfDNA and cfRNA can be analyzed by the methods described elsewhere in this specification.
[0106] A biological sample can be collected, obtained, or derived from a subject using a collection tube. The collection tube can be an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube and a CTC collection tube, or other blood collection tubes. The collection tube can contain additional reagents for stabilizing nucleic acid molecules or blood cells. The collection tube can enable the stabilization of nucleic acids or blood cells to minimize the degradation of the biological sample before the assay. The additional reagents can include buffer salts or chelating agents.
[0107] A biological sample can be obtained from or derived from a subject at various time points. A biological sample can be obtained from or derived from a subject before the subject undergoes cancer treatment. A biological sample can be obtained from or derived from a subject during cancer treatment. A biological sample can be obtained from or derived from a subject after cancer therapy. A biological sample can be taken at time points of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or over time points. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.
[0108] In various aspects described herein, a clinical intervention or treatment can be identified based at least in part on the identification of the presence of cancer or the presence of cancer parameters. The clinical intervention can be a plurality of clinical interventions. The clinical intervention can be selected from a plurality of clinical interventions. The clinical intervention can be surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, or a combination thereof. In some cases, the clinical intervention can be administered to a subject. After administration of the clinical intervention, a sample can be obtained from or be derived from the subject to monitor cancer or cancer parameters. Thus, the methods and systems disclosed herein can be repeatedly performed such that monitoring of cancer can be carried out. Further, by repeatedly performing the method or system, the treatment or clinical intervention can be updated based on the results of the method. Monitoring of cancer can also include an assessment and a difference in the assessment from a previously generated assessment. A difference in the assessment of cancer in a subject between multiple time points (or samples) can indicate one or more clinical applications such as a cancer diagnosis, a cancer prognosis, or the effectiveness or ineffectiveness of a treatment course for treating the subject's cancer. The prognosis can include a predicted progression-free survival (PFS), overall survival (OS), or other metrics related to cancer severity or survival rate.
[0109] The biological sample can be subjected to further reactions or conditions prior to the assay. For example, the biological sample can be subjected to conditions sufficient to isolate, concentrate, or extract nucleic acids such as cfDNA molecules or cfRNA molecules.
[0110] The methods disclosed herein may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reaction may include contacting the sample with one or more beads or bead sets. The enrichment reaction may include one or more hybridization reactions. For example, the enrichment reaction may include contacting the sample with one or more capture probes or bait molecules that hybridize to nucleic acid molecules of a biological sample. The enrichment reaction may include differential amplification of a set of nucleic acid molecules. The enrichment reaction may enrich a plurality of loci, or sequences corresponding to loci. For example, the enrichment reaction may enrich sequences corresponding to the genes of Table 1 or Table 2. The enrichment reaction may include the use of primers or probes that may be complementary to the sequence of the sequence to be enriched (or upstream or downstream sequences). For example, a capture probe may include sequence complementarity to a set of genomic loci and may enable enrichment of the genomic loci. The enrichment reaction may include a plurality of probes or primers. The plurality of probes may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.
[0111] The methods disclosed herein may include performing one or more isolation or purification reactions on one or more nucleic acid molecules in a sample. The isolation or purification reaction may include contacting the sample with one or more beads or bead sets. The isolation or purification reaction may include one or more hybridization reactions, concentration reactions, amplification reactions, sequencing reactions, or combinations thereof. The isolation or purification reaction may include the use of one or more separators. The one or more separators may comprise a magnetic separator. The isolation or purification reaction may include separating bead bound nucleic acid molecules from bead free nucleic acid molecules. The isolation or purification reaction may include separating nucleic acid molecules hybridized to capture probes from nucleic acid molecules that do not contain capture probes. The isolation reaction may include removing or separating a group of nucleic acid molecules from another group of nucleic acids.
[0112] The methods disclosed herein may include conduction extraction reactions on one or more nucleic acids in a biological sample. The extraction reaction may lyse cells or disrupt nucleic acid interactions with cells such that the nucleic acids can be isolated, purified, concentrated, or otherwise made available for other reactions.
[0113] The methods disclosed herein may include amplification or extension reactions. The amplification reaction may include polymerase chain reaction. The amplification reaction may include PCR-based amplification, non-PCR-based amplification, or combinations thereof. One or more PCR-based amplifications may include PCR, qPCR, nested PCR, linear amplification, or combinations thereof. One or more non-PCR-based amplifications may include multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, circle-to-circle amplification, or combinations thereof. The amplification reaction may include isothermal amplification.
[0114] The methods disclosed herein may include a barcoding reaction. The barcoding reaction may include the addition of a barcode or tag to a nucleic acid. The barcode may be a molecular barcode or a sample barcode. For example, the barcoded nucleic acid may include a barcode sequence that may be a degenerate n-mer. The sequence may be generated randomly or may be generated to synthesize a specific barcode sequence. The barcoded nucleic acid may be added to a sample to label nucleic acid molecules in the sample. The barcode may be specific to the sample. For example, multiple barcoded nucleic acids may be added to a sample that has the same barcode sequence. Upon barcoding of the nucleic acid, those derived from the same sample may have the same barcode sequence, enabling the nucleic acid to be identified as belonging to a particular or given sample. Molecular barcodes may be used such that each molecule (or molecules) in the same volume has a different molecular barcode. This barcode may be subjected to amplification such that all amplicons derived from the molecule have the same barcode. In this way, molecules derived from the same molecule may be identified. Sequence reads may be processed based on the barcode sequence. For example, the processing may reduce errors or may enable molecules to be tracked. The barcode sequence may be added or otherwise appended or incorporated into the sequence by various reactions, such as amplification, extension, or ligation reactions, and may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be overhang or blunt-end ligation, and the barcode may include complementarity to the nucleic acid being barcoded. This complementarity may be a sequence derived from a sample from the subject or may be a constant sequence generated via a reaction performed on the nucleic acid in the sample.
[0115] In some cases, a biological sample may contain multiple components. For example, the biological sample can be a whole blood sample. The biological sample can be subjected to reactions such as separating or fractionating the biological sample. For example, the whole blood sample can be fractionated to obtain cell-free nucleic acids. The whole blood sample can be fractionated using centrifugation so that blood cells can be separated from plasma (which may contain cell-free nucleic acids). The sample can be subjected to multiple separations or fractionations.
[0116] In various aspects described throughout this disclosure, nucleic acids can be subjected to a sequencing reaction. The sequencing reaction can be used for DNA, RNA, or other nucleic acid molecules. Examples of sequencing reactions that can be used include capillary sequencing, next-generation sequencing, Sanger sequencing, sequencing by synthesis, single molecule nanopore sequencing, sequencing by ligation, sequencing by hybridization, nanopore current-limited sequencing, or combinations thereof. Sequencing by synthesis can include reversible terminator sequencing, processive single molecule sequencing, sequential nucleotide flow sequencing, or combinations thereof. Sequential nucleotide flow sequencing can include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or combinations thereof. The sequencing reaction can include whole genome sequencing, whole exome sequencing, low-pass whole genome sequencing, targeted sequencing, methylation-aware sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing. The sequencing reaction can be transcriptome sequencing, mRNA-seq, totalRNA-seq, smallRNA-seq, exosome sequencing, or combinations thereof. Combinations of sequencing reactions can be used in the methods described elsewhere in this specification. For example, a sample can be subjected to whole genome sequencing and whole transcriptome sequencing. Since the sample can contain multiple types of nucleic acids (e.g., RNA and DNA), sequencing reactions specific to DNA or RNA can be used to obtain sequence reads related to the nucleic acid type.
[0117] Nucleic acid sequencing can generate sequencing read data. The sequencing reads can be processed to generate improved quality data. The sequencing reads can be generated using quality scores. The quality scores can indicate the accuracy of the sequence read, or a level or signal above a noise threshold, for a given base call. The quality scores can be used to filter the sequencing reads. For example, sequencing reads that do not meet a specific quality score threshold can be removed. The sequencing reads can be processed to generate a consensus sequence or a consensus base call. A given nucleic acid (or nucleic acid fragment) can be sequenced, and errors in the sequence can be generated due to reactions before or during sequencing. For example, amplification or PCR can generate errors in the amplicon such that the sequence is not identical to the parental sequence. Error correction can be performed using sample barcodes or molecular barcodes. Error correction can include identifying sequence reads that do not match other sequences from the same sample or the same original parental molecule. The use of barcodes can enable the identification of the same parent or sample. Further, the sequence reads can be processed by making single-stranded consensus calls or double-stranded consensus calls, thereby reducing or suppressing errors.
[0118] The methods disclosed herein can include determining an allele frequency or other cancer-related metric. The methods can include mutant allele frequencies of a set of somatic mutations among a set of biomarkers. The mutant allele frequencies can be used to determine the circulating tumor DNA (ctDNA) fraction of a subject's cancer. The plasma tumor mutation burden (pTMB) of a subject's cancer can be determined based at least in part on a set of mutant allele frequencies. Detection of microsatellite instability can also be used to determine the presence or absence of cancer or a cancer metric. The methylation state can be determined using the methods described herein and can be used to identify the presence of cancer or a cancer parameter.
[0119] In various aspects, a set of biomarkers is processed and data corresponding to the biomarkers is generated. The set of biomarkers can include quantitative measures from a set of cancer-related genomic loci. The cancer-related genomic loci can correspond to a set of genes. The cancer-related genomic loci can include one or more genes selected from Table 1. In some cases, the set of cancer-related genomic loci is selected from the group consisting of the genes listed in Table 1 and includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members. The cancer-related genomic loci can include one or more genes selected from Table 2. In some cases, the set of cancer-related genomic loci is selected from the group consisting of the genes listed in Table 2 and includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members.
[0120]
Table 1
[0121]
Table 2
[0122] A set of biomarkers can correspond to genetic abnormalities at loci. The genetic abnormalities can be tumor-related changes. The genetic abnormalities can be copy number alterations (CNA), copy number losses (CNL), single nucleotide variants (SNV), insertions or deletions (indels), and rearrangements. The set of biomarkers can be identified in various nucleic acid types. For example, tumor-related changes can be identified in cfDNA or cfRNA. Tumor-related changes can include allelic expression or changes in gene expression. The methods and systems disclosed herein can enable gene expression profiling and identification of changes to gene expression levels.
[0123] In various aspects, the method can include identifying the presence of cancer or a cancer parameter. The method can include determining the probability or likelihood that cancer or a cancer parameter is present. For example, instead of a binary output indicating presence or absence, an output indicating the probability that a subject has cancer can be generated. This probability can be determined based on algorithms described elsewhere herein. Similarly, the probability or likelihood of response to a particular treatment, or the probability of recurrence, can be output.
[0124] Increased cfRNA transcript expression of drug resistance-related gene changes or splicing variants can serve as predictive biomarkers for identifying response or resistance to treatment. Specifically, in the case of prostate cancer, increased cfRNA transcript expression of drug resistance-related AR variants such as W742C / L and F877L, or splicing variants such as AR-V7 or AR-V9, can serve as predictive biomarkers for identifying response or resistance to anti-androgen therapy.
[0125] Compared to the use of cfDNA, blood ctRNA-based variant detection (including fusions) can be used more effectively to identify known and novel variants, particularly fusions, in cancer. For example, blood cfRNA-based detection of TMPRSS2-ERG provides higher detection sensitivity in prostate cancer.
[0126] An increase in the ratio of blood-based cancer variants to urine-based cancer variants serves as a prognostic biomarker in GU cancers, indicates disease aggressiveness, and can guide clinical treatment decisions. Specifically, in the case of muscle-invasive bladder cancer (MIBC), increased levels of blood-based cancer variants relative to urine-based cancer variants can serve as a prognostic biomarker in patients with MIBC and can provide evidence for clinical decision-making. These cancer variants can include, among others, ctDNA, cfRNA, microRNA, and methylation.
[0127] Together with cfDNA-based variant detection by genomics and epigenomics, cfRNA and / or microRNA can also be used, alone or in combination with genomic and epigenomic biomarkers, for minimal residual disease (MRD) detection, treatment monitoring, and early cancer detection.
[0128] In various aspects, a set of biomarkers is processed using an algorithm. The algorithm can be a trained algorithm. A trained algorithm can use the set of biomarkers as input and generate an output regarding the presence or absence of cancer. The output can be specific to the type or subtype of cancer. For example, the output can indicate the presence of castration-resistant prostate cancer.
[0129] The trained algorithm can be trained on a plurality of samples. For example, the trained algorithm can be trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000 or more independent training samples. The trained algorithm can be trained using 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000 or fewer independent training samples. The training samples can be related to the presence or absence of cancer. The training samples can be related to cancer recurrence. The training samples can be related to cancer that is resistant to a particular drug or treatment. An individual training sample can be positive for a particular cancer. An individual training sample can be negative for a particular cancer. By using the training samples, the trained algorithm may be able to detect cancer, determine the probability of cancer recurrence or relapse, or determine whether a cancer including a set of biomarkers is resistant to treatment. The training samples can be associated with additional clinical health data of the subject. For example, the additional clinical health data can include the subject's gender, weight, height, or levels of metabolites or antibodies.Additional clinical health data may include indicators of other diseases, disorders, or disease states.
[0130] A trained algorithm can be trained using multiple sets of training samples. The sets can include training samples as described elsewhere in this specification. For example, the training can be performed using a first set of independent training samples associated with the presence of cancer and a second set of independent training samples associated with the absence of cancer. Similarly, the first set can be associated with relapse, and the second sample can be associated with the absence of relapse.
[0131] A trained algorithm can also process additional clinical health data of a subject. For example, the additional clinical health data can include the subject's gender, weight, height, or levels of metabolites or antibodies. The additional clinical health data can include indicators of other diseases, disorders, or disease states that the subject may have. By using the additional clinical health data in conjunction with biomarkers, a trained algorithm can output the presence or absence of cancer, the probability of recurrence, or resistance to drug treatment, which can be different from the output of an algorithm that does not process additional clinical health.
[0132] A trained algorithm can be an unsupervised machine learning algorithm. For example, an unsupervised machine learning algorithm can utilize cluster analysis to identify attributes of interest. A trained algorithm can be a supervised machine learning algorithm. For example, the algorithm can be input with training data to generate an expected output or a desired output. Supervised learning algorithms can include deep learning algorithms, support vector machines (SVMs), neural networks, or random forests. Through a machine learning algorithm, a trained algorithm may be able to identify the relationship of biomarkers to a specific cancer prognosis or diagnosis. Without a trained algorithm, it may be difficult to identify the relationship of biomarkers to accurately identify the presence of cancer or other parameters related to cancer.
[0133] In various aspects, the system and method may include the accuracy, sensitivity, or specificity of the detection of cancer or cancer parameters. For example, the method or system may include detecting the presence or absence of cancer (or the presence of cancer parameters such as recurrence, relapse, or drug resistance) in a subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of cancer parameters such as recurrence, relapse, or drug resistance) in a subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of cancer parameters such as recurrence, relapse, or drug resistance) in a subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of cancer parameters such as recurrence, relapse, or drug resistance) in a subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of cancer parameters such as recurrence, relapse, or drug resistance) in a subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0134] Computer system
[0135] The present disclosure provides a computer system programmed to implement the methods of the present disclosure. FIG. 17 shows a computer system (1701) programmed or otherwise configured to perform an analysis or process of a method, such as determining a tumor mutation burden, determining a copy number burden, or executing an algorithm. The computer system (1701) can coordinate various aspects of the methods and systems of the present disclosure, such as executing an algorithm, inputting training data, analyzing a set of biomarkers, or outputting results to a user regarding a tumor mutation burden or copy number burden. The computer system (1701) can be a user's electronic device or a computer system located remotely from the electronic device. The electronic device can be a mobile electronic device.
[0136] The computer system (1701) includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") (1705), which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (1701) also includes a memory or memory location (1710) (e.g., random access memory, read-only memory, flash memory), an electronic storage unit (1715) (e.g., hard disk), a communication interface (1720) (e.g., network adapter) for communicating with one or more other systems, and peripheral devices (1725) such as a cache, other memory, data storage, and / or an electronic display adapter. The memory (1710), storage unit (1715), interface (1720), and peripheral devices (1725) communicate with the CPU (1705) via a communication bus (solid lines) such as a motherboard. The storage unit (1715) can be a data storage unit (or data repository) for storing data. The computer system (1701) can be operably coupled to a computer network ("network") (1730) using the communication interface (1720). The network (1730) can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. The network (1730) is, in some cases, a telecommunications and / or data network. The network (1730) can include one or more computer servers that enable distributed computing such as cloud computing. The network (1730) can, in some cases, implement a peer-to-peer network using the computer system (1701) to enable devices coupled to the computer system (1701) to function as clients or servers.
[0137] The CPU (1705) can execute an array of machine-readable instructions that can be embodied in a program or software. The instructions can be stored at a memory location, such as in the memory (1710). The instructions can be directed to the CPU (1705), which can then program or otherwise configure the CPU (1705) to implement the methods of the present disclosure. Examples of operations performed by the CPU (1705) can include fetch, decode, execute, and write-back.
[0138] The CPU (1705) can be part of a circuit, such as an integrated circuit. One or more other components of the system (1701) can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0139] The storage unit (1715) can store files such as drivers, libraries, and saved programs. The storage unit (1715) can store user data, such as user preferences and user programs. The computer system (1701) can, in some cases, include one or more additional data storage units external to the computer system (1701), such as located on a remote server in communication with the computer system (1701) through an intranet or the Internet.
[0140] A computer system (1701) can communicate with one or more remote computer systems via a network (1730). For example, the computer system (1701) can communicate with a remote computer system of a user (e.g., a medical professional or a patient). Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. A user can access the computer system (1701) via the network (1730).
[0141] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system (1701), such as in a memory (1710) or an electronic storage unit (1715). The machine executable or machine readable code can be provided in the form of software. In use, the code can be executed by a processor (1705). In some cases, the code can be retrieved from the storage unit (1715) and stored in the memory (1710) for easy access by the processor (1705). In some situations, the electronic storage unit (1715) can be excluded and the machine executable instructions can be stored in the memory (1710).
[0142] The code can be pre-compiled, configured, or compiled at runtime for use on a machine having a processor adapted to execute the code. The code can be supplied in a programming language selected to enable the code to be executed in a pre-compiled or compiled-in-place manner.
[0143] Aspects of the systems and methods provided herein, such as computer system (1701), can be embodied in programming. Various aspects of the technology can typically be regarded as “products” or “articles of manufacture” in the form of machine (or processor) executable code, and / or associated data carried on or embodied in some type of machine-readable medium. The machine executable code can be stored in an electronic memory unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A “storage” type of medium can include any or all of tangible memories such as computers, processors, or their associated modules such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage at any time for software programming. All or part of the software can sometimes be communicated via the Internet or various other electrical communication networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, another type of medium that can carry software elements includes those used over physical interfaces between local devices, through wired and optical landline networks, as well as via various air links, including light, electricity, and electromagnetic waves. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be regarded as media that carry software. As used herein, the term such as computer or machine “readable medium” refers to any medium involved in providing instructions to a processor for execution, unless limited to non-transitory and tangible “storage” media.
[0144] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks such as any of the storage devices in any computer, such as those that can be used to implement a database shown in the drawings. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus in a computer system. Carrier wave transmission media can take the form of acoustic or light waves such as electrical signals or electromagnetic signals, or those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tape, any other physical storage media with patterns of holes, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carrier waves that carry data or instructions, cables or links that carry such carrier waves, or any other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0145] A computer system (1701) can include, or be communicable with, an electronic display (1735) that includes, for example, a user interface (UI) (1740) for providing visual output for input of biomarker or sequencing data, or for detection, diagnosis, or prognosis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0146] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit (1705). The algorithms can, for example, determine tumor mutation burden or copy number burden.
Example
[0147] Example 1: Analysis of cell-free tumor DNA (ctDNA) across the genome to evaluate bTMB and bCNB
[0148] CDK4 / 6 inhibition (CDK4 / 6i) in combination with endocrine therapy (ET) improves the survival of patients with hormone receptor-positive (HR+) / human epidermal growth factor receptor 2-negative (HER2-) metastatic breast cancer (MBC). However, there is a lack of clinical biomarkers to identify patients who may not respond. The methods and systems of the present disclosure were used to perform genome-wide circulating tumor DNA (ctDNA) analysis to identify features associated with resistance to ET and CDK4 / 6i.
[0149] At baseline and during treatment in a phase II trial of palbociclib in combination with letrozole or fulvestrant (NCT3007979), ctDNA was isolated from 216 plasma samples collected from 51 patients with HR+ / HER2- MBC. At baseline and at clinical progression, boosted whole-exome sequencing (WES) was performed to profile genomic changes, evaluate mutation signatures, and derive blood tumor mutation burden (bTMB). Low-pass whole-genome sequencing was performed at baseline, serial time points of treatment, and at clinical progression to evaluate blood copy number burden (bCNB).
[0150] Results were obtained that included that high bTMB and high bCNB were associated with lack of clinical benefit and significantly shorter progression-free survival (PFS) compared to patients with low bTMB or low bCNB (all P<0.05). For low bTMB (0 / 37, 0%), in the case of high bTMB (5 / 13, 38.5%), the dominant APOBEC signature was detected exclusively at baseline (P=0.0006). New changes reported previously were detected at baseline and at progression in relation to treatment resistance. Changes in ESR1 were enriched in samples with high bTMB (P=0.0005). There was a high correlation between bTMB determined by WES and bTMB determined using a 600-gene panel (R=0.98). During serial monitoring, increases in bCNB preceded radiographic progression in 12 / 18 (66.7%) patients.
[0151] The results showed that genomic complexity demonstrated by high bTMB and bCNB was associated with lack of response and poor outcomes for patients treated with ET and CDK4 / 6i. This subset of HR+ / HER2- patients requires investigation of novel treatment strategies, including immunotherapy-based combinations. Non-invasive monitoring in blood was performed to identify the emergence of resistance changes and early evidence of pre-imaging progression.
[0152] The combination of endocrine therapy (ET) and cyclin-dependent kinase 4 / 6 inhibition (CDK4 / 6i) has emerged as the standard-of-care, first-line treatment for patients with hormone receptor-positive (HR+) / HER2-negative metastatic breast cancer (MBC). This treatment indication is based on significant improvements in survival outcomes and an extended chemotherapy-free interval across all clinical and pathological subgroups [1-5]. Thus, outside of clinical trials or impending organ failure, patients in the United States and Europe are offered CDK4 / 6i and ET as first-line treatment. Despite this advance in the care of patients with HR+ / HER2-negative MBC, a subset of patients progress rapidly and lack biomarkers to predict efficacy and resistance.
[0153] Analysis of circulating tumor DNA (ctDNA) using next-generation sequencing (NGS) enables non-invasive assessment of genomic changes during tumor progression and can be used to identify biomarkers for predicting and monitoring response to treatment [6-10]. In 2019, the US Food and Drug Administration (FDA) approved a ctDNA-based companion diagnostic test for the detection of PIK3CA mutations to select patients for treatment with alpelisib, leading to an increase in the use of ctDNA tests in clinical practice [11, 12]. NGS profiling based on both tissue and blood identified individual changes associated with resistance in patients treated with ET together with CDK4 / 6i, including changes in CCNE1, FGFR1, FAT1, PTEN, and RB1 [13-18]. However, to date, there are no clinical, pathological, or genomic signatures identified as being predictable at baseline to define subsets of patients who will benefit from alternative treatment strategies. Therefore, a comprehensive NGS-based liquid biopsy approach that includes assessment of ctDNA mutations and copy number burden was developed to identify prognostic and predictive biomarkers in patients with HR+ / HER2-negative MBC and to track response to ET and CDK4 / 6i treatment. To achieve this, the inventors utilized a combination assay that provides targeted coverage of 600 cancer genes in addition to whole-exome sequencing (WES) to enable comprehensive genomic profiling, assessment of mutational signatures, and derivation of baseline and on-treatment burden of tumor mutational burden (bTMB). Furthermore, the inventors performed low-pass whole-genome sequencing (LP-WGS) to derive a novel measure of genome-wide copy number variants (CNVs).
[0154] Tumor mutational burden (TMB), as used herein, generally refers to a measure of the number of mutations per megabase of sequenced DNA that can be measured, for example, using whole exome sequencing (WES)
[19] . The rationale for initially developing TMB as a clinical biomarker derived from tissue is based on the observation that tumor types with high tissue TMB (tTMB) (e.g., non-small cell lung cancer (NSCLC) in smokers, ultraviolet-related melanoma, and mismatch repair deficient tumors) respond well to immune checkpoint inhibitor (ICI) therapy [20-23]. tTMB may be a potential surrogate biomarker for neoantigen load to predict response to ICI monotherapy and, in combination with PD-L1 expression in tumors or immune cells, may be promising as a non-overlapping biomarker
[24] . Blood tumor mutational burden (bTMB) can be explored as a non-invasive method for TMB determination in NSCLC, and in some cases, it is difficult to obtain appropriate tissue for sequencing. Studies conducted using NGS targeted cancer gene panels may show that NSCLC patients with high bTMB respond preferentially to ICI over chemotherapy [25-27]. However, applying WES to measure TMB in blood samples may face various technical challenges [28, 29]. Compared to other malignancies, the evaluation of TMB in breast cancer is not as extensive, and many studies have been conducted instead of evaluating tTMB. The evaluation of tTMB shows that patients with breast cancer have a relatively low median tTMB, but tTMB may be higher in metastatic tissue compared to primary tissue. Importantly, early data may show that a subset of breast cancer patients with high TMB benefit from PD-1 inhibitors, with or without anti-CTLA-4 [30-32].Furthermore, parallel evaluation of mutation signatures in hypermutated malignancies can reveal the presence of APOBEC (apolipoprotein B mRNA-editing enzyme catalytic polypeptide-like) mutation signatures in high tTMB patients, which may be associated with response to ICI [30, 33-35].
[0155] Blood copy number burden (bCNB) derived from the PredicineCNB™ assay is a comprehensive measure of CNV via LP-WGS, including amplifications and deletions across the genome. Current strategies for blood-based treatment response monitoring can mainly track individual ctDNA mutations or changes in allele frequencies to evaluate tumor response to systemic therapy, but the integration of copy number changes and genome-wide methylation can provide early signals of response for patients treated with various systemic therapies prior to standard-of-care imaging [36-40]. Considering that LP-WGS is less expensive than other NGS methods and thus more feasible to perform sequentially from a cost perspective, this technique can provide a clinical utility for monitoring dynamic changes in CNV during the course of treatment. However, studies evaluating this technique in patients with MBC treated with CDK4 / 6i are limited.
[0156] Using the systems and methods of the present disclosure, two novel genome-wide ctDNA assays that combine sequencing breadth and depth were used to profile patients with HR+ / HER2-negative MBC receiving a combination of ET and CDK4 / 6i in a prospective phase II interventional clinical trial. Resistance biomarkers were determined to evaluate which patients may be suitable candidates for novel treatment strategies and to explore the potential for serial ctDNA monitoring to predict early disease progression. The inventors' comprehensive approach identified bTMB and bCNB levels that predict poor patient outcomes, identified APOBEC signatures exclusively in hypermutated patients, defined an expanded list of candidate changes that may mediate resistance at baseline and at clinical progression, and demonstrated the use of bCNB for predicting and monitoring early disease progression.
[0157] The patient cohort was obtained as follows. ctDNA samples of patients were retrospectively analyzed from a prospective, single-arm, phase II study (NCT03007979) conducted at Washington University School of Medicine (St. Louis, MO) and the University of Nebraska Medical Center (Omaha, NE). HR+ / HER2-negative MBC patients treated with 0-1 line of prior systemic therapy without prior use of CDK4 / 6i were enrolled. Patients were administered palbociclib 125 mg daily on a continuous 5-days-on, 2-days-off schedule in combination with letrozole or fulvestrant (at the physician's choice), and goserelin was administered to premenopausal patients. Each treatment cycle was 28 days. At baseline, on day 15 of cycle 1 (C1D15), day 1 of cycle 2 (C2D1), and day 1 of cycle 4 (C4D1), and then every 3 cycles until disease progression (with tumor imaging) on D1, research blood samples were collected in Streck tubes. Fifty-four patients were enrolled in this study, of which 51 were evaluable for response and included in this analysis. At the data cut-off point, 29 patients were removed from the study due to disease progression, thus samples were available from 29 patients during disease progression. Plasma samples collected at the time point immediately prior to clinical progression for these patients were also included in this analysis. The primary endpoint (rate of grade 3 or 4 neutropenia) and the results of clinical response were obtained
[41] . The study was approved by the institutional review board of each facility, and written informed consent was obtained from all patients to permit correlative research on blood samples.
[0158] ctDNA analysis was performed as follows. Patient samples were analyzed using two comprehensive NGS platforms, PredicineWES+(TM) and PredicineCNB(TM) (Predicine, Inc., Hayward, CA), to generate genomic profiles, perform mutation signature and pathway analyses, and derive measurements of bTMB and bCNB. Briefly, cell-free DNA (cfDNA) extracted from patient plasma samples and germline DNA extracted from peripheral blood mononuclear cells (PBMC) were processed and subjected to library construction. The resulting DNA libraries were sequenced at 5x coverage depth by PredicineCNB(TM) LP-WGS or further enriched by hybrid capture, and sequencing was performed with PredicineWES+(TM), a combinatorial assay designed to sequence the entire exome at a sequencing depth of 2,500x (detection level (LOD) 1%), and boosted sequencing of 600 cancer genes covered by the PredicineATLAS(TM) target panel was also performed at a sequencing depth of 20,000x (0.25% LOD) (Table 3). Using PredicineWES+(TM) sequencing data, a landscape of genomic changes including single nucleotide variants (SNVs), insertions and deletions (indels), copy number variations (CNVs), and gene fusions was created, a bTMB score reporting the total number of somatic mutations detected per megabase of DNA was derived, and mutation signatures and involvement of oncogenic signaling pathways were analyzed. Also, to compare bTMB values generated by PredicineWES+(TM), bTMB scores were derived from sequencing data generated by analyses using the targeted 600-gene PredicineATLAS(TM) panel and the targeted 152-gene PredicineCARE(TM) panel (Table 4). PredicineCNB(TM) sequencing data were evaluated to generate a bCNB score representing a comprehensive genomic-scale measure of CNV, including amplifications and deletions across the entire genome.
[0159]
Table 3
[0160]
Table 4
[0161] Statistical analysis was performed as follows. The clinical benefit rate (CBR), defined as the percentage of patients with a complete response, partial response, or stable disease lasting at least 24 weeks according to RECIST (version 1.1), and the statistical associations between individual changes, bTMB, and bCNB were analyzed using the Wilcoxon test and the Kruskal–Wallis test. The frequency of changes across patient subgroups was compared using the Fisher’s Exact test. The comparison of the frequency of changes across patient subgroups at the baseline and clinical progression time points was performed using McNemar’s Test. The degree of association between variables was evaluated with the Spearman rank correlation coefficient. The Kaplan–Meier (K-M) method was applied to estimate the empirical survival probability, the K-M curve was used to illustrate the survival period, and the log-rank test was utilized to compare the differences in survival. The hazard ratio and 95% confidence interval were estimated from univariate Cox proportional hazards regression analysis. Different cutoffs, including unbiased cutoffs for the median and third quartile, were applied to bTMB for the analysis of its association with PFS, and the optimal cutoff was further investigated based on the Cox model settings for PFS and the Harrell’s C-index in the receiving operating characteristic (ROC) analysis for clinical benefit. The changes in bCNB were evaluated at consecutive time points and compared with the concurrent assessment of clinical progression based on RECIST 1.1.
[0162] Serial ctDNA samples were retrospectively analyzed from a prospective clinical trial as follows. The ctDNA assay was performed retrospectively on samples collected from a prospective clinical trial of palbociclib combined with ET (letrozole or fulvestrant)
[41] . Two hundred and sixty-five samples from 51 evaluable patients with HR+ / HER2− negative MBC were analyzed using the Predicine liquid biopsy NGS platform (Figure 1A). At the data cutoff time point, 29 patients had a median follow-up time of 16.4 months (range 1.4–50.9) during which the trial was ongoing. Only two (2 / 265) samples failed sequencing quality control, and more than 99% of the ctDNA samples were successfully sequenced. For 78 plasma samples collected at the baseline (n = 50) and clinical progression (n = 28) time points, PredicineWES+™ analysis was performed by combining 2,500X (1% LOD) whole-exome sequencing (WES) and boosted sequencing of 600 genes (20,000X, LOD 0.25%). PredicineCNB™ was performed at a sequencing depth of 5X for 216 plasma samples, including all available baseline (n = 51), C1D15 (n = 47), C2D1 (n = 51), pre-progression staging evaluation (n = 38), and clinical progression (n = 29) time points (Figures 1B–1C).
[0163] The clinical and pathological characteristics of the patients included in the trial are summarized in Table 5. The majority of the patients were postmenopausal (84.3%), received letrozole (72.6%), and the remaining patients received fulvestrant (27.5%). Based on the ESMO 2020 criteria, a total of 17 patients were de novo metastatic, 22 patients were classified as endocrine resistant, and 12 patients were endocrine sensitive
[42] .
[0164]
Table 5-1
[0165]
Table 5-2
[0166] Furthermore, a high baseline bTMB has been demonstrated to be associated with poor clinical outcomes. For 50 patients at the baseline time point, bTMB was evaluable (Figure 2A). The median bTMB was 1.85 mutations per megabase pair (MBp) [interquartile range (IQR) 1.01 - 3.86, range 0.1 - 71.7]. Patients with no clinical benefit who experienced disease progression within 6 months (N = 10) had significantly higher bTMB compared to patients with clinical benefit (N = 40) (median 8.90 mutations / MBp [IQR 2.30 - 31.2] vs median 1.63 [IQR 0.69 - 2.83], Wilcoxon, P = 0.012) (Figure 2B). Patients with ESR1 mutations at the baseline time point (N = 8) also had significantly higher baseline bTMB (Wilcoxon, P = 5.0×10 -4 ) (Figure 2C), and similar associations were detected for baseline mutations in the ARID1A, BSN, CDH1, DNAH3, DSP, MUC6, MUC16, PIK3CA, and USH2A genes (Figures 7A - 7C). Comparing patients based on the clinical classification of de novo, endocrine-resistant, or endocrine-sensitive MBC, no significant difference in baseline bTMB was observed (Figure 2D). However, the majority of patients with high bTMB at the baseline time point were in the endocrine-resistant cohort. ROC analysis was performed and the optimal cut-off point for bTMB in our dataset was determined to be 3.2 mutations / MBp, with an AUC of 0.76 (Figure 8), which closely aligns with the bTMB value of 3.8 that dichotomizes patients above and below the third quartile. Higher bTMB at the baseline time point was significantly associated with worse PFS (median 13.8 months vs 32.1 months) based on the median of the samples (hazard ratio (HR) 2.62 [confidence interval (CI) [1.21 - 5.66], P = 0.011) (Figure 2E). A similar significant difference between high and low bTMB was observed at the third quartile of the samples (median 6.5 months vs 32.1 months) (HR 4.87 [CI 2.19 - 10.81], P = 2.27×10 -5)(Figure 2F), and a cutoff of 10 mutations / MBp (median 3.8 months vs. 22.3 months) (HR 7.15 [CI 2.82 - 18.13], P = 1.94×10 -6 )(Figure 2G). When evaluating survival in the endocrine-resistant cohort, a significant difference was observed in high baseline bTMB that was significantly associated with worsened PFS (HR 4.76 [CI 1.52 - 14.97], P = 0.004) (Figure 2H). In summary, high baseline bTMB was determined to be significantly associated with lack of clinical benefit and shorter PFS as measured using multiple pre-determined and experimentally determined cutoff points.
[0167] Furthermore, the inventors demonstrated that bTMB scores generated from targeted sequencing panels and WES are highly correlated. Using the PredicineWES™ assay, bTMB levels generated from 50 baseline samples were compared to values obtained using the targeted 600-gene PredicineATLAS™ and 152-gene PredicineCARE™ sequencing assays. bTMB values obtained by PredicineWES+™ were highly correlated with levels derived from PredicineATLAS™ (R = 0.98) (Figure 3), and PredicineCARE™ (R = 0.93) (Figure 9) (Spearman's rank test). In summary, these comparisons indicate that accurate bTMB scores can also be generated from shorter NGS gene panels.
[0168] Furthermore, the inventors demonstrated that high cfDNA yields were associated with significantly shorter PFS based on the median (HR 2.36 [CI 1.12–4.98], P = 0.021) and third quartile (HR 2.96 [CI 1.34–6.54], P = 0.006) cut-offs of the samples (Figure 10). In the endocrine-resistant cohort, high cfDNA yields were also associated with significantly shorter PFS based on the median (HR 3.45 [CI 1.18–10.14], P = 0.017) and third quartile (HR 4.35 [CI 1.29–14.63], P = 0.01) cut-offs. Endocrine sensitivity versus endocrine resistance and disease site in the clinical presentation (e.g., bone or visceral) did not predict CBR or PFS in patients treated with palbociclib combined with ET (Table 5). Similarly, there was no significant association between the sites of metastatic spread at baseline and bTMB (Figure 11). In summary, it was determined that high baseline cfDNA levels are associated with poor clinical outcomes but not with the commonly evaluated clinicopathological features.
[0169] Furthermore, the inventors demonstrated that the dominant APOBEC mutation signature is exclusively present in patients with high bTMB. The off-target activity of the APOBEC family of mutagenic enzymes can generate somatic mutations throughout the genome, resulting in distinct mutation signatures associated with the development and progression of multiple cancers [22, 43-44]. To evaluate the contribution of these mutation signatures to the genomic landscape of high bTMB patients versus low bTMB patients in this cohort, sequencing data obtained by PredicineWES™ from 50 baseline patients were evaluated for single nucleotide substitution (SBS) patterns and compared to 94 curated reference SBS mutation signatures available in the COSMIC database. The dominant APOBEC signature was exclusively identified in high bTMB patients, while other signatures were observed across both the high bTMB and low bTMB groups (Figure 4A). The dominant APOBEC signature was detected in 5 / 13 (38.5%) patients with high bTMB versus 0 / 37 (0%) patients with low bTMB (P = 0.0006, Fisher's exact test) (Figure 4B). The median bTMB score was significantly higher in the dominant APOBEC signature (34.8 Mbp) compared to patients with other signatures (1.7 Mbp) (P = 0.00048, Wilcoxon rank sum test) (Figure 4C). Collectively, these findings indicate the determination of a subset of hypermutated HR+ / HER2-negative patients with MBC.
[0170] Furthermore, the inventors showed that in high bTMB patients and high bCNB patients, specific cancer signaling pathways are more frequently altered. To compare the relative proportions of changes within the major oncogenic signaling pathways in high bTMB patients versus low bCNB patients, the frequencies of changes identified across all breast cancer driver genes present in 12 pathways were compared [34-35, 45]. Cell cycle (P = 0.04), DNA damage repair (DDR) (P = 0.02), Hippo (P = 0.009), NOTCH (P = 0.003), PI3K (P = 2.9x10 -05) and across breast cancer driver genes in the oncogenic signaling pathway of receptor tyrosine kinase (RTK)-RAS (P = 0.005), changes (including SNV and CNV) were observed at significantly higher frequencies in high bTMB patients compared to low bTMB patients (P = 0.04) (Fisher's exact test) (Figure 12). Also, in the cell cycle (P = 0.009), DDR (P = 0.001), Hippo (P = 0.04), Notch (P = 0.04), RTK-RAS (P = 0.007), and TP53 (P = 0.002) pathways, significantly higher frequencies of changes were observed in high bCNB patients compared to low bCNB patients (Fisher's exact test) (Figure 12). These findings indicate that the relevant signaling pathways involved in driving ET and CDK4 / 6i resistance are preferentially activated in genomically complex breast cancer patients with excessive mutations in this cohort.
[0171] Furthermore, the inventors demonstrated that comprehensive profiling expanded the detection of clinically relevant ctDNA changes at baseline and detected the enrichment of novel ctDNA changes during progression. The PredicineWES+(TM) assay was performed on 50 / 51 samples collected at baseline and 28 / 29 samples collected during progression. One of the 51 baseline samples was sequenced using the PredicineATLAS(TM) assay instead of the PredicineWES+(TM) assay, and one of the samples during progression failed library yield quality control. Across all 51 patients at baseline, the most frequently observed changes were PIK3CA (45%), TP53 (31%), and ESR1 (20%) (Figure 13). Baseline changes (SNVs and CNVs) in 17 genes were significantly associated with worse PFS, including AURKA, AKT3, ATM, BRCA2, CCND1, CCNE2, DDR2, DSP, ESR1, MYC, MUC16, PIK3CA, PLCG1, RB1, RUNX1T1, USH2A, and ZFHX3 (Table 6). Five of these genes (CCNE2, DSP, MUC16, PLCG1, and USH2A) are not targeted by the PredicineATLAS(TM) panel and may not be involved in CDK4 / 6i and ET resistance, thereby representing novel resistance changes detected by WES.
[0172]
Table 6
[0173] The comparison of the most frequently changed genes detected at progression with the baseline was performed across all evaluable samples from the patients (28 / 29) who were profiled at the time of analysis (Figure 5A). The most frequently observed changes across all 28 patients at baseline were PIK3CA (54%), TP53 (39%), AKT3 (32%), DDR2 (29%), ATM (29%), AURKA (25%), ESR1 (25%), BRCA2 (21%) and EGFR (21%), whereas the most frequently observed changes at progression were TP53 (50%), PIK3CA (43%), RB1 (36%), AURKA (32%), CCND1 (32%), ESR1 (32%), BRCA2 (29%), ATM (25%), and MUC12 (25%). Changes in RB1 were significantly enriched at progression (36% vs 14%) (P = 0.04) (McNemar's test) (Figure 5B). Some additional changes previously implicated in CDK4 / 6i and / or ET resistance showed non-significant enrichment at progression compared to baseline, including AR (18% vs 7%), AURKA (32% vs 25%), CCND1 (32% vs 18%), CDKN2A (21% vs 14%), ESR1 (32% vs 25%), FGFR1 (18% vs 14%), PTEN (21% vs 11%), MYC (18% vs 14%) and TP53 (50% vs 39%). Additional changes not commonly implicated in CDK4 / 6i and ET resistance, including BRCA2 (29% vs 21%), CBL (18% vs 7%), CDH1 (18% vs 11%), KMT2D (18% vs 14%), MUC12 (25% vs 14%) and PREX2 (21% vs 14%), showed non-significant enrichment at progression (Figure 5B). Two of the enriched genes (MUC12 and PREX2) were not targeted by the PredicineATLAS™ panel, thereby representing novel changes detected by WES. In contrast to increases in the levels of some individual variants, no significant difference was observed between the median bTMB or bCNB levels detected at baseline and at progression (Figures 14A - 14B).In summary, the extension of boosted WES sequencing to blood samples taken at baseline and during progression enabled the determination of additional candidate biomarkers for de novo and acquired resistance.
[0174] Furthermore, the inventors showed that the bCNB score predicts poor clinical outcomes and increases before clinical progression is detected by imaging. From LP-WGS data generated from all 51 baseline samples, 47 C1D15 samples, 51 C2D1 samples, 38 staging samples, and 29 progression samples, a bCNB score reflecting genome-wide CNV assessment was derived (Figure 1B). The median bCNB score at baseline was 9.36 [IQR 5.78–13.2]. The baseline bCNB score was significantly higher in patients who experienced progression within 6 months (no clinical benefit) (Wilcoxon test, P = 0.036) (Figure 6A), and a high baseline bCNB score defined by a cutoff of 5.6 was significantly associated with shorter PFS (HR 3.14 [CI 1.34–7.74] (P = 0.009)) (Figure 6B). The baseline bCNB score was also correlated with the baseline bTMB score (R = 0.68, P = 6.6 × 10 -08) was significantly correlated (Figure 15). Serial analysis of bCNB during treatment revealed a decrease at C1D15 and / or C2D1 relative to baseline levels in 38 / 51 (74.5%) patients. Staging samples were collected concomitantly with imaging studies performed every three months for assessment of tumor response. For 18 of the 29 progressive patients, 1 - 3 time points (3 - 9 months) immediately prior to progression were available. Analysis of these samples revealed an increase (above the previous nadir level) at least three months prior to radiographic detection of progressive disease in 12 / 18 (66.7%) of the patients, as shown in Figure 6C. An increase in bCNB preceded radiographic detection of progressive disease up to six months in six patients and up to nine months in one patient. Comparison of dynamic bCNB patterns using orthogonal measures of ctDNA fraction from PredicineATLAS™ profiling in a subset of four patients revealed patterns similar to the ctDNA fraction levels (Figure 16). In summary, the inventors have shown that high baseline bCNB is significantly associated with poor clinical outcomes and that an increase in bCNB precedes radiographic detection of clinical progression in two - thirds of the cases monitored over the course of treatment.
[0175] As described, the inventors reported a comprehensive ctDNA NGS analysis including a plasma-based boosted WES assay, LP-WGS, and a bioinformatics pipeline to determine bTMB and bCNB to enable genome-wide evaluation of novel resistance mechanisms and clonal evolution in patients with HR+ / HER2− MBC who receive ET in combination with CDK4 / 6i. Specifically, the inventors identified a subset of patients defined by hypermutation (high bTMB) and increased copy number polymorphisms (high bCNB) as being associated with poor outcomes requiring novel treatment strategies. Furthermore, PredicineWES™ enabled the extended detection of genomic changes associated with resistance at baseline and progression. The inventors also demonstrated that dynamic changes in LP-WGS-derived bCNB scores over the course of treatment precede radiographic response and clinical progression in a subset of patients, identifying potential utility for response monitoring. These studies using non-invasive blood-based sequencing represent a comprehensive assessment of genome-wide ctDNA in this patient population, leading to the generation of biological insights and potential treatment hypotheses to improve clinical outcomes.
[0176] Importantly, bTMB and bCNB were determined using one 8-ml tube of whole blood in all patients with evaluable samples, demonstrating feasibility from the perspective of clinical application. As expected, the median bTMB was relatively low in this cohort (less than 2 MBp), a finding that could be consistent with the evaluation of tTMB in breast cancer patients, particularly those with HR+ MBC
[30] . Based on the lack of clinical benefit and the observed association of high bTMB with shorter PFS, the inventors demonstrated a stratification tool with treatment implications. tTMB can predict the response of patients treated with ICI monotherapy in other tumor types
[20] . However, it can still be difficult to define an optimal cut-off point based on different sequencing platforms, bioinformatics techniques, and the use of methods to determine tTMB. Thus, the optimal tTMB threshold is likely to vary across different tumor types
[46] . Therefore, the inventors did not use a priori bTMB thresholds in the outcome analysis. Instead, multiple bTMB thresholds, including the median (1.9 MBp), the third quartile (3.8 MBp), and the FDA-approved threshold of 10 MBp in tissue, were significantly associated with PFS. These findings reinforce the consistency in defining a hypermutated, resistant subset of patients with higher bTMB.
[0177] Interestingly, in the cohort of the present inventors, high bTMB patients were enriched for the dominant APOBEC mutation signature. The APOBEC family of DNA editing enzymes generates mutations during various normal biological processes, including innate and adaptive immune responses
[47] . However, the upregulated "off target" activity of APOBEC enzymes can be a major source of somatic mutations in several cancers that result in characteristic mutation signatures [22, 43 - 44]. APOBEC signatures can be observed in various hypermutated malignancies and may be associated with response to ICI [30 - 31, 33, 48]. Within this cohort, the enrichment of these signatures in high bTMB HR+ / HER2 - negative patients compared to low bTMB HR+ / HER2 - negative patients further emphasizes the identification of a biomarker - defined subset of patients who may benefit from the incorporation of ICI therapy. The inventors' study also identified several oncogenic pathways associated with high bTMB (e.g., cell cycle, DDR, NOTCH, PI3K, and RTK - RAS) as potential drug targets.
[0178] The data of the present inventors also demonstrate an overlap between patients with high bTMB and endocrine resistance as defined by the ESMO criteria. Patients with clinically defined endocrine resistance had a similar median bTMB compared to patients with de novo MBC or patients with endocrine-sensitive disease, but the majority of high bTMB cases were present in the endocrine resistance cohort at baseline. Importantly, the bTMB score stratified PFS in a subgroup of patients with clinically defined endocrine resistance. Furthermore, patients with ESR1 mutations at baseline had higher bTMB scores compared to patients with wild-type ESR1. Clinically defined endocrine resistance, the site of metastatic disease on imaging, and other pathological variables did not stratify baseline patients with poor prognosis, further supporting the need for novel biomarkers for risk stratification. In summary, these findings demonstrate the use of the bTMB score to define a subgroup of patients who are likely to be less responsive to standard first-line therapy with CDK4 / 6i and ET, and these findings can be used to determine alternative combination treatment strategies, including ICI, to administer to patients.
[0179] The findings of the present inventors indicate that novel treatment strategies are needed for patients with high bTMB and high bCNB at baseline. The high association of high tTMB (defined by a threshold of over 10 mutations / MBp) with response to ICI based on the tissue agnostic approval of pembrolizumab for patients with high tTMB indicates a potential treatment approach
[49] . In the TAPUR and NIMBUS studies, a subset of patients with MBC and high tTMB across subtypes were durable responders [32, 50]. However, in other non-biomarker selected populations, there was no improvement in outcomes when ICI was added to chemotherapy
[51] . For this reason, it is necessary to evaluate the potential of incorporating ICI for HR+ / HER2-negative patients with high bTMB, either as monotherapy or in combination. Preclinical data indicate that CDK4 / 6i may enhance T cell activation, increase tumor infiltration, and have a synergistic effect with ICI therapy
[52] . Chemotherapy for patients with HR+ / HER2-negative MBC may be reserved for impending organ failure or endocrine refractory disease, but the optimal use of chemotherapy in this biologically defined cohort may be investigated, and these patients may benefit from earlier incorporation of cytotoxic therapy. In addition, since bCNB may precede the clinical detection of disease recurrence, interventional studies may be conducted to determine whether early switching of treatment based on molecular progression of the disease can improve clinical outcomes, as opposed to imaging of progression.
[0180] The use of a genome-wide approach also identifies many individual resistance changes, validates previously associated mechanisms, and leads to the discovery of novel candidate genes. Similar to the novel baseline changes in DSP, MUC16, PLCG1, USH2A, and ZFHX3, baseline changes in other genes associated with de novo resistance to RB1 and ET CDK4 / 6i therapies were determined to be associated with shorter PFS. Median levels of bTMB and bCNB did not increase significantly at the time of clinical progression, but the inventors observed enrichment of individual changes previously involved in endocrine and / or CDK4 / 6i treatment resistance, including AR, AURKA, CCND1, CDKN2A, ESR1, FGFR1, MYC, and RB1
[53] . The inventors also observed significant changes in genes not commonly associated with ET and CDK4 / 6i treatment resistance, including CBL [54-56], a member of the RING finger ubiquitin ligase family that regulates receptor tyrosine kinase signaling, KMT2D
[57] , a methyltransferase involved in estrogen receptor mobilization and activation, MUC12 [58-60], a glycosylated transmembrane protein in the mucin family involved in the regulation of proliferation, invasion, and metastatic potential, and PREX2 [61-62], a guanine nucleotide exchange factor that regulates cancer cell motility and invasion, which encode various oncogenic proteins. Many of these novel alterations are not covered by targeted sequencing panels, highlighting the value of extended WES for identifying diverse resistance mechanisms.
[0181] Performing WES extends the gold standard of TMB measurement to blood samples, enabling the discovery of novel candidate resistance mechanisms, which may further enable clinical applications such as the administration of appropriate therapeutic agents. However, the inventors observed a high correlation between bTMB measurements obtained by WES and a shorter targeted sequencing panel, demonstrating the utility of using a cost-effective test to measure bTMB in the clinic. Furthermore, the inventors demonstrated that bCNB derived from the more cost-effective PredicineCNB™ showed a high degree of concordance with baseline bTMB and was associated with poor patient outcomes. The inventors also demonstrated the utility of serial bCNB evaluations for monitoring the dynamic changes in ctDNA during treatment. bCNB decreased early, within 2 weeks after treatment initiation, thereby providing an early signal of molecular response to treatment (e.g., earlier than detectable by imaging). Furthermore, when comparing simultaneous imaging and bCNB evaluations, an increase in bCNB preceded clinical progression of the disease in two-thirds of the patients. Thus, serial blood-based molecular evaluations may serve as a surrogate for PFS, which can be clinically evaluated via imaging.
[0182] Since some of the patients in the study did not progress, the landscape of changes at the time of progression can be investigated to assess the extent to which it reflects patients with a long-term response to treatment. Also, alternative dosing regimens of palbociclib can be used, which may affect the occurrence of resistance changes. Furthermore, for the TMB test, blood biopsies and tissue biopsies can be performed simultaneously.
[0183] In summary, our study demonstrates the potential utility of blood-based bCNB and bTMB assessment in treatment decisions for patients with HR+ / HER2-negative MBC. Furthermore, this study demonstrates the utility of performing whole-genome ctDNA analysis to comprehensively define the molecular mechanisms of baseline and continuous resistance to CDK4 / 6i in combination with ET for patients with HR+ / HER2-negative MBC. The results identify a subset of patients at baseline with poor outcomes to first-line standard therapies and also demonstrate a non-invasive approach for detecting early blood-based progression. Using the results of bTMB and bCNB, optimal treatments can be selected for these hypermutated genomically complex patients, such as the incorporation of initial ICI and combination therapies.
[0184] Reference
[0185] 1. Sledge GW, Jr., Toi M, Neven P et al. The Effect of Abemaciclib Plus Fulvestrant on Overall Survival in Hormone Receptor-Positive, ERBB2-Negative Breast Cancer That Progressed on Endocrine Therapy - MONARCH 2: A Randomized Clinical Trial. JAMA Oncol 2020;6:116 - 124 is incorporated herein by reference in its entirety.
[0186] 2. Giuliano M, Schettini F, Rognoni C et al. Endocrine treatment versus chemotherapy in postmenopausal women with hormone receptor-positive, HER2-negative, metastatic breast cancer: a systematic review and network meta-analysis. Lancet Oncol 2019;20:1360-1369 is incorporated herein by reference in its entirety.
[0187] 3. Turner NC, Slamon DJ, Ro J et al. Overall Survival with Palbociclib and Fulvestrant in Advanced Breast Cancer. N Engl J Med 2018;379:1926-1936 is incorporated herein by reference in its entirety.
[0188] 4. Im SA, Lu YS, Bardia A et al. Overall Survival with Ribociclib plus Endocrine Therapy in Breast Cancer. N Engl J Med 2019;381:307-316 is incorporated herein by reference in its entirety.
[0189] 5. Slamon DJ, Neven P, Chia S et al. Overall Survival with Ribociclib plus Fulvestrant in Advanced Breast Cancer. N Engl J Med 2020;382:514-524 is incorporated herein by reference in its entirety.
[0190] 6. Turner NC, Kingston B, Kilburn LS et al. Circulating tumour DNA analysis to direct therapy in advanced breast cancer (plasmaMATCH): a multicentre, multicohort, phase 2a, platform trial. Lancet Oncol 2020;21:1296-1308 is incorporated by reference in its entirety.
[0191] 7. Alix-Panabieres C, Pantel K. Clinical Applications of Circulating Tumor Cells and Circulating Tumor DNA as Liquid Biopsy. Cancer Discov 2016;6:479-491 is incorporated by reference in its entirety.
[0192] 8. Davis AA, Jacob S, Gerratana L et al. Landscape of circulating tumour DNA in metastatic breast cancer. EBioMedicine 2020;58:102914 is incorporated by reference in its entirety.
[0193] 9. Wan JCM, Massie C, Garcia-Corbacho J et al. Liquid biopsies come of age.: towards the implementation of circulating tumour DNA. Nat Rev Cancer 2017;17: 223-238 is incorporated by reference in its entirety.
[0194] 10. O’Leary B, Cutts RJ, Liu Y et al. The Genetic Landscape and Clonal Evolution of Breast Cancer Resistance to Palbociclib plus Fulvestrant in The PALOMA-3 Trial. Cancer Discov 2018;8:1390-1403 is hereby incorporated by reference in its entirety.
[0195] 11. Ander F, Ciruelos E, Rubovszky G et al. Alpelisib for PIK3CA-Mutated, Hormone Receptor-Positive Advanced Breast Cancer. N Engl J Med 2019;380:1929-1940 is hereby incorporated by reference in its entirety.
[0196] 12. The FDA has approved alpelisib for metastatic breast cancer. May 24, 2019, https: / / www.fda.gov / drugs / resources-information-approved-drugs / fda-approves-alpelisib-metastatic-breast-cancer is hereby incorporated by reference in its entirety.
[0197] 13. Turner NC, Liu Y, Zhu Z et al. Cyclin E1 Expression and Palbociclib Efficacy in Previously Treated Hormone Receptor-Positive Metastatic Breast Cancer. J Clin Oncol 2019;37:1169-1178 is incorporated by reference in its entirety.
[0198] 14. Li Z, Razavi P, Li Q et al. Loss of the FAT1 Tumor Suppressor Promotes Resistance to CDK4 / 6 Inhibitors via the Hippo Pathway. Cancer Cell 2018;34:893-905 e898 is incorporated by reference in its entirety.
[0199] 15. Formisano L, Lu Y, Servetto A et al. Aberrant FGFR signaling mediates resistance to CDK4 / 6 inhibitors in ER breast cancer. Nat Commun 2019;10:1373 is incorporated by reference in its entirety.
[0200] 16. Turner N, Pearson A, Sharpe R et al. FGFR1 amplification drive endocrine therapy resistance and a therapeutic target in breast cancer. Cancer Res 2010;70:2085-2094 is incorporated by reference in its entirety.
[0201] 17. Costa C, Wang Y, Ly A et al. PTEN Loss Mediates Clinical Cross-Resistance to CDK4 / 6 and PI3Kalpha Inhibitors in Breast Cancer. Cancer Discov 2020;10:72-85 is incorporated by reference in its entirety.
[0202] 18. Condorelli R, Spring L, O’Shaughnessy J et al. Polyclonal RB1 mutations and acquired resistance to CDK 4 / 6 inhibitors in patients with metastatic breast cancer. Ann Oncol 2018;29:640-645 is incorporated by reference in its entirety.
[0203] 19. Fancello L, Gandini S, Pelicci PG, Mazzarella L. Tumor mutational burden quantification from targeted gene panels: major advancements and challenges. J Immunother Cancer 2019;7:183 is incorporated by reference in its entirety.
[0204] 20. Rizvi NA, Hellmann MD, Snyder A et al. Cancer immunology. Mutational landscape determine sensitivity to PD-1 blockade in non-small cell lung cancer. Science 2015;348:124-128 is incorporated by reference in its entirety.
[0205] 21. Snyder A, Makarov V, Merghoub T et al. Genetic basis for clinical response to CTLA-4 blockade in melanoma. N Engl J Med 2014;371:2189-2199 is incorporated by reference in its entirety.
[0206] 22. Alexandrov LB, Nik-Zainal S, Wedge DC et al. Signatures of mutational processs in human cancer. Nature 2013;500:415-421 is incorporated herein by reference in its entirety.
[0207] 23. Chalmers ZR, Connelly CF, Fabrizio D et al. Analysis of 100,000 human cancer genomes reveals the landscape of tumor mutational burden. Genome Med 2017;9:34 is incorporated herein by reference in its entirety.
[0208] 24. Yarchoan M, Albacker LA, Hopkins AC et al. PD-L1 expression and tumor mutational burden are independent biomarkers in most cancers. JCI Insight 2019;4 is incorporated herein by reference in its entirety.
[0209] 25. Gandara DR, Paul SM, Kowanetz M et al. Blood-based tumor mutational burdens as a predictor of clinical benefit in non-small-cell lung cancer patients with atezolizumab. Nat Med 2018;24:1441-1448 is incorporated herein by reference in its entirety.
[0210] 26. Wang Z, Duan J, Cai S, et al. Assessment of Blood Tumor Mutational Burden as a Potential Biomarker for Immunotherapy in Patients With Non-Small Cell Lung Cancer With Use of a Next-Generation Sequencing Cancer Gene Panel. JAMA Oncol 2019;5:696-702 is incorporated herein by reference in its entirety.
[0211] 27. Kim ES, Velcheti V, Mekhail T, et al. Blood-based tumor mutational burden as a biomarker for atezolizumab in non-small cell lung cancer: the phase 2 B-F1RST trial. Nat Med 2022 is incorporated herein by reference in its entirety.
[0212] 28. Bos MK, Angus L, Nasserinejad K, et al. Whole exome sequencing of cell-free DNA - A systematic review and Bayesian individual patient data meta-analysis. Cancer Treat Rev 2020;83:101951 is incorporated herein by reference in its entirety.
[0213] 29. Koeppel F, Blanchard S, Jovelet C, et al. Whole exome sequencing for determination of tumor mutation load in liquid biopsy from advanced cancer patients. PLoS One 2017;12:e0188174 is incorporated herein by reference in its entirety.
[0214] Barroso-Sousa R, Jain E, Cohen O, et al. Prevalence and mutational determinants of high tumor mutation burden in breast cancer. Ann Oncol 2020;31:387-394 is incorporated herein by reference in its entirety.
[0215] Wang R, Yang Y, Ye WW, et al. Case Report: Significant Response to Immune Checkpoint Inhibitor Camrelizumab in a Heavily Pretreated Advanced ER+ / HER2- Breast Cancer Patient With High Tumor Mutational Burden. Front Oncol 2020;10:588080 is incorporated herein by reference in its entirety.
[0216] Barroso-Sousa R, Li T, Reddy S, et al. Abstract GS2-10: Nimbus: A phase 2 trial of nivolumab plus ipilimumab for patients with hypermutated her2-negative metastatic breast cancer (MBC). Cancer Research. February 2022 is incorporated herein by reference in its entirety.
[0217] Wang S, Jia M, He Z, Liu XS. APOBEC3B and APOBEC mutational signature as potential predictive markers for immunotherapy response in non-small cell lung cancer. Oncogene 2018;37:3924-3936 is incorporated herein by reference in its entirety.
[0218] 34. Dietlein F, Weghorn D, Taylor-Weiner A et al. Identification of cancer driver genes based on nucleotide context. Nat Genet 2020;52:208-218 is hereby incorporated by reference in its entirety.
[0219] 35. Martinez-Jimenez F, Muinos F, Sentis I et al. A compendium of mutational cancer driver genes. Nat Rev Cancer 2020;20:555-572 is hereby incorporated by reference in its entirety.
[0220] 36. Dawson SJ, Tsui DW, Murtaza M et al. Analysis of circulating tumor DNA to monitor metastatic breast cancer. N Engl J Med 2013;368:1199-1209 is hereby incorporated by reference in its entirety.
[0221] 37. O’Leary B, Cutts RJ, Huang X et al. Circulating Tumor DNA Markers for Early Progression on Fulvestrant With or Without Palbociclib in ER Advanced Breast Cancer. J Natl Cancer Inst 2021;113:309-317 is hereby incorporated by reference in its entirety.
[0222] 38. Jacob S, Davis AA, Gerratana L, et al. The use of serial circulating tumor DNA (ctDNA) to detect resistance alterations in progressive metastatic breast cancer. Clin Cancer Res 2020 is incorporated by reference in its entirety.
[0223] 39. Davis AA, Iams WT, Chan D, et al. Early Assessment of Molecular Progression and Response by Whole-genome Circulating Tumor DNA in Advanced Solid Tumors. Mol Cancer Ther 2020;19:1486-1496 is incorporated by reference in its entirety.
[0224] 40. Jongbloed EM, Deger T, Sleijfer S, et al. A Systematic Review of the Use of Circulating Cell-Free DNA Dynamics to Monitor Response to Treatment in Metastatic Breast Cancer Patients. Cancers (Basel) 2021;13 is incorporated by reference in its entirety.
[0225] 41. Krishnamurthy J, Luo J, Suresh R, et al. A phase II trial of an alternative schedule of palbociclib and embedded serum TK1 analysis. NPJ Breast Cancer 2022;8:35 is incorporated by reference in its entirety.
[0226] 42. Cardoso F, Paluch-Shimon S, Senkus E et al. 5th ESO-ESMO International Consensus Guidelines for Advanced Breast Cancer (ABC 5). Ann Oncol 2020;31:1623-1649 is hereby incorporated by reference in its entirety.
[0227] 43. Alexandrov LB, Kim J, Haradhvala NJ et al. The repertoire of mutational signatures in human cancer. Nature 2020;578:94-101 is hereby incorporated by reference in its entirety.
[0228] 44. Granadillo Rodriguez M, Flath B, Chelico L. The interesting relationship between APOBEC3 deoxycytidine deaminases and cancer: a long road ahead. Open Biol 2020;10:200188 is hereby incorporated by reference in its entirety.
[0229] 45. Sanchez-Vega F, Mina M, Armenia J et al. Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell 2018;173:321-337 e310 is hereby incorporated by reference in its entirety.
[0230] 46. Samstein RM, Lee CH, Shoushtari AN et al. Tumor mutational load predicts survival after immunotherapy across multiple cancer types. Nat Genet 2019;51:202-206 is hereby incorporated by reference in its entirety.
[0231] 47. Knisbacher BA, Gerber D, Levanon EY. DNA Editing by APOBECs: A Genomic Preserver and Transformer. Trends Genet 2016;32:16 - 28 is incorporated by reference in its entirety.
[0232] 48. Miao D, Margolis CA, Vokes NI et al. Genomic correlates of response to immune checkpoint blockade in microsatellite - stable solid tumors. Nat Genet 2018;50:1271 - 1281 is incorporated by reference in its entirety.
[0233] 49. Marabelle A, Le DT, Ascierto PA et al. Efficacy of Pembrolizumab in Patients With Noncolorectal High Microsatellite Instability / Mismatch Repair - Deficient Cancer: Results From the Phase II KEYNOTE - 158 Study. J Clin Oncol 2020;38:1 - 10 is incorporated by reference in its entirety.
[0234] 50. Alva AS, Mangat PK, Garrett - Mayer E et al. Pembrolizumab in Patients With Metastatic Breast Cancer With High Tumor Mutational Burden: Results From the Targeted Agent and Profiling Utilization Registry (TAPUR) Study. J Clin Oncol 2021;39:2443 - 2451 is incorporated by reference in its entirety.
[0235] 51. Keenan TE, Guerriero JL, Barroso-Sousa R et al. Molecular correlates of response to eribulin and pembrolizumab in hormone receptor-positive metastatic breast cancer. Nat Commun 2021;12:5563 is incorporated by reference in its entirety.
[0236] 52. Deng J, Wang ES, Jenkins RW et al. CDK4 / 6 Inhibition Augments Antitumor Immunity by Enhancing T-cell Activation. Cancer Discov 2018;8:216 - 233 is incorporated by reference in its entirety.
[0237] 53. Asghar US, Kanani R, Roylance R, Mittnacht S. Systematic Review of Molecular Biomarkers Predictive of Resistance to CDK4 / 6 Inhibition in Metastatic Breast Cancer. JCO Precis Oncol 2022;6:e2100002 is incorporated by reference in its entirety.
[0238] 54. Daniels SR, Liyasova M, Kales SC et al. Loss of function Cbl-c mutations in solid tumors. PLoS One 2019;14:e0219143 is incorporated by reference in its entirety.
[0239] 55. Wang Y, Dai J, Zeng Y et al. E3 Ubiquitin Ligases in Breast Cancer Metastasis: A Systematic Review of Pathogenic Functions and Clinical Implications. Front Oncol 2021;11:752604 is incorporated herein by reference in its entirety.
[0240] 56. Xu L, Zhang Y, Qu X et al. E3 Ubiquitin Ligase Cbl-b Prevents Tumor Metastasis by Maintaining the Epithelial Phenotype in Multiple Drug-Resistant Gastric and Breast Cancer Cells. Neoplasia 2017;19:374-382 is incorporated herein by reference in its entirety.
[0241] 57. Toska E, Osmanbeyoglu HU, Castel P et al. PI3K pathway regulates ER-dependent transcription in breast cancer through the epigenetic regulator KMT2D. Science 2017;355:1324-1330 is incorporated herein by reference in its entirety.
[0242] 58. Gao SL, Yin R, Zhang LF et al. The oncogenic role of MUC12 in RCC progression depends on c-Jun / TGF-beta signalling. J Cell Mol Med 2020;24:8789-8802 is incorporated herein by reference in its entirety.
[0243] 59. Mukhopadhyay P, Chakraborty S, Ponnusamy MP et al. Mucins in the pathogenesis of breast cancer: implications in diagnosis, prognosis and therapy. Biochim Biophys Acta 2011;1815:224-240 is incorporated herein by reference in its entirety.
[0244] 60. van Putten JPM, Strijbis K. Transmembrane Mucins: Signaling Receptors at the Intersection of Inflammation and Cancer. J Innate Immun 2017;9:281-299 is incorporated herein by reference in its entirety.
[0245] 61. Mense SM, Barrows D, Hodakoski C et al. PTEN inhibits PREX2-catalyzed activation of RAC1 to restrain tumor cell invasion. Sci Signal 2015;8:ra32 is incorporated herein by reference in its entirety.
[0246] 62. Pandiella A, Montero JC. Molecular pathways: P-Rex in cancer. Clin Cancer Res 2013;19:4564-4569 is incorporated herein by reference in its entirety.
[0247] Supplementary Methods
[0248] Blood collection and cfDNA / gDNA extraction were performed as follows. Each single blood draw of 8 mL of whole blood was collected in a Streck tube, followed by two-step centrifugation to separate the plasma and buffy coat fractions. Aliquot samples were stored at -80 °C for batch processing. Cell-free DNA (cfDNA) was extracted from plasma samples using the QIAamp Circulating Nucleic Acid Kit (Qiagen, Hilden, Germany). The quantity and quality of the purified cfDNA were checked using a Qubit fluorometer (ThermoFisher Scientific, Waltham, Massachusetts, USA), and a Bioanalyzer 2100 (Agilent Technologies, California, USA). For cfDNA samples with significant genomic contamination from peripheral blood cells, bead-based size selection was performed to remove large genomic fragments (AMPure XP beads, Beckman Coulter, California, USA). Genomic DNA (gDNA) was extracted from matched peripheral blood mononuclear cells (PBMCs) using the QIAamp DNA Blood Mini Kit (Qiagen), then enzymatically fragmented and purified.
[0249] Library preparation, hybrid capture, and sequencing were performed as follows. 5 - 30 ng of extracted cfDNA or 30 - 50 ng of fragmented PBMC gDNA was processed for library construction, including end repair, dA tailing, and adapter ligation. Ligated library fragments with appropriate adapters were amplified by PCR. The amplified DNA libraries were then further checked using a Bioanalyzer 2100, and samples with sufficient yields were advanced to hybrid capture.
[0250] Hybrid capture was performed using biotin-labelled DNA probes. Briefly, each library was hybridized overnight with the Predicine NGS panel and paramagnetic beads. Unbound fragments were washed away and the enriched fragments were amplified by PCR. The purified product was checked on a Bioanalyzer 2100 and then loaded onto an Illumina NovaSeq 6000 (San Diego, CA, USA) for NGS sequencing using a paired-end 2x150bp sequencing kit.
[0251] Analysis of NGS data from cfDNA was performed as follows. NGS data from cfDNA were analyzed using the Predicine DeepSea NGS analysis pipeline, which started from raw sequencing data (BCL files) and output final variant calls. Briefly, the pipeline first performed adapter trimming, barcode checking, and correction. The cleaned paired FASTQ files were aligned to the human reference genome build hg19 using the BWA alignment tool. Next, a consensus bam file was derived by merging paired-end reads from the same molecule as the single-stranded fragment (based on mapping position and unique molecular identifier). Single-stranded fragments from the same double-stranded DNA molecule were further merged as double-stranded. By performing error suppression (as described, for example, in [Newman, 2016]), both sequencing and PCR errors were mostly corrected during this process.
[0252] Somatic mutations were identified as follows. Candidate variants were called by comparing them to a local variant background (defined based on plasma samples and historical data from healthy donors). Variants were further filtered by log odds (LOD) thresholds [Cibulskis 2013], base and mapping quality thresholds, repetitive regions, and other quality metrics. Generally, variants identified in cfDNA were considered candidates for somatic mutations only if (i) at least three distinct fragments (one of which must be double-stranded) contained the mutation, (ii) the variant allele frequency was higher than 0.25% or 0.1% for hotspot mutations, and (iii) the ctDNA variant fragments were significantly overrepresented compared to matched PBMC samples using Fisher's exact test.
[0253] Candidate somatic mutations were further filtered based on gene annotation to identify those occurring in protein-coding regions. Intronic and silent changes were excluded, but mutations resulting in missense, nonsense, frameshift, or splice-site changes were retained. Mutations annotated as benign or likely benign were also excluded based on the ClinVar database or as common germline variants in databases including 1000 Genomes, ExAC, gnomAD, and KAVIAR with population allele frequency > 0.5%. Finally, previously described hematopoietic expansion-related variants, including those in specific changes within DNMT3A, ASXL1, TET2, and ATM (residue 3008), GNAS (residues 201, 202), or JAK2 (residue 617), were marked as CHIP-related mutations.
[0254] Germline DNA analysis was performed as follows. Germline variants were determined by co-sequencing buffy coat PBMCs. Variants that were candidates with low base quality, mapping scores, and other poor quality metrics were filtered. Candidate variants with an allele frequency of less than 5% or having less than 8 distinct reads containing the variant were excluded. Unknown variants in repetitive regions were also excluded. Details of the analysis workflow are provided in "Analyses of NGS data generated from cfDNA" above.
[0255] Copy number analysis using a target panel was performed as follows. Copy number polymorphisms were estimated at the gene level. The pipeline calculates on-target unique fragment coverage based on the consensus bam file, first corrects for GC bias, and then adjusts for probe-level bias (estimated from the pooled reference). Each adjusted coverage profile is first self-normalized (assuming a diploid state for each sample) and then compared to the corresponding adjusted coverage from a group of normal reference samples to estimate the significance of copy number variants. To call gene amplifications or deletions, the absolute z-score and the copy number change need to exceed a minimum threshold.
[0256] Analysis of DNA rearrangements was performed as follows. DNA rearrangements were detected by identifying alignment breakpoints based on the bam file prior to the consensus step. Suspect alignments were filtered based on repetitive regions, local entropy calculations, and the similarity between the reference alignment and alternative alignments. To report a DNA fusion, more than 3 unique alignments (at least one of which should be in a double-stranded form) are required.
[0257] The analysis of the ctDNA fraction was performed as follows. The ctDNA fraction was estimated based on the allelic fraction of autosomal somatic mutations (e.g., as described in [Vandekerkhove, 2017]). Briefly, the mutant allelic fraction (MAF) and the ctDNA fraction are related as MAF = (ctDNA * 1) / [(1 - ctDNA) * 2 + ctDNA * 1], and thus ctDNA = 2 / ((1 / MAF) + 1). Somatic mutations in genes with detectable copy number changes were excluded from the ctDNA fraction estimation.
[0258] The bTMB score was estimated as follows. Blood-based tumor mutational burden (bTMB) was defined as the number of somatic coding SNVs, including synonymous and non-synonymous variants, within the panel target regions. The bTMB score was then normalized by the total effective target panel size within the coding region [Gandara, 2018]. Since TMB estimation considers all variants (including synonymous and non-whitelist variants), higher variant call specificity is required for TMB estimation. As a result, a more stringent cut-off was used for variant calling, and only variants with an allelic frequency ≥ 0.35% were used for TMB estimation. Samples with a maximum somatic allelic fraction (MSAF) < 0.7% were excluded from bTMB estimation. Variants in common CHIP genes (DNMT3A, TET2, ASXL1, and JAK2) were excluded in TMB estimation.
[0259] Copy number burden analysis using low-pass whole genome sequencing was performed as follows. Low-pass whole genome sequencing (LP-WGS) with an average coverage of 5× was performed on patient samples. The ichorCNA algorithm [Adalsteinsson, 2017] was applied to GC- and mappability-normalized reads to estimate plasma copy number variations using a hidden Markov model (HMM). First, the inventors measured the copy number deviation at the segment level (1 Mb genomic regions) as the log2 ratio of the normalized reads between the sample and the mean of a group of normal plasma samples as background. Next, the inventors quantified the arm-level CNV deviation as the mean of the segment CNVs across each chromosomal arm. Finally, the inventors calculated the sample-level copy number burden (CNB score) as the sum of the absolute z-scores of the arm-level CNV deviations, where higher / lower CNB scores indicate higher / lower CNV abnormalities compared to a normal background. A CNB score cutoff of 5.6 was defined as being 3 standard deviations away from the population mean of the normal plasma CNB scores.
[0260] The somatic mutation signature analysis was performed as follows. Somatic mutation analysis was carried out, and the pattern of single nucleotide substitutions (SBS) was compared with previously reported SBS signatures [Alexandrov, 2013, 2020] available in the COSMIC database using the maftools package (version 2.4.15) in R (version 3.6.3). Briefly, each of the 96 possible mutational substitution types is defined by one of six substitution types (T>A, T>C, T>G, C>A, C>G, C>T) and the bases immediately 5’ and 3’ of the mutated base. For each sample, the number of mutational substitutions was counted. Non-negative matrix factorization was used to decompose the count matrix into n signatures. The number of signatures (n) that best fit the data was estimated using Cophenetic correlation. The signatures were compared to 78 available SBS COSMIC signatures (v3.2 - March 2021, cancer.sanger.ac.uk / signatures / sbs / ) using cosine similarity. The dominant signature was defined as the signature with the maximum signature score in each sample.
[0261] The oncogenic signaling pathway analysis was performed as follows. To compare the relative proportions of mutations within important oncogenic signaling pathways between high bTMB patients and low bTMB patients, the inventors filtered a gene list [Sanchez-Vega, 2018] that describes oncogenic signaling pathways to include only those identified as breast cancer driver genes [Dietlein, 2020; Martinez-Jiminez, 2020]. The resulting gene list is shown in Table 7. The frequency of SNVs across these genes was compared between high bTMB patients and low bTMB patients, and statistical significance was evaluated using Fisher's exact test.
[0262]
Table 7
[0263] Reference
[0264] 1. Patel, P. G., et al., Preparation of Formalin-fixed Paraffin-embedded Tissue Cores for both RNA and DNA Extraction. J Vis Exp, 2016(114) is incorporated herein by reference in its entirety.
[0265] 2. Newman, A. M., et al., Integrated digital error suppression for improved detection of circulating tumor DNA. Nat Biotechnol, 2016.34(5): p. 547-555 is incorporated herein by reference in its entirety.
[0266] 3. Cibulskis, K., et al., Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat Biotechnol, 2013.31(3): p. 213-9 is incorporated herein by reference in its entirety.
[0267] 4. Vandekerkhove, G., et al., Circulating Tumor DNA Reveals Clinically Actionable Somatic Genome of Metastatic Bladder Cancer. Clin Cancer Res, 2017.23(21): p. 6487-6497 is incorporated herein by reference in its entirety.
[0268] 5. Gandara, D. R., et al., Blood-based tumor mutational burden as a predictor of clinical benefit in non-small-cell lung cancer patients treated with atezolizumab. Nat Med, 2018. 24(9): p. 1441-1448 is incorporated herein by reference in its entirety.
[0269] 6. Adalsteinsson, V. A., et al., Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat Commun, 2017. 8(1): p. 1324 is incorporated herein by reference in its entirety.
[0270] 7. Alexandrov, L. B., et al., The repertoire of mutational signatures in human cancer. Nature, 2020. 578(7793): p. 94-101 is incorporated herein by reference in its entirety.
[0271] 8. Alexandrov, L. B., et al., Signatures of mutational processes in human cancer. Nature, 2013. 500(7463): p. 415-21 is incorporated herein by reference in its entirety.
[0272] 9. Sanchez-Vega, F., et al., Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell, 2018. 173(2): p. 321-337 e10 is incorporated herein by reference in its entirety.
[0273] 10. Dietlein, F., et al., Identification of cancer driver genes based on nucleotide context. Nat Genet, 2020. 52(2): p. 208-218 is hereby incorporated by reference in its entirety.
[0274] 11. Martinez-Jimenez, F., et al., A compendium of mutational cancer driver genes. Nat Rev Cancer, 2020. 20(10): p. 555-572 is hereby incorporated by reference in its entirety.
[0275] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. Although described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, modifications, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative ratios described herein, which may vary depending on various conditions and variables. It should be understood that various alternative forms of the embodiments of the present invention described herein may be employed in practicing the present invention. Accordingly, it is intended that the present invention will further encompass any such alternatives, modifications, variations, or equivalents thereof. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. It is a method, (a) A step of assaying a cell-free deoxyribonucleic acid (cfDNA) molecule obtained from or derived from a biological sample from a subject, wherein the subject has cancer, has previously had cancer, or is suspected to have cancer, and the assay step includes sequencing at least a portion of the cfDNA molecule or its derivatives to generate a set of sequencing reads, wherein the sequencing includes at least one of whole exome sequencing (WES) and whole genome sequencing (WGS). (b) A step of determining at least one of the target tumor mutation load and copy number load, at least in part, based on processing the set of sequencing reads. Methods that include...
2. The method according to claim 1, wherein the biological sample is selected from the group consisting of plasma samples, serum samples, buffy coat samples, urine samples, saliva samples, tissue biopsy samples, pleural fluid samples, ascites samples, amniotic fluid samples, cerebrospinal fluid samples, lymph fluid samples, sweat samples, tear fluid samples, semen samples, derivatives thereof, and combinations thereof.
3. The method according to claim 2, wherein the biological sample is obtained from or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, another blood collection tube, or a circulating tumor cell (CTC) collection tube.
4. The method according to claim 1, further comprising the step of subjecting the biological sample to conditions sufficient to isolate, concentrate, or extract the cfDNA molecule.
5. The method according to claim 1, further comprising the step of fractionating the whole blood sample of the subject in order to obtain the cfDNA molecule.
6. The method according to claim 1, wherein the WGS further comprises a low-pass WGS.
7. The method according to claim 1, wherein the sequence determination in (a) further comprises next-generation sequencing, low-pass sequencing, target sequencing, methylation-recognition sequencing, bisulfite sequencing, or a combination thereof.
8. The method according to claim 1, further comprising (a) amplifying at least a portion of the cfDNA molecule, genomic DNA (gDNA), or a derivative thereof.
9. The method according to claim 8, wherein the amplification further comprises polymerase chain reaction (PCR) or isothermal amplification.
10. The method according to claim 1, further comprising using a microarray in (a).
11. The method according to claim 1, wherein the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, liver cancer, and combinations thereof.
12. The method according to claim 11, wherein the cancer includes breast cancer.
13. The method according to claim 12, wherein the breast cancer is metastatic breast cancer, hormone receptor-positive (HR+) breast cancer, HER2-negative (HER2-) breast cancer, or a combination thereof.
14. The method according to claim 1, wherein the processing in (b) further comprises using a trained machine learning algorithm.
15. The method according to claim 14, wherein the trained machine learning algorithm further comprises a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.
16. The method according to claim 1, further comprising the step of determining the recurrence or relapse of the cancer in the subject, the cancer's resistance to drug treatment, and the prognosis of the cancer in the subject, based at least in part on at least one of the tumor mutation load and the copy number load of the subject.
17. The method according to claim 16, wherein the prognosis includes the possibility of progression-free survival, the length of progression-free survival, the possibility of overall survival, the length of overall survival, or a combination thereof.
18. The method of claim 1, wherein the biological sample is obtained from or derived from the subject (i) before the subject receives the clinical intervention for the cancer, (ii) while the subject is receiving the clinical intervention for the cancer, (iii) after the subject has received the clinical intervention for the cancer, or a combination thereof.
19. The method according to claim 18, wherein the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deficiency therapy, and combinations thereof.
20. The method according to claim 1, further comprising the step of providing a clinical intervention to the subject based at least partially on at least one of the tumor mutation load and the copy number load of the subject.
21. The method according to claim 20, wherein the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiotherapy, immunotherapy, endocrine therapy, adjuvant therapy, neoadjuvant therapy, androgen deficiency therapy, and combinations thereof.
22. The method according to claim 1, wherein the set of sequencing reads includes a quantitative measure of a set of cancer-related genomic loci.
23. The method according to claim 22, wherein the set of cancer-related genomic loci includes one or more members selected from the group consisting of the genes listed in Table 3, the genes listed in Table 4, the genes listed in Table 6, and the genes listed in Table 7.
24. The method according to claim 1, further comprising the step of using a nucleic acid primer or probe configured to selectively concentrate the biological sample with respect to DNA molecules corresponding to a set of genomic loci.
25. The method according to claim 1, further comprising the step of monitoring at least one of the tumor mutation load and the copy number load of the subject, wherein the monitoring step comprises evaluating at least one of the tumor mutation load and the copy number load of the subject at each of a plurality of time points.
26. The method according to claim 1, further comprising the processing in (b) detecting tumor-related changes selected from the group consisting of copy number variations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements.
27. A step of filtering at least a subset of the set of sequencing reads based on a quality score, A step of performing error correction on the set of sequencing reads using a sample barcode or molecular barcode attached to at least one of the cfDNA molecules, or The method according to claim 1, further comprising the step of performing at least one of single-strand consensus calling and double-strand consensus calling on the set of sequencing reads, thereby suppressing sequencing and PCR errors in the set of sequencing reads.
28. The method according to claim 1, further comprising the step of determining the frequency of mutated alleles in a set of somatic mutations.
29. A system comprising one or more computer processors and computer memory connected thereto, wherein the computer memory, when executed by the one or more computer processors, (a) A step of obtaining or extracting a biological sample from a subject, wherein the subject has cancer, has had cancer in the past, or is suspected to have cancer, (b) A step of assaying a cell-free deoxyribonucleic acid (cfDNA) molecule obtained from or derived from the biological sample, wherein the assay step includes sequencing at least a portion of the cfDNA molecule or its derivatives to generate a set of sequencing reads, and the sequencing includes at least one of whole exome sequencing (WES) and whole genome sequencing (WGS), (c) A step of determining at least one of the target tumor mutation load and copy number load, at least in part, based on processing the set of sequencing reads. A computer memory containing machine-executable code that implements the method, and A system that includes this.
30. A non-temporary computer-readable medium, which, when executed by one or more computer processors, (a) A step of obtaining or extracting a biological sample from a subject, wherein the subject has cancer, has had cancer in the past, or is suspected to have cancer, (b) A step of assaying a cell-free deoxyribonucleic acid (cfDNA) molecule obtained from or derived from the biological sample, wherein the assay step includes sequencing at least a portion of the cfDNA molecule or its derivatives to generate a set of sequencing reads, and the sequencing includes at least one of whole exome sequencing (WES) and whole genome sequencing (WGS), (c) A step of determining at least one of the target tumor mutation load and copy number load, at least in part, based on processing the set of sequencing reads. A non-temporary computer-readable medium comprising machine-executable code that implements a method including the following.