Methods of cancer detection by discordant methylation in cfdna
By analyzing discordant methylation patterns in cfDNA, the method effectively distinguishes between benign and malignant plasma cell disorders and predicts multiple myeloma progression, enhancing cancer detection and monitoring without invasive procedures.
Patent Information
- Application Number
- PCT/IL2025/050282
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-02
AI Technical Summary
Current methods for detecting cancer through liquid biopsies struggle to distinguish between benign and malignant plasma cell disorders, such as MGUS and MM, and lack effective markers for predicting progression, relying on invasive procedures and insufficient blood tests.
A method involving the analysis of methylation patterns in cell-free DNA (cfDNA) by quantifying discordant methylation reads at specific genomic loci to detect and quantify plasma cell DNA, using deep sequencing to identify at least two methylated and two unmethylated sites within a 300-base pair sequence, enabling differentiation between benign and malignant forms and predicting progression to multiple myeloma.
This approach allows for non-invasive, accurate detection and prediction of multiple myeloma by identifying specific methylation signatures in cfDNA, reducing the need for invasive procedures and improving early detection and monitoring.
Smart Images

Figure IL2025050282_02102025_PF_FP_ABST
Abstract
Description
METHODS OF CANCER DETECTION BY DISCORDANT METHYLATION INCFDNACROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 569,750 filed on March 26, 2024, the contents of which are all incorporated herein by reference in their entirety.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (HUJI-HDST-P-0118-PCT.xml; Size: 47,610 bytes; and Date of Creation: March 12, 2025) is herein incorporated by reference in its entirety.FIELD OF INVENTION
[0003] The present invention is in the field of cancer detection by liquid biopsy.BACKGROUND OF THE INVENTION
[0004] Detecting and monitoring cancer poses a significant clinical challenge as cancer cells can be hidden, remote or inaccessible. Liquid biopsies are a common and evolving practice in cancer surveillance. A notable epigenetic marker for cancer is DNA methylation disruption. When cancer cells die, they release small fragments of DNA to the bloodstream called tumor derived cell free DNA (cfDNA). Epigenetic liquid biopsies have emerged as a promising avenue for cancer detection and monitoring, particularly using cfDNA methylation patterns. However, distinguishing cfDNA originating from healthy cells versus tumors is a challenge. Changes in DNA methylation patterns are hallmarks of cancer, known for their stochastic nature, but typically vary among individual tumors and therefore cannot be used as a universal marker.
[0005] Plasma cell disorders (PCD) encompass a range of monoclonal plasma cell proliferation disorders, most of which are premalignant and asymptomatic. A small fraction(10% within 10 years) of patients with monoclonal gammopathy of undetermined significance (MGUS) and a significant fraction (50% within 5 years) of patients with smoldering multiple myeloma (SMM) will progress to the highly malignant and morbid form of multiple myeloma (MM). In addition, almost all patients with MM will have MGUS and SMM prior to their active disease. Moreover, MGUS is very abundant (1-10% of the adult population, respective to age). The discrimination between benign forms (MGUS and SMM) and malignant forms (AL amyloidosis and MM) is often difficult, as routine serum / urine tests for monoclonal protein are insufficient, as the plasma cells in all the PCD forms secrete monoclonal proteins. To date, there is no blood test that can differentiate among the PCD, and at present, MGUS and SMM patients are monitored using a combination of blood markers (monoclonal protein and free light chain ratio), imaging and laborious and invasive measures such as bone marrow biopsies. Foreseeing progression from benign to malignant is important as it allows physicians to focus on patients at risk and start life-saving treatments early and preserve target organ function; however, prediction or early detection of progression have proven to be challenging and currently there are no specific markers for prediction. Additionally, longitudinal assessments are the mainstay of monitoring for progression, and while clinical risk models are of aid, their applicability is limited, and prompt more frequent monitoring in clinical routine are sometimes necessitated. Designing and finding a simple peripheral blood (PB) related, accessible pattern to discriminate among benign and malignant forms of PCD, and furthermore to predict patients who will progress to the malignant MM is of high importance.
[0006] DNA methylation patterns are a unique characteristic of each cell type, controlling gene expression, and can serve as a definitive biomarker for the presence of DNA derived from a given cell type. This is the basis for epigenetic liquid biopsies that detect the tissue origins of circulating cell-free DNA (cfDNA) as an indication of increased cell turnover, which is often associated with pathology. A comprehensive atlas of human DNA methylation was recently published, allowing for the identification of methylation biomarkers unique to major human cell types (Loyfer et al., 2023, “A DNA Methylation Atlas of Normal Human Cell Types.” Nature 2023 1-10, the contents of which are hereby incorporated by reference in their entirety).
[0007] Beyond the programmed, stable patterns of DNA methylation characteristic of normal differentiated cells, the methylome undergoes specific changes in cancer, which may contribute to tumor evolution. One universal feature of DNA methylation in cancer is theloss of order: genomic loci containing multiple cytosine-guanosine (CpGs) sites that normally are homogenously methylated or demethylated, tend to show a disordered methylation pattern. Such disorder can be quantified using Shannon's entropy (Jenkinson et al., 2017, “Potential Energy Landscapes Identify the Information-Theoretic Nature of the Epigenome.” Nature Genetics 2017 49:5 49(5):719-29; Koldobskiy et al., 2020, “A Dysregulated DNA Methylation Landscape Linked to Gene Expression in MLL-Rearranged AML.” Epigenetics 15(8): 841—58), termed the proportion of discordant reads (PDR) (Landau et al., 2014, “Locally Disordered Methylation Forms the Basis of Intratumor Methylome Variation in Chronic Lymphocytic Leukemia.” Cancer Cell 26(6): 813-25), and epi-polymorphism (Landan et al., 2012, "Epigenetic polymorphism and the stochastic formation of differentially methylated regions in normal and cancerous tissues.” Nat Genet Nov;44(l 1): 1207-14). Genome-wide assessment of disordered methylation has been applied to biopsy cellular material of solid and hematological tumors (leukemia) and shown to correlate with disease progression and outcomes (Derrien et al., 2021, “The DNA Methylation Landscape of Multiple Myeloma Shows Extensive Inter- and Intrapatient Heterogeneity That Fuels Transcriptomic Variability.” Genome Medicine 13(1):1— 21; Landau et al., 2014). Its application to liquid biopsies remains challenging, given the low fraction of tumor DNA in typical cfDNA samples and the low coverage of whole genome bisulfite sequencing (WGBS). Indeed, methods for using these phenomena have focused on genomic DNA derived directly from the tumors. In addition, these methods used whole - methylome analysis to infer the degree of disorder typical of cancer, thus necessitating significant sequencing resources, which limits applicability. A new accurate method for cancer detection and monitoring based on discordant methylation, that does not require biopsy, is therefore greatly needed.SUMMARY OF THE INVENTION
[0008] The present invention provides methods of detecting cancer in a subject in need thereof, comprising ascertaining the methylation status of at least four methylation sites in the same double- stranded cell-free DNA (cfDNA) molecule and detecting the presence of at least two sites that are methylated and at least two sites that are unmethylated in the at least four methylation sites indicating that the subject suffers from cancer. Methods of quantifying molecules of cfDNA and also provided, as are methods of detecting and quantifying plasmacell DNA. Methods of diagnosing multiple myeloma or predicting progression of smoldering multiple myeloma (SMM) or monoclonal gammopathy of undetermined significance (MGUS) to multiple myeloma in a subject are also provided.
[0009] According to a first aspect, there is provided a method of detecting multiple myeloma (MM) in a subject in need thereof, the method comprising a. receiving a fluid sample from the subject comprising cell free DNA (cfDNA); b. quantifying a number of cfDNA with discordant methylation reads in the sample, wherein the quantifying comprises ascertaining the methylation status of at least four methylation sites in the same double- stranded cell- free DNA (cfDNA) molecule, wherein the at least four methylation sites are in a contiguous nucleotide sequence which comprises no more than 300 base pairs; and wherein the presence of at least two sites that are methylated and at least two sites that are unmethylated in the at least four methylation sites indicating a cfDNA with discordant methylation; and c. comparing the percentage of cfDNA molecule in the sample with discordant methylation reads with the percentage of cfDNA molecules with discordant methylation reads in a control sample, wherein a higher percentage of discordant methylation reads in the received sample as compared to the control sample indicates the subject suffers from MM; thereby detecting MM in a subject.
[0010] According to some embodiments, the contiguous nucleotide sequence is a genomic locus which is completely methylated or completely unmethylated in non-cancerous cells.[Oi l] According to some embodiments, the contiguous nucleotide sequence is a genomic locus selected from the group consisting of: chrl5:101, 168, 755-101, 168, 812, chr7:134, 597, 197-134, 597, 250, chrl5:91,l 19,606-91,119,642, chr2:85,005,386-85,005,443, chrl5:41, 072, 702-41, 072, 749, chr7:23, 150, 939-23, 151, 010, chr5:81,968,704- 81,968,773, chrll:5, 225, 201-5, 225, 231, chrll:107, 747, 505-107, 747, 531 and chr6:80, 356, 706-80, 356, 741, wherein the coordinates are with respect to human genome build HG19.
[0012] According to some embodiments, the contiguous nucleotide sequence is a sequence selected from the group consisting of SEQ ID NO: 8-17.
[0013] According to some embodiments, the detecting the presence of at least two sites comprises detecting all cfDNA molecules comprising the genomic locus and determining the percentage of the cfDNA molecules comprising the genomic locus that comprise at least two sites that are methylated and at least two sites that are unmethylated thereby determining a percentage of discordant reads (PDR) for the genomic locus.
[0014] According to some embodiments, the method comprises determining a PDR for at least two of the genomic loci, calculating an average PDR for the at least two of the genomic loci and wherein an average PDR above a predetermined threshold indicates the subject suffers from cancer.
[0015] According to some embodiments, the at least two genomic loci are all ten of the genomic loci.
[0016] According to some embodiments, the ascertaining is affected by contacting the cfDNA molecule in the sample with bisulfite to generate single- stranded DNA molecules of which demethylated cytosines of the single-stranded DNA molecules are converted to uracils, further contacting the single-stranded DNA with amplification primers under conditions that generate amplified DNA from the single- stranded DNA following the contacting with the bisulfite and sequencing the amplified DNA.
[0017] According to some embodiments, the sequencing is deep sequencing, massively parallel sequence or next-generation sequencing.
[0018] According to some embodiments, the control sample is a sample from a healthy subject, or a sample from a subject suffering from smoldering multiple myeloma (SMM) or monoclonal gammopathy of undetermined significance (MGUS).
[0019] According to some embodiments, the control sample is a sample from a subject suffering from SMM or MGUS and the method is a method of diagnosing or predicting progression to MM in subject with SMM or MGUS.
[0020] According to another aspect, there is provided a method for quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a multiple myeloma (MM) cell in a fluid sample derived from a subject, comprising: a. identifying a DNA sequence of 10-300 base pairs that has at least four methylation sites, wherein the methylation pattern of the at least four methylation sites in a MM cell is different as compared to the methylationpattern of the at least four methylation sites in the same DNA sequence in non-MM cells, wherein the methylation pattern in a MM cell comprises at least two sites that are methylated and at least two sites that are unmethylated and the methylation pattern in non-MM cells comprises all sites being methylated or all sites being unmethylated; b. amplifying the identified DNA sequence in the cfDNA of the sample to obtain amplified DNA molecules; c. using deep sequencing, massively parallel sequencing or next-generation sequencing to obtain the sequence of individual molecules of the amplified DNA molecules; d. ascertaining from the sequenced amplified DNA molecules the methylation status of each of the at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecules of the sequenced amplified DNA molecules of the sample; and e. quantifying the number of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a MM cell that are present in the total of the amplified DNA molecules; thereby quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a MM cell.
[0021] According to some embodiments, the DNA sequence is a genomic locus selected from the group consisting of: chrl5:101, 168, 755-101, 168, 812, chr7: 134,597, 197- 134,597,250, chrl5:91,l 19,606-91,119,642, chr2:85, 005, 386-85, 005, 443, chrl5:41, 072, 702-41, 072, 749, chr7:23, 150, 939-23, 151, 010, chr5:81, 968, 704-81, 968, 773, chrll:5, 225, 201-5, 225, 231, chrll:107, 747, 505-107 ,747, 531 and chr6:80,356,706- 80,356,741, wherein the coordinates are with respect to human genome build Hgl9.
[0022] According to some embodiments, the DNA sequence is a sequence selected from the group consisting of SEQ ID NO: 8-17.
[0023] According to some embodiments, the fluid sample is selected from the group consisting of blood, plasma, sperm, milk, urine, saliva and cerebral spinal fluid.
[0024] According to some embodiments, the fluid sample is a plasma sample.
[0025] According to some embodiments, the method further comprises contacting the DNA of the sample with bisulfite to convert demethylated cytosines of the DNA to uracils prior to the amplifying step (b).
[0026] According to some embodiments, the sequence is no more than 167 base pairs.
[0027] According to some embodiments, in step (c), the sequences of at least 10,000 individual molecules are obtained.
[0028] According to some embodiments, the MM cells are selected from SMM cells and MGUS cells.
[0029] According to some embodiments, the method further comprises administering an anti-MM drug to a subject detected to have cancer.
[0030] According to another aspect, there is provided a method of detecting DNA from a plasma cell in a sample, the method comprising: a. receiving DNA methylation measurements of DNA from a sample in at least one genomic region selected from the group consisting of: chr2: 128090328-128090485, chr6:53139872-53139964, chr6: 109220925- 109220980, chr20:24995553-24995799, chr6:7880318-7880470, chrl3:113246637-113246717, and chrl9: 14522665- 14522768, wherein the coordinates are with respect to human genome build Hgl9; and b. assigning a DNA molecule as being from a plasma cell when the genomic region comprises unmethylation of all CpGs; thereby detecting DNA from a plasma cell is a sample comprising DNA.
[0031] According to another aspect, there is provided a method for quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a plasma cell in a sample derived from a subject, comprising: a. amplifying a DNA sequence selected from the group consisting of: chr2: 128090328-128090485, chr6:53139872-53139964, chr6: 109220925- 109220980, chr20:24995553-24995799, chr6:7880318-7880470, chrl3:113246637-113246717, and chrl9: 14522665- 14522768, wherein the coordinates are with respect to human genome build Hgl9, in the cfDNA of the sample to obtain amplified DNA molecules;b. using deep sequencing, massively parallel sequencing or next-generation sequencing to obtain the sequence of individual molecules of the amplified DNA molecules; c. ascertaining from the sequenced amplified DNA molecules the methylation status of at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecules of the sequenced amplified DNA molecules of the sample; and d. quantifying the amount of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a plasma cell that is present in the total of the amplified DNA molecules, wherein the methylation pattern in a plasma cell comprises all of the at least four sites being unmethylated; thereby quantifying molecule of cell-free DNA (cfDNA) having the methylation pattern of a plasma cell.
[0032] According to some embodiments, the at least one genomic region or the DNA sequence is selected from the group consisting of: SEQ ID NO: 1-7.
[0033] According to some embodiments, the at least one genomic region is all seven regions.
[0034] According to another aspect, there is provided a method of diagnosing multiple myeloma or predicting progression to multiple myeloma from SMM or MGUS in a subject in need thereof, the method comprising: a. receiving DNA methylation measurements of cfDNA from a sample from the subject; and b. quantifying the number of cfDNA molecules from a plasma cell in the sample by a method of the invention, wherein a quantity of cfDNA molecules from a plasma cell in the sample above a predetermined threshold indicates the subject suffers from multiple myeloma; thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject.
[0035] According to another aspect, there is provided a method of predicting progression to multiple myeloma in a subject suffering from smoldering multiple myeloma (SMM) or monoclonal gammopathy of undetermined significance (MGUS), the method comprising:a. receiving DNA methylation measurements of cfDNA from a sample from the subject; and b. quantifying the number of cfDNA molecules from a plasma cell in the sample by a method of the invention, wherein a quantity of cfDNA molecules from a plasma cell in the sample above a predetermined threshold indicates the subject is predicted to progress to multiple myeloma; thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
[0036] According to another aspect, there is provided a method of diagnosing multiple myeloma or predicting progression to multiple myeloma from SMM or MGUS in a subject in need thereof, the method comprising receiving measurements of 6 features in the fluid sample from the subject, wherein the 6 features are: a. the number of cfDNA molecules from a plasma cell in the sample; b. the percentage of discordant methylation reads (PDR) at genomic location chrl 1:107, 747, 505- 107 ,747, 531 wherein the coordinates are with respect to human genome build HG19; c. the entropy of the methylation at genomic location chrl 1:107,747,505- 107,747,531 wherein the coordinates are with respect to human genome build HG19; d. the entropy of the methylation at genomic location chr6:80,356,706- 80,356,741 wherein the coordinates are with respect to human genome build HG19; e. the entropy of the methylation at genomic location chrl5:41,072,702- 41,072,749 wherein the coordinates are with respect to human genome build HG19; and f. whether the free light chain (FLC) ratio is greater than or less than 20; applying a trained machine learning model to the 6 features, wherein the trained machine learning model is trained on the 6 features from subjects with MM and the 6 features from subjects with SMM, subjects with MGUS or both, and outputs a diagnosis of multiple myeloma or not or a diagnosis of progression to multiple myeloma or not;thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
[0037] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figures 1A-1C: Identification of specific plasma cell DNA methylation markers. (1A) A heat map representing methylation states across 38 cell types based on Whole- Genome Bisulfite Sequencing (WGBS). Blue denotes methylation, yellow indicates unmethylation. (IB) Bar graph of 7 targeted specific plasma cell methylation markers, assessed on genomic DNA from multiple tissues, including memory B-cells and plasma cells from normal bone marrow, MGUS and MM. (1C) Spike-in experiment assessing assay sensitivity. Plasma cell DNA was spiked in leukocyte DNA in decreasing quantities. X-axis denotes the percentage of plasma cell DNA mixed within leukocytes. Y-axis is the average (%) of plasma DNA measured by plasma cell methylation-based markers.
[0039] Figures 2A-2J: Plasma cell derived cfDNA is elevated in multiple myeloma compared to smoldering myeloma and MGUS. (2A) Correlation between the percentage of methylation-based plasma cell markers and flow cytometry of plasma cells from bone marrow (Pearson r=0.99, P<0.0001). (2B) Dot plot comparative analysis of plasma-derived cell-free DNA (cfDNA) levels across healthy controls (N=86), MGUS (N=54), SMM (N=28) and MM (N=36) (p-value < 0.0001, Kruskal-Wallis). (2C) Receiver operating characteristic (ROC) curves for discriminating multiple myeloma compared to healthy (left), MGUS middle and SMM (right) using the concentration of plasma cell-derived cfDNA levels within patient plasma samples (GE / ml). (2D) Correlation between the percentage of methylation-based plasma cell markers in cfDNA of plasma samples from MGUS, SMM and MM and flow cytometry of plasma cells from bone marrow. (2E) Correlation between % of aberrant plasma cells i n the bone marrow determined by flow cytometry (Spearman’s r=0.68, P<0.0001). (2F) Correlation between the percentage of methylation-based plasmacell markers and pathology estimations of plasma cell percentages in the bone marrow (Pearson r=0.43, P=049). (2G) Correlation between pathology estimations of plasma cell percentages in the bone marrow and flow cytometry of plasma cells in the bone marrow (Pearson r=0.4, P=073). (2H) Comparative analysis of total cfDNA concentration (ng / ml) across healthy controls (N=86), MGUS (N=54), SMM (N=28) and MM (N=36) (p- value<0.0001, Kruskal-Wallis). (21) Comparative analysis of plasma cell derived cfDNA (%) (p-value<0.0001, Kruskal- Wallis). (2J) Receiver operating characteristic (ROC) curve for discriminating multiple myeloma compared to healthy (left) MGUS (middle) and SMM (right) using the percentage of plasma cell-derived cfDNA levels within patient plasma samples.
[0040] Figures 3A-3M: Targeted cancer specific genomic loci of local methylation discordance. (3A) Cartoon showing the proportion of discordant reads as proposed by Landau et al., 2014 as all epialleles that are not fully methylated or fully unmethylated (left). The upgraded Proportion of Discordant Reads (PDR) metric, used herein for cfDNA, includes only epialleles that show inconsistency in two or more CpG sites, considering the occasional variation at a single CpG as a normal biological variation. Black circles- methylated CpGs; white circles-non methylated CpGs. (3B) A workflow for finding specific PDR regions from deep WGBS data and designing targeted markers accordingly. (3C) A heat map representing methylation discordant reads across 26 cell types based on Whole- Genome Bisulfite Sequencing (WGBS) in 30X depth. Normal cfDNA as well as MGUS, SMM and MM plasma cells are also included. Blue denotes low PDR orange indicates high PDR. (3D) Bar graph comparing PDR in 10 specific genomic loci and 12 other cell types including normal plasm cells. (3E) Comparative analysis of PDR markers in cfDNA across healthy controls (N=60), MGUS (N=55), SMM (N=28) and MM (N=38). (3F) Correlation between PDR markers and aberrant plasma cells in bone marrow indicated by flow cytometry. (3G) Receiver operating characteristic (ROC) curve for discriminating multiple myeloma compared to MGUS (left) and SMM (right) using targeted PDR markers in cfDNA from patients' plasma samples. (3H) PDR specific markers were compared between DNA extracted from whole bone marrow aspirations of healthy controls, MGUS, SMM and MM. (31) Spike-in experiment assessing PDR assay sensitivity. Myeloma plasma cell DNA from MM patient was spiked in leukocyte DNA in decreasing quantities. X-axis denotes the percentage of plasma cell DNA mixed within leukocytes. Y-axis is the fraction of PDR measured by positive PDR methylation-based markers. Y-axis begins at the value that isestablished by the blank sample (0%). (3J) Comparative analysis of each single PDR marker in cfDNA of MGUS (N=55), SMM (N=28) and MM (N=38). (3K) Cumulative bar plots of all 10 PDR markers in MGUS, SMM, and MM, with each bar representing one individual. (3L-3M) Receiver operating characteristic (ROC) curves for discriminating (3L) multiple myeloma compared to MGUS and (3M) multiple myeloma compared to SMM using individual PDR markers in cfDNA.
[0041] Figures 4A-4P: Plasma cell-derived cfDNA and PDR in cfDNA predicts progression to multiple myeloma. (4A-4C) Kaplan-Meier plot illustrating the correlation between clinical progression to multiple myeloma and (4A) the IMWG “2 / 20 / 20” risk stratification model, (4B) plasma cell cfDNA levels (GE / ml), and (4C) PDR fraction. (4D- E) XY scatter plots of clinically progressed patients with MGUS and SMM. Showing the correlation between the levels of (4D) plasma cell-derived cfDNA (GE / ml) and (4E) PDR fraction and the time it took to progress from the sample collection (time 0, Y axis). (4F-4G) Kaplan-Meier plot illustrating a new scoring model where bone marrow aspiration is substituted with (4F) PC-cfDNA or (4G) PDR. (4H-4J) Kaplan-Meier plot illustrating the correlation between biochemical progression and (4H) the IMWG “2 / 20 / 20” risk stratification model, (41) plasma cell cfDNA levels (GE / ml), and (4J) PDR fraction. (4K- 4L) XY scatter plots of biochemically progressed patients with MGUS and SMM. Showing the correlation between the levels of (4K) plasma cell-derived cfDNA (GE / ml) and (4L) PDR fraction and the time it took to progress from the sample collection (time 0, Y axis). (4M-4N) Kaplan-Meier plot illustrating a new scoring model where bone marrow aspiration is substituted with (4M) PC-cfDNA or (4N) PDR. (4O-P) Scatter plots of (40) PC cfDNA and (4P) PDR cfDNA of two groups -patients that biochemically progressed within 24 months and those that did not (p=0.067, p=0.045, Mann-Whitney).
[0042] Figures 5A-5G: Entropy increases in multiple myeloma and serves as a predictor for progression-indicating clonal variation and turnover. (5A) An illustrated scheme depicting the relationship between entropy and PDR, the left panel portrays normal healthy plasma cells. Regions prone to PDR in multiple myeloma are predominantly fully methylated (depicted as black circles), resulting in minimal discordant reads and low entropy levels. As cancer develops (illustrated in the middle and right panels), methylation disturbances become more prevalent. If a dominant epiclone emerges, both PDR and entropy levels rise in similar proportions. However, in scenarios where multiple clones exist, the dissonance between PDR and entropy is significantly amplified. (5B) Comparative analysisof the entropy of PDR markers in cfDNA across healthy controls (N=60), MGUS (N=55), SMM (N=28) and MM (N=38). Right image illustrates clonal diversity with two samples exhibiting levels of entropy in the extremities of the scale. Each color represents a different epiclone. (5C) Receiver operating characteristic (ROC) curve for discriminating multiple myeloma compared to healthy (left), MGUS (middle), and SMM (right) using entropy levels in cfDNA. (5D) Kaplan-Meier plots illustrating the correlation between entropy in cfDNA and biochemical and clinical progression. (5E) A correlation matrix between PDR and entropy measurements in the BM and cfDNA (Spearmen's). (5F) Heatmap of Spearman’s correlation coefficients for each "epiclone" fraction in the bone marrow relative to cfDNA. The heatmap is ordered by descending average correlation per patient. (5G) Schematic representation of a specific genomic locus, illustrating the distribution of different epiclone fractions in the bone marrow compared to cfDNA from the same patient. Left panel: MM 52; Right panel: MM 98.
[0043] Figures 6A-6G: Validation of performance of plasma cell and PDR markers using predictive modeling and an independent cohort. Shapley analysis of cfDNA and clinical features. Evaluation of the contribution of each feature to the model’s prediction of time to progression of MM (only top 10 shown). (6A) Shapley value distributions (6B), The average absolute SHAP value for each individual feature. (6C), Repeated 5-fold cross- validation results on the best feature set. Bar plot of metrics for cfDNA features (Dark purple; 'Entropy-SH3BGRL2', 'Entropy-DNAJC17', 'Entropy-PDR9', 'PDR-PDR9', 'plasma-cell- cfDNA-%’), clinical laboratory features (Orange, Risk by IMWG), combined (orange & purple stripes; clinical + cell-free DNA) and all non-invasive measures (Light purple; cell- free + ‘M-Protein’ + ’FLC ratio’) (6D) An independent validation cohort comprising 16 MGUS patients, 11 SMM patients, and 18 MM patients. Comparative analysis of plasma- derived cell-free DNA (cfDNA) levels MGUS (N=16), SMM (N=l l) and MM (N=18) (p- value<0.0001, Kruskal-Wallis). (6E) Receiver operating characteristic (ROC) curve for discriminating multiple myeloma compared to MGUS (left panel) and SMM (right panel) using the concentration of plasma cell-derived cfDNA levels within patient plasma samples (GE / ml). (6F) Comparative analysis of PDR markers in cfDNA of patients with MGUS, SMM and MM (p-value=0.005, Kruskal-Wallis). (6G) Receiver operating characteristic (ROC) curve for discriminating multiple myeloma compared to MGUS (left) and SMM (right) using targeted PDR markers in cfDNA from patients' plasma samples.DETAILED DESCRIPTION OF THE INVENTION
[0044] The present invention, in some embodiments, provides methods of detecting cancer in a subject in need thereof, comprising ascertaining the methylation status of at least four methylation sites in the same double- stranded cell-free DNA (cfDNA) molecule and detecting the presence of at least two sites that are methylated and at least two sites that are unmethylated in the at least four methylation sites indicating that the subject suffers from cancer. Methods of quantifying molecules of cfDNA having a methylation pattern are also provided. Methods of detecting or quantifying DNA from a plasma cell are also provided as are methods of diagnosing multiple myeloma (MM) or predicting progression to MM comprising quantifying DNA from plasma cells.
[0045] The invention is based, at least in part, on the surprising finding that cfDNA methylation pattern analysis can distinguish between MGUS, SMM and MM and even furthermore can predict progression. To examine the utility of epigenetic liquid biopsies for MM, samples relevant to plasma cell disorders, namely plasma cells from healthy individuals and from patients with MGUS, SMM and MM were examined. Methylation markers of plasma cells as well as cancer-specific markers of disordered methylation were extracted and were tested in PB plasma samples from patients with PCD. Presented herein is a unique and reliable system, showing how these makers can non-invasively distinguish different types of PCD, predict progression to MM, and inform on clonal turnover dynamics in MM.
[0046] In a first aspect, there is provided a method of identifying a methylation signature for a cancer cell comprising identifying in cell free DNA (cfDNA) a continuous sequence of no more than 300 nucleotides which comprise at least 4 methylation sites, wherein at least two sites are methylated and at least two sites are unmethylated, thereby identifying the methylation signature for the cancer cell.
[0047] In another aspect, there is provided a method of detecting cancer cfDNA, the method comprising ascertaining the methylation status of at least four methylation sties in the same cfDNA molecule and detecting the presence of at least two sites that are methylated and at least two sites that are unmethylated in the at least four methylation sties, thereby indicating that the cfDNA molecule is cancer cfDNA.
[0048] In another aspect, there is provided a method of quantifying molecule of DNA having a methylation pattern of a cancer cell, the method comprising:a. identifying a DNA sequence comprising at least four methylation sites, wherein the methylation pattern of the at least four methylation sites in a cancer cell is different as compared to the methylation pattern of the at least four methylation sites in the same DNA sequence in a non-cancerous cell, wherein the methylation pattern in a cancer cell comprises at least two sites that are methylated and at least two sites that are unmethylated and the methylation pattern in a non-cancerous cell comprises all sites being methylated or all sites being unmethylated; b. amplifying the identified DNA sequence to obtain amplified DNA molecules; c. using deep sequencing, massively parallel sequencing or next-generation sequencing to obtain the sequence of individual molecules of the amplified DNA molecules; d. ascertaining from the sequenced amplified DNA molecules the methylation status of each of the at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecule of the sequenced amplified DNA molecules; and e. quantifying the number of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a cancer cell that are present in the amplified DNA molecules; thereby quantifying molecules of DNA having the methylation pattern of a cancer cell.
[0049] In some embodiments, the methylation signature is a signature of a cancerous cell. In some embodiments, a methylation pattern is a methylation signature. In some embodiments, the methylation signature is a cancer signature. In some embodiments, the method is a method of diagnosing cancer. In some embodiments, the diagnosing is in a subject. In some embodiments, the cfDNA is from a sample. In some embodiments, the sample is from the subject. In some embodiments, the cfDNA is from the subject. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a diagnostic method. In some embodiments, the method is a prognostic method. In some embodiments, the method is acomputerized method. In some embodiments, the sample is provided. In some embodiments, the sample is obtained from the subject. In some embodiments, a sample from the subject is provided.
[0050] As used herein, the term “methylation site” refers to a cytosine residue adjacent to a guanine residue (CpG site) that has a potential of being methylated.
[0051] In some embodiments, the DNA molecule is a continuous sequence. In some embodiments, the DNA molecule is a double stranded DNA molecule. In some embodiments, the DNA molecule is a cfDNA molecule. In some embodiments, the sequence is no longer than 300 nucleotides, 295 nucleotides, 290 nucleotides, 285 nucleotides, 280 nucleotides, 275 nucleotides, 270 nucleotides, 265 nucleotides, 260 nucleotides, 255 nucleotides, 250 nucleotides, 245 nucleotides, 240 nucleotides, 235 nucleotides, 230 nucleotides, 225 nucleotides, 220 nucleotides, 215 nucleotides, 210 nucleotides, 205 nucleotides, 200 nucleotides, 195 nucleotides, 190 nucleotides, 185 nucleotides, 180 nucleotides, 175 nucleotides, 170 nucleotides, 167 nucleotides, 165 nucleotides, 160 nucleotides, 155 nucleotides, 150 nucleotides, 145 nucleotides, 140 nucleotides, 135 nucleotides, 130 nucleotides, 125 nucleotides, 120 nucleotides, 115 nucleotides, 110 nucleotides, 105 nucleotides, 100 nucleotides, 95 nucleotides, 90 nucleotides, 85 nucleotides, 80 nucleotides, or 75 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the sequence is no longer than 300 nucleotides. In some embodiments, the sequence is no longer than 200 nucleotides. In some embodiments, the sequence is no longer than 167 nucleotides. In some embodiments, the sequence is no longer than 150 nucleotides. In some embodiments, the sequence is no longer than 100 nucleotides. In some embodiments, the sequence is no longer than 75 nucleotides. In some embodiments, the sequence is no longer than 71 nucleotides. In some embodiments, the sequence is between 5-500, 5-450, 5-400, 5-350, 5-300, 5-250, 5-200, 5-175, 5-167, 5- 150, 5-125, 5-100, 5-75, 10-500, 10-450, 10-400, 10-350, 10-300, 10-250, 10-200, 10-175, 10-167, 10-150, 10-125, 10-100, 10-75, 15-500, 15-450, 15-400, 15-350, 15-300, 15-250, 15-200, 15-175, 15-167, 15-150, 15-125, 15-100, 15-75, 20-500, 20-450, 20-400, 20-350, 20-300, 20-250, 20-200, 20-175, 20-167, 20-150, 20-125, 20-100, 20-75, 25-500, 25-450, 25-400, 25-350, 25-300, 25-250, 25-200, 25-175, 25-167, 25-150, 25-125, 25-100, or 25-75 nucleotides. Each possibility represents a separate embodiment of the invention. In some embodiments, the sequence is between 10-300 nucleotides. In some embodiments, the sequence is between 10-167 nucleotides. In some embodiments, the sequence is between 10-100 nucleotides. In some embodiments, the sequence is between 25-300 nucleotides. In some embodiments, the sequence is between 25-167 nucleotides. In some embodiments, the sequence is between 25-100 nucleotides. In some embodiments, the sequence is between 10- 75 nucleotides. In some embodiments, the sequence is between 25-75 nucleotides. In some embodiments, the sequence is a continuous sequence. In some embodiments, the sequence is a DNA sequence. In some embodiments, a nucleotide is a base pair.
[0052] In some embodiments, the sequence is within a coding region. In some embodiments, the sequence is within a non-coding region. In some embodiments, the sequence is in a gene. In some embodiments, the sequence is in an intragenic region.
[0053] In some embodiments, the sequence or molecule comprises at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 methylation sites. Each possibility represents a separate embodiment of the invention. In some embodiments, the sequence or molecule comprises a plurality of methylation sites. In some embodiments, the sequence or molecule comprises at least 3 methylation sites. In some embodiments, the sequence or molecule comprises at least 4 methylation sites. In some embodiments, the sequence or molecule comprises at least 5 methylation sites. In some embodiments, all methylation sites in the sequence or molecule are part of the signature. In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 methylation sites in the sequence or molecule are part of the signature. Each possibility represents a separate embodiment of the invention. In some embodiments, a plurality of methylation sites in the sequence or molecule are part of the signature. In some embodiments, at least 4 methylation sites in the sequence or molecule are part of the signature. In some embodiments, at least 5 methylation sites in the sequence or molecule are part of the signature. In some embodiments, the methylation sites are adjacent. In some embodiments, the methylation sites are sequential. In some embodiments, the at least 4 methylation sites are not intervened by other methylation sites.
[0054] In some embodiments, a methylation signature for a plasma cell comprises all the methylation sites unmethylation. In some embodiments, a methylation signature for a plasma cell comprises a plurality of methylation sites unmethylated. In some embodiments, a methylation signature for a plasma cell comprises a majority of methylation sites unmethylated. In some embodiments, a methylation signature for a plasma cell comprises at least all methylation sites but one unmethylated. In some embodiments, a methylation signature for a plasma cell comprises all methylation sites or all methylation sites but one unmethylated.
[0055] In some embodiments, a methylation signature for a cancer cell comprises at least two methylation sites methylated and at least two methylation sites unmethylated. In some embodiments, a methylation signature for a cancer cell comprises within a sequence of at least 4 methylation sites at least two methylation sites methylated and at least two methylation sites unmethylated. In some embodiments, the at least 4 methylation sites are sequential. In some embodiments, the at least 4 methylation sites are adjacent. In some embodiments, the at least two methylated sites and at least two unmethylated sites are adjacent. In some embodiments, the at least two methylated sites and at least two unmethylated sites are sequential. As used herein, the terms “sequential” and “adjacent” refer to CpGs that have no intervening CpGs between them. Thus, adjacent / sequential CpGs need not be directly adjacent in a sequence (i.e., no intervening sequence whatsoever), rather between them there must not be any other CpGs (i.e., no intervening CpGs).
[0056] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is in need of a method of the invention. In some embodiments, the subject is in need of diagnosis. In some embodiments, the subject is in need of prognosis. In some embodiments, the subject is at risk of developing cancer. In some embodiments, the subject is suspected of having cancer. In some embodiments, the subject suffers from monoclonal gammopathy of undetermined significance (MGUS). In some embodiments, the subject suffers from smoldering multiple myeloma (SMM). In some embodiments, the subject is at risk for developing multiple myeloma (MM). In some embodiments, the subject is healthy.
[0057] In some embodiments, the sample is a bodily fluid sample. In some embodiments, a bodily fluid is a fluid from the subject. In some embodiments, the sample is a bone marrow sample. In some embodiments, the bone marrow sample is a bone marrow aspirate. In some embodiments, the bone marrow sample is a bone marrow biopsy. In some embodiments, the sample is a tumor biopsy. In some embodiments, the sample comprises cancer cells. In some embodiments, the fluid is selected from at least one of: blood, serum, plasma, gastric fluid, intestinal fluid, saliva, bile, tumor fluid, breast milk, semen, sperm, urine, interstitial fluid, cerebral spinal fluid, bone marrow aspirate and stool. In some embodiments, the fluid is selected from at least one of: blood, serum, plasma, gastric fluid, intestinal fluid, saliva, bile, tumor fluid, breast milk, semen, sperm, urine, interstitial fluid, cerebral spinal fluid and stool. In some embodiments, the sample is selected from the group consisting of blood, plasma, serum, sperm, milk, urine, saliva and cerebral spinal fluid. In some embodiments, the sampleis selected from the group consisting of blood, plasma, serum, sperm, milk, urine, saliva, bone marrow aspirate and cerebral spinal fluid. In some embodiments, the sample is a blood sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a serum sample. Blood, plasma and serum may be isolated according to methods well known in the art. DNA may be isolated from the sample by any method known in the art. Kits for DNA isolation from blood, serum, plasma, feces, urine and many other bodily fluids are well known in the art and are commercially available. Any such method or kit may be used. DNA may be isolated from the blood immediately or within 1 hour, 2 hours, 3 hours, 4 hours, 5 hours or 6 hours. Optionally the blood is stored at temperatures such as 4 °C, or at -20 °C prior to isolation of the DNA. In some embodiments, a portion of the blood sample is used in accordance with the invention at a first instance of time whereas one or more remaining portions of the blood sample (or fractions thereof) are stored for a period of time for later use.
[0058] In some embodiments, the DNA is cellular DNA. Methods of DNA extraction are well-known in the art. A classical DNA isolation protocol is based on extraction using organic solvents such as a mixture of phenol and chloroform, followed by precipitation with ethanol (J. Sambrook et al., "Molecular Cloning: A Laboratory Manual", 1989, 2nd Ed., Cold Spring Harbour Laboratory Press: New York, N.Y.). Other methods include: salting out DNA extraction (P. Sunnucks et al., Genetics, 1996, 144: 747-756; S. M. Aljanabi and I. Martinez, Nucl. Acids Res. 1997, 25: 4692-4693), trimethylammonium bromide salts DNA extraction (S. Gustincich et al., BioTechniques, 1991, 11: 298-302) and guanidinium thiocyanate DNA extraction (J. B. W. Hammond et al., Biochemistry, 1996, 240: 298-300). There are also numerous versatile kits that can be used to extract DNA from tissues and bodily fluids and that are commercially available from, for example, BD Biosciences Clontech (Palo Alto, Calif.), Epicentre Technologies (Madison, Wis.), Gentra Systems, Inc. (Minneapolis, Minn.), MicroProbe Corp. (Bothell, Wash.), Organon Teknika (Durham, N.C.), and Qiagen Inc. (Valencia, Calif.). User Guides that describe in great detail the protocol to be followed are usually included in all these kits. Sensitivity, processing time and cost may be different from one kit to another. One of ordinary skill in the art can easily select the kit(s) most appropriate for a particular situation.
[0059] In some embodiments, the DNA is cell free DNA (cfDNA). As used herein, the term “cell free DNA” or “cfDNA” refers to any DNA that is no longer within a cell membrane. In some embodiments, cfDNA is DNA that was released by a cell when it died. When usingcfDNA, cell lysis is not performed on the sample. Methods of isolating cell-free DNA from body fluids are also known in the art. For example, Qiaquick kit, manufactured by Qiagen may be used to extract cell-free DNA from plasma or serum.
[0060] The sample may be processed before the method is carried out, for example DNA purification may be carried out following the extraction procedure. The DNA in the sample may be cleaved either physically or chemically (e.g. using a suitable enzyme). Processing of the sample may involve one or more of: filtration, distillation, centrifugation, extraction, concentration, dilution, purification, inactivation of interfering components, addition of reagents, and the like.
[0061] It will be understood by a skilled artisan that reference to contiguous sequence refers to the sequence being within one DNA molecule, that is within the same DNA molecule. Some sequencing and methyl-sequencing techniques make use of pooling of DNA molecules. That is the average methylation at a site is calculated or the consensus base is calculated. Such methods of sequencing are not compatible with the instant methods that require ascertaining the methylation status of multiple sites on the same piece of DNA (i.e., a continuous sequence).
[0062] Methods of determining the methylation status of a methylation site are known in the art and include the use of bisulfite. In some embodiments, DNA is treated with bisulfite which converts cytosine residues to uracil (which are converted to thymidine following PCR), but leaves 5-methylcytosine residues unaffected. Thus, bisulfite treatment introduces specific changes in the DNA sequence that depend on the methylation status of individual cytosine residues, yielding single- nucleotide resolution information about the methylation status of a segment of DNA. Various analyses can be performed on the altered sequence to retrieve this information. The objective of this analysis is therefore reduced to differentiating between single nucleotide polymorphisms (cytosines and thymidine) resulting from bisulfite conversion. In some embodiments, the method comprises bisulfite treating the DNA. In some embodiments, bisulfite treating is bisulfite converting.
[0063] During the bisulfite reaction, care should be taken to minimize DNA degradation, such as cycling the incubation temperature. Bisulfite sequencing relies on the conversion of every single unmethylated cytosine residue to uracil. If conversion is incomplete, the subsequent analysis will incorrectly interpret the unconverted unmethylated cytosines as methylated cytosines, resulting in false positive results for methylation. Only cytosines insingle- stranded DNA are susceptible to attack by bisulfite, therefore denaturation of the DNA undergoing analysis is critical. It is important to ensure that reaction parameters such as temperature and salt concentration are suitable to maintain the DNA in a single- stranded conformation and allow for complete conversion.
[0064] According to a particular embodiment, an oxidative bisulfite reaction is performed. 5-methylcytosine and 5 -hydroxy methylcytosine both read as a C in bisulfite sequencing. Oxidative bisulfite reaction allows for the discrimination between 5-methylcytosine and 5- hydroxymethylcytosine at single base resolution. The method employs a specific chemical oxidation of 5 -hydroxy methylcytosine to 5-formylcytosine, which subsequently converts to uracil during bisulfite treatment. The only base that then reads as a C is 5-methylcytosine, giving a map of the true methylation status in the DNA sample. Levels of 5 -hydroxy methylcytosine can also be quantified by measuring the difference between bisulfite and oxidative bisulfite sequencing.
[0065] Prior to analysis (or concomitant therewith), the bisulfite-treated DNA sequence which comprises the methylation sites may be subjected to an amplification reaction. In some embodiments, the DNA is amplified. In some embodiments, the bisulfite treated DNA is amplified. In some embodiments, the bisulfite treatment occurs before the amplification. If amplification of the sequence is required care should be taken to ensure complete desulfonation of pyrimidine residues. This may be effected by monitoring the pH of the solution to ensure that desulfonation is complete.
[0066] As used herein, the term "amplification" refers to a process that increases the representation of a population of specific nucleic acid sequences in a sample by producing multiple (i.e., at least 2) copies of the desired sequences. Methods for nucleic acid amplification are known in the art and include, but are not limited to, polymerase chain reaction (PCR) and ligase chain reaction (LCR). In a typical PCR amplification reaction, a nucleic acid sequence of interest is often amplified at least fifty thousand fold in amount over its amount in the starting sample. A "copy" or "amplicon" does not necessarily mean perfect sequence complementarity or identity to the template sequence. For example, copies can include nucleotide analogs such as deoxyinosine, intentional sequence alterations (such as sequence alterations introduced through a primer comprising a sequence that is hybridizable but not complementary to the template), and / or sequence errors that occur during amplification.
[0067] A typical amplification reaction is carried out by contacting a forward and reverse primer (a primer pair) to the sample DNA together with any additional amplification reaction reagents under conditions which allow amplification of the target sequence. The terms "forward primer" and "forward amplification primer" are used herein interchangeably, and refer to a primer that hybridizes (or anneals) to the target (template strand). The terms "reverse primer" and "reverse amplification primer" are used herein interchangeably, and refer to a primer that hybridizes (or anneals) to the complementary target strand. The forward primer hybridizes with the target sequence 5' with respect to the reverse primer.
[0068] The term "amplification conditions" , as used herein, refers to conditions that promote annealing and / or extension of primer sequences. Such conditions are well-known in the art and depend on the amplification method selected. Thus, for example, in a PCR reaction, amplification conditions generally comprise thermal cycling, i.e., cycling of the reaction mixture between two or more temperatures. In isothermal amplification reactions, amplification occurs without thermal cycling although an initial temperature increase may be required to initiate the reaction. Amplification conditions encompass all reaction conditions including, but not limited to, temperature and temperature cycling, buffer, salt, ionic strength, pH, and the like.
[0069] As used herein, the term "amplification reaction reagents", refers to reagents used in nucleic acid amplification reactions and may include, but are not limited to, buffers, reagents, enzymes having reverse transcriptase and / or polymerase activity or exonuclease activity, enzyme cofactors such as magnesium or manganese, salts, nicotinamide adenine dinuclease (NAD) and deoxy nucleoside triphosphates (dNTPs), such as deoxyadenosine triphospate, deoxyguanosine triphosphate, deoxycytidine triphosphate and thymidine triphosphate. Amplification reaction reagents may readily be selected by one skilled in the art depending on the amplification method used.
[0070] According to this aspect of the present invention, the amplifying may be effected using techniques such as polymerase chain reaction (PCR), which includes, but is not limited to Allele- specific PCR, Assembly PCR or Polymerase Cycling Assembly (PCA), Asymmetric PCR, Helicase-dependent amplification, Hot-start PCR, Intersequence-specific PCR (ISSR), Inverse PCR, Ligation-mediated PCR, Methylation- specific PCR (MSP), Miniprimer PCR, Multiplex Ligation-dependent Probe Amplification, Multiplex- PCR, Nested PCR, Overlap-extension PCR, Quantitative PCR (Q-PCR), Reverse Transcription PCR (RT-PCR), Solid Phase PCR: encompasses multiple meanings, includingPolony Amplification (where PCR colonies are derived in a gel matrix, for example), Bridge PCR (primers are covalently linked to a solid-support surface), conventional Solid Phase PCR (where Asymmetric PCR is applied in the presence of solid support bearing primer with sequence matching one of the aqueous primers) and Enhanced Solid Phase PCR (where conventional Solid Phase PCR can be improved by employing high Tm and nested solid support primer with optional application of a thermal 'step' to favour solid support priming), Thermal asymmetric interlaced PCR (TAIL-PCR), Touchdown PCR (Step-down PCR), PAN- AC and Universal Fast Walking.
[0071] The PCR (or polymerase chain reaction) technique is well-known in the art and has been disclosed, for example, in K. B. Mullis and F. A. Faloona, Methods Enzymol., 1987, 155: 350-355 and U.S. Pat. Nos. 4,683,202; 4,683,195; and 4,800,159 (each of which is incorporated herein by reference in its entirety). In its simplest form, PCR is an in vitro method for the enzymatic synthesis of specific DNA sequences, using two oligonucleotide primers that hybridize to opposite strands and flank the region of interest in the target DNA. A plurality of reaction cycles, each cycle comprising: a denaturation step, an annealing step, and a polymerization step, results in the exponential accumulation of a specific DNA fragment ("PCR Protocols: A Guide to Methods and Applications", M. A. Innis (Ed.), 1990, Academic Press: New York; "PCR Strategies", M. A. Innis (Ed.), 1995, Academic Press: New York; "Polymerase chain reaction: basic principles and automation in PCR: A Practical Approach", McPherson et al. (Eds.), 1991, IRL Press: Oxford; R. K. Saiki et al., Nature, 1986, 324: 163-166). The termini of the amplified fragments are defined as the 5' ends of the primers. Examples of DNA polymerases capable of producing amplification products in PCR reactions include, but are not limited to: E. coli DNA polymerase I, Klenow fragment of DNA polymerase I, T4 DNA polymerase, thermostable DNA polymerases isolated from Thermus aquaticus (Taq), available from a variety of sources (for example, Perkin Elmer), Thermus thermophilus (United States Biochemicals), Bacillus stereo thermophilus (BioRad), or Thermococcus litoralis ("Vent" polymerase, New England Biolabs). RNA target sequences may be amplified by reverse transcribing the mRNA into cDNA, and then performing PCR (RT-PCR), as described above. Alternatively, a single enzyme may be used for both steps as described in U.S. Pat. No. 5,322,770.
[0072] The duration and temperature of each step of a PCR cycle, as well as the number of cycles, are generally adjusted according to the stringency requirements in effect. Annealing temperature and timing are determined both by the efficiency with which a primer isexpected to anneal to a template and the degree of mismatch that is to be tolerated. The ability to optimize the reaction cycle conditions is well within the knowledge of one of ordinary skill in the art. Although the number of reaction cycles may vary depending on the detection analysis being performed, it usually is at least 15, more usually at least 20, and may be as high as 60 or higher. However, in many situations, the number of reaction cycles typically ranges from about 20 to about 45.
[0073] The denaturation step of a PCR cycle generally comprises heating the reaction mixture to an elevated temperature and maintaining the mixture at the elevated temperature for a period of time sufficient for any double- stranded or hybridized nucleic acid present in the reaction mixture to dissociate. For denaturation, the temperature of the reaction mixture is usually raised to, and maintained at, a temperature ranging from about 85 °C. to about 100 °C, usually from about 90 °C to about 98 °C, and more usually from about 93 °C. to about 96 °C. for a period of time ranging from about 3 to about 120 seconds, usually from about 5 to about 30 seconds.
[0074] Following denaturation, the reaction mixture is subjected to conditions sufficient for primer annealing to template DNA present in the mixture. The temperature to which the reaction mixture is lowered to achieve these conditions is usually chosen to provide optimal efficiency and specificity, and generally ranges from about 50 °C to about °C, usually from about 55 °C to about 70 °C, and more usually from about 60 °C to about 68 °C. Annealing conditions are generally maintained for a period of time ranging from about 15 seconds to about 30 minutes, usually from about 30 seconds to about 5 minutes.
[0075] Following annealing of primer to template DNA or during annealing of primer to template DNA, the reaction mixture is subjected to conditions sufficient to provide for polymerization of nucleotides to the primer's end in a such manner that the primer is extended in a 5' to 3' direction using the DNA to which it is hybridized as a template, (i.e., conditions sufficient for enzymatic production of primer extension product). To achieve primer extension conditions, the temperature of the reaction mixture is typically raised to a temperature ranging from about 65°C to about 75 °C, usually from about 67 °C to about 73 °C, and maintained at that temperature for a period of time ranging from about 15 seconds to about 20 minutes, usually from about 30 seconds to about 5 minutes.
[0076] The above cycles of denaturation, annealing, and polymerization may be performed using an automated device typically known as a thermal cycler or thermocycler. Thermalcyclers that may be employed are described in U.S. Pat. Nos. 5,612,473; 5,602,756; 5,538,871; and 5,475,610 (each of which is incorporated herein by reference in its entirety). Thermal cyclers are commercially available, for example, from Perkin Elmer-Applied Biosystems (Norwalk, Conn.), BioRad (Hercules, Calif.), Roche Applied Science (Indianapolis, Ind.), and Stratagene (La Jolla, Calif.).
[0077] According to one embodiment, the primers which are used in the amplification reaction are methylation independent primers. These primers flank the first and last of the at least four methylation sites (but do not hybridize directly to the sites) and in a PCR reaction, are capable of generating an amplicon which comprises all four or more methylation sites.
[0078] The methylation-independent primers of this aspect of the present invention may comprise adaptor sequences which include barcode sequences. The adaptors may further comprise sequences which are necessary for attaching to a flow cell surface (P5 and P7 sites, for subsequent sequencing), a sequence which encodes for a promoter for an RNA polymerase and / or a restriction site. The barcode sequence may be used to identify a particular molecule, sample or library. The barcode sequence may be between 3-400 nucleotides, more preferably between 3-200 and even more preferably between 3-100 nucleotides. Thus, the barcode sequence may be 6 nucleotides, 7 nucleotides, 8, nucleotides, nine nucleotides or ten nucleotides. The barcode is typically 4-15 nucleotides.
[0079] The methylation-independent oligonucleotide of this aspect of the present invention need not reflect the exact sequence of the target nucleic acid sequence (i.e. need not be fully complementary), but must be sufficiently complementary so as to hybridize to the target site under the particular experimental conditions. Accordingly, the sequence of the oligonucleotide typically has at least 70 % homology, preferably at least 80 %, 90 %, 95 %, 97 %, 99 % or 100 % homology, for example over a region of at least 13 or more contiguous nucleotides with the target sequence. The conditions are selected such that hybridization of the oligonucleotide to the target site is favored and hybridization to the non-target site is minimized.
[0080] Various considerations must be taken into account when selecting the stringency of the hybridization conditions. For example, the more closely the oligonucleotide (e.g. primer) reflects the target nucleic acid sequence, the higher the stringency of the assay conditions can be, although the stringency must not be too high so as to prevent hybridization of the oligonucleotides to the target sequence. Further, the lower the homology of theoligonucleotide to the target sequence, the lower the stringency of the assay conditions should be, although the stringency must not be too low to allow hybridization to non-specific nucleic acid sequences.
[0081] The DNA may be sequenced using any method known in the art - e.g. massively parallel DNA sequencing, sequencing-by-synthesis, sequencing-by-ligation, 454 pyrosequencing, cluster amplification, bridge amplification, and PCR amplification, although preferably, the method comprises a high throughput sequencing method. Typical methods include the sequencing technology and analytical instrumentation offered by Roche 454 Life SciencesTM, Branford, Conn., which is sometimes referred to herein as "454 technology" or "454 sequencing."; the sequencing technology and analytical instrumentation offered by Illumina, Inc, San Diego, Calif, (their Solexa Sequencing technology is sometimes referred to herein as the "Solexa method" or "Solexa technology"); or the sequencing technology and analytical instrumentation offered by ABI, Applied Biosystems, Indianapolis, Ind., which is sometimes referred to herein as the ABI-SOLiDTM platform or methodology.
[0082] Other known methods for sequencing include, for example, those described in: Sanger, F. et al., Proc. Natl. Acad. Sci. U.S.A. 75, 5463-5467 (1977); Maxam, A. M. & Gilbert, W. Proc Natl Acad Sci USA 74, 560-564 (1977); Ronaghi, M. et al., Science 281, 363, 365 (1998); Lysov, 1. et al., Dokl Akad Nauk SSSR 303, 1508-1511 (1988); Bains W. & Smith G. C. J. Theor Biol 135, 303-307 (1988); Dmanac, R. et al., Genomics 4, 114-128 (1989); Khrapko, K. R. et al., FEBS Lett 256.118-122 (1989); Pevzner P. A. J Biomol Struct Dyn 7, 63-73 (1989); and Southern, E. M. et al., Genomics 13, 1008-1017 (1992). Pyrophosphate-based sequencing reaction as described, e.g., in U.S. Patent Nos. 6,274,320, 6,258,568 and 6,210,891, may also be used.
[0083] The Illumina or Solexa sequencing is based on reversible dye-terminators. DNA molecules are typically attached to primers on a slide and amplified so that local clonal colonies are formed. Subsequently one type of nucleotide at a time may be added, and nonincorporated nucleotides are washed away. Subsequently, images of the fluorescently labeled nucleotides may be taken and the dye is chemically removed from the DNA, allowing a next cycle. The Applied Biosystems' SOLiD technology, employs sequencing by ligation. This method is based on the use of a pool of all possible oligonucleotides of a fixed length, which are labeled according to the sequenced position. Such oligonucleotides are annealed and ligated. Subsequently, the preferential ligation by DNA ligase for matchingsequences typically results in a signal informative of the nucleotide at that position. Since the DNA is typically amplified by emulsion PCR, the resulting bead, each containing only copies of the same DNA molecule, can be deposited on a glass slide resulting in sequences of quantities and lengths comparable to Illumina sequencing. Another example of an envisaged sequencing method is pyro sequencing, in particular 454 pyro sequencing, e.g. based on the Roche 454 Genome Sequencer. This method amplifies DNA inside water droplets in an oil solution with each droplet containing a single DNA template attached to a single primer-coated bead that then forms a clonal colony. Pyro sequencing uses luciferase to generate light for detection of the individual nucleotides added to the nascent DNA, and the combined data are used to generate sequence read-outs. A further method is based on Helicos' Heliscope technology, wherein fragments are captured by polyT oligomers tethered to an array. At each sequencing cycle, polymerase and single fluorescently labeled nucleotides are added and the array is imaged. The fluorescent tag is subsequently removed, and the cycle is repeated. Further examples of sequencing techniques encompassed within the methods of the present invention are sequencing by hybridization, sequencing by use of nanopores, microscopy-based sequencing techniques, microfluidic Sanger sequencing, or microchip-based sequencing methods. The present invention also envisages further developments of these techniques, e.g. further improvements of the accuracy of the sequence determination, or the time needed for the determination of the genomic sequence of an organism etc.
[0084] In some embodiments, the sequencing is deep sequencing. In some embodiments, the sequencing is massively parallel sequencing. In some embodiments, the sequencing is next generation sequencing. As used herein, the term “deep sequencing” and variations thereof refers to the number of times a nucleotide is read during the sequencing process. Deep sequencing indicates that the coverage, or depth, of the process is many times larger than the length of the sequence under study. In some embodiments, deep sequencing comprises at least a 30X depth. In some embodiments, deep sequencing comprises at least a 20X depth. In some embodiments, deep sequencing comprises at least a 10X depth.
[0085] In some embodiments, the sequencing comprises sequencing at least 100, 500, 1000, 1500, 200, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500 or 10000 individual molecules. Each possibility represents a separate embodiment of the invention. In some embodiments, the sequence of at least 100, 500, 1000, 1500, 200, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500,9000, 9500 or 10000 individual molecules are obtained. Each possibility represents a separate embodiment of the invention. In some embodiments, at least 10000 individual molecules are sequenced. In some embodiments, at least 5000 individual molecules are sequenced.
[0086] It will be appreciated that any of the analytical methods described herein can be embodied in many forms. For example, it can be embodied on a tangible medium such as a computer for performing the method operations. It can be embodied on a computer readable medium, comprising computer readable instructions for carrying out the method operations. It can also be embodied in an electronic device having digital computer capabilities arranged to run the computer program on the tangible medium or execute the instruction on a computer readable medium.
[0087] Computer programs implementing the analytical method of the present embodiments can commonly be distributed to users on a distribution medium such as, but not limited to, CD-ROMs or flash memory media. From the distribution medium, the computer programs can be copied to a hard disk or a similar intermediate storage medium. In some embodiments of the present invention, computer programs implementing the method of the present embodiments can be distributed to users by allowing the user to download the programs from a remote location, via a communication network, e.g., the internet. The computer programs can be run by loading the computer instructions either from their distribution medium or their intermediate storage medium into the execution memory of the computer, configuring the computer to act in accordance with the method of this invention. All these operations are well-known to those skilled in the art of computer systems.
[0088] Additional methods which rely on the use of bisulfite that may be used to analyze the methylation pattern as described herein are described herein below:
[0089] Methylation- sensitive single-nucleotide primer extension: DNA is bisulfite - converted, and bisulfite- specific primers are annealed to the sequence up to the base pair immediately before the CpG of interest. The primer is allowed to extend one base pair into the C (or T) using DNA polymerase terminating dideoxynucleotides, and the ratio of C to T is determined quantitatively. A number of methods can be used to determine this C:T ratio, such as the use of radioactive ddNTPs as the reporter of the primer extension, fluorescencebased methods or Pyro sequencing can also be used. Matrix-assisted laser desorption ionization / time-of-flight (MAEDI-TOF) mass spectrometry analysis can be used todifferentiate between the two polymorphic primer extension products can be used, in essence, based on the GOOD assay designed for SNP genotyping. Ion pair reverse-phase high-performance liquid chromatography (IP-RP-HPLC) can also be used to distinguish primer extension products.
[0090] Base-specific cleavage / MALDI-TOF: This method takes advantage of bisulfite - conversions by adding a base- specific cleavage step to enhance the information gained from the nucleotide changes. By first using in vitro transcription of the region of interest into RNA (by adding an RNA polymerase promoter site to the PCR primer in the initial amplification), RNase A can be used to cleave the RNA transcript at base-specific sites. As RNase A cleaves RNA specifically at cytosine and uracil ribonucleotides, base-specificity is achieved by adding incorporating cleavage-resistant dTTP when cytosine- specific (C-specific) cleavage is desired, and incorporating dCTP when uracil- specific (U-specific) cleavage is desired. The cleaved fragments can then be analyzed by MALDI-TOF. Bisulfite treatment results in either introduction / removal of cleavage sites by C-to-U conversions or shift in fragment mass by G-to-A conversions in the amplified reverse strand. C-specific cleavage will cut specifically at all methylated CpG sites. By analyzing the sizes of the resulting fragments, it is possible to determine the specific pattern of DNA methylation of CpG sites within the region.
[0091] The present inventors further contemplate analyzing the methylation status of the at least four sites including the use of methylation-dependent oligonucleotides. Methylation dependent oligonucleotides hybridize to either the methylated form of the at least one methylation site or the unmethylated form of the at least one methylation site.
[0092] According to one embodiment, the methylation dependent oligonucleotide is a probe. In one embodiment, the probe hybridizes to the methylated site to provide a detectable signal under experimental conditions and does not hybridize to the non-methylated site to provide a detectable signal under identical experimental conditions. In another embodiment, the probe hybridizes to the non-methylated site to provide a detectable signal under experimental conditions and does not hybridize to the methylated site to provide a detectable signal under identical experimental conditions. The probes of this embodiment of this aspect of the present invention may be, for example, affixed to a solid support (e.g., arrays or beads).
[0093] According to another embodiment, the methylation dependent oligonucleotide is a primer which when used in an amplification reaction is capable of amplifying the targetsequence, when the methylation site is methylated. According to another embodiment, the methylation dependent oligonucleotide is a primer which when used in an amplification reaction is capable of amplifying the target sequence, when the methylation site is unmethylated - see for example International PCT Publication No. W02013131083, the contents of which are incorporated herein by reference.
[0094] The methylation-dependent oligonucleotide of this aspect of the present invention need not reflect the exact sequence of the target nucleic acid sequence (i.e. need not be fully complementary), but must be sufficiently complementary so as to distinguish between a methylated and non-methylated site under the particular experimental conditions. Accordingly, the sequence of the oligonucleotide typically has at least 70 % homology, preferably at least 80 %, 90 %, 95 %, 97 %, 99 % or 100 % homology / identity, for example over a region of at least 13 or more contiguous nucleotides with the target sequence. The conditions are selected such that hybridization of the oligonucleotide to the methylated site is favored and hybridization to the non-methylated site is minimized (and vice versa).
[0095] By way of example, hybridization of short nucleic acids (below 200 bp in length, e.g. 13-50 bp in length) can be effected by the following hybridization protocols depending on the desired stringency; (i) hybridization solution of 6 x SSC and 1 % SDS or 3 M TMAC1, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 qg / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature of 1 - 1.5 °C below the Tm, final wash solution of 3 M TMAC1, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS at 1 - 1.5 °C below the Tm (stringent hybridization conditions) (ii) hybridization solution of 6 x SSC and 0.1 % SDS or 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 qg / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature of 2 - 2.5 °C below the Tm, final wash solution of 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS at 1 - 1.5 °C below the Tm, final wash solution of 6 x SSC, and final wash at 22 °C (stringent to moderate hybridization conditions); and (iii) hybridization solution of 6 x SSC and 1 % SDS or 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 qg / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature at 2.5-3 °C below the Tm and final wash solution of 6 x SSC at 22 °C (moderate hybridization solution).
[0096] Oligonucleotides of the invention may be prepared by any of a variety of methods (see, for example, J. Sambrook et al., "Molecular Cloning: A Laboratory Manual", 1989, 2.sup.nd Ed., Cold Spring Harbour Laboratory Press: New York, N.Y.; "PCR Protocols: A Guide to Methods and Applications", 1990, M. A. Innis (Ed.), Academic Press: New York, N.Y.; P. Tijssen "Hybridization with Nucleic Acid Probes— Laboratory Techniques in Biochemistry and Molecular Biology (Parts I and II)", 1993, Elsevier Science; "PCR Strategies", 1995, M. A. Innis (Ed.), Academic Press: New York, N.Y.; and "Short Protocols in Molecular Biology", 2002, F. M. Ausubel (Ed.), 5. sup. th Ed., John Wiley & Sons: Secaucus, N.J.). For example, oligonucleotides may be prepared using any of a variety of chemical techniques well-known in the art, including, for example, chemical synthesis and polymerization based on a template as described, for example, in S. A. Narang et al., Meth. Enzymol. 1979, 68: 90-98; E. L. Brown et al., Meth. Enzymol. 1979, 68: 109-151; E. S. Belousov et al., Nucleic Acids Res. 1997, 25: 3440-3444; D. Guschin et al., Anal. Biochem. 1997, 250: 203-211; M. J. Blommers et al., Biochemistry, 1994, 33: 7886-7896; and K. Frenkel et al., Free Radic. Biol. Med. 1995, 19: 373-380; and U.S. Pat. No. 4,458,066.
[0097] For example, oligonucleotides may be prepared using an automated, solid-phase procedure based on the phosphoramidite approach. In such a method, each nucleotide is individually added to the 5'-end of the growing oligonucleotide chain, which is attached at the 3 '-end to a solid support. The added nucleotides are in the form of trivalent 3'- phosphoramidites that are protected from polymerization by a dimethoxytriyl (or DMT) group at the 5'-position. After base-induced phosphoramidite coupling, mild oxidation to give a pentavalent phosphotriester intermediate and DMT removal provides a new site for oligonucleotide elongation. The oligonucleotides are then cleaved off the solid support, and the phosphodiester and exocyclic amino groups are deprotected with ammonium hydroxide. These syntheses may be performed on oligo synthesizers such as those commercially available from Perkin Elmer / Applied Biosystems, Inc. (Foster City, Calif.), DuPont (Wilmington, Del.) or Milligen (Bedford, Mass.). Alternatively, oligonucleotides can be custom made and ordered from a variety of commercial sources well-known in the art, including, for example, the Midland Certified Reagent Company (Midland, Tex.), ExpressGen, Inc. (Chicago, Ill.), Operon Technologies, Inc. (Huntsville, Ala.), and many others.
[0098] Purification of the oligonucleotides of the invention, where necessary or desirable, may be carried out by any of a variety of methods well-known in the art. Purification ofoligonucleotides is typically performed either by native acrylamide gel electrophoresis, by anion-exchange HPLC as described, for example, by J. D. Pearson and F. E. Regnier (J. Chrom., 1983, 255: 137-149) or by reverse phase HPLC (G. D. McFarland and P. N. Borer, Nucleic Acids Res., 1979, 7: 1067-1080).
[0099] The sequence of oligonucleotides can be verified using any suitable sequencing method including, but not limited to, chemical degradation (A. M. Maxam and W. Gilbert, Methods of Enzymology, 1980, 65: 499-560), matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry (U. Pieles et al., Nucleic Acids Res., 1993, 21: 3191-3196), mass spectrometry following a combination of alkaline phosphatase and exonuclease digestions (H. Wu and H. Aboleneen, Anal. Biochem., 2001, 290: 347-352), and the like.
[0100] In certain embodiments, the detection probes or amplification primers or both probes and primers are labeled with a detectable agent or moiety before being used in amplification / detection assays. In certain embodiments, the detection probes are labeled with a detectable agent. Preferably, a detectable agent is selected such that it generates a signal which can be measured and whose intensity is related (e.g., proportional) to the amount of amplification products in the sample being analyzed.
[0101] The association between the oligonucleotide and detectable agent can be covalent or non-covalent. Labeled detection probes can be prepared by incorporation of or conjugation to a detectable moiety. Labels can be attached directly to the nucleic acid sequence or indirectly (e.g., through a linker). Linkers or spacer arms of various lengths are known in the art and are commercially available, and can be selected to reduce steric hindrance, or to confer other useful or desired properties to the resulting labeled molecules (see, for example, E. S. Mansfield et al., Mol. Cell. Probes, 1995, 9: 145-156).
[0102] Methods for labeling nucleic acid molecules are well-known in the art. For a review of labeling protocols, label detection techniques, and recent developments in the field, see, for example, L. J. Kricka, Ann. Clin. Biochem. 2002, 39: 114-129; R. P. van Gijlswijk et al., Expert Rev. Mol. Diagn. 2001, 1: 81-91; and S. Joos et al., J. Biotechnol. 1994, 35: 135-153. Standard nucleic acid labeling methods include: incorporation of radioactive agents, direct attachments of fluorescent dyes (L. M. Smith et al., Nucl. Acids Res., 1985, 13: 2399-2412) or of enzymes (B. A. Connoly and O. Rider, Nucl. Acids. Res., 1985, 13: 4485-4502); chemical modifications of nucleic acid molecules making them detectableimmunochemically or by other affinity reactions (T. R. Broker et al., Nucl. Acids Res. 1978, 5: 363-384; E. A. Bayer et al., Methods of Biochem. Analysis, 1980, 26: 1-45; R. Langer et al., Proc. Natl. Acad. Sci. USA, 1981, 78: 6633-6637; R. W. Richardson et al., Nucl. Acids Res. 1983, 11: 6167-6184; D. J. Brigati et al., Virol. 1983, 126: 32-50; P. Tchen et al., Proc. Natl. Acad. Sci. USA, 1984, 81: 3466-3470; J. E. Landegent et al., Exp. Cell Res. 1984, 15: 61-72; and A. H. Hopman et al., Exp. Cell Res. 1987, 169: 357-368); and enzyme-mediated labeling methods, such as random priming, nick translation, PCR and tailing with terminal transferase (for a review on enzymatic labeling, see, for example, J. Temsamani and S. Agrawal, Mol. Biotechnol. 1996, 5: 223-232). More recently developed nucleic acid labeling systems include, but are not limited to: ULS (Universal Linkage System), which is based on the reaction of mono-reactive cisplatin derivatives with the N7 position of guanine moieties in DNA (R. J. Heetebrij et al., Cytogenet. Cell. Genet. 1999, 87: 47-52), psoralen-biotin, which intercalates into nucleic acids and upon UV irradiation becomes covalently bonded to the nucleotide bases (C. Levenson et al., Methods Enzymol. 1990, 184: 577-583; and C. Pfannschmidt et al., Nucleic Acids Res. 1996, 24: 1702-1709), photoreactive azido derivatives (C. Neves et al., Bioconjugate Chem. 2000, 11: 51-55), and DNA alkylating agents (M. G. Sebestyen et al., Nat. Biotechnol. 1998, 16: 568-576).
[0103] Any of a wide variety of detectable agents can be used in the practice of the present invention. Suitable detectable agents include, but are not limited to, various ligands, radionuclides (such as, for example, 32P, 35S, 3H, 14C, 1251, 1311, and the like); fluorescent dyes (for specific exemplary fluorescent dyes, see below); chemiluminescent agents (such as, for example, acridinium esters, stabilized dioxetanes, and the like); spectrally resolvable inorganic fluorescent semiconductor nanocrystals (i.e., quantum dots), metal nanoparticles (e.g., gold, silver, copper and platinum) or nanoclusters; enzymes (such as, for example, those used in an ELISA, i.e., horseradish peroxidase, beta-galactosidase, luciferase, alkaline phosphatase); colorimetric labels (such as, for example, dyes, colloidal gold, and the like); magnetic labels (such as, for example, DynabeadsTM); and biotin, dioxigenin or other haptens and proteins for which antisera or monoclonal antibodies are available.
[0104] In certain embodiments, the inventive detection probes are fluorescently labeled. Numerous known fluorescent labeling moieties of a wide variety of chemical structures and physical characteristics are suitable for use in the practice of this invention. Suitable fluorescent dyes include, but are not limited to, fluorescein and fluorescein dyes (e.g., fluorescein isothiocyanine or FITC, naphthofluorescein, 4',5'-dichloro-2',7'-dimethoxy-fluorescein, 6 carboxyfluorescein or FAM), carbocyanine, merocyanine, styryl dyes, oxonol dyes, phycoerythrin, erythrosin, eosin, rhodamine dyes (e.g., carboxytetramethylrhodamine or TAMRA, carboxyrhodamine 6G, carboxy-X-rhodamine (ROX), lissamine rhodamine B, rhodamine 6G, rhodamine Green, rhodamine Red, tetramethylrhodamine or TMR), coumarin and coumarin dyes (e.g., methoxycoumarin, dialkylaminocoumarin, hydroxycoumarin and aminomethylcoumarin or AMCA), Oregon Green Dyes (e.g., Oregon Green 488, Oregon Green 500, Oregon Green 514), Texas Red, Texas Red-X, Spectrum Red.TM., Spectrum Green. TM., cyanine dyes (e.g., Cy-3.TM., Cy-5.TM., Cy-3.5.TM., Cy- 5.5.TM.), Alexa Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 660 and Alexa Fluor 680), BODIPY dyes (e.g., BODIPY FL, BODIPY R6G, BODIPY TMR, BODIPY TR, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591, BODIPY 630 / 650, BODIPY 650 / 665), IRDyes (e.g., IRD40, IRD 700, IRD 800), and the like. For more examples of suitable fluorescent dyes and methods for linking or incorporating fluorescent dyes to nucleic acid molecules see, for example, "The Handbook of Fluorescent Probes and Research Products", 9th Ed., Molecular Probes, Inc., Eugene, Oreg. Fluorescent dyes as well as labeling kits are commercially available from, for example, Amersham Biosciences, Inc. (Piscataway, N.J.), Molecular Probes Inc. (Eugene, Oreg.), and New England Biolabs Inc. (Beverly, Mass.).Another contemplated method of analyzing the methylation status of the sequences is by analysis of the DNA following exposure to methylation-sensitive restriction enzymes - see for example US Application Nos. 20130084571 and 20120003634, the contents of which are incorporated herein.
[0105] In some embodiments, the nucleotide sequence is a genomic locus. In some embodiments, the nucleotide sequence comprises the genomic locus. In some embodiments, the nucleotide sequence consists of the genomic locus. In some embodiments, the cfDNA molecule comprises a genomic locus. In some embodiments, the methylation sites are within a genomic locus. In some embodiments, the at least four methylation sites are within a genomic locus. In some embodiments, the genomic locus is within a gene. In some embodiments, the genomic locus is in an intergenic region.
[0106] In some embodiments, the genomic locus is selected from those provided in Table 2. In some embodiments, the genomic locus is chrl5:101, 168, 755-101, 168, 812. In some embodiments, the genomic locus is chr7:134, 597, 197-134, 597, 250. In some embodiments, the genomic locus is chrl5:91,l 19,606-91,119,642. In some embodiments, the genomiclocus is chr2:85, 005, 386-85, 005, 443. In some embodiments, the genomic locus is chrl5:41, 072, 702-41, 072, 749. In some embodiments, the genomic locus is chr7:23, 150,939- 23,151,010. In some embodiments, the genomic locus is chr5:81, 968, 704-81, 968, 773. In some embodiments, the genomic locus is chrl 1 :5, 225, 201-5, 225, 231. In some embodiments, the genomic locus is chrl l:107, 747, 505-107, 747, 531. In some embodiments, the genomic locus is chr6:80, 356, 706-80, 356, 741. In some embodiments, the genomic location is with respect to human genome build Hgl9. In some embodiments, location is coordinates. In some embodiments, the locus is in human genome build Hgl9.
[0107] In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 8. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 9. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 10. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 11. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 12. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 13. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 14. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 15. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 16. In some embodiments, the nucleotide sequence comprise or consists of SEQ ID NO: 17. In some embodiments, the nucleotide sequence comprise a sequence selected from the group consisting of SEQ ID NO: 8-17. In some embodiments, the nucleotide sequence consists of a sequence selected from the group consisting of SEQ ID NO: 8-17. In some embodiments, the nucleotide sequence is a DNA sequence. In some embodiments, the genomic locus comprise a sequence selected from the group consisting of SEQ ID NO: 8-17. In some embodiments, the genomic locus consists of a sequence selected from the group consisting of SEQ ID NO: 8-17.
[0108] In some embodiments, the methylation status of at least one genomic locus is ascertained. In some embodiments, the methylation status of a plurality of genomic loci is ascertained. In some embodiments, the methylation status of at least two genomic loci is ascertained. In some embodiments, the methylation status of at least three genomic loci is ascertained. In some embodiments, the methylation status of at least four genomic loci is ascertained. In some embodiments, the methylation status of at least five genomic loci is ascertained. In some embodiments, the methylation status of at least six genomic loci is ascertained. In some embodiments, the methylation status of at least seven genomic loci isascertained. In some embodiments, the methylation status of at least eight genomic loci is ascertained. In some embodiments, the methylation status of at least nine genomic loci is ascertained. In some embodiments, the methylation status of all ten genomic loci are ascertained.
[0109] In some embodiments, the ascertaining is detecting. In some embodiments, detecting comprises detecting all molecules comprising the genomic locus. In some embodiments, all molecules comprising the locus comprises at least 100, 500, 1000, 1500, 200, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500 or 10000 individual molecules. Each possibility represents a separate embodiment of the invention. In some embodiments, all molecules are all cfDNA molecules. In some embodiments, all molecules are all molecules in the sample. In some embodiments, the detecting comprises amplifying the genomic locus. In some embodiments, the detecting comprises determining the methylation status of all the molecules comprising the genomic locus. In some embodiments, the detecting comprises determining the percentage of the molecules comprising the genomic locus that comprises at least two sites that are methylated and at least two sites that are unmethylated. In some embodiments, a molecule with at least two sites that are methylated and at least two sites that are unmethylated is a discordant read. In some embodiments, a discordant read is a discordant epiallele. In some embodiments, the percentage of discordant reads (PDR) is determined. In some embodiments, determined is calculated. In some embodiments, the PDR of the genomic locus is calculated.
[0110] In some embodiments, the PDR is determined for a plurality of genomic loci. In some embodiments, the PDR is determined for at least two genomic loci. In some embodiments, the PDR is determined for at least three genomic loci. In some embodiments, the PDR is determined for at least four genomic loci. In some embodiments, the PDR is determined for at least five genomic loci. In some embodiments, the PDR is determined for at least six genomic loci. In some embodiments, the PDR is determined for at least seven genomic loci. In some embodiments, the PDR is determined for at least eight genomic loci. In some embodiments, the PDR is determined for at least nine genomic loci. In some embodiments, the PDR is determined for all ten genomic loci.
[0111] In some embodiments, a PDR above a predetermined threshold indicates the subject suffers from cancer. In some embodiments, the average PDR from a plurality of genomic loci above a predetermined threshold indicates the subject suffers from cancer. In some embodiments, the average PDR from all loci for which a PDR is determined above apredetermined threshold indicates the subject suffers from cancer. In some embodiments, the threshold is the PDR in a sample from a healthy subject. In some embodiments, the threshold is the average PDR in a sample from a healthy subject. In some embodiments, the PDR is for the genomic locus or plurality of genomic loci. In some embodiments, the predetermined threshold is specific to the type of sample.
[0112] In some embodiments, the predetermined threshold is a PDR of 0.005, 0.01, 0.015, 0.02, 0.025, 0.03, 0.035, 0.04, 0.045 or 0.05. Each possibility represents a separate embodiment of the invention. In some embodiments, the predetermined threshold is a PDR of 0.015. In some embodiments, the sample is a blood / serum / plasma sample and the predetermined threshold is a PDR of 0.015. In some embodiments, the sample is bone marrow aspirate and the predetermined threshold is a PDR of 0.015. In some embodiments, the predetermined threshold is a PDR of 0.03. In some embodiments, the sample is a blood / serum / plasma sample and the predetermined threshold is a PDR of 0.03. In some embodiments, the sample is bone marrow aspirate and the predetermined threshold is a PDR of 0.03.
[0113] In some embodiments, the cancer is a solid cancer. In some embodiments, the cancer is a tumor. As used herein "cancer" or "pre-malignancy" are diseases associated with cell proliferation. Non-limiting types of cancer include carcinoma, sarcoma, lymphoma, leukemia, blastoma and germ cells tumors. In one embodiment, carcinoma refers to tumors derived from epithelial cells including but not limited to breast cancer, prostate cancer, lung cancer, pancreas cancer, and colon cancer. In one embodiment, sarcoma refers of tumors derived from mesenchymal cells including but not limited to sarcoma botryoides, chondrosarcoma, ewings sarcoma, malignant hemangioendothelioma, malignant schwannoma, osteosarcoma and soft tissue sarcomas. In one embodiment, lymphoma refers to tumors derived from hematopoietic cells that leave the bone marrow and tend to mature in the lymph nodes including but not limited to hodgkin lymphoma, non-hodgkin lymphoma, multiple myeloma and immunoproliferative diseases. In one embodiment, leukemia refers to tumors derived from hematopoietic cells that leave the bone marrow and tend to mature in the blood including but not limited to acute lymphoblastic leukemia, chronic lymphocytic leukemia, acute myelogenous leukemia, chronic myelogenous leukemia, hairy cell leukemia, T-cell prolymphocytic leukemia, large granular lymphocytic leukemia and adult T-cell leukemia. Examples of cancer include, but are not limited to, hepato-biliary cancer, cervical cancer, urogenital cancer (e.g., urothelial cancer), testicularcancer, prostate cancer, thyroid cancer, ovarian cancer, nervous system cancer, ocular cancer, lung cancer, soft tissue cancer, bone cancer, pancreatic cancer, bladder cancer, skin cancer, intestinal cancer, hepatic cancer, rectal cancer, colorectal cancer, esophageal cancer, gastric cancer, gastroesophageal cancer, breast cancer (e.g., triple negative breast cancer), renal cancer (e.g., renal carcinoma), skin cancer, head and neck cancer, leukemia and lymphoma. In some embodiments, the cancer is selected from the group consisting of: multiple myeloma, smoldering multiple myeloma, prostate cancer, head and neck cancer, colon cancer, pancreatic cancer and breast cancer. In some embodiments, the head and neck cancer is laryngeal cancer. In some embodiments, the cancer myeloma. In some embodiments, the myeloma is multiple myeloma. In some embodiments, the cancer is multiple myeloma. In some embodiments, the cancer is smoldering multiple myeloma. In some embodiments, the multiple myeloma is smoldering multiple myeloma.
[0114] In some embodiments, the quantifying is quantifying molecules of cfDNA from cancer cells. In some embodiments, the quantifying is quantifying cancer cfDNA. In some embodiments, quantifying molecules is quantifying reads. In some embodiments, the DNA is cfDNA. In some embodiments, quantifying DNA molecules is quantifying cfDNA molecules. In some embodiments, the quantifying is quantifying molecules in a sample. In some embodiments, molecules in a sample are cfDNA molecules in a sample. In some embodiments, the sample is from a subject. In some embodiments, from a subject is derived from a subject.
[0115] In some embodiments, a non-cancerous cell is a control cell. In some embodiments, a non-cancerous cell is a healthy cell. In some embodiments, a non-cancerous cell is a plurality of non-cancerous cells. In some embodiments, a non-cancerous cell is non- cancerous cells. In some embodiments, the plurality of non-cancerous cells are from a plurality of cell types or tissues. In some embodiments, the non-cancerous cell is a brain cell. In some embodiments, a brain cell is a neuron. In some embodiments, the brain cell is an oligodendrocyte. In some embodiments, the non-cancerous cell is a breast cell. In some embodiments, the non-cancerous cell is a colon cell. In some embodiments, the non- cancerous cell is a cardiac cell. In some embodiments, a cardiac cell is a heart cell. In some embodiments, a cardiac cell is a cardiomyocyte. In some embodiments, the non-cancerous cell is a pancreas cell. In some embodiments, the non-cancerous cell is a kidney cell. In some embodiments, a kidney cell is a renal cell. In some embodiments, the non-cancerous cell is an esophageal cell. In some embodiments, the non-cancerous cell is a liver cell. In someembodiments, a liver cell is a hepatic cell. In some embodiments, a liver cell is a hepatocyte. In some embodiments, the non-cancerous cell is a muscle cell. In some embodiments, the muscle is skeletal muscle. In some embodiments, the muscle is smooth muscle. In some embodiments, the non-cancerous cell is a thyroid cell. In some embodiments, the non- cancerous cell is a leukocyte. In some embodiments, the non-cancerous cell is a granulocyte. In some embodiments, the non-cancerous cell is natural killer (NK) cell. In some embodiments, the non-cancerous cell is a T cell. In some embodiments, the T cell is a CD3 positive cell. In some embodiments, the T cell is a CD4 positive T cell. In some embodiments, the T cell is a CD8 positive T cell. In some embodiments, the non-cancerous cell is a B cell. In some embodiments, the non-cancerous cell is an erythrocyte progenitor cell. In some embodiments, the non-cancerous cell is a platelet. In some embodiments, the non-cancerous cell is a plasma cell. In some embodiments, the plasma cell is a circulating plasma cell. In some embodiments, the plasma cell is a bone marrow plasma cell. In some embodiments, the non-cancerous cell is an endothelial cell. In some embodiments, the non- cancerous cell is a bone cell. In some embodiments, the bone cell is an osteocyte. In some embodiments, the non-cancerous cell is an adipocyte. In some embodiments, the non- cancerous cell is an endothelial cell. In some embodiments, the non-cancerous cell is a bladder cell. In some embodiments, the non-cancerous cell is a keratinocyte. In some embodiments, the non-cancerous cell is fallopian cell. In some embodiments, a fallopian cell is a fallopian tube cell. In some embodiments, the non-cancerous cell is gallbladder cell. In some embodiments, the non-cancerous cell is a cholangiocyte. In some embodiments, the non-cancerous cell is lung cell. In some embodiments, the lung cell is a lung alveolar cell. In some embodiments, the non-cancerous cell is an ovary cell. In some embodiments, the non-cancerous cell is a prostate cell. In some embodiments, the non-cancerous cell is white blood cell (WBC). In some embodiments, the bladder cell, breast cell, colon cell, esophageal cell, fallopian cell, gallbladder cell, kidney cell, larynx cell, lung cell, ovary cell, prostate cell, and / or thyroid cell is an epithelial cell.
[0116] In some embodiments, the method further comprises administering an anticancer therapy to a subject determined to have cancer. In some embodiments, determined to have cancer is detected to have cancer. In some embodiments, determined to have cancer is diagnosed with cancer. In some embodiments, anticancer therapy is anti-MM therapy. In some embodiments, the therapy is a drug. In some embodiments, the therapy is radiotherapy. In some embodiments, the therapy is surgery. In some embodiments, the therapy isimmunotherapy. In some embodiments, the drug is chemotherapy. In some embodiments, the chemotherapy is an alkylating agent. In some embodiments, the alkylating agent is cyclophosphamide. In some embodiments, the immunotherapy is an immune checkpoint inhibitor (ICI). In some embodiments, the immunotherapy is an immunomodulator. In some embodiments, the immunomodulator is lenalidomide. In some embodiments, the immunomodulator is thalidomide. In some embodiments, the therapy is bone marrow transplant. In some embodiments, bone marrow transplant is bone marrow replacement. In some embodiments, the therapy is a proteasome inhibitor. In some embodiments, the proteasome inhibitor is bortezomib. In some embodiments, the proteasome inhibitor is bortezomib. In some embodiments, the proteasome inhibitor is carfizomib. In some embodiments, the proteasome inhibitor is ixazomib. In some embodiments, the proteasome inhibitor is selected from bortezomib, carfizomib and ixazomib. In some embodiments, the therapy is a target specific monoclonal antibody. In some embodiments, the antibody is an anti-CD38 antibody. In some embodiments, the anti-CD38 antibody is daratumumab. In some embodiments, the anti-CD38 antibody is isatuximab. In some embodiments, the daratumumab is administered in combination with hyaluronidase. In some embodiments, the therapy is administered in combination with a steroid. In some embodiments, the steroid is dexamethasone. In some embodiments, the therapy is a combination of bortezomib and lenalidomide. In some embodiments, the therapy is a combination of bortezomib, lenalidomide and dexamethasone. In some embodiments, the therapy is a combination of bortezomib, thalidomide and dexamethasone. In some embodiments, the therapy is a combination of bortezomib, cyclophosphamide and dexamethasone. In some embodiments, the therapy is a combination of daratumumab, isatuximab, bortezomib and at least one of lenalidomide, thalidomide and cyclophosphamide. In some embodiments, the therapy is a combination of daratumumab, isatuximab, bortezomib, at least one of lenalidomide, thalidomide and cyclophosphamide and dexamethasone.
[0117] As used herein, the terms “administering,” “administration,” and like terms refer to any method which, in sound medical practice, delivers a composition containing an active agent to a subject in such a manner as to provide a therapeutic effect. One aspect of the present subject matter provides for oral administration of a therapeutically effective amount of a composition of the present subject matter to a patient in need thereof. Other suitable routes of administration can include parenteral, subcutaneous, intravenous, intramuscular, or intraperitoneal.
[0118] The dosage administered will be dependent upon the age, health, and weight of the recipient, kind of concurrent treatment, if any, frequency of treatment, and the nature of the effect desired.
[0119] As used herein, the terms “treatment” or “treating” of a disease, disorder, or condition encompasses alleviation of at least one symptom thereof, a reduction in the severity thereof, or inhibition of the progression thereof. Treatment need not mean that the disease, disorder, or condition is totally cured. To be an effective treatment, a useful composition or method herein needs only to reduce the severity of a disease, disorder, or condition, reduce the severity of symptoms associated therewith, or provide improvement to a patient or subject’s quality of life.
[0120] By another aspect, there is provided a method of detecting DNA from a plasma cell, the method comprising: a. receiving DNA methylation measurements at least one genomic region selected from the group consisting of: chr2: 128090328-128090485, chr6:53139872-53139964, chr6: 109220925-109220980, chr20:24995553- 24995799, chr6:7880318-7880470, chrl3:113246637-113246717, and chrl9: 14522665- 14522768, wherein said coordinates are with respect to human genome build Hgl9; and b. assigning a DNA molecule as being from a plasma cell when said genomic region comprises unmethylation of all CpGs; thereby detecting DNA from a plasma cell.
[0121] By another aspect, there is provided a method for quantifying molecules of cfDNA having a methylation pattern of a plasma cell, comprising: a. amplifying a DNA sequence selected from the group consisting of: chr2: 128090328-128090485, chr6:53139872-53139964, chr6: 109220925- 109220980, chr20:24995553-24995799, chr6:7880318-7880470, chrl3:113246637-113246717, and chrl9: 14522665- 14522768, wherein said coordinates are with respect to human genome build Hgl9, in the cfDNA of said sample to obtain amplified DNA molecules; b. using sequencing to obtain the sequence of individual molecules of the amplified DNA molecules;c. ascertaining from the sequenced amplified DNA molecules the methylation status of at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecules of the sequenced amplified DNA molecules; and d. quantifying the amount of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a plasma cell that is present in the total of the amplified DNA molecules, wherein said methylation pattern in a plasma cell comprises all of the at least four sites being unmethylated; thereby quantifying molecule of cfDNA having the methylation pattern of a plasma cell.
[0122] In some embodiments, the DNA is in a sample. In some embodiments, the detecting is in a sample. In some embodiments, the quantifying is in a sample. In some embodiments, the sample is from the subject. In some embodiments, the sample is derived from the subject. In some embodiments, the amplifying is in the sample. In some embodiments, the cfDNA is of the sample. In some embodiments, the cfDNA is cfDNA in the sample.
[0123] In some embodiments, the at least one genomic region is SEQ ID NO: 1. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 1. In some embodiments, the at least one genomic region is SEQ ID NO: 2. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 2. In some embodiments, the at least one genomic region is SEQ ID NO: 3. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 3. In some embodiments, the at least one genomic region is SEQ ID NO: 4. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 4. In some embodiments, the at least one genomic region is SEQ ID NO: 5. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 5. In some embodiments, the at least one genomic region is SEQ ID NO: 6. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 6. In some embodiments, the at least one genomic region is SEQ ID NO: 7. In some embodiments, the at least one genomic region comprises or consists of SEQ ID NO: 7.
[0124] By another aspect, there is provided a method of diagnosing myeloma in a subject, the method comprising: a. receiving DNA methylation measurements from a sample from the subject; andb. quantifying the number of cfDNA molecule from a plasma cell in the sample, wherein a quantity of cfDNA molecules from a plasma cell in the sample above a predetermined threshold indicates the subject suffers from myeloma; thereby diagnosing myeloma in a subject.
[0125] By another aspect, there is provided a method of predicting progression to multiple myeloma in a subject suffering from SMM or MGUS, the method comprising: a. receiving DNA methylation measurements of cfDNA from a sample from the subject; and b. quantifying the number of cfDNA molecules from a plasma cell in the sample, wherein a quantity of cfDNA molecule form a plasma cell above a predetermined threshold indicates the subject is predicted to progress to MM; thereby predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
[0126] By another aspect, there is provided a method of diagnosing myeloma in a subject, the method comprising: a. receiving measurements of at least one feature from a sample from the subject, wherein the feature is selected from the following 6 features: i. the number of cfDNA molecules from a plasma cell in the sample; ii. the percentage of discordant methylation reads (PDR) at genomic location chrl l:107, 747, 505-107 ,747, 531 wherein the coordinates are with respect to human genome build HG19; iii. the entropy of the methylation at genomic location chrl 1:107,747,505- 107,747,531 wherein said coordinates are with respect to human genome build HG19; iv. the entropy of the methylation at genomic location chr6:80, 356, 706-80, 356, 741 wherein said coordinates are with respect to human genome build HG19;v. the entropy of the methylation at genomic location chrl5:41, 072, 702-41, 072, 749 wherein said coordinates are with respect to human genome build HG19; and vi. whether the free light chain (FLC) ratio is greater than or less than 20; and b. applying a trained machine learning (ML) model to the at least one feature, wherein the trained machine learning model is trained on the at least one feature from subjects with MM and the at least one feature from control subjects; thereby diagnosing myeloma in a subject.
[0127] By another aspect, there is provided a method of predicting progression to multiple myeloma in a subject suffering from SMM or MGUS, the method comprising: a. receiving measurements of at least one feature from a sample from the subject, wherein the feature is selected from the following 6 features: i. the number of cfDNA molecules from a plasma cell in the sample; ii. the percentage of discordant methylation reads (PDR) at genomic location chrl l:107, 747, 505-107 ,747, 531 wherein the coordinates are with respect to human genome build HG19; iii. the entropy of the methylation at genomic location chrl 1:107,747,505- 107,747,531 wherein said coordinates are with respect to human genome build HG19; iv. the entropy of the methylation at genomic location chr6:80, 356, 706-80, 356, 741 wherein said coordinates are with respect to human genome build HG19; v. the entropy of the methylation at genomic location chrl5:41, 072, 702-41, 072, 749 wherein said coordinates are with respect to human genome build HG19; and vi. whether the free light chain (FLC) ratio is greater than or less than 20; andb. applying a trained machine learning (ML) model to the at least one feature, wherein the trained machine learning model is trained on the at least one feature from subjects with MM and the at least one feature from subjects with SMM, subjects with MGUS or both; thereby predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
[0128] In some embodiments, quantifying is quantifying the number of molecules having the methylation pattern of a plasma cell. In some embodiments, the quantifying detecting DNA from a plasma cell. In some embodiments, the quantifying is by a method of the invention. In some embodiments, measurements are quantifications. In some embodiments, the method comprises quantifying the at least one feature. In some embodiments, the method comprises receiving a fluid sample from the subject. In some embodiments, the method comprises measuring the at least one feature in the received fluid sample.
[0129] In some embodiments, at least one feature is a plurality of features. In some embodiments, at least one feature is at least 3 features. In some embodiments, at least one feature is at least 4 features. In some embodiments, at least one feature is at least 5 features. In some embodiments, at least one feature is all 6 features.
[0130] In some embodiments, a machine learning model is a machine learning algorithm. ML models and algorithms are well known in the art and a skilled artisan can select the appropriate ML model / algorithm to use. In some embodiments, the algorithm is a supervised learning algorithm. In some embodiments, the algorithm is an unsupervised learning algorithm. In some embodiments, the algorithm is a reinforcement learning algorithm. In some embodiments, the machine learning model is a Convolutional Neural Network (CNN). In some embodiments, the at least one hardware processor trains a machine learning model. In some embodiments, the model is based, at least in part, on a training set. In some embodiments, the model is based on a training set. In some embodiments, the model is trained on a training set. In some embodiments, the at least one hardware processor applies the machine learning model to a factor expression level from a subject.
[0131] In some embodiments, the training set comprises the at least one feature from subject with MM. In some embodiments, the training set comprises the at least one feature from control subjects. In some embodiments, control subjects are healthy subjects. In some embodiments, control subjects are subjects with SMM. In some embodiments, controlsubjects are subjects with MGUS. In some embodiments, the training set comprises the at least one feature from subjects with SMM. In some embodiments, the training set comprises the at least one feature from subject with MGUS. In some embodiments, the training set comprises the at least one feature from subjects with SMM and subject with MGUS. In some embodiments, the training set comprises the at least one feature from subjects with SMM, subjects with MGUS and healthy subjects. In some embodiments, the training set comprises labels indicating the subject that provided the sample.
[0132] In some embodiments, the ML model outputs a diagnosis. In some embodiments, the ML model outputs a prediction. In some embodiments, the diagnosis is MM or not MM. In some embodiments, not MM is no MM. In some embodiments, not MM is an absence of MM. In some embodiments, the prediction is progression to MM. In some embodiments, the prediction is not progression to MM. In some embodiments, progression is progressing. In some embodiments, the prediction is a diagnosis.
[0133] In some embodiments, the method further comprises treating a subject diagnosed with MM. In some embodiments, the method further comprises treating a subject predicted of progressing to MM. In some embodiments, the treating is with an anticancer therapy. In some embodiments, the anticancer therapy is an anti-MM therapy. In some embodiments, treating comprises administering the therapy to the subject.
[0134] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.
[0135] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.
[0136] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" wouldinclude but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0137] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0138] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.
[0139] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES
[0140] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "CurrentProtocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I- III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes LIII Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.MATERIALS AND METHODS
[0141] Subject enrollment: This study was conducted according to protocols approved by the Institutional Review Board, with procedures performed in accordance with the Declaration of Helsinki. Clinical data, blood, bone marrow and tissue samples were obtained from donors who have provided written informed consent. Patient characteristics are presented in Table 3. MGUS, SMM, and newly diagnosed MM, were diagnosed according to the The International Myeloma Working Group (IMWG) criteria. Patients were allowed to have been in routine follow up for the clinical data. Patients with concomitant second primary malignancies were excluded. The patient cohort data was collected prospectively. Most PB cfDNA samples were taken in parallel to the BM samples. Biochemical progression was calculated utilizing criteria in accordance with the IMWG criteria. Clinical progression was defined as patients necessitating treatment.
[0142] Table 3: Patient baseline characteristics - other non-plasma cell related disorders and laboratory data.Disease p.valueHypertension 28 (48.3) 12 (42.9) 14 (38.9) 0.66Dyslipidemia 16 (27.6) 4 (14.3) 10 (27.8) 0.35Diabetes Mellitus 9 (15.5) 6 (21.4) 10 (27.8) 0.35Chronic Kidney 10 (17.2) 1 (3.6) 8 (22.2) 0.09DiseaseRheumatological / 8 (13.8) 7 (25) 2 (5.6) 0.09Immunological DiseaseS / P Oncology 4 (6.9) 5 (17.9) 7 (19.4) 0.13Creatinine mg / dL 0.94 (0.6-2.88) 0.93 (0.61-1.5) 0.99 (0.54-13.24) 0.57Albumin g / dL 4.2 (3.5-5.1) 4.2 (3.8-4.7) 4.2 (2.9-5.2) 0.91HGB g% 13.6 (9.1-18.9) 13.4 (9.4-15.6) 12 (6.8-15.5) <0.0001
[0143] Plasma cell isolation: Plasma cells were isolated from bone marrow aspiration of a healthy individual that underwent hip replacement surgery, and from the routine BM aspirations of MGUS, SMM and newly diagnosed MM patients. Plasma cells from patient samples were enriched using CD138+-beads and autoMACSpro separator (Miltenyi biotec)
[0144] Sample collection and processing: Blood samples were collected by routine venipuncture in 10 mL EDTA Vacutainer® tubes or Streck® blood collection tubes and stored at room temperature for up to 4 hours or 5 days, respectively. Tubes were centrifuged at l,500xg for 10 minutes at 4°C (EDTA tubes) or at room temperature (Streck tubes). The supernatant was transferred to a fresh 15 mL conical tube without disturbing the cellular layer and centrifuged again for 10 minutes at 3000xg. The supernatant was collected and stored at -80°C. cfDNA was extracted from 2-4 mL of plasma using the QIAsymphony liquid handling robot (Qiagen). cfDNA concentration was determined using Qubit double-strand molecular probes kit (Invitrogen) according to the manufacturer’s instructions.
[0145] According to the manufacturer's instructions, DNA derived from all samples was treated with bisulfite using EZ DNA Methylation-Gold™ (Zymo Research) and eluted in 24pl elution buffer.
[0146] WGBS: Up to 75 ng of sheared genomic DNA from sorted plasma cells was subjected to bisulfite conversion using the EZ-96 DNA Methylation Kit (Zymo Research), with liquid handling on a MicroLab STAR (Hamilton). Dual-indexed sequencing libraries were prepared using Accel-NGS Methyl-Seq DNA library preparation kits (Swift BioSciences) and custom liquid handling scripts executed on the Hamilton MicroLab STAR. Libraries were quantified using KAPA Library Quantification Kits for Illumina Platforms(Kapa Biosystems). Four uniquely dual-indexed libraries, along with the 10% PhiX v.3 library (Illumina), were pooled and clustered on an Illumina NovaSeq 6000 S2 flow cell followed by 150 bp, paired-end sequencing.
[0147] WGBS computational processing: Paired-end FASTQ files were mapped to the human (hgl9, hg38), lambda, pUC19 and viral genomes using bwa-meth (v.0.2.0) then converted to BAM files using SAMtools (v.1.9). Duplicated reads were marked by Sambamba (v.0.6.5) with parameters ‘-1 1 -t 16 —sort-buffer-size 16000 — overflow-list-size 10000000’. Reads with low mapping quality, duplicated or not mapped in a proper pair were excluded using SAMtools view with parameters ‘-F 1796 -q 10’. Reads were stripped from nonCpG nucleotides and converted to PAT files using wgbstools (v.0.1.0).
[0148] Next-Generation Sequencing and analysis: Pooled PCR products were subjected to multiplex next-generation sequencing (NGS) using the NextSeq 500 / 550 v2 Reagent Kit (Illumina). Sequenced reads were separated by barcode, aligned to the target sequence, and analyzed using custom scripts written and implemented in R. Reads were quality filtered based on Illumina quality scores. Reads were identified as having at least 80% similarity to the target sequences and containing all the expected CpGs. CpGs were considered methylated if “CG” was read and unmethylated if “TG” was read. Proper bisulfite conversion was assessed by analyzing methylation of non-CpG cytosines. We then determined the fraction of molecules in which all CpG sites were unmethylated. The fraction obtained was multiplied by the concentration of cfDNA measured in each sample, to obtain the concentration of tissue-specific cfDNA from each donor. Given that the mass of a haploid human genome is 3.3 picograms, the concentration of cfDNA could be converted from units of ng / ml to haploid genome equivalents / ml (GE / ml) by multiplying by a factor of 303.
[0149] Plasma cell marker identification: Plasma cell- specific methylation candidate biomarkers were identified through comparative methylome analysis using the previously published human methylation atlas (Loyfer et al., 2023) and unpublished WGBS data from plasma cells derived from healthy individuals, as well as MGUS, SMM, and MM samples. Genomic loci containing more than five CpG sites within a 150 bp region, with an average methylation value of <0.3 in plasma cells and >0.8 in over 90% of other tissues and immune cell types were prioritized. As was previously published (Loyfer et al., 2023), cell-specific methylation genomic loci are far more abundant in an unmethylated form in the genome than loci that are methylated. Through this approach, approximately 50 CpG blocks that were uniquely unmethylated in plasma cells while being methylated in all other major immunecells and tissues were identified. From these ten loci were arbitrarily selected as candidate markers for further validation. Three markers were excluded due to technical limitations, leaving seven robust plasma cell-specific markers, as detailed in Table 1.
[0150] PDR markers identification: To identify myeloma- specific markers, genome-wide methylome analysis of sorted plasma cells from healthy individuals, MGUS, SMM, and multiple myeloma patients was performed alongside other immune cell types and solid organ tissues. Candidate genomic loci were considered eligible myeloma markers if they contained >4 CpG sites within a 160-base-pair region. Loci with sequencing coverage below 10X were excluded.
[0151] The proportion of discordant reads (PDR) was calculated as previously defined by Landau et al., 2014, the contents of which are hereby incorporated by reference in its entirety:*Discordant reads were initially defined as DNA molecules that were not entirely methylated or demethylated. Genomic loci were considered potential markers if they demonstrated a PDR difference of a fraction > 0.3 between myeloma plasma cells and all other cell types. This analysis identified 19 loci meeting these criteria.
[0152] To enable analysis of these loci in cell-free DNA (cfDNA), targeted primers were designed to amplify fragments <160 base pairs in length using a multiplex two-step PCR amplification method.
[0153] While the PDR definition by Landau et al. provided high specificity, it was associated with substantial baseline noise, primarily from molecules with a single unmethylated CpG site. To address this, the definition of discordant reads was refined as: #number of discordant readsPDR= - number of total reads#molecules with >2 CpG site discrepancies within a block, allowing for a single CpG demethylation error while still categorizing the molecule as concordant.
[0154] Clonal evolution characterization using entropy: Entropy as opposed to PDR takes into account the different patterns in each loci. Shannon’s entropy is measured by:Entropy = =- pilog2pi
[0155] An “epiclone” was defined as a distinct methylation pattern that contains a locally disrupted pattern as defined - molecules with >2 CpG site discrepancies within a block. Patterns that were considered non-malignant / concordant were excluded from the analysis. The fraction of each epiclone was quantified relative to the total malignant clones, accounting for differences in the composition of healthy cells within the bone marrow and cfDNA.
[0156] Statistics: To assess the correlation between flow cytometry quantification in the bone marrow and plasma cell-specific methylation quantification Pearson's correlation test was employed, assuming a linear relationship. The correlation between cfDNA and bone marrow measurements was assessed using Spearman's correlation. To determine the significance of differences between groups a non-parametric two-tailed Mann-Whitney test was used. For multiple comparisons, a Kruskal-Wallis multiple comparison test was used. P-value was considered significant when <0.05. All Statistical analyses were performed with GraphPad Prism 10.4.0.
[0157] Kaplan Meier: To evaluate the predictive power of information-theoretic and tissuespecific cfDNA measurements, as well as biochemical measurements in relation to clinically diagnosed rate of progression, classical machine learning techniques similar to Avni et la., 2024, “Chronic Graft-versus-Host Disease Detected by Tissue-Specific Cell-Free DNA Methylation Biomarkers.” The Journal of Clinical Investigation 134(2), the contents of which are hereby incorporated by reference in its entirety, were used. The class of models chosen is accelerated failure time (AFT) models, as there is strong basic science evidence indicating that AFT models are the most suitable model for biological survival processes. Multiple models (namely Weibull, Log-normal, Log-logistic and exponential) were compared on the dataset, and very similar results were obtained with each of them, with Weibull yielding slightly better results than the rest. This was expected since these distributions have very similar shapes, i.e. can approximate each other very well. In addition, using a Generalized Gamma model was tried, however, the predictions it gave posited an infinite survival time for all patients, rendering it unusable.
[0158] Furthermore, Weibull AFT (WAFT) emerged as the most fitting estimator based on the following considerations:(a) The size of the data does not support models with a large number of parameters — WAFT bears a single parameter per measurement, reducing the risk of overfitting.(b) WAFT inference naturally provides a survival curve, allowing for any percentile to be calculated easily.
[0159] It was hypothesized that information-theoretic measurements (PDR and entropy) of cfDNA and blood biochemical values possess significant predictive potential for the progression rate from MGUS and SMM to MM.
[0160] Feature Selection: The feature selection was done very similarly to Avni et la., 2024, with several key differences:• Application of PCA to eliminate collinearity of the input features: Since p-values tend to explode in the presence of strong feature collinearity, PCA is crucial in order to observe the true statistical significance of the model given a set of features. Metrics for model selection: Instead of AUC and accuracy, model selection was done based on the following metrics: o Concordance - the fraction of pairs predicted in the correct order, calculated by 4-fold cross validation with stratified partitions. o P-values - The p-values calculated for the WAFT model by the “lifelines” library. Since the number of positive samples in the dataset is exceedingly small (12), we consider a model with p-values smaller than 0.1 as viable. o Normalized Error (NE)- The relative error of the model predictions.o Prospective Accuracy (PA) - Partitioning the potential time of progression into 4 intervals (0-6, 6-12, 12-24 and >24 months), we calculate the accuracy with which the model classifies the patients time to progression as the correct category.
[0161] A total of 62 samples were used for the analysis, 12 of which were clinically diagnosed with MM at some point after day 0. Shapley analysis on a collection of 37 features (comprising of PDR and entropy for each marker individually, as well as age, sex, M-Protein, FLC Ratio, total cfDNA concentration, plasma cell cfDNA (% and GE / ml) was employed. Next, to robustly validate the predictive potential of the features, Repeated K-Fold cross- validation was utilized. Repeated 3-fold cross-validation was conducted across the feature sets to ensure an equal number of positive samples in each fold. Ranking the features byaverage absolute SHAP values, WAFT models trained using the top 1-10 features were evaluated, and the ones with p-values greater than 10% were rejected. All analyses were done using Python 3.12.
[0162] Using the model, concordance on the train set, cross-validation concordance, normalized error (NE) and prospective accuracy (PA) were measured. Both concordance metrics were calculated to observe the measure of overfitting, while the NE and PA measure the effective predictive potential of the model.Example 1: DNA methylation markers of plasma cells
[0163] To identify genomic regions that are differentially methylated in plasma cells, plasma cells were sorted from the bone marrow (BM) of one healthy individual (N=l) and patients with MGUS (N=4), SMM (N=5) and MM (N=5), and deep whole genome bisulfite sequencing (WGBS) was performed. The resulting methylation profiles were then compared with those of all other cell types available in the previously reported reference methylation atlas (Loyfer et al., 2023), including other B-cell derived subpopulations such as memory B- cells. Regions that are uniquely demethylated or methylated in healthy plasma cells and that remain so across all plasma cell disorders were searched for (Fig. 1A). As previously established for other cell type-specific markers, such regions were typically hypomethylated (i.e. hypomethylated in plasma cells and methylated elsewhere), reflecting the characteristic of such markers as tissue-specific gene enhancers. Seven genomic regions were selected (Table 1) and PCR primers were designed to co-amplify and sequence them, creating unique pan-plasma cell specific markers, to be assayed on the various cells as well as the peripheral blood (PB) cell free DNA (cfDNA). The amplicons selected were shorter than 160bp (to capture the typical nucleosome size of cfDNA fragments) and contained at least five CpGs (to maximize specificity. The specificity of these markers to plasma cells was validated using genomic DNA from multiple sorted cell types and tissues (Fig. IB). Additionally, assay sensitivity was determined by spiking DNA from sorted plasma cells into leukocyte DNA. The assay was able to identify the presence of plasma cell DNA even when comprising 0.12% of the total, or 1 genome equivalents of plasma cells (Fig. 1C).
[0164] Table 1: Plasma cell informative genomic lociExample 2: Liquid biopsy of plasma cell-derived cfDNA is elevated in multiple myeloma
[0165] The performance of these unique plasma cell methylation markers (ACSS1 chr20:24995553-24995799; ARMC2chr6: 109220925-109220980; BMP6 chr6:7880318- 7880470; DDX39A chrl9: 14522665- 14522768; ELOVL5 chr6:53139872-53139964; MAP3K2 chr2: 128090328- 128090485; Plasma 2 chrl3:113246637-113246717 in Hgl9) were tested on bone marrow (BM) aspirates, in which the percentage of plasma cells was quantified using flow cytometry or trephine biopsy immunohistochemical (IHC) pathological analysis (n=9 MGUS samples, 11 SMM and 9 MM). The methylation-based markers showed an almost full correlation with plasma cell analysis via flow cytometry (r=0.98, p-value<0.0001, Pearson’s correlation) (Fig. 2A). The correlation with pathological IHC estimations was weaker (r=0.43, p-value=0.049, Pearson’s correlation) (Fig. 2F), which is expected due to the subjective nature of human estimation and the patchy nature of the disease. Notably, there was a similarly weak correlation between pathology estimates and flow cytometry (r=0.4, p-value=0.073, Pearson’s correlation) (Fig. 2G).
[0166] It was hypothesized that plasma cell -derived cfDNA analysis could be used to monitor plasma cell turnover in the bone marrow of patients with plasma cell disorders, even though plasma cells are typically absent from the systemic circulation. cfDNA isolated from the PB-plasma of collected blood samples from healthy controls (N=86), MGUS (N=54), SMM (N=28) and MM (N=36) was analyzed. The patients baseline characteristics are shown in Table 3. MGUS, SMM and MM patients were differentiated by their diseasecharacteristics as expected, but not by other baseline characteristics or other medical history characteristics Table 3.
[0167] For control purposes, the total concentration of cfDNA was measured, and was found to be higher in MM patients as compared with MGUS and healthy controls (p-value<0.0001, Kruskal-Wallis), but not as compared with SMM (Fig. 2H). When measuring plasma cell cfDNA, by averaging the seven methylation-based plasma cell markers, minimal levels were found in healthy individuals (median=0.7) and MGUS patients (median=0.88). SMM (median=3.48) and MM patients (median=54.5) showed a gradual and significant elevation, with the highest levels in MM patients (Fig. 2B, and 21). Importantly, MM could be separated from pre-malignant states based on the concentration of plasma cell-derived cfDNA, as indicated by ROC curve analysis (healthy vs MM AUC=0.94, p-value<0.0001, MGUS vs MM AUC=0.9, p-value=0.0001; SMM vs MM AUC=0.81, p-value=0.0001) (Fig. 2C). The percentage of plasma cell cfDNA, a more direct measurement, also distinguished MM from health and from premalignant PCD (Fig. 2J).
[0168] To evaluate the potential of plasma cfDNA with non-invasive estimates, and quantify the plasma cells residing in the BM, the percentage of plasma cell cfDNA was correlated with the percentage of plasma cells in the BM, as assessed by flow cytometry on BM aspirates or pathological examination. The correlation between cfDNA and flow cytometry was strong (r=0.68, p<0.0001, Spearman’s correlation, Fig. 2D), and the correlation between cfDNA and pathological assessment was weaker yet still significant (r=0.43, p<0.049, Fig. 2F). The correlation between cfDNA and flow cytometry surpassed the correlation between pathological assessment and flow cytometry (Fig. 2G). Similar findings were observed when plasma cell-derived cfDNA was compared to the percentage of aberrant cells, as defined (Arroz et al., 2016, “Consensus Guidelines on Plasma Cell Myeloma Minimal Residual Disease Analysis and Reporting.” Cytometry Part B: Clinical Cytometry 90(l):31— 39, the contents of which are hereby incorporated by reference in their entirety) and measured by flow cytometry (r=0.68, p<0.0001, Spearman’s correlation) (Fig, 2E), thus correlating also with the phenotypic malignant portion of BM plasma cells. In summary, targeted methylation markers of plasma cells provide a non-invasive cfDNA biomarker that reliably reports on plasma cell involvement of the BM, correlates with their aberrant phenotype, and distinguishes between healthy individuals and patients with MGUS, SMM and MM.Example 3: Targeted cancer-specific markers based on disordered methylation
[0169] The plasma cell-specific methylation markers described above rely on the regional nature of DNA methylation: rather than assessing the methylation status of individual CpG sites, they take advantage of adjacent CpG sites that are concordantly methylated or demethylated in all plasma cells, as a pan-plasma cell signature of all PCDs. As we have shown, this allows one to define markers with extreme specificity to the individual cell types. It was desired to identify markers that will distinguish and be specific to the DNA of the malignant form (MM), rather than to all plasma cells of healthy people or patients with MGUS or SMM. Within the WGBS, there was a failure to identify methylation markers that characterize only the samples of multiple myeloma; this is to be expected, since methylation changes in cancer do not obey a strict developmental program. However, methylation alterations in cancer do show a universal feature: the loosening of ordered methylation (PDR). The pioneering study by Landau et al (Landau et al., 2014) defined a key feature of disrupted methylation in leukemia cells, namely the proportion of sequence reads that show a disordered pattern of methylation. The feature developed in this study, using tumor WGBS, was the PDR, i.e., the fraction of epialleles that contain both methylated and demethylated CpGs in the same sequence read (Fig. 3A, left panel). The same PDR definition was applied to WGBS data from plasma cell disorders, and a search was performed for genomic regions that 1) contain at least 4 CpGs within a range of approximately 160 base pairs; 2) exhibit high PDR in MM; and 3) exhibit low PDR in plasma cells from MGUS and SMM and all other cell types in the reference atlas (APDR>0.3) (Fig. 3B).
[0170] Applying the original PDR definition, the signal -to-noise ratio was low due to high PDR in normal samples, typically caused by one CpG in a region that behaved in a discordant manner. It was hypothesized that a stricter definition of PDR, i.e., the proportion of reads where at least 2 CpG are discordant, may better distinguish cancer cells' methylation patterns from normal cells' methylation patterns (Fig. 3A, right panel). Approximately 20 genomic regions meeting these stringent criteria were identified (Fig. 3C), that were all predominantly methylated. Henceforth, PDR is defined per these criteria.
[0171] Primer pairs were made to test all 20 genomic regions by an ultra-deep PCR- sequencing of these genomic loci, to allow further specific "pan-tumorous marker" analysis of cells and further assay the cfDNA in the PB for PDR. However, only 10 pairs were functional in this non-in silico validation (Table 8). The targeted markers were validated on DNA from multiple tissues and cell types, showing that PDR is higher in isolated plasma cells from PCD patients compared with all other samples (Fig. 3D). Note that some PDRmarkers are high in erythrocyte progenitors, which is expected given the global demethylation occurring during erythrocyte development. As has been reported previously, erythrocyte progenitors contribute about 1% of cfDNA, potentially confounding cfDNA findings when using PDR markers. However, some PDR markers (e.g. SH3BGRL2 and DNAJC17) were not elevated in erythroblasts (Fig. 3D), enabling a correction for the contribution of DNA from these cells.
[0172] To examine assay sensitivity, DNA from sorted plasma cells from a MM patient was spiked into normal leukocyte DNA in decreasing amounts (Fig. 31). PDR markers were also validated in whole bone marrow aspirations. MM patients had higher PDR levels compared to healthy individuals, MGUS and SMM (Kruskal Wallis p-value=0.0018) (Fig. 3H).
[0173] Table 8: Primers usedExample 4: Myeloma-specific disordered methylation as a liquid biopsy
[0174] The new PDR markers (Table 2) were applied to cfDNA from various plasma cell disorders. This revealed a substantial elevation in MM patients (median PDR 0.027; interquartile range 0.016-0.036) compared to healthy individuals and those withpremalignant states (healthy controls 0.013; interquartile range 0.011-0.017, MGUS 0.014; interquartile range 0.01-0.016, SMM 0.014; interquartile range 0.011-0.016; Kruskal-Wallis p<0.0001). Thus, unlike normal plasma-cell derived cfDNA, PDR was similar in healthy samples, MGUS and SMM, suggesting that it marks specifically the DNA originating in malignant myeloma cell turnover rather than the proportion of plasma cells in the bone marrow. Consistent with this idea, it was found that PDR in cfDNA correlated with the proportion of aberrant cells measured by flow cytometry in bone marrow biopsies (Spearman’s correlation r=0.62; Fig. 3F).
[0175] PDR effectively distinguished MM from other premalignant states with good sensitivity and specificity, as indicated by the ROC curves (MM vs healthy AUC=0.83, p<0.0001; MM vs MGUS AUC=0.84, p<0.0001; MM vs SMM AUC=0.8, p<0.0001) (Fig. 3G). While individual markers could separate MM from MGUS and SMM to varying degrees (Fig. 3J, 3L, 3M), combining them yielded the best results (Fig. 3K).
[0176] While there was a strong correlation found between cfDNA methylation patterns, the PDR patterns and among the disease states MGUS, SMM and MM, no correlation was found between each of the baseline disease characteristics and the methylation patterns. Thus, cfDNA methylation patterns and the PDR are highly sensitive disease markers correlating and differentiating among the PCDs.
[0177] Table 2: PDR informative genomic loci, potentially methylated cytosines are boldExample 5: Plasma Cell cfDNA and PDR in cfDNA Predict progression to multiple myeloma
[0178] It was hypothesized that utilizing plasma cell markers and PDR markers will allow prediction of progression. Sixty-six patients with MGUS or SMM were followed prospectively and included in the analysis. At data-cut-off, 21 (31.8%) patients had experienced a 'biochemical progression', and 12 (18.2%) had a 'clinical progression' as defined in the Methods section (Table 4). The association between patient baseline characteristics and progression (PD) to treatment revealed that IgM levels (p=0.004), M protein levels (p =0.046), FLC ratio (p=0.004) and bone marrow plasma cell (BMPC) involvement (p=0.005) (Mann-Whitney) were significantly correlating with PD, this is in line, and was corroborated with the commonly used "2 / 20 / 20" risk stratification criteria (Mateos et al., 2020, “International Myeloma Working Group Risk Stratification Model for Smoldering Multiple Myeloma (SMM).” Blood Cancer Journal 2020 10: 10 10(10): 1-11, the contents of which are hereby incorporated by reference in its entirety) which also strongly correlated with progression. Other risk stratification biomarkers such as cytogenetic risk, presence of immunoparesis and aberrant cell ratio were not correlating with progression. Neither did any other patient baseline characteristics (Table 4) or previous medical history. Assessing plasma cell cfDNA and PDR cfDNA in relation to progression revealed mostly PDR as a strong predictor for both clinical (p=0.03) and biochemical (p=0.045) progression). Applying the “2 / 20 / 20” model to the cohort revealed that those with one or more risk factors were 10-fold more prone to progress to treatment than those with no clinical risk factors (HR=10.06, p=0.0001) (Fig. 4A, 4H, Table 5).
[0179] Table 4: Patient characteristics by means of Clinical and Biochemical Progression.Clinical Progression Biochemical ProgressionPD n=12 No PD n=54 p.value PD n=21 No PD n=45 p.valueGender M,F8<67 / 40.4416<67) / 7 32<7‘ / 130.71(33) (33) (29)MGUS 2 (17) 39 (72) 6 (29) 135 (78)PCD 0.001 0.0001SMM 10 (83) 15 (28) 15 (71) 10 (22)Age, years 71 (35-91) 68 (40-93) 0.75 74 (35-93) 68 (40-89) 0.24IgA 2 (17) 8 (15) 0.86 5 (24) 5 (11) 0.435M- IgG 7(58) 28 (52) 10(48) 25 (56)ProteinIgM0 3 (6) 0 3 (7) lypeLC 3 (25) 15 (28) 6 (29) 12 (37)LC K 8 (73) 35 (65) 13 (62) 30(68)TyPe1 0.622 3 (27) 19 (35) 8 (38) 14 (32)M-Protein g / dL 1.65(0-5) 0.9 (0-3) 0.043 1.2 (0-5) 0.85(0-3) 0.211437 (13- 170120-IgA mg / dL 96(13-1600) 170(19-1373) 0.09 0.831600) 1373)M-Protein >2 5(42) 6(11) 0.038 6(29) 5(11) 0.15FLC Ratio >20 6(50) 6(11) 0.049 7(33) 5(11) 0.042BMPC % >20 2(18) 2(5) 0.477 2(11) 2(6) 0.6BMPC % (IHC) 15 (1-50) 4(1-65) 0.021 10(1-65) 3(1-50) 0.006BMPC % (FC) 4(0-16) 2(0-11) 0.037 3(0-16) 1 (0-8) 0.023FISH CA HR 4(36) 6(17) 0.051 7(41) 4(13) 0.02Immunoparesis 9(75) 23 (43) 0.079 12(57) 20(44) 0.336Shaded: "2 / 20 / 20" risk stratification. Abbreviations: PCD- Plasma cell disorders, FLC Ratio- Free light chain involved to uninvolved ratio, dFLC- Free light chain involved to uninvolved difference, BMPC- Bone marrow plasma cells, IHC- Immunohistochemistry (Pathology), FC- Flowcytometry, LC- Light Chain, HR- High Risk.
[0180] Table 5: cfDNA methylation patterns by progressionProgression To Treatment Biochemical ProgressionPD n=9 No PD n=57 p.value PD n=21 No PD n=45 p.valuePC-cfDNA 7.15 0.59 0.055 3.63 0.497 0.067GE / ml(0-88.21) (0-270.14) (0-26.83) (0-270)PDR 0.015 0.011 0.003 0.014 0.011 0.045(0.013-0.021) (0.003-0.04) (0.006-0.03) (0.003-0.04)Entropy 0.23 0.235 0.77 0.235 0.23 0.74(0.216-0.29) (0.19-0.35) (0.21-0.29) (0.19-0.35)Abbreviations: PC-cfDNA- Plasma cell- cell free DNA, PDR- Proportion of Discordant Reads
[0181] Based on the ROC curve analysis, 6 GE / ml plasma cell markers and 0.013 PDR levels were selected, respectively, as the optimal binary cutoffs for progression over time (identifying patients that progressed to MM with 0.5 sensitivity and 0.8 specificity (plasma cfDNA) and 0.83 sensitivity and 0.7 specificity(PDR cfDNA)). Plasma cell cfDNA and PDR were able to predict both clinical progression (HR=6.81, p<0.0001 and HR=7.4, p=0.0016, respectively) (Fig. 4B-4C) and biochemical progression (HR=2.6, p=0.019 and HR=5.82, p<0.0001, respectively) (Fig. 4I-4J) of MGUS and SMM patients. Although specificity was moderate, the sensitivity of both tests was high (Table 6), translating into a remarkable negative predictive value of 0.95-1.0 (p<0.0001) and a high hazard ratio. A strong negative correlation was found between plasma cell cfDNA levels in those that experienced a biochemical progression (Pearson r=-0.75 p=0.0001) (Fig. 4K-4L) and a clinical progression (Pearson r=-0.63; p=0.03) and the time to progression from the PB sampling time (Fig. 4D-4E). The PDR showed also a high correlation to time for biochemical progression (Pearson r= 0.5; p=0.025) but not for the time for clinical progression (Fig. 4E, 4M-4N).
[0182] Table 6: cfDNA methylation patterns risk stratification utilizing binary thresholds.Sensitivity Specificity PPV NPV Log-rank HR (C.I) p.valueBiochemical ProgressionPC-cfDNA> 3.057 0.57 0.76 0.52 0.79 <0.0001 8.3 (2.8-24.28)PDR> 0.01304 0.67 0.69 0.5 0.82 0.037 2.46 (0.99-6.1)Entropy>0.216 0.1 0.18 0.36 1 0.07Clinical ProgressionPC-cfDNA> 3.057 0.78 0.72 0.3 0.95 0.0001 19.44 (2.37-159.6)PDR> 0.01308 1 0.7 0.35 1.0 0.0004Entropy>0.224 0.56 0.26 0.11 0.79 0.2 0.44 (0.12-1.64)Abbreviations: PC-cfDNA- Plasma cell- cell free DNA, PDR- Proportion of Discordant Reads
[0183] Next, the implications of plasma cell cfDNA and PDR levels for the rate of biochemical progression as evident by longitudinal outpatient clinical assessment was examined. The patients were divided into two groups - those with a faster progression (i.e., progression within one year of follow-up) and those with a slower progression. Plasma cell cfDNA levels showed a trend to discriminate among progressing and non-progressing patients, however this was not significant (p=0.067) (Fig. 40). PDR levels significantly distinguished between fast progressors and slow or non-progressors (p=0.045) (Mann- Whitney) (Fig. 4P).
[0184] The plasma cell and PDR methylation biomarkers were combined with the "2 / 20 / 20" risk stratification model (Mateos et al., 2020). Upon replacement of BMPC involvement with either marker, the prediction maintained its sustainability, not only discriminating between biochemically and clinically progressing patients, but also improving the positive predictive values, without compromising the remarkable negative predictive values (Table 7). Moreover, it was more powerful than the "2 / 20 / 20" alone (Fig. 4F-4G). This suggests that cfDNA methylation markers can be combined with the routine blood predictive MM related biomarkers (i.e. M protein and FLC ratio), to generate a non-invasive risk stratification model avoiding the need for invasive bone marrow biopsies.
[0185] Table 7: Comparisons between 2 / 20 / 20 rule (IR & HR) and replacement of BMPC component with cfDNA methylation patterns.Sensitivity J Sp 1 ecificity J PPV NPV pI.o^‘ valruaenkHRBiochemical Progression2 / 20 / 20 0.53 0.78 0.56 0.76 0.038 2.41 (0.98-5.95)2 / 20 / cfDNA 0.62 0.64 0.45 0.78 0.006 3.16 (1.3-7.72)2 / 20 / PDR 0.71 0.64 0.48 0.83 0.044 2.46 (0.95-6.35)2 / 20 / Entropy 1 0.13 0.35 1 0.177Clinical Progression2 / 20 / 20 0.88 0.77 0.39 0.97 0.001 14.52 (1.78-118.15)2 / 20 / cfDNA 0.89 0.63 0.28 0.97 0.001 14.15 (1.74-114.78)2 / 20 / PDR 1 0.65 0.31 1 0.0012 / 20 / Entropy 0.89 0.21 0.15 0.92 0.61 1.72 (0.21-13.75)Abbreviations: IR- Intermediate Risk, HR- High Risk, PC-cfDNA- Plasma cell- cell free DNA, PDR- Proportion of Discordant Reads
[0186] Taken together, these results reveal plasma cell cfDNA and PDR, taken in a single blood sampling at diagnosis, as powerful noninvasive biomarkers that can distinguish between PCD states and also predict biochemical and more importantly clinical progression.Example 6: Methylation entropy in targeted cfDNA markers predicts progression and informs clonal turnover dynamics
[0187] The definition of PDR sums up information on locally disordered methylation that effectively captures the presence of tumorigenic DNA in a sample. However, this summation makes PDR blind to the richness of variants of disordered methylation. In fact, PDR would provide the same value when a tumor consists of a single clone (in which all cells show exactly the same pattern of methylation disorder), or multiple clones each with distinct patterns of disorder. It was hypothesized that capturing the full richness of methylation patterns among the PDR markers would inform on the epi-clonal composition of tumors and the turnover dynamics of individual clones. For such measurements, Shannon’s entropy was used (Jenkinson et al., 2017, “Potential Energy Landscapes Identify the Information- Theoretic Nature of the Epigenome.” Nature Genetics 2017 49:5 49(5):719-29; Koldobskiy et al., 2020, “A Dysregulated DNA Methylation Landscape Linked to Gene Expression inMLL-Rearranged AML.” Epigenetics 15(8):841— 58) - a means of assessing methylation disorder that considers the fraction of each methylation pattern in a mixture (schematic Fig. 5A).
[0188] The PDR markers showed elevated entropy in the cfDNA of MM patients (median=0.18) compared to premalignant plasma cell disorders (median=0.09-0.11; p<0.0001, Kruskal -Wallis) (Fig. 5B). Samples with high methylation entropy exhibited a variety of epi-clones, while those with low entropy had fewer, dominant epi-clones (Fig. 5B). With a cutoff of 0.091 (found by ROC curve, Fig. 5C), entropy in cfDNA was able to predict biochemical (HR 4.8, p=0.0037) and clinical progression (HR 5.27, p=0.0013) (Fig. 5D), although not as effectively as PDR and plasma cell cfDNA. As expected, there was a strong correlation between the PDR and entropy values, both in bone marrow samples and in cfDNA (Spearman’s r=0.99-l).
[0189] Strikingly, the correlation between entropy in the bone marrow and cfDNA of the same patients was weak, as was the correlation between PDR in bone marrow and cfDNA (Spearman’s r=0.32-0.35) (Fig. 5E). It was speculated that this reflected differential turnover dynamics of distinct clones within the tumor. Indeed, in some patients there was a similar distribution of epi-clones in the bone marrow and in cfDNA, while in others cfDNA was dominated by one disordered methylation pattern (Fig. 5F). This concept is further illustrated with two MM patients who had both BM and cfDNA samples analyzed. Each epi-clone is color-coded to track clonal dynamics. In MM Patient 52, cfDNA closely mirrored the clonal composition observed in the BM (Spearman’s R = 0.92). In contrast, MM Patient 98 showed a weak correlation (Spearman’s R = 0.11), suggesting that clonal turnover in cfDNA did not reflect the actual fraction sizes in the BM. Notably, the black-colored clone exhibited a dominant turnover rate in cfDNA, despite being minimally represented in the BM, highlighting differential clonal dynamics (Fig. 5G). These findings suggest that the diversity of epi-clones represented in cfDNA (and quantified by entropy relative to PDR) reflects differential turnover dynamics of distinct epi-clones in the bone marrow. Thus, an epi-clone that makes a small fraction of tumor DNA in the bone marrow but a large fraction of tumor DNA in cfDNA is likely undergoing expansion, providing a surprising non-invasive insight into intra-tumoral dynamics.
[0190] In summary, methylation entropy as quantified in cfDNA allows for distinguishing among myeloma clones, refining the assessment of MGUS and SMM biochemical and clinical progression, and provides an insight into the epigenetic landscape of MM subclones.Example 7: Validation of performance of plasma cell and PDR markers
[0191] Two independent sets of experiments were performed to validate these findings. First, statistical 'k-fold' cross-validation was performed on the prospective cohort, to assess the validity of predictions of progression.
[0192] Shapley analysis was employed on 37 clinical and cfDNA features (see Methods) to yield maximal SHAP values for the following 10 features: SH3BGRL2 entropy, DNAJC17 entropy, PDR9 entropy, ASB7 entropy, PDR3 entropy, PDR9 PDR, PDR3 PDR, FLC Ratio>20 (binary), FLC Ratio (continuous) and plasma cell cfDNA percentage. Figure 6A shows distribution graphs of the Shapley values of each feature with respect to the time to progression to MM, and Figure 6B shows the average absolute Shapley values for each feature. 4-fold cross-validation was conducted across these 10 feature sets, starting from all 10, and sequentially removing the feature with the lowest average Shapley value. Notably, the only set resulting in a maximal p-value <0.1 was the top 6, consisting of SH3BGRL2 entropy, DNAJC17 entropy, PDR9 entropy, PDR9 PDR, FLC Ratio>20 (binary) and plasma cell cfDNA percentage. Therefore these 6 were used as the cfDNA features for machine learning (ML) analysis.
[0193] Finally, 4 sets of features were compared:• Cell-free: the top 6 features in cfDNA markers chosen above (Max p-value=0.08).• Clinical: Risk as traditionally calculated by clinicians (2 / 20 / 20 risk stratification) (Max p-value=0.2).• All: Concatenation of the previous 2 sets (Max p-value=0.9).• Non-invasive: The top 6 features + M-Protein+ FLC ratio (Max p-value=0.8) - excluding PC % in bone marrow.
[0194] The train concordance, cross-validation concordance, normalized error (NE) and prospective accuracy (PA) for using each set as ML parameters (see Material and Methods) are shown in Figure 6C. The ‘cell-free’, ‘non-invasive’ and ‘all’ feature sets preform similarly in overall metrics, with ‘cell-free’ outperforming the rest on PA (0.75). Furthermore, the ‘clinical’ feature set is slightly more robust in terms of overfitting (train set concordance is identical to cross-validation concordance), however, it yields significantly worse results on NE (0.68) and PA (0.5).
[0195] Second, plasma samples from an independent cohort of patients with PCD were analyzed using the same biomarkers to differentiate MM from MGUS and SMM. This cohort included 16 MGUS, 11 SMM, and 18 MM patients. Both PC cfDNA methylation-based markers and PDR cfDNA markers were applied, and the initial results were confirmed: PC cfDNA and PDR cfDNA levels were significantly elevated in MM compared to MGUS and SMM (p < 0.0001, p = 0.005, Kruskal-Wallis test) (Fig. 6D-6G).
[0196] Applying the optimal PC cfDNA cutoff from the training cohort (Youden’s Index: PC > 7 copies / ml), 100% sensitivity and 60% specificity was achieved in the validation cohort. For PDR cfDNA, applying a cutoff of PDR fraction > 0.0176 yielded a sensitivity of 78% and a specificity of 64%, further supporting the robustness of these markers in distinguishing MM from other premalignant conditions.
[0197] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
Claims
CLAIMS:
1. A method of detecting multiple myeloma (MM) in a subject in need thereof, the method comprising a. receiving a fluid sample from said subject comprising cell free DNA (cfDNA); b. quantifying a number of cfDNA with discordant methylation reads in said sample, wherein said quantifying comprises ascertaining the methylation status of at least four methylation sites in the same double-stranded cell-free DNA (cfDNA) molecule, wherein said at least four methylation sites are in a contiguous nucleotide sequence which comprises no more than 300 base pairs; and wherein the presence of at least two sites that are methylated and at least two sites that are unmethylated in said at least four methylation sites indicating a cfDNA with discordant methylation; and c. comparing the percentage of cfDNA molecule in said sample with discordant methylation reads with the percentage of cfDNA molecules with discordant methylation reads in a control sample, wherein a higher percentage of discordant methylation reads in said received sample as compared to said control sample indicates said subject suffers from MM; thereby detecting MM in a subject.
2. The method of claim 1, wherein said contiguous nucleotide sequence is a genomic locus which is completely methylated or completely unmethylated in non-cancerous cells.
3. The method of claim 1 or 2, wherein said contiguous nucleotide sequence is a genomic locus selected from the group consisting of: chrl5:101, 168, 755-101, 168, 812, chr7:134, 597, 197-134, 597, 250, chrl5:91, 119,606-91, 119,642, chr2:85,005,386-85,005,443, chrl5:41, 072, 702-41, 072, 749, chr7:23, 150, 939-23, 151, 010, chr5:81, 968, 704-81, 968, 773, chrll:5, 225, 201-5, 225, 231, chrl 1:107,747,505-107,747,531 and chr6:80, 356, 706-80, 356, 741, wherein said coordinates are with respect to human genome build HG19.
4. The method of claim 3, wherein said contiguous nucleotide sequence is a sequence selected from the group consisting of SEQ ID NO: 8-17.
5. The method of claim 3 or 4, wherein said detecting the presence of at least two sites comprises detecting all cfDNA molecules comprising said genomic locus and determining the percentage of said cfDNA molecules comprising said genomic locusthat comprise at least two sites that are methylated and at least two sites that are unmethylated thereby determining a percentage of discordant reads (PDR) for said genomic locus.
6. The method of claim 5, comprising determining a PDR for at least two of said genomic loci, calculating an average PDR for said at least two of said genomic loci and wherein an average PDR above a predetermined threshold indicates said subject suffers from cancer.
7. The method of claim 6, wherein said at least two genomic loci are all ten of said genomic loci.
8. The method of any one of claims 1 to 7, wherein said ascertaining is affected by contacting the cfDNA molecule in the sample with bisulfite to generate singlestranded DNA molecules of which demethylated cytosines of said single- stranded DNA molecules are converted to uracils, further contacting said single- stranded DNA with amplification primers under conditions that generate amplified DNA from said single- stranded DNA following said contacting with said bisulfite and sequencing said amplified DNA.
9. The method of claim 8, wherein said sequencing is deep sequencing, massively parallel sequence or next-generation sequencing.
10. The method of any one of claims 1 to 9, wherein said control sample is a sample from a healthy subject, or a sample from a subject suffering from smoldering multiple myeloma (SMM) or monoclonal gammopathy of undetermined significance (MGUS).
11. The method of claim 10, wherein said control sample is a sample from a subject suffering from SMM or MGUS and said method is a method of diagnosing or predicting progression to MM in subject with SMM or MGUS.
12. A method for quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a multiple myeloma (MM) cell in a fluid sample derived from a subject, comprising: a. identifying a DNA sequence of 10-300 base pairs that has at least four methylation sites, wherein the methylation pattern of said at least four methylation sites in a MM cell is different as compared to the methylation pattern of said at least four methylation sites in the same DNA sequence in non-MM cells, wherein said methylation pattern in a MM cell comprises at least two sites that are methylated and at least two sites that are unmethylatedand said methylation pattern in non-MM cells comprises all sites being methylated or all sites being unmethylated; b. amplifying said identified DNA sequence in the cfDNA of said sample to obtain amplified DNA molecules; c. using deep sequencing, massively parallel sequencing or next-generation sequencing to obtain the sequence of individual molecules of said amplified DNA molecules; d. ascertaining from said sequenced amplified DNA molecules the methylation status of each of said at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecules of said sequenced amplified DNA molecules of said sample; and e. quantifying the number of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a MM cell that are present in the total of said amplified DNA molecules; thereby quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a MM cell.
13. The method of claim 12, wherein said DNA sequence is a genomic locus selected from the group consisting of: chrl5:101, 168, 755-101, 168, 812, chr7: 134,597, 197- 134,597,250, chrl5:91,l 19,606-91,119,642, chr2:85, 005, 386-85, 005, 443, chrl5:41, 072, 702-41, 072, 749, chr7:23, 150, 939-23, 151, 010, chr5:81,968,704-81,968,773, chrll:5, 225, 201-5, 225, 231, chrll:107, 747, 505-107, 747, 531 and chr6:80, 356, 706-80, 356, 741, wherein said coordinates are with respect to human genome build Hgl9.
14. The method of claim 13, wherein said DNA sequence is a sequence selected from the group consisting of SEQ ID NO: 8-17.
15. The method of any one of claims 1 to 14, wherein said fluid sample is selected from the group consisting of blood, plasma, sperm, milk, urine, saliva and cerebral spinal fluid.
16. The method of claim 15, wherein said fluid sample is a plasma sample.
17. The method of any one of claims 12 to 16, further comprising contacting the DNA of the sample with bisulfite to convert demethylated cytosines of the DNA to uracils prior to said amplifying step (b).
18. The method of any one of claims 1 to 17, wherein said sequence is no more than 167 base pairs.
19. The method of any one of claims 12 to 18, wherein, in step (c), the sequences of at least 10,000 individual molecules are obtained.
20. The method of any one of claims 12 to 19, wherein said MM cells are selected from SMM cells and MGUS cells.
21. The method of any one of claims 1 to 20, further comprising administering an anti- MM drug to a subject detected to have cancer.
22. A method of detecting DNA from a plasma cell in a sample, the method comprising: a. receiving DNA methylation measurements of DNA from a sample in at least one genomic region selected from the group consisting of: chr2: 128090328- 128090485, chr6:53139872-53139964, chr6: 109220925-109220980, chr20:24995553-24995799, chr6:7880318-7880470, chrl3: 113246637-113246717, and chrl9: 14522665- 14522768, wherein said coordinates are with respect to human genome build Hgl9; and b. assigning a DNA molecule as being from a plasma cell when said genomic region comprises unmethylation of all CpGs; thereby detecting DNA from a plasma cell is a sample comprising DNA.
23. A method for quantifying molecules of cell-fee DNA (cfDNA) having the methylation pattern of a plasma cell in a sample derived from a subject, comprising: a. amplifying a DNA sequence selected from the group consisting of: chr2: 128090328-128090485, chr6:53139872-53139964, chr6: 109220925- 109220980, chr20:24995553-24995799, chr6:7880318-7880470, chrl3:113246637-113246717, and chrl9: 14522665-14522768, wherein said coordinates are with respect to human genome build Hgl9, in the cfDNA of said sample to obtain amplified DNA molecules; b. using deep sequencing, massively parallel sequencing or next-generation sequencing to obtain the sequence of individual molecules of said amplified DNA molecules; c. ascertaining from said sequenced amplified DNA molecules the methylation status of at least four methylation sites per DNA molecule, thereby obtaining the methylation pattern of each of the individual DNA molecules of said sequenced amplified DNA molecules of said sample; andd. quantifying the amount of amplified DNA molecules having the specific methylation pattern of a DNA molecule of a plasma cell that is present in the total of said amplified DNA molecules, wherein said methylation pattern in a plasma cell comprises all of said at least four sites being unmethylated; thereby quantifying molecule of cell-free DNA (cfDNA) having the methylation pattern of a plasma cell.
24. The method of claim 22 or 23, wherein said at least one genomic region or said DNA sequence is selected from the group consisting of: SEQ ID NO: 1-7.
25. The method of any one of claims 22 to 24, wherein said at least one genomic region is all seven regions.
26. A method of diagnosing multiple myeloma or predicting progression to multiple myeloma from SMM or MGUS in a subject in need thereof, the method comprising: a. receiving DNA methylation measurements of cfDNA from a sample from said subject; and b. quantifying the number of cfDNA molecules from a plasma cell in said sample by a method of any one of claims 23 to 25, wherein a quantity of cfDNA molecules from a plasma cell in said sample above a predetermined threshold indicates said subject suffers from multiple myeloma; thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject.
27. A method of predicting progression to multiple myeloma in a subject suffering from smoldering multiple myeloma (SMM) or monoclonal gammopathy of undetermined significance (MGUS), the method comprising: a. receiving DNA methylation measurements of cfDNA from a sample from said subject; and b. quantifying the number of cfDNA molecules from a plasma cell in said sample by a method of any one of claims 23 to 25, wherein a quantity of cfDNA molecules from a plasma cell in said sample above a predetermined threshold indicates said subject is predicted to progress to multiple myeloma; thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
8. A method of diagnosing multiple myeloma or predicting progression to multiple myeloma from SMM or MGUS in a subject in need thereof, the method comprising receiving measurements of 6 features in a fluid sample from said subject, wherein said 6 features are: a. the number of cfDNA molecules from a plasma cell in said sample; b. the percentage of discordant methylation reads (PDR) at genomic location chrl 1:107, 747, 505- 107 ,747, 531 wherein the coordinates are with respect to human genome build HG19; c. the entropy of the methylation at genomic location chrl 1:107,747,505- 107,747,531 wherein the coordinates are with respect to human genome build HG19; d. the entropy of the methylation at genomic location chr6:80,356,706- 80,356,741 wherein the coordinates are with respect to human genome build HG19; e. the entropy of the methylation at genomic location chrl5:41,072,702- 41,072,749 wherein the coordinates are with respect to human genome build HG19; and f. whether the free light chain (FLC) ratio is greater than or less than 20; applying a trained machine learning model to said 6 features, wherein said trained machine learning model is trained on said 6 features from subjects with MM and said 6 features from subjects with SMM, subjects with MGUS or both and outputs a diagnosis of multiple myeloma or not or a diagnosis of progression to multiple myeloma or not; thereby diagnosing multiple myeloma or predicting progression to multiple myeloma in a subject suffering from SMM or MGUS.
Citation Information
Patent Citations
Identification of source of DNA samples
US20120003634A1
Methylation profiling of DNA samples
US20130084571A1
Process for preparing polynucleotides
US4458066A
Test for Huntington's disease
US4666828A
Process for amplifying, detecting, and / or-cloning nucleic acid sequences
US4683195A