Methods to detect methylation status of ultrashort single-stranded and mononucleosomal cell-free DNA
By ligating 5mC-protected adaptors and performing bisulfite conversion on uscfDNA and mncfDNA, the method addresses the degradation issues in current techniques, enabling precise methylation profiling for early cancer detection.
Patent Information
- Application Number
- PCT/IB2025/050843
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-31
AI Technical Summary
Current methods for determining the methylation profile of ultrashort single-stranded cell-free DNA (uscfDNA) are inadequate, as they often result in DNA degradation and fail to accurately capture the methylation patterns of this crucial biomarker, which is essential for early cancer detection and diagnosis.
A method involving the ligation of 5mC-protected adaptors to uscfDNA and mncfDNA followed by bisulfite conversion, which minimizes DNA degradation and allows for the generation of a bisulfite-converted library for sequencing, enabling accurate methylation profiling.
This approach preserves the integrity of ultrashort DNA fragments, providing a more accurate methylation profile that can be used to identify biomarkers for diseases like non-small cell lung cancer, enhancing early detection and diagnosis.
Smart Images

Figure IB2025050843_31072025_PF_FP_ABST
Abstract
Description
Attorney Docket No.206030-0306-00WO METHODS TO DETECT METHYLATION STATUS OF ULTRASHORT SINGLE-STRANDED AND MONONUCLEOSOMAL CELL-FREE DNA CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No.63 / 624,499, filed January 24, 2024, which is hereby incorporated by reference herein in its entirety. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under CA264398 awarded by the National Institute of Health. The government has certain rights in the invention. REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0003] This application contains a Sequence Listing, which is submitted electronically via EFS-Web as an XML Document formatted sequence listing with a file name “206030-0306- 00WO Sequence Listing.xml”, a creation date of January 16, 2025, and having a size of 8,751 bytes. The sequence listing submitted via EFS-Web is part of the specification and is herein incorporated by reference in its entirety. BACKGROUND OF THE INVENTION
[0004] Cancer contributes to substantial morbidity and mortality worldwide, with 19.3 million new cancer cases and 10.0 million deaths in 2020 (Sung, H. et al., (2021) CA. Cancer J. Clin., 71, 209–249). Early cancer detection remains the best strategy to improve patient prognosis (Ignatiadis, M. et al., (2015) Clin. Cancer Res., 21, 4786–4800). A liquid biopsy strategy where biomolecules are analyzed within biofluids can provide noninvasive, easily repeated, and real-time insights into a patient’s tumor burden and treatment response (Wan, J.C.M. et al., (2017) Nat. Rev. Cancer, 17, 223–238).
[0005] Although many biomolecules (e.g., circulating tumor cells, exosomes, proteins, RNAs, or metabolites) are viable for liquid biopsy, there has been a focus on cell-free DNA(cfDNA). Examining methylation characteristics inherent within cfDNA fragments is a promising approach (Alexander et al., 2023, PLOS ONE, 18, e0283001; Chen et al., 2020, Nat. Commun., 11, 3475.) Dysregulated epigenetic control in tumor cells contributes to tumorigenesis and genetic instability. Both Global hypomethylation (Jones et al., 2002, Nat. Rev. Genet., 3, 415–428) and regional hypermethylation (Weber et al., 2005, Nat. Genet., 37, 853–862) are hallmarks of cancer cells observed in cfDNA. Since 60-80% of the 28 million CpG sites are methylated in humans (Lister et al., 2009, Nature, 462, 315–322), any deviations in the global profile could suggest aberrant methylation from cells, not of blood origin, which make up the majority of cfDNA (Mattox et al., 2023, Cancer Discov., 13, 2166–2179). Additionally, unlike somatic mutations, which may only be present heterogeneously in a tumor cluster, DNA methylation patterns are more consistent amongst all cells within cancer cells (Moss et al., 2018, Nat. Commun., 9, 5068).
[0006] The observed cytosine methylation, fragmentation, and strandedness of cfDNA are influenced by both its biological origins and the laboratory workflow. Previously it has been shown that preprocessing cfDNA using the Broad-Range Cell-free DNA Sequencing pipeline (BRcfDNA-Seq) reveals the presence of ultrashort single-stranded cell-free DNA (uscfDNA) sized around 50nt in addition to conventionally reported ~167bp mononucleosomal-cell free DNA (mncfDNA) (Cheng et al., 2022, iScience, 25, 104554; Cheng et al., 2023, Clin. Chem., 10.1093 / clinchem / hvad131). BRcfDNA-Seq uses enhanced isopropanol extraction coupled with single-stranded DNA library preparation to capture and incorporate short single-stranded or nicked cfDNA fragments routinely lost by double-stranded library kits. This ultrashort population has also been observed by other groups using similar extraction and library protocols (Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121; Hisano et al., 2021, BMC Biol., 19, 225; Cheng et al., 2022, iScience, 25, 105046; Miura et al., 2023, Sci. Rep., 13, 13913). Although mncfDNA has been thoroughly investigated as a cfDNA biomarker for tissue-of-origin deconvolution and cancer detection (Moss et al., 2018, Nat. Commun., 9, 5068; Luo et al., 2011, Trends Mol. Med., 27, 482–500), the epigenetic properties of uscfDNA have yet to be examined.
[0007] Whole genome bisulfite sequencing (BS-Seq) is a bisulfite-based sequencing technique that informs the methylation state of all the cytosines in the DNA sample at a single- base pair resolution (Luo et al., 2011, Trends Mol. Med., 27, 482–500). Alternative methods such as reduced representation bisulfite sequencing (RBBS) (Stackpole et al., 2022, Nat.Commun., 13, 5566) and methylated DNA immunoprecipitation (MeDIP-Seq) (Pv et al., 2020, Nat. Med., 26) demonstrate promising performance in cfDNA but will only provide information on about ~10% of the whole genome. In an unexplored biological context such as uscfDNA, it may be important to first gather a genome-wide impression of methylation behavior before limiting the scope to specific regions of interest.
[0008] For low-input samples typical of cfDNA, BS-Seq requires initial treatment with sodium bisulfite prior to the ligation of adapters and subsequent amplification. One drawback of bisulfite conversion is that it damages DNA, resulting in the creation of artificial smaller fragments (Moss et al., 2018, Nat. Commun., 9, 5068). The degree of degradation has been reported to be inversely proportional to the fragment size, with large genomic DNA being the most susceptible (Munson et al., 2007, Nucleic Acids Res., 35, 2893–2903; Kint et al., 2018, PLoS ONE, 13, e0199091). It has been reported that fragments up to 131bp will typically undergo a 20% loss due to degradation, while those sized around 62bp will experience a 10% loss (Munson et al., 2007, Nucleic Acids Res., 35, 2893–2903; Werner et al., 2019, PLoS ONE, 14, e0224338). The impact of artificially generated ultrashort fragments from bisulfite degradation should be minimized, as many cfDNA studies are contingent on the accurate measurements of size profiles of fragments. Thus, avoiding bisulfite-induced degradation reduces the potential masking of actual signal from natively ultrashort cfDNA fragments.
[0009] Thus, there is a need in the art for improved Next-generation Sequencing (NGS) methods to determine the methylation profile of ultra-short single-stranded cell-free DNA (uscfDNA). This invention satisfies this unmet need. SUMMARY OF THE INVENTION
[0010] The present invention provides a method for generating a bisulfite converted library from ultrashort single-stranded cell-free DNA (uscfDNA) and / or mononucleosomal cell- free DNA (mncfDNA), the method comprising the steps of: ligating a 5mC-protected adaptor to the 5’ and 3’ ends of the uscfDNA and / or mncfDNA; and treating the adaptor-conjugated uscfDNA and / or mncfDNA with bisulfite, thereby generating bisulfite converted uscfDNA and / or mncfDNA.
[0011] In some embodiments, the method further comprises preparing a sequence library from the bisulfite converted uscfDNA and / or mncfDNA.
[0012] In some embodiments, the method further comprises the step of sequencing the library of bisulfite converted uscfDNA and / or mncfDNA.
[0013] In some embodiments, the sample is a biological fluid sample.
[0014] In some embodiments, the sample is selected from the group consisting of a blood sample, a plasma sample, a serum sample, a saliva sample, a sweat sample, a sputum sample, a urine sample, an amniotic fluid sample, a peritoneal fluid sample, a pleural fluid sample, and a liquid biopsy sample.
[0015] In some embodiments, the method further comprises analyzing the methylation profile of the mncfDNA, uscfDNA, or the combination thereof, to identify biomarkers of a disease or disorder.
[0016] In some embodiments, the biomarker is an increase or decrease in the total amount of methylation of uscfDNA, mncfDNA, or a combination thereof in a test sample as compared to a control sample.
[0017] In some embodiments, the biomarker is an increase or decrease in the amount of methylation in a specific region of uscfDNA or mncfDNA in a test sample as compared to a control sample.
[0018] In some embodiments, the invention provides a method of diagnosing a disease or disorder in a subject in need thereof, the method comprising obtaining a sample from the subject, isolating uscfDNA, mncfDNA, analyzing the total amount of methylation or the amount of methylation of a specific region of the mncfDNA, uscfDNA, or a combination thereof to detect a methylation biomarker of a disease or disorder, and diagnosing the subject as having or at risk of the disease or disorder associated with the identified biomarker.
[0019] In some embodiments, analyzing the total amount of methylation or the amount of methylation of a specific region of the mncfDNA, uscfDNA, or a combination thereof, comprises the steps of: ligating a 5mC-protected adaptor to the 5’ and 3’ ends of the uscfDNA and / or mncfDNA; and treating the adaptor-conjugated uscfDNA and / or mncfDNA with bisulfite, thereby generating bisulfite converted uscfDNA and / or mncfDNA.
[0020] In some embodiments, the biomarker is an increase or decrease in the total amount of methylation of uscfDNA, mncfDNA, or a combination thereof in a test sample as compared to a control sample.
[0021] In some embodiments, the biomarker is an increase or decrease in the amount of methylation associated with a specific region of uscfDNA, or mncfDNA, in a test sample as compared to a control sample.
[0022] In some embodiments, the disease or disorder is cancer.
[0023] In some embodiments, the method further comprises administering a treatment based on the diagnostic outcome for the disease or disorder. In some embodiments, the disease or disorder is non-small cell lung cancer.
[0024] In some embodiments, the specific region is selected from the group consisting of: chromosome 1, bases 125183700-125183714; chromosome 1, bases 143214715- 143214803; chromosome 1, bases 143253126- 143253234; chromosome 4, bases 49137126- 49137132; chromosome 10, bases 38868465- 38868511; chromosome 10, bases 42080517-42080648; chromosome 10, bases 132804122-132804137; chromosome 11, bases 402680- 402832; chromosome 16, bases 34586783- 34586856; chromosome 16, bases 34588514- 34588812; chromosome 16, bases 34593300- 34593324; chromosome 16, bases 46394166- 46394407; chromosome 20, bases 29877880- 29877930; chromosome 20, bases 31061227- 31061429; chromosome 21, bases 7941250- 7941311; and chromosome 21, bases 8208928- 8209038.
[0025] In some embodiments, the method further comprises administering a treatment based on the diagnostic outcome for non-small cell lung cancer. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The following detailed description of embodiments of the invention will be better understood when read in conjunction with the appended drawings. It should be understood that the invention is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.
[0027] Figure 1A through Figure 1F depict data demonstrating that merging paired-end reads demonstrates a more similar profile to untreated non-targeted sequencing. Schematic pre- merging pipeline paired reads prior to alignment. (Figure 1A) Situations where reads are accepted for downstream analysis. In the uscfDNA (~50 bases) scenario, the consensus sequence of read 1 and 2 should have 100% overlap, whereas for mncfDNA (~167 bases), there will be a 150 base perfect overlap of read 1 read 2 with 17 bases with no overlap. These reads will still be accepted. (Figure 1B) Potential Scenarios where reads fail to merge and are discarded fromdownstream analysis. (Figure 1C) BRcfDNA-Seq libraries show little difference in pattern with and without merging per processing pipeline. (Figure 1D) BS-Seq libraries show the difference in the pattern when reads are merged prior to alignment (dip at 150 bases). MAPQ scores for binned reads of 10 bases for BRcfDNA-Seq (Figure 1E) and BS-Seq (Figure 1F) libraries for both paired ends and merged bioinformatic preprocessing. Vertical lines in (Figure 1C & Figure 1D) indicate SEM from the mean of five subjects. Some error bars may not be observable in Figure 1C & Figure 1D due to their length being smaller than the size of the data point.
[0028] Figure 2A through Figure 2F depict data demonstrating that the majority of reads are excluded during the initial merging of reads. The percent of total reads is shown after each bioinformatic preprocessing step comparing BRcfDNA-Seq libraries processed using both the Paired-End Reads (Figure 2A) and Merged Reads (Figure 2B) protocol. BS-Seq libraries are shown comparing Paired-End (Figure 2D) and Merged Reads (Figure 2E) processing. Comparison of final read count of individual samples between Pair-End Reads vs. Merged Read pipelines for BRcfDNA-Seq (Figure 2C) and BS-Seq (Figure 2F). Error bars indicate SEM.
[0029] Figure 3A and Figure 3F depict data demonstrating that the 5mCAdpBS-Seq protocol reduces the inclusion of DNA degradation into the ultrashort region of the final library. (Figure 3A) Schematic of routine BS-Seq workflow incorporates degraded cell-free DNA or genomic DNA, which enters the library, potentially masking the uscfDNA methylation signal. (Figure 3B) 5mCAdpBS-Seq protocol in which permethylated adapters are attached prior to bisulfite conversion, preventing degraded DNA from entering the final library. BRcfDNA-Seq (black), BS-Seq (pink), and 5mCAdpBS-Seq (green) protocols generate different fragment profiles for reads aligning to nuclear (Figure 3C) and mitochondria genomes (Figure 3D). CpG Density (Figure 3E) and G-Quad Density (Figure 3F) of 5mCAdpBS-Seq resemble the BRcfDNA-Seq profile compared to BS-Seq Protocol. The reads that align to mitochondria contributed to the minority of total sequence reads, averaging 0.35 ± 0.006%, 0.0135 ± 0.002%, and 0.0664 ± 0.012% for BRcfDNA-Seq, BS-Seq, and 5mCAdpBS-Seq respectively. Data represents SEM and the mean of 5 paired non-cancer subjects.
[0030] Figure 4A and Figure 4B depict data demonstrating that the comparative read attrition between BS-Seq and BS-Seq (5-mC Adp) during bioinformatic processing. (Figure 4A) Comparison of BS-Seq vs 5mCAdpBS-Seq protocols in their read attrition during each step ofthe preprocessing pipeline prior to downstream analysis. (Figure 4B) Final remaining reads for five individual plasma samples that underwent each protocol.
[0031] Figure 5A and Figure 5F depict data demonstrating that the genomic coverage linear correlation for BS-Seq and 5mCAdPBS-Seq compared to BRcfDNA-Seq for uscfDNA (Figure 5A and Figure 5B) and mncfDNA (Figure 5D and Figure 5E). For uscfDNA fragments, the linear correlation coefficient for samples processed with the 5mCAdpBS-Seq protocol is higher than those prepared for the BS-Seq. Data shows the average of 5 samples. Paired t-test for uscfDNA (Figure 5C) and mncfDNA (Figure 5F) fragments show on average a higher R2 value for uscfDNA fragments processed by the 5mCAdpBS-Seq protocol.
[0032] Figure 6A and Figure 6E depict data demonstrating that the 5mCAdpBS-Seq protocol resembles BRcfDNA-Seq for coverage of genomic elements and epigenetic marks. (Figure 6A) Schematic of intersection methodology to determine where uscfDNA or mncfDNA bases overlap with genomic element and epigenetic mark regions from reference .bed files from genomic or CHIP-seq databases. (Figure 6B) Genomic regions derived from epigenetic marks, including methylation patterns and histone modifications, are associated with active or repressed gene activity. % composition of genomic elements of each CpG-site was compared between BRcfDNA-Seq (Black), BS-Seq (Pink), and 5mCAdpBS-Seq (Green) protocols for uscfDNA (Figure 6C) bins (40-70 bases) and mncfDNA (Figure 6D) bins (120-250 bases). SINE: short interspersed nuclear element, LINE: long interspersed nuclear element, TTS: transcription termination site, 5’UTR: 5’ untranslated region, 3’UTR: 3’ untranslated region. (Figure 6E) Observed ratio (% of intersecting bases of .bed file to % intersection bases of randomly shuffled control bed files) for each epigenetic mark for uscfDNA and mncfDNA bins. Randomly shuffled bed files were generated for each sample to act as a control for intersection locations. The horizontal dotted line represents the observed ratio of 1.0. Data represents SEM and mean of 5 paired non-cancer subjects. Stars indicate p-values with * p <0.05, ** p <0.01, *** p< 0.001 after Tukey’s multiple comparison test after 2way ANOVA. Only comparisons with BRcfDNA- Seq are shown.
[0033] Figure 7 depicts data demonstrating that the observed ratio (% of intersecting bases of .bed file to % intersection bases of randomly shuffled control bed files) for other epigenetic marks not shown in main figure for uscfDNA and mncfDNA bins. Randomly shuffled bed files were generated for each sample to act as a control for intersection locations. Thehorizontal dotted line represents the observed ratio of 1.0. Data represents SEM and mean of 5 paired non-cancer subjects. Stars indicate p-values with * p <0.05, ** p <0.01, *** p< 0.001 after Tukey’s multiple comparison test after 2way ANOVA. Only comparisons with BRcfDNA- Seq are shown.
[0034] Figure 8A through Figure 8F depict data demonstrating that, compared to the BS- Seq protocol, the 5mCAdpBS-Seq protocol portrays that nuclear uscfDNA fragments are globally hypomethylated. CpG methylation (Figure 8A to Figure 8C) and non-CpG methylation (Figure 8D to Figure 8F) for nuclear, mitochondria, and lambda genome spike-in respectfully for BS-Seq and 5mCAdpBS-Seq protocols in 10 bases increment bins from 40-200. Lambda spike- in control indicates the inherent noise of bisulfite conversion methodology. Samples are from five paired samples undergoing both protocols. Error bars indicate SEM from the mean of five subjects.
[0035] Figure 9A through Figure 9E depict data demonstrating the zoom in scale for low CpG and non-CpG methylation bins in nuclear (Figure 9A), mitochondria (Figure 9B and Figure 9C), and lambda spike-in (Figure 9D and Figure 9E) reads. Samples are from five paired samples undergoing both protocols. Vertical lines and error bars indicate SEM from the mean of five subjects.
[0036] Figure 10A through Figure 10C depict data demonstrating that CpG positions of uscfDNA differ from mncfDNA. (Figure 10A) Karyograms averaged from 5 non-cancer subjects of % coverage for 1 million base-sized bins across chromosome 1 for the uscfDNA and mncfDNA. The ratio is calculated by dividing the mean uscfDNA %coverage by mncfDNA %coverage. (Figure 10B) Karyograms of all chromosomes. (Figure 10C) Intra-sample count of common and unique CpG site counts between cfDNA populations. Values above bars indicate the count of common CpG sites. Data represents SEM and the mean of 5 paired non-cancer subjects processed with 5mcAdpBS-Seq. Stars indicate unadjusted p-values with * p <0.05, ** p <0.01 , *** p< 0.001, and **** p <0.0001.
[0037] Figure 11A through Figure 11C depict data demonstrating that CpG methylation patterns differ between uscfDNA and mncfDNA fragments. (Figure 11A) Differences in % composition of different methylated CpG reads categories between uscfDNA and mncfDNA fragments. (Figure 11B) % composition of select genomic elements of different methylated read categories along select elements of typical gene structure. (Figure 11C) The average CpGmethylation % patterns from 5000 bases upstream and 5000 bases downstream from the body of the element for uscfDNA and mncfDNA sized reads. Lines show five separate non-cancer samples processed with the 5mCAdpBS-Seq protocol.
[0038] Figure 12A through Figure 12E depict data demonstrating that the pattern of enrichment of CpG fragments -1000 bases upstream and +1000 bases downstream from the transcription start site (TSS) differ amongst uscfDNA and mncfDNA fragments and correlate to gene activity. (Figure 12A) uscfDNA fragments are enriched upstream from the average TSS compared to mncfDNA fragments. The enrichment of uscfDNA (Figure 12B) and mncfDNA (Figure 12D) 0% methylated fragments is positively correlated to TSS of high expression genes, whereas the 75-100% CpG methylated fragments are negatively correlated uscfDNA (Figure 12C) and mncfDNA (Figure 12E) TSS categories were based on the RNA expression activity from RNA-Seq experiments of the buffy coat from previous literature. High expression was considered (>41.07 RPKM), medium (15.36-41.06 RPKM), low (1-15.36 RPKM), and silent (<0 RPKM). Lines show five separate non-cancer samples processed with the 5mCAdpBS-Seq protocol. Enrichment was normalized to all samples in the comparison.
[0039] Figure 13A through Figure 13D depict data demonstrating that the pattern of enrichment of CpG fragments -1000 bases upstream and +1000 bases downstream from the transcription start site for differentially expressed genes for uscfDNA and mncfDNA fragments with 0< to 24% CpG methylation (Figure 13A and Figure 13B) and 25-74% CpG methylation (Figure 13C and Figure 13D). TSS categories were based on the RNA expression activity from RNA-Seq experiments of the buffy coat from previous literature. High expression was considered (>41.07 RPKM), medium (15.36-41.06 RPKM), low (1-15.36 RPKM), and silent (<0 RPKM). Lines show five separate non-cancer samples processed with the 5mCAdpBS-Seq protocol. Enrichment was normalized to all samples in the comparison.
[0040] Figure 14A and Figure 14B depict data demonstrating that differentially methylated regions show differences in genes and cell-of-origin. (Figure 14A) Differentially methylated region analysis between merged uscfDNA and mncfDNA .bam files from five samples show 68 significant DMRs (q-value <0.01) and the closest gene. Only candidates with q-value <1.0 are shown. (Figure 14B) Box and whisker plots of CelFiE deconvolution prediction of blood cell tissue of origin signal from the methylation patterns in the uscfDNA and mncfDNA. Prediction reveals that uscfDNA and mncfDNA are derived from blood cells. Non-paired multiple paired t-tests were used to compare the % contribution of cell type between uscfDNA and mncfDNA. Stars represent unadjusted p-values with * p <0.05. Errors bars show min and max from five non-cancer samples that underwent 5mCAdpBS-Seq protocol.
[0041] Figure 15A and Figure 15B depict data demonstrating that genomic and methylation profiles differ between Non-Cancer and NSCLC samples processed by 5mCAdpBS- Seq. (Figure 15A) Fragment size distribution profile comparing non-caner and NSCLC cohorts. (Figure 15B) CpG Methylation of % of uscfDNA region of Non-Small Cell Lung Carcinoma Samples are elevated compared to non-cancer subjects. Plots represent the mean and SEM from 5 paired non-cancer plasma and 4 NSCLC, which 5mCAdpBS-Seq protocol.
[0042] Figure 16 depicts data demonstrating that the majority of DMRs between uscfDNA and mncfDNA are in close vicinity to TSS. Only DMR candidates between merged uscfDNA and mncfDNA .bam files from five samples with q-value <1.0 are shown. The total count was 573.
[0043] Figure 17A through Figure 17I depict data demonstrating that CpG coverage and methylation patterns differ between non-cancer and NSCLC samples. The composition of different genomic element category locations where CpG-site containing reads aligned were compared between the NSCLC and Non-Cancer samples for uscfDNA (Figure 17A) and mncfDNA (Figure 17B). The pattern of enrichment of CpG fragments -1000 bases upstream and +1000 bases downstream and average CpG methylation % patterns from 5000 bases upstream and 5000 bases downstream the body of the exon and LINE elements differ amongst non-cancer and non-cancer for uscfDNA (Figure 17C) and mncfDNA (Figure 17D) fragments. Differential methylated region analysis between merged uscfDNA and mncfDNA .bam files from 5 non- cancer subjects and 4 NSCLC subjects reveal significant DMRs in the uscfDNA(Figure 17E) and mncfDNA(Figure 17F) bin (q-value <0.01, only candidates with q-value <1.0 are shown). (Figure 17G) uscfDNA has a higher proportion of significant DMRs compared to the mncfDNA. Box and whiskers plot of CelFiE deconvolution algorithm suggests changes in cell type composition between non-cancer and NSCLC samples of the uscfDNA (Figure 17H) and the mncfDNA (Figure 17I) sized bins. Stars indicate p-values with * p <0.05, ** p <0.01 , *** p< 0.001 after Tukey’s multiple comparison test after 2way ANOVA (A and B). For the CelFiE deconvolution, error bars represent min and max positions with individual samples and unadjusted non-paired student t-tests. Data represents SEM and the mean of 5 paired non-cancerand 4 NSCLC plasma subjects. Stars indicate unadjusted p-values are presented with * p <0.05, ** p <0.01, *** p< 0.001.
[0044] Figure 18A and Figure 18B depict data demonstrating that CpG methylation patterns differ between non-cancer and NSCLC samples. The average CpG methylation % patterns from 5000 bases upstream and 5000 bases downstream are plotted for each genomic element for (Figure 18A) uscfDNA and (Figure 18B) mncfDNA-sized reads. SINE: short interspersed nuclear element, LINE: long interspersed nuclear element, TTS: transcription termination site, 5’UTR: 5’ untranslated region, 3’UTR: 3’ untranslated region. Lines show 5 paired plasms samples that underwent 5mCAdpBS-Seq protocol.
[0045] Figure 19A through Figure 19C depict data demonstrating that G-Quad methylation% and epigenetic mark overlap% are potential NSCLC biomarkers. (Figure 19A) G-Quad density is decreased in the uscfDNA regions (40-70 bases) in NSCLC. (Figure 19B) CpG methylation % significantly increased in G-Quad-containing fragments in uscfDNA. (Figure 19C) Normalized % of intersecting bases for three epigenetic marks (H3K27ac, H3K4me3, and hypomethylated regions) decreased in NSCLC samples in uscfDNA and mncfDNA bins. Observed ratio (% of intersecting bases of .bed file to % intersection bases of randomly shuffled control bed files) for each epigenetic mark for uscfDNA and mncfDNA bins. The horizontal dotted line represents the observed ratio of 1.0. Data represents SEM and mean of 5 paired non-cancer subjects and 4 NSCLC plasma subjects. Stars indicate p-values with * p <0.05, ** p <0.01, *** p< 0.001 after student t-test (Figure 19B) and Tukey’s multiple comparison test after 2way ANOVA (Figure 19C).
[0046] Figure 20 depicts data demonstrating that the normalized % of intersecting bases for epigenetic marks (H3K27me, H3K36me3, H3K5me1, H3k9me3, and hypermethylated regions. Minor changes are observed between NSCLC samples and non-cancer samples in both uscfDNA and mncfDNA bins. % intersection was normalized to control shuffled bed files. The horizontal dotted line represents the observed ratio of 1.0. Data represents SEM and mean of 5 paired non-cancer subjects and 4 NSCLC plasma subjects. Stars indicate p-values with * p <0.05, ** p <0.01, *** p< 0.001 after Tukey’s multiple comparison test after 2way ANOVA.
[0047] Figure 21 depicts data demonstrating that the enzymatic conversion of 5mC did not generate sufficient libraries. Electrophoresis gel of the comparison of bisulfite and enzyme conversion protocols for extracted cell-free DNA from 2mL of non-cancer plasma after single-stranded library preparation. BS-Seq and 5mCAdpBS-Seq protocols are shown. The enzyme conversion protocol generates libraries with only adapter dimers or cell-free DNA-sized bands with low concentrations.
[0048] Figure 22 depicts data demonstrating that filtering for proper pairs (R1 and R2 with orientation towards each other) the pair-end reads pipeline mimics the size-distribution curve of the merged reads pipeline.
[0049] Figure 23 depicts a schematic of the routine BS-Seq (5mC-AdP) workflow. DETAILED DESCRIPTION Definitions
[0050] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0051] As used herein, each of the following terms has the meaning associated with it in this section.
[0052] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0053] “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, or ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.
[0054] “Amplification,” as used herein, refers to any in vitro process for increasing the number of copies of a nucleotide sequence or sequences, i.e., creating an amplification product which may include, by way of example additional target molecules, or target-like molecules or molecules complementary to the target molecule, which molecules are created by virtue of the presence of the target molecule in the sample. These amplification processes include but are not limited to polymerase chain reaction (PCR), multiplex PCR, Rolling Circle PCR, ligase chain reaction (LCR) and the like, in a situation where the target is a nucleic acid, an amplification product can be made enzymatically with DNA or RNA polymerases or transcriptases. Nucleicacid amplification results in the incorporation of nucleotides into DNA or RNA. As used herein, one amplification reaction may consist of many rounds of DNA replication. PCR is an example of a suitable method for DNA amplification. For example, one PCR reaction may consist of 2-40 “cycles” of denaturation and replication.
[0055] “Amplification products,” “amplified products” “PCR products” or “amplicons” comprise copies of the target sequence and are generated by hybridization and extension of an amplification primer. This term refers to both single stranded and double stranded amplification primer extension products which contain a copy of the original target sequence, including intermediates of the amplification reaction.
[0056] A “barcode”, as used herein, refers to a nucleotide sequence that serves as a means of identification for sequenced polynucleotides of the present invention. Barcodes of the present invention may comprise at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more bases in length.
[0057] “Nucleic acid” or “oligonucleotide” or “polynucleotide” or “nucleic acid fragment” as used herein may mean at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary strand, or the sequence of a molecule that hybridizes to at least a portion of the single strand sequence. Thus, a nucleic acid also encompasses the complementary strand of a depicted single strand as well as probes, primers or oligonucleotide sequences having complementarity to at least a portion of the strand. Many variants of a nucleic acid may be used for the same purpose as a given nucleic acid. Thus, a nucleic acid also encompasses substantially identical nucleic acids and complements thereof. A single strand provides a probe that may hybridize to a target sequence. Thus, a nucleic acid also encompasses a probe that hybridizes under appropriate hybridization conditions.
[0058] Nucleic acids may be single stranded or double stranded, or may contain portions of both double stranded and single stranded sequence. The nucleic acid may be DNA, both genomic and cDNA, RNA, or a hybrid, where the nucleic acid may contain combinations of deoxyribo- and ribo-nucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine and isoguanine. Nucleic acids may be obtained by chemical synthesis methods or by recombinant methods. As used herein, the term nucleic acids includes both natural and non-natural nucleic acids. Non-natural nucleic acids include, but are not limited to, 2′F, 2′-fluoro; 2′OMe, 2′-O-methyl; LNA, locked nucleic acid;FANA, 2′-fluoro arabinose nucleic acid; HNA, hexitol nucleic acid; 2′MOE, 2′-O-methoxyethyl; ribuloNA, (1′-3′)-β-L-ribulo nucleic acid; TNA, α-L-threose nucleic acid; tPhoNA, 3′-2′ phosphonomethyl-threosyl nucleic acid; dXNA, 2′-deoxyxylonucleic acid; PS, phosphorothioate; phNA, alkyl phosphonate nucleic acid; and PNA, peptide nucleic acid.
[0059] “Primer” as used herein refers to a single-stranded oligonucleotide or a single- stranded polynucleotide that is extended on its 3’ end by covalent addition of nucleotide monomers during amplification. Nucleic acid amplification often is based on nucleic acid synthesis by a nucleic acid polymerase. Many such polymerases require the presence of a primer that can be extended to initiate such nucleic acid synthesis.
[0060] As used herein, “sample” or “test sample,” may refer to any source used to obtain nucleic acids for examination using the compositions and methods of the invention. A test sample is typically anything suspected of containing a target sequence.
[0061] As used herein “endogenous” refers to any material from or produced inside an organism, cell, tissue or system.
[0062] As used herein, the term “exogenous” refers to any material introduced from or produced outside an organism, cell, tissue or system.
[0063] As used herein, the term “fragment,” as applied to a nucleic acid, refers to a subsequence of a larger nucleic acid. A “fragment” of a nucleic acid can be at least about 36 nucleotides in length; for example, at least about 40 nucleotides to about 50 nucleotides; at least about 50 to about 60 nucleotides, at least about 60 to about 70 nucleotides; at least about 70 nucleotides to about 80 nucleotides; about 80 nucleotides to about 90 nucleotides; or about 100 nucleotides (and any integer value in between). As used herein, the term “fragment,” as applied to a protein or peptide, refers to a subsequence of a larger protein or peptide. A “fragment” of a protein or peptide can be at least about 12 amino acids in length; for example, a fragment of SEQ ID NO:1 can be at least about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length.
[0064] The term “functionally equivalent” as used herein refers to a polypeptide according to the invention that retains at least one biological function or activity of the specific amino acid sequence of either the first or second peptide.
[0065] “Homologous” refers to the sequence similarity or sequence identity between two polypeptides or between two nucleic acid molecules. When a position in both of the twocompared sequences is occupied by the same base or amino acid monomer subunit, e.g., if a position in each of two DNA molecules is occupied by adenine, then the molecules are homologous at that position. The percent of homology between two sequences is a function of the number of matching or homologous positions shared by the two sequences divided by the number of positions compared X 100. For example, if 6 of 10 positions in two sequences are matched or homologous then the two sequences are 60% homologous. By way of example, the DNA sequences ATTGCC and TATGGC share 50% homology. Generally, a comparison is made when two sequences are aligned to give maximum homology.
[0066] “Isolated” means altered or removed from the natural state. For example, a nucleic acid or a peptide naturally present in a living animal is not “isolated,” but the same nucleic acid or peptide partially or completely separated from the coexisting materials of its natural state is “isolated.” An isolated nucleic acid or protein can exist in substantially purified form, or can exist in a non-native environment such as, for example, a host cell.
[0067] The terms “patient,” “subject,” “individual,” and the like are used interchangeably herein, and refer to any animal, or cells thereof whether in vitro or in situ, amenable to the methods described herein. In certain non-limiting embodiments, the patient, subject, or individual is a human.
[0068] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.
[0069] As used herein, the term “bind” or “binding” refers to the specific association or other specific interaction between two molecular species, such as, but not limited to, protein- DNA / RNA interactions and protein-protein interactions, for example, the specific association between proteins and their DNA / RNA targets, receptors and their ligands, enzymes and their substrates, etc. Such binding may be specific or non-specific, and can involve variousnoncovalent interactions, such as hydrogen bonding, metal coordination, hydrophobic forces, van der Waals forces, pi-pi interactions, and / or electrostatic effects.
[0070] The term “cancer” as used herein is defined as disease characterized by the rapid and uncontrolled growth of aberrant cells. Cancer cells can spread locally or through the bloodstream and lymphatic system to other parts of the body. Examples of various cancers include, but are not limited to, lung cancer, including non-small cell lung cancer, breast cancer, prostate cancer, ovarian cancer, cervical cancer, skin cancer, pancreatic cancer, colorectal cancer, renal cancer, liver cancer, brain cancer, lymphoma, leukemia, oral cancer and the like.
[0071] The terms “cells” and “population of cells” are used interchangeably and refer to a plurality of cells, i.e., more than one cell. The population may be a pure population comprising one cell type. Alternatively, the population may comprise more than one cell type. In the present invention, there is no limit on the number of cell types that a cell population may comprise.
[0072] As used herein, “methylation” refers to methylation of cytosine at position C5 or N4 of cytosine, position N6 of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is usually unmethylated because in general in vitro DNA amplification methods do not preserve the methylation pattern of the amplified template. However, “unmethylated DNA” or “methylated DNA” may refer to amplified DNA in which the original template is unmethylated or methylated, respectively.
[0073] “Bisulfite conversion” as used herein refers to treating DNA with, for example, bisulfite, disulfite, hydrogen sulfite, or combinations thereof, thereby converting unmethylated cytosine nucleotides to uracil, while methylated cytosine and other bases remain unchanged, thus distinguishing, for example, CpG methylated and unmethylated cytidines in the nucleotide sequence. Description
[0074] The present invention is based on an optimized bisulfite sequencing protocol for profiling the methylation of ultrashort single-stranded cell-free DNA (uscfDNA) and mononucleosomal cell-free DNA (mncfDNA) in a sample. In one embodiment, the methylation profile of cell-free DNA is used for the detection of a disease or disorder in a subject.
[0075] In one embodiment, the sample is selected from the group consisting of plasma, serum, blood, a blood fraction, urine, sweat, sputum / oral fluid, saliva, amniotic fluid, or fineneedle biopsy samples (e.g., surgical biopsy, fine needle biopsy, etc.), peritoneal fluid, pleural fluid, and the like. In one embodiment, the sample is plasma.
[0076] In one embodiment, the disease or disorder is associated with differential total methylation of uscfDNA and / or mncfDNA. In one embodiment, the disease or disorder is associated with differential methylation of a specific region of uscfDNA or mncfDNA. In one embodiment, the disease or disorder is cancer. In one embodiment, the disease or disorder is non- small cell lung cancer. Methods for obtaining and using uscfDNA and mncfDNA
[0077] The invention is based, in part, on the development of a new pipeline for methylation sequencing of uscfDNA and mncfDNA. The baseline process may comprise the following steps: a) collect or obtain a sample of a subject b) extract uscfDNA and / or mncfDNA from the sample using an extraction method optimized for uscfDNA and / or mncfDNA, c) prepare a bisulfite sequencing library from the extracted uscfDNA and / or mncfDNA and d) perform next generation sequencing on the sequencing library.
[0078] In some embodiments, the methods of the invention comprise a step of obtaining a plasma fraction of the whole blood sample, wherein the plasma fraction comprises the uscfDNA and mncfDNA. In some embodiments, uscfDNA and / or mncfDNA are isolated from a sample using the miRNA protocol of the QIAamp Circulating Nucleic Acid Kit. Library preparation
[0079] In some embodiments, the method of the invention comprises the preparation of a bisulfite sequencing library from cell-free DNA. In some embodiments, the method of the invention comprises attaching 5mC protected sequencing adapters to ends of ultrashort single- stranded cell-free DNA fragments or mononucleosomal cell-free DNA fragments before bisulfite conversion, thereby preparing a sequencing library comprising library fragments having the sequencing adapters attached to either end of the cell-free DNA fragments.
[0080] In some embodiments, the steps of the method of the invention comprise: a) ligation of 5mC protected adapters to uscfDNA or mncfDNA, b) bisulfite conversion, c) single- strand library preparation. In some embodiments, the 5mC protected adapters are SRSLY™adapters. In some embodiments, bisulfite conversion is performed using Zymo Research DNA Methylation-Lightning™ kit. In some embodiments, the elution volume after bisulfite conversion is 15μL. In some embodiments, library preparation is conducted using the single- strand library preparation protocol provided by SRSLYTM PicoPlus DNA NGS Library Preparation Base Kit wherein during the final index PCR, the Index PCR Master Mix is substituted with the Kapa HIFI HotStart Uracil+ ReadyMix and the Bisulfite PCR protocol is as follows: 98°C for 3 minutes, [98°C for 30 seconds, 60°C for 30 seconds, 72°C for 1 minute] for 11 cycles, 72°C for 1 minute, and low molecular weight retention purification protocol is followed for all bead clean-up steps. Multiplex sequencing
[0081] The large number of sequencing reads that can be obtained per sequencing run permits the analysis of pooled samples i.e. multiplexing, which maximizes sequencing capacity and reduces workflow. For example, the massively parallel sequencing of eight libraries performed using the eight-lane flow cell of the Illumina Genome Analyzer, and Illumina's HiSeq Systems, can be multiplexed to sequence two or more samples in each lane such that 16, 24, 32 etc. or more samples can be sequenced in a single run. Parallelizing sequencing for multiple samples i.e. multiplex sequencing, requires the incorporation of sample-specific index sequences, also known as barcodes, during the preparation of sequencing libraries. Sequencing indexes are distinct base sequences of about 5, about 10, about 15, about 20 about 25, or more bases that are added at the 3' end of the genomic and marker nucleic acid. The multiplexing system enables sequencing of hundreds of biological samples within a single sequencing run. The preparation of indexed sequencing libraries for sequencing clonally amplified sequences can be performed by incorporating an index sequence into a PCR primer used for cluster amplification. Alternatively, the index sequence can be incorporated into the adaptor, which is ligated to the cell-free DNA prior to the PCR amplification. Sequencing of the uniquely marked indexed nucleic acids provides index sequence information that identifies samples in the pooled sample libraries, and sequence information of marker molecules correlates sequencing information of the genomic nucleic acids to the sample source. In embodiments wherein the multiple samples are sequenced individually i.e. singleplex sequencing, marker and cell-free DNA of each sample need only bemodified to contain the adaptor sequences as required by the sequencing platform and exclude the indexing sequences. Samples
[0082] In some embodiments, the sample containing cell-free DNA is derived from a biological fluid, cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids comprising at least one cell-free DNA molecule. Such samples include, but are not limited to plasma, serum, blood, a blood fraction, urine, sweat, sputum / oral fluid, saliva, amniotic fluid, or fine needle biopsy samples (e.g., surgical biopsy, fine needle biopsy, etc.), peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (e.g., patient), the assays can be from any mammal, including, but not limited to, dogs, cats, horses, goats, sheep, cattle, pigs, etc.
[0083] The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample. Such “treated” or “processed” samples are still considered to be biological samples with respect to the methods described herein. Applications
[0084] Methylation profile information generated as described herein can be used for any number of applications. Exemplary applications include identifying methylation biomarkers for diseases or disorders using uscfDNA or mncfDNA. The methods and apparatus described herein may employ next generation sequencing technology (NGS) as described elsewhere herein. In certain embodiments, clonally amplified mncfDNA or uscfDNA molecules are sequenced in a massively parallel fashion within a flow cell (e.g. as described in Volkerding et al., 2009, ClinChem, 55:641-658; Metzker, 2010, Nature Rev, 11:31-46). In addition to high-throughput methylation information, NGS provides quantitative information, in that each sequence read is a countable “sequence tag” representing an individual clonal DNA template or a single DNA molecule. In some embodiments, the methods and apparatus disclosed herein may employ some or all of the operations from the following: obtain a nucleic acid test sample from a subject (typically by a non-invasive procedure); process the test sample in preparation for sequencing as described herein (ligate 5mC protected adaptors and perform bisulfite conversion); sequence nucleic acids from the test sample to produce numerous reads (e.g., at least 10,000); align the reads to portions of a reference sequence / genome and determine the amount of DNA (e.g., the number of reads) that map to defined portions the reference sequence (e.g., to defined chromosomes or chromosome segments); calculate a dose of one or more of the defined portions by normalizing the amount of methylated DNA mapping to the defined portions with an amount of unmethylated DNA mapping to the defined portion; determining whether the dose indicates that the defined portion is “affected” (e.g., hypermethylated or hypomethylated); reporting the determination and optionally converting it to a diagnosis; using the diagnosis or determination to develop a plan of treatment, monitoring, or further testing for the patient. In some embodiments, the biological sample is obtained from a subject and comprises a mixture of nucleic acids contributed by different subjects. Diagnostic Assays
[0085] In some embodiments, use of the methods described herein in the diagnosis, and / or monitoring, and or treating pathologies is contemplated. For example, the methods can be applied to determining the presence or absence of a disease, to monitoring the progression of a disease and / or the efficacy of a treatment regimen. Methylation biomarkers associated with these diseases and disorders can be identified in uscfDNA or mncfDNA enriched samples generated according to the methods of the invention. In some embodiments, the disease or disorder associated with differential methylation of uscfDNA and / or mncfDNA is cancer. In sone embodiments, the disease or disorder associated with differential methylation of uscfDNA and / or mncfDNA is non-small cell lung cancer.
[0086] In some embodiments, blood, plasma and serum DNA from cancer patients contains measurable quantities of tumor DNA, that can be identified using the methods of theinvention to identify the type or stage of the tumor. Identification of dysregulated methylation associated with cancers that can be determined in the circulating uscfDNA and mncfDNA in cancer patients is a potential diagnostic and prognostic tool. In one embodiment, methods described herein are used to determine a methylation biomarker of one or more region(s) of interest in a sample, e.g., a sample comprising a mixture of nucleic acids derived from a subject that is suspected or is known to have cancer. In one embodiment, the sample is a plasma sample derived (processed) from peripheral blood that may comprise a mixture of uscfDNA and / or mncfDNA derived from normal and cancerous cells.
[0087] The following are non-limiting examples of cancers that can be diagnosed by the disclosed methods: acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, appendix cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, brain and spinal cord tumors, brain stem glioma, brain tumor, breast cancer, bronchial tumors, burkitt lymphoma, carcinoid tumor, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumors, central nervous system lymphoma, cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, cerebral astrocytotna / malignant glioma, cervical cancer, childhood visual pathway tumor, chordoma, chronic lymphocytic leukemia, chronic myelogenous leukemia, chronic myeloproliferative disorders, colon cancer, colorectal cancer, craniopharyngioma, cutaneous cancer, cutaneous t-cell lymphoma, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, ewing family of tumors, extracranial cancer, extragonadal germ cell tumor, extrahepatic bile duct cancer, extrahepatic cancer, eye cancer, fungoides, gallbladder cancer, gastric (stomach) cancer, gastrointestinal cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (gist), germ cell tumor, gestational cancer, gestational trophoblastic tumor, glioblastoma, glioma, hairy cell leukemia, head and neck cancer, hepatocellular (liver) cancer, histiocytosis, hodgkin lymphoma, hypopharyngeal cancer, hypothalamic and visual pathway glioma, hypothalamic tumor, intraocular (eye) cancer, intraocular melanoma, islet cell tumors, kaposi sarcoma, kidney (renal cell) cancer, langerhans cell cancer, langerhans cell histiocytosis, laryngeal cancer, leukemia, lip and oral cavity cancer, liver cancer, lung cancer, lymphoma, macroglobulinemia, malignant fibrous histiocvtoma of bone and osteosarcoma, medulloblastoma, medulloepithelioma, melanoma, merkel cell carcinoma, mesothelioma, metastatic squamous neck cancer with occult primary, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis,myelodysplastic syndromes, myelodysplastic / myeloproliferative diseases, myelogenous leukemia, myeloid leukemia, myeloma, myeloproliferative disorders, nasal cavity and paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, non-hodgkin lymphoma, non-small cell lung cancer, oral cancer, oral cavity cancer, oropharyngeal cancer, osteosarcoma and malignant fibrous histiocytoma, osteosarcoma and malignant fibrous histiocytoma of bone, ovarian, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, pancreatic cancer, papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal parenchymal tumors of intermediate differentiation, pineoblastoma and supratentorial primitive neuroectodermal tumors, pituitary tumor, plasma cell neoplasm, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma, primary central nervous system cancer, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell (kidney) cancer, renal pelvis and ureter cancer, respiratory tract carcinoma involving the nut gene on chromosome 15, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma, sezary syndrome, skin cancer (melanoma), skin cancer (nonmelanoma), skin carcinoma, small cell lung cancer, small intestine cancer, soft tissue cancer, soft tissue sarcoma, squamous cell carcinoma, squamous neck cancer , stomach (gastric) cancer, supratentorial primitive neuroectodermal tumors, supratentorial primitive neuroectodermal tumors and pineoblastoma, T-cell lymphoma, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, transitional cell cancer, transitional cell cancer of the renal pelvis and ureter, trophoblastic tumor, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, visual pathway and hypothalamic glioma, vulvar cancer, waldenstrom macroglobulinemia, and wilms tumor.
[0088] In some embodiments, use of the methods described herein in the diagnosis, and / or monitoring, and or treating non-small cell lung cancer is contemplated. Methylation biomarkers associated with non-small cell lung cancer can be identified in uscfDNA or mncfDNA enriched samples generated according to the methods of the invention. In some embodiments, the methylation biomarker associated with non-small cell lung cancer is selected from the group consisting of: chromosome 1, bases 125183700-125183714; chromosome 1, bases 143214715- 143214803; chromosome 1, bases 143253126- 143253234; chromosome 4, bases 49137126- 49137132; chromosome 10, bases 38868465- 38868511; chromosome 10, bases 42080517-42080648; chromosome 10, bases 132804122-132804137; chromosome 11,bases 402680- 402832; chromosome 16, bases 34586783- 34586856; chromosome 16, bases 34588514- 34588812; chromosome 16, bases 34593300- 34593324; chromosome 16, bases 46394166- 46394407; chromosome 20, bases 29877880- 29877930; chromosome 20, bases 31061227- 31061429; chromosome 21, bases 7941250- 7941311; and chromosome 21, bases 8208928- 8209038. Data Processing
[0089] After isolating uscfDNA and / or mncfDNA as described herein, the uscfDNA and / or mncfDNA may be detected and / or analyzed by any suitable method and any suitable detection device. One or more target nucleic acids in the uscfDNA and / or mncfDNA may be detected and / or analyzed. In some embodiments, the uscfDNA and / or mncfDNA may contain methylated markers that can be used to identify a disease or disorder. In some embodiments, the uscfDNA and / or mncfDNA may also be useful for as a global biomarker in which its increase concentration may be diagnostic of aberrations in the patient’s condition. Therefore, in some embodiments, the invention includes methods of diagnosing subjects based on the identification of a methylation biomarker in uscfDNA and / or mncfDNA.
[0090] In some embodiments, a diagnosis or the presence or absence of an outcome can be determined from the detection and / or analysis results. In some embodiments, the term “outcome” as used herein can refer to the presence, absence, or amount of a methylation biomarker in a population of uscfDNA and / or mncfDNA nucleic acids in the sample. In some embodiments, the term “outcome” as used herein can refer to an increase or decrease in the proportion of total methylation in uscfDNA and / or mncfDNA nucleic acids in the sample. In some embodiments, the term “outcome” as used herein can refer to identification of a disease, disorder or condition associated with a methylation biomarker or total methylation of uscfDNA and / or mncfDNA nucleic acids in the sample. A non-limiting example of an outcome includes presence or absence of a cancer, for example, non-small cell lung cancer.
[0091] As described herein, algorithms, software, processors and / or machines, for example, can be utilized to (i) process detection data pertaining to uscfDNA and / or mncfDNA nucleic acid, and / or (ii) identify the presence or absence of an outcome.
[0092] The presence or absence of an outcome may be determined for all samples tested, or in some embodiments, the presence or absence of an outcome is determined in a subset of thesamples (e.g., samples from individual subjects). An outcome may be determined for about 60, 65, 70, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99%, or greater than 99%, of samples analyzed in a set. A set of samples can include any suitable number of samples, and in some embodiments, a set has about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 samples, or more than 1000 samples. The set may be considered with respect to samples tested in a particular period of time, and / or at a particular location. The set may be otherwise defined by, for example, age and / or ethnicity. The set may be comprised of a sample which is subdivided into subsamples or replicates all or some of which may be tested. The set may comprise a sample from the same subject collected at two different times. An outcome may be determined about 60% or more of the time for a given sample analyzed (e.g., about 65, 70, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99%, or more than 99% of the time for a given sample). Analyzing a higher number of characteristics (e.g., sequence variations) that discriminate alleles can increase the percentage of outcomes determined for the samples (e.g., discriminated in a multiplex analysis). One or more fluid samples (e.g., one or more blood samples) may be provided by a subject. One or more uscfDNA and / or mncfDNA enriched samples, or two or more replicate uscfDNA and / or mncfDNA enriched samples, may be isolated from a single fluid sample, and analyzed by methods described herein.
[0093] Presence or absence of an outcome can be expressed in any suitable form, and in conjunction with any suitable variable, collectively including, without limitation, ratio, deviation in ratio, frequency, distribution, probability (e.g., odds ratio, p-value), likelihood, percentage, value over a threshold, or risk factor, associated with the presence of a outcome for a subject or sample. An outcome may be provided with one or more variables, including, but not limited to, sensitivity, specificity, standard deviation, probability, ratio, coefficient of variation (CV), threshold, score, probability, confidence level, or combination of the foregoing, in certain embodiments.
[0094] One or more of ratio, sensitivity, specificity and / or confidence level may be expressed as a percentage. The percentage, independently for each variable, may be greater than about 90% (e.g., about 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99%, or greater than 99% (e.g., about 99.5%, or greater, about 99.9% or greater, about 99.95% or greater, about 99.99% or greater)). Coefficient of variation (CV) in some embodiments is expressed as a percentage, and sometimesthe percentage is about 10% or less (e.g., about 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1%, or less than 1% (e.g., about 0.5% or less, about 0.1% or less, about 0.05% or less, about 0.01% or less)). A probability (e.g., that a particular outcome determined by an algorithm is not due to chance) in certain embodiments is expressed as a p-value, and sometimes the p-value is about 0.05 or less (e.g., about 0.05, 0.04, 0.03, 0.02 or 0.01, or less than 0.01 (e.g., about 0.001 or less, about 0.0001 or less, about 0.00001 or less, about 0.000001 or less)).
[0095] For example, scoring or a score may refer to calculating the probability that a particular outcome is actually present or absent in a subject / sample. The value of a score may be used to determine for example the variation, difference, or ratio of amplified nucleic detectable product that may correspond to the actual outcome. For example, calculating a positive score from detectable products can lead to an identification of an outcome, which is particularly relevant to analysis of single samples.
[0096] Simulated (or simulation) data can aid data processing for example by training an algorithm or testing an algorithm. Simulated data may for instance involve hypothetical various samples of different concentrations of methylated uscfDNA and / or mncfDNA in serum, plasma, saliva and the like. Simulated data may be based on what might be expected from a real population or may be skewed to test an algorithm and / or to assign a correct classification based on a simulated data set. Simulated data also is referred to herein as “virtual” data. Simulations can be performed in most instances by a computer program. One possible step in using a simulated data set is to evaluate the confidence of the identified results, i.e. how well the selected positives / negatives match the sample and whether there are additional variations. A common approach is to calculate the probability value (p-value) which estimates the probability of a random sample having better score than the selected one. As p-value calculations can be prohibitive in certain circumstances, an empirical model may be assessed, in which it is assumed that at least one sample matches a reference sample (with or without resolved variations). Alternatively other distributions such as Poisson distribution can be used to describe the probability distribution.
[0097] An algorithm can assign a confidence value to the true positives, true negatives, false positives, and false negatives calculated. The assignment of a likelihood of the occurrence of an outcome can also be based on a certain probability model.
[0098] Simulated data often is generated in an in silico process. As used herein, the term “in silico” refers to research and experiments performed using a computer. In silico methods include, but are not limited to, molecular modeling studies, karyotyping, genetic calculations, biomolecular docking experiments, and virtual representations of molecular structures and / or processes, such as molecular interactions.
[0099] As used herein, a “data processing routine” refers to a process that can be embodied in software that determines the biological significance of acquired data (i.e., the ultimate results of an assay). For example, a data processing routine can determine the amount of each nucleotide sequence species based upon the data collected. A data processing routine also may control an instrument and / or a data collection routine based upon results determined. A data processing routine and a data collection routine often are integrated and provide feedback to operate data acquisition by the instrument, and hence provide assay-based judging methods provided herein.
[0100] As used herein, software refers to computer readable program instructions that, when executed by a computer, perform computer operations. Typically, software is provided on a program product containing program instructions recorded on a computer readable medium, including, but not limited to, magnetic media including floppy disks, hard disks, and magnetic tape; and optical media including CD-ROM discs, DVD discs, magneto-optical discs, and other such media on which the program instructions can be recorded.
[0101] Different methods of predicting abnormality or normality can produce different types of results. For any given prediction, there are four possible types of outcomes: true positive, true negative, false positive or false negative. The term “true positive” as used herein refers to a subject correctly diagnosed as having a outcome. The term “false positive” as used herein refers to a subject wrongly identified as having a outcome. The term “true negative” as used herein refers to a subject correctly identified as not having a outcome. The term “false negative” as used herein refers to a subject wrongly identified as not having a outcome. Two measures of performance for any given method can be calculated based on the ratios of these occurrences: (i) a sensitivity value, the fraction of predicted positives that are correctly identified as being positives (e.g., the fraction of nucleotide sequence sets correctly identified by level comparison detection / determination as indicative of outcome, relative to all nucleotide sequence sets identified as such, correctly or incorrectly), thereby reflecting the accuracy of the results indetecting the outcome; and (ii) a specificity value, the fraction of predicted negatives correctly identified as being negative (the fraction of nucleotide sequence sets correctly identified by level comparison detection / determination as indicative of chromosomal normality, relative to all nucleotide sequence sets identified as such, correctly or incorrectly), thereby reflecting accuracy of the results in detecting the outcome.
[0102] The term “sensitivity” as used herein refers to the number of true positives divided by the number of true positives plus the number of false negatives, where sensitivity (sens) may be within the range of 0 ≤ sens ≤ 1. Ideally, method embodiments herein have the number of false negatives equaling zero or close to equaling zero, so that no subject is wrongly identified as not having at least one outcome when they indeed have at least one outcome. Conversely, an assessment often is made of the ability of a prediction algorithm to classify negatives correctly, a complementary measurement to sensitivity. The term “specificity” as used herein refers to the number of true negatives divided by the number of true negatives plus the number of false positives, where sensitivity (spec) may be within the range of 0 ≤ spec ≤ 1. Ideally, methods embodiments herein have the number of false positives equaling zero or close to equaling zero, so that no subject wrongly identified as having at least one outcome when they do not have the outcome being assessed. Hence, a method that has sensitivity and specificity equaling one, or 100%, sometimes is selected.
[0103] One or more prediction algorithms may be used to determine significance or give meaning to the detection data collected under variable conditions that may be weighed independently of or dependently on each other. The term “variable” as used herein refers to a factor, quantity, or function of an algorithm that has a value or set of values. For example, a variable may be the design of a set of amplified nucleic acid species, the number of sets of amplified nucleic acid species, type of outcome assayed, and the like.
[0104] Any suitable type of method or prediction algorithm may be utilized to give significance to the data of the present technology within an acceptable sensitivity and / or specificity. For example, prediction algorithms such as Mann-Whitney U Test, binomial test, log odds ratio, Chi-squared test, z-test, t-test, ANOVA (analysis of variance), regression analysis, neural nets, fuzzy logic, Hidden Markov Models, multiple model state estimation, and the like may be used. One or more methods or prediction algorithms may be determined to give significance to the data having different independent and / or dependent variables of the presenttechnology. And one or more methods or prediction algorithms may be determined not to give significance to the data having different independent and / or dependent variables of the present technology. One may design or change parameters of the different variables of methods described herein based on results of one or more prediction algorithms (e.g., number of sets analyzed, types of nucleotide species in each set).
[0105] Several algorithms may be chosen to be tested. These algorithms then can be trained with raw data. For each new raw data sample, the trained algorithms will assign a classification to that sample (e.g., trisomy or normal). Based on the classifications of the new raw data samples, the trained algorithms' performance may be assessed based on sensitivity and specificity. Finally, an algorithm with the highest sensitivity and / or specificity or combination thereof may be identified.
[0106] Provided are methods for identifying the presence or absence of an outcome that comprise: (a) providing a system, wherein the system comprises distinct software modules, and wherein the distinct software modules comprise a signal detection module, a logic processing module, and a data display organization module; (b) detecting signal information indicating the presence, absence or amount of enriched nucleic acid; (c) receiving, by the logic processing module, the signal information; (d) calling the presence or absence of an outcome by the logic processing module; and (e) organizing, by the data display organization model in response to being called by the logic processing module, a data display indicating the presence or absence of the outcome.
[0107] Provided also are methods for identifying the presence or absence of an outcome, which comprise providing signal information indicating the presence, absence or amount of enriched nucleic acid; providing a system, wherein the system comprises distinct software modules, and wherein the distinct software modules comprise a signal detection module, a logic processing module, and a data display organization module; receiving, by the logic processing module, the signal information; calling the presence or absence of an outcome by the logic processing module; and, organizing, by the data display organization model in response to being called by the logic processing module, a data display indicating the presence or absence of the outcome.
[0108] Provided also are methods for identifying the presence or absence of an outcome, which comprise providing a system, wherein the system comprises distinct software modules,and wherein the distinct software modules comprise a signal detection module, a logic processing module, and a data display organization module; receiving, by the logic processing module, signal information indicating the presence, absence or amount of enriched nucleic acid; calling the presence or absence of an outcome by the logic processing module; and, organizing, by the data display organization model in response to being called by the logic processing module, a data display indicating the presence or absence of the outcome.
[0109] By “providing signal information” is meant any manner of providing the information, including, for example, computer communication means from a local, or remote site, human data entry, or any other method of transmitting signal information. The signal information may be generated in one location and provided to another location.
[0110] By “obtaining” or “receiving” signal information is meant receiving the signal information by computer communication means from a local, or remote site, human data entry, or any other method of receiving signal information. The signal information may be generated in the same location at which it is received, or it may be generated in a different location and transmitted to the receiving location.
[0111] By “indicating” or “representing” the amount is meant that the signal information is related to, or correlates with, for example, the amount of enriched nucleic acid or presence or absence of enriched nucleic acid. The information may be, for example, the calculated data associated with the presence or absence of enriched nucleic acid as obtained, for example, after converting raw data obtained by mass spectrometry.
[0112] Also provided are computer program products, such as, for example, a computer program products comprising a computer usable medium having a computer readable program code embodied therein, the computer readable program code adapted to be executed to implement a method for identifying the presence or absence of an outcome, which comprises (a) providing a system, wherein the system comprises distinct software modules, and wherein the distinct software modules comprise a signal detection module, a logic processing module, and a data display organization module; (b) detecting signal information indicating the presence, absence or amount of enriched nucleic acid; (c) receiving, by the logic processing module, the signal information; (d) calling the presence or absence of an outcome by the logic processing module; and, organizing, by the data display organization model in response to being called by the logic processing module, a data display indicating the presence or absence of the outcome.
[0113] Also provided are computer program products, such as, for example, computer program products comprising a computer usable medium having a computer readable program code embodied therein, the computer readable program code adapted to be executed to implement a method for identifying the presence or absence of an outcome, which comprises providing a system, wherein the system comprises distinct software modules, and wherein the distinct software modules comprise a signal detection module, a logic processing module, and a data display organization module; receiving signal information indicating the presence, absence or amount of enriched nucleic acid; calling the presence or absence of an outcome by the logic processing module; and, organizing, by the data display organization model in response to being called by the logic processing module, a data display indicating the presence or absence of the outcome.
[0114] Also provided are methods identifying the presence or absence of an outcome that comprises: (a) detecting signal information, wherein the signal information indicates presence, absence, or amount of methylation of uscfDNA and / or mncfDNA; (b) transforming the signal information into identification data, wherein the identification data represents the presence or absence of the outcome, whereby the presence or absence of the outcome is identified based on the signal information; and (c) displaying the identification data.
[0115] Also provided are methods for identifying the presence or absence of an outcome that comprises:
[0116] (a) providing signal information indicating the presence, absence, or amount of methylation of uscfDNA and / or mncfDNA; (b) transforming the signal information representing into identification data, wherein the identification data represents the presence or absence of the outcome, whereby the presence or absence of the outcome is identified based on the signal information; and (c) displaying the identification data.
[0117] Also provided are methods for identifying the presence or absence of an outcome that comprises:
[0118] (a) receiving signal information indicating the presence, absence, or amount of methylation of uscfDNA and / or mncfDNA; (b) transforming the signal information into identification data, wherein the identification data represents the presence or absence of the outcome, whereby the presence or absence of the outcome is identified based on the signal information; and (c) displaying the identification data.
[0119] For purposes of these, and similar embodiments, the term “signal information” indicates information readable by any electronic media, including, for example, computers that represent data derived using the present methods. For example, “signal information” can represent the amount of methylation at a given region of uscfDNA and / or mncfDNA. Signal information, such as in these examples, that represents physical substances may be transformed into identification data, such as a visual display that represents other physical substances, such as, for example, a DNA methylation profile. Identification data may be displayed in any appropriate manner, including, but not limited to, in a computer visual display, by encoding the identification data into computer readable media that may, for example, be transferred to another electronic device (e.g., electronic record), or by creating a hard copy of the display, such as a printout or physical record of information. The information may also be displayed by auditory signal or any other means of information communication. In some embodiments, the signal information may be detection data obtained using methods to detect methylation of uscfDNA and / or mncfDNA.
[0120] Once the signal information is detected, it may be forwarded to the logic- processing module. The logic-processing module may “call” or “identify” the presence or absence of an outcome.
[0121] The term “identifying the presence or absence of an outcome” or “an increased risk of an outcome,” as used herein refers to any method for obtaining such information, including, without limitation, obtaining the information from a laboratory file. A laboratory file can be generated by a laboratory that carried out an assay to determine the presence or absence of an outcome. The laboratory may be in the same location or different location (e.g., in another country) as the personnel identifying the presence or absence of the outcome from the laboratory file. For example, the laboratory file can be generated in one location and transmitted to another location in which the information therein will be transmitted to the subject. The laboratory file may be in tangible form or electronic form (e.g., computer readable form), in certain embodiments.
[0122] The term “transmitting the presence or absence of the outcome to the subject” or any other information transmitted as used herein refers to communicating the information to the subject, or family member, guardian or designee thereof, in a suitable medium, including, without limitation, in verbal, document, or file form.
[0123] Also provided are methods for providing to a subject a medical prescription based on genetic information, which comprise identifying the presence or absence of an outcome, wherein the presence or absence of the outcome has been determined from the presence, absence, or amount of methylation of uscfDNA and / or mncfDNA from a sample from the subject; and providing a medical prescription based on the presence or absence of the outcome to the subject.
[0124] The term “providing a medical prescription based on genetic information” refers to communicating the prescription to the subject, or family member, guardian, or designee thereof, in a suitable medium, including, without limitation, in verbal, document or file form. The medical prescription may be for any course of action determined by, for example, a medical professional upon reviewing the uscfDNA and / or mncfDNA methylation information. For example, the medical prescription may be for the subject to undergo additional testing or confirmatory testing. In yet another example, the medical prescription may be medical advice to not undergo further testing.
[0125] Also provided are files, such as, for example, a file comprising the presence or absence of outcome for a subject, wherein the presence or absence of the outcome has been determined from the presence, absence, or amount of methylation of uscfDNA and / or mncfDNA in a sample from the subject. The file may be, for example, but not limited to, a computer readable file, a paper file, or a medical record file.
[0126] Computer program products include, for example, any electronic storage medium that may be used to provide instructions to a computer, such as, for example, a removable storage device, CD-ROMS, a hard disk installed in hard disk drive, signals, magnetic tape, DVDs, optical disks, flash drives, RAM or floppy disk, and the like.
[0127] The systems discussed herein may further comprise general components of computer systems, such as, for example, network servers, laptop systems, desktop systems, handheld systems, personal digital assistants, computing kiosks, and the like. The computer system may comprise one or more input means such as a keyboard, touch screen, mouse, voice recognition or other means to allow the user to enter data into the system. The system may further comprise one or more output means such as a CRT or LCD display screen, speaker, FAX machine, impact printer, inkjet printer, black and white or color laser printer or other means of providing visual, auditory or hardcopy output of information.
[0128] The input and output means may be connected to a central processing unit which may comprise among other components, a microprocessor for executing program instructions and memory for storing program code and data. In some embodiments the methods may be implemented as a single user system located in a single geographical site. In other embodiments methods may be implemented as a multi-user system. In the case of a multi-user implementation, multiple central processing units may be connected by means of a network. The network may be local, encompassing a single department in one portion of a building, an entire building, span multiple buildings, span a region, span an entire country or be worldwide. The network may be private, being owned and controlled by the provider, or it may be implemented as an Internet based service where the user accesses a web page to enter and retrieve information.
[0129] The various software modules associated with the implementation of the present products and methods can be suitably loaded into the computer system as desired, or the software code can be stored on a computer-readable medium such as a floppy disk, magnetic tape, or an optical disk, or the like. In an online implementation, a server and web site maintained by an organization can be configured to provide software downloads to remote users. As used herein, “module,” including grammatical variations thereof, means, a self-contained functional unit which is used with a larger system. For example, a software module is a part of a program that performs a particular task. Thus, provided herein is a machine comprising one or more software modules described herein, where the machine can be, but is not limited to, a computer (e.g., server) having a storage device such as floppy disk, magnetic tape, optical disk, random access memory and / or hard disk drive, for example.
[0130] The present methods may be implemented using hardware, software or a combination thereof and may be implemented in a computer system or other processing system. An example computer system may include one or more processors. A processor can be connected to a communication bus. The computer system may include a main memory, sometimes random-access memory (RAM), and can also include a secondary memory. The secondary memory can include, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a magnetic tape drive, an optical disk drive, memory card etc. The removable storage drive reads from and / or writes to a removable storage unit in a well- known manner. A removable storage unit includes, but is not limited to, a floppy disk, magnetic tape, optical disk, etc. which is read by and written to by, for example, a removable storagedrive. As will be appreciated, the removable storage unit includes a computer usable storage medium having stored therein computer software and / or data.
[0131] Alternatively, secondary memory may include other similar means for allowing computer programs or other instructions to be loaded into a computer system. Such means can include, for example, a removable storage unit and an interface device. Examples of such can include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM, or PROM) and associated socket, and other removable storage units and interfaces which allow software and data to be transferred from the removable storage unit to a computer system.
[0132] The computer system may also include a communications interface. A communications interface allows software and data to be transferred between the computer system and external devices. Examples of communications interface can include a modem, a network interface (such as an Ethernet card), a communications port, a PCMCIA slot and card, etc. Software and data transferred via communications interface are in the form of signals, which can be electronic, electromagnetic, optical, or other signals capable of being received by communications interface. These signals are provided to communications interface via a channel. This channel carries signals and can be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, an RF link and other communications channels. Thus, in one example, a communications interface may be used to receive signal information to be detected by the signal detection module.
[0133] In a related aspect, the signal information may be input by a variety of means, including but not limited to, manual input devices or direct data entry devices (DDEs). For example, manual devices may include keyboards, concept keyboards, touch sensitive screens, light pens, mouse, tracker balls, joysticks, graphic tablets, scanners, digital cameras, video digitizers and voice recognition devices. DDEs may include, for example, bar code readers, magnetic strip codes, smart cards, magnetic ink character recognition, optical character recognition, optical mark recognition, and turnaround documents. In one embodiment, an output from a gene or chip reader may serve as an input signal. Determining Effectiveness of Therapy or Prognosis
[0134] In one aspect, the level of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, in a biological sample of a patient is used to monitor the effectiveness of treatment or the prognosis of disease. In some embodiments, the level of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, in a test sample obtained from a treated patient can be compared to the level from a reference sample obtained from that patient before initiation of a treatment. Clinical monitoring of treatment typically entails that each patient serves as his or her own baseline control. In some embodiments, test samples are obtained at multiple time points following administration of the treatment. In these embodiments, measurement of the level of methylation of one or more uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, in the test samples provides an indication of the extent and duration of in vivo effect of the treatment.
[0135] Measurement of the total level of methylation of one or more uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, may allow for the course of treatment of a disease to be monitored. The effectiveness of a treatment regimen for a disease can be monitored by detecting methylation of one or more uscfDNA and / or mncfDNA in an effective amount from samples obtained from a subject over time and comparing the detected level of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein. For example, a first sample can be obtained before the subject receives treatment and one or more subsequent samples are taken after or during treatment of the subject. Changes in methylation profile of uscfDNA and / or mncfDNA across the samples may provide an indication as to the effectiveness of the therapy.
[0136] In some embodiments, the disclosure provides a method for monitoring the total levels of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, in response to treatment. For example, in certain embodiments, the disclosure provides for a method of determining the efficacy of treatment in a subject, by measuring the total levels of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, as described herein. In some embodiments, the level of methylation profile of the uscfDNA and / or mncfDNA can be measured over time, where the level at one timepoint after the initiation of treatment is compared to the level at another timepoint after the initiation of treatment. In some embodiments, the level of methylation of the uscfDNA and / or mncfDNA, or a methylationbiomarker identified therein, can be measured over time, where the level at one timepoint after the initiation of treatment is compared to the level before initiation of treatment.
[0137] In some embodiments, uscfDNA and / or mncfDNA total methylation levels or methylation biomarkers can be used to identify therapeutics or drugs that are appropriate for a specific subject. For example, a test sample from the subject can be exposed to a therapeutic agent or a drug, and the level of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, can be determined. The level of methylation of uscfDNA and / or mncfDNA, or a methylation biomarker identified therein, can be compared to a sample derived from the subject before and after treatment or exposure to a therapeutic agent or a drug or can be compared to samples derived from one or more subjects who have shown improvements relative to a disease as a result of such treatment or exposure. Thus, in one aspect, the disclosure provides a method of assessing the efficacy of a therapy with respect to a subject comprising taking a first measurement of uscfDNA and / or mncfDNA methylation in a first sample from the subject; effecting the therapy with respect to the subject; taking a second measurement of the uscfDNA and / or mncfDNA methylation in a second sample from the subject and comparing the first and second measurements to assess the efficacy of the therapy.
[0138] Accordingly, treatments or therapeutic regimens for use in can be selected based on the amounts of a specific methylation biomarker in uscfDNA and / or mncfDNA or total methylation of uscfDNA and / or mncfDNA in samples obtained from the subjects and compared to a reference value. Two or more treatments or therapeutic regimens can be evaluated in parallel to determine which treatment or therapeutic regimen would be the most efficacious for use in a subject to delay onset, or slow progression of a disease. In various embodiments, a recommendation is made on whether to initiate or continue treatment of a disease.
[0139] A prognosis may be expressed as the amount of time a patient can be expected to survive. Alternatively, a prognosis may refer to the likelihood that the disease goes into remission or to the amount of time the disease can be expected to remain in remission. Prognosis can be expressed in various ways; for example, prognosis can be expressed as a percent chance that a patient will survive after one year, five years, ten years, or the like. Alternatively, prognosis may be expressed as the number of years, on average, that a patient can expect to survive as a result of a condition or disease. The prognosis of a patient may be considered as an expression of relativism, with many factors affecting the ultimate outcome. For example, forpatients with certain conditions, prognosis can be appropriately expressed as the likelihood that a condition may be treatable or curable, or the likelihood that a disease will go into remission, whereas for patients with more severe conditions, prognosis may be more appropriately expressed as likelihood of survival for a specified period of time. Additionally, a change in a clinical factor from a baseline level may impact a patient's prognosis, and the degree of change in level of the clinical factor may be related to the severity of adverse events. Statistical significance is often determined by comparing two or more populations and determining a confidence interval and / or a p value.
[0140] Multiple determinations of uscfDNA and / or mncfDNA methylation can be made, and a temporal change in uscfDNA and / or mncfDNA methylation can be used to determine a prognosis. For example, comparative measurements are made of the uscfDNA and / or mncfDNA methylation in a patient at multiple time points, and a comparison of the uscfDNA and / or mncfDNA methylation at two or more time points may be indicative of a particular prognosis.
[0141] In certain embodiments, other prognostic factors may be combined with the uscfDNA and / or mncfDNA methylation level or other biomarkers in the algorithm to determine prognosis with greater accuracy. Exemplary additional prognostic factors may include one or more prognostic factors selected from the group consisting of cytogenetics, performance status, age, gender, and contemporary diagnosis. Treatments
[0142] In one aspect, the disclosure provides a method of diagnosing, treating, or preventing a disease or disorder associated with a methylation biomarker identified from analysis of uscfDNA and / or mncfDNA or a general increase or decrease of total uscfDNA and / or mncfDNA methylation. In some embodiments, the method comprises administering to the subject an effective amount of a pharmaceutical agent for the treatment of a disease or disorder identified associated with a methylation biomarker identified from analysis of uscfDNA and / or mncfDNA or a general increase or decrease of total uscfDNA and / or mncfDNA methylation.
[0143] In some embodiments, the disease or disorder is cancer. In some embodiments, the disease or disorder is non-small cell lung cancer. Therapy for cancer includes, but is not limited to, surgery, chemotherapy, chemotherapeutic agent, radiation therapy, or hormonal therapy or a combination thereof.
[0144] Chemotherapeutic agents include cytotoxic agents (e.g., 5-fluorouracil, cisplatin, carboplatin, methotrexate, daunorubicin, doxorubicin, vincristine, vinblastine, oxorubicin, carmustine (BCNU), lomustine (CCNU), cytarabine USP, cyclophosphamide, estramucine phosphate sodium, altretamine, hydroxyurea, ifosfamide, procarbazine, mitomycin, busulfan, cyclophosphamide, mitoxantrone, carboplatin, cisplatin, interferon alfa-2a recombinant, paclitaxel, teniposide, and streptozoci), cytotoxic alkylating agents (e.g., busulfan, chlorambucil, cyclophosphamide, melphalan, or ethylesulfonic acid), alkylating agents (e.g., asaley, AZQ, BCNU, busulfan, bisulphan, carboxyphthalatoplatinum, CBDCA, CCNU, CHIP, chlorambucil, chlorozotocin, cis-platinum, clomesone, cyanomorpholinodoxorubicin, cyclodisone, cyclophosphamide, dianhydrogalactitol, fluorodopan, hepsulfam, hycanthone, iphosphamide, melphalan, methyl CCNU, mitomycin C, mitozolamide, nitrogen mustard, PCNU, piperazine, piperazinedione, pipobroman, porfiromycin, spirohydantoin mustard, streptozotocin, teroxirone, tetraplatin, thiotepa, triethylenemelamine, uracil nitrogen mustard, and Yoshi-864), antimitotic agents (e.g., allocolchicine, Halichondrin M, colchicine, colchicine derivatives, dolastatin 10, maytansine, rhizoxin, paclitaxel derivatives, paclitaxel, thiocolchicine, trityl cysteine, vinblastine sulfate, and vincristine sulfate), plant alkaloids (e.g., actinomycin D, bleomycin, L-asparaginase, idarubicin, vinblastine sulfate, vincristine sulfate, mitramycin, mitomycin, daunorubicin, VP-16- 213, VM-26, navelbine and taxotere), biologicals (e.g., alpha interferon, BCG, G-CSF, GM-CSF, and interleukin-2), topoisomerase I inhibitors (e.g., camptothecin, camptothecin derivatives, and morpholinodoxorubicin), topoisomerase II inhibitors (e.g., mitoxantron, amonafide, m-AMSA, anthrapyrazole derivatives, pyrazoloacridine, bisantrene HCL, daunorubicin, deoxydoxorubicin, menogaril, N,N-dibenzyl daunomycin, oxanthrazole, rubidazone, VM-26 and VP-16), and synthetics (e.g., hydroxyurea, procarbazine, o,p'-DDD, dacarbazine, CCNU, BCNU, cis- diamminedichloroplatimun, mitoxantrone, CBDCA, levamisole, hexamethylmelamine, all-trans retinoic acid, gliadel and porfimer sodium).
[0145] Antiproliferative agents are compounds that decrease the proliferation of cells. Antiproliferative agents include alkylating agents, antimetabolites, enzymes, biological response modifiers, miscellaneous agents, hormones and antagonists, androgen inhibitors (e.g., flutamide and leuprolide acetate), antiestrogens (e.g., tamoxifen citrate and analogs thereof, toremifene, droloxifene and roloxifene), Additional examples of specific antiproliferative agents include, butare not limited to levamisole, gallium nitrate, granisetron, sargramostim strontium-89 chloride, filgrastim, pilocarpine, dexrazoxane, and ondansetron.
[0146] Other anti-tumor agents include cytotoxic / antineoplastic agents and anti- angiogenic agents. Cytotoxic / anti-neoplastic agents are defined as agents which attack and kill cancer cells. Some cytotoxic / anti-neoplastic agents are alkylating agents, which alkylate the genetic material in tumor cells, e.g., cis-platin, cyclophosphamide, nitrogen mustard, trimethylene thiophosphoramide, carmustine, busulfan, chlorambucil, belustine, uracil mustard, chlomaphazin, and dacabazine. Other cytotoxic / anti-neoplastic agents are antimetabolites for tumor cells, e.g., cytosine arabinoside, fluorouracil, methotrexate, mercaptopuirine, azathioprime, and procarbazine. Other cytotoxic / anti-neoplastic agents are antibiotics, e.g., doxorubicin, bleomycin, dactinomycin, daunorubicin, mithramycin, mitomycin, mytomycin C, and daunomycin. There are numerous liposomal formulations commercially available for these compounds. Still other cytotoxic / anti-neoplastic agents are mitotic inhibitors (vinca alkaloids). These include vincristine, vinblastine and etoposide. Miscellaneous cytotoxic / anti-neoplastic agents include taxol and its derivatives, L-asparaginase, anti-tumor antibodies, dacarbazine, azacytidine, amsacrine, melphalan, VM-26, ifosfamide, mitoxantrone, and vindesine.
[0147] Anti-angiogenic agents are well known to those of skill in the art. Suitable anti- angiogenic agents for use in the methods and compositions of the present disclosure include anti- VEGF antibodies, including humanized and chimeric antibodies, anti-VEGF aptamers and antisense oligonucleotides. Other known inhibitors of angiogenesis include angiostatin, endostatin, interferons, interleukin 1 (including alpha and beta) interleukin 12, retinoic acid, and tissue inhibitors of metalloproteinase-1 and -2. (TIMP-1 and -2). Small molecules, including topoisomerases such as razoxane, a topoisomerase II inhibitor with anti-angiogenic activity, can also be used.
[0148] Other anti-cancer agents that can be used include, but are not limited to: acivicin; aclarubicin; acodazole hydrochloride; acronine; adozelesin; aldesleukin; altretamine; ambomycin; ametantrone acetate; aminoglutethimide; amsacrine; anastrozole; anthramycin; asparaginase; asperlin; azacitidine; azetepa; azotomycin; batimastat; benzodepa; bicalutamide; bisantrene hydrochloride; bisnafide dimesylate; bizelesin; bleomycin sulfate; brequinar sodium; bropirimine; busulfan; cactinomycin; calusterone; caracemide; carbetimer; carboplatin; carmustine; carubicin hydrochloride; carzelesin; cedefingol; chlorambucil; cirolemycin;cisplatin; cladribine; crisnatol mesylate; cyclophosphamide; cytarabine; dacarbazine; dactinomycin; daunorubicin hydrochloride; decitabine; dexormaplatin; dezaguanine; dezaguanine mesylate; diaziquone; docetaxel; doxorubicin; doxorubicin hydrochloride; droloxifene; droloxifene citrate; dromostanolone propionate; duazomycin; edatrexate; eflornithine hydrochloride; elsamitrucin; enloplatin; enpromate; epipropidine; epirubicin hydrochloride; erbulozole; esorubicin hydrochloride; estramustine; estramustine phosphate sodium; etanidazole; etoposide; etoposide phosphate; etoprine; fadrozole hydrochloride; fazarabine; fenretinide; floxuridine; fludarabine phosphate; fluorouracil; fluorocitabine; fosquidone; fostriecin sodium; gemcitabine; gemcitabine hydrochloride; hydroxyurea; idarubicin hydrochloride; ifosfamide; ilmofosine; interleukin II (including recombinant interleukin II, or rIL2), interferon alfa-2a; interferon alfa-2b; interferon alfa-n1; interferon alfa-n3; interferon beta- I a; interferon gamma-I b; iproplatin; irinotecan hydrochloride; lanreotide acetate; letrozole; leuprolide acetate; liarozole hydrochloride; lometrexol sodium; lomustine; losoxantrone hydrochloride; masoprocol; maytansine; mechlorethamine hydrochloride; megestrol acetate; melengestrol acetate; melphalan; menogaril; mercaptopurine; methotrexate; methotrexate sodium; metoprine; meturedepa; mitindomide; mitocarcin; mitocromin; mitogillin; mitomalcin; mitomycin; mitosper; mitotane; mitoxantrone hydrochloride; mycophenolic acid; nocodazole; nogalamycin; ormaplatin; oxisuran; paclitaxel; albumin-bound paclitaxel; pegaspargase; peliomycin; pentamustine; peplomycin sulfate; perfosfamide; pipobroman; piposulfan; piroxantrone hydrochloride; plicamycin; plomestane; porfimer sodium; porfiromycin; prednimustine; procarbazine hydrochloride; puromycin; puromycin hydrochloride; pyrazofurin; riboprine; rogletimide; safingol; safingol hydrochloride; semustine; simtrazene; sparfosate sodium; sparsomycin; spirogermanium hydrochloride; spiromustine; spiroplatin; streptonigrin; streptozocin; sulofenur; talisomycin; tecogalan sodium; tegafur; teloxantrone hydrochloride; temoporfin; teniposide; teroxirone; testolactone; thiamiprine; thioguanine; thiotepa; tiazofurin; tirapazamine; toremifene citrate; trestolone acetate; triciribine phosphate; trimetrexate; trimetrexate glucuronate; triptorelin; tubulozole hydrochloride; uracil mustard; uredepa; vapreotide; verteporfin; vinblastine sulfate; vincristine sulfate; vindesine; vindesine sulfate; vinepidine sulfate; vinglycinate sulfate; vinleurosine sulfate; vinorelbine; vinorelbine tartrate; vinrosidine sulfate; vinzolidine sulfate; vorozole; zeniplatin; zinostatin; zorubicin hydrochloride. Other anti-cancer drugs include, but are not limited to: 20-epi-1,25 dihydroxyvitamin D3; 5-ethynyluracil; abiraterone; aclarubicin; acylfulvene; adecypenol; adozelesin; aldesleukin; ALL- TK antagonists; altretamine; ambamustine; amidox; amifostine; aminolevulinic acid; amrubicin; amsacrine; anagrelide; anastrozole; andrographolide; angiogenesis inhibitors; antagonist D; antagonist G; antarelix; anti-dorsalizing morphogenetic protein-1; antiandrogen, prostatic carcinoma; antiestrogen; antineoplaston; antisense oligonucleotides; aphidicolin glycinate; apoptosis gene modulators; apoptosis regulators; apurinic acid; ara-CDP-DL-PTBA; arginine deaminase; asulacrine; atamestane; atrimustine; axinastatin 1; axinastatin 2; axinastatin 3; azasetron; azatoxin; azatyrosine; baccatin III derivatives; balanol; batimastat; BCR / ABL antagonists; benzochlorins; benzoylstaurosporine; beta lactam derivatives; beta-alethine; betaclamycin B; betulinic acid; bFGF inhibitor; bicalutamide; bisantrene; bisaziridinylspermine; bisnafide; bistratene A; bizelesin; breflate; bropirimine; budotitane; buthionine sulfoximine; calcipotriol; calphostin C; camptothecin derivatives; canarypox IL-2; capecitabine; carboxamide- amino-triazole; carboxyamidotriazole; CaRest M3; CARN 700; cartilage derived inhibitor; carzelesin; casein kinase inhibitors (ICOS); castanospermine; cecropin B; cetrorelix; chlorins; chloroquinoxaline sulfonamide; cicaprost; cis-porphyrin; cladribine; clomifene analogues; clotrimazole; collismycin A; collismycin B; combretastatin A4; combretastatin analogue; conagenin; crambescidin 816; crisnatol; cryptophycin 8; cryptophycin A derivatives; curacin A; cyclopentanthraquinones; cycloplatam; cypemycin; cytarabine ocfosfate; cytolytic factor; cytostatin; dacliximab; decitabine; dehydrodidemnin B; deslorelin; dexamethasone; dexifosfamide; dexrazoxane; dexverapamil; diaziquone; didemnin B; didox; diethylnorspermine; dihydro-5-azacytidine; dihydrotaxol, 9-; dioxamycin; diphenyl spiromustine; docetaxel; docosanol; dolasetron; doxifluridine; droloxifene; dronabinol; duocarmycin SA; ebselen; ecomustine; edelfosine; edrecolomab; eflornithine; elemene; emitefur; epirubicin; epristeride; estramustine analogue; estrogen agonists; estrogen antagonists; etanidazole; etoposide phosphate; exemestane; fadrozole; fazarabine; fenretinide; filgrastim; finasteride; flavopiridol; flezelastine; fluasterone; fludarabine; fluorodaunorunicin hydrochloride; forfenimex; formestane; fostriecin; fotemustine; gadolinium texaphyrin; gallium nitrate; galocitabine; ganirelix; gelatinase inhibitors; gemcitabine; glutathione inhibitors; hepsulfam; heregulin; hexamethylene bisacetamide; hypericin; ibandronic acid; idarubicin; idoxifene; idramantone; ilmofosine; ilomastat; imidazoacridones; imiquimod; immunostimulant peptides; insulin-like growth factor-1 receptor inhibitor; interferon agonists; interferons; interleukins; iobenguane; iododoxorubicin;ipomeanol, 4-; iroplact; irsogladine; isobengazole; isohomohalicondrin B; itasetron; jasplakinolide; kahalalide F; lamellarin-N triacetate; lanreotide; leinamycin; lenograstim; lentinan sulfate; leptolstatin; letrozole; leukemia inhibiting factor; leukocyte alpha interferon; leuprolide+estrogen+progesterone; leuprorelin; levamisole; liarozole; linear polyamine analogue; lipophilic disaccharide peptide; lipophilic platinum compounds; lissoclinamide 7; lobaplatin; lombricine; lometrexol; lonidamine; losoxantrone; lovastatin; loxoribine; lurtotecan; lutetium texaphyrin; lysofylline; lytic peptides; maitansine; mannostatin A; marimastat; masoprocol; maspin; matrilysin inhibitors; matrix metalloproteinase inhibitors; menogaril; merbarone; meterelin; methioninase; metoclopramide; MIF inhibitor; mifepristone; miltefosine; mirimostim; mismatched double stranded RNA; mitoguazone; mitolactol; mitomycin analogues; mitonafide; mitotoxin fibroblast growth factor-saporin; mitoxantrone; mofarotene; molgramostim; monoclonal antibody, human chorionic gonadotrophin; monophosphoryl lipid A+myobacterium cell wall sk; mopidamol; multiple drug resistance gene inhibitor; multiple tumor suppressor 1- based therapy; mustard anticancer agent; mycaperoxide B; mycobacterial cell wall extract; myriaporone; N-acetyldinaline; N-substituted benzamides; nafarelin; nagrestip; naloxone+pentazocine; napavin; naphterpin; nartograstim; nedaplatin; nemorubicin; neridronic acid; neutral endopeptidase; nilutamide; nisamycin; nitric oxide modulators; nitroxide antioxidant; nitrullyn; O6-benzylguanine; octreotide; okicenone; oligonucleotides; onapristone; ondansetron; ondansetron; oracin; oral cytokine inducer; ormaplatin; osaterone; oxaliplatin; oxaunomycin; paclitaxel; paclitaxel analogues; paclitaxel derivatives; palauamine; palmitoylrhizoxin; pamidronic acid; panaxytriol; panomifene; parabactin; pazelliptine; pegaspargase; peldesine; pentosan polysulfate sodium; pentostatin; pentrozole; perflubron; perfosfamide; perillyl alcohol; phenazinomycin; phenylacetate; phosphatase inhibitors; picibanil; pilocarpine hydrochloride; pirarubicin; piritrexim; placetin A; placetin B; plasminogen activator inhibitor; platinum complex; platinum compounds; platinum-triamine complex; porfimer sodium; porfiromycin; prednisone; propyl bis-acridone; prostaglandin J2; proteasome inhibitors; protein A-based immune modulator; protein kinase C inhibitor; protein kinase C inhibitors, microalgal; protein tyrosine phosphatase inhibitors; purine nucleoside phosphorylase inhibitors; purpurins; pyrazoloacridine; pyridoxylated hemoglobin polyoxyethylene conjugate; raf antagonists; raltitrexed; ramosetron; ras farnesyl protein transferase inhibitors; ras inhibitors; ras- GAP inhibitor; retelliptine demethylated; rhenium Re 186 etidronate; rhizoxin; ribozymes; RIIretinamide; rogletimide; rohitukine; romurtide; roquinimex; rubiginone B1; ruboxyl; safingol; saintopin; SarCNU; sarcophytol A; sargramostim; Sdi 1 mimetics; semustine; senescence derived inhibitor 1; sense oligonucleotides; signal transduction inhibitors; signal transduction modulators; single chain antigen binding protein; sizofuran; sobuzoxane; sodium borocaptate; sodium phenylacetate; solverol; somatomedin binding protein; sonermin; sparfosic acid; spicamycin D; spiromustine; splenopentin; spongistatin 1; squalamine; stem cell inhibitor; stem- cell division inhibitors; stipiamide; stromelysin inhibitors; sulfinosine; superactive vasoactive intestinal peptide antagonist; suradista; suramin; swainsonine; synthetic glycosaminoglycans; tallimustine; tamoxifen methiodide; tauromustine; tazarotene; tecogalan sodium; tegafur; tellurapyrylium; telomerase inhibitors; temoporfin; temozolomide; teniposide; tetrachlorodecaoxide; tetrazomine; thaliblastine; thiocoraline; thrombopoietin; thrombopoietin mimetic; thymalfasin; thymopoietin receptor agonist; thymotrinan; thyroid stimulating hormone; tin ethyl etiopurpurin; tirapazamine; titanocene bichloride; topsentin; toremifene; totipotent stem cell factor; translation inhibitors; tretinoin; triacetyluridine; triciribine; trimetrexate; triptorelin; tropisetron; turosteride; tyrosine kinase inhibitors; tyrphostins; UBC inhibitors; ubenimex; urogenital sinus-derived growth inhibitory factor; urokinase receptor antagonists; vapreotide; variolin B; vector system, erythrocyte gene therapy; velaresol; veramine; verdins; verteporfin; vinorelbine; vinxaltine; vitaxin; vorozole; zanoterone; zeniplatin; zilascorb; imilimumab; mirtazapine; BrUOG 278; BrUOG 292; RAD0001; CT-011; folfirinox; tipifarnib; R115777; LDE225; calcitriol; AZD6244; AMG 655; AMG 479; BKM120; mFOLFOX6; NC-6004; cetuximab; IM-C225; LGX818; MEK162; BBI608; MEDI4736; vemurafenib; ipilimumab; ivolumab; nivolumab; panobinostat; leflunomide; CEP-32496; alemtuzumab; bevacizumab; ofatumumab; panitumumab; pembrolizumab; rituximab; trastuzumab; STAT3 inhibitors (e.g., STA-21, LLL-3, LLL12, XZH-5, S31-201, SF-1066, SF-1087, STX-0119, cryptotanshinone, curcumin, diferuloylmethane, FLLL11, FLLL12, FLLL32, FLLL62, C3, C30, C188, C188-9, LY5, OPB-31121, pyrimethamine, OPB-51602, AZD9150, etc.); hypoxia inducing factor 1 (HIF-1) inhibitors (e.g., LW6, digoxin, laurenditerpenol, PX-478, RX-0047, vitexin, KC7F2, YC-1, etc.) zinostatin stimalamer, Lynparza (olaparib), talazoparib, niraparib, and rucaparib.
[0149] In certain embodiments, immunosuppressive agents may be used to treat cancer, such as cyclosporin, azathioprine, methotrexate, mycophenolate, and FK506, antibodies, or other immunoablative agents such as CAM PATH, anti-CD3 antibodies or other antibody therapies,cytoxin, fludaribine, cyclosporin, FK506, rapamycin, mycophenolic acid, steroids, FR901228, cytokines, and irradiation. These drugs inhibit either the calcium dependent phosphatase calcineurin (cyclosporine and FK506) or inhibit the p70S6 kinase that is important for growth factor induced signaling (rapamycin) (Liu et al., Cell 66:807-815, 1991; Henderson et al., Immun.73:316-321, 1991; Bierer et al., Curr. Opin. Immun.5:763-773, 1993).
[0150] In some embodiments, cancer treatment includes bone marrow transplantation, T cell ablative therapy using either chemotherapy agents such as, fludarabine, external-beam radiation therapy (XRT), cyclophosphamide, or antibodies such as OKT3 or CAMPATH, B-cell ablative therapy such as agents that react with CD20, e.g., Rituxan, Ospemifene, Tamoxifen, Raloxifene, or other drugs such as ICI 182,780 and RU 58668. Tamoxifen and Raloxifene may act as partial antiestrogens, and the drugs such as ICI 182,780 and RU 58668 may act as full antiestrogens. In another embodiment, cancer treatment includes aromatase inhibitors. Non- limiting examples of aromatase inhibitors include Exemestane, Letrozole, and Anastrozole. In one embodiment, the therapeutic agent is gemcitabine.
[0151] In certain embodiments, cancer is treated with drugs that target DNA repair factors (i.e. PARP1, PARG, ATM, ATR, DNApk, RAD51, CHK1, WEE1, topoisomerase I, topoisomerase II) and / or act as genotoxic agents (i.e. chemotherapies and radiation / radiotherapy, proton therapy) and induce DNA damage. In some embodiments, the cancer drugs that target DNA repair factors are one or more selected from the group consisting of DNA damage agents, platinum agents, DNA damage response inhibitors, ATR inhibitors, WEE1 inhibitors, NDA-PK inhibitors, and ATM inhibitors. EXPERIMENTAL EXAMPLES
[0152] The invention is further described in detail by reference to the following experimental examples. These examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following examples, but rather, should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.
[0153] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize thepresent invention and practice the claimed methods. The following working examples therefore are not to be construed as limiting in any way the remainder of the disclosure. Example 1: Single-stranded Premethylated 5mC Adapters Uncovers the Methylation Profile of Plasma Ultrashort Single-Stranded Cell-Free DNA
[0154] In this report, an optimized library preparation protocol is described for cfDNA in which single-stranded 5mC premethylated adapters are ligated to heat-denatured DNA fragments prior to bisulfite conversion and sequencing (5mCAdpBS-Seq). This method improves the accuracy of downstream analysis by preventing bisulfite conversion degraded DNA from being incorporated into the final library and masking the methylation signal of uscfDNA. Using the 5mCAdpBS-Seq protocol, the CpG methylation % of uscfDNA was observed to be 60% compared to 70-80% in mncfDNA. The unique methylation patterns, genomic location, and strandedness (Cheng et al., 2022, iScience, 25, 104554; Cheng et al., 2023, Clin. Chem., 10.1093 / clinchem / hvad131) suggest that uscfDNA could originate through a different mechanism than mncfDNA, which is worth further exploring. Lower levels of DNA methylation observed in uscfDNA could be due to the inherently lower methylation levels found due to expression activity in genomic DNA, or because they are subject to further enzymatic modifications (such as TETs-mediated DNA demethylation) after being “detached” from genomic DNA or when entering circulation.
[0155] Methodologically, several strategies for directly detecting 5mC in cfDNA are limited to low throughput single molecule techniques such as nanopore sequencing (Rand et al., 2017, Nat. Methods, 14, 411–413) or single-molecule polymerase fluorescent labeling (Flusberg et al., 2010, Nat. Methods, 7, 461–465). Most methylation workflows require pretreatment of the DNA fragments to indicate the CpG site methylation status for downstream analysis. The main methods of pretreatment are bisulfite treatment (performed in this paper), restriction enzyme digestion prior to bisulfite treatment (e.g., RRBS, MRE-BS), affinity enrichment, or other combinatory methods (Table 1). Targeted sequencing coupled with bisulfite conversion and microarray-based methods were not explored in this study as its primary focus was the genome- wide profile of uscfDNA.Table 1. Summary Methylation Analysis Techniques for cfDNA Technique Single Portrays DNA Optimized for Nucleotide non-CpG Degradation uscfDNA Resolution sites BS-Seq (used in this paper) Yes Yes Yes Not currently 5mCAdpBS-Seq Yes Yes Yes Yes (developed in this paper) Reduced representation bisulfite Yes Yes* Yes Not currently sequencing (RRBS) Circulating free methylated DNA No No No Not currently immunoprecipitation sequencing (cfMeDIP-Seq) Methyl-CpG binding domain No No No Not currently protein capture sequencing (MBD- Seq) Enzymatic Methyl Seq Yes Yes No Not currently (EM-Seq) Targeted Sequencing BS Yes Yes* Yes Not currently Microarray-based Yes No Yes Not currently *only in the enriched regions
[0156] Reduced representation bisulfite sequencing (RRBS) is based on the digestion of genomic DNA by methylation-insensitive restriction enzymes (such as MspI), with the intent to enrich CpG-dense regions (Wang et al., 2013, BMC Genomics, 14, 11). RBBS has been adapted to cfDNA analysis (Stackpole et al., 2022, Nat. Commun., 13, 5566). Since RRBS is based on double-stranded DNA-cutting enzymes, it is not compatible with uscfDNA methylation profiling unless a prior second-strand synthesis is incorporated. Other enzyme-based approaches, such as MRE-seq and MRE-BS-seq, suffer from the same problems as RRBS. Another strategy is to use 5mC-specific antibodies (meDIP-Seq) or methyl-binding proteins (MBD-Seq) to enrich the content of methylated DNA. In cfmeDIP-Seq, a monoclonal antibody against 5mC is immunoprecipitated with heat-denatured DNA and assessed with PCR, sequencing, or an array (Shen et al., 2018, Nature, 563, 579–583). Alternatively, MethylCap uses GST-MBD fusion protein to capture methylated CpG-containing molecules (Brinkman et al., 2010, Methods SanDiego Calif, 52, 232–236). However, these techniques do not have single nucleotide resolution, and the small size of uscfDNA fragments can affect the immunoprecipitation efficiency (fewer CpG sites / fragment).
[0157] Enzymatic conversion promises lower DNA degradation and improved library yield but is time-consuming and may not have the equivalent conversion efficiency (Zheng et al., 2022, 10.1101 / 2022.01.12.475986). Preliminary experiments with the enzymatic conversion did not generate libraries of sufficient quality to sequence (Figure 21). Enzymatic conversion may not be optimized for single-stranded DNA (Vaisvila et al., 2021, Genome Res., 31, 1280– 1289) since the initial ten-eleven translocation2 (TET2) oxidation step has a preference for double-stranded DNA compared to single-stranded DNA or RNA (Leddin et al., 2019, Adv. Protein Chem. Struct. Biol., 117, 91–112). Ten-eleven translocation (TET)-assisted pyridine borate sequencing (TAPS) is another method that manipulates the identity of methylated CpG sites.5mC and 5hmC are oxidized to 5-carboxylcytosine (5caC), and using pyridine, borane is reduced to dihydrouracil (DHU). During a final PCR step, DHU is converted to thymine. To evaluate the applicability to uscfDNA, these conversion methods may need further optimization (DeNizio et al., 2019, Biochemistry, 58, 411–421). For those reasons, a bisulfite-based methodology was used for the first foray into studying the methylation profile of uscfDNA.
[0158] Bioinformatically, the two-peaked mononucleosomal profile seen from the paired-end processing (Figure 1D) could be explained by orphan reads of various lengths leading to a disproportioned accumulation of fragments up to a maximum of 150 bases. This pattern, up to the 150-base demarcation, matches the size distribution pattern of the merged protocol (Figure 1C and D). Beyond 150 bases, the pattern also matches but with a decreased abundance, suggesting an artifactual proportional decrease. By filtering for only properly paired reads, the size distribution pattern resultingly resembles the merged pipeline (Figure 22). Therefore, the merged reads protocol not only “fixes” the double-peak by removing these peaks but also filters for fragments of high confidence since both paired reads must match to proceed with alignment.
[0159] Using unmethylated non-human lambda spike-in, the conversion efficiency was shown to be >99% and >98.5% for CpG and non-CpG methylation (Figure 3C and F). Interestingly, there was a slight increase in cytosine methylation levels in the 5mCAdpBS-Seq protocol for the digested lambda reads (still lower than 0.8%) (Figure 8C and Figure 9D). Allexperiments should use CpG methylation of lambda as a quality control of bisulfite conversion efficiency.
[0160] Since the mitochondrial genome has been described to contain low or absent CpG% methylation (Liu et al., 2016, Sci. Rep., 6, 23421; Mechta et al., 2017, Front. Genet., 8, 166), it can act as a biological internal negative control for the 5mCAdpBS-Seq Protocol. Low levels (<2%) of both CpG and non-CpG methylation were observed in mitochondria cfDNA fragments (30-75 bases), suggesting that the workflow did not artificially over-represent methylation levels. There was a pattern of increasing methylation variability in fragments in bins >150 bases, potentially due to the lower number of reads in this fraction (Figure 8B, 8E and Figure 5B and 5C).
[0161] Regarding bisulfite-induced degradation, higher molecular weight cfDNA has been documented to be more susceptible compared to mncfDNA (Werner et al., 2019, PLoS ONE, 14, e0224338). Therefore, during the BS-Seq protocol, the observed degraded DNA likely originated from these larger fragments of cfDNA. The CpG residues of genomic DNA are reportedly 70-80% hypermethylated (Strichman-Almashanu et al., 2002, Genome Res., 12, 543– 554), and both sources of degraded fragments (either genomic DNA or high molecular weight DNA, which also derives from apoptosis) would still be expected to carry these characteristics. This would explain why, during the BS-Seq protocol, the “bleeding” of the degraded DNA into the 40-100 bases fraction skewed the average of CpG methylation% towards higher levels (closer to 80%). By contrast, except for neurons and stem cells, non-CpG methylation is considered indistinguishable from non-conversion rates for most cell types, reflecting the low non-CpG methylation observed in this study (Titcombe et al., 2022, Epigenetics, 17, 653–664) (Figure 8D and Figure 9A).
[0162] Both the BS-Seq and 5mCAdpBS-Seq protocols indicated that uscfDNA appears to have a lower CpG methylation %. It is unclear if the lower CpG methylation % of uscfDNA is from an alternative mechanism separate from mncfDNA undergoing further fragmentation or if uscfDNA is disproportionally derived from genomic regions that take on hypomethylated states during cell activity. Supporting the latter hypothesis, the uscfDNA fragments were enriched occupancy of regions categorized as simple repeat, promoters, exon, 5UTR, and CpG-Island elements regions, whereas the mncfDNA bin was increased in SINE and intergenic elements. The enrichment in promoters of uscfDNA was previously demonstrated in BRcfDNA-Seq andsimilar studies (Cheng et al., 2022, iScience, 25, 104554; Hisano et al., 2021, BMC Biol., 19, 225).
[0163] Additionally, the uscfDNA fragments had the highest enrichment in H3K4me3 and hypomethylated regions when compared to the control (random genomic regions) (Figure 6C). These genome regions may exhibit a more accessible chromatin organization to nucleases, generating hypomethylated uscfDNA in circulation (Teif et al., 2014, Genome Res., 24, 1285– 1295; Domcke et al., 2015, Nature, 528, 575–579). Another study has reported that the pattern of cfDNA fragmentation of H3K4me3 resembles the fragmentation pattern of regions of housekeeping genes in contrast to H3K9me3, which matches repressed genes (Guo et al., 2020, BMC Genomics, 21, 473). That report did not include uscfDNA analysis, which might have demonstrated an even more distinct fragment pattern between active and non-active regions of the genome. To confirm these findings, ChIP assays could be performed on plasma to determine if uscfDNA is immunoprecipitated with the proteins assayed. Another intriguing possibility is that DNA structures themselves, such as G-Quads, could confer protection from circulating nucleases. Functionally, G-Quad structures have been associated with open chromatin regions near promoters and are linked with increased transcription through specific recognition by transcription factors (Lago et al., 2021, Nat. Commun., 12, 3885; Esnault et al., 2023, Nat. Genet., 55, 1359–1369).
[0164] Further support for the hypothesis that a subpopulation of uscfDNA can report gene expression activity, the data was able to recapitulate the prior observation that ultrashort fragments are enriched along the transcription start sites versus mncfDNA, which has decreased coverage (Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121; Snyder et al., 2016, Cell, 164, 57–68; Ulz et al., 2016, Nat. Genet., 48, 1273–1278) (Figure 12A). Hypomethylated uscfDNA fragments showed an increased enrichment through the TSSs of highly expressed hemopoietic genes compared to the opposite inflection in mncfDNA fragments (Figure 12B). Creative attempts to study the dimension of cfDNA fragmentation patterns to infer gene expression can potentially be applied with uscfDNA. It appears this correlation with expression is most pronounced with the 0% CpG methylation subpopulation of fragments of uscfDNA alongside the 75< to 100% fragments of mncfDNA (Figure 12B and 12E). Target enrichment of the TSS region may allow for single-gene expression resolution and another strategy for deconvoluting micro signals within the noise of global activity in cell-free DNA.
[0165] Examining the common CpG regions for DMRs, there were regions where uscfDNA showed higher levels of CpG methylation than mncfDNA. However, the majority of significantly different DMRs were from regions of decreased methylation in uscfDNA, reflecting the global trend of the two circulating DNA populations. Additionally, most DMRs were near TSSs, suggesting potential differences in gene regulation related to uscfDNA and mncfDNA sequences (Figure 15).
[0166] The deconvolution attempt predicted that the mncfDNA derived from an assortment of blood cells that matched with expected cell type levels in the blood (Razavi et al., 2019, Nat. Med., 25, 1928–1937) and other prior cell-free DNA studies that used methylation DMRs for deconvolution (Moss et al., 2018, Nat. Commun., 9, 5068; Guo et al., 2017, Nat. Genet., 49, 635–642) (Figure 14B). The uscfDNA showed significant enrichment in eosinophils which are reported to exhibit efficient DNA repair machinery for both double-strand and single- strand breaks (Salati et al., 2007, Haematologica, 92, 1311–1318). One possibility is that because uscfDNA is enriched in simple repeats, which are predisposed to double-strand break damage (Gadgil et al., 2020, J. Biol. Chem., 295, 15378–15397), the efficient repair process in blood cells (such as eosinophils) might lead to the generation of circulating uscfDNA by-products. Eosinophils are also reported to release DNA-based extracellular traps into circulation, which is another potential source of uscfDNA (Mukherjee et al., 2018, Front. Immunol., 9, 2763; Aoki et al., 2021, Allergol. Int., 70, 3–8).
[0167] Various CpG-related cfDNA characteristics could be useful biomarkers to differentiate between non-cancer and NSCLC samples. When CpG methylation ratios for each size fragment were considered, the NSCLC samples appeared more hypermethylated in size bins <140 bases (Figure 16). This observation contrasted with genome-wide hypomethylation typically observed in cancer cells compared to healthy cells (Jones et al., 2002, Nat. Rev. Genet., 3, 415–428). However, the regions covered by cfDNA, particularly uscfDNA, do not faithfully represent the genome in its entirety, as uscfDNA appears to be enriched in regulatory regions (Cheng et al., 2022, iScience, 25, 104554; Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121; Hisano et al., 2021, BMC Biol., 19, 225; Cheng et al., 2022, iScience, 25, 105046). Hypomethylation of transcriptionally active regions seems to occur less frequently in lung cancer (Rauch et al., 2008, Proc. Natl. Acad. Sci. U. S. A., 105, 252–257; Hoffmann et al., 2005, Biochem. Cell Biol. Biochim. Biol. Cell., 83, 296–321; Pfeifer et al., 2009, Semin.Cancer Biol., 19, 181–187). Additionally, cfDNA is composed of DNA predominantly from blood cells more so than cancer tissue exclusively, which can explain the discrepancy.
[0168] It has previously been reported that cancer-specific promoters and CpG Islands may become hypermethylated (Harden et al., 2003, Clin. Cancer Res. Off. J. Am. Assoc. Cancer Res., 9, 1370–1375). In the sample set, both NSCLC uscfDNA and mncfDNA demonstrated substantial hypermethylation in the promoter, 5’UTR, CpG Islands, and exon elements compared to non-cancer subjects (Figure 17C and 17D and Figure 18).
[0169] In contrast to the other elements, which were either hypermethylated or variable, it was observed that the LINE and SINE elements of NSCLC subjects trended toward a hypomethylated state. In the genome, LINE and SINE elements have been described to undergo hypomethylation in cancer (Rauch et al., 2008, Proc. Natl. Acad. Sci. U. S. A., 105, 252–257). These high variability traces may indicate micro instability in the epigenetic regulation of these elements. The greater separation in mncfDNA may be due to the greater contribution of tumor- derived fragments, which have been shown to be enriched at 90-150 bases (Mouliere et al., 2018, Sci. Transl. Med., 10). It is unclear if the changes in methylation patterns originate from an increasing load of tumor-cfDNA or adjustments in activity from the immune system.
[0170] The limited number of DMRs for uscfDNA resulted from the overlap between the two uscfDNA fractions. Regardless, the data shows that both uscfDNA and mncfDNA bins could be a valuable source of DMR candidates (Figure 14E). The Increased expression of PK3P is associated with various types of cancer, including colon, lung, and bladder cancer (Ruan et al., 2021, J. Healthc. Eng., 2021, 9391104; Furukawa et al., 2005, Cancer Res., 65, 7102–7110). CPLX1 is one of several factors that are able to influence the activity of cyclin B1 (CCNB1), which is highly expressed in lung adenocarcinoma and associated with poor prognosis (Li et al., 2022, Oncol. Lett., 24, 441). CPLX1 has been documented to promote malignancy in gastric cancer (Tanaka et al., 2022, J. Gastroenterol., 57, 640–653). The expression of COL26A1 has been observed to be downregulated in subjects who respond well to PD-L1 inhibitors in transformed small-cell lung carcinoma. For mncfDNA candidates, mutations in ZNF595 have been indicated as a potential germline mutation in familial lung cancer (Kanwal et al., 2018, Gene, 641, 94–104) and region for prevalent somatic mutations in gastric cancer (Cui et al., 2015, Int. J. Cancer, 137, 86–95). The non-pseudo gene version of MLLT10 has been documented to be a promoter of tumor cell proliferation, migration, and invasion in NSCLC celllines (Tian et al., 2020, Cancer Manag. Res., 12, 5749–5758) and MLLT10P1 is commonly mutated in breast cancer subjects (Pongor et al., 2015, Genome Med., 7, 104). NERUDO2 is hypermethylated in adenocarcinoma in situ tissues, contrasting the hypomethylation seen in this study (Selamat et al., 2011, PLOS ONE, 6, e21443). Despite the potential biological rationale discussed, these DMRs are not currently validated. However, this approach shows the merit of DMR discovery, which could give rise to useful targets for future cancer detection.
[0171] Surprisingly, the deconvolution prediction did not indicate a signal from lung cell tissues despite the samples coming from NSCLC. Despite late-stage cases, most cfDNA is still from blood cell origin (Razavi et al., 2019, Nat. Med., 25, 1928–1937). For uscfDNA, the starkest change was a decrease in eosinophils % and a trend in increased neutrophils. Increased eosinophils have been associated with improved prognosis in lung cancer (Costello et al., 2005, Rev. Med. Interne, 26, 479–484; Davis et al., 2014, Cancer Immunol. Res., 2, 1–8). In the mncfDNA, the megakaryocytes were increased, which has also been described to be associated with cancer (Huang et al., 2015, J. Clin. Oncol. Off. J. Am. Soc. Clin. Oncol., 33, 836–845; Soares et al., 1992, J. Clin. Pathol., 45, 140–142; Dejima et al., 2018, Lung Cancer Amst. Neth., 125, 128–135) in the literature.
[0172] Using whole-genome sequencing, other investigators have reported that uscfDNA predicted to contain G-Quad secondary structures is decreased in cancer subjects (Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121). This pattern was also observed in NSCLC samples that have undergone bisulfite conversion (Figure 12E). Interestingly, in NSCLC subjects, fragments that contained potential G-Quad structures showed increased CpG methylation levels compared to non-cancer subjects (Figure 19A and 19B). Within the genome, G-Quad has been described to regulate methylation behavior at CpG Islands (Mao et al., 2018, Nat. Struct. Mol. Biol., 25, 951–957; Mukherjee et al., 2019, Trends Genet. TIG, 35, 129–144). It is possible that although there is a decrease in G-Quad structures present in the plasma, it reflects changes in altered CpG methylation and subsequent changes in transcription factors or chromosomal inaccessibility.
[0173] Mutations in chromatin-bound proteins frequently occur in cancer (Shen et al., 2013, ell, 153, 38–55). It was observed that %intersection with epigenetic marks was also altered in NSCLC subjects with the greatest decreases in %intersection of H3K27ac and H3K4me3 for both uscfDNA and mncfDNA (Figure 19C). As these two marks are associatedwith genes with high expression, their decrease in the NSCLC samples seen in this study may be suggestive of dysregulation in cancer and a potential viable global indicator.
[0174] Therefore, the 5mCAdpBS-Seq single-stranded DNA library preparation is advantageous for uscfDNA methylation profile investigation due to the preservation of the native fragment length and methylation level in each size bin. Using this protocol, the methylation characteristics of uscfDNA appear distinctly different than mncfDNA, further illustrating that it should be considered a separate cfDNA molecule. As a methylation-based cancer biomarker, potentially useful features of uscfDNA are global CpG% methylation changes, genome element profiles, CpG-methylation traces for specific elements, DMRs, tissue-of-origin deconvolution, G-Quad signature changes, and epigenetic mark association. Although the data presented herein have focused on cfDNA from plasma, the 5mCAdpBS-Seq protocol is useful for any contexts where very short DNA templates are present. This can include analysis of other biofluids with fragmented DNA (saliva and urine (Brooks et al., 2023, PLOS ONE, 18, e0285214; Chen et al., 2022, PLoS Genet., 18, e1010262)), cell-culture conditioned media environments (Silver et al., 2023, eLife, 12, e83532), or theoretically in any in-vitro intracellular work where the accurate methylation analysis of short single-stranded DNA is required. Therefore, if investigators are interested in examining the methylation profile of a DNA sample with heterogeneous sizes, the 5mCAdpBS-Seq protocol should be considered. The Materials and Methods are now described. Subjects In Study Table 2 Paired BS-Seq and 5mCAdpBS-Seq Number Lot Number Age Sex 1 666 38 M 2 668 52 M 3 681 18 M 4 698 26 F 5 700 35 F NSCLC Samples Number Lot Number Age Stage Sex 1 120E 47 3A F 2 147E 75 4 F 3 161E 79 3B F4 231E 62 3A M
[0175] Plasma from healthy donors (Table 2) was commercially purchased from Innovative Research (IPLASK2E10ML) in K2EDTA tubes. According to vendor instructions, whole blood was spun at 5000xG for 15 minutes and plasma was removed using a plasma extractor.
[0176] Source of the NSCLC Plasma Samples. Plasma from late-stage NSCLC patients (Table 2) were obtained. Biopsy specimens were examined histologically, and the presence of EGFR mutations were determined using the Therascreen EGFR RGQ PCR Kit (EGFR IVD Kit) (Syed, Y.Y. (2016) Mol. Diagn. Ther., 20, 191–198). The staging criteria used were those from the American Joint Committee on Cancer (AJCC) TNM system (Huang,S.H. et al., (2015) J. Clin. Oncol. Off. J. Am. Soc. Clin. Oncol., 33, 836–845). Lambda DNA Control Restriction Enzyme Reactions
[0177] For all reactions 1.5µL (1µg) of unmethylated lambda (Promega, D1521) was used. After the restriction enzyme reaction, the DNA was purified by combining 20µL of reaction mixture, 60µL of SPRI-select beads, and 60µL of 100% isopropanol and incubated for 10 minutes. The tube was placed on a magnetic rack for five minutes to allow the beads to migrate. The supernatant was discarded, and the beads were washed twice with 200µL of 80% ethanol. Once the second ethanol wash was removed, the beads were left to air dry for 10 minutes. The beads were then resuspended in 20µL of Qiagen elution buffer.
[0178] CviKL Restriction Enzyme: 1.5µL of Lambda DNA was combined with 2µl (10x) rCutSmart Buffer, 1µl CviKL enzyme (NEB, R0710S) and 15.5µl H2O. The mixture was heated to 25ºC for 60 minutes and then the enzyme deactivated by heating it to 65ºC for 20 minutes.
[0179] NlaIII Restriction Enzyme: 1.5µL of Lambda DNA was combined with 2µl (10x) rCutSmart Buffer, 1µL Nlalll (NEB, R0125S), and 15.5µl H2O. The mixture was heated to 37ºC for 15 minutes and then the enzyme deactivated by heating it to 65ºC for 20 minutes.
[0180] AluI Restriction Enzyme: 1.5µL of Lambda DNA was combined with 2µl (10x) rCutSmart Buffer, Alul (NEB, R0137S), and 15.5µl H2O. The mixture was heated to 37ºC for 60 minutes and then the enzyme deactivated by heating it to 65ºC for 20 minutes.Nucleic Acid Extraction.
[0181] cfDNA was extracted from plasma using the QIAmp Circulating Nucleic Acid Kit (Qiagen, cat# 55114) following the manufacturer protocol “Purification of Circulating microRNA from 1ml of Plasma” (QiaM). For BRcfDNA-Seq (Cheng et al., 2022, iScience, 25, 104554), cfDNA was extracted from 1 mL of plasma, whereas, for the methylation pipeline, cfDNA was extracted from 2 mL of plasma. Proteinase-K digestion was carried out as instructed. For BRcfDNA-Seq, BS-Seq, and 5mCAdpBS-Seq, carrier RNA was not used, and the ATL Lysis buffer (Qiagen, 19076) was used as indicated in the microRNA protocol. The final elution volume for all protocols was 20µl. Broad Range Cell-free DNA Seq (BRcfDNA-Seq) Library Preparation.
[0182] Single-stranded DNA library preparation was performed using the SRSLYTM PicoPlus DNA NGS Library Preparation Base Kit with the SRSLY 12 UMI-UDI Primer Set, UMI Add-on Reagents, and purified with Clarefy Purification Beads using the low molecular weight protocol. (Claret Bioscience, cat #CBS-K250B-24, CBS-UM-24, CBS-UR-24, CBS-BD- 24). Since there is currently no optimized method to measure uscfDNA specifically, 18µL of extracted cfDNA was used as input and heat-denatured as instructed. In experiments including digested lambda DNA, 50pg was added to the library preparation as spike-in. The index PCR was performed as specified in the manual for 11 cycles. All bead clean-up steps followed the low molecular weight retention purification protocol (Troll et al., 2019, BMC Genomics, 20, 1023). Specifically, after adapter ligation, the 50ul reaction was combined with 48ul of water, 12uL of 100% isopropanol, and 65.2ul of Clarefy beads. After the UMI extension, the 40ul reaction was combined with 80ul of Clarefy beads. After the index PCR, the 50ul reaction was combined with 75ul of Clarefy beads. 5mCAdpBS-Seq Library Preparation.
[0183] The first step of the single-stranded library preparation (premethylated single- stranded adapter ligation) was performed on extracted cfDNA prior to bisulfite conversion. Custom 5mC-protected SRSLY adapters were provided by Claret Bioscience and are identical tothose found in the regular SRSLY kit, with the exception that all cytosine residues on the adapter strands of the duplexed splint adapters are pre-methylated (5mC). The adapter sequences are as follows 5’-Adapter (5’ -A5mCA 5mCT5mC TTT 5mC5mC5mC TA5mC A5mCG A5mCG 5mCT5mC TT5mC 5mCGA T5mCT – 3’) (SEQ ID NO: 1) and 3’Adpater (5’ - AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC - 3’) (SEQ ID NO: 2). used in place of the regular adapters in the adapter ligation step and, after bead clean-up, resuspended to 20µL. Then 20µl of adapter-ligated DNA underwent bisulfite conversion using Zymo Research DNA Methylation Lightning kit (Zymogen, cat# D5030) with an elution volume of 15µL into the UMI-UDI step of the single-strand library preparation protocol. The remaining steps (Addition of UMI by Primer extension and Index PCR) were performed as described in the “BRcfDNA- Seq Library Preparation” section. During the final index PCR, the Index PCR Master Mix was substituted with the Kapa HIFI HotStart Uracil+ ReadyMix. The PCR protocol is as follows: 98oC for 3 minutes, [98°C for 30 seconds, 60°C for 30 seconds, 72°C for 1:00] for 11 cycles, 72°C for 1 minute then hold at 12°C. All bead clean-up steps followed the low molecular weight retention purification protocol (Troll et al., 2019, BMC Genomics, 20, 1023). Specifically, after adapter ligation, the 50µl reaction was combined with 48ul of water, 12µL of 100% isopropanol, and 65.2µl of Clarefy beads. After the UMI extension, the 40µl reaction was combined with 80µl of Clarefy beads. After the index PCR, the 50µl reaction was combined with 75µl of Clarefy beads.
[0184] Final Library Concentration and Quality Control. Library concentrations were measured using the Qubit Fluorometer (ThermoFisher Scientific, cat# Q33327), and quality was assessed using the Tapestation 4200 using D1000 High-Sensitivity Assay (Agilent Technologies, cat# 5067-5584). Samples were pooled to a final molarity of 5nM.
[0185] Sequencing. Pooled libraries were sequenced 150bp x 2 on NovaSeq6000 either on an SP or an S1 flow cell, aiming at 40 million reads per sample. Data Analysis.
[0186] Paired reads were merged into single-end reads using BBMerge (Bushnell et al., 2017, PLOS ONE, 12, e0185056), in order to obtain one .fastq file per sample. Each .fastq file was trimmed with fastp using the adapter sequence AGATCGGAAGAGCACACGTCTGAACTCCAGTCA (r1; SEQ ID NO: 3) andAGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (r2; SEQ ID NO: 4) and filtered for a Phred score of >15 (Chen et al., 2018, Bioinforma. Oxf. Engl., 34, i884–i890). Standard (unconverted) and BS-treated libraries were aligned against the combined human [GenBank:GCA_000001305.2, GRCh38] and lambda phage [GenBank: J02459.1] reference genomes using BWA-mem (Li et al., 2009, Bioinforma. Oxf. Engl., 25, 1754–1760) and BSBolt’s default setting (Farrell et al., 2021, GigaScience, 10, giab033), respectively. Sequence reads were demultiplexed using the SRSLYumi python package (SRSLYumi 0.4 version, Claret Bioscience). The duplicated reads were removed using Picard Toolkit (broadinstitute.github.io / picard / ) with VALIDATION_STRINGENCY set as LENIENT and REMOVE_DUPLICATES as TRUE. Soft and hard clipped reads with removed samtools (samtools 1.9 version). Quality control was performed with Qualimap Version 2.2.2c (García- Alcalde et al., 2012, Bioinforma. Oxf. Engl., 28, 2678–2679). Samtools view was used to isolate reads from mitochondrial DNA. Each of the bam files aligning to human, mitochondria, and lambda phage was binned in increments of 10 bases from 20 to 200 using alignmentSieve (deepTools 3.5.0) (Ramírez et al., 2016, Nucleic Acids Res., 44, W160-165). CpG and Non-CpG% Methylation.
[0187] BSBolt CallMethylation was used on of each 10-base bin to determine the %CG methylation and %CHH methylation. MapQ scores were calculated from each size bin by Qualimap (version 2.2.2c) (García-Alcalde et al., 2012, Bioinforma. Oxf. Engl., 28, 2678–2679). G-Quad signatures.
[0188] G-Quad percentage was calculated by first converting binned .bam files to .bed using bamtobed (bedtools ver 2.18) and then from .bed to .fasta using getfasta (bedtools ver 2.18)(Quinlan and Hall 2010). fastaRegexFinder.py was used to analyze the sequences in the reads (github.com / dariober / bioinformatics-cafe / tree / master / fastaRegexFinder). In general, this python pipeline examines if the sequences contain this pattern in this equation “([gG]{3,}\w{1,7}){3,}[gG]{3,}”. This translates to identifying three or more G nucleotides followed by 1 to 7 of any other bases and must be repeated three or more times and ending with three or more Gs. The G-Quad counts were divided by the total read counts to identify the G-Quad percentage. A normalized ratio was calculated by dividing the read counts by the median value of the bin length (e.g., 20-29 is 25). Only primary fragments that contained G-Quad sequences were counted (e.g., complementary sequences that contained G-Quads were excluded). The coordinates of the G-Quad sequences were used to generate G-Quad Only bam files for CpG methylation % calling. Linear Correlation Analysis.
[0189] The .bam files were split into genomic bins of 100k bases along the genome (e.g., Chr1:1-100,000 for two in-silico categories: uscfDNA (40-70 bases) and mncfDNA (120-250 bases). The % of total coverage was calculated for each bin and then the linear correlation was calculated by comparing the signal from at each genomic bin position with Graphpad Prism 9. Both samples processed with BS-Seq and 5mCAdpBS-Seq was compared to the paired equivalent BRcfDNA-Seq sample. Genome Wide Ideogram.
[0190] Bed files containing the location of each CpG Site were split into genomic bins of 1 million bases along the genome (e.g., Chr1:1-1,000,000) for two in-silico categories: uscfDNA (40-70 bases) and mncfDNA (120-250 bases). Karyograms were self-normalized so that the legend reflects the intrasample dynamic range. Ideograms were constructed from the average of 5 samples showing the CpG site frequency for each 1 million bases bin using the rideogram R package (Hao et al., 2020, PeerJ Comput. Sci., 6, e251). CpG Intersection Positions with Genomic Elements.
[0191] Analyzed by converting CG.map files from 5mCAdpBS-Seq libraries processed by BSBolt into .bed files. The converted .bed files were then intersected using bedtools (version 2.30.0) intersect (Quinlan et al., 2010, Bioinformatics, 26, 841–842) with .bed for eleven genomic elements (SINEs, LINEs, Simple Repeats, Exons, Introns, Intergenic, Promoters, CpG Islands, 5’ Untranslated Region, 3’ Untranslated Region, and Transcription Termination Site (TTS). The intersection counts were used to calculate the % of CpG-containing fragments. Notethat certain elements were potentially counted in duplicate due to overlap in regions (e.g., Promoters and CpG-Islands). Different CpG Methylation Fragment Overlap with Genomic Elements.
[0192] Fragments were binned by their CpG methylation status based on the SAM flag status into four categories (0%, 0%< to 25%, 25< to 75%, and 75< to 100%). The different bins were then intersected using bedtools (version 2.30.0) intersect (Quinlan et al., 2010, Bioinformatics, 26, 841–842) with the -wo argument and the .bed for elements related to genes (CpG Shelf, CpG Shore, CpG Island, promoter, 5’ Untranslated Region, first exon, all exons, introns, 3’ Untranslated Region, Transcription Termination Site (TTS) and intergenic regions. Epigenetic Marks.
[0193] Epigenetic marks were calculated using the bedtools intersect function with the - wo argument (version 2.30.0) (Quinlan et al., 2010, Bioinformatics, 26, 841–842). Intersected base counts were divided by the total base counts of the bed file. Control bed files were generated using bedtools shuffle with human reference [GenBank:GCA_000001305.2, GRCh38] and by sourcing size fragments count and distribution from the uscfDNA or mncfDNA bed file from each respective subject. The observed ratio was calculated by dividing the % of intersecting bases of the .bed file by the % intersection bases of randomly shuffled control bed files for each epigenetic mark for uscfDNA and mncfDNA bins. Experiment reference files were retrieved from the BLUEPRINT Data Analysis Portal (Fernández et al., 2016, Cell Syst., 3, 491-495.e5). Specific subjects used were the following: eosinophil (S006XE53 and S006XEH2), macrophage (C005VG51 and C005VGH1), monocyte (C000S5A1b and C000S5H2), and neutrophil (C0010KA1bs and C0010KH2). The % intersected base pairs were normalized against control shuffled bed files to compare non-cancer and NSCLC samples. Enrichment of Different CpG Methylation % Bins Along TSS.
[0194] The pattern of enrichment of CpG fragments -1000 bases upstream and +1000 bases downstream from the transcription start site differ amongst uscfDNA and mncfDNA fragments and correlate to gene activity. Plots were generated using SeqMonk (version 1.48.1,bioinformatics.babraham.ac.uk / projects / seqmonk / ). Bed files from different methylation bins (0%, 0%< to 24%, 25 to 74%, and 75-100%) were loaded, and probes were defined using “Feature Probe Generator” and designed around .bed files of all TSS or different sets of TSS of genes of different expression levels. Probes were designed over features from -1000 bp to +1000 bp, quantified for their enrichment, and plotted using probe trend plot with “Force plot to be relative” (Figure 5) or “scale within each data store” (Figure 8). Sets of TSSs were categorized according RNA-Seq data of gene the PBMC from the buffy coat as described (Esfahani et al., 2022, Nat. Biotechnol., 40, 585–597). High expression was considered (>41.07 RPKM), medium (15.36-41.06 RPKM), low (1-15.36 RPKM), and silent (<0 RPKM). The list of TSSs were taken from Homer hg38.bed files (Heinz et al., 2010, Mol. Cell, 38, 576–589). CpG Methylation % Quantification Trend Plots.
[0195] Plots were generated using SeqMonk (version 1.48.1, bioinformatics.babraham.ac.uk / projects / seqmonk / ). The .CGmap.gz files generated by BSBolt were converted to cg.bismark.cov.gz files. The files were imported, and SeqMonk in-silico probes were defined using Feature Probe Generator for the different genomic elements of interest. Probes were defined as over features from -5000 bases to +5000 bases. Probes (or window intervals) were then defined using the “Running Window Generator.” Probe size | step as 100 and 100 over the active probes. After, define quantification was selected with “features to quantitate” as existing probes, minimum count to include position as 1, minimum observations to include feature was set to 1, and combined value to report as mean. Next, quantification trend plots were constructed for the genomic elements. The option to remove exact duplicates was checked, and probes were made from -5000 bases to +5000 bases upstream and downstream from the body of the feature. Differentially Methylated Regions.
[0196] Samples were aggregated using metilene_input.pl from metilene package (Jühling et al., 2016, Genome Res., 26, 256–262) using a minimum coverage of 1. DMRs between samples were identified using metilene (Jühling et al., 2016, Genome Res., 26, 256–262) using the settings --mincpgs 3 –maxdist 100 –minMethDiff 0.1 –valley 0.7. Closest gene was analyzedusing bedtools closest with default settings with a hg38 gene reference from UCSC RefSeq (refgene) from genome.ucsc.edu / cgi-bin / hgTables.
[0197] Deconvolution of Tissue-of-Origin. Samples were analyzed using CelFiE (CEL Free DNA decomposition Expectation maximization) algorithm with default parameters as described (Caggiano et al., 2021, Nat. Commun., 12, 2717). Statistical Analysis
[0198] For genomic element profiles and epigenetic mark observed ratios, Tukey’s multiple comparison test was performed after 2way ANOVA. Individual student t-tests were performed for the different tissue types for the % predicated tissue deconvolution and for CpG methylation of G-Quad containing fragments. Error bars represent SEM from the average of 5 non-cancer and 4 NSCLC subjects. Stars indicate p-values with * p <0.05, ** p <0.01 , *** p< 0.001. The Experimental Results are now described. Merging paired-end reads prior to alignment impacts fragment length profile of BS-treated cfDNA libraries.
[0199] Depending on the length of the cfDNA fragment, paired-end sequencing cycles may only report a fraction of the DNA sequence (Figure 1A), whereas certain circumstances lead to fragments being excluded from further processing (Figure 1B). A paired-end mapping pipeline was compared to merging paired reads prior to alignment (Figure 1C-1F). For BS-Seq libraries, in the mncfDNA region, two distinct peaks were present (150 bases and 167 bases) (Figure 1D). Conversely, the BS-Seq libraries showed a different pattern when merged pre-alignment compared to the standard paired-end processing.
[0200] The BRcfDNA-Seq libraries had comparable MAPQ scores in the two analysis modes, with the merged-reads protocol being slightly lower (Supplementary Figures 1E and F). For the BS-Seq, both default and merged processing had a lower MAPQ score for bins from 30- 39 bases but stabilized for the bins >40 bases. The merged reads were slightly lower compared to the paired-end processing.
[0201] The percentage of total reads for different workflows (BRcfDNA-Seq or BS-Seq) were compared along the bioinformatic pipeline (Figures 2A and 2D). A proportion of reads (72.6% ± 3.2%) were kept after the initial merging step (Figure 2B and 2E). When compared to unmerged analysis, merged processing universally resulted in lower final usable reads (51.9 ± 4.7% (Unmerged) vs 46.6 ± 3.6% (Merged)) (Figure 2C and 2F). However, subsequent steps after merging (quality control and alignment) retained more reads compared to the paired-end pipeline. All sequenced samples were then processed with the merged pipeline in addition to alleviating the artificial double-peak generated in the mononucleosomal region (Figure 1D). The 5mCAdpBS-Seq protocol reduces the inclusion of degraded DNA into the uscfDNA region of the final library.
[0202] Since the initial BS-Seq experiments indicated differences in the size-distribution compared to non-BS BRcfDNA-Seq (Figure 1C and 1D), without being bound by theory, it was hypothesized that the increased representation of fragments in the 70-130 bases region derived from larger-sized cell-free DNA degradation during the bisulfite treatment process (Figure 3A) (Tanaka et al., 2007, Bioorg. Med. Chem. Lett., 17, 1912–1915). To this end, it was tested if ligating single-stranded 5mC-protected adapters prior to bisulfite treatment (5mCAdpBS-Seq) would reduce the incorporation of degraded DNA, which could be misclassified as uscfDNA (Figure 3B).
[0203] The in silico read loss was assessed between the BS-Seq and 5mCAdpBS-Seq protocols. Compared to the BS-Seq, the 5mCAdpBS-Seq protocol showed a greater read loss in most processing steps of the protocol (Merging, Quality Control, and Alignment) (67.8 ± 3.4 % for BS-Seq vs 54.6 ± 5.3% for 5mCAdpBS-Seq reads remaining). However, after the removal of PCR-duplicated reads (deduplication based on UMI – unique molecular identifiers), the remaining reads between both protocols were comparable (46.6 ± 3.6% for BS-Seq vs 45 ± 4.7% for 5mCAdpBS-Seq reads remaining) (Figure 4A). In some cases, the % of remaining reads for the 5mCAdpBS-Seq protocol was even higher than the BS-Seq protocol (Figure 4B).
[0204] Reads obtained from the 5mCAdpBS-Seq protocol after aligning to nuclear DNA showed substantial differences in the fragment profile compared to the BS-Seq while closely resembling the BRcfDNA-Seq protocol (Figure 3C). In particular, the region from 70 to 130bases is largely absent from the DNA degradation seen in the BS-Seq protocol profiles (Figure 3C, Figure 1D).
[0205] The % total genomic coverage at each 100K base bin along the genome was compared between the three protocols (Figure 5 A, B, D, and E). It revealed for uscfDNA fragments, on average, the 5mCAdpBS-Seq protocol showed a larger R2 coefficient compared to BS-Seq, demonstrating a closer similarity to BRcfDNA-Seq (Figure 5C) but no trend for mncfDNA (Figure 5F).
[0206] Alongside nuclear DNA, the mitochondrial genome (mitDNA) also contributes to the pool of cell-free DNA in circulation (An et al., 2019, Precis. Clin. Med., 2, 131–139). However, the reads that align to mitDNA only represent a minor fraction of total sequence reads, averaging 0.35 ± 0.006%, 0.0135 ± 0.002%, and 0.0664 ± 0.012% for BRcfDNA-Seq, BS-Seq, and 5mCAdpBS-Seq, respectively. The reads aligned to the mitDNA of 5mCAdpBS-Seq closely resembled the BRcfDNA-Seq fragment distribution pattern with a slight peak shift to the left (Figure 3D). By comparison, the BS-Seq mitDNA profile had a peak at 57 bases, with most fragments occupying between 40 to 75 bases. Compared to the nuclear uscfDNA, the mitDNA fragment curve is not symmetrical, with a more prominent shoulder greater than 60 bases. Genomic Characteristics of the 5mCAdpBS-Seq protocol closely resemble BRcfDNA-Seq
[0207] The pattern of CpG density was examined at each bin of fragment size and observed that the 5mCAdpBS-Seq protocol reads followed the same pattern as those for BRcfDNA-Seq (Figure 3E), whereas the BS-Seq protocol had a lower peak at 50nt but an elevated CpG density from 70-130 bases. This pattern resembled that of the fragment size distribution, in line with the hypothesis that bisulfite treatment leads to fragmented genomic and mncfDNA contributing to the uscfDNA region (Figure 3C).
[0208] Previous reports have shown that the uscfDNA is enriched in G-rich sequences that can potentially form G-Quad secondary structures (Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121). G-Quad signatures were enriched in the BRcfDNA-Seq and 5mCAdpBS-Seq protocol, with the highest peak in the 40-49 bases bin (Figure 3F). In the same samples processed with BS-Seq, however, this enrichment was absent in the ultrashort region.BRcfDNA-Seq and 5mCAdpBS-Seq protocols show that uscfDNA maps to regions associated with active genes compared to mncfDNA
[0209] The intersection profiles of presenting CpG site and cfDNA fragments (Figure 6A) with genomic elements and epigenetic parks profiles (Figure 6B) of all three protocols were compared for both uscfDNA and mncfDNA populations (Figure 6C and 6D). For both cfDNA populations, the 5mCAdpBS-Seq protocol more accurately recapitulated the genomic element profile compared to BS-Seq, for SINE elements, promoters, exon, and CpG Islands. Additionally, uscfDNA fragments appear significantly enriched in promoters, exons, and CpG island locations, whereas mncfDNA is more enriched in SINE and intergenic regions.
[0210] Since uscfDNA appears enriched in promoters (Figure 6C) (Cheng et al., 2022, iScience, 25, 104554; Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121; Hisano et al., 2021, BMC Biol., 19, 225; Cheng et al., 2022, iScience, 25, 105046), it was examined if the 5mCAdpBS-Seq protocol resembled BRcfDNA-Seq in genomic regions associated with increased gene activity (Figure 6B) (Zhang et al., 2015, EMBO Rep., 16, 1467–1481). It was observed that the mapping patterns for epigenetic marks of the 5mCAdpBS-Seq protocol closely reflected the patterns of BRcfDNA-Seq compared to BS-Seq (Figure 6E and Figure 7) for both uscfDNA and mncfDNA fragments. uscfDNA demonstrated a higher intersection % (versus a matched shuffled position control) with active gene epigenetic marks (H3K4m1, H3K4m3, H3K27ac modifications, and the hypomethylated regions), whereas mncfDNA showed the opposite trend. In contrast, for H3K27me, both uscfDNA and mncfDNA showed an increased intersection % in H3K27me, H3K9me3, and hypermethylated regions.
[0211] Based on these findings of apparent similarity between 5mCAdpBS-Seq and BRcfDNA-Seq, all the subsequent analyses were performed on the data obtained with the premethylated adapter protocol. The 5mCAdpBS-Seq protocol portrays that nuclear uscfDNA fragments are globally hypomethylated
[0212] UscfDNA fragments had lower mean CpG methylation% in the 5mCAdpBS-Seq profiles compared to the BS-Seq protocol (63.6-64.6% vs 76.8-77.1%) (Figure 8A). By contrast, mncfDNA fragments (120-200 bases) had a similar CpG methylation % between both protocols (80.2-80.9% vs 80.5-82.5%) and were higher than the uscfDNA population. The nuclear non-CpG methylation in both protocols was approximately 1% (Figure 8D and Supplementary Figure 9A).
[0213] Reads aligning to the mitochondrial genome (Figure 8B and Figure 9B) were used as a biological control since the mitochondria genome is expected to be hypomethylated (Liu et al., 2016, Sci. Rep., 6, 23421; Mechta et al., 2017, Front. Genet., 8, 166). For both protocols, the CpG and the non-CpG methylation levels were below 5%, with fluctuations for bins >130 bases (Figures 8B and 8E and Figure 9B and 9C). However, there were few reads beyond 130 bases.
[0214] As a negative control, enzymatically sheared Lambda phage DNA was spiked into plasma samples undergoing both the BS-Seq and 5mCAdpBS-Seq protocol to determine the bisulfite conversion efficiency (Figure 8C, 8F and Figure 9D and 9E). The mean CpG methylation % for both the BS-Seq and 5mCAdpBS-Seq protocols was similar and <1% for CG% and <1.5% for non-CpG% methylation, albeit a little higher for the latter. CpG Methylation levels in uscfDNA are lower than mncfDNA, with differing patterns for genomic elements
[0215] Using the 5mCAdpBS-Seq protocol, karyograms were constructed showing differences in % coverage for CpG sites of uscfDNA and mncfDNA (Figure 10A and 10B). Of the uscfDNA CpG site positions, 41.4 ± 5% of uscfDNA sites could be found within the mncfDNA population, but most sites were unique to mncfDNA (Figure 10C).
[0216] Both uscfDNA and mncfDNA fragments were subdivided into four CpG methylation categories (0%, 0< to 25%, 25< to 75%, and 75< to 100%) (Figure 11A). Recapitulating the global CpG methylation, uscfDNA fragments had a greater proportion of 0% methylation and a subsequent lower proportion of 75< to 100% CpG methylation fragments compared to mncfDNA. When intersected against gene regulatory elements, a larger proportion of uscfDNA fragments of all methylation statuses converged around CpG elements and promoters (Figure 11B). Interestingly, despite contributing a <0.5% of total fragments, the 0< to 25% CpG methylation fragments were enriched in CpG elements (island, shore, or shelf).
[0217] The behavior of various genomic elements of interest was examined by plotting the CpG methylation % from 5000 bases upstream from the center of the element to 5000 bases downstream from the body of the element (Figure 11C). In general, the CpG methylation % of uscfDNA fragments was 10-20% lower than that for mncfDNA over the same regions, mirroringthe genome-wide CpG methylation state (Figure 8A). The general patterns of the CpG% methylation distribution were similar, although uscfDNA had more variance, most likely caused by reduced coverage. The three most distinct methylation patterns between the two cfDNA populations were those for Simple Repeats, LINE, Intergenic, and Exons. uscfDNA fragments are enriched directly upstream to the transcription start sites and reflect gene expression activity
[0218] It has previously been shown that while mncfDNA fragments demonstrate a decreased coverage over transcript start sites (TSSs), ultrashort cfDNA fragments show an enrichment (Hudecova et al., 2021, Genome Res., 10.1101 / gr.275691.121; Snyder et al., 2016, Cell, 164, 57–68). When examining only CpG-containing fragments, a similar pattern was observed for uscfDNA fragments showing an upward inflection (Figure 12A). The pattern of uscfDNA and mncfDNA fragments binned by 0% or 75< to 100% CpG Methylation amongst TSSs grouped by hemopoietic cell gene expression (Figure 12B to 12E). The enrichment of 0% CpG methylated uscfDNA towards the TSS was positively correlated with highly expressed genes (Figure 12B), and the 75< to 100% CpG methylated fragments were negatively correlated. In contrast, the 0% CpG methylated mncfDNA showed more pronounced depression towards the TSSs of genes with high expression (Figure 12D). In general, the profile of fragment enrichment for TSSs of genes of low expression was more horizontally stable along the TSS regardless of CpG methylation state (Figure 12B to 12E). These patterns were less pronounced in the 0< to 25% and 25< to 75% methylated fragments (Figure 13). Differentially methylated regions exist between uscfDNA and mncfDNA
[0219] Since there was a minor overlap between uscfDNA and mncfDNA CpG sites, the .bam files were aggregated from the uscfDNA and mncfDNA from the five subjects and analyzed regions that were differentially methylated (DMR) between the two cfDNA populations (Figure 14A). Sixty-eight significant DMRs were found, where the majority were hypomethylated in uscfDNA compared to mncfDNA. Additionally, the majority of DMRs were in close vicinity to TSSs (Figure 15).Deconvolution suggests that uscfDNA mainly derives from peripheral blood cells.
[0220] Next, the fragments were deconvoluted from the uscfDNA and mncfDNA populations into their cell / tissue-of-origin using the CpG methylation patterns (Figure 14B). The CelFiE algorithm (Caggiano et al., 2021, Nat. Commun., 12, 2717) is constructed as an expectation-maximization algorithm, which iteratively finds the vector of tissue proportions with the greatest maximum likelihood using information from both the reference data supplied at the time of running the algorithm and the regions of the genome that are variable between tissues using whole genome bisulfite sequencing data obtained from ENCODE and Blueprint (Fernández et al., 2016, Cell Syst., 3, 491-495.e5 ; Dunham et al., 2012, Nature, 489, 57–74). Using this algorithm which was designed to deconvolute signal from low input cfDNA samples, it was confirmed that the major tissue of origin for both uscfDNA and mncfDNA is blood, as expected. This methylation tissue-of-origin analysis indicated uscfDNA contribution from eosinophils, erythroblasts, monocytes, neutrophils, and T-cells. Moreover, CelFiE indicated that uscfDNA has significantly more fragments originating from eosinophils compared to the mncfDNA. uscfDNA CpG mapping patterns and methylation characteristics can discriminate non-cancer subjects from late-stage NSCLC.
[0221] As a proof of concept, the methylation profile of uscfDNA was examined to determine if it would be an effective biomarker for cancer detection. To that end, four late-stage NSCLC samples were processed with the 5mCAdpBS-Seq protocol and compared with non- cancer samples. The global fragment patterns showed an elevated uscfDNA peak in the NSCLC samples and a lower rightward shoulder in the mncfDNA regions of 175 to 200 bases (Figure 16A). For reads mapping to the nuclear genome, in the 10-bases bins between 40 bases and 140 bases, it appeared that the NSCLC samples had higher levels of CpG% methylation (4-6%) compared to the non-cancer samples in sizes below 140 bases (Figure 16B).
[0222] 8 types of genomic regions were identified that were differentially represented by uscfDNA (Figure 17A). By contrast, only 4 regions showed differential representation in mncfDNA (Figure 17B). In the uscfDNA bin, there were significant changes in the proportion of SINE, Simple Repeats, promoters, introns, intergenic, 5’UTR, and CpG-Islands. In themncfDNA, however, promoters, exons, 5’UTR, and CpG-island proportion appeared statistically different. The methylation pattern of genomic elements is altered in NSCLC
[0223] When the enrichment was examined around TSS and CpG methylation % patterns for uscfDNA (Figure 17C and Figure 18A) and mncfDNA bins (Figure 17D and Figure 18B). The TSS was altered in the NSCLC samples for both 0% and 75< to 100% CpG methylated fragments. The NSCLC samples showed greater methylation variability in the cancer samples, whereas the non-cancer samples were more uniform. For promoters, 5’UTR, exons, and LINE elements, the methylation towards the body of the element was hypermethylated in NSCLC samples compared to the non-cancer, but this observation was more evident in the mncfDNA bins. For the LINE elements, the NSCLC samples appeared more hypomethylated compared to the non-cancer subjects, but this was more apparent in the mncfDNA bin (Figure 18A). By contrast, the flanking regions of SINE elements were hypomethylated in NSCLC, which was more prominent in the uscfDNA bin compared to the mncfDNA bin. Introns, 3’UTR, and TTS elements were globally more hypermethylated in mncfDNA. In the uscfDNA, the methylation profile of the non-cancer samples was more uniform compared to the highly variable NSCLC traces. Differentially Methylated Regions and Deconvolution Are Potential Biomarkers for NSCLC Detection
[0224] DMR analysis between the CpG% methylation of NSCLC and non-cancer samples revealed that both the uscfDNA and mncfDNA fragments had significant DMR candidates (Figure 17E and 17F). UscfDNA fragments had 12 significant DMRs out of 18160 tested regions (0.066% significant) compared to mncfDNA, which had had 302 significant DMRs out of 1223476 tested regions (0.025% significant) (Figure 17G). For both uscfDNA and mncfDNA, significant DMRs had lower methylation in NSCLC compared to non-cancer subjects (Figure 17E and 17F). Some examples for the uscfDNA DMR candidate nearest genes were plakophilin3 (PKP3), complexin 1 (CPLX1), and collagen type XXVI alpha 1 chain (COL26A1). For mncfDNA, the top candidates based on q value were zinc finger protein 595(ZNF595), myeloid / lymphoid or mixed-lineage leukemia translocated to pseudogene 1 (MLLT10P1), and neuronal differentiation 2 (NEUROD2).
[0225] The CelFiE deconvolution prediction algorithm suggested differences in the tissue of origin profiles between the two cohorts (Figure 17H and 17I). In the uscfDNA fragment bin, the eosinophil signal appeared significantly decreased in NSCLC samples, matching the level found in the mncfDNA fraction. Whereas in mncfDNA, there was an increase in megakaryocyte signal in some NSCLC samples (not significant). G-Quad containing uscfDNA fragments show an increased CpG methylation % in NSCLC Samples
[0226] NSCLC samples had a decreased G-Quad signature % in the uscfDNA region compared to non-cancer samples (Figure 19A). When the G-Quad-containing fragments were filtered out and analyzed for CpG methylation %, NSCLS samples were observed to be significantly hypermethylated compared to non-cancer samples in the uscfDNA fraction, while in the mncfDNA fraction the difference is not significant (Figure 19B). cfDNA overlap with cell type-specific epigenetic marks is altered in NSCLC
[0227] Next, it was examined if the normalized % of intersecting base pairs for epigenetic marks were altered in NSCLC. Three epigenetic marks (H3K27ac, H3K4me3, and hypomethylated regions) had significantly decreased overlaps in NSCLC samples in uscfDNA and mncfDNA fractions (Figure 19C). A decrease in %intersection was also observed in uscfDNA and mncfDNA in H3K27me, H3K36me, and H3K4me1 epigenetic marks but to a lower extent (Figure 20). There did not appear to be any differences in % intersection of H3K9me3 or hypermethylated regions in NSCLC samples (Figure 20). Table 3: DMR Discovery Between Non-Cancer and NSCLC for the uscfDNA Bin mean mean mean minus log10 methylation # methylation methylation chr start stop q-value q value difference CPGS healthy cancer chr11 402680 402832 0.001402 9.47829794 -0.245267 14 0.64899 0.89426 chr10 132804122 132804137 0.01767 5.82255415 0.318306 2 0.83081 0.5125chr1 125183700 125183714 0.027901 5.16353936 -0.187348 2 0.75714 0.94449 chr16 34593300 34593324 0.03634 4.78229777 0.264651 2 0.87356 0.60891 chr6 94511494 94511527 0.063903 3.96797253 -0.132516 4 0.62093 0.75345 chr5 49601536 49601567 0.071974 3.79638035 0.338325 2 0.8205 0.48218 chr19 613638 613663 0.07314 3.77319556 -0.302462 2 0.62633 0.92879 chr16 34594616 34594640 0.082733 3.59539329 -0.234542 2 0.59879 0.83333 chr4 49228855 49228882 0.12775 2.9686048 0.186043 2 0.25042 0.064374 chr20 61193119 61193145 0.15166 2.72108747 0.325609 2 0.67699 0.35138 chr21 8220145 8220154 0.30136 1.73044016 -0.258124 2 0.31965 0.57778 chr1 1141256 1141311 0.59523 0.74848085 0.283472 2 0.61472 0.33125 chr1 1141146 1141180 1 0 0.174557 2 0.41743 0.24288 chr1 125183787 125183789 1 0 -0.198365 2 0.65594 0.85431 chr1 228069111 228069115 1 0 -0.132379 2 0.50887 0.64125 chr1 228070027 228070213 1 0 0.107526 2 0.67081 0.56328 chr1 876344 876385 1 0 0.101027 5 0.80133 0.7003 chr10 38813193 38813224 1 0 0.156329 2 0.41126 0.25493 chr11 402855 402872 1 0 -0.168222 3 0.68659 0.85481 chr11 402893 402904 1 0 -0.222321 2 0.68753 0.90985 chr11 402908 402989 1 0 -0.12939 5 0.699 0.82839 chr12 114032617 114032643 1 0 0.118735 2 0.46348 0.34474 chr14 44798925 44798934 1 0 0.148887 2 0.8286 0.67971 chr16 34586832 34586846 1 0 0.140762 2 0.8944 0.75364 chr16 34586855 34587009 1 0 -0.109335 3 0.81451 0.92385 chr16 34592823 34593120 1 0 -0.107201 3 0.8928 1 chr16 34594883 34594894 1 0 0.164808 2 0.62684 0.46203 chr16 46390518 46390523 1 0 0.196395 2 0.60549 0.40909 chr16 46390552 46390670 1 0 -0.288445 2 0.62822 0.91667 chr16 46390757 46390768 1 0 -0.194491 2 0.80551 1 chr16 46394547 46394662 1 0 -0.236182 2 0.68622 0.9224 chr16 46394885 46394951 1 0 -0.214747 2 0.74359 0.95833 chr16 46400889 46400900 1 0 -0.136344 2 0.84082 0.97716 chr16 767942 767956 1 0 -0.208814 2 0.64404 0.85285 chr19 1049856 1049868 1 0 0.152073 2 0.88074 0.72867 chr19 47966403 47966418 1 0 0.221572 2 0.55877 0.3372 chr2 10469 10484 1 0 0.151878 3 0.48832 0.33644 chr2 10507 10514 1 0 0.14165 2 0.77112 0.62947 chr20 64094838 64094848 1 0 0.220635 2 0.88293 0.6623 chr22 11058149 11058169 1 0 0.222284 2 0.54839 0.32611chr4 1783495 1783507 1 0 0.147136 2 0.83482 0.68768 chr5 114929125 114929143 1 0 0.125147 2 0.67546 0.55031 chr5 49658133 49658180 1 0 -0.100111 2 0.43333 0.53344 chr8 1508433 1508464 1 0 0.135302 3 0.6475 0.5122 chr9 134764003 134764096 1 0 -0.113597 2 0.69554 0.80914 chrX 151562820 151562836 1 0 0.14694 2 0.61648 0.46954 Table 4: DMR Discovery Between Non-Cancer and NSCLC for the mncfDNA Bin minus mean log10 q methylation mean mean chr start stop q-value value difference # CPGS healthy cancer chr21 8208928 8209038 1.55E-17 1.68E+01 -0.440315 21 0.22635 0.66667 chr16 34586783 34586856 1.15E-05 4.94E+00 -0.109144 13 0.83421 0.94335 chr10 42080517 42080648 0.00067616 3.17E+00 -0.11688 7 0.84741 0.96429 chr10 38868465 38868511 0.0007295 3.14E+00 0.230502 4 0.89717 0.66667 chr16 46394166 46394407 0.0024709 2.61E+00 0.42982 4 0.83489 0.40507 chr16 34588514 34588812 0.0051581 2.29E+00 -0.179268 9 0.76147 0.94074 chr1 143214715 143214803 0.0076111 2.12E+00 0.150575 5 0.80058 0.65 chr20 29877880 29877930 0.013116 1.88E+00 -0.398665 3 0.53294 0.93161 chr20 31061227 31061429 0.016676 1.78E+00 -0.280577 4 0.61484 0.89542 chr21 7941250 7941311 0.016801 1.77E+00 -0.31839 4 0.54173 0.86012 chr1 143253126 143253234 0.030152 1.52E+00 -0.128313 10 0.80078 0.92909 0.08738 chr4 49137126 49137132 0.045945 1.34E+00 -0.312611 2 9 0.4 7.50E- chr21 8438824 8438830 0.058329 1.23E+00 0.23624 3 0.23624 07 chr16 34592887 34592898 0.058953 1.23E+00 0.114126 2 0.96961 0.85548 chr10 41859769 41859799 0.059415 1.23E+00 -0.336613 2 0.46339 0.8 chr10 38527782 38528353 0.059557 1.23E+00 -0.226058 7 0.72774 0.95379 chr4 49105934 49105986 0.0621 1.21E+00 0.418846 3 0.66885 0.25 chr16 34585653 34585751 0.065298 1.19E+00 -0.123719 7 0.83329 0.957 chr22 10728556 10728577 0.068429 1.16E+00 -0.226072 2 0.22165 0.44772 chr21 10748591 10748958 0.068963 1.16E+00 0.179979 4 0.34069 0.16071 chr2 90392160 90392171 0.095581 1.02E+00 0.280534 2 0.86387 0.58333 chr22 10729136 10729157 0.095891 1.02E+00 0.331825 2 0.88183 0.55 chr20 31060322 31060325 0.095901 1.02E+00 0.328603 2 0.6536 0.325 chr21 7946575 7946642 0.097489 1.01E+00 0.292159 4 0.78874 0.49658chr5 49660732 49660734 0.098754 1.01E+00 0.532763 2 0.79367 0.26091 chr16 34574029 34574040 0.10758 9.68E-01 0.231068 2 0.87393 0.64286 chr1 16726773 16726780 0.10764 9.68E-01 0.322407 2 0.94741 0.625 chr17 25074391 25074462 0.13661 8.65E-01 0.182518 4 0.82538 0.64286 chr10 38912715 38912801 0.13969 8.55E-01 0.176016 2 0.37602 0.2 chr21 7931424 7931511 0.15872 7.99E-01 0.131464 3 0.84813 0.71667 chr10 41890143 41890194 0.16073 7.94E-01 0.579855 2 0.70486 0.125 chr10 41882615 41882637 0.17594 7.55E-01 -0.260212 2 0.57312 0.83333 chr16 34594593 34594617 0.20712 6.84E-01 -0.180158 3 0.77307 0.95323 chr21 9329996 9330078 0.26434 5.78E-01 -0.160634 4 0.77687 0.9375 chr22 10717337 10717408 0.26452 5.78E-01 -0.315554 2 0.58445 0.9 chr21 10715494 10715609 0.28333 5.48E-01 0.405713 2 0.78071 0.375 chr16 46393142 46393421 0.31642 5.00E-01 0.370875 2 0.78159 0.41071 chr21 7940271 7940437 0.31776 4.98E-01 0.170938 7 0.74237 0.57143 chr20 31074188 31074379 0.33329 4.77E-01 -0.24736 3 0.63042 0.87778 chr2 90385864 90385973 0.37332 4.28E-01 -0.161875 6 0.71313 0.875 chr1 125180222 125180233 0.43158 3.65E-01 0.34081 2 0.95823 0.61742 chr1 125176962 125176972 0.48138 3.18E-01 0.102851 2 0.92791 0.82506 chr1 143213928 143213965 0.49546 3.05E-01 -0.15956 4 0.77794 0.9375 chr7 62309053 62309103 0.49989 3.01E-01 -0.254514 4 0.62049 0.875 chr21 7917130 7917367 0.51636 2.87E-01 -0.182619 5 0.74129 0.92391 7.50E- chr21 9163577 9163646 0.60549 2.18E-01 0.388641 2 0.38864 07 chr3 75669164 75669168 0.62906 2.01E-01 -0.260843 2 0.50999 0.77083 7.50E- chr21 6372913 6372938 0.63346 1.98E-01 0.234933 3 0.23493 07 7.50E- chr21 10751598 10751684 0.63796 1.95E-01 0.412206 2 0.41221 07 chr1 125178261 125178282 0.653 1.85E-01 0.12539 2 0.95479 0.8294 chr16 34572731 34572742 0.68616 1.64E-01 -0.107792 2 0.86642 0.97421 chr21 7916179 7916210 0.68693 1.63E-01 -0.275017 2 0.62498 0.9 chr21 7935285 7935290 0.68826 1.62E-01 -0.14153 2 0.66204 0.80357 chr2 90383807 90383821 0.68883 1.62E-01 0.159679 2 0.90968 0.75 chr17 25920336 25920338 0.69294 1.59E-01 0.253822 2 0.85994 0.60612 chr16 34591790 34591827 0.69593 1.57E-01 -0.293028 2 0.66393 0.95696 chr10 41896873 41896934 0.70037 1.55E-01 0.270684 2 0.58021 0.30952 chr1 125085197 125085224 0.70343 1.53E-01 0.418257 2 0.51826 0.1 chr10 41877875 41877897 0.70838 1.50E-01 0.246562 2 0.62156 0.375 chr21 7921599 7921660 0.71996 1.43E-01 -0.257375 2 0.64262 0.9 chr16 34588481 34588505 0.72411 1.40E-01 0.196505 2 0.82151 0.625chr10 38528522 38528553 0.73368 1.34E-01 0.211774 2 0.7368 0.52503 chr21 9248818 9248914 0.80044 9.67E-02 -0.297277 2 0.41701 0.71429 chr21 10706635 10706707 0.80272 9.54E-02 -0.243295 3 0.4948 0.7381 chr16 46399251 46399253 0.8234 8.44E-02 -0.166799 2 0.80195 0.96875 chr21 7937928 7937948 0.82387 8.41E-02 -0.355583 2 0.51942 0.875 0.00711 chr5 49602075 49602077 0.82494 8.36E-02 0.111563 2 0.11868 51 chr21 7948805 7948842 0.83624 7.77E-02 -0.217013 2 0.59965 0.81667 chr16 34593387 34593398 0.89785 4.68E-02 0.313873 3 0.85654 0.54267 chr1 210653990 210654035 0.89895 4.63E-02 -0.270461 2 0.60454 0.875 chr1 125080639 125080663 1 0.00E+00 0.406262 2 0.65626 0.25 chr1 125080721 125080761 1 0.00E+00 -0.266182 3 0.56715 0.83333 chr1 125177091 125177093 1 0.00E+00 -0.112063 2 0.84112 0.95318 chr1 125177658 125177669 1 0.00E+00 0.173168 2 0.96823 0.79506 chr1 125178154 125178230 1 0.00E+00 -0.11642 4 0.72485 0.84127 chr1 125181945 125181975 1 0.00E+00 0.215852 2 0.71784 0.50199 chr1 143213246 143213325 1 0.00E+00 0.10594 3 0.81518 0.70924 chr1 143232005 143232055 1 0.00E+00 0.124398 2 0.8994 0.775 chr1 143234417 143234515 1 0.00E+00 0.126088 5 0.80609 0.68 chr1 143264616 143264633 1 0.00E+00 0.257335 2 0.75734 0.5 chr1 224013818 224013879 1 0.00E+00 0.267218 2 0.76722 0.5 chr10 38527008 38527039 1 0.00E+00 -0.131849 2 0.78482 0.91667 7.50E- chr10 38527082 38527093 1 0.00E+00 0.407391 2 0.40739 07 chr10 38527208 38527229 1 0.00E+00 -0.192386 2 0.80761 1 chr10 38528632 38528718 1 0.00E+00 -0.19801 2 0.7656 0.96361 chr10 38905764 38905809 1 0.00E+00 -0.179407 2 0.69559 0.875 chr10 39967545 39967596 1 0.00E+00 0.159284 2 0.78428 0.625 chr10 41859559 41859605 1 0.00E+00 0.169705 2 0.7947 0.625 chr10 41859834 41859890 1 0.00E+00 -0.107987 2 0.58487 0.69286 chr10 41881946 41882386 1 0.00E+00 0.225548 5 0.72555 0.5 0.09978 chr10 41883518 41883574 1 0.00E+00 -0.103386 2 5 0.20317 chr10 41883604 41883714 1 0.00E+00 -0.286691 2 0.21331 0.5 chr10 41904413 41904464 1 0.00E+00 0.186191 2 0.63619 0.45 chr10 42066804 42066827 1 0.00E+00 -0.268399 2 0.67605 0.94444 chr10 42070047 42070097 1 0.00E+00 -0.117953 3 0.88205 1 chr10 42070637 42070888 1 0.00E+00 -0.12876 5 0.87124 1 chr11 70371426 70371454 1 0.00E+00 0.196758 2 0.44676 0.25 0.08333 chr13 18178211 18178263 1 0.00E+00 0.348895 2 0.43223 4chr16 34571977 34572021 1 0.00E+00 -0.171737 2 0.60207 0.77381 chr16 34572846 34572870 1 0.00E+00 -0.170897 2 0.52586 0.69676 chr16 34573246 34573257 1 0.00E+00 -0.158317 2 0.72745 0.88576 chr16 34576594 34576611 1 0.00E+00 -0.149402 2 0.7506 0.9 chr16 34582315 34582374 1 0.00E+00 -0.148842 2 0.7048 0.85365 chr16 34584643 34584765 1 0.00E+00 -0.150218 5 0.84978 1 chr16 34584788 34584815 1 0.00E+00 0.290393 3 0.70706 0.41667 chr16 34585842 34585894 1 0.00E+00 0.278959 4 0.79026 0.5113 chr16 34585952 34585954 1 0.00E+00 -0.133237 2 0.75909 0.89233 chr16 34586187 34586191 1 0.00E+00 0.16938 2 0.78942 0.62004 chr16 34586327 34586354 1 0.00E+00 -0.143192 2 0.85681 1 chr16 34587405 34587407 1 0.00E+00 0.417077 2 0.85097 0.43389 chr16 34587416 34587455 1 0.00E+00 -0.164556 2 0.71125 0.87581 chr16 34588290 34588292 1 0.00E+00 -0.124351 2 0.7419 0.86625 chr16 34593493 34593730 1 0.00E+00 0.10769 2 0.76178 0.65409 chr16 34593788 34593861 1 0.00E+00 -0.126675 2 0.82201 0.94868 chr16 34594556 34594558 1 0.00E+00 0.11049 2 0.89443 0.78394 chr16 34594580 34594593 1 0.00E+00 0.215403 2 0.83024 0.61483 chr16 34594876 34594878 1 0.00E+00 -0.138673 2 0.78888 0.92755 chr16 34594892 34594894 1 0.00E+00 -0.160688 2 0.61023 0.77092 chr16 34595064 34595075 1 0.00E+00 -0.15563 2 0.84437 1 chr16 46380745 46380821 1 0.00E+00 -0.127954 2 0.85465 0.98261 chr16 46380835 46380846 1 0.00E+00 0.172926 2 0.76848 0.59555 chr16 46380868 46380944 1 0.00E+00 -0.140709 4 0.81384 0.95455 chr16 46386423 46386425 1 0.00E+00 -0.162311 2 0.76342 0.92573 chr16 46386537 46386539 1 0.00E+00 0.181582 2 0.80031 0.61872 chr16 46388845 46388856 1 0.00E+00 -0.194981 2 0.47588 0.67086 chr16 46388991 46389002 1 0.00E+00 -0.206517 2 0.71659 0.92311 chr16 46389637 46389639 1 0.00E+00 0.102369 2 0.90722 0.80486 chr16 46390156 46390206 1 0.00E+00 -0.161819 2 0.55532 0.71714 chr16 46390551 46390553 1 0.00E+00 -0.111727 2 0.66271 0.77444 chr16 46390918 46391077 1 0.00E+00 -0.140124 3 0.79973 0.93985 chr16 46391221 46391280 1 0.00E+00 0.132804 4 0.83775 0.70495 chr16 46391391 46391459 1 0.00E+00 0.105216 4 0.61086 0.50565 chr16 46394459 46394461 1 0.00E+00 0.141387 2 0.85196 0.71057 chr16 46394625 46394627 1 0.00E+00 -0.166228 2 0.65478 0.82101 chr16 46394973 46394975 1 0.00E+00 -0.15478 2 0.78361 0.93839 chr16 46395457 46395530 1 0.00E+00 0.205914 4 0.63901 0.43309chr16 46395910 46395937 1 0.00E+00 -0.137991 2 0.842 0.97999 chr16 46398293 46398304 1 0.00E+00 -0.157653 2 0.74235 0.9 chr16 46398342 46398353 1 0.00E+00 -0.175796 2 0.8242 1 chr16 46398760 46398787 1 0.00E+00 0.140199 2 0.82196 0.68176 chr16 46398812 46398834 1 0.00E+00 -0.103138 6 0.82488 0.92802 chr16 46399123 46399157 1 0.00E+00 -0.101672 5 0.68873 0.7904 chr16 46399157 46399230 1 0.00E+00 0.274707 3 0.645 0.37029 chr16 46400481 46400483 1 0.00E+00 0.160636 2 0.8773 0.71667 chr16 46400498 46400557 1 0.00E+00 0.214904 2 0.7899 0.575 chr16 46400579 46400581 1 0.00E+00 0.253424 2 0.71369 0.46027 chr16 46400678 46400692 1 0.00E+00 0.148969 2 0.802 0.65304 chr16 46401046 46401334 1 0.00E+00 0.227702 2 0.79012 0.56242 chr17 21863905 21863910 1 0.00E+00 0.256785 2 0.65679 0.4 chr17 21870725 21870815 1 0.00E+00 -0.228162 2 0.65645 0.88462 chr17 21883845 21883906 1 0.00E+00 0.157838 4 0.4879 0.33006 chr17 21910481 21910497 1 0.00E+00 -0.1206 2 0.16273 0.28333 chr17 21910517 21910602 1 0.00E+00 0.147472 4 0.43914 0.29167 chr17 21969264 21969450 1 0.00E+00 -0.138104 4 0.4244 0.5625 chr17 21973125 21973196 1 0.00E+00 0.245432 2 0.67723 0.4318 chr17 25389877 25389879 1 0.00E+00 0.441141 2 0.68353 0.24239 chr17 25389915 25389917 1 0.00E+00 0.13636 2 0.48823 0.35187 chr17 26594862 26594872 1 0.00E+00 0.109843 2 0.89299 0.78314 chr17 26619473 26619650 1 0.00E+00 -0.116817 5 0.81641 0.93323 chr17 26785326 26785338 1 0.00E+00 0.100391 3 0.43372 0.33333 chr17 26840643 26840684 1 0.00E+00 -0.29742 2 0.51537 0.81279 chr17 26842714 26842745 1 0.00E+00 -0.114142 2 0.27604 0.39018 chr17 26842994 26843019 1 0.00E+00 0.204334 2 0.56148 0.35714 chr17 26856066 26856086 1 0.00E+00 -0.133034 2 0.45808 0.59111 chr17 26885479 26885481 1 0.00E+00 0.159664 2 0.88675 0.72708 chr17 26938833 26938839 1 0.00E+00 0.176406 2 0.5541 0.37769 chr17 26938927 26938988 1 0.00E+00 -0.156556 2 0.21487 0.37143 chr18 110772 110840 1 0.00E+00 0.187815 2 0.43781 0.25 chr18 110840 110922 1 0.00E+00 -0.311018 2 0.43898 0.75 chr19 50123811 50123820 1 0.00E+00 0.103577 2 0.91123 0.80766 chr19 50136098 50136138 1 0.00E+00 -0.15499 2 0.84501 1 chr2 117077819 117077840 1 0.00E+00 -0.163593 5 0.83641 1 chr2 161279006 161279014 1 0.00E+00 -0.110331 2 0.7256 0.83593 chr2 161279071 161279196 1 0.00E+00 -0.134104 2 0.69818 0.83228chr2 161279292 161279376 1 0.00E+00 -0.166553 2 0.50068 0.66724 chr2 161280368 161280452 1 0.00E+00 -0.116311 4 0.42536 0.54167 chr2 89814416 89814528 1 0.00E+00 0.133501 6 0.57001 0.43651 chr2 89830993 89831050 1 0.00E+00 -0.21778 2 0.30706 0.52484 chr2 89841042 89841104 1 0.00E+00 -0.187181 2 0.38001 0.5672 chr2 90390473 90390533 1 0.00E+00 -0.138346 2 0.73538 0.87373 chr2 90390558 90390608 1 0.00E+00 0.120863 2 0.83566 0.7148 chr2 90390765 90390775 1 0.00E+00 0.148844 2 0.92206 0.77321 chr2 90390790 90390792 1 0.00E+00 -0.106946 2 0.82698 0.93393 chr2 90392134 90392148 1 0.00E+00 -0.216369 3 0.78363 1 chr2 90392209 90392243 1 0.00E+00 -0.255041 2 0.66163 0.91667 chr2 90396784 90396821 1 0.00E+00 0.122782 2 0.78783 0.66505 chr2 90397620 90397650 1 0.00E+00 -0.142881 2 0.78122 0.9241 chr2 90397863 90397913 1 0.00E+00 -0.157186 2 0.79406 0.95125 chr2 90398585 90398587 1 0.00E+00 0.237104 2 0.85829 0.62118 chr2 90398657 90398675 1 0.00E+00 0.184443 2 0.79962 0.61518 chr20 28861090 28861098 1 0.00E+00 -0.169279 2 0.58786 0.75714 chr20 29263569 29263571 1 0.00E+00 -0.108541 2 0.68139 0.78993 chr20 29263714 29263736 1 0.00E+00 -0.210085 2 0.78992 1 chr20 29877965 29878008 1 0.00E+00 0.287911 2 0.5022 0.21429 chr20 29884014 29884057 1 0.00E+00 0.191823 3 0.35849 0.16667 chr20 29896923 29896949 1 0.00E+00 -0.105308 3 0.47802 0.58333 0.06279 chr20 30813371 30813397 1 0.00E+00 0.434855 2 0.49765 3 chr20 31054746 31055102 1 0.00E+00 0.166922 5 0.61327 0.44635 chr20 31055415 31055427 1 0.00E+00 -0.211555 2 0.41344 0.625 chr20 31058189 31058200 1 0.00E+00 0.145211 2 0.71664 0.57143 chr20 31060363 31060473 1 0.00E+00 0.151769 2 0.5808 0.42904 chr20 31060938 31060968 1 0.00E+00 0.290034 2 0.68543 0.39539 chr20 31061578 31061658 1 0.00E+00 0.168019 2 0.71327 0.54525 chr20 31061662 31061769 1 0.00E+00 -0.144271 2 0.68906 0.83333 chr20 31064108 31064364 1 0.00E+00 -0.141045 5 0.69229 0.83333 chr20 31065258 31065309 1 0.00E+00 -0.163713 2 0.74273 0.90644 chr20 31066925 31066996 1 0.00E+00 -0.163236 2 0.75343 0.91667 chr20 31068409 31068634 1 0.00E+00 0.102024 6 0.6734 0.57138 chr20 31072498 31072799 1 0.00E+00 -0.141171 4 0.71694 0.85811 chr20 31075735 31075762 1 0.00E+00 -0.10971 2 0.64029 0.75 chr20 31185141 31185192 1 0.00E+00 -0.207154 2 0.18048 0.38763 chr20 31186825 31186901 1 0.00E+00 0.163977 2 0.27755 0.11358chr21 10416728 10416730 1 0.00E+00 0.22439 2 0.67439 0.45 chr21 10473034 10473113 1 0.00E+00 0.268217 2 0.53347 0.26525 chr21 10702757 10702899 1 0.00E+00 -0.182901 2 0.4421 0.625 chr21 10707516 10707526 1 0.00E+00 0.297741 2 0.79774 0.5 chr21 10732612 10732622 1 0.00E+00 0.206798 2 0.7068 0.5 chr21 10749893 10749901 1 0.00E+00 -0.126577 2 0.87342 1 chr21 10751343 10751363 1 0.00E+00 -0.143866 2 0.5228 0.66667 chr21 10762194 10762281 1 0.00E+00 -0.217656 2 0.44901 0.66667 chr21 6368452 6368477 1 0.00E+00 -0.168442 2 0.83156 1 chr21 6375851 6375931 1 0.00E+00 0.202305 2 0.61897 0.41667 chr21 7916622 7916713 1 0.00E+00 0.26817 2 0.86817 0.6 chr21 7916727 7916789 1 0.00E+00 -0.142359 2 0.85764 1 chr21 7916847 7916893 1 0.00E+00 0.345366 3 0.6787 0.33333 chr21 7917455 7917477 1 0.00E+00 0.146232 2 0.57126 0.42503 chr21 7919605 7919627 1 0.00E+00 0.341162 3 0.84116 0.5 chr21 7921415 7921435 1 0.00E+00 -0.274944 2 0.56077 0.83571 chr21 7921479 7921525 1 0.00E+00 0.168931 2 0.75226 0.58333 chr21 7925840 7925856 1 0.00E+00 -0.175687 2 0.73453 0.91021 chr21 7927396 7927511 1 0.00E+00 -0.307219 4 0.69278 1 chr21 7928614 7928630 1 0.00E+00 -0.107893 2 0.80877 0.91667 chr21 7929549 7929625 1 0.00E+00 -0.198113 3 0.65109 0.84921 chr21 7929724 7929880 1 0.00E+00 0.150299 2 0.7003 0.55 chr21 7930014 7930030 1 0.00E+00 0.292045 2 0.90871 0.61667 chr21 7930459 7930585 1 0.00E+00 0.177958 4 0.73396 0.556 chr21 7930734 7930825 1 0.00E+00 -0.281599 3 0.63507 0.91667 chr21 7930869 7931035 1 0.00E+00 -0.143698 6 0.59797 0.74167 chr21 7931824 7931900 1 0.00E+00 0.521365 4 0.77137 0.25 chr21 7931914 7931960 1 0.00E+00 -0.194043 2 0.80596 1 chr21 7932129 7932141 1 0.00E+00 -0.13205 2 0.78462 0.91667 chr21 7932174 7932220 1 0.00E+00 -0.110708 2 0.73133 0.84204 chr21 7932255 7932270 1 0.00E+00 0.137851 2 0.6589 0.52105 chr21 7932454 7932564 1 0.00E+00 -0.187684 2 0.48732 0.675 chr21 7934392 7934487 1 0.00E+00 -0.196611 3 0.80339 1 chr21 7935919 7935965 1 0.00E+00 0.295278 2 0.87028 0.575 chr21 7937090 7937097 1 0.00E+00 0.395843 2 0.64584 0.25 chr21 7938792 7938833 1 0.00E+00 0.230069 2 0.54435 0.31429 chr21 7939345 7939377 1 0.00E+00 -0.147731 2 0.75227 0.9 chr21 7944437 7944463 1 0.00E+00 0.235015 2 0.48502 0.25chr21 7947608 7947650 1 0.00E+00 -0.164374 2 0.46461 0.62898 chr21 7949155 7949166 1 0.00E+00 -0.345451 2 0.47273 0.81818 chr21 7950686 7950731 1 0.00E+00 -0.184094 2 0.56556 0.74966 chr21 7950745 7950761 1 0.00E+00 -0.136457 2 0.80104 0.9375 chr21 7952534 7952545 1 0.00E+00 0.148339 2 0.87782 0.72948 chr21 7952569 7952585 1 0.00E+00 0.117009 2 0.81366 0.69665 chr21 7955588 7955634 1 0.00E+00 -0.196728 2 0.6983 0.89503 chr21 7957297 7957311 1 0.00E+00 -0.195505 2 0.50449 0.7 chr21 8219812 8219815 1 0.00E+00 -0.28714 2 0.33823 0.62537 chr21 8220145 8220154 1 0.00E+00 -0.185471 2 0.30323 0.4887 0.03374 chr21 8438766 8438775 1 0.00E+00 0.114877 4 0.14862 6 chr21 8452312 8452338 1 0.00E+00 0.136667 2 0.93455 0.79788 chr21 8453224 8453322 1 0.00E+00 0.167182 2 0.87683 0.70964 chr21 8466915 8466933 1 0.00E+00 0.260253 2 0.91403 0.65378 chr21 9100045 9100067 1 0.00E+00 0.217405 2 0.76741 0.55 chr21 9130669 9130725 1 0.00E+00 -0.257272 2 0.58414 0.84141 chr21 9141177 9141204 1 0.00E+00 -0.181758 2 0.53363 0.71538 chr21 9164670 9164704 1 0.00E+00 -0.195718 2 0.80428 1 chr21 9166749 9166800 1 0.00E+00 0.265234 2 0.50914 0.2439 chr21 9246461 9246599 1 0.00E+00 -0.23424 2 0.33189 0.56613 chr21 9250565 9250599 1 0.00E+00 -0.210351 2 0.65697 0.86732 chr21 9252236 9252454 1 0.00E+00 -0.121779 4 0.46521 0.58699 chr21 9252855 9252971 1 0.00E+00 -0.184739 2 0.31526 0.5 chr21 9317912 9317930 1 0.00E+00 -0.179722 2 0.44528 0.625 chr21 9330111 9330184 1 0.00E+00 0.275351 2 0.76948 0.49413 chr21 9330455 9330466 1 0.00E+00 0.227308 2 0.82499 0.59769 chr21 9330609 9330715 1 0.00E+00 -0.226609 3 0.62739 0.85399 chr21 9371916 9372097 1 0.00E+00 -0.231912 2 0.62654 0.85845 chr21 9372207 9372271 1 0.00E+00 0.155581 2 0.80693 0.65135 chr21 9819076 9819099 1 0.00E+00 -0.159169 2 0.43012 0.58929 chr21 9819099 9819145 1 0.00E+00 0.124383 2 0.52438 0.4 0.09549 chr22 10724136 10724266 1 0.00E+00 0.229167 4 0.32466 8 chr22 10727369 10727396 1 0.00E+00 -0.144663 2 0.63034 0.775 chr22 10727781 10727867 1 0.00E+00 0.332383 2 0.58238 0.25 chr22 10728116 10728142 1 0.00E+00 -0.154314 2 0.72069 0.875 chr22 10728496 10728532 1 0.00E+00 -0.169113 4 0.39339 0.5625 chr22 10728661 10728807 1 0.00E+00 0.175367 2 0.62537 0.45 chr22 10729041 10729112 1 0.00E+00 -0.146428 2 0.47857 0.625chr22 10738903 10739180 1 0.00E+00 0.100969 2 0.41347 0.3125 chr22 11065310 11065344 1 0.00E+00 -0.170573 4 0.78776 0.95833 chr22 11927687 11927779 1 0.00E+00 -0.226474 2 0.15879 0.38526 chr22 11930076 11930265 1 0.00E+00 -0.23116 3 0.41698 0.64814 chr22 12175074 12175180 1 0.00E+00 0.175536 4 0.80054 0.625 chr22 12176187 12176190 1 0.00E+00 -0.456297 2 0.5437 1 chr22 12176464 12176487 1 0.00E+00 -0.256294 2 0.61871 0.875 chr22 18730660 18730703 1 0.00E+00 -0.190826 3 0.77855 0.96937 chr3 75669278 75669293 1 0.00E+00 0.117728 2 0.61483 0.4971 chr4 49092741 49092780 1 0.00E+00 -0.215694 2 0.53431 0.75 chr4 49096167 49096248 1 0.00E+00 -0.146549 3 0.52012 0.66667 chr4 49098598 49098683 1 0.00E+00 -0.157281 2 0.51772 0.675 chr4 49099614 49099626 1 0.00E+00 0.114656 2 0.71777 0.60312 chr4 49102579 49102626 1 0.00E+00 -0.135322 2 0.368 0.50332 chr4 49102664 49102700 1 0.00E+00 0.188847 2 0.68089 0.49204 chr4 49103851 49103867 1 0.00E+00 -0.197 2 0.2373 0.4343 chr4 49104340 49104485 1 0.00E+00 0.147447 4 0.38638 0.23894 0.04195 chr4 49108352 49108358 1 0.00E+00 0.165907 2 0.20786 2 chr4 49108396 49108432 1 0.00E+00 0.144729 2 0.74167 0.59694 7.50E- chr4 49113385 49113391 1 0.00E+00 0.133888 2 0.13389 07 chr4 49118682 49118753 1 0.00E+00 0.200734 2 0.60281 0.40207 chr4 49119076 49119121 1 0.00E+00 0.282381 2 0.74933 0.46695 7.50E- chr4 49119486 49119492 1 0.00E+00 0.107718 2 0.10772 07 chr4 49120890 49120916 1 0.00E+00 0.139953 2 0.67745 0.5375 chr4 49120929 49121005 1 0.00E+00 -0.1177 2 0.62448 0.74218 chr4 49121115 49121135 1 0.00E+00 -0.149421 2 0.48721 0.63663 chr4 49149680 49149765 1 0.00E+00 0.176521 2 0.45569 0.27917 chr4 49150119 49150255 1 0.00E+00 -0.139981 2 0.1153 0.25528 chr4 49150403 49150450 1 0.00E+00 -0.342021 2 0.46112 0.80314 7.50E- chr4 49150729 49150760 1 0.00E+00 0.21614 2 0.21614 07 chr4 49511528 49511653 1 0.00E+00 0.112048 7 0.6391 0.52706 chr4 49512747 49512812 1 0.00E+00 0.16893 4 0.66893 0.5 chr4 49512824 49512873 1 0.00E+00 -0.1258 3 0.70753 0.83333 7.50E- chr4 49651475 49651481 1 0.00E+00 0.16619 2 0.16619 07 chr4 49657539 49657843 1 0.00E+00 0.186727 4 0.55037 0.36364 chr4 49710751 49710912 1 0.00E+00 -0.288228 2 0.32644 0.61467 0.08562 chr4 49711435 49711437 1 0.00E+00 -0.216487 2 4 0.30211chr4 67457 67656 1 0.00E+00 -0.244816 2 0.72096 0.96578 chr5 46434567 46434762 1 0.00E+00 -0.107339 2 0.89266 1 chr5 46617383 46617556 1 0.00E+00 -0.122514 2 0.87749 1 chr5 49599881 49599883 1 0.00E+00 0.151952 2 0.83571 0.68376 chr5 49599921 49599923 1 0.00E+00 -0.176164 2 0.6344 0.81056 chr5 49600599 49600601 1 0.00E+00 0.178286 2 0.30354 0.12525 chr5 49601640 49601642 1 0.00E+00 0.12218 2 0.61037 0.48819 chr5 49601731 49601758 1 0.00E+00 -0.102694 2 0.79731 0.9 chr5 49601842 49601936 1 0.00E+00 0.212912 2 0.37577 0.16285 chr5 49602046 49602066 1 0.00E+00 0.206297 2 0.81164 0.60534 chr5 49602751 49602903 1 0.00E+00 -0.344109 3 0.65589 1 chr5 49657282 49657303 1 0.00E+00 0.150721 2 0.5629 0.41218 chr5 49657587 49657657 1 0.00E+00 -0.13316 2 0.74106 0.87422 chr5 49658203 49658235 1 0.00E+00 0.200404 2 0.55813 0.35772 chr5 49658961 49658988 1 0.00E+00 0.146772 2 0.70458 0.55781 chr5 49659152 49659239 1 0.00E+00 -0.271672 2 0.6569 0.92857 chr5 49661097 49661400 1 0.00E+00 -0.116271 4 0.56449 0.68076 chr5 49661517 49661547 1 0.00E+00 0.106489 2 0.51523 0.40874 chr5 49666598 49666619 1 0.00E+00 -0.197268 2 0.55273 0.75 chr7 100958759 100958773 1 0.00E+00 0.238312 2 0.48831 0.25 chr7 152405669 152405695 1 0.00E+00 0.203154 2 0.32815 0.125 chr7 58064281 58064344 1 0.00E+00 -0.184265 3 0.61574 0.8 chr7 60920866 60920871 1 0.00E+00 -0.148103 2 0.6019 0.75 chr9 63823834 63824040 1 0.00E+00 -0.203969 2 0.29603 0.5 chr9 63824063 63824364 1 0.00E+00 0.141751 2 0.70997 0.56822 chrX 49346320 49346362 1 0.00E+00 -0.255015 2 0.74498 1 Example 2: Next-Generation Sequencing Pipeline To Detect Methylation Status Of Ultrashort Single-Stranded Cell-Free DNA
[0228] Described is a Next-generation Sequencing (NGS) pipeline to determine the methylation profile of ultrashort single-stranded cell-free DNA (uscfDNA) within a biofliud sample for disease detection. This NGS pipeline is unique in that it uses 5mC-protected adapters on single-stranded DNA prior to bisulfite conversion and subsequent incorporation into a final library. It will report the methylation status of ultrashort cell-free ssDNA of 25-75bp in additionto the prototypical ~150bp mononucleosomal cfDNA (mncfDNA). This pipeline can be used for clinical applications or for molecular biology research purposes (Figure 23).
[0229] The disclosures of each and every patent, patent application, and publication cited herein are hereby incorporated herein by reference in their entirety. While this invention has been disclosed with reference to specific embodiments, it is apparent that other embodiments and variations of this invention may be devised by others skilled in the art without departing from the true spirit and scope of the invention. The appended claims are intended to be construed to include all such embodiments and equivalent variations.
Claims
CLAIMS 1. A method for generating a bisulfite converted library from ultrashort single-stranded cell- free DNA (uscfDNA) and / or mononucleosomal cell-free DNA (mncfDNA), the method comprising the steps of: a) ligating a 5mC-protected adaptor to the 5’ and 3’ ends of the uscfDNA and / or mncfDNA; and b) treating the adaptor-conjugated uscfDNA and / or mncfDNA with bisulfite, thereby generating bisulfite converted uscfDNA and / or mncfDNA.
2. The method of claim 1 further comprising preparing a sequence library from the bisulfite converted uscfDNA and / or mncfDNA.
3. The method of claim 2, further comprising the step of sequencing the library of bisulfite converted uscfDNA and / or mncfDNA.
4. The method of claim 1, wherein the sample is a biological fluid sample.
5. The method of claim 4, wherein the sample is selected from the group consisting of a blood sample, a plasma sample, a serum sample, a saliva sample, a sweat sample, a sputum sample, a urine sample, an amniotic fluid sample, a peritoneal fluid sample, a pleural fluid sample, and a liquid biopsy sample.
6. A method of identifying biomarkers for a disease or disorder comprising obtaining mncfDNA, uscfDNA, or a combination thereof, from a sample according to the method of any one of claims 1-5 and analyzing the methylation profile of the mncfDNA, uscfDNA, or the combination thereof, to identify biomarkers of a disease or disorder.
7. The method of claim 6, wherein the biomarker is an increase or decrease in the total amount of methylation of uscfDNA, mncfDNA, or a combination thereof in a test sample as compared to a control sample.
8. The method of claim 6, wherein the biomarker is an increase or decrease in the amount of methylation in a specific region of uscfDNA or mncfDNA in a test sample as compared to a control sample.
9. A method of diagnosing a disease or disorder in a subject in need thereof, the method comprising obtaining a sample from the subject, isolating uscfDNA, mncfDNA, analyzing the total amount of methylation or the amount of methylation of a specific region of the mncfDNA, uscfDNA, or a combination thereof to detect a methylation biomarker of a disease or disorder, and diagnosing the subject as having or at risk of the disease or disorder associated with the identified biomarker.
10. The method of claim 9 wherein analyzing the total amount of methylation or the amount of methylation of a specific region of the mncfDNA, uscfDNA, or a combination thereof, comprises the steps of: a) ligating a 5mC-protected adaptor to the 5’ and 3’ ends of the uscfDNA and / or mncfDNA; b) treating the adaptor-conjugated uscfDNA and / or mncfDNA with bisulfite, thereby generating bisulfite converted uscfDNA and / or mncfDNA.
11. The method of claim 9, wherein the biomarker is an increase or decrease in the total amount of methylation of uscfDNA, mncfDNA, or a combination thereof in a test sample as compared to a control sample.
12. The method of claim 9, wherein the biomarker is an increase or decrease in the amount of methylation associated with a specific region of uscfDNA, or mncfDNA, in a test sample as compared to a control sample.
13. The method of claim 9, wherein the disease or disorder is cancer.
14. The method of claim 13, further comprising administering a treatment based on the diagnostic outcome for the disease or disorder.
15. The method of claim 9, wherein the disease or disorder is non-small cell lung cancer.
16. The method of claim 15, wherein the specific region is selected from the group consisting of: a) chromosome 1, bases 125183700-125183714; b) chromosome 1, bases 143214715- 143214803; c) chromosome 1, bases 143253126- 143253234; d) chromosome 4, bases 49137126- 49137132; e) chromosome 10, bases 38868465- 38868511; f) chromosome 10, bases 42080517-42080648; g) chromosome 10, bases 132804122-132804137; h) chromosome 11, bases 402680- 402832; i) chromosome 16, bases 34586783- 34586856; j) chromosome 16, bases 34588514- 34588812; k) chromosome 16, bases 34593300- 34593324; l) chromosome 16, bases 46394166- 46394407; m) chromosome 20, bases 29877880- 29877930; n) chromosome 20, bases 31061227- 31061429; o) chromosome 21, bases 7941250- 7941311; and p) chromosome 21, bases 8208928- 8209038.
17. The method of claim 15, further comprising administering a treatment based on the diagnostic outcome for non-small cell lung cancer.