Classification of samples using methylation enrichment
By enriching for nucleic acids with CpG sites using a targeted adapter ligation and nuclease cutting method, the challenges of low and fragmented DNA in clinical DNA methylation profiling are addressed, achieving accurate and comprehensive methylation profiling even in challenging samples.
Patent Information
- Application Number
- PCT/US2024/056111
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods for clinical DNA methylation profiling face challenges with low and fragmented DNA input, low tumor fraction, and minimal overlap between methylation array probes and cell-type specific markers, particularly in samples like cell-free DNA or degraded DNA from FFPE samples.
The method involves ligating a first adapter to target double-stranded nucleic acids, cutting with a nuclease that targets CpG sites, and adding a second adapter to enrich for nucleic acids with CpG sites, allowing for improved methylation status assessment even in challenging samples.
This approach enables accurate methylation profiling in samples with low and fragmented DNA, achieving high correlation with gold standard methods and providing comprehensive methylation data, including in liquid biopsies and small biopsy samples.
Smart Images

Figure US2024056111_22052025_PF_FP_ABST
Abstract
Description
CLASSIFICATION OF SAMPLES USING METHYLATION ENRICHMENTPRIORITY STATEMENT
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 599,650, filed November 16, 2023, the entire contents of which are incorporated herein by reference for all purposes.STATEMENT REGARDING FEDERAL FUNDING
[0002] This invention was made with Government support under contract CA230156 awarded by the National Institutes of Health. The Government has certain rights in the invention.SEQUENCE LISTING
[0003] The text of the computer readable sequence listing filed herewith, titled “STDU2-42579- 601_SQL.xml”, created November 14, 2024, having a file size of 12,139 bytes, is hereby incorporated by reference in its entirety.FIELD
[0004] The present disclosure relates to methods of analyzing nucleic acids in a sample. In particular, provided herein are methods of analyzing a sample involving enriching for nucleic acids containing a CpG site in a sample and assessing methylation status in enriched nucleic acids.BACKGROUND
[0005] Clinical DNA methylation profiling is being rapidly developed and adopted as a diagnostic test for tumor classification. This progressive approach, especially alongside conventional histopathological and molecular method, can enhance tumor assessments and, sometimes even result in a change in the final diagnosis. However, DNA methylation classification remains a challenge in samples with low and fragmented DNA input, as in the caseof cell-free (of) or degraded DNA from archived formalin-fixed paraffin-embedded (FFPE) samples, and in samples with low tumor fraction due to the high immune or stromal cell background. Moreover, there is minimal overlap (11%) between probes commonly used derived from methylation arrays and markers specific for cell types. Accordingly, there is a need for improved methods for clinical DNA methylation profiling that are effective in challenging samples, such as those with low and / or fragmented DNA input or a low tumor fraction, such as liquid biopsy or small biopsy samples.SUMMARY
[0006] Embodiments of the present disclosure include methods of preparing nucleic acids for analysis. In some embodiments, provided herein is a method of preparing nucleic acids for analysis, comprising: a) ligating a first adapter to target double-stranded nucleic acids in a sample, wherein the first adapter comprises a primer binding strand containing a first primer binding site; b) adding a nuclease that cuts at a motif containing a CpG site, if present in the target double-stranded nucleic acids in the sample, thereby producing a subpopulation of nucleic acid fragments ligated to the first adapter, wherein at least some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer binding strand of the first adapter; and c) adding a second adapter to the sample, wherein a portion of the second adapter comprising a second primer landing site ligates to the nucleic acid fragments ligated the first adapter, thereby forming nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site.
[0007] In some embodiments, the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6).
[0008] In some embodiments, the nuclease cuts at a 5’ location to the CpG site in the motif.
[0009] In some embodiments, the nuclease cut between the first and second residues of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.
[0010] In some embodiments, the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
[0011] In some embodiments, the target double- stranded nucleic acid comprises a first strand and a second strand, and a 5’ end of the primer binding strand ligates to a 3’ end of a first strand of the target double-stranded nucleic acid. In some embodiments, the primer binding strand comprises a 5’ phosphate.
[0012] In some embodiments, the first adapter comprises the primer binding strand and a blocking strand. In some embodiments, the primer binding strand ligates to a first strand of the target double-stranded nucleic acid in the sample and the blocking strand ligates to a second strand of the target double- stranded nucleic acid in the sample.
[0013] In some embodiments, the primer binding strand ligates to a first strand of the target double-stranded nucleic acids in the sample, and a second strand of the target double- stranded nucleic acids in the sample does not ligate to the first adapter. For example, in some embodiments the adapter comprises a primer binding strand that ligates to the first strand of the target double-stranded nucleic acids, and a blocking strand that does not ligate to the second strand of the target double- stranded nucleic acids. As another example, in some embodiments a single- stranded ligation of the primer binding strand to the first strand of the target doublestranded nucleic acid occurs. In some embodiments, the first primer binding site of the first adapter comprises cytosine residues resistant to a methylation conversion or no cytosine residues.
[0014] The second adapter may be double- stranded or single-stranded. In some embodiments, the second adapter is double- stranded. In some embodiments, the second adapter is singlestranded.
[0015] In some embodiments, the cut site flanking sequence results from the nuclease cutting a single motif containing a CpG site in a target nucleic acid strand.
[0016] In some embodiments, the nuclease cuts at a motif containing a CpG site in each strand of the double- stranded target nucleic acid, if present in the sample, such that the subpopulation of nucleic acid fragments ligated to the first adapter are substantially double-stranded. In some embodiments, for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of asecond strand of the substantially double-stranded nucleic acid fragment is ligated to the blocking strand of the first adapter. In some embodiments, for the substantially double-stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double- stranded nucleic acid fragment is not ligated to or is otherwise disconnected from the first adapter. For example, the second strand may be initially ligated to the first adapter but is subsequently disconnected from the first adapter. As another example, the second strand may initially ligated to the first adapter but the ligation is subsequently disrupted by degradation of either the second strand or the portion of the first adapter to which the second strand was initially ligated.
[0017] In some embodiments, the method comprises performing a conversion step to convert cytosine residues in the sample depending on their methylation status to a different nucleic acid besides cytosine. In some embodiments, the conversion step selectively converts unmethylated cytosine residues in the sample to uracil residues. In some embodiments, the conversion step is an enzymatic conversion or bisulfite conversion. In some embodiments, the enzymatic conversion step comprises contacting the sample with an APOBEC -related enzyme or a TET- related enzyme. In some embodiments, the conversion step is performed in between adding the nuclease and adding the second adapter to the sample. In some embodiments, the conversion step is performed after adding the nuclease and adding the second adapter to the sample.
[0018] In some embodiments, the nuclease is MspI, Taql-v2, AccII, Acil, AspLEI, HPall, HpyCH4IV, or a CRISPR / Cas nuclease.
[0019] In some embodiments, the cut site flanking sequence is at least 20 base pairs in length. In some embodiments, the cut site flanking sequence is 20 to 300 base pairs in length. For example, in some embodiments the cut site flanking sequence is about 70 to about 150 base pairs in length.
[0020] In some embodiments, the method further comprises selectively amplifying and / or sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site. In some embodiments, the method further comprises generating a methylation profile for the sample based upon said sequencing. In some embodiments, the method comprises sequencing the nucleic acid strands comprising the first primer landing site, the cute site flanking sequence, and the second primer landing site by nanopore sequencing. In some embodiments, no enzymatic conversion step to selectivelyconvert cytosine residues in the sample depending on their methylation status is performed prior to said sequencing nucleic acid strands by nanoporc sequencing
[0021] Further embodiments of the present disclosure provide methods of classifying methylation status of a sample. In some embodiments, provided herein is a method of classifying methylation status of a sample, comprising: a) ligating a first adapter to target double-stranded nucleic acids in a sample, wherein the first adapter comprises a double- stranded end comprising a primer binding strand containing a first primer binding site; b) adding a nuclease that cuts at a motif containing a CpG site, if present in target nucleic acids in the sample, thereby producing a subpopulation of nucleic acid fragments ligated to the first adapter, wherein at least some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer binding strand of the first adapter; c) adding a second adapter to the sample, wherein a portion of the second adapter comprising a second primer landing site ligates to the nucleic acid fragments ligated to the first adapter, thereby forming nucleic acid strands comprising the first primer binding site, the cut site flanking sequence, and the second primer binding site; and d) generating a methylation profile for the sample by sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site.
[0022] In some embodiments, the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TOGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6).
[0023] In some embodiments, the nuclease cuts at a 5’ location to the CpG site in the motif.
[0024] In some embodiments, the nuclease cut between the first and second residues of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.
[0025] In some embodiments, the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
[0026] In some embodiments, the method further comprises performing a conversion step to selectively convert cytosine residues in the sample depending on their methylation status to adifferent nucleic acid besides cytosine. The second adapter may be added to the sample before or after performing the conversion step. For example, in some embodiments the second adapter is added to the sample before performing the conversion step. As another example, in some embodiments the second adapter is added to the sample after performing the conversion step.
[0027] In some embodiments, the conversion step is an enzymatic conversion or bisulfite conversion. In some embodiments, the enzymatic conversion step comprises contacting the sample with an APOBEC-related enzyme or a TET-related enzyme.
[0028] In some embodiments, said sequencing the nucleic acid strands comprises nanopore sequencing. In some embodiments, no enzymatic conversion step to selectively convert cytosine residues in the sample depending on their methylation status is performed prior to said sequencing nucleic acid strands by nanopore sequencing
[0029] In some embodiments, the target double- stranded nucleic acid comprises a first strand and a second strand, and a 5’ end of the primer binding strand ligates to a 3’ end of a first strand of the target double-stranded nucleic acid. In some embodiments, the primer binding strand comprises a 5’ phosphate.
[0030] In some embodiments, the first adapter comprises the primer binding strand and a blocking strand. In some embodiments, the primer binding strand ligates to a first strand of the target double-stranded nucleic acid in the sample and the blocking strand ligates to a second strand of the target double- stranded nucleic acid in the sample. In some embodiments, the primer binding strand ligates to a first strand of the target double- stranded nucleic acids in the sample, and a second strand of the target double-stranded nucleic acids in the sample does not ligate to or is otherwise disconnected from the first adapter. For example, the second strand may be initially ligated to the first adapter but is subsequently disconnected from the first adapter. As another example, the second strand may initially ligated to the first adapter but the ligation is subsequently disrupted by degradation of either the second strand or the portion of the first adapter to which the second strand was initially ligated.
[0031] In some embodiments, the first primer binding site of the first adapter comprises cytosine residues resistant to a methylation conversion or no cytosine residues.
[0032] The second adapter may be single-stranded or double-stranded. For example, in some embodiments the second adapter is double-stranded. As another example, in some embodiments the second adapter is single-stranded.
[0033] In some embodiments, the cut site flanking sequence results from the nuclease cutting a single motif containing a CpG site in a target nucleic acid strand.
[0034] In some embodiments, the nuclease cuts at a motif containing a CpG site in each strand of the double- stranded target nucleic acid, if present in the sample, such that the subpopulation of nucleic acid fragments ligated to the first adapter are substantially double-stranded. In some embodiments, for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double-stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is ligated to the blocking strand of the first adapter. In some embodiments, for the substantially double-stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is not ligated to or is otherwise disconnected from the first adapter.
[0035] In some embodiments, the nuclease is MspI, Taql-v2, AccII, Acil, AspLEI, HPall, HpyCH4IV, or a CRISPR / Cas nuclease.
[0036] In some embodiments, the cut site flanking sequence is at least 20 base pairs in length. In some embodiments, the cut site flanking sequence is 20 to 300 base pairs in length. For example, in some embodiments the cut site flanking sequence is about 70 to about 150 base pairs in length.
[0037] In some embodiments, the method generates a methylation profile of at least 100,000 loci. In some embodiments, method generates a methylation profile of at least 1,000,000 loci.
[0038] For any of the methods described herein, the sample may be a liquid biopsy sample. For any of the methods described herein, the sample may comprise less than 250 ng of DNA. The sample may be obtained from a subject having or suspected of having cancer.BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIGS. 1A-1B show a schematic of an exemplary FLEXseq design and workflow.FIG.1A shows the design goal and workflow of FLEXseq. The design aims to sequence adjacent regions flanking motifs containing a CpG site, such as CCGG (SEQ ID NO: 1) motifs, while preserving the methylation markers at the motif. Information-poor regions are suppressed, while information-rich flanking regions are amplified. In the workflow, input DNA fragments are firstligated with a semi-permissive Adapter A, which serves both as a blocker for untargeted DNA and a piece of targeted DNA. In this exemplary embodiment, the exemplary nuclease MspI then cuts at the CCGG motif, followed by the ligation of Adapter B. Only molecules with both Adapters A and B are amplified and sequenced, excluding untargeted DNA with only one type of adapter. In this exemplary method, and enzymatic conversion step is performed. As described herein, depending on the sequencing method used this enzymatic conversion step may be performed or not performed (e.g. for nanopore sequencing). The term “FLEXseq” does not indicate that an enzymatic conversion step must be performed. Although the schematic shows targeting of a CCGG (SEQ ID NO: 1) motif, this CCGG (SEQ ID NO: 1) motif is only to be an exemplary motif containing a CpG site described herein, but other motifs are also suitable. For example, the motif comprising a CpG site may comprise CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6). The term “FLEXseq” does not necessarily indicate that the nuclease cuts at the CCGG cut site, but rather the nuclease may cut at any suitable motif containing a CpG site including any of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. FIG. IB shows the overall workflow and analyses. Specimen inputs are sheared genomic (g) DNA from cells, cfDNA from body fluid and plasma, and fragmented DNA from FFPE tissues. Analyses include copy number detection for malignant aneuploidy, machine learning (ML) classification, and deconvolution of cell types.
[0040] FIGS. 2A-2H show FLEXseq enrichment improves coverage at low methylation bias. FIG. 2A shows theoretical coverage of cell type markers (top 100 markers of each cell type) that overlap with different methods. FIG. 2B shows theoretical coverage of CpGs from CCGG flanks overlapped by different markers. The CNS classifier indicates the 32,000 most differentiated probes from the CNS tumor and control references, and the TCGA classifier indicates the 60,000 most differentiated probes from TCGA tumor and control references. FIG.2C shows Correlation of methylation beta values between inter-run replicates (Pearson’s r = 0.98). FIG. 2D shows correlation of beta values between FLEXseq and the WGBS gold standard (Pearson’s r = 0.97). FIG. 2E shows correlation between EPIC array and FLEXseq (Pearson’s r = 0.96). FIG. 2F shows CfDNA input titrations sourced from a pleural fluid sample (BF3713) and assayed by FLEXseq. All non-10 ng inputs were correlated with 10 ng input (Pearson’s r > 0.89). FIG. 2G shows enrichment of CCGG flanking 50 bp regions, the top 250 cell type-specific markers acrossdifferent cell types, and all CpGs with coverage >10x for WGBS (blue) and FLEXseq (red). Fold enrichment is based on 5Gbp of WGBS and FLEXseq. FIG. 2H shows a representative area compares equal base pairs (30 Gbp) between WGBS and FLEXseq. Unmethylated (blue) and methylated (red) CpGs are highlighted. Gene and cell type marker tracks are shown at the bottom.
[0041] FIGS. 3A-3D shows in silico coverage of CpGs across genomic regions by different methods. FIG. 3A shows CpG percentages across genomic regions by different methylation detection methods. CCGG flanks arc the 50-300 bp regions flanking CCGG motifs. RRBS regions are two CCGG motifs within a maximum distance of 50-300 bp. Cell type markers are defined across 39 purified cell types from a human DNA methylation atlas. MeDIP-seq markers are defined as CpGs in hypermethylated regions derived from the above cell type- specific markers. The ranges between dashed lines indicate the 50%, 75%, and 95% distribution of CpGs, respectively. CDS, coding sequence; UTR, untranslated region, including 5’ and 3’ UTR. FIG. 3B, 3C, and 3D show the theoretical distance from the CCGG motifs for all CpGs covered by EPIC array, CCGG flanks, and cell type markers.
[0042] FIG. 4A, FIG. 4B, and FIG. 4C show schematics of an exemplary process performed in accordance with the methods described herein, including exemplary sequences for adapters, primers, and target nucleic acids. FIGS. 5A-5E show correlations of methylation beta values across different methods, sample types, and DNA inputs. FIG. 5A shows FLEXseq is highly correlated with the EPIC array, as observed in K562 and three FFPE tumor tissue samples (TF211, TF640, and TF642). FIG. 5B shows density plots of beta values of the same four samples in (a). FIG. 5C shows correlation between WGBS, FLEXseq, and RRBS / XRBS of K562 gDNA. All correlations are based on individual CpG methylation beta values. FIG. 5D shows FFPE tissue (TF640) DNA input titrations assayed by FLEXseq. Inputs of 1, 5, and 50 ng were correlated with 10 ng input (Pearson’s r >= 0.90), though the number of covered CpGs decreased from 5 million with 10 ng to 0.2 million with 1 ng. FIG. 5E shows K562 gDNA input titrations assayed by FLEXseq. Inputs of 1 ng and 100 ng were correlated with 10 ng input (Pearson’s r = 0.98). FIG. 5F shows K562 gDNA input of EPICv2 array at 100 ng was less correlated with the recommended input of 250 ng (Pearson’s r = 0.82).
[0043] FIG. 6A-6C show classification of cancer cell line titrations using microarray-based ML classifiers. FIG. 6A is a schematic of microarray-based ME classification. FIG. 6B shows t-SNEdimensionality reduction of the 2,508 TCGA tumor and control references across 45 classes. FIG. 6C shows / -SNE dimensionality reduction of tumor cell line titrations into immune cell DNA. Cell lines are MCF-7 (BRCA), CL-40 (COAD), and DBTRG (GBM).
[0044] FIGS. 7A-7F show computational inference of methylation haplotypes. FIG. 7A is a schematic for inferring methylation beta values through haplotypes between disparate datasets. The left panel illustrates that individual CpG sites are matched one-to-one between array and sequencing data, leaving some CpGs unmatched and unused, while the right panel presents an approach that leverages haplotypes (from a human DNA methylation atlas of purified benign cell types) to synchronize CpG behavior across the genome, maximizing molecular depth and expanding CpG coverage. FIG. 7B shows a theoretical example of inferring CpG methylation values through haplotypes. Four CpG sites measured by FLEXseq were intersected with one haplotype that includes six CpGs, two of which were array probes. The haplotype’s beta value is calculated as 0.27 by integrating all sequencing data covering this segment. Subsequently, the beta value of CpG-4 was changed from 0.25 to 0.27, and the beta value of CpG-5 was inferred as 0.27 from the haplotype segment. FIG. 7C shows original beta values are highly concordant with the average beta values after intersecting with the haplotypes. FIG. 7D shows FLEXseq intersected beta values are concordant with the EPIC array data. All benchmarks were performed using the K562 cell line. FIG. 7E shows frequencies and percentages of all CpGs with 5x, lOx, 20x, and 30x coverage by FLEXseq. FIG. 7F shows frequencies and percentages of CpG overlap between FLEXseq and tumor classifiers, with coverage >5x and >10x. The red bar indicates increased CpGs after the intersection, and the blue bar indicates the original CpGs. The dot indicates the percentage based on the right y-axis label.
[0045] FIG. 8 is an Illustrative flowchart of development of the random forest classifier. New classifiers were trained using CpG markers that overlap between FLEXseq and array reference datasets. The training samples were used to train the random forest model and conducted score calibration using a ridge regression model. Then the test samples were fit to the random forest model to get the predicted scores. The predicted scores were translated to meaningful class probabilities in the ridge regression.
[0046] FIG. 9 shows performance of the TCGA ML classifier (2,508 reference samples, with 41 subclasses).
[0047] FIG. 10 shows performance of the CNS ML classifier (2,801 reference samples, with 91 subclasses).
[0048] FIGS. 11A-11E show classification of body fluid and FFPE tissue samples. FIG. 11A shows a summary of copy number analysis and ML classification using the TCGA and CNS ML classifiers of 79 non-CSF samples. ML classifications were categorized as ‘Matched’ (confirmation of diagnosis with classifier score >0.3), ‘Misleading profile’ (mismatched cancer with score >0.3), or ‘Indeterminate’ (positive tumors with score <0.3 or predicted as control). FIG. 11B shows sample sources. BF - body fluid; ABDO - abdominal / peritoneal / ascitic fluid; LN - lymph node; OVAR - ovarian cyst fluid; PELV - pelvic wash fluid; PLEU - pleural fluid. FIG. 11C shows a classifier score distributions of FLEXseq samples predicted as tumors, including 25 non-CSF body fluids, 20 FFPE brain tissues, and six gDNA samples of tumor cell line titrations from the TCGA ML classifier. FIG. 1 ID and FIG. 1 IE show cases with misleading ML classification (BF3090 and TF112) classified to the gold standard in the Z-SNE;HNSC / squamous (head and neck squamous cell carcinoma and squamous family) and MCF MB SHH (medulloblastoma, subclass SHH). The black open triangle indicates the FLEXseq sample.
[0049] FIGS. 12A-12B show tumor classification using FLEXseq ML classifier. FIG. 12A is a schematic of FLEXseq classification based only on FLEXseq references. FIG. 12B shows classifier scores for the microarray- and FLEXseq-based classifier across 12 matching tumors and six control samples. The grey line in the box plot connects the same sample in both groups. FIG. 12C shows t-SNE dimensionality reduction of 57 FLEXseq samples colored by pathological classes. FIG. 12D shows t-SNE dimensionality reduction colored by sample types (FFPE or body fluid cfDNA).
[0050] FIGS. 13A-13D deconvolution of in silico titrations, physical DNA titrations, and plasma. FIG. 13A shows cell type proportions deconvoluted from in silico mixtures using the WGBS references (WGBS reads intersected with CCGG flanks). CpG- (blue) and fragment-level deconvolution methods (red) are shown. Error bars indicate the SD at each titration level. FIG. 13B shows deconvolution of gDNA of B cells, monocytes, neutrophils, and T cells titrated into mixtures of three immune cell types, respectively. FIG. 13C shows fragment-level deconvolution of BRCA (tumor cell-of-origin, breast luminal epithelium), COAD (tumor cell-of-origin, colon epithelium), and GBM (approximate tumor cell-of-origin lineage, neuron and oligodendrocyte) cell line DNA titrated into mixtures of the same four immune cell types in (a). The dark blue lineindicates deconvolution with references exclusive to the approximate tumor cell type and background immune cells. The dark red line indicates deconvolution with a broader set of reference cell types encompassing common brain metastases. FIG. 13D shows deconvolution of healthy plasma samples (WBGS2), WBGS2 intersected with CCGG flanks to simulate FLEXseq data, and FLEXseq from a healthy donor (P2).
[0051] FIG. 14A-14B show establishing deconvolution classifier based on deconvoluted cell type proportions. FIG. 14A is a schematic of deconvolution classification. FIG. 14B shows Z- scores of the cell-of-origin of three common CNS tumors: 20 LUADs (vs. 36 non-LUADs), 21 DLBCs (vs. 35 non-DLBCs), and six BRCAs (vs. 50 non-BRCAs). The boxplot x-axis indicates the target (solid dots) / non-target tumors (solid triangles), and the x-axis in the scatter plots indicates the z-score rankings of the target and non-target tumors. The black dashed lines are at z-scores of 2. The ‘x’ indicates tumor samples where the cell type is not ranked first. LU AD, lung adenocarcinoma; DLBC, B-cell lymphoma; BRCA, breast carcinoma.
[0052] FIGS. 15A-15D show deconvolution of CSF samples. FIG. 15 A shows deconvoluted cell type proportions in CSF negative controls (n = 50). Five outlier controls had higher T-cell percentages. FIG. 15B shows Z-scores of the cell-of-origin in less common CNS metastases and primary CNS tumors (labeled Brain) in 56 CSF tumor samples. FIG. 15C shows Z-scores of all 20 deconvoluted cell types across 56 CSF tumor samples as normalized by the negative controls. FIG. 15D shows Z-scores of all 20 cell types, but normalized against non-target tumors.
[0053] FIG. 16A-16E show a CSF case-control study. FIG. 16A shows a summary of ML (random forest model), deconvolution, and composite classification and copy number analysis of 106 CSF cfDNA samples by FLEXseq. The triangle indicates CNA-positive results based on WGS. Five references used in the ML classifier were excluded due to high percentages of T cells. FIG. 16B shows a schematic of the workflow using a representative CSF case (BF3741) that leads to copy number and classification analyses. FIG. 1C shows the composite classification workflow integrates the copy number analysis (evaluating tumor presence and purity) with the ML and deconvolution classification. FIG. 16D shows ROC curves for ML (blue dotted line), deconvolution (red dashed line), and composite (black solid line) classification for LU AD, DLBC, and BRCA. The curves only include samples analyzed by all three classifiers, excluding 53 CNA-negatives and seven indeterminates from deconvolution. FIG. 16E shows a confusion matrix for all 106 CSF samples based on the composite classification workflow in (c).#The case is classified with the pathological group in the t-SNE plot and misclassified in the random forest model. * One indeterminate case was not included under each tumor type.
[0054] FIG. 17A-17D shows CfDNA inputs of CSF and non-CSF body fluids. FIG. 17A shows the plasma cfDNA input titration curve used to calculate sample input. The curve is based on a constant quantity of lambda phage DNA spiked into every sample before library preparation. The titration is calibrated by a linear pre-quantified healthy plasma sample (P2) (R2= 0.99). FIG. 17B shows the correlation between CSF cfDNA input (ng) extracted at Stanford (n = 152) and Ct from qPCR. Samples that have multiple library preparations are represented separately. Dots are shifted within ±0.25 Ct of the actual Ct value visually to see all data points without overlaps. The equation for the best fit log-linear regression (R2= 0.84) is shown. The Y-axis is logged. FIG. 17c shows CfDNA concentrations calculated based on the cfDNA input titration curve in (a).The 92 CSF, 10 FNA saline wash fluid, nine abdominal / peritoneal / ascitic fluid (ABDO), and 18 pleural fluid (PLEU) samples were shown, which were extracted at Stanford and met quality metrics (Ct <14 and >30 million deduplicated paired-end reads). CSF cfDNA concentration is lower than abdominal (adjusted p = 0.008) and pleural fluids (adjusted p < 0.001). Only significant p-values are labeled. The Y-axis is logged. FIG. 17D shows CSF cfDNA concentrations of all CSF samples extracted at Stanford (n = 117, no quality metric filtering) based on diagnostic categories: autoimmune / autoinflammatory (n - 25), cancer (n - 59), infection (n ~ 13), other (n = 7), and unknown cause (n = 13). Inputs are highest in the infection group, followed by the cancer and autoimmune / autoinflammatory group (infection vs. cancer, adjusted p = 0.04; cancer vs. autoimmune / autoinflammatory. adjusted p = 0.006). The Y-axis is logged.
[0055] FIG. 18 shows ML Classification and deconvolution of CSF cfDNA. a, The t-SNE plot (top) and copy number plot (top right) of a misclassified ease (BF3027) that is classified with the matching pathological group (IDH-mutant glioma) and has the lp / 19q-codeletion. The t-SNE plots (bottom) of six cases with indeterminate ML classifications that are classified with their matching pathological group, b, Scatter plot between tumor purity and z-scores of CSF CNA- positives (n = 53). The correlation is weak when tumor purity is above 50% (p > 0.05).
[0056] FIG. 19 shows the on-target rate of using the Taql-v2 nuclease (NEB, part number R0149) targeting the 'TCGA' motif rather than Mspl. Cutsmart buffer was used. The on-target rate was estimated by dividing reads starting with CGA or TGA by the total reads on the side cut by the nuclease. Each on-target read is guaranteed to contain CpG methylation data based on the first position.DETAILED DESCRIPTION1. Definitions
[0057] To facilitate an understanding of the present technology, a number of terms and phrases are defined below. Additional definitions are set forth throughout the detailed description.
[0058] The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. The singular forms “a,” “and” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of’ and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.
[0059] For the recitation of numeric ranges herein, each intervening number therebetween with the same degree of precision is explicitly contemplated. For example, for the range of 6-9, the numbers 7 and 8 arc contemplated in addition to 6 and 9, and for the range 6.0-7.0, the number 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated.
[0060] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. For example, any nomenclature used in connection with, and techniques of, cell and tissue culture, biochemistry, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those that are well known and commonly used in the art. The meaning and scope of the terms should be clear; in theevent, however, of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.
[0061] The term "amino acid" refers to natural amino acids, unnatural amino acids, and amino acid analogs, all in their D and L stereoisomers, unless otherwise indicated, if their structures allow such stereoisomeric forms.
[0062] The terms “complementary” and “complementarity” refer to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick base-paring or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences can be indicated by the percentage of nucleotides in a nucleic acid sequence which can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary). Two nucleic acid sequences are “perfectly complementary” if all the contiguous nucleotides of a nucleic acid sequence will hydrogen bond with the same number of contiguous nucleotides in a second nucleic acid sequence. Two nucleic acid sequences are “substantially complementary” if the degree of complementarity between the two nucleic acid sequences is at least 60% (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%. 97%, 98%, 99%, or 100%) over a region of at least 8 nucleotides (e.g., 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides), or if the two nucleic acid sequences hybridize under at least moderate, preferably high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37° C in a solution comprising 20% formamide, 5xSSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5xDenhardt’s solution, 10% dextran sulfate, and 20 mg / ml denatured sheared salmon sperm DNA, followed by washing the filters in IxSSC at about 37-50° C., or substantially similar conditions, e.g., the moderately stringent conditions described in Sambrook et al., infra. High stringency conditions are conditions that use, for example (1) low ionic strength and high temperature for washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50° C, (2) employ a denaturing agent during hybridization, such as formamide, for example, 50% (v / v) formamide with 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42° C., or (3) employ 50% formamide, 5xSSC (0.75 M NaCl, 0.075 M sodiumcitrate), 50 mM sodium phosphate (pH 6.8), 0.1 % sodium pyrophosphate, 5xDenhardt’s solution, sonicated salmon sperm DNA (50 pg / ml), 0.1% SDS, and 10% dextran sulfate at 42° C., with washes at (i) 42° C. in 0.2xSSC, (ii) 55° C. in 50% formamide, and (iii) 55° C. in O.lxSSC (preferably in combination with EDTA). Additional details and an explanation of stringency of hybridization reactions are provided in, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, N.Y. (2001); and Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and John Wiley & Sons, New York (1994).
[0063] As used herein, a “nucleic acid” or a “nucleic acid sequence” refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine (C), thymine (T), and uracil (U), and adenine (A) and guanine (G), respectively. The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogenous or homogenous in composition and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single- stranded or double- stranded form, including homoduplex, heteroduplex, and hybrid states. In some embodiments, a nucleic acid or nucleic acid sequence comprises other kinds of nucleic acid structures such as, for instance, a DNA / RNA helix, peptide nucleic acid (PNA), morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14): 4503-4510 (2002)) and U.S. Pat. No. 5,034,506, incorporated herein by reference), locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci.U.S.A., 97: 5633-5638 (2000), incorporated herein by reference), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: 8595-8602 (2000), incorporated herein by reference), and / or a ribozyme. Hence, the term “nucleic acid” or “nucleic acid sequence” may also encompass a chain comprising non-natural nucleotides, modified nucleotides, and / or non- nucleotide building blocks that can exhibit the same function as natural nucleotides (i.e., “nucleotide analogs”); further, the term “nucleic acid sequence” as used herein refers to an oligonucleotide, nucleotide or polynucleotide, and fragments or portions thereof, and to DNA or RNA of genomic or synthetic origin, which may be single or double-stranded, and represent the sense or antisense strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and“oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, cither dcoxyribonuclcotidcs or ribonucleotides, or analogs thereof.
[0064] A “peptide” or “polypeptide” is a linked sequence of two or more amino acids linked by peptide bonds. The peptide or polypeptide can be natural, synthetic, or a modification or combination of natural and synthetic. Polypeptides include proteins such as binding proteins, receptors, and antibodies. The proteins may be modified by the addition of sugars, lipids or other moieties not included in the amino acid chain. The terms “polypeptide” and “protein,” are used interchangeably herein.
[0065] As used herein, the term “percent sequence identity” refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid sequence, or amino acids in an amino acid sequence, that is identical with the corresponding nucleotides or amino acids in a reference sequence after aligning the two sequences and introducing gaps, if necessary, to achieve the maximum percent identity. Hence, in case a nucleic acid according to the technology is longer than a reference sequence, additional nucleotides in the nucleic acid, that do not align with the reference sequence, are not taken into account for determining sequence identity. Methods and computer programs for alignment are well known in the art, including BLAST, Align 2, and FASTA.2. Methods
[0066] In some aspects, provided herein are methods. In some embodiments, provided herein are methods of preparing nucleic acids for analysis. In some embodiments, the methods provide herein are used to prepare nucleic acids in a sample for analyses such as amplification (e.g. PCR), sequencing, or methylation analyses. In some embodiments, provided herein are methods of classifying methylation status of a sample.
[0067] The methods provided herein involve contacting a sample with a first adapter that ligates to target nucleic acids in the sample, cutting the adapter-ligated nucleic acids with a nuclease that cuts at a motif containing a CpG site, if present, and adding a second adapter to the sample that ligates (e.g. at the cute site) to the subpopulation of adapted-ligated nucleic acid fragments resulting from cutting. Accordingly, the methods provided herein enrich for nucleic acids cut at a motif containing a CpG site, which provide valuable information from the cute site flanking region regarding methylation status of a sample.
[0068] In some embodiments, the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6). In some embodiments, the nuclease cuts at a 5’ location to the CpG site in the motif. The nuclease cutting at a 5’ location to the CpG site in the motif produces a sequence referred to herein as a “flanking sequence” or a “cut site flanking sequence”, which contains the CpG site. In some embodiments, the nuclease cut between the first and second residues of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. In some embodiments, the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
[0069] The methods described herein produce nucleic acids containing a cut site flanking sequence covalently bonded at one end to a portion of the first adapter and covalently bonded at a second end to a portion of the second adapter. In some embodiments, the first adapter and the second adapter each comprise a primer landing site, which facilitates amplification and / or sequencing of the nucleic acids. In some embodiments, the methods comprise performing a conversion step to selectively convert cytosine residues in the sample depending on their methylation status to a different nucleic acid besides cytosine, and generation a methylation profile for the sample by sequencing. In some embodiments, the methods comprise sequencing the nucleic acids containing the cut site flanking sequence covalently bonded to the first adapter and to the second adapter by nanopore sequencing. In embodiments wherein nanopore sequencing is used, no enzymatic conversion step to selectively convert cytosine residues in the sample depending on their methylation status is needed. In embodiments wherein nanopore sequencing is used, the adapters need not contain primer landing sites. For example, in some embodiments the first adapter ligated to the target double- stranded nucleic acid in the sample and / or the second adapter ligated to the population of nucleic acid fragments (e.g. produced by the nuclease that cuts at a motif containing a CpG site in the target double-stranded nucleic acids in the sample) do not contain a primer landing site and instead contain a motor protein which helps to pass a given nucleic acid through a given nanopore of a nanopore sequencing device / system.
[0070] In some embodiments, the method comprises ligating a first adapter to target doublestranded nucleic acids in a sample. In some embodiments, the first adapter comprises a primer binding strand containing a first primer landing site. The primer binding strand of the firstadapter ligates to a first strand of the target nucleic acids in the sample. This first strand of the target double-stranded nucleic acid in the sample is subsequently cut by a nuclease that cuts at the motif containing a CpG site, if present in the first strand, thereby producing a nucleic acid fragment. This nucleic acid fragment is subsequently ligated to a second adapter containing a second primer landing site, and is therefore ultimately able to be substantially amplified due to the presence of both the first primer landing site and the second primer landing site. In contrast, a second strand of the target double- stranded nucleic acids in the (e.g. the strand complementary to the first strand) is ultimately not able to be substantially amplified, for various possible reasons. Accordingly, the methods provided herein facilitate selective substantial amplification and sequencing of a first strand, but not a second strand, of target double- stranded nucleic acids in a sample. For example, in some embodiments the second strand of the target double-stranded nucleic acids is ligated to a blocking strand of the first adapter, thereby ultimately preventing substantial amplification of the second strand as described in more detail below. As another example, in some embodiments the second strand of the target nucleic acids does not ligate to the first adapter, and is therefore not able to be substantially amplified due to the absence of at least a first primer landing site.
[0071] In some embodiments, the first adapter comprises a double-stranded end portion. In some embodiments, the first adapter comprises a double- stranded end portion comprising (i) a primer binding strand containing a first primer landing site, and (ii) a blocking strand. The double stranded end portion comprises 5’ end of a DNA single strand and a 3’ end of a DNA single strand of the adapter that is substantially double-stranded. The adapter is not necessarily perfectly double-stranded at this end portion. For example, a double-stranded end portion may have an overhang (also referred to as a jagged end or a sticky end). In some embodiments, the double-stranded end portion is a blunt end (e.g. lacks an overhang). In some embodiments, the 5’ end of the adapter contains the first primer landing site. In some embodiments, a 5’ end of the primer binding strand ligates to a 3’ end of a first strand of the target double-stranded nucleic acid.
[0072] In some embodiments, the first adapter comprises a primer binding strand. In some embodiments, the double stranded end portion comprises a primer binding strand. As used herein, the term “primer binding strand” indicates that the strand contains a primer landing site and thus facilitates downstream amplification of the nucleic acid containing the primer landingsite in the presence of the primer. The landing site may be designed to facilitate direct or indirect primer binding. In some embodiments, the landing site facilitates direct primer binding. Direct primer binding indicates that a first primer is added to the sample and directly binds to the primer landing site on the primer binding strand of the first adapter. In such embodiments, the primer landing site of the first adapter is complementary to the first primer. In other embodiments, the landing site facilitates indirect primer binding. Indirect primer binding indicates that a second primer added to the sample binds to second (e.g. separate) primer landing site (e.g., on the second adapter) and during subsequent extension of the primer-bound nucleic acid, the primer landing site for a first primer is created. In such embodiments, the primer landing site of the first adapter is the reverse complement of the primer sequence for the first primer. In some embodiments, the landing site, if it has cytosines, have nucleic acid modifications that will resist conversion to another nucleic acid upon methylation conversion.
[0073] In some embodiments, the double stranded end portion further comprises a blocking strand. As used herein, the term “blocking strand” indicates that the strand contains features which prevent substantial amplification of the strand or lacks essential features for substantial amplification of the strand (e.g. during PCR, during sequencing). Accordingly, the blocking strand is not substantially amplified in downstream PCR or sequencing steps. The mechanism by which the blocking strand is not substantially amplified or sequenced can be active (e.g. the blocking strand may contain a moiety that prevents substantial amplification and / or sequencing) or passive (e.g. the blocking strand may lack a moiety that would otherwise facilitate amplification and / or sequencing). In some embodiments, a blocking strand contains a moiety to prevents ligation of the blocking strand to the target nucleic acid yet allows for the ligation of the primer landing strand to the target nucleic acid. In some embodiments, a blocking strand contains a moiety that prevents attachment of primers to the blocking strand for subsequent amplification and / or sequencing steps. In some embodiments, a blocking strand contains a moiety that interferes with the ability of a polymerase during amplification to transverse across the molecule. In some embodiments, a blocking strand contains a moiety that allows the strand to be disconnected. In some embodiments, a blocking strand contains a moiety allows for at least a portion of the blocking strand to be degraded. In some embodiments, a blocking strand lacks a primer landing site (e.g. lacks a direct primer landing site and / or lacks an indirect primer landingsite). In some embodiments, the blocking strand comprises a nucleotide that is modified during a conversion step such that primer landing to the blocking strand is disrupted upon conversion.In other embodiments, a single-stranded ligation of the primer binding strand of the first adapter to a target double stranded nucleic acid occurs. In such embodiments, the primer binding strand of the first adapter ligates to a first strand of the target double- stranded nucleic acid, and the second strand of the target nucleic acid does not ligate to the first adapter. In such embodiments, the first adapter comprises a primer binding strand and does not comprise a blocking strand.
[0074] In some embodiments, the first adapter is substantially double-stranded in its entirety. In some embodiments, the first adapter is a Y-shaped or a U-shaped adapter. For example, in some embodiments the first adapter is a Y-shaped or a U-shaped adapter comprising a substantially double-stranded end containing the primer binding strand and the blocking strand, and a singlestranded portion at the non-ligating end of the adapter. In some embodiments, a Y-shaped or U- shaped adapter has non-complementary (therefore non-double stranded) strands at the nonligating end of the adapter.
[0075] The primer binding strand of the first adapter may comprise any suitable sequence. Generally, the primer binding strand of the first adapter comprises a first primer landing site having a sequence complementary to a first primer sequence, and the second primer landing site (e.g. of the second adapter) has a sequence is complementary to a second primer sequence. The sequences of the first primer, first primer landing site, second primer and the second primer landing site may vary so long as sufficient complementarity between the first primer and the first primer landing site, and the second primer and the second primary landing site, is achieved for exponential amplification. In some embodiments, the primer binding strand is not originally present, but through methylation conversion, can be created by swapping nucleic acids (e.g. cytosine for uracil). Here, the original sequence that becomes the primer binding strand is also regarded indirectly as a form of the primer binding strand.
[0076] In some embodiments, the primer binding strand of the first adapter comprises the sequence 5’GATCGGAAGAGCGTCGT-3’. (SEQ ID NO: 7). In some embodiments, the primer binding strand of the first adapter comprises a sequence having at least 88% sequence identity with SEQ ID NO: 7. The blocking strand of the first adapter is substantially complementary to (e.g. is substantially the reverse complement of) the primer binding strand of the first adapter, thus facilitating the formation of the double- stranded end portion of the adapter. In someembodiments, the blocking strand of the first adapter comprises the sequence 5’- ACGCTCTTCCGATCT -3’ (SEQ ID NO: 8). In some embodiments, the blocking strand of the first adapter comprises a sequence having at least 88% sequence identity with SEQ ID NO: 8. In some embodiments, the primer binding strand of the first adapter comprises a sequence having at least 88% sequence identity with SEQ ID NO: 7 and the blocking strand of the first adapter comprises a sequence having at least 88% sequence identity with SEQ ID NO: 8. In some embodiments, the primer binding strand of the first adapter comprises the sequence of SEQ ID NO: 7 and the blocking strand of the first adapter comprises the sequence of SEQ ID NO: 8.
[0077] In some embodiments, the primer binding strand of the first adapter comprises the sequence 5’-GATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT-3’ (SEQ ID NO: 9). In some embodiments, the primer binding strand of the first adapter comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 9. In some embodiments, the cytosines (C) in the primer binding strand sequence are resistant to methylation conversion (e.g. they are methylated). In some embodiments, the blocking strand of the first adapter comprises the sequence 5’- ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ ID NO: 10). In some embodiments, the blocking strand of the first adapter comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 10. In some embodiments, the primer binding strand of the first adapter comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 9 and the blocking strand of the first adapter comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 10. In some embodiments, the primer binding strand comprises the sequence of SEQ ID NO: 9 and the blocking strand comprises the sequence of SEQ ID NO: 10.
[0078] In some embodiments, the 5’ end of the primer binding strand of the first adapter covalently attaches to the 3’ end of a target nucleic acid in the sample. In some embodiments, the5’ end of the primer binding strand of the first adapter comprises a moiety that facilitates covalent attachment of the primer binding strand to target DNA sequences during ligation. For example, in some embodiments the 5’ end of the primer binding sequence comprises a phosphate. In some embodiments, a 5’ phosphate is not added to the sample prior to ligation to prevent covalent attachment of the first adapter’ s blocking strand to the target nucleic acids in the sample. In some embodiments, 5' phosphates are removed from the sample prior to ligation to prevent covalent attachment of the first adapter’ s blocking strand to the target nucleic acids in the sample. In some embodiments, the first adapter sequence is partially modified such that the cytosine residues present in at least the primer binding strand are not converted to other residues during subsequent conversion (e.g. enzymatic conversion, bisulfite conversion) steps. For example, in some embodiments the cytosine residues in the primer binding strand are methylated or hydroxymethylated to resist conversion (e.g. to uracil, to thymine) upon addition of a converting enzyme. Whether the residues are methylated or hydroxymethylated depends on the specific enzyme(s) to be added during a subsequent conversion step.
[0079] In some embodiments, the blocking strand lacks a moiety (e.g. a phosphate group) that would otherwise facilitate covalent attachment of the blocking strand to a target nucleic acid in the sample.
[0080] In some embodiments, the first adapter sequence is partially modified such that the cytosine residues present in at least the primer binding strand are not converted to other residues during subsequent conversion (e.g. enzymatic conversion, bisulfite conversion) steps. For example, in some embodiments the cytosine residues in the primer binding strand are methylated or hydroxymethylated to resist conversion (e.g. to uracil, to thymine) upon addition of a converting enzyme or chemical. Whether the residues are changed due to its methylated or hydroxymethylated status depends on the specific enzyme(s) to be added during a subsequent conversion step. In some embodiments, cytosine residues in the blocking strand are not modified (e.g. not methylated, not hydroxymethylated). The protection of one adapter strand and not the complementary strand is referred to herein as a “hemimethylated” adapter. A hemimethylated adapter refers to an adapter wherein cytosine residues in the binding strand are methylated or hydroxymethylated, and cytosine residues in the blocking strand are not.
[0081] In some embodiments, the primer binding strand ligates (e.g. covalently) to a strand of the double- stranded target nucleic acid in the sample. In some embodiments, the 5’ portion ofthe primer binding strand ligates to a 3’ end of the target nucleic acid. In some embodiments, the primer binding strand and the blocking strand each ligate a strand of the double-stranded target nucleic acid in the sample. For example, in some embodiments the primer binding strand of the first adapter ligates to a first strand of the double-stranded target nucleic acid, and the blocking strand of the first adapter ligates to a second strand of the double- stranded target nucleic acid in the sample. In such embodiments, a population of substantially double-stranded target nucleic acids are produced wherein one strand is bound to the primer binding strand and the other strand is bound to the blocking strand. As described above, only the strand of the double- stranded nucleic acid that is bound to the primer binding strand of the first adapter is ultimately able to be amplified and / or sequenced, whereas the strand of the double-stranded nucleic acid bound to the blocking strand is prevented from being substantially amplified and / or sequenced, either actively (e.g. by the blocking strand comprising a moiety which prevents primer binding) or passively (e.g. by the blocking strand lacking a primer landing site for a primer used during amplification and / or sequencing). In some embodiments, the primer binding strand binds to a strand of the double-stranded target nucleic acid in the sample, and the blocking strand does not bind to target nucleic acid. In some embodiments, the primer binding strand binds to a strand of the doublestranded target nucleic acid in the sample, and the other strand of the double- stranded target nucleic acid does not bind to either the primer binding strand or the blocking strand of the first adapter and is thus ultimately not substantially amplified or sequenced due to lacking the first primer landing site. In some embodiments, the first adapter comprises a primer binding strand that binds to a first strand of the double- stranded target nucleic acid in the sample, and the second strand of the double-stranded target nucleic acid does not bind to the first adapter and is thus ultimately not substantially amplified or sequenced. In some embodiments, following ligating the first adapter to the target nucleic acid, the method comprises adding a nuclease that cuts at a motif containing a CpG site, if present in the target nucleic acids in the sample. In some embodiments, the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6). In some embodiments, the nuclease cuts at a 5’ location to the CpG site in the motif. For example, in some embodiments the nuclease cuts at a 5’ location to the CpG site in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. In some embodiments, the nuclease cut between the first and second residues of SEQ ID NO:1 , SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. The cute site flanking sequence contains the CpG site. In some embodiments, the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
[0082]
[0083] In some embodiments, the motif containing a CpG site comprises CCGG (SEQ ID NO: 1). In some embodiments, the nuclease cuts between the first C and second C in the CCGG sequence. Accordingly, in such embodiments the nuclease produces a first nucleic acid fragment wherein the CCGG cut site flanking sequence begins with CGG.
[0084] Any suitable nuclease may be used to cut at the motif containing a CpG site. In some embodiments, the nuclease is MspI, Taql-v2, AccII, Acil, AspLEI, HPall, HpyCH4IV. In some embodiments, the nuclease is a type II nuclease. In some embodiments, the nuclease has impaired cleavage when the CpG site at the motif is methylated. In some embodiments, the nuclease is MspI and the motif containing a CpG site comprises CCGG (SEQ ID NO: 1). In some embodiments, the nuclease is Hpall and the motif containing a CpG site comprises CCGG (SEQ ID NO: 1). In some embodiments the nuclease is Taql-v2 and the motif containing a CpG site comprises TCGA (SEQ ID NO: 2). In some embodiments, the nuclease is AccII and the motif containing a CpG site comprises CGCG (SEQ ID NO: 3). In some embodiments the nuclease is Acil and the motif containing a CpG site comprises CCGC (SEQ ID NO: 4). In some embodiments, the nuclease is AspLEI and the motif containing a CpG site comprises GCGC (SEQ ID NO: 5). In some embodiments, the nuclease is HpyCH4IV and the motif containing a CpG site comprises ACGT (SEQ ID NO: 6).
[0085] In some embodiments, the nuclease is a CRISPR / Cas nuclease (e.g. CRISPR / Cas9, CRISPR / Casl2, CRISPR / Cas 13, etc.). In some embodiments, the nuclease is a CRISPR / nuclease highly specific for targeted sites. CRISPR-Cas9 nuclease conveniently requires only a subset of the CCGG motif — the ‘GG’ — on the PAM site to make a cut. In some embodiments, CRISPR / nuclease targets a plurality of sites, including ones that contain CCGG. Accordingly, Cas9 targeting is compatible with the methods described herein.
[0086] In some embodiments, residual phosphates are removed using a suitable phosphatase prior to adding the nuclease to the sample. For example, in some embodiments, the method incubates the sample with a suitable phosphatase to remove residual phosphates and subsequently adding the nuclease to the sample.
[0087] Cutting at the motif containing a CpG site produces a subpopulation of nucleic acid fragments ligated to the first adapter. In some embodiments, at least some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer landing strand of the first adapter. In some embodiments, some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer landing strand of the first adapter, and some nucleic acid fragments comprise a first end containing a cut site flanking sequence and a second end ligated to the blocking strand of the first adapter. For example, in some embodiments, the target double-stranded nucleic acid comprises a first strand that ligates to the primer binding strand of the first adapter and a second strand that ligates to the blocking strand of the first adapter. Accordingly, in some embodiments following cutting at the motif containing a CpG site in each strand, the first strand produces a fragment comprising a first end containing a cut site flanking sequence and a second end ligated to the primer landing strand of the first adapter, whereas the second strand produces a fragment comprising a first end containing a cut site flanking sequence and a second end ligated to the blocking strand of the first adapter.
[0088] In some embodiments, the target nucleic acid is double stranded, and the nuclease cuts a motif containing a CpG site in each strand in the target nucleic acid. This is referred to herein as a double-strand cut. In some embodiments, a double-strand cut creates a substantially doublestranded nucleic acid fragment. In some embodiments, one end of a first strand in the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter. In some embodiments, one end of a first strand in the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and one end of a second strand in the substantially double- stranded nucleic acid fragment is ligated to the blocking strand of the first adapter. In some embodiments, one end of a first strand in the substantially double- stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and one end of a second strand in the substantially double- stranded nucleic acid fragment is not ligated to the first adapter.
[0089] In some embodiments, the method further comprises adding a second adapter to the sample. The second adapter is designed to selectively ligate to nucleic acid fragments generated as a result of cleavage of target nucleic acids at motif containing a CpG site. Accordingly, the second adapter does not substantially ligate to uncleaved nucleic acids (e.g. nucleic acids in thesample lacking a motif containing a CpG site). Therefore, the methods provided herein produce a population of nucleic acid that do not contain a motif containing a CpG site and arc therefore ligated to the first adapter, but not ligated to the second adapter. These nucleic acids are not able to be substantially amplified, due to the absence of the second primer landing site, as described in more detail below.
[0090] In some embodiments, the second adapter contains a second primer landing site. In some embodiments, a portion of the second adapter comprising the second primer landing site selectively ligates to the nucleic acid fragments ligated the first adapter. Ligation of the second adapter to the nucleic acid fragments ligated to the first adapter forms a population of nucleic acid strands that are ligated (e.g. covalently bound) to both the first adapter and the second adapter. In some embodiments, ligation of the second adapter to the nucleic acid fragments ligated to the first adapter produces a population of nucleic acid strands wherein one end of the strand is ligated (e.g. covalently bound) to the primer binding strand of the first adapter (which contains the first primer landing site) and the other end of the strand is ligated (e.g. covalently bound) to a portion of the second adapter comprising the second primer landing site. The strand comprises the cut site flanking sequence, which is sandwiched between the primer binding strand of the first adapter and the portion of the second adapter comprising the second primer landing site. Accordingly, in some embodiments ligation of the second adapter to the nucleic acid fragments ligated to the first adapter produces a population of nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site. These nucleic acid strands, which contain the first primer landing site and the second primer landing site, can be subsequently amplified and / or sequenced using appropriate primers. Such strands are referred to herein as “amplifiable nucleic acid strands” or “amplifiable strands”.
[0091] In some embodiments, in addition to producing a population of amplifiable nucleic acid strands, attempted ligation of the second adapter produces a population of nucleic acid strands that do not contain at least one of the first primer landing site or the second primer landing site and are therefore not able to be substantially amplified and / or sequenced. Such strands are referred to herein as “non-amplifiable”, even though a low degree of linear amplification can occur with a single primer landing site. For example, in some embodiments ligation of the second adapter to the nucleic acid fragments ligated to the first adapter produces a population of amplifiable nucleic acid strands (e.g. strands wherein one end of the strand is ligated to theprimer binding strand of the first adapter and the other end of the strand is ligated to a portion of the second adapter comprising the second primer landing site, as described above) and a population of non-amplifiable nucleic acid strands wherein one end of the strand is ligated to the blocking strand of the first adapter and the other end of the strand is ligated to a portion of the second adapter comprising the second primer landing site. As another example, in some embodiments ligation of the second adapter to the nucleic acid fragments ligated to the first adapter produces a population of amplifiable nucleic acid strands (e.g. strands wherein one end of the strand is ligated to the primer binding strand of the first adapter and the other end of the strand is ligated to a portion of the second adapter comprising the second primer landing site, as described above) and a population of non-amplifiable nucleic acid strands wherein one end of the strand is not ligated to either strand of the first adapter and the other end of the strand is ligated to a portion of the second adapter comprising the second primer landing site.
[0092] In some embodiments, the second adapter is double- stranded and is intended for double stranded ligation. In some embodiments, the adapter is intended for single stranded ligation. In some embodiments, the portion of the second adapter comprising a second primer landing site comprises the sequence 5’-GACTGGAGTTCAGACGTGTGCTCT TCCGATCT-3’. (SEQ ID NO: 12). In some embodiments, the portion of the second adapter comprising a second primer landing site comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 12. In some embodiments, the second adapter is a single-stranded adapter comprising a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 12. In some embodiments, the cytosines (C) in the sequence are resistant to methylation conversion (e.g. they are methylated). In some embodiments, the second adapter is a double- stranded adapter, wherein a first strand of the adapter comprises a sequence having at least 80% sequence identity (e.g. at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) with SEQ ID NO: 12 and the second strand is substantially complementary to (e.g. is substantially the reverse complement of) the sequence of the first strand. Accordingly, in some embodiments the second adapter is a double- stranded adapterwherein the first strand of the adapter comprises a sequence having at least 80% sequence identity with SEQ ID NO: 12 and the second strand of the adapter comprises a sequence having at least 80% sequence identity with 5’-GATCGGAAGAGCACACGTCTGAACTCCAGTC-3’ (SEQ ID NO: 13).
[0093] In some embodiments, the first adapter and / or the second adapter does not comprise a primer landing site. For example, in some embodiments the methods herein involve nanopore sequencing to determine the sequence of a nucleic acid strand containing a cut site flanking sequence ligated at one end to the first adapter and ligated at the second end to the second adapter. In such embodiments, the first adapter and / or the second adapter need not contain a primer landing site, as no PCR amplification will be performed prior to nanopore sequencing. In lieu of the primer landing site, a motor protein may be utilized in the first adapter and / or in the second adapter. Such a motor protein guides a given nucleic acid strand through the nanopore, facilitating sequencing of the strand. In some embodiments, the method comprises ligating a first adapter to target double- stranded nucleic acids in the sample, wherein the first adapter is a nanopore sequencing adapter, thereby producing a population of nucleic acid strands ligated to the first adapter. In some embodiments, the method comprises adding a nuclease that cuts at a motif containing a CpG site, if present in the target double-stranded nucleic acids in the sample, thereby producing a subpopulation of nucleic acid fragments comprising a cut site flanking sequence and at least one end ligated to the first adapter. In some embodiments, the method further comprises adding a second adapter to the sample that ligates to the nucleic acid fragments ligated to the first adapter, thereby producing nucleic acid strands comprising the cut site flanking sequence ligated at one end to the first adapter and ligated at the other end to the second adapter. The first and / or second adapters may comprise motor proteins that guide the nucleic acid strands through a nanopore of a nanopore sequencing device. Such a nanopore is able to determine sequencing information about the strand, including methylation status. Accordingly, such embodiments need not include primer landing sites on the adapters, as no PCR amplification is necessary prior to nanopore sequencing. Moreover, as nanopore sequencing is able to determine methylation status of a given strand, no enzymatic conversion step to selectively convert residues is needed prior to nanopore sequencing.
[0094] In some embodiments, the cut site flanking sequence results from the nuclease cutting a single motif containing a CpG site in a target nucleic acid. In other words, in some embodimentsthe nuclease cuts a single motif containing a CpG site in the target nucleic acid, which produces a cut site flanking sequence only in the 5’ or the 3’ direction from the cut site. This is in contrast to methods (e.g. RRBS) wherein only target nucleic acid with two cut sites are produced by a nuclease, producing a flanking sequenced sandwiched between the two cut sites. Such methods, involving two cut sites, may produce fragments that are too small or are lost completely prior to subsequent sequencing analysis. In contrast, methods that produce a cut site flanking sequence from a nuclease cutting a single motif containing a CpG site in the target nucleic acid result in flanking sequences of sufficient length for downstream analyses (e.g. methylation analyses) without substantial loss. In some embodiments, the second adapter contains the second primer landing site and the first primer landing site on the opposite strand and can accommodate target nucleic acid with two cut sites in addition to target nucleic acid with one cut site.
[0095] The cut site flanking sequence may be any suitable number of nucleotides in length. In some embodiments, the cut site flanking sequence is 10-2000 bases in length. In some embodiments, the cut site flanking sequence comprises is at least 20 bases in length. In some embodiments, the cut site flanking sequence is 20 to 300 bases in length (e.g. about 20 to 300 bases, about 25 to 250 bases, about 30 to 200 bases, about 35 to about 175 bases, about 40 to about 150 bases, or about 50 to about 100 bases. In some embodiments, the cut site flanking sequence is about 50 base pairs in length. In some embodiments, the cut site flanking sequence is about 70 to about 150 base pairs in length. The results presented herein demonstrate that for all CCGG flanking sequence lengths (e.g. 50 bases, 100 bases, 150 bases, 200 bases, and 300 bases) investigated, the percent coverage of 320,838 known differential methylation markers from the human methylation atlas significantly outperformed the coverage for other methods, including 450k array, EPIC array, RRBS, and meDIPSeq.
[0096] In some embodiments, the methods provided herein further performing a conversion step to selectively convert cytosine residues in the sample depending on their methylation status to a different nucleic acid besides cytosine. The conversion step may be performed in between adding the nuclease and adding the second adapter to the sample, or may be performed after adding the nuclease and adding the second adapter to the sample. The conversion step may comprise bisulfite treatment or an enzymatic treatment. Cytosine methylation (5-methylcytosine, 5mC) and hydroxymethylation (5-hydroxylmethylcytosine, 5hmC) arc the most common epigenetic marks in the eukaryotic genome and can be used to evaluate methylation status of a sample. Insome embodiments, unmodified cytosine residues (e.g. unmethylated cytosine residues and unhydroxymcthylatcd residues) arc converted. In some embodiments, methylated cytosine (5mC) residues and / or hydroxymethylated (5hmC) residues are converted. In some embodiments, the desired cytosine residues are converted to uracil or another base that is detectably dissimilar to cytosine. Exemplary methods for enzymatic conversion are described in PMID 34140313, 30804537, 37322153.
[0097] In some embodiments, the conversion step selectively converts unmethylated cytosine residues in the sample to uracil residues. For example, bisulfite conversion can be used, wherein a sample is treated with bisulfite to selectively convert unmethylated, but not methylated or hydroxymethylated, cytosine bases to uracil. Exemplary methods for bisulfite conversion are described in (Frommer et al., 1992, Proc Natl Acad Sci USA 89:1827-31; Olek, 1996, Nucleic Acids Res 24:5064-6; EP 1394172).
[0098] In some embodiments, methylated cytosine (5mC) residues and / or hydroxymethylated (5hmC) residues arc converted using an enzymatic treatment. The enzymatic treatment may be performed using one enzyme or multiple enzymes. For example, in some embodiments the conversion step is an enzymatic conversion utilizing a ten-eleven translocation (TET)- related enzyme. For example, TET methylcytosine dioxygenases comprise three enzymes that catalyze the hydroxylation of DNA methyl cytosine (5mC) into 5-hydroxymethylcytosine (5hmC) and then further oxidize 5hmC to form 5-formylcytosine (5fC) and 5-carboxycytosine (5cAC).These oxidation products can be further reduced or converted into a suitable base detectably dissimilar from cytosine. An exemplary method involving use of a TET-related enzyme is described in Liu, Y., Siejka-Zieliriska, P., Velikova, G. et al. Bisulfite-free direct detection of 5- methylcytosine and 5-hydroxymethylcytosine at base resolution. Nat Biotechnol 37, 424-429 (2019). https: / / doi.org / 10.1038 / s41587-019-0041-2. The method involves contacting a sample with TET enzymes such that 5mC and 5hmC are oxidised to 5-carboxylcytosine (5caC) and subsequently reduced to dihydrouracil (DHU) by pyridine borane. DHU is then amplified and sequenced as thymine (T) during final sequencing.
[0099] In some embodiments, the conversion step is an enzymatic conversion using an apolipoprotein B mRNA-editing catalytic polypeptide- like (APOBEC) enzyme. APOBEC enzymes readily deaminate unmodified cytosine and 5mC, but not cytosines with other modifications. One drawback of this approach is that it can only be used for measuring 5hmC,and does not distinguish 5mC from unmodified cytosine. In some embodiments, APOBEC enzymes can be used in conjunction with other suitable enzymatic modifications to facilitate investigation of both 5mC and 5hmC, as opposed to 5hmC alone. For example, the commercially available NEBNext Enzymatic Methyl-seq Kit (EM-seq™) uses TET2 to first oxidize 5mC and 5hmC, which provides protection of these bases during the subsequent addition of APOBEC. APOBEC is then applied to deaminate the unmodified cytosines to uracils. By using these two enzymatic steps in combination, the modified bases 5rnC and 5hmC can be detected without harsh chemical treatments such as bisulfite.
[0100] In some embodiments, the method further comprising selectively amplifying and / or sequencing the nucleic acid strands comprising the first adapter, the cut site flanking sequence, and the second adapter. For example, in some embodiments the method comprises selectively sequencing, by nanopore sequencing, the nucleic acid strands comprising the first adapter, the cut site flanking sequence, and the second adapter (as described above, the first adapter and / or the second adapter need not contain primer landing sites but instead may contain motor protein(s) that assist in nanopore sequencing. In some embodiments, the method comprises selectively amplifying and / or sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site (also referred to herein as the “amplifiable strands”). In some embodiments, amplification is performed by polymerase chain reaction (PCR) using amplification primers. In some embodiments, amplification is performed using a first primer containing a sequence that is complementary to the primer landing site or the reverse complement thereof in the first adapter, and using a second primer containing a sequence that is complementary to the second primer landing site or the reverse complement thereof in the second adapter. The sequence of the primers can vary depending on the sequence of the primer landing sites present within the first and second adapters. As described above, the primer landing sites can be designed such that direct primer binding occurs (e.g. the primer is complementary to the sequence of the primer landing site and thus directly binds to the primer landing site) or indirect primer binding occurs (e.g. the primer is complementary to the reverse complement of the primer landing site generated during synthesis of the reverse complementary strand).
[0101] In some embodiments, amplification is performed using a first primer (P5) comprising the sequence 5’-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ IDNO: 10). Tn some embodiments, a primer comprising the sequence of SEQ TD NO: 10 directly binds to the primer landing site of the primer binding strand of the first adapter. In some embodiments, the first primer comprises the sequence AATGATACGGCGACCACCGAGATCTACACCGAATACGACACTCTTTCCCTACACGAC GCTCTTCCGATCT (SEQ ID NO: 11).
[0102] In some embodiments, amplification is performed using a second primer (P7) comprising the sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 14). In some embodiments, a primer comprising the sequence of SEQ ID NO: 14 binds to the reverse complement of second primer landing site (e.g. formed after the reverse complement of the second adapter is made). In some embodiments, the second primer comprises the sequence CAAGCAGAAGACGGCATACGAGATGTCGGTAAGTGACTGGAGTTCAGACGTGTGCT CTTCCGATCT (SEQ ID NO: 15).
[0001] In some embodiments, the method further comprises sequencing the amplified nucleic acids. Given that the amplified nucleic acids contain a cut site flanking sequence, which contains the CpG site, sequencing provides valuable information about the methylation status of cut site flanking sequences obtained by such a method. A variety of suitable sequencing methods and technologies may be used to determine the sequence of the nucleic acid strands. For example, the sequencing method may be a next generation sequencing technology. The term next generation sequencing, or “NGS”, refers to a variety of sequencing techniques that permit simultaneous sequencing of millions of nucleic acid sequences, and is otherwise referred to as high-through put sequencing or massively parallel sequencing. Suitable NGS technologies are reviewed in, for example, Zhong et al., Ann Lab Med. 2021 Jan; 41(1): 25-43, and Slatko et al., Curr Protoc Mol Biol. 2018 Apr; 122(1): e59., the entire contents of each of which are incorporated herein by reference. Suitable NGS technologies include, for example, second generation sequencing technologies such as pyro sequencing, ion torrent sequencing, and bridge PCR-based amplification methods. In general, pyrosequencing methods captures pyrophosphate (PPi) release and uses it as an indicator of specific base incorporation. Ion torrent sequencing methods rely on hydrogen ion detection technology, which detects the release of protons during incorporation of nucleotides into the nucleic acid strand during synthesis. Suitable bridge PCR- based amplification technologies include various Illumina platforms, such as MiSeq, MiniSeq, MiSeq, HiSeq, and NextSeq platforms. Additional suitable NGS technologies include singlemolecule real-time (SMRT) technology, nanoball sequencing technology, and sequencing-by- binding technology, and scqucncing-by-hybridization technology. In some embodiments, sequencing comprises nanopore sequencing.
[0103] In some embodiments, the method further comprises generating a methylation profile for the sample based upon said sequencing. A “methylation profile” provides the methylation status of residues within the cut site flanking sequence. In some embodiments, a methylation profile provides a methylation status of CpG sites in the cut site flanking sequence. In some embodiments, a methylation profile identifies differentially methylated regions (DMRs). Identifying a methylation profile, methylation status, or differentially methylated region is inclusive of identifying / evaluating methylated and / or hydroxymethylated residues. The methylation profile can provide valuable information about a disease state in a subject from which the sample was obtained. For example, the methylation profile can be used to classify tumors, including central nervous system (CNS) tumors, in a sample obtained from a subject. DNA methylation is associated with cancer, autoimmune disease, metabolic disorders, and neurological disorders. Accordingly, the methylation profile can be used to diagnose or classify cancer, autoimmune disease, a metabolic disorder, or a neurological disorder in a subject from which a sample was obtained. In some embodiments, the method generates a methylation profile of at least 100,000 loci. In some embodiments, the method generates a methylation profile of at least 1,000,000 loci.
[0104] The methods provided herein find use in classifying tumors in a subject from which a sample was obtained. The methods provided herein are particularly advantageous over existing methods, in that the methods can be used in samples with low tumor fraction, fragmented DNA, and to classify as opposed to merely diagnose cancer in a subject. The methods provided herein are thus particularly useful for early monitoring of tumor and selection of appropriate tumor treatment based upon the classification of the tumor. Moreover, the methods can be used to classify hard-to-diagnose tumors (e.g. brain tumors, sarcoma, undifferentiated tumors, and spindle tumors), atypical presentations, carcinoma of unknown primary, and from small or liquid biopsies thus without requiring surgical means to obtain a tissue sample.
[0105] In some embodiments, the sample is a liquid biopsy sample (e.g. blood, serum, plasma urine, cerebrospinal fluid, or another biological fluid obtained from the subjectcontaining tumor-derived entities such as circulating tumor cells, circulating tumor DNA, tumor extracellular vesicles, etc.). In some embodiments, the sample comprises cell free DNA (cfDNA). In some embodiments, the sample comprises less than lOOng DNA. In some embodiments, the sample comprises less than lOOng, less than 90ng, less than 80ng, less than 70ng, less than 60ng, less than 50ng, less than 40ng, less than 30ng, less than 25ng, less than 20ng, less than 15ng, or less than lOng DNA.
[0106] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0107] Preferred embodiments are described herein, including the best mode known to the inventors for carrying out the methods described herein. Variations of those preferred embodiments may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for the methods to be practiced otherwise than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.EXAMPLESExample 1
[0108] Clinical DNA methylation profiling is being rapidly developed and adopted as a diagnostic test for tumor classification. This progressive approach, especially alongside conventional histopathological and molecular methods, can enhance tumor assessments and, sometimes even result in a change in the final diagnosis.
[0109] Despite its potential, DNA methylation classification faces some challenges: 1) low and fragmented DNA input, as in the case of cell-free (cf) or degraded DNA from archivedformalin-fixed paraffin-embedded (FFPE) samples, 2) low tumor fraction due to the high immune or stromal cell background, and 3) a minimal overlap (11%) between the currently mostly used probes derived from methylation arrays and the markers specific for cell types.
[0110] Whole genome bisulfite sequencing (WGBS), a technique for multiplexed methylation study, is rarely used due to large cost, which practically rules out large studies and deep sequencing. Reduced representation bisulfite sequencing (RRBS), an adaption of wholegenome bisulfite sequencing (WGBS), can enrich CpG-dense regions, but only from intact genomic DNA, not from fragmented DNA, which is commonly encountered in clinical samples (e.g., cfDNA, FFPE DNA). Furthermore, bisulfite conversion (modify unmethylated cytosines to uracils) and the following random priming steps can further shorten and impede the alignment of fragmented molecules, effectively rendering them invisible in the analysis. Recently, an alternative technique, extended representation bisulfite sequencing (XRBS), expands the coverage of 50% of all CpGs by allowing for isolated MspI sites in comparison to RRBS. However, XRBS does not perform well for fragmented DNA either.
[0111] Presented herein is a technique referred to as “FLEXseq” (Fragment Ligation Exclusive methylation sequencing), a methylation profiling method that targets the adjacent flanks of CCGG motifs in the genome that associate with cell type-specific markers, enhancers, and other epigenetic functional elements (Fig. la). Extended-representation bisulfite sequencing (XRBS) also targets CCGG flanks (Shareef, S. J. et al. Nat Biotechnol (2021) doi:10.1038 / s41587-021-00910), but FLEXseq provides superior results using samples such as fragmented and low-input DNA found in clinical specimens (e.g. cfDNA, FFPE tissue DNA). To design for fragmented DNA, a semi-permissive adapter was introduced that selectively blocks the free ends of non-target DNA, allowing FLEXseq to deliver a highly on-target and accurate methylome.
[0112] To demonstrate the power of FLEXseq for clinical testing, a case-control study of 106 cerebrospinal fluids (CSF) and a combined case series of 42 non-CSF body fluids and 37 FFPE tissue blocks across different specimen types was performed (Fig. lb). Copy number analysis, tumor classification, and cell type deconvolution analysis can be performed by leveraging FLEXseq’ s substantial overlap with past reference datasets.RESULTS
[0113] 1. Design and benchmarks of FLEXseq - enriched methylation profiling
[0114] FLEXseq - theoretical coverage
[0115] The need to detect and classify suspected tumors motivated the design of a genome-wide methylation profiler that overlaps with cell type-specific markers (Loyfer, N. et al. Nature 613, 355-364 (2023) and references for tumor classification (Capper, D. et al. Nature 555, 469-474 (2018); Maros, M. E. et al. Nat Protoc 15, 479-512 (2020), Koelsche, C. et al. Nat Commun 12, 498 (2021)). The theoretical overlap of various approaches with past references (Table 1) was first calculated, focusing specifically on short cfDNA (-160 bp) and FFPE (-100- 500 bp) DNA in clinical specimens.
[0116] An in silico analysis showed that the 100 bp regions flanking CCGG motifs cover 9 million CpGs with the highest yield closest to CCGG and 46% of CpGs in cell type-specific markers (Fig. 2a, Fig. 3a-c). In contrast, other methylation profiling methods, including methylation microarrays, RRBS, and methylated DNA immunoprecipitation sequencing (MeDIP-Seq) only cover 3%-8% of those CpGs. CCGG flanks also cover 36%-37% of CpGs from the machine learning (ME) classifiers based on methylation array data (central nervous system [CNS] or TCGA) (Fig. 2b). CCGG flanks cover more enhancers, CpG shores / shelves, open sea, and introns than other methods (Fig. 3d).
[0117] FLEXseq - design
[0118] FEEXseq was designed to target motifs containing a CpG site, output accurate methylation data, and be compatible with fragmented DNA input. FEEXseq involves two ligations and a double-stranded cut between the ligations, followed by methylation conversion (Fig. la). In the first step, both sides of input DNA molecules ligate on a semi-permissive adapter (Adapter A) that will serve as a blocker for untargeted background molecules and a primer landing site for targeted molecules. A nuclease is then used to make double- stranded cuts, creating new and unblocked DNA ends. After an end repair, these newly exposed ends at the cut site are targeted through a second ligation with a second adapter (Adapter B) that is also needed for sequencing library formation. Because each of the two adapters at opposite ends has a primer landing site (i.e. A+B are both needed), only targeted and cut DNA molecules are eventually sequenced while non-target DNA is ignored. Sequencing excludes non-target DNA with two A adapters or two B adapters. All DNA molecules have a free end associated with Adapter B that allows PCR duplicate removal and the potential of fragmentomics. The distinct free end is a major advantage compared to most profiling methods besides WGBS. Enzymatic conversionwas then used instead of bisulfite conversion to avoid breaking the adapter-ligated molecules. Conversion is followed by amplification and sequencing of the target molecules.
[0119] In the exemplary FLEXseq method shown in FIG. 1, the schematic shows targeting of a CCGG (SEQ ID NO: 1) motif. Although a CCGG (SEQ ID NO: 1) is often referred to herein as targeted by FLEXseq, this CCGG (SEQ ID NO: 1) motif is only intended to be an exemplary motif containing a CpG site described herein, other motifs are also suitable. For example, the motif comprising a CpG site may comprise CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6). The nuclease used depends on the motif targeted. The exemplary nuclease MspI is shown in FIG. 1 to target the CCGG (SEQ ID NO: 1), motif, but other suitable nucleases may be used to target other CpG containing motifs, as described herein. The term “FLEXseq” does not necessarily indicate that the nuclease cuts at the CCGG cut site, but rather the nuclease may cut at any suitable motif containing a CpG site including any of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. In the schematic shown in FIG. 1, an enzymatic conversion step is performed. However, as described herein, depending on the sequencing method used this enzymatic conversion step may be performed or not performed (e.g. for nanopore sequencing). The term “FLEXseq” does not indicate that an enzymatic conversion step must be performed.
[0120] A schematic of an exemplary process performed in accordance with the methods described herein, including exemplary sequences for adapters, primers, and target nucleic acids, is shown in FIG. 4.
[0121] FLEXseq - analytical performance
[0122] To assess reproducibility and bias in the methylation output, K562 replicates were correlated and compared with WGBS as the gold standard. Two inter-run replicates were highly correlated (Pearson’s r = 0.98, Fig. 2c), demonstrating reproducibility. FLEXseq was highly correlated with WGBS, demonstrating biological fidelity (Pearson’s r = 0.97, Fig. 2d), and methylation array (Pearson’s r = 0.96, Fig. 2e). Three representative FFPE tissue samples were also correlated between FLEXseq and the EPIC version 2 (EPICv2) array (Pearson’s r = 0.93- 0.95, FIG. 5a-b). Moreover, FLEXseq had similar or higher correlations with WGBS than the public RRBS (Zhang, J. et al. Nat Commun 11, 3696 (2020) and XRBS (Shareef et al) data (Fig.5c). These high correlations suggest that FLEXseq data is compatible with past datasets such as ones derived from the microarrays.
[0123] The DNA input limit of detection was next assessed by titrating cfDNA from pleural fluid (BF3713) and gDNA (K562). CpG methylation from titrated cfDNA was correlated to 10 ng input and had similar coverage down to 1 ng (Fig. 2f). Further dilutions down to 250 pg decreased CpG coverage, but bias in methylation remained low (Pearson’s r > 0.90). Similarly, there was high correlation between 1-50 ng of FFPE tissue DNA (Pearson’s r > 0.90, Fig. 5d) and 1-100 ng K562 gDNA (Pearson’s r = 0.98, Fig. 5e). In contrast, the EPICv2 methylation array titrations of K562 gDNA had a weaker correlation (Pearson’s r = 0.82) between 250 ng and 100 ng (Fig. 51).
[0124] The median on-target rate of K562 DNA input titrations was 96% (IQR 95%- 96%) based on reads starting with (C / T)GG. Each on-target read is guaranteed to contain methylation data based on the first position. This high on-target rate yields an 18-, 6-, and 5-fold enrichment of CCGG flanks, cell type markers, and all CpGs compared with WGBS on a nucleotide basis (Fig. 2g). FLEXseq had the deepest coverage around the targeted CCGG cut sites (Fig. 2h, Fig. 3c).
[0125] 2. Analysis - copy number aberrations (CNA)
[0126] FLEXseq provided chromosomal copy number output after normalizing against control diploid references from CSF cfDNA. FLEXseq and whole genome sequencing (WGS) results were comparable with the exception of one case.
[0127] Tumor DNA was detected by interpreting deviations from diploid as CNAs and indicators of clonal aneuploidy. Tumor purity was estimated based on the difference between the log2 ratios of chromosomal segments and the baseline at gains and losses. Both measures were used in decision-making for the classifiers in the following sections.
[0128] 3. Analysis - ML classification of tumors
[0129] DNA methylation profiling is being adopted clinically to adjudicate the diagnosis of CNS and solid tumors through tissue microarray analysis. Herein, this approach was optimized to FLEXseq while reducing both DNA input requirements and cost.
[0130] Development ofTCGA and CNS ML classifiers
[0131] A solid tumor (TCGA) ML classifier comprising 2,508 samples and 24,346 markers was developed (FIG. 6, FIG. 7, FIG. 8, and FIG. 9). A parallel CNS tumor ML classifierconsists of 2,801 samples and 16,429 markers (Fig. 10). The estimated error rate of the ML calibrated scores was 10.2% for the TCGA and 1.0% for the CNS ML classifier. The 45 tumor and control groups of the TCGA ML classifier are visualized using / -distributed stochastic neighbor embedding (t-SNE) dimensionality reduction (Fig. 6b).
[0132] Classification of tumor cell line DNA titrations
[0133] Both ML classifiers were applied to DNA titrations of three tumor cell lines (breast invasive carcinoma [BRCA], colon carcinoma [COAD], and glioblastoma [GBM]) mixed into primary immune cell DNA background (B cells, T cells, monocytes, and neutrophils). Samples with higher tumor purities were accurately classified, while low-purity samples were classified with the control references (Fig. 6c).
[0134] Classification of tumor samples across sample types
[0135] FEEXseq’s versatility across sample types was demonstrated by analyzing 106 CSF cfDNA samples and a case series of 42 other body fluid cfDNA and 37 FFPE tissue samples (Fig. 11). Non-CSF body fluid tumor cases with low (<50%) purity were more often classified as controls or had lower classifier scores, leading to an indeterminate classification (84.6% vs. 31.2%, P < 0.001). Four cases were misclassified with ME, but half were accurately classified in r-SNE plots (BF3090 and TF112, Fig. 1 Id-e).
[0136] Classification with only FLEXseq references
[0137] Classifying FLEXseq data with microarray-based classifiers resulted in lower classifier scores due to differences in specimen type and methylation readout. These constraints motivated the assembly of a preliminary classifier using only FLEXseq data, comprised of 3.2 million CpG markers and 57 references (Fig. 12a). This classifier had matching classifications for all 18 hold-out samples with >50% tumor purity. The FLEXseq-based classifier achieved higher scores (mean 0.89 ± 0.10 standard deviation [SD]) than the microarray-based classifiers (P = 0.012, Fig. 12b). In the new classifier, different reference sample types (FFPE and body fluid) co-clustered in the z-SNE (Fig. 12c-d).
[0138] 4. Analysis - deconvolution
[0139] The ML tumor classification’s accuracy and classifier score decrease with tumor purity <50-70%, limiting the approach’s broader use in small tissue specimens and liquid biopsies. To overcome this limitation, an orthogonal classifier was developed based on the deconvolution of cell type proportions using cell type-specific markers.
[0140] Deconvolution of in silico titrations
[0141] To assess the performance of deconvolution in silico, WGBS data was subsampled from purified cell types and intersected them with CCGG flanks to create seven titrations. Two deconvolution approaches were tested to predict cell type proportions for each titration: i) CelFiE CpG-level deconvolution using the top 30 cell type-specific markers that excluded the titration data and ii) UXM fragment- level deconvolution using the top 250 markers. Both methods yielded similar predicted and actual proportions (root- mean- square error [RMSE] <0.01, Fig. 13a), with UXM fragment-level deconvolution remaining accurate down to 0.3% purity of target cell types.
[0142] Deconvolution of DN A titrations and plasma
[0143] Physical titrations were continued by mixing DNA from purified primary immune cell types. Fragment-level deconvolution detected target cell types down to 1% purity, and this approach was used for all further deconvolutions (Fig. 13b). The titrations of DNA from cancer cell lines were deconvoluted into DNA from immune cell mixtures. The deconvolution underestimated the tumor cell type proportion, but the rank order of the proportions was reliable (Fig. 13c). Plasma cfDNA from healthy donors was deconvoluted using WGBS and FLEXseq data (Fig. 13d). Cell type proportions were similar between WGBS and WGBS intersected with CCGG flanks (RMSE = 0.02).
[0144] Deconvolution Classifier
[0145] These results were used to build a CSF-specific deconvolution classifier (Fig. 14a) that normalizes deconvoluted cell type proportions against negative controls (n = 50, Fig. 15a) via a z-score. Ranked ordered z-scores are then used to classify a tumor based on its cell-of- origin. The 20 cell type references used for deconvolution were based on common tumor types associated with CSF samples (see Methods). Excluding noisy cell types unrelated to common tumors, our tumor classification threshold was first ranked z-scores above two (Fig. 3b, Extended Data Fig. 8b-c). Alternatively, we normalized z-scores against non-target tumors (Extended Data Fig. 8d), but low tumor purity cases increased the background variance. The deconvolution classifier yielded particularly high z-scores for LU AD (n = 20, cell-of-origin: lung alveolar epithelium) and DLBC (n = 21, cell-of-origin: B cell) in the context of low background (Fig.3b).
[0146] 5. CSF Case-control Study
[0147] Current methods for tumor detection and classification in CSF are insensitive, and a brain biopsy is significantly more invasive. FLEXscq was applied on CSF to demonstrate tumor classification. In this study, 548 consecutively collected CSF samples were obtained. Based on selection criteria, 106 samples were used for tumor classification in the case-control study (56 cases and 50 negative controls, Fig. 16a). Patients in the case group (median 59.5, IQR 52.8-67.0) were older than the negative controls (median 41.0, IQR 31.8-60.0, P < 0.001) and similar by sex (P - 0.122). CSF cfDNA was sequenced at a median of 154 million 2x50 reads (IQR 116-181 million). The CCGG on-target rate remained high at 98%, similar to the initial K562 benchmark. Fig. 17 show additional specimen characteristics and cfDNA inputs.
[0148] Tumor classification of CSF cfDNA
[0149] Copy number analysis, ML classifier, deconvolution classifier, and a composite classifier that integrated all analyses were used (Fig. 16a-b). None of the negative controls showed detectable CNAs. Among 56 cases, 53 were CNA-positive with a median tumor purity of 50% (IQR 30%-65%). Of the three CNA-negatives, one (BF3269) was cytology atypical results and flow cytometry positive, another (BF3811) was cytology positive without concurrent flow cytometry, and the third (BF3723) was the only case CNA-positive by WGS but not by FLEXseq. FLEXseq has a detection rate of 94.6% (53 / 56) compared to 67.9% by the clinical testing (cytology / flow cytometry, 38 / 56) among CSF tumor cases.
[0150] The ML classifier had an overall accuracy of 94% for tumors >50% purity, excluding 10 indeterminates. Cases with tumor purity <50% (30 of 56) tended to be categorized as indeterminate due to low classifier scores or classified to controls (73.3% vs. 34.6%, P = 0.004). The TCGA and CNS microarray-based ML classifiers were used sequentially to maximize the coverage of different tumor entities. One of two misclassified cases (BF3027) and six of ten indeterminates were classified with their pathological diagnoses in the Z-SNE plots (Fig. 18a).
[0151] The deconvolution classifier (described above) had an overall accuracy of 96% (excluding seven indeterminates) and was 100% for LU AD, DLBC, and BRCA, respectively. Of the two misclassified cases, BF3683 was classified with STAD in the Z-SNE but was estimated to have a renal origin by deconvolution. Further charting showed that this patient had a concurrent renal cell carcinoma (RCC) not known to be in the CSF. The other case, BF3369, was alsopredicted to have a predominate renal origin, but the gold standard was also unclear based on unusual histology and non-matching clinical methylation classification testing of the tumor tissue
[0152] A composite tumor classifier integrating ML and deconvolution classification was developed (Fig. 16c). The ML classifier was first used for its broad range of references, then deferred to the deconvolution classifier when cases had either low tumor purity (<50%, Fig. 18b), low classifier scores (<0.3), or classification as controls. The deconvolution classifier improved the classification of low tumor purity and indeterminate cases that performed poorly in the ML classifier (Fig. 16d). The composite classifier’s overall accuracy was 98% (excluding three indeterminates) and reached 100% accuracy for the three most prevalent tumor types: LUAD, DLBC, and BRCA (Fig. 16d-e).
[0153] Provided herein is an assay referred to as FLEXseq, a methylation enrichment profiling assay. FLEXseq is demonstrated herein to have clinical relevance through copy number detection, ML classification, and deconvolution across different liquid biopsies and FFPE tissues. This method leverages both the breadth of references with ML classification and the lower tumor purity requirements with deconvolution classification. FLEXseq achieves an 18- fold enrichment covering CCGG flanking regions with an on-target rate of 96%-98%, guaranteeing methylation data from nearly every sequenced DNA molecule. FLEXseq is highly concordant with the WGBS gold standard, accurately reflecting biological truth. Its genomewide coverage yields low-noise copy number plots. Its broad coverage allows for integration with past datasets and facilitates both ML and deconvolution classification with external references.
[0154] TET, BGT, and APOBEC enzymatic conversions were used herein, but alternatives include other non-destructive conversion strategies such as TAPS, DM-seq, or SEM- seq. Bisulfite conversion is possible, albeit with DNA loss. Because the enrichment occurs before PCR, FLEXseq is also compatible with direct methylation readout using nanopore sequencing. Moreover, the design of FLEXseq is not limited to the MspI nuclease, which is restricted to CCGG motifs. Other nucleases like CRISPR-Cas9 can target alternative flanking regions. For example, the Taql-v2 enzyme was used to cut at the CpG containing motif TCGA (SEQ ID NO: 2). Results are shown in FIG. 19. The on-target rate of using the TaqLv2 nuclease (NEB, pail number R0149) targeting the 'TCGA' motif rather than MspI. Cutsmart buffer was used. The on-target rate was estimated by dividing reads starting with CGA or TGAby the total reads on the side cut by the nuclease. Each on-target read is guaranteed to contain CpG methylation data based on the first position.
[0155] There are numerous methylation profiling assays, but each has limitations in clinical use. WGBS is the gold standard, but its high sequencing and computational costs bar large-scale studies or clinical testing. Methylation microarrays are commonly used but require 250 ng of DNA input and cover only 2%-4% of all CpGs and 3%-8% of cell type markers (Fig. 2a), limiting deconvolution accuracy. Building array reference sets is costly due to the need to scale to thousands of samples. The EPICv2 array costs $265 per sample and $70 for FFPE repair. In this study, 160M reads amount to $114 per sample (NovaseqX 25B kit). As research tools, microarrays are restricted to human and mouse genomes. RRBS is effective for CpG islands but covers fewer enhancers and cell type-specific markers (Fig. 3). XRBS targets CCGG flanks like FEEXseq but was not designed for fragmented DNA found in clinical specimens. MeDIP-seq targets methylated cytosines but imprecisely detects hypomethylated markers, encompassing 98% of cell type-specific markers.
[0156] As a cost-efficient yet broad profiling technology, FEEXseq opens future possibilities. By preserving one free end on every DNA molecule, it has the potential to be utilized in fragmentomics. FEEXseq captures half of the known methylation aging markers associated with epigenetic clocks and is poised to broaden the identification of aging markers. FLEXseq can be used to explore the methylome in non-human organisms. Genotypes and mutations can be phased with the methylation status in over 95% of reads. Finally, this data is compatible with metagenomics, which could potentially provide information relevant to infectious diseases.
[0157] In conclusion, FLEXseq represents a clinical tool for low-input and low-purity samples from small and / or liquid biopsies. As a research tool, it offers base-pair resolution of the methylome, providing an economical yet comprehensive alternative to whole genome sequencing and microarrays.
Claims
Claims:
1. A method of preparing nucleic acids for analysis, comprising: a) ligating a first adapter to target double-stranded nucleic acids in a sample, wherein the first adapter comprises a primer binding strand containing a first primer binding site; b) adding a nuclease that cuts at a motif containing a CpG site , if present in the target double-stranded nucleic acids in the sample, thereby producing a subpopulation of nucleic acid fragments ligated to the first adapter, wherein at least some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer binding strand of the first adapter; and c) adding a second adapter to the sample, wherein a portion of the second adapter comprising a second primer landing site ligates to the nucleic acid fragments ligated the first adapter, thereby forming nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site.
2. The method of claim 1, wherein the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6).
3. The method of claim 1 or claim 2, wherein the nuclease cuts at a 5’ location to the CpG site in the motif.
4. The method of claim 2 or claim 3, wherein the nuclease cut between the first and second residues of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.
5. The method of claim 4, wherein the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
6. The method of any one of claims 1 -5, wherein the target double- stranded nucleic acid comprises a first strand and a second strand, and wherein a 5’ end of the primer binding strand ligates to a 3’ end of a first strand of the target double-stranded nucleic acid.
7. The method of claim 6, wherein the primer binding strand comprises a 5’ phosphate.
8. The method of any one of claims 1-7, wherein the first adapter comprises the primer binding strand and a blocking strand.
9. The method of claim 8, wherein the primer binding strand ligates to a first strand of the target double-stranded nucleic acid in the sample and wherein the blocking strand ligates to a second strand of the target double-stranded nucleic acid in the sample.
10. The method of any one of claims 1-9, wherein the primer binding strand ligates to a first strand of the target double- stranded nucleic acids in the sample, and wherein a second strand of the target double- stranded nucleic acids in the sample does not ligate to the first adapter.
11. The method of any one of claims 1-10, wherein the first primer binding site of the first adapter comprises cytosine residues resistant to a methylation conversion or no cytosine residues.
12. The method of any one of claims 1-11, wherein the second adapter is double- stranded.
13. The method of any one of claims 1-12, wherein the second adapter is single-stranded.
14. The method of any one of claims 1-13, wherein the cut site flanking sequence results from the nuclease cutting a single motif containing a CpG site in a target nucleic acid strand.
15. The method of any one of claims 1-14, wherein the nuclease cuts at a motif containing a CpG site in each strand of the target double-stranded target nucleic acid, if present in the sample, such that the subpopulation of nucleic acid fragments ligated to the first adapter are substantially double- stranded.
16. The method of claim 15, wherein for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double-stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is ligated to the blocking strand of the first adapter.
17. The method of claim 15, wherein for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double-stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is not ligated to or is otherwise disconnected from the first adapter.
18. The method of any one of the preceding claims, further comprising performing a conversion step to selectively convert cytosine residues in the sample depending on their methylation status to a different nucleic acid besides cytosine.
19. The method of claim 18, wherein the conversion step selectively converts unmethylated cytosine residues in the sample to uracil residues.
20. The method of claim 18 or claim 19, wherein the conversion step is an enzymatic conversion or bisulfite conversion.
21. The method of claim 20, wherein the enzymatic conversion step comprises contacting the sample with an APOBEC-related enzyme or a TET-related enzyme.
22. The method of any one of claims 18-21 , wherein the conversion step is performed in between adding the nuclease and adding the second adapter to the sample.
23. The method of any one of claims 18-21, wherein the conversion step is performed after adding the nuclease and adding the second adapter to the sample.
24. The method of any one of the preceding claims, wherein the nuclease is MspI, Taql-v2, AccII, Acil, AspLEI, HPall, HpyCH4IV, or a CRISPR / Cas nuclease.
25. The method of any one of the preceding claims, wherein the cut site flanking sequence is at least 20 base pairs in length.
26. The method of claim 25, wherein the cut site flanking sequence is 20 to 300 base pairs in length.
27. The method of claim 22, wherein the CCGG cut site flanking sequence is about 70 to 150 base pairs in length.
28. The method of any one of the preceding claims, further comprising selectively amplifying and / or sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site.
29. The method of claim 28, further comprising generating a methylation profile for the sample based upon said sequencing.
30. The method of claim 28, comprising sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site by nanopore sequencing.
31. The method of claim 30, wherein no enzymatic conversion step to selectively convert cytosine residues in the sample depending on their methylation status is performed priorto said sequencing nucleic acid strands by nanopore sequencing.
32. A method of classifying methylation status of a sample, comprising: a) ligating a first adapter to target double-stranded nucleic acids in a sample, wherein the first adapter comprises a double- stranded end comprising a primer binding strand containing a first primer binding site; b) adding a nuclease that cuts at a motif containing a CpG site, if present in target nucleic acids in the sample, thereby producing a subpopulation of nucleic acid fragments ligated to the first adapter, wherein at least some nucleic acid fragments in the subpopulation comprise a first end containing a cut site flanking sequence and a second end ligated to the primer binding strand of the first adapter; c) adding a second adapter to the sample, wherein a portion of the second adapter comprising a second primer landing site ligates to the nucleic acid fragments ligated to the first adapter, thereby forming nucleic acid strands comprising the first primer binding site, the cut site flanking sequence, and the second primer binding site; and generating a methylation profile for the sample by sequencing the nucleic acid strands comprising the first primer landing site, the cut site flanking sequence, and the second primer landing site.
33. The method of claim 32, wherein the motif containing a CpG site comprises CCGG (SEQ ID NO: 1), TCGA (SEQ ID NO: 2), CGCG (SEQ ID NO: 3), CCGC (SEQ ID NO: 4), GCGC (SEQ ID NO: 5), or ACGT (SEQ ID NO: 6).
34. The method of claim 32 or claim 33, wherein the nuclease cuts at a 5’ location to the CpG in the motif.
35. The method of claim 33 or claim 34, wherein the nuclease cut between the first and second residues of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.
36. The method of claim 35, wherein the cut site flanking sequence comprises CGG, CGA, GCG, CGC, or CGT.
37. The method of any one of claims 32-36, further comprising performing a conversion step to selectively convert cytosine residues in the sample depending on their methylation status to a different nucleic acid besides cytosine, wherein the conversion step is performed prior to said sequencing, and wherein the second adapter is added to the sample before or after performing the conversion step.
38. The method of claim 37, wherein the conversion step is an enzymatic conversion or bisulfite conversion.
39. The method of claim 38, wherein the enzymatic conversion step comprises contacting the sample with an APOB EC-related enzyme or a TET-related enzyme.
40. The method of any one of claims 32-36, wherein said sequencing the nucleic acid strands comprises nanopore sequencing.
41. The method of claim 40, wherein no enzymatic conversion step to selectively convert cytosine residues in the sample depending on their methylation status is performed prior to said sequencing nucleic acid strands by nanopore sequencing42. The method of any one of claims 32-41, wherein the target double-stranded nucleic acid comprises a first strand and a second strand, and wherein a 5’ end of the primer binding strand ligates to a 3’ end of a first strand of the target double-stranded nucleic acid.
43. The method of claim 42, wherein the primer binding strand comprises a 5’ phosphate.
44. The method of any one of claims 32-43, wherein the first adapter comprises the primer binding strand and a blocking strand.
45. The method of any one of claims 32-44, wherein the primer binding strand ligates to a first strand of the target double-stranded nucleic acid in the sample and wherein the blocking strand ligates to a second strand of the target double-stranded nucleic acid in the sample.
46. The method of any one of claims 32-44, wherein the primer binding strand ligates to a first strand of the target double-stranded nucleic acids in the sample, and wherein a second strand of the target double-stranded nucleic acids in the sample does not ligate to or is otherwise disconnected from the first adapter.
47. The method of any one of claims 32-46, wherein the first primer binding site of the first adapter comprises cytosine residues resistant to a methylation conversion or no cytosine residues.
48. The method of any one of claims 32-47, wherein the second adapter is double-stranded.
49. The method of any one of claims 32-47, wherein the second adapter is single-stranded.
50. The method of any one of claims 32-49, wherein the cut site flanking sequence results from the nuclease cutting a single motif containing a CpG site in a target nucleic acid strand.
51. The method of any one of claims 32-50, the nuclease cuts at a motif containing a CpG site in each strand of the target double-stranded target nucleic acid, if present in the sample, such that the subpopulation of nucleic acid fragments ligated to the first adapter are substantially double- stranded.
52. The method of claim 51, wherein for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double-stranded nucleic acid fragment is ligated to the primer binding strand of the firstadapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is ligated to the blocking strand of the first adapter.
53. The method of claim 52, wherein for the substantially double- stranded nucleic acid fragments ligated to the first adapter, the second end of a first strand of the substantially double-stranded nucleic acid fragment is ligated to the primer binding strand of the first adapter and the second end of a second strand of the substantially double-stranded nucleic acid fragment is not ligated to or is otherwise disconnected from the first adapter.
54. The method of any one of claims 32-53, wherein the nuclease is MspI, Taql-v2, AccII, Acil, AspLEI, HPall, HpyCH4IV, or a CRISPR / Cas nuclease.
55. The method of any one of claims 32-54, wherein the cut site flanking sequence is at least 20 base pairs in length.
56. The method of claim 55, wherein the cut site flanking sequence is 20 to 300 base pairs in length.
57. The method of claim 56, wherein the CCGG cut site flanking sequence is about 70 to about 150 base pairs in length.
58. The method of any one of claims 32-57, wherein the second adapter is double-stranded or single- stranded.
59. The method of any one of claims 32-58, wherein the method generates a methylation profile of at least 100,000 loci.
60. The method of claim 59, wherein the method generates a methylation profile of at least 1,000,000 loci.61 . The method of any one of the preceding claims, wherein the sample is a liquid biopsy sample.
62. The method of any one of the preceding claims, wherein the sample comprises less than 250 ng of DNA.
63. The method of any one of the preceding claims, wherein the sample is obtained from a subject having or suspected of having cancer.
Citation Information
Patent Citations
Methods and compositions for high-throughput bisulphite DNA-sequencing and utilities
US20090047680A1
Preparation of templates for methylation analysis
US20210277459A1
Methods and systems to improve the signal to noise ratio of DNA methylation partitioning assays
US20230130140A1
Methods and compositions for analyzing nucleic acid
US20230235320A1