Method for sequencing nucleic acid molecules and related methods
The method addresses the limitations of existing cfDNA sequencing technologies by using a methylation-sensitive restriction enzyme and TdT-mediated sequence identifiers, enhancing the precision and reliability of cfDNA analysis for cancer diagnostics.
Patent Information
- Application Number
- PCT/NL2024/050664
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-21
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Existing methodologies for sequencing nucleic acid molecules, particularly circulating free DNA (cfDNA), face challenges such as low abundance, fragmentation, and contamination with non-tumor-derived DNA, leading to compromised precision and reliability in cancer diagnostics.
A method involving the use of a restriction enzyme sensitive to methylation, followed by terminal deoxynucleotidyl transferase (TdT) treatment to add sequence identifiers, and rolling circle amplification to enhance the sequencing and methylation status analysis of cfDNA.
This method significantly improves the sensitivity, specificity, and efficiency of cfDNA analysis, enabling more accurate profiling and detection of cancer-specific mutations and methylation patterns.
Smart Images

Figure 00000049_0000 
Figure 00000050_0000 
Figure 00000051_0000
Abstract
Description
[0001]P136423PC00 Title: Method for sequencing nucleic acid molecules and related methods. The present invention relates to the field of molecular diagnostics, more specifically to methods and kits for barcoding, amplifying, and sequencing a plurality of nucleic acid molecules and the use thereof for diagnostic purposes, particularly in the field of cancer. Circulating free DNA, often released into the bloodstream by apoptotic and necrotic cells, serves as a rich source of genetic information, including mutations associated with cancer. The accurate identification and characterization of this circulating free DNA can be used in tailoring precise and personalized cancer treatment strategies. Regrettably, prevailing methodologies are impeded by challenges such as the low abundance of tumor-derived cfDNA, its fragmentation, and the presence of non-tumor-derived DNA, thereby compromising the precision and reliability of sequencing results. This patent application seeks to introduce a comprehensive solution to the limitations tied to existing techniques, thereby propelling the field of liquid biopsy diagnostics and enhancing the personalized care of cancer patients. The means and methods not only provide enhanced sensitivity, specificity, and overall efficiency of cfDNA analysis in cancer patients but also overcome the hurdles associated with accurate profiling of cfDNA in diverse cancer contexts. Currently employed technologies for cfDNA sequencing primarily include next-generation sequencing (NGS) methods. These encompass approaches such as targeted sequencing, whole exome sequencing, and whole genome sequencing. Targeted sequencing focuses on specific genomic regions of interest, offering a cost-effective and focused approach, while whole exome sequencing and whole genome sequencing provide a broader scope, enabling comprehensive analysis of the entire exome or genome, respectively. Despite their utility, existing technologies face limitations in accurately capturing low- abundance tumor-derived cfDNA, and the patent application herein discloses a novel sequencing method that significantly augments the precision and reliability of cfDNA analysis in cancer patients. The subsequent sections provide a detailed exposition of the inventive steps and distinctive features underpinning our improved sequencing methodology. Circulating free DNA (also known as cell-free DNA) are degraded DNA fragments released to body fluids such as blood plasma, urine, cerebrospinal fluid, etc. refers to DNA that is found outside of cells in the bloodstream or other bodily fluids. It is released into the circulation through processes such as cell death (apoptosis and necrosis) and is being increasingly studied for its potential as a non-invasive biomarker for various medical conditions, including cancer. cfDNA is typically small because it originates from the fragmentation of DNA released from cells. The size of cfDNA fragments is influenced by various biological processes, and the primary reasons for the small size of cfDNA include Apoptosis and Necrosis. During apoptosis, DNA is cleaved into fragments as part of the cellular breakdown. Necrosis, which is uncontrolled cell death often due to injury or disease, can also result in the release of DNA fragments. Various enzymes, such as nucleases, are present in the cellular environment, the extracellular environment, and the various fluid streams, such as the bloodstream and urine. These enzymes can further degrade DNA into smaller fragments. DNA fragments can be released into the bloodstream as part of normal physiological processes, such as tissue turnover and repair. Other processes, such as the clearance from the body, also affect the level and the size distribution of cfDNA fragments. The combination of these factors results in the presence of small DNA fragments in the bloodstream. In the context of liquid biopsy and diagnostic applications, the small size of cfDNA is advantageous because it allows for easier extraction and analysis. The analysis of cfDNA has become a valuable tool in medical research and clinical applications, such as the detection of genetic mutations, monitoring of cancer, and prenatal testing. The fragments of cfDNA are often double-stranded and relatively short, typically a few hundred base pairs in length. These fragments can have different termini (ends), including blunt ends or staggered ends with overhangs. The size distribution of cell-free DNA (cfDNA) can vary, but there are general trends observed in its fragmentation. The size range of cfDNA is commonly described in terms of its median or average size and its broader range. Keep in mind that these values can be influenced by factors such as the underlying biology of the source tissue, the specific extraction methods, and the analytical techniques used for measurement. The median size of cfDNA is often reported to be around 160-180 base pairs (bp). This means that, on average, half of the cfDNA fragments are shorter than this size, and half are longer. The broader size range of cfDNA fragments can extend from around 70 to several hundred base pairs. In healthy individuals, cfDNA fragments may vary in size, but they typically fall within this range. In the context of cfDNA, the organization of histones is disrupted due to the fragmentation of DNA. However, certain studies have explored the possibility of analyzing nucleosome- sized fragments within cfDNA, which could potentially provide information about the chromatin structure of the cells from which the cfDNA originated. The idea being that the sensitivity of DNA to cleavage and fragmentation is not uniformly distributed over the chromosomes and can vary over time, differentiation status, activation status, mutation status, and other factors. Although many applications are being developed with cfDNA and its usefulness has been established, much still can be improved with respect to retrieving the information that is present in cfDNA. It is not easy to distill the above-mentioned information from the sequences of cfDNA determined with the present methods. For instance, it is difficult to get good information on the exact ends of the cfDNA or the relative abundance of specific cfDNA. Albeit suited for the analysis of cfDNA, the methods as described herein are also very suited for the analysis of genomic DNA or other complex sources of DNA such as fecal, urine, sewage, environmental, animal and other sources of DNA. SUMMARY OF THE INVENTION The invention provides a method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of preferably linear DNA molecules; - incubating said sample with a restriction enzyme that is sensitive to methylation at the recognition site for said restriction enzyme and subsequently sequencing the DNA molecules in the sample thereby determining sequences of the DNA molecules; and determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules. The invention also provides a method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of linear DNA molecules; - contacting said sample with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of the linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single-stranded DNA molecules comprising the sequence identifiers are self-ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products are produced each comprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier. The invention also provides a kit for sequencing a plurality of DNA molecules, wherein the kit comprises a TdT, a single-strand DNA ligase, a polymerase with strand displacement activity, and at least two nucleotides. The invention also allows for the detection of circular DNA in cfDNA samples. Circular DNA in human cell-free DNA (cfDNA) samples is a noteworthy topic in molecular biology and genetics. Circular DNA, unlike the more common linear DNA, forms a closed loop and is often associated with extrachromosomal DNA elements. In humans, cfDNA primarily originates from the breakdown of cells and is present in bodily fluids like blood. The presence of circular DNA next to linear DNA in cfDNA samples is significant because it can provide valuable information about various physiological and pathological states. For instance, circular DNA can be used as a biomarker for detecting and monitoring diseases, including cancers, where it may reflect tumor-specific genetic alterations. Additionally, the study of circular DNA in cfDNA can shed light on aging processes, as its accumulation in cells has been linked to cellular senescence and age-related diseases. Therefore, the presence and analysis of circular DNA in human cfDNA samples have substantial implications for diagnostics, prognostics, and understanding human biology at the molecular level. Circular DNA present in the cfDNA fraction will not be provided with a sequence identifier by the TdT but it is still present in the reaction, and it will be amplified. The absence of a TdT-synthesized tag in the sequencing reads will then indicate the reads that were generated starting from a circular DNA source. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1: Schematic representation of classic barcoding strategies and TdT strategies. Figure 2: Schematic representation of a method of the invention and the ability to distinguish between circular and linear DNA in a sample. Figure 3: Schematic representation of a method for selective “scarring / marking” of non- methylated DNA to enhance the detection of DNA methylation. It starts with methylation- sensitive restriction enzymes that target methylation hotspots, particularly in CpG islands, digesting only the non-methylated DNA. The DNA pool treated this way can be directly sequenced. The analysis of the sequencing data reveals “scars / marks” on the reads that can be used to determine the methylation status of the underlying DNA. On the left hand side of the top panel two double stranded fragments are shown with the same sequence. One fragment contains a methylated restriction enzyme site, the other fragment contains the same restriction site but this one is not methylated. The fragments are incubated with a restriction enzyme (RE) that only digests unmethylated restriction sites. Methylated versions (m5-C) of the restriction enzyme site are not digested. Sequencing of the result of the digestion yields reads that end in the restriction enzyme site representing fragments that were digested by the enzyme and reads that contain the entire restriction enzyme site representing fragments where the enzyme did not cut (-). This is elaborated further in the lower panel where the left part schematically represents sequence reads of fragments before digestion (Std reads), the middle part schematically represents sequence reads of fragments that were digested by the methylation-sensitive enzyme (Me- reads) and the right part schematically represents sequence reads of fragments that were not digested by the methylation-sensitive enzyme (Me+ reads). Figure 4: Schematic representation of the technology for selective removal of non- methylated DNA to enhance the detection of DNA methylation. It starts with methylation- sensitive restriction enzymes that target methylation hotspots, particularly in CpG islands, digesting only the non-methylated DNA. Exonucleases then further break down these non- methylated fragments, enriching the sample in methylated DNA. The final step involves sequencing the enriched samples, where the analysis focuses on the coverage of CpG islands compared to surrounding areas. Higher coverage signifies preserved methylation, indicating a successful enrichment process and enabling precise methylation pattern analysis. Figure 5: Electrophoretic traces of a proof of concept (POC) experiment showing the selective depletion of non-methylated DNA. Three DNA sequences were used for this experiment: M-P* (lane H1) is the MYC promoter sequence with methylated BstUI sites, M- P (lane A2) is a non-methylated MYC promoter and M-T (lane B2) is the MYC terminator sequence which lack the BstUI site. Lanes B1, C1, and D1 show the effect of BstUI digestion on the input DNA; as expected, only the non-methylated MYC promoter sequence (M-P) was digested. Following digestion, the DNA was treated with Lambda and P1 exonucleases, which fully deplete the digested M-P sequence while M-P* and M-T are left untouched. Figure 6: Overview of the synthetic sequences used for the proof of concept (POC) experiment. The figure depicts Integrative Genomics Viewer (IGV) tracks showing the eventual presence of a CpG island (“CpG_GRCh38.bed” track), The eventual presence of a BstUI restriction site (“BstUI +” track), and the location of the DNA sequences used for the POC experiment (“selected_targets.bed” track). On the top panel, the MYC-Promoter sequence, which is part of a known CpG island, has 2 distinct BstUI restriction sites. In contrast, on the bottom panel, the MYC-Terminator sequence is not part of any CpG islands, and it does not contain any BstUI site. In this context, BstUI is expected to digest only a non-methylated MYC-Promoter sequence. The MYC-Terminator sequence will not be digested as it doesn’t contain the restriction site, while methylation on the restriction site would block BstUI from digesting the MYC- Promoter sequence. Figure 7: IGV screenshot showing the result of a proof-of-concept experiment. Reads from methylated DNA (top track) and non-methylated DNA (bottom track) are compared. Non- methylated DNA was correctly digested in correspondence with the two restriction sites highlighted by the blue bars below the sequence track. The alignment shows the “scar” corresponding to the non-methylated restriction site. Figure 8: IGV screenshot showing the result of methylation detection using a methylation sensitive restriction enzyme and methylation detection using bisulfite conversion. The figure represents a typical IGV screenshot used in this application to show examples of results obtained with a method of the invention. IGV visualizes the alignment of sequencing reads versus a reference genome. The reference genome used is listed in the top bar, left corner (1), followed by the selected chromosome (2) and the genetic coordinates of the zoomed-in area (3). Below the top bar is a graphical representation of the selected chromosome, its banding pattern (4), and the length of the zoomed-in area (5). This figure compares the sequencing data obtained with the methods of the invention (6) and bisulfite sequencing (7). The image visualizes the coverage (8) associated with the sequencing reads aligned to the reference genome (9). At the bottom are the reference sequence (10), genes and transcripts annotations (11), and a track highlighting the presence of CpG islands (12). Mutations with respect to the reference sequence that are detected in the sequence reads are indicated (13). The figure shows a comparison of the methylation detection technology (top panel) with conventional bisulfite sequencing (bottom panel) in analyzing DNA methylation and single nucleotide polymorphisms (SNPs). This figure depicts a region of Chromosome 5 of the human genome. In the top panel, sequencing reads generated using the proposed technology are aligned. The uniformity of coverage across the region is clearly visible, demonstrating the method's ability to preserve the quality of sequencing data while maintaining the integrity of the DNA's chemical structure. In contrast, the bottom panel illustrates the results of conventional bisulfite sequencing. The aggressive chemical modification required for this approach significantly disrupts the uniformity of coverage. As a result, while methylation information may still be obtained although not as straightforward as a method of the invention, the ability to detect SNPs in the surrounding areas is compromised, limiting the overall utility of this method. Figure 9: IGV screenshot comparing sequencing reads generated by a method of the invention (top) with bisulfite sequencing (bottom). This picture shows the alignments near a CpG island on Chromosome 9. The CpG island appears to be non-methylated, as evidenced by the sudden drop in coverage. The rest of the region does not include other CpG islands. Compared to the bisulfite sequencing method, the proposed technology delivers more mappable reads leading to a higher and more uniform coverage. Notably, a distinct drop in coverage is explicitly observed at the methylation site, indicating the absence of methylation at this location. This precise detection of methylation patterns is achieved without the need for aggressive chemical treatments, enabling simultaneous identification of SNPs in the surrounding regions. Figure 10: IGV screenshot (chromosome 20) comparing sequencing reads generated by a method of the invention (top) with bisulfite sequencing (bottom). This picture shows the alignments near four CpG islands on Chromosome 20. The first three CpG islands are non- methylated, resulting in a sudden drop in coverage in the aligned reads generated with the proposed technology. The right-most CpG island, instead, appears to be methylated, thus protected by the effect of the particular restriction enzyme, resulting in regular coverage. Again, the comparison vs bisulfite sequencing shows a striking difference in the quality of the coverage and the alignments and the ease of detection.Figure 11: IGV screenshot (chromosome 15) comparing sequencing reads generated by theproposed technology (top) with bisulfite sequencing (bottom) near CpG-88 on Chromosome 15. Notably, the proposed technology can correctly detect the methylation status of the CpG islands (non-methylated) while keeping a more uniform coverage on non-CpG sites while generating high-quality reads. The bisulfite sequencing alignments show less regular coverage and abundant noise on the mapped reads. Figure 12: IGV screenshot (chromosome 5) comparing sequencing reads generated by a method of the invention (top) with bisulfite sequencing (bottom). The IGV screenshot shows the same features as figure 8, as detailed in the legend thereto, but now for an exemplary region on chromosome 5. The method used preserves the native DNA sequence, enabling accurate SNP calling, even in GC-rich regions. In contrast, bisulfite sequencing involves the chemical conversion of cytosines to thymines (and guanines to adenines), leading to numerous false SNPs visible as colored mismatches in the coverage track. These artifacts highlight the limitations of bisulfite sequencing for SNP detection, which are effectively avoided with the proposed technology. The proposed method preserves the native DNA sequence, enabling accurate SNP calling, even in GC-rich regions. In contrast, bisulfite sequencing involves chemical conversion of cytosines to thymines (and guanines to adenines), leading to numerous false SNPs visible as grayscale mismatches in the coverage track. These artifacts highlight the limitations of bisulfite sequencing for SNP detection, which are effectively avoided with the proposed technology. DETAILED DESCRIPTION OF THE INVENTION A plurality of nucleic acid molecules refers to a diverse set or a multitude of DNA (deoxyribonucleic acid) or RNA (ribonucleic acid) molecules. The term "plurality" in this context emphasizes the presence of a variety or multiple types of nucleic acid molecules. These molecules can differ in their sequences, lengths, and functions, reflecting the genetic diversity within a biological system. In a genomic context, a plurality of nucleic acid molecules may encompass the entire genetic material of an organism or be a subset thereof, representing a part of the complexity and heterogeneity of the genetic information encoded within that collection of molecules. The present inventions have various advantages such as but not limited to obtaining sequencing information in a single assay. They allow the determination of the nucleotide sequence of a target pool of DNA molecules; the determination of the nucleotide sequence of the DNA edges, i.e. the complete sequence of the 3’ and 5’ ends of the target DNA; the determination of the methylation status of a pool of DNA; and the determination of the molecular structure of the DNA molecules, i.e. circular and linear DNA can be distinguished. Further advantages are that they do not require the modification of the nucleic acid of which the sequence is to be determined. Some embodiments add nucleotides to an end of nucleic acid molecules in the sample. However, also these embodiments do not change the sequence of the nucleic acid sequence that was in the sample. No preparation of the DNA is required, particularly if circulating free DNA is used, albeit that incubation with an enzyme that removes 5’-phospohate groups can improve sensitivity in some embodiments. The methods of the present invention are simple and allow to obtain sequence information of whole genomes in one go. The methods of the invention do not require the use blunting enzymes, PCR or ligation-based barcoding strategies. It also avoids the chemical conversion of methylated bases, like bisulfite conversion. Blunting enzymes such as exo- and endonucleases remove nucleotides digest single stranded DNA and remove overhangs from the 5’- or 3’-ends of the target DNA thereby removing the possibility to determine these end sequences in the target DNA. Polymerases that are typically used in PCR have the inherent problem of in accuracy. Inaccuracy can lead to changes in the nucleotide sequence that pop-up in sequence reads of the target molecule. Techniques are available with which one can determine whether a difference with respect to a reference sequence is due to a PCR error or already present in the target DNA, but such techniques typically require additional steps rendering the methods more complex and elaborate. Ligation-based barcoding involves the use of DNA ligases to attach short sequences of nucleotides (barcodes) to DNA fragments. These barcodes serve as unique identifiers that can be used to differentiate between samples during downstream analyses. In general, barcoding enables the simultaneous analysis of multiple samples, enhancing throughput and efficiency. However, ligation-based barcoding strategies are typically complex, requiring careful design and validation of barcodes to avoid overlapping sequences. Also, the ligation reaction may have varying efficiencies, potentially leading to incomplete ligation and the additional reaction steps increase the risk of cross-contamination between samples, which can complicate data analysis. When DNA is treated with bisulfite, unmethylated cytosine residues are converted to uracil, whereas methylated cytosines remain unchanged. The conversion is readily detected for instance using sequencing techniques. One of the problems of using the bisulfite method for detecting methylation is that the treatment also introduces other changes in the DNA rendering a concomitant sequence determination less accurate. A method of the invention does not introduce significant changes in the target sequence and also leave these ends of the target sequence intact. Thus allowing accurate detection the sequence of the target molecule, the ends of the target molecule and the methylation status of the target molecule. However, if desired one may add one or more of these steps at the discretion of the user but they are not necessary and / or ideal for obtaining advantages mentioned above. The barcoding methods of the invention, the methylation sensitive restriction digestion and the subsequent sequencing of the modified DNA can be achieved without changing the reaction vessel. This renders the methods of the invention robust and less prone to contamination than other methods of barcoding, methylation detection and sequencing. Embodiments of the present invention provide novel approaches for barcoding, using terminal deoxynucleotidyl transferase (TdT), and / or methylation detection, using selective digestion of methylated or non-methylated DNA. The methylation status of the target DNA can be unveiled either by alignment of DNA reads after digestion of non-methylated DNA, (method depicted in Figure 3) or for instance by depletion of non-methylated DNA (method depicted in Figure 4). The exact sequence of 3’ and 5’ ends of the target DNA, the barcoding of the same DNA, and the distinction between circular and linear DNA can be accomplished using methods depicted in Figure 2. When using methods for the determination or the detection of the methylation status of DNA in the sample it is preferred that the DNA is digested with a restriction enzyme that digests non-methylated versions of the restriction enzyme digestion site. The methods for methylation detection of the present invention mark a significant breakthrough in both epigenetics and cancer diagnostics. Unlike traditional approaches that often rely on bisulfite conversion and amplification—processes that chemically modify DNA and can hinder mutation detection—the present invention enables direct detection of methylation from native DNA strands. In some embodiments, this is achieved through a selective depletion step that removes unmethylated CpG islands while preserving methylated DNA. This selective depletion involves two key enzymatic processes. First, a restriction enzyme is used to specifically target and cleave non-methylated CpG sites. Following this, an exonuclease digests the cleaved DNA fragments, a process facilitated by the presence of phosphate groups introduced during the restriction enzyme digestion. As a result, unmethylated regions are effectively degraded, while methylated DNA, which remains intact due to its resistance to restriction enzyme cleavage, is preserved. In other embodiments, the depletion step is not necessary. Instead, the digestion alone can reveal the presence of methylated or non-methylated sites directly through sequencing data analysis. After sequencing, the reads are mapped to the reference genome, and the digested sites are identified by examining the starting and ending positions of the reads. For example: If a site is methylated, the restriction enzyme cannot cleave it, and sequencing reads will span the site, indicating its protection. If a site is non-methylated, it will be cleaved by the restriction enzyme. In this case, sequencing reads will start or end in proximity to the digestion site, and no reads will span the site. In cases of partial methylation, where one allele is methylated and the other is not, a mix of reads spanning the site and reads ending at the site will be observed in approximately equal ratios. There are also suitable restriction enzymes that cut the site if it is methylated and not when it is not methylated. In such cases the sequence reads will span the site if it is not methylated and there will be no reads etc in the case where the site was methylated. This analysis allows for a detailed assessment of methylation patterns while maintaining sequence integrity, enabling simultaneous mutation detection and methylation analysis. The accompanying IGV screenshots illustrate this principle: the top image, obtained using the method of the invention, demonstrates superior coverage and preserved sequence integrity, while the bottom image, generated using conventional bisulfite sequencing, exhibits degraded regions corresponding to unmethylated CpG islands. This highlights the ability of the invention to perform efficient, high-resolution methylation detection in a single assay. A method of the invention is advantageous for several reasons such as but not limited to the possibility of obtaining information on the methylation status of essentially all occurrences of the selected restriction enzyme. The enzymatic methods described herein are much more accurate than chemical methods for the detection of methylation making it easier to detect the presence of mutations in the plurality of DNA molecules as opposed to errors induced by the methods used. Another advantage is the frequency of useful sequence reads that is obtained with a method of the invention relative chemical methods for the detection of methylation. Another advantage is the low number of false positive with enzymatic digestion as chemical method can also have very low yields of sequence reads for positions on the DNA that are unrelated to the methylation status at the positions. Terminal transferases are enzymes that catalyze the addition of nucleotides to the 3' end (terminus) of a DNA or RNA molecule. Terminal deoxynucleotidyl transferase (TdT) is an enzyme particularly known for its ability to add nucleotides to the 3' ends of DNA in a template-independent manner, meaning that it does not require a DNA template to guide the addition of nucleotides. In humans, the protein is encoded by the DNTT gene. Aliases for the gene or protein in humans are DNA Nucleotidylexotransferase; TDT; Terminal Deoxynucleotidyltransferase; Terminal Addition Enzyme; Terminal Transferase; or EC 2.7.7.31. For the present invention, it is preferred to use the commercially available human TdT. However, when suitably available, it is possible to substitute this for the ortholog in other species. It is also possible to substitute the enzyme for another transferase with the same functional properties in kind, not necessarily in amount. TdT is utilized in the labeling of DNA fragments for techniques such as DNA sequencing and in the addition of labeled nucleotides for various labeling experiments. Important applications of TdT are in the cloning and other molecular biology techniques. By adding a homopolymeric tail to the 3' end of a DNA fragment, researchers can create a template for priming the synthesis of a complementary DNA strand using a primer complementary to the homopolymeric tail. Homopolymeric tailing can be achieved by contacting a DNA molecule with TdT in the presence of only one type of nucleotide. TdT will add the nucleotides to the 3’ end of the DNA molecule. Because only one type of nucleotide is added, the resulting sequence is known and can be used to design a primer to synthesize a reverse complementary strand of the DNA molecule. In the present invention, the TdT is incubated under conditions that allow the addition of two or more different nucleotides to the 3’ termini of the plurality of single- stranded DNA molecules. In such circumstances, TdT can add nucleotides depending on the availability of the particular nucleotide in the mixture. The two or more different nucleotides are typically added in equimolar amounts. In such a case, when the two or more different nucleotides are present in equimolar amounts, for instance, the inclusion in the nascent sequence is apparently random. The complexity of newly formed sequences increases with the number of different nucleotides in the mixture and the number of nucleotides that is added to the 3’ termini. When TdT is allowed to add one nucleotide to the 3’ end, the complexity of the new sequences added is very low. When two different nucleotides, for instance A and T are used there are only two different ends, specifically; ends with an A and ends with a T. The complexity rapidly increases with the number of nucleotides that are added to the 3’ ends and the number of different nucleotides in the mixture. For two different nucleotides the number of different sequences formed is 2nwhere n is the average length of the newly formed sequences. For three different nucleotides there are 3ndifferent sequences whereas for 4 or more the complexity is (4 or more)n. The two or more different nucleotides are preferably three or more, preferably four, different nucleotides. The nucleotides are preferably A, C, G and T. The complexity (or number of different sequences) does not usually need to exactly match the number of sequences in the plurality of single-stranded DNA molecules. Discrimination of different molecules in the mixture is aided by the increasing complexity of sequences in the plurality. The combination of an identifier and the sequence of the DNA molecule also provides identification information. So even when there are two or more identifiers with the same sequence, the linkage to target DNA molecules with a substantially different sequence allows adequate discrimination and subsequent identification of the target DNA molecules in the plurality of single-stranded DNA molecules. The complexity (number of different) sequence identifiers generated can be increased with an increased number of different nucleotides and increased length of the sequence identifier. The complexity can be increased to such an extent that each sequence identifier attached to the 3’ end of the plurality of DNA molecules is unique. Such sequence identifiers are also referred to as unique sequence identifiers. This is a nice feature but, as mentioned herein above, not an essential feature. A method of the invention is particularly suited to sequence the entirety of linear DNA fragments, including the ends of the fragments. Barcoding target DNA or providing DNA with sequence identifiers is typically accomplished by capturing the DNA with a backbone or with adaptors by means of ligation. This typically involves manipulation of the ends of the DNA to allow the efficient ligation of the barcodes, causing the loss of terminal sequence. Other common methods rely on PCR to generate barcoded amplicons. This clearly leads to the loss of sequences outside the amplification range. Instead, by using the method of this invention, a terminal transferase, such as preferably TdT, to produce unique identifiers and circularizing the resulting DNA, it is possible to sequence the target DNAs completely, including the ends of the fragments (Figure 1). Conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of the DNA molecules by the TdT are known to the person skilled in the art. The human enzyme is well known, and various commercial sources are available. Conditions encompass a suitable buffer for the enzyme to work in, a suitable temperature, and sufficient time to produce sequence identifiers of sufficient complexity at the 3’ termini. Upon provision with sequence identifiers, the DNA molecules comprising the sequence identifiers at their 3’ termini are contacted with a single-stranded DNA ligase under conditions in which the ends of the DNA molecules comprising the sequence identifiers are self-ligated, thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier. In a preferred embodiment, the single-stranded DNA ligase is a CircLigaseTM. CircLigaseTMis a thermostable ligase that catalyzes the intramolecular ligation (i.e. circularization) of ssDNA and ssRNA templates. Self-ligation can further be enhanced over intermolecular ligation by reducing the concentration of DNA molecules in solution by increasing the volume of the ligation reaction. The ligation reaction is performed using buffer conditions that support efficient ligation of the single strands. If necessary the DNA molecules comprising the sequence identifiers at their 3’ termini are provided with a phosphate group at the 5’-end prior to the ligation. This can be done with a suitable polynucleotide kinase (PNK) using the protocol of the manufacturer. Methods of the invention are particularly suited to produce DNA circles that each have one single-stranded DNA molecule from the plurality of single-stranded DNA molecules and one sequence identifier. The plurality of circular single-stranded DNA molecules is subsequently contacted with a polymerase with strand displacement activity and one or more primers whereupon a rolling circle amplification reaction is performed, thereby generating a plurality of amplification products, each comprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules. In embodiments of the invention, leftover linear DNA, if any, is preferably removed prior to the rolling circle amplification. Performing a rolling circle amplification after removing linear DNA typically produces more high molecular weight concatemers of backbone and target DNA. Methods further include subjecting DNA circles that are produced in the ligation reaction to an amplification reaction, preferably a rolling circle amplification (RCA). Rolling circle amplification produces an ordered array of copies of at least two of said DNA circles. Rolling circle amplification produces DNA molecules of high molecular weight. Which is suited for sequencing. The multiple copies of the same circle and, thus, the same DNA molecule from the plurality and the same sequence identifier allow for a more accurate sequence determination and reduce sequencing errors. Rolling circle amplification has recently been reviewed by Mohsen and Kool (2016) Acc Chem Res. Vol 49(11): pp 2540–2550; Published online 2016 Oct 24. doi: 10.1021 / acs.accounts.6b00417. The terms rolling circle amplification and rolling circle replication are sometimes used interchangeably in the art. In other instances, rolling circle replication is used to refer to the replication of naturally occurring plasmid and virus genomes. The terms refer to a similar underlying principle, i.e. the repeated copying of the same circular DNA producing a longer nucleic acid molecule with an ordered array of backbone-target nucleic acid copies. Present techniques for rolling circle amplification enable the production of large arrays containing many copies of the produced DNA circles. Concatemers can have 2 or more copies, preferably 4 or more copies of the produced circles. Rolling circle amplification is performed by a polymerase with strand displacement activity and requires the usual priming sequence to generate the start. Particular polymerases with strand displacement activity and with high processivity are available to produce concatemers of considerable length. Polymerases with high processivity are polymerases that can polymerize a thousand nucleotides or more without dissociating from the DNA template. They can preferably polymerize two, three, four thousand nucleotides or more without dissociating from the DNA template. Polymerases with high processivity are among others discussed in Kelman et al; 1998: Structure Vol 6; pp 121-125. Rolling circle amplification can yield very high molecular weight concatemers using polymerases with high processivity and strand-displacement capacity, such as phi29 polymerase. This polymerase can polymerize 10 kb or more. High processivity polymerases are, therefore, preferably polymerases that polymerize 10 kb or more without dissociating from the DNA template (Blanco et al; 1999. J. Biol. Chem. 264 (15): 8935–40). The polymerization can be started on a nick in the double-strand DNA, or the DNA can be melted and annealed in the presence of one or more suitable primers. Examples of suitable primers are random hexamer primers, one or more backbone-specific primers, one or more target nucleic acid- specific primers, or a combination thereof. Random primers are typically preferred when target nucleic acid sequences are not known or when a variety of target nucleic acid sequences are to be sequenced. One or more specific primers can be used to sequence-specific target nucleic acids of which the basis sequence is known. A variant is one or moreprimers that are specific for one or more particular sequences in a plurality of DNA molecules. Such primers can be used in different situations, for instance, in the focused sequencing of particular sets of sequences, for instance, sequences known to be associated with certain cancers or certain pathogens. The polymerase is preferably Phi29 DNA Polymerase, Bst DNA Polymerase, or Vent (exo-) DNA Polymerase. In a preferred embodiment, the polymerase is Phi29 DNA Polymerase. The Phi29 DNA Polymerase is preferably EquiPhi29 DNA polymerase. Both are available from Thermofisher. EquiPhi29 DNA polymerase is a mutant of Phi29 DNA polymerase with particularly favorable properties. Sequencing methods have evolved over time. The old Sanger sequencing method has been replaced by the now common next-generation sequencing (NGS) methods. These methods have recently been reviewed by Goodwin et al. (2016; Nature Reviews | Genetics Volume 17:pp 333-351: doi: 10.1038 / nrg.2016.49). The most common NGS methods rely on the sequencing of short stretches of DNA. Sequencing techniques for short stretches of DNA suffer from inherent error profiles. Errors are reduced by independently sequencing multiple copies of the same target sequence. However, for each individual sequence read it is impossible to determine whether a change represents an error or a true mutation. The cumulative evidence across several independent sequence reads allows for the filtering of mutations introduced during amplification and errors in sequencing. Longer target DNAs can also be sequenced with short-read methods. This is typically done by sequencing overlapping fragments that can be aligned to create an assembled longer sequence. This so- called short read paired-end technique has been very successful in the sequencing of large target nucleic acid and has been instrumental in the various genome projects. The genome projects have revealed that genomes are highly complex with many long repetitive elements, copy number alterations and structural variations. Many of these elements are so long that short-read paired-end technologies are insufficient to resolve them. Long-read sequencing delivers reads in excess of several kilobases and allows for the resolution of these large structural features in whole genomes. Two popular platforms for long read sequencing are the Pacific Biosciences systems (the RSII and the Sequel) and the Oxford Nanopore systems (MinION and PromethION). Both are single-molecule sequencers. Both platforms allow reads in excess of 55 kb and longer. However, these systems have even higher error rates than next (second) generation sequencers. These errors can be reduced by increasing the number of times the same target nucleic acid is sequenced (Goodwin et al 2016; doi: 10.1038 / nrg.2016.49). For the present invention it is preferred to use a short read sequencing method, preferably the Illumina Sequencing Platform (Illumina sequencing). Illumina sequencing, also known as sequencing by synthesis, is a widely adopted high- throughput DNA sequencing technology. Illumina sequencing is characterized by its high accuracy, high throughput, and relatively short read lengths. It is widely used for various applications, including whole-genome sequencing, resequencing, metagenomics, and transcriptomics. The technology has played a pivotal role in advancing genomics research and has become a cornerstone in many laboratories for its efficiency and cost-effectiveness. Determined sequences can be aligned based on the sequence identifiers and, optionally, the determined sequence of a single-stranded DNA molecule. The sequences from the aligned different reads can be used to filter out errors and determine the sequence of one single-stranded DNA molecule from the plurality of single-stranded DNA molecules. The plurality of linear single-stranded DNA molecules can be derived from any source. When the source is or contains linear double-stranded DNA, the source can be denatured to create the plurality of single-stranded DNA molecules. The denaturing step can be done before or after the step wherein the sequence identifier is added to the 3’-end of the DNA molecules. The denaturing step is preferably done after the sequence identifier is added. With denaturing is meant that the two strands of a double stranded DNA are separated from each other. This can be done in various ways such as but not limited to a pH shift or a heat treatment. In the present invention it is preferred that the strands are separated by a heat treatment, such as an incubation at 980C for 10 minutes. An RNA source can be copied into DNA and made single-stranded to become a plurality of single-stranded DNA molecules. To this end, it is preferred that a plurality of RNA molecules is contacted with a reverse transcriptase under conditions in which RNA molecules are copied into cDNA molecules, thereby producing said plurality of single- stranded DNA molecules. A plurality of single-stranded nucleic acid molecules is a plurality of linear single- stranded nucleic acid molecules. A plurality of single-stranded DNA molecules is a plurality of linear single-stranded DNA molecules. A plurality of single-stranded RNA molecules is a plurality of linear single-stranded RNA molecules. The plurality of single-stranded molecules preferably comprises circulating free DNA. DNA that circulates freely or that is associated with cellular particles in the blood or other bodily fluid samples is typically smaller than 400 nucleotides. Target nucleic acid molecules of such lengths are particularly suited to the methods of the invention. Other samples with relatively small nucleic acid molecules are some types of forensic samples, fossil samples, samples of nucleic acid isolated from environments that are inherently hostile to nucleic acid molecule integrity, such as stool samples, surface water samples, and other samples rich in microbial organisms. The sample can be any sample comprising one or more nucleic acid molecules of which the sequence is to be determined. One example is a sample comprising tumor DNA. The plurality of single-stranded DNA molecules can also be circulating tumor DNA (ctDNA) or cell-free DNA (cfDNA) present in liquid biopsies, including but not limited to blood, saliva, pleural fluid, or ascites fluid. The plurality of single-stranded DNA molecules can also be single-stranded cDNA derived from messenger RNA, microRNA, CRISPR RNA, non-coding RNA, viral RNA, or other sources of RNA. The plurality of single-stranded DNA molecules can also be derived from genomic DNA, PCR products, plasmid DNA, viral DNA, or other sources. The means and methods of the present invention are particularly suited for the sequencing of short DNA. Preferably 400 base pairs or smaller. DNA in the plurality of single-stranded DNA molecules is preferably 400 base pairs or less, more preferably 300 base pairs or less, more preferably 200 base pairs or less, more preferably 150 or less. The lower limit of the target DNA is preferably 20 base pairs, more preferably 30 base pairs, more preferably 40 base pairs and more preferably 50 base pairs. Any lower limit can be combined with any upper limit. The size of DNA fragments is given in nucleotides here. This refers, of course, to the number in one strand. The invention further provides a kit for sequencing a plurality of DNA molecules, wherein the kit comprises a TdT, a single-strand DNA ligase, a polymerase with strand displacement activity, and at least two nucleotides. The kit preferably comprises CircLigase as a ligase and a Phi29 Polymerase as the polymerase. CfDNA holds significant diagnostic potential due to its unique properties and the wealth of information it carries. Released into the fluid streams such as the bloodstream and urine from various cells, including tumor cells, apoptotic cells, and normal cells, cfDNA represents a non-invasive source of genetic material. Its diagnostic utility lies, among others, in detecting specific genetic alterations, such as mutations, methylation patterns, and copy number variations, reflecting the genomic landscape of tissues of origin. In cancer diagnostics, the analysis of cfDNA allows for the identification of tumor-specific mutations, enabling early detection, monitoring of treatment response, and the assessment of minimal residual disease. Furthermore, cfDNA can be indicative of other pathological conditions, such as prenatal abnormalities or autoimmune disorders. The ability to obtain genetic information through a simple blood draw makes cfDNA a promising biomarker for diverse diagnostic applications, emphasizing its potential to revolutionize personalized medicine and enhance the precision of disease diagnosis and monitoring The sample comprising circulating fluid is preferably a blood sample such as a serum sample, a urine sample, a cerebrospinal fluid sample, a lymph fluid sample, or a saliva sample. The sample is preferably a blood sample, preferably a serum sample. The present application also provides inventions as worded in the following aspects and description thereof. The terms used in the various aspects are the same as used in other parts of the description and have the same meaning unless otherwise specifically specified. Aspects 1. A method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of preferably linear DNA molecules; - incubating said sample with a restriction enzyme that is sensitive to methylation at the recognition site for said restriction enzyme and subsequently sequencing the DNA molecules in the sample thereby determining sequences of the DNA molecules; and determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules. 2. The method of aspect 1, wherein sequences comprising said restriction enzyme site or a part thereof at a terminal end are aligned. 3. The method of aspect 1 or aspect 2, wherein the methylation status of more than one recognition enzyme site for said restriction enzyme is determined in said plurality of DNA molecules. 4. The method of any one of aspects 1-3, further comprising removing DNA molecules that were cut by said restriction enzyme from said sample prior to said sequencing the DNA molecules in the sample. 5. The method of aspect 4, wherein the majority of the plurality of DNA molecules do not have a phosphate group at their 5’ ends prior to said incubation with said restriction enzyme. 6. The method of aspect 5, wherein phosphate groups at the 5’-ends of the plurality of DNA molecules have been removed by treating the plurality of DNA molecules with a phosphatase that removes phosphate groups from the 5’-ends of DNA molecules. 7. The method of any one of aspects 4-6, further comprising incubating the sample with a 5’-phosphate dependent exonuclease subsequent to and / or at the same time as said incubation with said restriction enzyme. 8. The method of aspect 7, wherein the 5’-phosphate dependent exonuclease is Lambda exonuclease. 9. The method of any one of aspects 4-8, further comprising incubating the sample with a single strand nucleotide specific nuclease subsequent to and / or at the same time as said incubation with said restriction enzyme. 10. The method of aspect 9, wherein the single strand nucleotide specific nuclease is Nuclease P1. 11. The method of any one of aspects 1-10, further comprising adding sequence identifiers to DNA molecules in the sample prior to sequencing of the DNA molecules. 12. The method of aspect 11, wherein the sequence identifiers are added by providing a sample comprising the plurality of linear DNA molecules; - contacting said sample with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single- stranded DNA molecules comprising the sequence identifiers are self-ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products are produced each comprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - and wherein the sequencing of the DNA molecules comprising sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - and wherein determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules comprises identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier. 13. The method of aspect 11 or aspect 12, further comprising aligning respective sequences of a single-stranded DNA molecule using the sequence identifier. 14. The method of any one aspects 11-13, wherein the provided DNA sample comprises a plurality of linear single-stranded DNA molecules and a plurality of circular single- stranded DNA molecules. 15. The method of any one of aspects 11-14, wherein the sequence identifiers are unique sequence identifiers. 16. The method of any one of aspects 1-15, wherein the sample comprises circulating free DNA (cfDNA). 17. The method of any one of aspects 1-16, wherein the sequencing is performed using a short read sequencing method, preferably a next-generation sequencing method involving massively parallel sequencing, more preferably using the Illumina Sequencing Platform. 18. The method of any one of aspects 12-17, wherein said two or more different nucleotides are three or more, preferably four, different nucleotides. 19. The method of any one of aspects 12-18, wherein the single-strand DNA ligase is a CircLigaseTM. 20. The method of any one of aspects 12-19, wherein the polymerase with strand displacement activity is a DNA polymerase with high processivity. 21. The method of aspect any one of aspects 12-20, wherein the polymerase is selected from the group consisting of Phi29 DNA Polymerase, Bst DNA Polymerase, and Vent (exo-) DNA Polymerase. 22. The method of aspect 21, wherein the Phi29 DNA Polymerase is EquiPhi29TM DNA polymerase. 23. The method of any one of aspects 12-22, wherein the primer for the RCA is a specific primer. 24. A kit for sequencing a plurality of DNA molecules, wherein the kit comprises a TdT, a single-strand DNA ligase, a polymerase with strand displacement activity, and at least two nucleotides. 25. The kit of aspect 24, wherein the ligase is a CircLigaseTM, and the polymerase with strand displacement activity is a Phi29 Polymerase. - sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier. 26. A method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of linear DNA molecules; - contacting said sample with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of the linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single-stranded DNA molecules comprising the sequence identifiers are self-ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products is produced each comprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier. 27. The method of aspect 26, further comprising aligning respective sequences of a single-stranded DNA molecule using the sequence identifier. 28. The method of aspect 26 or aspect 27, wherein the provided DNA sample comprises a plurality of linear single-stranded DNA molecules and a plurality of circular single- stranded DNA molecules. 29. The method of any one of aspects 26-28, wherein the sequence identifiers are unique sequence identifiers. 30. The method of any one of aspects 26-29, comprising contacting a sample comprising a plurality of linear RNA molecules with a reverse transcriptase under conditions in which RNA molecules are copied into cDNA molecules, thereby producing said sample comprising a plurality of linear single-stranded DNA molecules. 31. The method of any one of aspects 26-29, wherein the sample comprises circulating free DNA (cfDNA). 32. The method of any one of aspects 26-31, wherein the sequencing is performed using a short read sequencing method, preferably a next-generation sequencing method involving massively parallel sequencing, more preferably using the Illumina Sequencing Platform. 33. The method of any one of aspects 26-32, wherein said two or more different nucleotides are three or more, preferably four, different nucleotides. 34. The method of any one of aspects 26-33, wherein the single-strand DNA ligase is a CircLigaseTM. 35. The method of any one of aspects 34, wherein the polymerase with strand displacement activity is a DNA polymerase with high processivity. 36. The method of any one of aspects 26-35, wherein the polymerase is selected from the group consisting of Phi29 DNA Polymerase, Bst DNA Polymerase, and Vent (exo-) DNA Polymerase. 37. The method of aspect 36, wherein the Phi29 DNA Polymerase is EquiPhi29TMDNA polymerase. 38. The method of any one of aspects 26-37, wherein the primer for the RCA is a specific primer. 39. The method of aspect 11, further comprising sequencing nucleic acid molecules according to a method of any one of the aspects 26-37. 40. The method of aspects 1-39, comprising detecting the methylation status of essentially all occurrences of said restriction enzyme site for said restriction enzyme on said plurality of DNA molecules. Aspects 1-10 are method wherein the methylation status of one or more DNA molecules is determined by determining the sequence of the DNA molecules in a situation wherein these had been previously exposed to a restriction enzyme that is sensitive to the presence or absence of methylation on the DNA. Aspects 25-39 are directed towards sequencing DNA molecules using sequence identifiers to be able to, among others, distinguish sequences derived from different DNA molecules with essentially the same sequence. Essentially identical sequences with different sequence identifiers are typically derived from different DNA molecules. This feature may, for instance, be useful in quantification methods. A non-limiting example of such usefulness is in determining the prevalence mutant DNA molecules. This is, for instance, of interest in circulating free DNA where prevalence may potentially be used to correlate a more abundant presence of copies to certain disease states such as cancer or infection or alternatively the tissue of origin. These theoretical examples are only given as hypothetical examples to demonstrate that prevalence is an interesting and useful feature in itself. Another non-limiting example of its usefulness is in the quantification of methylation at a certain position on DNA molecules in a sample. This useful feature is addressed in aspects 11-25 where the methylation detection is combined with adding sequence identifiers to the DNA molecules prior to sequencing. In aspect 11, any type of sequence identifiers is claimed as any type will do to distinguish different molecules. In aspects 12-25 and 38 the methylation detection aspects are combined with the unique methods for adding sequence identifiers DNA molecules of aspects 26-37. The sample may be a sample directly derived from a clinical or biological source or the sample may be a processing stage thereof for instance such as the samples contacted and processed in the various steps of the claims and / or aspects. The sample of aspect 26 can, for example, be directly derived from a clinical or biological source but the sample can also be the result of exposure of such a sample to a methylation sensitive restriction enzyme such as described in aspects 1-10 prior to determining the sequence thereof. The fact the sample typically comprises a plurality of linear DNA molecules does not exclude the fact that circular DNA molecules may also be present. Restriction enzymes, also known as restriction endonucleases, are enzymes that recognize specific DNA sequences and cleave the DNA at or near these sequences. The enzymes play a crucial role in molecular biology by cutting DNA molecules at precise locations, enabling scientists to manipulate and study DNA in various ways. The sequences that restriction enzymes recognize are typically palindromic, meaning they read the same backward and forward. Such sequences are referred to as restriction enzyme sites. After cutting the DNA, the resulting fragments can be analyzed, manipulated, or used in various molecular biology techniques, such as cloning and DNA sequencing. Methylation-sensitive restriction enzymes are a subclass of restriction enzymes that recognize specific DNA sequences and cleave DNA, but their activity is influenced by the methylation status of the DNA. Methylation is a chemical modification where a methyl group (CH3) is added to certain bases in the DNA molecule, typically cytosine bases in CpG dinucleotides. Methylation-sensitive restriction enzymes can be divided into two main categories, e.g. enzymes that cleave DNA only when specific sites are methylated. If the recognition site is unmethylated, these enzymes do not cleave the DNA and enzymes that cleave DNA at specific sequences only when specific methylation at the recognition site is absent. Methylation of the restriction site protects the DNA from cleavage by these enzymes. Non-limiting examples of methylation-sensitive restriction enzymes are: -HpaII: This enzyme recognizes the sequence 5'-CCGG-3' but only cleaves if theinternal cytosine (C) is unmethylated; -MspI: Similar to HpaII, MspI recognizes the sequence 5'-CCGG-3' and cleaves ifthe internal cytosine is unmethylated; -BstUI: This enzyme recognizes the sequence 5'-CGCG-3' and cleaves if thecytosine (C) is unmethylated. -HhaI: HhaI recognizes the sequence 5'-GCGC-3' and cleaves only if the internalcytosine (C) is unmethylated. -SacII: SacII recognizes the sequence 5'-CCGCGG-3' and cleaves DNA at this site.Methylation of the cytosine (C) within the recognition site inhibits cleavage by SacII. The methylation-sensitive restriction enzyme in methods of the invention is preferably a methylation-sensitive restriction enzyme that digests non-methylated versions of the restriction enzyme digestion site, such as , but not limited to: HpaII; MspI; BstUI, HhaI or SacII. DNA methylation is a fundamental epigenetic mechanism in which methyl groups are added to the DNA molecule. This modification doesn't change the DNA sequence itself but can affect how genes are expressed. Methylation typically occurs on cytosine bases, specifically at cytosine-guanine dinucleotide sequences (CpG sites), although it can also occur at non-CpG sites in certain contexts, such as during development. The methylation status of DNA refers to the presence or absence of methylation at a certain position; to the extent of such methylation at a certain site; or to the pattern and extent of methylation across the genome of an organism or a specific set of genes. It is sometimes described as methylated or unmethylated, or hyper- or hypo-methylated, or the like. The methylation status of DNA can be dynamic and can change in response to various environmental factors, developmental stages, and disease conditions. It plays critical roles in regulating gene expression, genomic imprinting, X-chromosome inactivation, and maintaining genome stability. Additionally, aberrant DNA methylation patterns have been implicated in various diseases, including cancer and neurological disorders. The sequences can be used to determine where the restriction enzyme has cut or not. The presence of a complete copy of the restriction enzyme site in the determined sequence indicates that the methylation-sensitive restriction enzyme did not cut this particular occurrence of the site in the sample. The presence of the resulting parts of digestion of the restriction enzyme site by the restriction enzyme at the ends of determined sequences indicates that the methylation-sensitive restriction enzyme did cut this particular occurrence of the site in the sample. These ends, when aligned with other sequences with these terminal parts, results in a concentration of sequencing stops at the cutting site resulting in what is sometimes referred to as a scar in the alignment (see Figure 3). It is possible to enrich the sample for occurrences where the methylation-sensitive restriction enzyme did not cut the recognition site. In a particular embodiment this is done by providing the plurality of DNA molecules as a plurality of unphosphorylated DNA molecules. In this embodiment the majority of the plurality of DNA molecules do not have a phosphate group at their 5’ ends prior to said incubation with said restriction enzyme. Non- limiting examples of such a plurality of unphosphorylated DNA molecules are cfDNA; and sheared genomic DNA, or restriction enzyme digested DNA that have undergone a 5’- phosphate removal step. Shearing or digestion with a restriction enzyme can be done to reduce the size of the fragments when for instance genomic DNA is tested. Various enzymatic and non-enzymatic methods exist for the specific removal of 5’-phosphate groups, such a hot sodium chloride treatment, calf intestinal alkaline phosphatase (CIP): CIP is a widely used phosphatase in molecular biology, particularly for dephosphorylating DNA molecules; antarctic phosphatase (AP): an enzyme derived from a strain of Antarctic bacteria or shrimp alkaline phosphatase (SAP). In a preferred embodiment the phosphate groups at the 5’-ends of the plurality of DNA molecules have been removed by treating the plurality of DNA molecules with a phosphatase that removes phosphate groups from the 5’- ends of DNA molecules. The phosphatase is preferably CIP. This treatment is typically performed before incubation with the methylation- sensitive restriction enzyme. Cutting of the DNA with the enzyme will result in new 5’-ends that do have a phosphate group. This difference is exploited to enrich the sample for DNA fragments that were not cut by the methylation-sensitive restriction enzyme. The sample is preferably incubated with a 5’-phosphate dependent exonuclease subsequent to and / or at the same time as said incubation with said restriction enzyme. This will result in digestion of DNA molecules that have a 5’-phosphate group and thus enrich the sample for resistant molecules that were not digested with by the restriction enzyme. This embodiment is particularly useful for cases where the majority of the plurality of DNA molecules in the sample comprises on average only one recognition site for the methylation-sensitive restriction enzyme. If genomic or other large DNA fragments are used these can be fragmented to obtain suitably sized fragments. For instance as indicated herein above by shearing or sonication or by digestion with another restriction enzyme. Exposure to the methylation-sensitive restriction enzyme will result in undigested fragments that have only unphosphorylated 5’-ends and result in digested fragments that are unphosphorylated on one 5’-end and phosphorylated on the 5’-end that is the result of the successful digestion by the enzyme. These latter ends are sensitive for digestion by the 5’-phosphate dependent exonuclease and will be removed. The removed fragments will not be sequenced. The result will be that the sample is enriched for DNA fragments that were not cut by the methylation-sensitive restriction enzyme. The exonuclease step not only enriches for certain types of fragments but also creates a gap in the sequence reads that is more easily noticed than without the exonuclease treatment. The 5’-phosphate dependent exonuclease is preferably Lambda exonuclease. This enzyme digests only strands that have a 5’- phosphate. Digestion by Lambda exonuclease of a double stranded DNA fragment that has one phosphorylated 5’-end and one unphosphorylated 5’-end will result in digestion of the strand with the 5’-phosphate group while the other unphosphorylated strand is left intact. This single strand is preferably also removed thereby increasing the enrichment further. The enrichment is preferably further improved by incubating the sample with a single strand nucleotide specific nuclease subsequent to and / or at the same time as said incubation with said restriction enzyme. This will remove the single strand and enrich the sequences for sequences of DNA fragments where not digested by the methylation-sensitive restriction enzyme. The single strand nucleotide specific nuclease is preferably Nuclease P1. Specially preferred samples for the enrichment are sample of small DNA fragments such as but not limited to cfDNA. In another embodiment, the DNA can be treated with a methylation-sensitive restriction enzyme, without prior de-phosphorylation and without subsequent exonuclease treatment. The DNA can then optionally be tagged with random nucleotides using TdT, undergo melting to generate single-stranded DNA, and self-circularized using CircLigase. The circular DNA is then amplified using rolling circle amplification and sequenced. The effect of the methylation-sensitive restriction enzyme can then be deducted by the analysis of the aligned reads. Reads coming from digested DNA will terminate at the restriction site, while undigested reads will span the restriction site. Figure 3. Sequencing results are more easily interpreted by adding sequence identifiers to DNA molecules in the sample prior to sequencing of the DNA molecules. It is therefore preferred to add sequence identifiers to DNA molecules in the sample prior to sequencing of the DNA molecules. Sequence identifiers can be used to distinguish sequences of the same molecule from an identical sequence originating from a different molecule because such a sequence will typically have a different sequence identifier. Sequence identifiers can also be used to improve the accuracy of the sequence information for instance in determining difference between a mutation in the original sample from a mutation / error that occurred during the sequencing reaction. Sequence identifiers can also be used to economize and sequence several samples at the same time. In such a case the sequence identifier can contain a part that is the same for all DNA molecules in a sample but different from other samples. In a preferred embodiment wherein the sequence identifiers are added by - contacting a sample comprising the plurality of linear DNA molecules with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single-stranded DNA molecules comprising the sequence identifiers are self- ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products are produced each comprising one or more copies of a circular single- stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - and wherein the sequencing of the DNA molecules comprising sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - and wherein determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules comprises identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier. The method preferably further comprises aligning respective sequences of a single- stranded DNA molecule using the sequence identifier. The provided DNA sample preferably comprises a plurality of linear single-stranded DNA molecules and a plurality of circular single-stranded DNA molecules. The sequence identifiers are preferably unique sequence identifiers. The sample preferably comprises circulating free DNA (cfDNA). The sequencing is preferably performed using a short read sequencing method, preferably a next-generation sequencing method involving massively parallel sequencing, more preferably using the Illumina Sequencing Platform. Said two or more different nucleotides are preferably three or more, preferably four, different nucleotides. The single- strand DNA ligase is preferably a CircLigaseTM. The polymerase with strand displacement activity is preferably a DNA polymerase with high processivity. The polymerase is preferably selected from the group consisting of Phi29 DNA Polymerase, Bst DNA Polymerase, and Vent (exo-) DNA Polymerase. The Phi29 DNA Polymerase is preferably EquiPhi29TM DNA polymerase. The primer for the RCA is preferably a specific primer. The kit for sequencing a plurality of DNA molecules, is preferably a kit that comprises a TdT, a single-strand DNA ligase, a polymerase with strand displacement activity, and at least two nucleotides. In a preferred embodiment the ligase is a CircLigaseTM, and the polymerase with strand displacement activity is a Phi29 Polymerase. For the purpose of clarity and a concise description features are described herein as part of the same or separate embodiments, however, it will be appreciated that the scope of the invention may include embodiments having combinations of all or some of the features described. EXAMPLES Example 1: Sequencing Complete cfDNA Fragments for Cancer Diagnostics The method of the present invention was applied to sequence cfDNA fragments obtained from a group of cancer patients. The objective was to capture the complete sequences of these fragments, including their terminal regions, which are often challenging to sequence but hold crucial information for cancer diagnostics. Procedure: cfDNA Extraction: cfDNA was extracted from blood samples of patients using a standard cfDNA extraction kit (QIA kit) using the manufacturers protocol. The extracted cfDNA predominantly consisted of short, double-stranded DNA fragments. Terminal Deoxynucleotidyl Transferase (TdT) Treatment: The cfDNA was treated with TdT (obtained from New England Biolabs) and a mixture of four different nucleotides (A, T, C, G; Sigma) using the manufacturers protocol. This step randomly added a sequence of nucleotides to the 3’ ends of each cfDNA fragment, serving as unique sequence identifiers. Preparation of Single-Stranded DNA: The double-stranded cfDNA was denatured for 10 minutes at 980C to generate single-stranded DNA molecules. Circularization: The modified cfDNA fragments were then treated with the single-strand DNA ligase CircLigase (Biosearch technologies) using the manufacturers protocol, leading to the circularization of these molecules. Rolling Circle Amplification (RCA): The circularized DNA fragments underwent RCA, using Phi29 DNA Polymerase (Thermofischer) using the manufacturers protocol, to produce a linear DNA molecules comprising multiple copies of a cfDNA fragment linked to a sequence identifier. Sequencing: The amplified products were sequenced using the Illumina Sequencing Platform, allowing for the identification of both the original cfDNA sequences and the unique sequence identifiers. This procedure is expected to result in: Complete Sequencing of cfDNA Fragments: The method enabled the sequencing of entire cfDNA fragments, including their terminal regions. This comprehensive approach revealed mutations and alterations at the edges of cfDNA fragments, which are typically missed by conventional methods. Utility of Terminal Sequences: The terminal regions of cfDNA provided valuable information, such as the fragmentation patterns characteristic of certain types of cancers, and helped in distinguishing between cfDNA from tumor and non-tumor origins. Integration with AI for Enhanced Diagnostics The sequences obtained, including the terminal regions, can be fed into an AI-based analytical model designed to identify patterns and mutations associated with various cancers. The AI model can then be trained to recognize cancer-specific signatures in the cfDNA sequences, including those present in the sequence edges. Including these features is expected to bring: Improved Diagnostic Accuracy: The AI model is expected to use the end-sequence information, together with other features, to distinguish between benign and malignant cfDNA samples. Personalized Treatment Strategies: The model also provided insights into the genetic makeup of individual tumors, facilitating personalized treatment planning. Conclusion The described method represents a significant advancement in the field of cancer diagnostics through liquid biopsies. The ability to sequence complete cfDNA fragments, including their terminal regions, combined with AI-based analysis, offers a promising approach to early cancer detection and personalized treatment strategies. Materials and Methods Materials ●Terminal Deoxynucleotidyl Transferase (NEB)● CircLigase (Biosearch)● EquiPhi29 DNA Polymerase (Thermofisher)● Exo-resistant Random Primers (Thermofisher)● CoCl2 2.5mM● dNTPs 10mM● ATP 10mM● MnCl2 50mM● DTT 100mM● TdT Buffer (10X)○ 500 mM Potassium Acetate○ 200 mM Tris-acetate○ 100 mM Magnesium Acetate● CircLigaseBuffer (10X)○ MOPS 500 mM○ KCl 100 mM○ MgCl2 50 mM○ DTT 10 mM● CutSmart Buffer (10X)○ 500 mM Potassium Acetate○ 200 mM Tris-acetate○ 100 mM Magnesium Acetate○ 1000 µg / ml BSADetailed Protocol 1) TdT tagging Reagent Volume (ul)cfDNA (0.5 ng / ul) 10 (5ng)TdT Buffer 1.5CoCl2 (2.5mM) 1.5dNTPs (10mM) 1.5Terminal Transferase 0.5 Incubate for 20 min at 370C Heat Inactivate for 10 min at 700C 2) cfDNA denaturing Reagent Volume (ul)Previous Reaction 15CircLigase Buffer 10X 3.5H2O 11.5Incubate for 10 min at 980C. Fast Cool Down immediately on ice. 3) ssDNA self-circularization To the previous reaction, add: Reagent Volume (ul)Previous Reaction 30ATP (10mM) 0.5MnCl2 (50mM) 3.5CircLigase 1Incubate for 1h at 600C. Inactivate at 800C for 10 min. 4) Rolling Circle Amplification Add 5ul of Exo-resistant Random Hexamers (10uM). Incubate for 10 min at 980C. Slow Cool Down to 220C Pre-mix the following reagents and then add them to the previous reaction Reagent Volume (ul)H2O 73CutSmart 10dNTPs (10mM) 10 DTT (100mM) 2EquiPhi29 Poly. 5Incubate at 420C for 3h, then keep at 40C Inactivate at 700C for 10 min 5) Sequencing This part differs based on the platform used, any platform can be used. The RCA product can be readily sequenced in case of long-read sequencing, or it can be fragmented to generate short reads. Follow the manufacturer's methods as they fit. Example 2 Figure 5 is the size fraction result of a Proof-of-concept experiment. TapeStation results. In this experiment, synthetic DNA was used to demonstrate the technology. The input DNA is composed of 3 non-phosphorylated sequences derived from the human MYC gene. The input DNA is composed of a methylated MYC-Promoter (M-P*), a non- methylated MYC-Promoter (M-P), and a MYC-Terminator without any CpG sites (M-T). Lanes H1 A2 B2. The input DNA is digested with a methylation-blocked restriction enzyme, BstUI. Lanes B1 C1 D1. BstUI digests only the unmethylated MYC-Promoter, as supposed to, thus generating 5'- Phosphorylated ends only on that DNA. Finally, a blend of exonucleases is used to deplete the digested DNA. In this case, Lambda and P1 nucleases were used. Lambda Exonuclease digests 5'-P-dsDNA into ssDNA, and P1 nuclease digests all ssDNA. As expected, after the 2-step process, only the DNA containing unmethylated CpG islands (M-P) was depleted. Figure 6 depicts the synthetic sequences used and figure 7 is a screenshot of the sequencing result of the POC experiment. Materials and Methods Materials ●Terminal Deoxynucleotidyl Transferase (NEB)● T4 Polynucleotide Kinase (NEB)● CircLigase (Biosearch)● EquiPhi29 DNA Polymerase (Thermofisher)● Exo-resistant Random Primers (Thermofisher)● CoCl2 2.5mM● dNTPs 10mM● ATP 10mM● MnCl2 50mM● DTT 100mM● TdT Buffer (10X)○ 500 mM Potassium Acetate○ 200 mM Tris-acetate○ 100 mM Magnesium Acetate● CircLigaseBuffer (10X)○ MOPS 500 mM○ KCl 100 mM○ MgCl2 50 mM○ DTT 10 mM● CutSmart Buffer (10X)○ 500 mM Potassium Acetate○ 200 mM Tris-acetate○ 100 mM Magnesium Acetate○ 1000 µg / ml BSADetailed Protocol 1) TdT tagging Reagent Volume (ul)cfDNA (0.5 ng / ul) 10 (5ng)TdT Buffer 1.5CoCl2 (2.5mM) 1.5dNTPs (10mM) 1.5Terminal Transferase 0.5Incubate for 20 min at 370C Heat Inactivate for 10 min at 700C 2) Phosphorylation Reagent Volume (ul)Previous Reaction 15DTT (100mM) 1ATP (10mM) 1.5T4 Polynucleotide Kinase 1Incubate for 20 min at 370C. 3) cfDNA melting Reagent Volume (ul)Previous Reaction 18.5CircLigase Buffer 10X 3.5H2O 8Incubate for 10 min at 980C. Fast Cool Down immediately on ice. 4) ssDNA self-circularization To the previous reaction, add: Reagent Volume (ul)Previous Reaction 30ATP (10mM) 0.5MnCl2 (50mM) 3.5CircLigase 1Incubate for 1h at 600C. Inactivate at 800C for 10 min. 5) Rolling Circle Amplification Add 5ul of Exo-resistant Random Hexamers (10uM). Incubate for 10 min at 980C. Slow Cool Down to 220C Pre-mix the following reagents and then add them to the previous reaction Reagent Volume (ul)H2O 73CutSmart 10 dNTPs (10mM) 10 DTT (100mM) 2EquiPhi29 Poly. 5Incubate at 420C for 3h, then keep at 40C Inactivate at 700C for 10 min 5) Sequencing This part differs based on the platform used, any platform can be used. The RCA product can be readily sequenced in case of long-read sequencing, or it can be fragmented to generate short reads. Follow the manufacturer's methods as they fit. Example 3 Methylation detection method comparison. Genomic DNA was extracted from the blood of a donor. The DNA was fragmented via sonication to 300bp. The overhangs of the ends of the resulting molecules were repaired (blunt ended) by means of end repair module from New England Biolabs and subsequently dephosphorylated. The fragmented DNA was then split into two aliquots. One aliquot was sent to a third-party company for bisulfite sequencing, while the other was processed using the method disclosed in Example 2. The sequencing data obtained was analyzed by direct comparison on IGV, where images 8- 12 were taken as a screenshot. Figure 8 shows a zoom-in of figure 12. It shows the difference in accuracy of the resulting sequences from the method of the invention when compared to bisulfite methylation detection. Figure 9 show the sequence gaps obtained after treatment with the methylation sensitive enzyme BstUI and the gaps obtained with the bisulfite treatment (see legend). Figure 9 shows the result for a part of the chromosome X. Figures 10 and 11 show the results of the same experiment but now for parts of chromosomes 20 and 15 respectively.
Claims
CLAIMS 1. A method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of preferably linear DNA molecules; - incubating said sample with a restriction enzyme that is sensitive to methylation at the recognition site for said restriction enzyme and subsequently sequencing the DNA molecules in the sample thereby determining sequences of the DNA molecules; and determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules.
2. The method of claim 1, comprising detecting the methylation status of essentially all occurrences of said restriction enzyme site for said restriction enzyme on said plurality of DNA molecules.
3. The method of claim 1 or claim 2, wherein sequences comprising said restriction enzyme site or a part thereof at a terminal end are aligned.
4. The method of any one of claims 1-3, further comprising removing DNA molecules that were cut by said restriction enzyme from said sample prior to said sequencing the DNA molecules in the sample.
5. The method of claim 4, wherein the majority of the plurality of DNA molecules do not have a phosphate group at their 5’ ends prior to said incubation with said restriction enzyme.
6. The method of claim 5, wherein phosphate groups at the 5’-ends of the plurality of DNA molecules have been removed by treating the plurality of DNA molecules with a phosphatase that removes phosphate groups from the 5’-ends of DNA molecules.
7. The method of any one of claims 4-6, further comprising incubating the sample with a 5’-phosphate dependent exonuclease subsequent to and / or at the same time as said incubation with said restriction enzyme.
8. The method of claim 7, wherein the 5’-phosphate dependent exonuclease is Lambda exonuclease.
9. The method of any one of claims 4-8, further comprising incubating the sample with a single strand nucleotide specific nuclease subsequent to and / or at the same time as said incubation with said restriction enzyme.
10. The method of claim 9, wherein the single strand nucleotide specific nuclease is Nuclease P1.
11. The method of any one of claims 1-10, further comprising adding sequence identifiers to DNA molecules in the sample prior to sequencing of the DNA molecules.
12. The method of claim 11, further comprising sequencing nucleic acid molecules according to a method of any one of claims 15-29.
13. The method of claim 11, wherein the sequence identifiers are added by providing a sample comprising the plurality of linear DNA molecules; - contacting said sample with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single- stranded DNA molecules comprising the sequence identifiers are self-ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products are produced eachcomprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - and wherein the sequencing of the DNA molecules comprising sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and - and wherein determining from said sequences the methylation status of a restriction enzyme site for said restriction enzyme on said plurality of DNA molecules comprises identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier.
14. The method of any one of claims 1-6, wherein the sample comprises circulating free DNA (cfDNA).
15. A method for sequencing nucleic acid molecules, comprising: - providing a sample comprising a plurality of linear DNA molecules; - contacting said sample with a terminal deoxynucleotidyl transferase (TdT) and a mixture of two or more different nucleotides under conditions in which nucleotides in the mixture are sequentially and randomly added to the 3’ termini of the linear DNA molecules by the TdT and denaturing, when present, double stranded DNA in said sample, thereby generating linear single-stranded DNA molecules comprising sequence identifiers at their 3’ termini; - contacting the linear single-stranded DNA molecules comprising the sequence identifiers with a single-stranded DNA ligase under conditions in which the ends of linear single-stranded DNA molecules comprising the sequence identifiers are self-ligated thereby generating a plurality of circular single-stranded DNA molecules each comprising a DNA molecule and a sequence identifier; - contacting the plurality of circular single-stranded DNA molecules with a polymerase with strand displacement activity and a primer and performing a rolling circle amplification (RCA) under conditions in which a plurality of amplification products is produced each comprising one or more copies of a circular single-stranded DNA molecule from the plurality of circular single-stranded DNA molecules; - sequencing the plurality of amplification products thereby generating a plurality of sequences of the plurality of amplification products; and- identifying one or more sequences of a single-stranded DNA molecule in the sample using a sequence identifier.
16. The method of claim 15, further comprising aligning respective sequences of a single-stranded DNA molecule using the sequence identifier.
17. The method of claim 15 or claim 16, wherein the provided DNA sample comprises a plurality of linear single-stranded DNA molecules and a plurality of circular single- stranded DNA molecules.
18. The method of any one of claims 15-17, wherein the sequence identifiers are unique sequence identifiers.
19. The method of any one of claims 15-18, comprising contacting a sample comprising a plurality of linear RNA molecules with a reverse transcriptase under conditions in which RNA molecules are copied into cDNA molecules, thereby producing said sample comprising a plurality of linear single-stranded DNA molecules.
20. The method of any one of claims 15-19, wherein the sample comprises circulating free DNA (cfDNA).
21. The method of any one of claims 15-20, wherein the sequencing is performed using a short read sequencing method, preferably a next-generation sequencing method involving massively parallel sequencing, more preferably using the Illumina Sequencing Platform.
22. The method of any one of claims 15-21, wherein said two or more different nucleotides are three or more, preferably four, different nucleotides.
23. The method of any one of claims 15-22, wherein the single-strand DNA ligase is a CircLigaseTM.
24. The method of any one of claims 15-23, wherein the polymerase with strand displacement activity is a DNA polymerase with high processivity.
25. The method of claim any one of claims 15-24, wherein the polymerase is selected from the group consisting of Phi29 DNA Polymerase, Bst DNA Polymerase, and Vent (exo-) DNA Polymerase.
26. The method of claim 25, wherein the Phi29 DNA Polymerase is EquiPhi29TMDNA polymerase.
27. The method of any one of claims 15-26, wherein the primer for the RCA is a specific primer.
28. A kit for sequencing a plurality of DNA molecules, wherein the kit comprises a TdT, a single-strand DNA ligase, a polymerase with strand displacement activity, and at least two nucleotides.
29. The kit of claim 28, wherein the ligase is a CircLigaseTM, and the polymerase with strand displacement activity is a Phi29 Polymerase.
Citation Information
Patent Citations
A high throughput sequencing method and kit
EP3798318A1
Methods and compositions for generating and amplifying DNA libraries for sensitive detection and analysis of DNA methylation
US20180030527A1
Method of estimating the amount of a methylated locus in a sample
WO2016024182A1
Method for nucleic acid analysis directly from an unpurified biological sample
WO2016053638A1
Compositions and methods for detecting rare sequence variants
WO2018035170A1