Method for preparing nucleic acid molecules for sequencing

By connecting compatible double-stranded DNA molecules and rolling loop amplification technology, the accuracy and error rate of long DNA sequences in existing sequencing methods are solved, and efficient and accurate long read sequencing is achieved, suitable for sequencing of multiple nucleic acid sources.

CN120350093APending Publication Date: 2025-07-22UMC UTRECHT HLDG BV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510490261.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-12-11
Filing Date
2018-12-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing next-generation sequencing methods have high error rates, making it difficult to accurately determine structural variations and complex repeat structures in long DNA sequences, and short read technology is difficult to correct sequencing errors.

Method used

By providing ligation of compatible double-stranded backbone DNA molecules with target DNA molecules, a restriction enzyme and a recombinant enzyme are ligated to form a DNA loop, and then a multi-conjunctive DNA molecule is generated by rolling loop amplification for sequencing, increasing the number of sequencings to correct the error.

Benefits of technology

It improves the accuracy and length of sequencing, can effectively reduce the error rate, and is suitable for long-read sequencing, especially single-molecular sequencers such as Pacific Biosciences and Oxford Nanopore systems, and is suitable for sequencing of multiple nucleic acid sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005365358440000161
    Figure BDA0005365358440000161
  • Figure BDA0005365358440000162
    Figure BDA0005365358440000162
  • Figure BDA0005365358440000163
    Figure BDA0005365358440000163
Patent Text Reader

Abstract

The present invention relates to means and methods for preparing double stranded target DNA molecules for sequencing. In an embodiment, a double-stranded backbone DNA molecule is provided comprising a 5'end and a 3 'end, the 5' end and the 3 'end being provided to be ligation-compatible with the 5' end and the 3 'end of the target DNA; forming a first restriction enzyme recognition site when self-ligation; and the device is in a self-connection form. The method may include, if not yet present, providing the target DNA with a 5'end and a 3 'end such that the 5' end and the 3 'end are in a form that prevents self-ligation and are ligated compatible with the 5' end and the 3 'end of the backbone DNA. The method further includes linking the target DNA to the backbone DNA in the presence of a ligase and a first restriction enzyme that cleaves the first restriction enzyme recognition site, thereby producing at least one DNA loop comprising a backbone DNA molecule and a target DNA molecule. The linear DNA may be removed at this time, and then a multi-conjoined DNA molecule comprising an ordered array of copies of the at least one DNA loop is produced by rolling circle amplification, which multi-conjoined DNA molecule may be sequenced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to means and methods for determining the sequence of a nucleic acid molecule. Specifically, the present invention relates to a method that utilizes rolling circle amplification of a nucleic acid molecule whose sequence is to be determined. Background Art

[0002] Over time, sequencing methods have been evolving. The old Sanger sequencing method has been replaced by the currently commonly used next-generation sequencing (NGS) methods. Recently, these methods have been reviewed in the literature (Goodwin et al 2016; Nature Reviews|Genetics Volume 17: pp 333-351: doi:10.1038 / nrg.2016.49). The most commonly used NGS methods rely on the sequencing of short stretches of DNA. The sequencing techniques for short stretches of DNA have inherent errors. The errors can be reduced by independently sequencing multiple copies of the same target sequence. However, for each individual sequence read, it is not possible to determine whether a change represents an error or a true mutation. Cumulative evidence spanning several independent sequence reads allows filtering of mutations introduced during amplification and errors in sequencing longer target DNA (which can also be sequenced using short-read methods). This is typically done by sequencing overlapping fragments that can be aligned to generate an assembled longer sequence. This so-called short-read paired-end technique has been very successful in sequencing large target nucleic acids and is very useful for various genome projects. Genome projects have revealed that genomes are highly complex, with many long repetitive elements, copy number alterations, and structural variations. Many of these elements are so long that the short-read paired-end technique is defective in resolving them. Long-read sequencing provides read lengths of more than several kilobases and allows the resolution of these large structural features throughout the genome. Two popular platforms for long-read sequencing are the Pacific Biosciences systems (RSII and Sequel) and the Oxford Nanopore systems (MK1 MinION and PromethION). Both are single-molecule sequencers. Both platforms allow read lengths of more than 55 kb and even longer. However, these systems have error rates even higher than those of next-generation (second-generation) sequencers. These errors can be reduced by increasing the number of times the same target nucleic acid is sequenced (Goodwin et al 2016; doi:10.1038 / nrg.2016.49).

[0003] The present invention provides a new protocol for preparing nucleic acid molecules for sequencing. Summary of the Invention

[0004] Embodiments of the present invention provide a method for preparing a double-stranded target DNA molecule for sequencing, comprising:

[0005] - Provide a double-stranded backbone DNA molecule comprising a 5' end and a 3' end, the double-stranded backbone DNA molecule:

[0006] - Is ligation-compatible with the 5' end and the 3' end of the target DNA;

[0007] - Forms a first restriction enzyme recognition site when self-ligating;

[0008] - Is in a form capable of self-ligating; and

[0009] - If not already present, provide 5' and 3' ends for the target DNA in a form that prevents self-ligation and is ligation-compatible with the 5' and 3' ends of the backbone DNA;

[0010] The method further comprises:

[0011] - Ligate the target DNA to the backbone DNA in the presence of a ligase and a first restriction enzyme that cleaves the first restriction enzyme recognition site, thereby producing at least one DNA loop comprising a backbone DNA molecule and a target DNA molecule;

[0012] - Optionally remove linear DNA;

[0013] - Generate concatemeric DNA molecules comprising an ordered array of copies of the at least one DNA loop by rolling circle amplification; and

[0014] - Sequence the at least one concatemer.

[0015] - Sequence the at least one concatemer.

[0016] There is also provided a collection of DNA molecules (backbones) having a length of 50 to 1000 nucleotides, the DNA molecules comprising a 5' end and a 3' end, the 5' end comprising a portion of a first restriction enzyme recognition site at its outermost end, the 3' end comprising another portion of the first restriction enzyme recognition site at its outermost end, and the 5' end and the 3' end being ligation-compatible with each other and capable of forming a restriction enzyme (first restriction enzyme) recognition site when self-ligating, and wherein each backbone comprises:

[0017] A linker;

[0018] Optionally an identifier sequence (barcode) that is different from the sequences of the identifiers of other backbones in the collection;

[0019] Optionally a second identifier unique to the collection of backbone molecules; and

[0020] Optionally a restriction enzyme cleavage site for a nicking enzyme.

[0021] There is further provided a method for determining the sequences of a collection of nucleic acid molecules, comprising:

[0022] - Provide a double-stranded target DNA molecule having a 5' end and a 3' end, wherein protruding adenine residues are present at the 3' ends of both strands of the DNA molecule;

[0023] - Provide a collection of double-stranded backbone DNA molecules, said double-stranded backbone DNA molecules comprising 5' and 3' ends that are ligation-compatible with the 5' and 3' ends of the target DNA;

[0024] The method further comprises:

[0025] - Ligating the target DNA to the backbone in the presence of a ligase, thereby generating a DNA circle comprising the backbone and the target DNA molecule;

[0026] - Optionally removing linear DNA;

[0027] - Generating a concatemer comprising an ordered array of copies of at least two of said DNA circles by rolling circle amplification; and

[0028] - Sequencing the concatemer.

[0029] There is further provided a method for determining the sequence of a collection of nucleic acid molecules, comprising:

[0030] - Provide a double-stranded target DNA molecule having recombinase recognition sites specific for a target site-specific recombinase at the 5' end and the 3' end;

[0031] - Provide a backbone comprising said recognition sites separated by DNA comprising a linker;

[0032] - Incubate the target DNA molecule with the backbone in the presence of the target site-specific recombinase, preferably Cre recombinase, FLP recombinase or bacteriophage λ (lambda) integrase, thereby generating a DNA circle comprising the backbone and the target DNA molecule;

[0033] - Optionally removing linear DNA; and

[0034] - Generating a concatemer comprising an ordered array of copies of at least two of said DNA circles by rolling circle amplification; and

[0035] - Sequencing the concatemer.

[0036] In a preferred embodiment, the backbone is a loop comprising two recombinase recognition sites, the recombinase recognition sites being separated on one side by DNA comprising a linker and on the other side by DNA encoding a restriction enzyme recognition site, and wherein the restriction enzyme cleavage site is the only recognition site for the restriction enzyme in the backbone. In this embodiment, the method preferably further comprises digesting the DNA with the restriction enzyme after the recombination, and subsequently removing the linear DNA, followed by generating the concatemer.

[0037] There is further provided a method for determining the sequence of a collection of nucleic acid molecules, comprising:

[0038] - providing a double-stranded target DNA molecule having at its 5' and 3' ends recombinase recognition sites specific for a target site-specific recombinase;

[0039] - providing a collection of double-stranded circular backbone DNA molecules comprising the recombinase recognition sites and a linker;

[0040] The method further comprises:

[0041] - incubating the target DNA molecule with the backbone in the presence of a target site-specific recombinase for the recognition sites, thereby generating a DNA loop comprising the backbone and the target DNA molecule;

[0042] - optionally removing linear DNA; and

[0043] - generating a concatemer comprising an ordered array of copies of at least two of the DNA loops by rolling circle amplification; and

[0044] - sequencing the concatemer.

[0045] There is further provided a kit comprising one or more backbones. Detailed embodiments

[0046] The means and methods described herein can determine the sequence of the same target DNA molecule multiple times. This can be used as a means of error correction. This is different from the classical second-generation sequencing method of correcting errors by sequencing multiple independent molecules covering the same genomic locus. In this case, each read typically represents one sequencing event for one molecule. With the method of the present invention, a single (target) molecule is repeatedly replicated, so one read represents multiple sequencing events for the same molecule.

[0047] The target nucleic acid is usually double-stranded DNA. Single-stranded DNA or RNA whose sequence needs to be determined can be easily converted into double-stranded DNA by methods known in the art. Such methods include, but are not limited to, cDNA synthesis, reverse transcriptase (RT) polymerase chain reaction (PCR), PCR, random primer extension, etc. Before performing the method, the target DNA is linear or made linear.

[0048] The backbone is usually double-stranded DNA. In the method of ligating the target DNA to the backbone using restriction enzymes, the backbone is usually linear or made linear before or during the method. In the method of inserting the target DNA into the backbone using a target site-specific recombinase, the backbone can be linear or made circular before or during the method.

[0049] Herein, self-ligation is defined as ligating the 5' end of one and the same nucleic acid molecule to the 3' end.

[0050] The 5' and 3' ends of the target DNA are selected such that the 5' and 3' ends of the target DNA are ligation-compatible with the 5' and 3' ends of the backbone used in the reaction. Ligation-compatible means that the ligation of the ends to each other produces double-stranded DNA with correctly paired nucleotides and no gap at the ligation junction. Of course, a gap can be introduced later to allow initiation of the RCA reaction. Blunt ends are ligation-compatible with other blunt ends. DNA with sticky (also referred to as "cohesive") ends is ligation-compatible with other sticky ends if the protruding DNA strands can be annealed together without unpaired bases. This is usually the case when the ends have complementary sequences. "Ligation-compatible ends" are also referred to in the art as "compatible ends" or "compatible cohesive ends" or "compatible sticky ends".

[0051] The double-stranded target DNA molecule includes the sequence of the nucleic acid molecule whose sequence needs to be determined. The nucleic acid molecule whose sequence needs to be determined can already be double-stranded DNA with 5' and 3' ends that are ligation-compatible with the 5' and 3' ends of the backbone to be used. Sometimes the nucleic acid needs to be made into double-stranded DNA, for example, in the case of cDNA or mRNA. The target DNA can already have suitable 5' and 3' ends. For example, various polymerases produce blunt-ended fragments. Such blunt-ended fragments are ligation-compatible with a backbone having blunt 5' and 3' ends. The target nucleic acid can also provide suitable 5' and 3' ends, for example, by digestion with an appropriate one or more restriction enzymes, or by adding deoxynucleotides with a terminal transferase. Suitable 5' and 3' ends can also be introduced by the insertion of restriction enzyme sites, recombinase recognition sites, and / or homology regions. For example, by ligating an adaptor containing the site to the target DNA, or by amplifying the target DNA with primers containing restriction enzyme sites, recombinase recognition sites, and / or homology regions.

[0052] Enzymes that generate ends compatible with the ends of the backbone for ligation but different from the nucleotides in the region immediately adjacent to the overhanging ends are available. In this embodiment, preferably, the recognition site of the enzyme is different from the restriction enzyme site of the first restriction enzyme. Thus, ligation of compatible ends does not generate a site that can be cleaved by the first restriction enzyme. If a restriction enzyme is used to provide a target nucleic acid with appropriate ends, preferably, the enzyme is one that generates blunt ends. In one embodiment, a target DNA molecule having 5' and 3' ends compatible with the 5' and 3' ends of the backbone to be used is provided by digestion with one or more restriction enzymes.

[0053] In one embodiment, ligating the ends of the target DNA to the ends of the backbone produces a target-backbone linker having a sequence that is not recognized / cut by a restriction enzyme that cleaves the (first) restriction enzyme site formed by self-ligation of the backbone.

[0054] In a preferred embodiment, the form that prevents self-ligation is the 5'-hydroxyl group of one DNA end and the 3'-hydroxyl group of the other DNA end, and the form that allows self-ligation is the 5'-phosphate group of one DNA end and the 3'-hydroxyl group of the other DNA end. Ligation requires the presence of a 5'-phosphate group. By appropriate phosphatase removal of the 5' end of the nucleic acid molecule, self-ligation and ligation to other similarly treated DNA molecules can be prevented. Even if the ends have ligatable-compatible ends, ligation is prevented.

[0055] In one embodiment, the backbone includes a recognition site for a nicking enzyme.

[0056] The target DNA molecule has 5' and 3' ends in a form that prevents self-ligation. Preferably, the target DNA is in a form that prevents ligation to other target DNA molecules. These two requirements can be met by providing ends in a dephosphorylated form or by adding nucleotides (3' overhang) to the 3' end of the target DNA molecule.

[0057] When the 5' and 3' ends of the target DNA are ligation-incompatible, self-ligation is inherently prevented. However, in these cases, preferably, ligation to other target molecules is also prevented. Thus, in such circumstances, preferably, the ends are also provided in a dephosphorylated form. Incompatible ends are, for example but not limited to, blunt ends and overhanging ends or overhanging ends where the overhanging nucleotides are incompatible overhanging ends.

[0058] Preventing self-ligation and / or preventing ligation to other target DNA molecules is not necessarily absolute. The process can / will occur to some extent. This is admissible in the methods of the present invention. Even if the ligation efficiency is low, good reads can be obtained.

[0059] The 5' and 3' ends of the backbone DNA can be compatible for ligation with each other. In such an embodiment, preferably, the 5' and 3' ends of the target DNA are also compatible for ligation with each other. Preferably, self-ligation of the ends of the backbone is not prevented. Preferably, the 5' end of the target DNA is dephosphorylated. Preferably, ligation is carried out in the presence of a restriction enzyme that recognizes and cuts the first restriction enzyme site.

[0060] In an embodiment of capturing double-stranded target DNA, the backbone is a double-stranded nucleic acid molecule. This backbone includes 5' and 3' ends that are compatible for ligation with the 5' and 3' ends of the target DNA. The 5' and 3' ends of the backbone can also be compatible for ligation with each other.

[0061] In an embodiment, the backbone includes one or more of the following parts:

[0062] - A first part encoding a first restriction enzyme site, preferably the 5' end of the first half of the first restriction enzyme site (e.g., see 1 in the schematic example below),

[0063] - One or more sites that allow nicking of the double-stranded backbone sequence (e.g., see 2 below),

[0064] - One or more type I or type II restriction enzyme sites (e.g., see 3 below),

[0065] - A secondary cloning site (e.g., see 4 below),

[0066] - A flexible DNA extension that enables efficient cyclization (bending) of the backbone molecule, 5) below

[0067] - A unique molecular barcode (identifier) sequence used to label each individual backbone molecule (e.g., see 6 below)

[0068] - Encoding another part, preferably the 3' end of the other half of the mentioned first restriction enzyme site.

[0069] - Phosphorylation at the 5' end of the backbone molecule and a hydroxyl group at the 3' end of the backbone molecule.

[0070] - A secondary barcode sequence that can be used to identify individual samples.

[0071] Schematic example of a double-stranded backbone sequence: (1)(2)(3)(4)(5)(6)(1)

[0073] 5’-GGGC..CCTCAGC..ATTTAAAT..GTCTTCGAGAAGAC..CATACTATCATG..(N)..GCCC-3’

[0074] 3’-CCCG..GGAGTCG..TAAATTTA..CAGAAGCTCTTCTG..GTATGATAGTAC..(N)..CGGG-5’

[0075] The dots represent 0 nucleotides; 1 nucleotide; 2 nucleotides or more nucleotides.

[0076] The sequences GGGC and GCCC represent half of a restriction enzyme site. The sequences form an SrfI site, but another restriction enzyme site will also work. In the case of SrfI (GCCC|GGGC), the advantage is that it is a blunt-end site. Another advantage is that it recognizes an 8-base-long site, while most commercially available alternatives recognize 6-base-long sites.

[0077] Preferably, the first restriction enzyme site does not occur at other positions in the backbone sequence.

[0078] Ligation of compatible ends can generate a restriction enzyme cleavage site. This occurs if the ends and the flanking sequences (if any) encode a restriction enzyme site when ligated to each other. As an example, the ends of a double-stranded DNA molecule with a single-stranded sequence of 5’-AATT… are ligation-compatible with a double-stranded DNA molecule with a single-stranded sequence of …TTAA-5’, where the dots indicate the double-stranded part and 5’ or 3’ indicates the free end of the respective molecule. Ligation of the two ends generates a molecule with the following double-stranded sequence:

[0079] …AATT…

[0080] …TTAA…

[0081] The overhangs are the same as those generated by the EcoRI restriction enzyme. Ligation generates a restriction enzyme cleavage site EcoRI only in some cases, i.e., when the bold nucleotides have the indicated bases:

[0082] …GAATTC…

[0083] …CTTAAG…

[0084] EcoRI cannot cleave when the bold nucleotides have different bases. For example, the following sequences are not cleaved by EcoRI:

[0085] …CAATTC…; or …AAATTC…; or …GAATTA…

[0086] …GTTAAG…; …TTTAAG…; …CTTAAT…

[0087] The sequence of the ends of the target DNA thus determines whether the ligation junction formed by ligation of compatible ends can be digested by the enzyme that cleaves the first restriction enzyme site.

[0088] In an embodiment, the insertion capture efficiency of the backbone can be optimized, where higher efficiency is reflected in the formation of products with more efficient cyclization and rolling circle amplification (RCA). The insertion capture efficiency of the backbone can be estimated by the number of multimers that can be formed.

[0089] In the method for sequencing target DNA described herein, preferably, the ligation of the target DNA to the backbone DNA does not generate a first restriction enzyme recognition site at the target / backbone DNA junction. In the present invention, preferably, the self-ligation of the backbone generates the first restriction enzyme site, and the ligation of the backbone to the target DNA does not generate said site. The preferred first restriction enzyme site is the enzyme that allows the most sequence variation in the ligation junction. Since the sequence of the backbone has a part, and preferably half, of the recognition sequence of the first restriction enzyme site, the variation comes from the sequence of the target ends. In the case where the first restriction enzyme site is an EcoRI site, the backbone sequence encoding the first restriction enzyme site has a 5'-end with the sequence 5'-AATTC. Depending on the bases of the nucleotides flanking the overhang at the target DNA, the junction with the target DNA can have one of four different sequences. Only when the target sequence has an end with the sequence 5'-AATTC.., the ligation junction with the backbone is cleavable by the EcoRI enzyme. Junctions with other sequences are not cleaved by EcoRI. By selecting enzymes that produce few or no overhangs and by selecting enzymes that require more specific bases at the recognition site, the variation in the junction is improved. The first restriction enzyme site preferably comprises 6 and more preferably 8 and preferably more bases. Therefore, the enzyme that cuts the first restriction enzyme site is preferably at least a 6-cutter, more preferably at least a 7-cutter, more preferably an 8-cutter. The number indicates the number of bases in the recognition site of the enzyme. For example, EcoRI is a 6-cutter; AluI that recognizes AGCT is a 4-cutter. There are also 5-cutters (e.g., AvaII), 7-cutters (e.g., BbvCI), 8-cutters (e.g., NotI), and even other restriction enzymes. Together with the preference for fewer or no overhangs, this ensures a high probability of sequence variation in the ligation junction, and this reduces the probability that the junction of the target sequence and the backbone sequence is the first restriction enzyme site. First restriction enzymes with more nucleotides in the recognition site are also preferred because such enzymes can allow larger target nucleic acid insertions. The method is suitable for a variety of target nucleic acid sources. The method of the present invention can be carried out with two or more backbones having different first restriction enzyme sites. Thus, more target molecules can be captured into DNA circles. In the case where the target DNA has two very close first restriction enzyme sites, the intermediate sequence can be effectively sequenced, for example, by capturing the intermediate sequence with a backbone having a different first restriction enzyme site. The mention of "first" in the context of the restriction enzyme cleavage site refers to the position (half of the site) of the site on the backbone. Restriction enzyme recognition sites at other positions in the backbone will be referred to as second, third restriction enzyme recognition sites, etc.

[0090] Preferred first restriction enzyme recognition sites are sites for restriction enzymes SrfI (GGGC|GCCC), PmeI (GTTT|AAAC), and SweI (ATTT|AAAT). A particularly preferred first restriction enzyme site is the site for restriction enzyme SrfI.

[0091] The 5' end of the backbone includes a portion at the very end of the first restriction enzyme recognition site. It may but need not contain additional nucleotides internally. The number of nucleotides at the end can be varied. The 5' end typically has between 2 and 15 nucleotides, preferably between 2 and 10 nucleotides, preferably between 2 and 8 nucleotides, more preferably 2, 3, 4, 5, 6, 7, or 8 nucleotides. In some embodiments, the 5' end is 3 or 4 nucleotides.

[0092] The 3' end of the backbone includes a portion at the very end of the first restriction enzyme recognition site. It may but need not contain additional nucleotides internally. The number of nucleotides at the end can be varied. The 3' end typically has between 2 and 15 nucleotides, preferably between 2 and 10 nucleotides, preferably between 2 and 8 nucleotides, more preferably 2, 3, 4, 5, 6, 7, or 8 nucleotides. In some embodiments, the 3' end is 3 or 4 nucleotides.

[0093] The 5' and 3' ends of the target DNA are preferably blunt ends. If self-ligation is not prevented, the 5' and 3' ends can also be sticky ends that can ligate together. The 5' and 3' ends of the target DNA are preferably provided in a dephosphorylated form to prevent self-ligation. The 5' and 3' ends of the target DNA can also be sticky ends that cannot ligate together, such as adenine overhangs added by terminal transferase.

[0094] Preferably, ligation is carried out in the presence of a ligase and a restriction enzyme (first restriction enzyme) that cuts the first restriction enzyme site. Ligation of the ends of the backbone to the ends of the target DNA produces a double-stranded DNA circle. In the methods of the present invention, self-ligation of the backbone is often not prevented. In the presence of a ligase, ligation of the two ends of the backbone to each other or to the ends of other backbones can impede capture of the target nucleic acid by the backbone. The presence of the first restriction enzyme can counteract ligation of the backbone ends. Since this ligation usually generates (regenerates) the first restriction enzyme site, the backbone is linearized and / or deconcatemerized. The ligation reaction is carried out under buffer conditions that support both efficient ligation and efficient cleavage by the first restriction enzyme. The methods of the present invention are particularly suitable for generating DNA circles having one backbone and one target nucleic acid.

[0095] In embodiments of the present invention, linear DNA, if present, is preferably removed prior to rolling circle amplification. Carrying out rolling circle amplification after removal of the linear DNA typically produces higher molecular weight concatemers of the backbone and the target DNA.

[0096] The method includes: subjecting DNA loops generated in a ligation reaction to rolling circle amplification (RCA). The rolling circle amplification generates an ordered array of copies of at least two of the DNA loops. The rolling circle amplification generates high molecular weight DNA molecules. The high molecular weight is suitable for sequencing, particularly suitable for long-read sequencing.

[0097] Rolling circle amplification has now been reviewed in the literature (Mohsen and Kool (2016) Acc Chem Res. Vol 49(11): pp2540-2550; published online Oct 24, 2016. doi: 10.1021 / acs.accounts.6b00417). The terms rolling circle amplification and rolling circle replication are sometimes used interchangeably in the art. In other cases, rolling circle amplification is used to refer to the replication of naturally occurring plasmids and viral genomes. These terms refer to similar basic principles, namely, the repeated copying of the same circular DNA to generate a longer nucleic acid molecule with an ordered array of backbone-target nucleic acid copies. The prior art for rolling circle amplification enables the generation of large arrays containing many copies of the DNA loops produced. The concatemer may have 2 or more copies, preferably 4 or more copies of the loop produced.

[0098] Rolling circle amplification is carried out by a polymerase and requires a common primer sequence to generate a promoter. A specific polymerase with high processive synthesis ability can be used to generate rather long concatemers. A polymerase with high processive synthesis ability is a polymerase that can polymerize one thousand or more nucleotides without dissociating from the DNA template. They can preferably polymerize two thousand, three thousand, four thousand or more nucleotides without dissociating from the DNA template. Polymerases with high processive synthesis ability are discussed in the literature (Kelman et al.; 1998: Structure Vol 6; pp 121-125). Rolling circle amplification can utilize a polymerase such as phi29 polymerase with high processive synthesis ability and strand displacement ability to generate concatemers of very high molecular weight. This polymerase can polymerize 10 kb or more. So a preferred polymerase with high processive synthesis ability is a polymerase that polymerizes 10 kb or more without dissociating from the DNA template (Blanco et al.; 1999. J. Biol. Chem. 264(15): 8935-40). In the presence of one or more suitable primers, polymerization can start at a nick in double-stranded DNA or the DNA can be denatured and annealed. Examples of suitable primers are hexamer random primers, one or more backbone-specific primers, one or more target nucleic acid-specific primers or combinations thereof. Random primers are generally preferred when the target nucleic acid sequence is unknown or when multiple target nucleic acid sequences are to be sequenced. One or more specific primers can be used for sequencing specific target nucleic acids with known underlying sequences. In one embodiment, one or more primers are specific to the backbone. Such primers can be used in different scenarios, for example, but not limited to, high-throughput systems with optimized backbones.

[0099] The advantage of having double-stranded circular DNA is that one of the strands can be used as a template for rolling circle amplification. For example, by using strand-specific primers to initiate the RCA reaction. Data analysis of Oxford nanopore sequencing results allows determination of the accuracy of base-calling and variant-calling for each strand separately. Specifically, we found that it is often difficult to distinguish between C and A bases because their raw current intensities are similar. However, the current signal from T is substantially different from all other bases and is easily correctly classified. For example, if A is expected to mutate on the positive strand, sequencing of the complementary strand will give a clearer result because A in the positive strand may be misjudged as G. Therefore, in this scenario, specific enrichment of the complementary strand would be advantageous. So in a preferred embodiment, the rolling circle initiation primer is a strand-selective primer.

[0100] Further optimization of obtaining strand-specific sequences may involve using (or additionally using) real-time selective sequencing methods, such as those described in the prior art (Loose et al. 2016. Nature methods. Real-time selective sequencing using nanopore technology).

[0101] The backbone is preferably 20 to 1000 nucleotides in length, preferably 20 to 800 nucleotides in length, preferably 50 to 800 nucleotides in length, more preferably 100 to 600 nucleotides in length, preferably 200 to 600 nucleotides in length. Depending on the application, the target nucleic acid is preferably 40 to 15000 nucleotides in length.

[0102] DNA that is free-circulating or bound to cell particles in a blood or other body fluid sample is typically less than 400 nucleotides. Target nucleic acid molecules of this length are particularly suitable for the methods of the present invention. Other samples with relatively small nucleic acid molecules are some types of forensic samples, fossil samples, nucleic acid samples isolated from environments that are inherently unfriendly to nucleic acid molecule integrity, such as fecal samples, surface water samples, and other microbiome-rich samples. For small target DNA (less than 100 nucleotides), it is preferred to use a larger backbone as described herein. The target nucleic acid can also be double-stranded circulating tumor DNA (ctDNA) or cell-free DNA (cfDNA) present in a liquid biopsy (including but not limited to blood, saliva, pleural fluid, or ascites). The target nucleic acid can also be double-stranded or single-stranded cDNA derived from messenger RNA, microRNA, CRISPR RNA, non-coding RNA, viral RNA, or other sources of RNA. The target nucleic acid can also be double-stranded DNA derived from genomic DNA, PCR products, plasmid DNA, viral DNA, or other sources of double-stranded DNA. The means and methods of the present invention are particularly suitable for capturing small DNA. Preferably 400 base pairs or less. The target DNA is captured in the backbone of the present invention. The captured DNA is also referred to as inserted DNA or target DNA. The target DNA is preferably 400 base pairs or less, more preferably 300 base pairs or less, more preferably 200 base pairs or less, more preferably 150 base pairs or less. The lower limit of the target DNA is preferably 20 base pairs, more preferably 30 base pairs, more preferably 40 base pairs, and more preferably 50 base pairs. Any lower limit can be combined with any upper limit.

[0103] Here, the size of the DNA fragment is given in nucleotides. This of course refers to the number in one strand. The size of double-stranded DNA can also be given in base pairs. So, a DNA of 400 nucleic acids is 400 base pairs long.

[0104] The generated concatemers can be sequenced using a variety of different methods. Among them, long-read sequencing methods are preferred. A person skilled in the art can use a variety of long-read sequencing methods. They share the characteristic that molecules greater than 200 nucleotides are generated in the sequencing reaction. Usually lengths greater than 500 nucleotides and even thousands of nucleotides. Two existing platforms for long-read, real-time, single-molecule sequencing are the Pacific Biosciences system (RSII and Sequel) and the Oxford Nanopore system (MK1 MinION, GridION, and PromethION). These allow read lengths in excess of 55 kb and even longer (Goodwin et al 2016; doi:10.1038 / nrg.2016.49). The long-read system is preferably a single-molecule real-time sequencing system. The single-molecule system does not rely on a clonal population of amplified DNA fragments to generate a detectable signal. These systems immobilize the sequence-determining protein at a specific location and allow the nucleic acid strand to advance through the protein. The existing Pacific Biosciences system uses polymerase, while the Oxford Nanopore system currently uses membrane channel proteins. In a preferred embodiment, the sequencing method is single-molecule real-time (SMRT) sequencing. The generated concatemers have an ordered array of at least one copy of the DNA loop, preferably at least 2, 3, 4, or preferably 5 of the DNA loops.

[0105] In some embodiments of the present invention, the backbone has an identifier. Such an identifier is also referred to as a barcode. The barcode or identifier is an extension of a nucleic acid whose sequence can differ between backbones. The barcode can be used to group the sequencing results of specific DNA loops. The barcode can identify the DNA loop. The barcode can be used to group the sequencing results of fragments of an ordered array of concatemers generated by RCA of the DNA loop. The method of using a backbone with a barcode typically has one or more backbone sets, where the backbones in the set have unique barcodes among other similar or identical backbones. Two or more backbone sets can be used, for example, to accommodate the different first restriction enzyme sites mentioned above, or to identify the sequencing results of different samples. The barcodes between sets can be the same because sequence differences in another part of the backbone identify the set. The backbone set can include more than one copy of the backbone containing the specific barcode of the backbone. The combination of the barcode and the specific overall target sequence can also unambiguously identify the nucleic acid as originating from a specific DNA loop, for example, when the target nucleic acid is complex and / or the number of identical barcodes in the backbone set is low. The sequencing results of the sequences of a set of DNA loops can be used to filter out errors such as amplification errors or polymerase errors. This is schematically illustrated in Figure 1 A and Figure 1 B. The backbone in the method as disclosed herein preferably includes at least two backbones with unique identifiers.

[0106] DNA loops are generated during the ligation step. Generally, the longer the molecule, the higher the efficiency of cyclization. Flexible molecules are more easily cyclized than rigid molecules. Small target nucleic acids (20 to 200 nucleotides) are captured by larger backbones, and the larger the backbone, the higher the capture efficiency. For small target nucleic acid backbones, it is preferably 200 nucleotides or more, preferably 300 nucleotides or more, more preferably 300 nucleotides or more, more preferably 400 nucleotides or more, preferably between 450 and 650 nucleotides. Smaller backbones generally allow more multimers per DNA loop. The average length of the target nucleic acid and the length of the backbone in the DNA loop are preferably 90 to 16000 nucleotides, preferably 200 to 12000 nucleotides, preferably 300 to 8000 nucleotides, preferably 400 to 4000 nucleotides, preferably 500 to 2000 nucleotides. The average length of the target nucleic acid plus the backbone nucleic acid is preferably about 1000 nucleotides.

[0107] The backbone DNA molecule preferably comprises the following sequences:

[0108] >BB1(199bp)

[0109] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0110] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0111] GTAGAGTACGACTGCGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC;

[0112] >BB2(259bp)

[0113] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0114] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0115] GTCACCGGGTCTTCGAGAAGACCTGTTTAGAGTACGACTGCAAATGGCTCTAGA

[0116] GGTACCCGTTACATAACTTACGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC;

[0117] >BB2_100 (341)

[0118] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0119] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0120] GTCACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATAT

[0121] ATATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTTAGAGTACGA

[0122] CTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGTGATCAGTGA

[0123] CTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC;或

[0124] >BBpX2(557bp)

[0125] GGGCATGCACAGATGTACACGAACGCCAGCAACGCGGCCTTTTTACGGTTCCTG

[0126] GCCTTTTGCTGGCCTTTTGCTCACATGTGAGGGCCTATTTCCCATGATTCCTTCAT

[0127] ATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTA

[0128] AACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGT

[0129] AGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTT

[0130] GAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGG

[0131] GTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGT

[0132] CCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTGTTTTAGAGCTA

[0133] GAAATAGCAAGTTAAAATAAGGCTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGC

[0134] AGACAAATGGCTCTAGAGGTACCCGTTACATAACTTATAGATGCTGCATGACATAGCCC。

[0135] >BB100_1(143bp)

[0136] GGGCATGCACAGATGTACACGATTCCCAACACACCGTGCGGGCCATCGACCTATG

[0137] CATACCGTACATATCATATATAAATCACATAATTTATTATACGTATGTCGCGCGGGTG

[0138] GCTGTGGGTAGATGCTGCATGACATAGCCC

[0139] >BB100_2(143bp)

[0140] GGGCATGCACAGATGTACACGCACTACATGCCAATGCCCAAGCAGTGCGCATATC

[0141] ACGTATCATATCTAATATATTATAATATTATGATAATGAGTATTTATTTAATTTGTTTG

[0142] TGTGAGGTAGATGCTGCATGACATAGCCC

[0143] >BB100_3(143bp)

[0144] GGGCATGCACAGATGTACACGCATTGGCCGTCTGTGCTGTCCATGGATCGTCTGA

[0145] TTGATATGATATCATATATTATAATTATACAGTAAGGTGATTGGGTATTGAGGGTTGT

[0146] GTGGTTGGTAGATGCTGCATGACATAGCCC

[0147] >BB100_4(145bp)

[0148] GGGCATGCACAGATGTACACGGTAGACATGCGAAGCGTGCGATGACAATCGATG

[0149] TGGACATCATGCATATATATGTTGTATAATTAAACAAATATGTGTAGTGTGTGAGGT

[0150] GGGTGTAGGAAGTAGATGCTGCATGACATAGCCC

[0151] >BB100_5(143 bp)

[0152] GGGCATGCACAGATGTACACGTTGTCATGGGAATTTGTGGTTATGAAATGAGTAT

[0153] GCGACGAATATGTATACATATATATTAAATTATAGAGTGATGTATGAGTTTGTGATG

[0154] TGTGGTGTATAGATGCTGCATGACATAGCCC

[0155] >BB200_1(243 bp)

[0156] GGGCATGCACAGATGTACACGGCGGCGCAAGATGATGTGCCGAACCTGACATGG

[0157] CATCGACTGGTATGGATCAATACTGATGCGATATCGATACCGGATAAATCATATATG

[0158] CATAATATCACATTATATTAATTATAATACATCGGCGTACATATACACGTACGCATCA

[0159] TTTCACTATCTATCGGTACTATACGTAGTGCCGGTCTGTTGGCCGGGCGACATAGA

[0160] TGCTGCATGACATAGCCC

[0161] >BB200_2(244 bp)

[0162] GGGCATGCACAGATGTACACGTGACGCAACGATGATGTTAGCTATTTGTTCAATG

[0163] ACAAATCTGGTATGATCAATACCGATGCGATATTGATATCTGATAACTCATATATGT

[0164] AGAATATCACATTATATTTATTATAATACATCGTCGAACATATACACAATGCATCTTA

[0165] TCTATACGTATCGGGATAGCGTTGGCATAGCACTGGATGGCATGACCCTCATTAGA

[0166] TGCTGCATGACATAGCCC

[0167] >BB200_3(244 bp)

[0168] GGGCATGCACAGATGTACACGAGACCGCAAGATGATGTTCATTCTTGAACATGAG

[0169] ATCGGATGGGTATGGATCAATACCGATGCGATATGATAACTGATAAATCATATATCT

[0170] ATAATATCACATTATATTAATTATAATACAGGATCGTTACATGCATACACAATGTATA

[0171] CTATACGTATTCGGTAGTTAGTGTACGGTCGGAATGGAGGTGGTGGCGGTGATAG

[0172] ATGCTGCATGACATAGCCC

[0173] >BB200_4(243 bp)

[0174] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGA

[0175] GATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGC

[0176] ATAATCTCACATTATATTTATTATAATAAATCATCGTAGATATACACAATGTGAATTG

[0177] TATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATG

[0178] CTGCATGACATAGCCC

[0179] >BB200_5(243 bp)

[0180] GGGCATGCACAGATGTACACGAATCCGTGAGATGACTATCTTATTTGTGACATTCA

[0181] TCGATCTGGATATGATCAATACCATGCGATATTGATTACTGATAAATCATATATGTAG

[0182] AATATCACATTATATTAATTATAATAAATCGTCGTACATATACATCCACAATTAGCTA

[0183] TGTATACTATCTATAGAGATGGTGCATCATCGTACTCCACCATTCCCACTAGATGCT

[0184] GCATGACATAGCCC

[0185] >BB300_1(348 bp)

[0186] GGGCATGCACAGATGTACACGCATAAGACCACAGGGTGCAAATCTGGATTGCGG

[0187] CATGGATGATTCATCATCGTGGCATATTCGCTATGGATATATCCATCATAATACATTG

[0188] ATACGTCATGCGTATAATCGCATTATATGTCGATATTGGTCATAGGGATACATCCGT

[0189] GTATACTATCGTATATGCGTGCAATGTAGCCATGTTAATCATGCTATAACCATAACA

[0190] TAAATATAATATATACAGATGGTGTATCTCTACTTATGTATGCTTGTATAGTAATGTC

[0191] GATACTGATGGGTCTCCGGCCCACTACACCACCTGGCCGCTCTAGATGCTGCATG

[0192] ACATAGCCC

[0193] >BB300_2(343 bp)

[0194] GGGCATGCACAGATGTACACGGGCAATCCGCCAGGGTTCAAATATGGATATGTGA

[0195] TGATCGATTCAACATGCACATATGCACGATATCATATATTACTCCAGATGTCATCAT

[0196] CGTCGTGCGTATATGAGATATGTATTTATGCATATAATCCACCATACATGGTAGCGA

[0197] TATTATAGTGCGATTATGTGTATATGACTATCATGGCTATTGTTAATATATAAATCATA

[0198] ACCATACCACTTCCACGCCTGGTATGGCGTATAGTATAGAGATATTGTGTGATGCC

[0199] CTATGTCGACCATGATGTGCCGTTGTACTGCCAATCCTAGATGCTGCATGACATAG

[0200] CCC

[0201] >BB300_3(344 bp)

[0202] GGGCATGCACAGATGTACACGTATCCATGCAGCTTATTGTAACTAGCGCATGCAC

[0203] GTGGTGATTCATCACATCTATATATACGATATGATATATTACACATATTTGCATAGTAT

[0204] CATCCGGTGTGATATCATCCGATATGCTCATACTTATTCATTGGTAGCATTGCATTG

[0205] ATGGATCAATAGTTATTATGACATCATGGCATGTACAATTATAAATAATACAACATA

[0206] CATAAATATACTATACACATCGTGTATGTGTTATACAGATCTGTGTGATGTATGATA

[0207] ATGTAATGGCGTCGAACACCACAAGGCAGTCCTATAATAGATGCTGCATGACATA

[0208] GCCC

[0209] >BB300_4(344 bp)

[0210] GGGCATGCACAGATGTACACGGTCCATTACAATCGAATCTATATCCCAATGTGTAT

[0211] CGATTATCACCACAATGACATAATACGATATCATATATTACTCCATATGCCTTACGTC

[0212] AGATCGTTATATGAGATATGTATTCATGCATATGATATCCCACAGTACACGTCGTCT

[0213] AATGCCATCATGAATGTATGACATATCTAGTCGATTATACATAATATAACATACCAAT

[0214] ATAACAATATCTATACACATTTGATGGCGTATAGTATAAAGATATTGTGGCAATGCC

[0215] CATACACCACTGACTGTCGCCGATCATTCCTACCACTAGATGCTGCATGACATAGC

[0216] CC

[0217] >BB300_5(344bp)

[0218] GGGCATGCACAGATGTACACGACCGACCGTGAAAGTGATTCAGAATGATGTGCA

[0219] TGAATGTTATCATGACATGATTTATGATGCACTGATATATGCATATTATAATATTGTA

[0220] CAATGTCGTATATACGACATATCTATACTATGAATTATGGCATCATGGACAATAGAT

[0221] GGTAAGGTATAGTACGATCTATATAGCATGTTGAAATGGGATATAAATTATCATAAA

[0222] CATACATACTTAACTAATATCAAGATGATATGTGTATGACATCAGAATGATAGTAGT

[0223] AATGAGTATTGTCAGATGTATGTACGAATATCACACGATTAGATGCTGCATGACAT

[0224] AGCCC

[0225] The backbone DNA molecule most preferably comprises the following sequence:

[0226] >BB200_4(243bp)

[0227] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGA

[0228] GATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGC

[0229] ATAATCTCACATTATATTTATTATAATAAATCATCGTAGATATACACAATGTGAATTG

[0230] TATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATG

[0231] CTGCATGACATAGCCC

[0232] The flexibility of the fixed-length backbone can be adjusted by customizing the sequence of the backbone. Different DNA molecules have different flexibilities according to the specific sequence of the molecule. Different sequences are provided by selecting different first restriction enzyme sites, different barcode sequences, and different sequences for other elements in the backbone. Preferably, the flexibility is adjusted by customizing the sequence of a dedicated portion of the backbone sequence. This dedicated portion is further referred to as a "linker". The linker preferably comprises 20 to 900 nucleotides, preferably 25 to 900 nucleotides, preferably 30 to 900 nucleotides, preferably 30 to 800 nucleotides, preferably 50 to 700 nucleotides, preferably 100 to 600 nucleotides, preferably 150 to 500 nucleotides. The linker can be a continuous sequence on the backbone or be divided into two, three, four or more continuous sequences. The linker is preferably a continuous sequence on the backbone or is divided into two, three, four continuous sequences, preferably one, two or three, preferably one or two, and more preferably one continuous sequence.

[0233] The free energy value (Breslauer et al. 1986) per base pair and the deviation of the twist angle (degrees) (Sarai et al. 1989) can be used to calculate the flexibility of any given DNA sequence. An example of such a calculation is as follows:

[0234] Flexibility calculation

[0235] The Python implementation of the TwistFlex algorithm (http: / / margalit.huji.ac.il / TwistFlex / ) (Menconi et al., 2015) can be used to calculate the DNA flexibility at the twist angles of the input sequence. The flexibility of each single dinucleotide is calculated based on the following angle table:

[0236]

[0237] Subsequently, the evolutionary algorithm for backbone optimization takes into account the average flexibility of the entire sequence when making selections. The average flexibility of a DNA sequence is calculated as the sum of all dinucleotide angles divided by the total number of dinucleotides. A suitable backbone has a flexibility score of 10 or greater, preferably 11 or greater, preferably 12 or greater, preferably 12.5 or greater (dinucleotide angles / dinucleotides in the backbone). A flexibility greater than 14 is generally not desired.

[0238] Entropy calculation for determining sequence complexity

[0239] The Shannon entropy of a string is defined as the minimum average number of bits per symbol required to encode that string. The formula for calculating Shannon entropy is:

[0240]

[0241] where p i is the probability that the character number i appears in the sequence.

[0242] This calculation can also be performed at http: / / www.shannonentropy.netmark.pl / .

[0243] The above formula is implemented using the following Python code:

[0244]

[0245]

[0246] A preferred backbone has a Shannon entropy value of 1.5 Sh or higher, preferably 1.5 or higher, preferably 2.5 or higher, more preferably 3.5 or higher.

[0247] Self-complementarity

[0248] The backbone core sequence preferably does not have 8 or more consecutive bases that are self-complementary in the same strand. Exceptions are the intentional insertion of one or more restriction enzyme sites or one or more other functional sequences. This sequence occasionally introduces self-complementary bases in the same strand. If possible, such bases greater than 8 should be avoided, but this is admissible in the functional backbone. However, when designing a new backbone, this sequence is preferably avoided if possible. The same is true for the kmers discussed hereinbelow.

[0249] Absence of repetitive motifs (kmers)

[0250] The backbone core sequence preferably does not have a motif of 6 bases that are repeated more than twice in the sequence.

[0251] The flexibility of the backbone and the Shannon index of the backbone can be adjusted by adding linkers to the backbone. The effect of the specific sequence of the linker on the flexibility and complexity score of the backbone can be easily calculated.

[0252] The linker preferably has one or more of the following characteristics: (i) The overall complexity of the linker sequence is preferably high. The above Shannon entropy formula is a method for determining the value of the complexity of a given sequence; (ii) Copies of DNA motifs longer than 5 bases preferably occur no more than twice, preferably no more than once, in the linker sequence. In a preferred embodiment, the linker does not include copies of DNA motifs longer than 5 bases (i.e., the linker sequence does not contain repetitive motifs where the motif is more than 6 consecutive bases); (iii) The linker preferably does not include more than two, preferably not more than one, preferably not more than 6 self-complementary sequences of nucleotides separated by less than 10 nucleotides (inverted repeats). The criteria mentioned help to avoid, generally speaking, the appearance of complex secondary structures in the single-stranded version of the linker. The possibility and strength of secondary structures can also be calculated by other means.

[0253] The backbone preferably includes a GC content of 30% - 60%; preferably 40% - 60%; preferably 40% - 50%; preferably 45% - 55%.

[0254] The backbone preferably has one or more of the following characteristics. The first restriction enzyme site is preferably a restriction enzyme that produces blunt ends. It has been observed that this improves the capture of the target nucleic acid. The backbone preferably includes a recognition site for a DNA nickase that produces a primer site for rolling circle amplification. Additional restriction enzyme recognition sites can be used for the sequential ligation of multiple short DNA molecules into a circular DNA. The backbone preferably includes a molecular identifier that differentiates the originally captured nucleic acid from its subsequent sequencing read length.

[0255] The method of the present invention can be used to sequentially capture two or more target nucleic acids on each backbone. Sometimes, two target nucleic acid molecules can be captured simultaneously in a single capture step. The probability of this occurring is deliberately low because measures are taken to prevent self-ligation. The sequential capture of two or more target nucleic acids can be a desirable feature. Additional restriction enzyme sites can be incorporated into the backbone. Once the first target nucleic acid has been captured, the method can be repeated by adding a restriction enzyme that cuts the additional restriction enzyme site. The DNA circle is cut and linearized by a second restriction enzyme and prepared for ligation to the target nucleic acid. If the second restriction enzyme produces the same type of ends as the first restriction enzyme (e.g., blunt ends), the reaction can continue to capture target nucleic acids that were not captured in the first iteration of the method. Optionally, the DNA circle can be purified (e.g., by removing linear DNA), and a new target nucleic acid can be added whose ends are ligation-compatible with the ends of the backbone produced by the second restriction enzyme. By adding further restriction enzyme sites to the backbone, this step can of course be repeated for sequentially capturing third, fourth, etc. target nucleic acids. When more than one target nucleic acid is captured, preferably, the second and first restriction enzyme sites are sites of enzymes that are infrequently cut. Such enzymes preferably have 8 or more cutters. Such enzymes preferably produce blunt ends of the enzyme. This sequential capture allows for the simultaneous sequencing of more than one nucleic acid. Different target nucleic acids can be identified based on their position in the backbone, i.e., based on the flanking backbone sequences into which they are inserted.

[0256] In addition, during data analysis, the backbone serves as a control sequence. Based on the read length of the backbone sequence, the error rate of each sequencing read can be inferred, enabling an accurate estimate of the likelihood of genetic variation within the captured nucleic acid sequence.

[0257] By-products can be produced in the method of the present invention. The number of single backbones and single target DNAs containing DNA circles is affected by, for example, the molar ratio of backbone / sample, which should promote the formation and insertion of molecules with backbones rather than unwanted by-products ( Figure 2 ), such as: (i) linear DNA formed by random concatenation of the backbone and sample DNA, (ii) circular DNA containing only the backbone or only the sample DNA, (iii) circular DNA containing an excess of backbone or sample DNA.

[0258] In an embodiment of the present invention, preferably, the molar ratio of backbone molecules to target nucleic acid molecules ranges from 1:10 to 10:1. Preferably, the ratio range of 1:5 to 5:1 is maintained, preferably, the ratio range of 1:2 to 2:1 is maintained. Preferably, the average ratio is 1:1.

[0259] The method described herein, including rolling circle amplification, is preferably carried out without using a switching container. The resulting concatemers can be sequenced in the same container or different containers.

[0260] The method of the present invention preferably produces concatemers that are as long (>10 Kb) as the linear dsDNA formed by multiple units composed of target nucleic acid-backbone copies. The concatemerization / polymerization of such units facilitates the detection of true genetic variations by distinguishing them from sequencing errors. In fact, in the case of rare genetic variations that occur at a frequency of less than 1% within a DNA molecule library, direct sequencing, such as short-read sequencing, may no longer be applicable due to the sequencing error rate being higher than the mutation frequency. Using the method described herein, the same rare sequence (genetic variation) is represented multiple times in the long concatemers, which provides a high confidence in the presence of the mutation, even if the mutation frequency is low in the original library of nucleic acid molecules.

[0261] The backbone includes a 3’ sequence encoding a part of the recognition site of a first restriction enzyme (restriction enzyme site) and a 5’ sequence encoding another part of the cleavage site of the first restriction enzyme. The backbone may contain further elements, such as one or more of the following: (i) one or more sites allowing nicking of the double-stranded backbone sequence; (ii) one or more type I or type II restriction enzyme cleavage sites; (iii) secondary cloning sites; (iv) flexible DNA extensions (linkers) that enable efficient cyclization (bending) of the backbone molecule; and (v) unique molecular barcode sequences used to label each individual backbone molecule.

[0262] The backbone preferably has 5’-phosphorylation at both ends of the backbone molecule.

[0263] The present invention also provides a collection of linear DNA molecules (backbones) with lengths between 20 and 1000 nucleotides, the linear DNA molecules including a 5’ end and a 3’ end, the 5’ end including a part of the recognition site of a first restriction enzyme at the outermost end, the 3’ end including another part of the recognition site of the first restriction enzyme at the outermost end, and the 5’ end and the 3’ end being ligatable to each other and forming a restriction enzyme (first restriction enzyme) recognition site when self-ligated, and wherein each strand of the backbone includes:

[0264] A linker;

[0265] An identifier sequence (barcode) that is different from the sequences of the identifiers of other backbones in the collection; and

[0266] Optionally, a restriction enzyme cleavage site for a nicking enzyme.

[0267] The backbone is the preferred backbone in the methods described herein. The first part and the second part of the first restriction enzyme cleavage site together form the complete recognition site of the first restriction enzyme cleavage site and are located on the molecule at positions where the two parts can form an operable connection for the first restriction enzyme cleavage site. Operable connection in the context refers to the effectiveness of cleavage by the first restriction enzyme. The backbone preferably further includes a second restriction enzyme cleavage site that is a type I or type II restriction enzyme site. The backbone preferably further includes a restriction enzyme site for a type II restriction enzyme that can generate non-palindromic overhangs (Golden-Gate cloning site). The linker is preferably the linker described above. The backbone preferably includes a nucleic acid molecule (captured nucleic acid molecule) in the first restriction enzyme cleavage site. The backbone preferably includes a library of captured nucleic acid molecules.

[0268] There is further provided a kit comprising a backbone as described herein. The kit preferably includes a collection of backbone molecules as described herein. The kit preferably further includes a polymerase with high processive synthesis ability and optionally one or more polymerization primers. The kit preferably further includes: a ligase and the first restriction enzyme; and / or the target site-specific recombinase. The kit preferably further includes a DNA exonuclease. The latter enzyme is suitable for removing linear DNA before generating the concatemer of the DNA circles.

[0269] In one aspect, the present invention provides a method for determining the sequence of a collection of nucleic acid molecules, the method comprising:

[0270] - providing a double-stranded target DNA molecule having a 5'-end and a 3'-end, the double-stranded target DNA molecule having protruding adenine residues at the 3'-ends of both strands of the DNA molecule;

[0271] - providing a collection of double-stranded backbone DNA molecules, the double-stranded backbone DNA molecules including 5'-ends and 3'-ends that are compatible for ligation with the 5'-end and 3'-end of the target DNA.

[0272] The method further comprises:

[0273] - ligating the target DNA to the backbone in the presence of a ligase, thereby generating a DNA circle comprising the backbone and the target DNA molecule;

[0274] - optionally removing linear DNA;

[0275] - generating a concatemer comprising an ordered array of at least two copies of the DNA circle by rolling circle amplification; and

[0276] - sequencing the concatemer.

[0277] The ends compatible with overhanging 3'-adenine ligation are ends having a 5'-overhanging thymidine base or an analogue thereof. This method differs from the method described above in that, since all ends have an a'-overhanging adenine base, intermolecular or intramolecular ligation of the targets is inherently inhibited. Thus, self-ligation of the target nucleic acid or ligation of one end to another target nucleic acid molecule is inherently impossible. The overhanging bases are nucleotides at the ends of the nucleic acid molecule and they do not base pair with bases on the opposite strand. There are no opposite bases for the overhanging bases. Such protrusions are also referred to as sticky ends or cohesive ends. The same is true for the backbone. They inherently prevent self-ligation. In this embodiment, the backbone does not have to have a portion of the first restriction enzyme site at the ends. The ends thus cannot be ligated to produce the first restriction enzyme site. Thus, the ligation does not have to be carried out in the presence of the first restriction enzyme. The remaining steps and definitions can be the same as those described elsewhere herein.

[0278] There is further provided a method for determining the sequences of a set of nucleic acid molecules, comprising:

[0279] providing a double-stranded target DNA molecule having recombinase recognition sites at the 5'-end and the 3'-end, the recombinase recognition sites being specific for a target-site specific recombinase;

[0280] providing a backbone comprising the recognition sites separated by DNA including a linker;

[0281] incubating the target DNA molecule with the backbone in the presence of the target-site specific recombinase, the target-site specific recombinase preferably being Cre recombinase, FLP recombinase or bacteriophage λ (lambda) integrase, thereby producing a DNA loop comprising the backbone and the target DNA molecule;

[0282] optionally removing linear DNA; and

[0283] generating a concatemer comprising an ordered array of copies of at least two of the DNA loops by rolling circle amplification; and

[0284] sequencing the concatemer.

[0285] In a preferred embodiment, the backbone is a loop comprising two recombinase recognition sites, which are separated on one side by DNA comprising a linker and on the other side by DNA encoding a further restriction enzyme recognition site, and wherein the further restriction enzyme cleavage site is the only recognition site for the restriction enzyme in the backbone. In this embodiment, the method preferably further comprises: digesting the DNA after recombination with the restriction enzyme and then removing the linear DNA before generating the concatemer. The further restriction enzyme cleavage site is preferably 6 or more cutters, preferably 7 or more cutters, preferably 8 cutters. The ends generated by enzymatic cleavage need not be blunt ends. In a preferred embodiment, the further restriction enzyme is not a blunt cutter.

[0286] Target site-specific recombinase

[0287] The target site-specific recombinase is a gene recombinase. Target site-specific DNA recombinases are widely used in multicellular organisms to manipulate genomic architecture and control gene expression. These enzymes, which are derived from bacteria and fungi, catalyze directionally sensitive DNA exchange reactions between short (30-40 nucleotide) target sequences specific to each recombinase. These reactions can perform four basic functional modules, excision / insertion, inversion, translocation, and cassette exchange. Non-limiting examples of recombinases are Cre recombinase, Hin recombinase, Tre recombinase, and FLP recombinase. Cre recombinase was the first widely used recombinase. It is a tyrosine recombinase derived from bacteriophage P1. The enzyme uses a topoisomerase I-like mechanism to perform site-specific recombination events. The enzyme (38 kDa) is a member of the integrase family of site-specific recombinases and is known to catalyze site-specific recombination events between two DNA recognition sites (LoxP sites). The 34 base pair (bp) LoxP recognition site consists of two 13 bp palindromic sequences flanking an 8 bp spacer region. The products of Cre-mediated recombination at the LoxP site depend on the location and relative orientation of the LoxP sites. Two separate DNA species each containing a LoxP site can be fused by Cre-mediated recombination. The DNA sequence found between two LoxP sites is referred to as "floxed".

[0288] Red / ET recombinase

[0289] Recombineering utilizes phage-derived protein pairs, either RecE / RecT from the Rac phage or Redα / Redβ from the λ phage, to assist in the cloning or subcloning of DNA fragments into vectors without the need for restriction enzyme sites or ligases. RecE / RecT, Redα / Redβ, and other similar protein pairs are further referred to as RecE / RecT protein pairs in this article. The limitation of the original homologous recombination technique was due to the fact that the bacterial RecBCD nuclease degrades linear DNA and the initial events had to be studied in RecBCD-deficient strains (7). This problem was solved because it was found that Redγ could assist Redα and Redβ, thereby inhibiting RecBCD nuclease activity and enabling this technique to be applied to Escherichia coli (E. coli) and other commonly used strains. In addition, the recombination efficiency was increased by 10 - 100 times. The combination of these three enzymes (α, β, and γ, or E, T, and γ) in a vector was named Red / ET recombination, and the basic principle of this method is that for recombination to occur in a linear fragment, at the two ends of a double-strand break (DSB), and in another linear or circular plasmid, two homologous regions greater than 15 bp, preferably greater than 20 bp, preferably greater than 30 bp, and preferably greater than 42 bp are required. Directed insertion can be achieved by flanking the target DNA and the insertion site with two different homologous regions. DSBs are essential, enabling RecE or Redα to bind and degrade one strand of the DNA (5' to 3') and simultaneously load RecT or Redβ onto the exposed single strand. The single-stranded DNA loaded with RecT or Redβ recombinase finds a perfectly matching sequence and joins the two sequences by strand invasion or annealing.

[0290] The insertion of homologous regions (HRs) is usually achieved by including them in the oligonucleotides used for the amplification of the product, which serves as the linear substrate for the recombination event. If the process requires longer DNA fragments, then the HRs can be inserted by traditional restriction / ligation techniques using plasmids or adaptors.

[0291] Restriction enzyme recognition sites

[0292] Restriction enzyme recognition sites are usually also simply referred to as restriction enzyme sites, restriction cleavage sites, or restriction recognition sites. They are positions on a DNA molecule containing a specific sequence of nucleotides that can be recognized by a restriction enzyme. These are usually palindromic sequences. A specific restriction enzyme can cut the sequence between two nucleotides within or near its recognition site. The enzyme usually cuts both strands of the DNA molecule, usually followed by the separation of the ends. So-called nicking enzymes can also recognize restriction cleavage sites but only cut one of the two strands. The resulting DNA molecule remains bound, but one of the two strands has a nick.

[0293] Restriction enzyme type

[0294] Naturally occurring restriction endonucleases (restriction enzymes) can be classified into four groups (Type I, II, III, and IV) based on their composition and cofactor requirements, the nature of their target sequences, and the position of their DNA cleavage sites relative to the target sequences. However, DNA sequence analysis of restriction enzymes has shown great variability, indicating more than four types. Enzymes of all types recognize specific short DNA sequences and effect endonucleolytic cleavage of DNA to produce specific fragments with 5'-phosphate termini.

[0295] Type I enzymes (EC 3.1.21.3) cleave at sites remote from the recognition site and require ATP and S-adenosyl-L-methionine to function. Their versatility lies in having both restriction enzyme and methylase (EC 2.1.1.72) activities.

[0296] Type II enzymes (EC 3.1.21.4) cleave within the recognition site or at a short specific distance from the recognition site. Most Type II enzymes require magnesium ions. They generally have a single function (restriction).

[0297] DNA phosphorylation

[0298] Single-stranded or double-stranded DNA with 5'-hydroxyl termini must have a 5'-phosphate group for efficient ligation. 5'-ends without such a phosphate group can be phosphorylated prior to ligation. Some polynucleotide kinases, including T4 PNK (NEB#M0201) and T4 PNK (negative 3'-phosphatase) (NEB#M0236), can be used to transfer the γ-phosphate of ATP to the 5'-end of DNA.

[0299] DNA dephosphorylation

[0300] Digested DNA usually has the 5'-phosphate group required for ligation. To prevent self-ligation, the 5'-phosphate can be removed prior to ligation. Dephosphorylation of the 5'-end inhibits self-ligation, allowing the technician to manipulate the DNA as needed prior to religation. Dephosphorylation can be accomplished using any of several phosphatases, including the Quick Dephosphorylation Kit (NEB#M0508), shrimp alkaline phosphatase (rSAP) (NEB#M0371), calf intestinal alkaline phosphatase (CIP) (NEB#M0290), and Antarctic phosphatase (NEB#M0289).

[0301] DNA ligation

[0302] DNA ligation is a central step in many modern molecular biology workflows. DNA ligases catalyze the formation of phosphodiester bonds between the 3’ hydroxyl and 5’ phosphate of adjacent DNA residues. In the laboratory, this reaction is used to join dsDNA fragments with blunt or sticky ends to form recombinant DNA plasmids and to add barcoded adapters to fragmented DNA in next-generation sequencing and many other applications. DNA ligase from bacteriophage T4 is the most commonly used ligase. It can ligate the sticky or cohesive ends of DNA, oligonucleotides, as well as RNA and RNA-DNA hybrids. It can also efficiently ligate blunt-ended DNA. Using Circ Ligase TM II ssDNA Ligase* (Epicentre) can efficiently ligate single-stranded DNA. This is a heat-resistant enzyme that can catalyze the intramolecular ligation (i.e., circularization) of ssDNA templates containing 5’-phosphate and 3’-hydroxyl. Circ Ligase TM II ssDNA Ligase* joins the ends of ssDNA in the absence of complementary sequences. Therefore, this enzyme is very useful in synthesizing circular ssDNA molecules from linear ssDNA. Circular ssDNA molecules can be used as substrates for rolling circle replication or circular transcription.

[0303] For clarity and conciseness of description, when referring to a step being performed on or with one or more substrates and which step is catalyzed by one or more enzymes, the step is carried out by contacting the substrate with the enzyme. This is typically done by adding the enzyme to the substrate in an appropriate buffer.

[0304] For clarity and conciseness of description, features are described herein as part of the same or separate embodiments. However, it is understood that the scope of the present invention may include embodiments having all or part of the combinations of the described features. BRIEF DESCRIPTION OF THE DRAWINGS

[0305] Figure 1 A) Schematic diagram of a method for capturing small nucleic acid molecules and generating concatemers by using a backbone and rolling circle amplification. B) Schematic diagrams of sequencing reactions using short read lengths and without a backbone and long read sequencing using a backbone.

[0306] Figure 2 、 Figure 1 Examples of possible linear and circular by-products of the cyclization reaction indicated in. The large circle shadow in the upper left corner of the circular by-product diagram represents the backbone sequence. The other shadows are target sequences.

[0307] Figure 3 、Schematic example of a double-stranded backbone sequence.

[0308] (1) Indicates the 5'-end and 3'-end sequences that together encode a restriction enzyme recognition site.

[0309] (2) Is a restriction enzyme cleavage site for the nicking enzyme BbvCI. Any other nicking site can also be used. However, the advantage of using BbvCI is that there are two forms of this enzyme available commercially, one that cleaves DNA on the positive strand and the other on the negative strand. The nicked DNA is an effective entry site for the rolling circle amplification (RCA) reaction. Depending on the situation, we may want to use the nicked DNA instead of a DNA primer to initiate the polymerization reaction.

[0310] (3) Is an additional blunt-end restriction enzyme cleavage site, which is the recognition site for SweI in the example case. The second blunt-end restriction enzyme cleavage site allows the capture of a second DNA fragment in a further cyclization reaction.

[0311] (4) Is a cloning site, which is a double-flip BbsI site in this example and can be used for the simple extension of the backbone via Golden-Gate or other types of cloning.

[0312] (5) Represents a flexible DNA extension (linker). Its length is variable and aids in efficient cyclization.

[0313] (6) The capital N indicates an extension of nucleic acid encoding a unique identifier. It is a barcode-like sequence. It can encode one or more (random) barcodes of any suitable size.

[0314] Element (1) is located at the outermost end. Elements 2 - 6 can have any order and may or may not be present depending on the situation.

[0315] Figure 4 We have developed a method for detecting gene fusions based on targeted cDNA synthesis, single-stranded DNA cyclization (ssDNA), and targeted rolling circle amplification. The cDNA generated by the reverse transcriptase step is shaded in gray in the left figure. The bottom cDNA is the fusion gene DNA, and there are two shadings indicating the part from one gene and the part from the other gene. Clearly, the RCA assay produces multimers of the fusion. Of course, this method can also be used to determine the sequence of one or more cDNAs with non-gene fusion results.

[0316] Figure 5 Schematic diagram of the MIP probe and method.

[0317] Figure 6 Plot the DNA content and predicted values for each band.

[0318] Figure 7 Determine the cyclization efficiency: A) Comparison of the insert before and after the reaction. B) Comparison of the cyclized product and the unreacted product.

[0319] Figure 8 Results of proof-of-concept experiments. (A) Gel image indicating RCA products for subsequent sequencing on nanopore MinION. (B) Nanopore read length distribution generated by MinION R9.4 runs using the samples indicated in (A) as input. (C) Pattern fraction distribution of 2083 read lengths greater than 10 kb. (D) Schematic of nanopore sequence reads, replacing inserts (green) and backbone (red), insert sequences aligned, and consensus sequence generated from aligned inserts. (E) Accuracy of consensus sequences.

[0320] Figure 9 Circularization of backbone 2 (BB2) and backbone 3 (BB3) with insert 17.2 in a 3:1 ratio. A) Comparison of BB2 and BB3. Red asterisks: correct circularization products. Multiple bands in condition 3: linear products vs. circular products, ligation of multiple backbones. Additional bands in condition 4: circularized backbones. B) Successful circularization using backbone 2 (BB2). Yellow asterisks: correct circularization products. In this gel image, the entire reaction was loaded into each lane. Lanes 1 and 2 represent circularization of BB2 with insert 17.2 in a 1:1 ratio before and after PlasmidSafe digestion. Lanes 3 and 4 represent the same circularization reaction using BB2 with insert 17.2 in a 3:1 ratio.

[0321] Figure 10 Backbone circularization efficiency at different backbone-to-insert ratios. Ligation products at different backbone (BB3)-to-insert (17.2) ratios were qualitatively and quantitatively assayed. (A) Agar gel showing circularization inputs and reaction products. PlasmidSafe digestion was used to remove linear products after the circularization reaction. Red asterisks: correct circularization products. Yellow asterisks: remnants of insert inputs. (B) Quantification of circularization efficiency. Quantification of circularization efficiency is defined as P / I * 100, where P is the amount (in moles) of correct backbone-insert products (red asterisks), and I is the amount (in moles) of input inserts. Band intensities and surface areas were measured using software ImageJ ( https: / / en.wikipedia.org / wiki / ImageJ ). Data were normalized using GeneRuler 50 bp DNA Ladder as a reference. See also Section 10 of Materials and Methods.

[0322] Figure 11 Circularization efficiency of BB2_100 (orange bars) with and without SrfI and HMGB1 added. Ligation was performed using the backbone-to-insert ratio. Blue bars represent control experiments of BB2 and BB3 ligated to the same insert without SrfI and HMGB1 added. Quantified circularization efficiency as described above ( Figure 10 legend).

[0323] Figure 12 Visual display of the reaction product of cyclization using the backbone BB2_100 and the inserted 17.2. Red asterisk: correct product. Orange box: position where the inserted residue is expected after cyclization. The complete disappearance of the inserted band after ligation indicates that the insertion is completely ligated. Cyclized 1: before Plasmid Safe DNase treatment. Cyclized 2: after Plasmid Safe DNase treatment.

[0324] Figure 13 Effect of adding the restriction enzyme SrfI in the cyclization reaction. The cyclization reaction was carried out using BB3 and the inserted 17.2. The reaction was carried out in the presence and absence of SrfI and Plasmid Safe DNase.

[0325] Figure 14 Barcode strategies useful for the described technology. (A) Marking individual DNA molecules with unique molecular identifiers to improve mutation discovery. (B) Marking individual samples with sample-specific barcodes for pooling in sequencing experiments.

[0326] Figure 15 RCA products using various DNA templates. RCA was carried out using circular DNA templates from various sources. (A) Cell-free DNA cyclized using the backbone BB2; (B) plasmid pX_Zeo; (C) ss-cDNA self-cyclized using Circ Ligase II (Epicentre #CL9021K); (D) PCR product 17.2 cloned into plasmid pJET. As a reference, a long-range 1Kb sequence ladder was used. The higher bands of the sequence ladder are 10Kb long and the RCA products are estimated to be between 20kb and 100kb long.

[0327] Figure 16 Number of read lengths containing 17.1 and 17.2. The ratio between reads containing 17.1 and reads containing 17.2 is 1:14, indicating significant enrichment of the target region due to site-directed RCA.

[0328] Figure 17 One Review of the reaction products of the steps in the pot reaction design. The insert (17.2) and the backbone (BB2_100) were cyclized to produce the product indicated in (1). The linear DNA product was digested using Plasmid Safe DNase as indicated in (2). Based on (2) as the input, the RCA reaction product was formed as indicated in (3).

[0329] Figure 18, Review of common sequence calling methods for short-read sequencing (left figure) and long-read sequencing (right figure).

[0330] Figure 19 , (A) Example of mapping insertion-derived Cyclomics sequencing reads with TP53 mutations (red box). (B) Graph showing the fraction of insertions supporting non-reference alleles in 558 read lengths with >4 insertions. Four reads showed a high fraction of non-reference alleles, and these contained insertions with the desired chr17:7578265, A->T mutation.

[0331] Figure 20 , Capture of DNA with a target-site specific recombinase. Recognition sites for the target-site specific recombinase are indicated by the letters A and B. The target DNA is indicated by the term "insert". Sites A and B can be introduced in various ways, such as by ligating an adaptor with the sites to the insert DNA or by amplifying the insert with primers that include sequences encoding said sites A and B. The backbone is indicated by the term "backbone". The backbone in the figure is a circular molecule including the DNA between two sites A and B. This intermediate DNA includes restriction enzyme cleavage sites unique to the entire backbone. Arrows indicate that the insert and the backbone are first recombined by adding the recombinase and then the restriction enzyme is added. The restriction enzyme will only cut the unreacted backbone and the backbone in which the linker has been replaced by the insert. The linearized DNA can be removed by adding a suitable exonuclease.

[0332] Figure 21 , Comparison of the efficiency between ligation reaction products and different backbone designs. Left: Ligation of different backbones with 250bp PCR amplicons. Right: Remaining circular products after digestion of linear DNA with Plasmid Safe deoxyribonuclease. Ligation to all members of the BB200 series showed high cyclization efficiency, as indicated by the formation of circular products including the PCR product and the backbone.

[0333] Figure 22 , Comparison of the ligation efficiency of backbones from the BB200 series. Left: Gel showing ligation products of 3 backbones. BB200_4 showed a brighter band indicating more products formed during the reaction. Right: Measurement of the brightness of the bands. From top to bottom: BB200_2, BB200_4 and BB200_5.

[0334] Figure 23 , Difference in the number of sequencing reads from RCA products formed by ligation of backbones from the BB200 series. Percentage of reads from different backbones in two independent experiments (red and blue). The backbones were initially mixed in a 1:1:1 ratio. BB200_4 with a higher number of reads compared to Figure 2Consistent with the higher ligation efficiency shown.

[0335] Figure 24 Inference of bases at specific positions of the gene TP53 (GRCh37 17:7577518). The Y-axis represents the distance (average match score) between the patterned nanopore signal corresponding to the reference sequence and the signal from the experimental sequence. The greater the distance, the more difficult it is to infer the correct base. The X-axis shows the number of insertions present in the read fragment. The inferred bases are indicated in different colors. The signal from the forward strand is not as clear as the signal measured on the reverse strand. This makes it difficult to distinguish the correct base (A, blue) from other possible bases even when the calculated distance is low.

[0336] Figure 25 Agar gel (S) depicting the product of the cyclization reaction between the backbone and the insert. Negative controls are indicated as C-. The band corresponding to the circular BB-I product is separated from the gel.

[0337] Figure 26 Agar gel showing an example product after rolling circle amplification.

[0338] Figure 27 BB200_4 (243 bp, indicated as BB in the figure) and S1_WT (158 bp, indicated as I in the figure) are circularized and amplified by RCA. When digesting the concatemer made by BB-I, we expect a band of approximately 400 bp, however if the concatemer consists only of BB, a band of approximately 250 bp should be produced. The concatemer formed only by I will not be digested, making the RCA band visible.

[0339] Examples

[0340] Example 1

[0341] Materials and methods

[0342] The method described herein involves sequencing nucleic acid sequences in a nucleic acid molecule (DNA) library to allow for the detection of gene mutations. Such nucleic acid molecules include, for example, ctDNA (circulating tumor DNA), cfDNA (cell-free DNA), genomic DNA, RNA, polymerase chain reaction (PCR) products, or other products. In this process, the "product" obtained is a long linear dsDNA (>10Kb), formed by multiple units consisting of nucleic acid-backbone copies. Such concatenation / polymerization of units is necessary to distinguish the detection of true gene mutations from sequencing errors. In fact, in the case of rare gene mutations occurring at a frequency of less than 1% within the DNA molecule library, direct sequencing, e.g., short-read sequencing, is no longer applicable due to the sequencing error rate being higher than the mutation frequency. Using the above method, the same rare sequence (gene mutation) is represented multiple times in the long concatemer, which provides a very high confidence in the presence of the mutation, even when the mutation frequency in the original nucleic acid molecule library is low.

[0343] The design of the backbone for capturing nucleic acids brings many advantages for efficient and specific nucleic acid acquisition and is crucial for the computational analysis of sequencing data.

[0344] Preferred backbone molecules are characterized by:

[0345] 1) Blunt-end restriction enzyme cleavage sites encoded at the very ends of the backbone to improve the ligation efficiency of short DNA.

[0346] 2) Recognition sites for DNA nicking enzymes, which are used to generate "single-stranded templates" for rolling circle amplification.

[0347] 3) Restriction enzyme target sites, which can be used to ligate multiple short DNA molecules continuously into a circular DNA.

[0348] 4) Molecular identifiers that enable the distinction between the originally captured nucleic acid and its subsequent sequencing reads.

[0349] In addition, during data analysis, the backbone serves as a control sequence. Based on the backbone sequence reads, the error rate of each sequencing read can be inferred, enabling an accurate estimation of the likelihood of gene mutations within the captured nucleic acid sequence.

[0350] 2. Materials and Methods for Cyclomics Technology

[0351] 2.1 Backbone Design

[0352] We have gone through an iterative method of replacing the backbone design, followed by experimental testing to find the best backbone that can achieve the most efficient cyclization (i.e., capture of double-stranded nucleic acid molecules). The basic design of our backbone can include one or more of the following parts:

[0353] 1) 3' sequence encoding a half restriction enzyme cleavage site.

[0354] 2) One or more sites allowing nicking of the double-stranded backbone sequence

[0355] 3) One or more type I or type II restriction enzyme cleavage sites

[0356] 4) Secondary cloning sites

[0357] 5) Flexible DNA extensions enabling efficient circularization (bending) of the backbone molecule

[0358] 6) Unique molecular barcode sequences for labeling each individual backbone molecule

[0359] 7) 5' sequence encoding the other half of the same blunt-end restriction enzyme cleavage site used in 1

[0360] 8) Phosphorylation of the 3' and 5' ends of the backbone molecule

[0361] The flexible DNA extensions are designed with the help of a custom evolution algorithm that utilizes various selection criteria from the following:

[0362] 1) High overall complexity of the sequence

[0363] 2) Absence of repetitive DNA motifs longer than 5 bases

[0364] 3) Absence of self-complementary sequences longer than 5 nucleotides

[0365] 4) Selection of the most flexible sequence in each design iteration loop

[0366] According to the above design, each sequence is manually inspected using the mFold server ( http: / / unafold.rna.albany.edu / ) and modified to minimize the formation of hairpins and generally complex secondary structures.

[0367] 2.2 Preparation of NDA templates for rolling circle amplification

[0368] The following protocol is for the preparation of templates suitable for RCA reactions. Any circular DNA is a suitable template, and different protocols for handling dsDNA or ssDNA are available and known in the art.

[0369] Examples of dsDNA that we can circularize are: cfDNA, ctDNA, sheared genomic DNA, PCR amplicons.

[0370] Examples of ssDNA include: cDNA, viral DNA.

[0371] 2.3 dsDNA circularization reaction

[0372] Here, the dephosphorylated dsDNA molecule is referred to as the "insert", which is ligated to both ends of the phosphorylated backbone to form a circular dsDNA product.

[0373] The reaction is carried out using both DNA ligase and restriction enzyme under suitable buffer conditions.

[0374] The buffer conditions have been optimized to allow ligation, digestion, and PlasmiSafe treatment in a one-pot reaction without the need for an intermediate DNA purification step.

[0375] Considering that the backbone in the example has SrfI half-sites at the outermost ends, the components of the reaction are as follows:

[0376] Buffer 1X

[0377] 50 mM potassium acetate

[0378] 20 mM Tris-acetate

[0379] 10 mM magnesium acetate

[0380] 100 μg / ml BSA

[0381] 1 mM ATP

[0382] 10 mM DTT

[0383] DNA and Enzymes

[0384] Backbone + insert in a 3:1 molar ratio

[0385] 1 unit of T4 DNA ligase

[0386] 1 unit of Srf1

[0387] 1 unit of HMGB1

[0388] Add H2O until the final volume is 20 to 50 μl (depending on the DNA loading amount), then incubate at 22 °C for 1 h, followed by heat inactivation at 65 °C for 15 min.

[0389] The presence of the restriction enzyme increases the total yield of the reaction, avoids the accumulation of backbone concatemers, and at the same time avoids the concatemerization of the insert through preventive dephosphorylation. HMGB1 (high mobility group protein 1) is used to promote the bending of short DNA, thus improving the cyclization efficiency.

[0390] The most abundant product of the above reaction is circular dsDNA containing one backbone and one insert.

[0391] Removal of Linear DNA

[0392] To remove residual linear dsDNA, our template was treated with 1 μl of Plasmid-Safe deoxyribonuclease at 37 °C for 15 min and then heat inactivated at 70 °C for 30 min.

[0393] 3. Materials and Methods for Detection of Fusion Genes

[0394] Based on RNA extracted from human cells, we developed a protocol using circularization and rolling circle amplification to detect fusion genes. In this case, ssDNA (as opposed to dsDNA) was used as the input for the circularization and amplification reactions. This protocol can be generalized to sequence any RNA of interest.

[0395] The first part of the protocol involves standard procedures for "RNA extraction" and "cDNA amplification", e.g., Trizol-based RNA isolation followed by synthesis of polyT primer cDNA using reverse transcriptase.

[0396] After digestion of the RNA of the RNA-DNA hybrid, we obtained linear ssDNA. At this point, we used an ssDNA ligase to self-circularize the input DNA.

[0397] Different ssDNA ligases are available on the market. We used Circ Ligase II (Epicentre) for proof-of-concept experiments. The circular ssDNA obtained according to the supplier's protocol has been successfully used as a template for the RCA reaction, which uses specific primers to direct the amplification of the fusion gene of interest.

[0398] The following protocol details all passages just after RNA isolation.

[0399] 3.1 Removal of Residual DNA

[0400] Buffers, enzymes, and inactivators (TURBO deoxyribonuclease kit) were purchased from Thermo Fisher.

[0401] Mix in a 0.5 ml tube:

[0402] 1 μl of 10x reaction buffer

[0403] 1 μg of extracted RNA

[0404] 0.5 μl of TURBO deoxyribonuclease

[0405] H2O to a final volume of 10 μl

[0406] The next solution was mixed and incubated at 37 °C for 30 min. 2 μl of inactivator was added to inactivate it. Mix for 5 minutes.

[0407] 3.2 cDNA Synthesis

[0408] We used the Super Script II kit from Invitrogen, and any other cDNA transcription kit can be used as an alternative.

[0409] Add to the previous reaction:

[0410] Primer 2 mM (random hexamer or specific) 1 μl The primer is phosphorylated

[0411] dNTP (10 mM each) 1 μl

[0412] Incubate at 65 °C for 5 min, then place on ice for 5 min. During this step, the primer is annealed to the template.

[0413] Next, add:

[0414] 5x First Strand Buffer 4 μl

[0415] 100 mM DTT 1 μl

[0416] Incubate at 42 °C for 2 min, then add 1 μl of SSII enzyme and incubate at 42 °C for 45 min. Finally, we inactivate the reaction at 70 °C for 15 min.

[0417] 3.3 RNA Removal

[0418] RNaseH was purchased from Thermo Fisher.

[0419] To the previous reaction, add 1 μl of RNaseH enzyme and incubate at 37 °C for 20 min, then heat inactivate at 70 °C for 10 min.

[0420] 3.4 Chelation of Divalent Metal Ions

[0421] We added this step to reduce the concentration of free Mg 2+ which would inhibit the ssCirc ligase II used in the next reaction.

[0422] To chelate all free Mg2+, add to the previous reaction:

[0423] 50 mM EDTA stock (0.9 EDTA + 50 ml H2O) 2 μl

[0424] 3.5 ssDNA Circularization

[0425] For this reaction, we used the ssDNA Circ ligase II kit from Epicentre.

[0426] 2 μl of 10x reaction buffer

[0427] 1 μl of MnCl2 (Manganese (Mn) should not be confused with Magnesium (Mg))

[0428] 4 μl of Betaine

[0429] 1 μl of Circ Ligase II

[0430] 10 pMoles of ss-cDNA

[0431] H2O to a final volume of 20 μl

[0432] Incubate at 60 °C for 1 - 2 hours and then heat inactivate the reaction at 70 °C for 10 min.

[0433] At this point, the reaction is treated with Plasmid Safe (optional) and used as a template for the RCA reaction (next step).

[0434] 3.6 Primer annealing

[0435] Depending on the situation, we use random primers or backbone-specific primers or target-specific primers.

[0436] In this step, the template DNA can also be single-stranded circular DNA, as in the case of self-cyclized cDNA.

[0437] The primers have two 3'-terminal phosphorothioate (PTO)-modified nucleotides that are resistant to the 3'→5' exonuclease activity of proofreading DNA polymerases such as phi29 DNA polymerase. They also have 5'-hydroxyl and 3'-hydroxyl termini.

[0438] If the circular DNA is in water, e.g., in the case of previous purification, add 11% of the volume of 10X annealing buffer.

[0439] 10X Annealing Buffer

[0440] 100 mM Tris, pH 7.5 - 8.0

[0441] 500 mM NaCl

[0442] 10 mM EDTA

[0443] Add concentrated primer (50 - 100 μM) to the reaction to a final concentration of 5 μM.

[0444] Adjust the reaction temperature to 98 °C and then slowly cool to room temperature.

[0445] Rolling Circle Amplification

[0446] For a 50 μl reaction, calculate the following volumes.

[0447] When the template is in annealing buffer, take 20 μl of the template and add to it:

[0448] 5 μl of Phi29 buffer (10x)

[0449] 1 μl of BSA

[0450] 1 μl of dNTPs (10 μM)

[0451] 2 μl of pyrophosphatase

[0452] 1 μl of Phi29 DNA polymerase

[0453] H2O to 50 μl

[0454] When the template is in ligation buffer, take 46 μl of the template and add to it:

[0455] 0.5 μl of Uracil-DNA glycosylase (for removing any deaminated cytosine from DNA)

[0456] 0.5 μl of formamidopyrimidine-DNA glycosylase (for removing 8-oxo-guanine products)

[0457] Reaction Conditions

[0458] Depending on the amount of input DNA (template), the reaction is run at:

[0459] >> 3 h @ 30 °C, if the template is 10 - 50 ng

[0460] >> 6 h @ 30 °C, if the template is 5 - 10 ng

[0461] >> 12 h @ 30 °C, if the template is 0.5 - 5 ng

[0462] 4. Materials and Methods for Targeted Cyclomics

[0463] To enable ultra-precise targeted sequencing of any double-stranded DNA molecule, we designed a workflow based on the existing molecular inversion probe (MIP) technology. The uniqueness of this design lies in the MIP capture backbone (minimizing backbone size, adding unique molecular barcodes, probe specificity, and distance), and the combination of assay and rolling circle amplification.

[0464] 4.1 Probe generation

[0465] Use PCR to Amplify Oligonucleotides Away from the Array (MIP Precursors): (2.5 hours)

[0466] 1. The MIP precursor oligonucleotides extracted by the array (a mixture of 100-mers obtained from Agilent) were dissolved in Tris-EDTA buffer with a pH between 8 and 0.1% to a final concentration of 100 nM.

[0467] 2. Prepare the following 400 μl PCR mixture in a 1.5 ml centrifuge tube.

[0468] Reagents Volume (μl) Final Concentration 2x iProof HFPCR Master Mix (Biorad) 200 1x Oligo_Fwd_Amp Primer (100 μM) 2 500 nM Oligo_Rev_Amp Primer (100 μM) 2 500 nM SYBRGreen I 100x (Invitrogen)** 1 0.2X Template (100 nM in 0.1% Tween) 1 250 pM <![CDATA[H2O]]> 194

[0469] It was divided into 8 x 50 μl reactions in 0.2 ml PCR tubes. One PCR preparation yielded approximately 1.5 μg of amplified DNA.

[0470] 3. Use the following PCR cycling program on a real-time thermal cycler such as a Biorad MJ Mini.

[0471] 1) 98 °C for 30 seconds

[0472] 2) 98 °C for 10 seconds

[0473] 3) 60 °C for 30 seconds

[0474] 4) 72 °C for 30 seconds (read plate)

[0475] 5) Repeat steps 2 to 4 for 25 cycles

[0476] 6) 4 °C for an indefinite time

[0477] 4. Combine and clean the PCR reaction on a column using the QIAquick PCR Purification Kit according to the manufacturer's instructions. Elute with 90 μl of elution buffer.

[0478] 5. Use the Qubit High Sensitivity dsDNA Assay Kit to quantify 1 μl of the amplified DNA.

[0479] Capture the exon with a molecular inversion probe.

[0480] 6. Analyze 1 μl of the amplified DNA on a 6% TBE PAGE gel (Invitrogen) to verify the amplification. When the primer is increased by 10 bp, a single band appears at 110 bp for the product.

[0481] Digest the PCR Product with Nicking Restriction Endonuclease to Generate 70-mer MIP (7.5 hours):

[0482] 1. Add 10 μl of NEB-2 (10x) and 5 μl of Nt.AlwI (10 U / μl; NEB) to 85 μl of the PCR product (total volume of 100 μl).

[0483] 2. Mix and aliquot into two tubes, 50 μl each. Incubate in a thermal cycler at 37 °C for 3 h, then at 80 °C for 20 min.

[0484] 3. Cool the temperature to 65 °C and hold for at least 1 min. Add 2.5 μl of Nb.BsrDI (2 U / μl; NEB) to each 50 μl reaction.

[0485] 4. Incubate at 65 °C for 3 h, then at 80 °C for 20 min.

[0486] 5. Purify the two 50 μl digestion reactions on one column using the QIAquick Nucleotide Removal Kit. Elute each column with 30 μl of elution buffer. The yield we observed for this step was 80 - 90%.

[0487] Quantify the Available Probes Using Denaturing Gel (2 hours):

[0488] 1. Accurate quantification of the available MIP in the digested probe mixture is important as it determines how much probe mixture to add to the capture reaction.

[0489] 2. Prepare two-fold dilutions of the NEB 100 bp DNA sequencing ladder (the dilutions we used ranged from 500 ng to 62 ng).

[0490] 3. Mix 2x TBE-Urea sample buffer (Invitrogen) with 1 μl of the digested probe and the above dilution solution.

[0491] 4. We denatured the DNA by heating it to 95 °C for 5 min and then immediately transferring it to ice.

[0492] 5. The samples were run on a precast 6% TBE-urea denaturing PAGE gel (Invitrogen) at 160 V for 1 h.

[0493] 6. Quantify the amount of available MIP in the digested mixture by comparing the sequencing ladder dilutions and the intensity of the 70 bp band. We used this MIP concentration when determining the volume of probe mixture to add to the capture reaction.

[0494] 4.2 Capturing Exons with Molecular Inversion Probes

[0495] Hybridize the Probes to Genomic DNA (37 hours):

[0496] 1. For each sample to be captured, we added the following reagents into a 0.2 ml PCR tube. The final capture reaction volume was 25 μl. Since there is no size selection for the 70 bp MIP, the volume of the probe mixture to be added was based on the concentration of the available MIP.

[0497]

[0498] 2. Denature at 95 °C for 10 minutes.

[0499] 3. Incubate at 60 °C for at least 36 h to hybridize MIP to gDNA.

[0500] Circularize the Captured Exons (1 day):

[0501] 1. Prepare a mixture of ligase and polymerase to be added to each capture reaction:

[0502]

[0503] Place this mixture on ice and keep it cooled before adding 4.7 μl to the capture reaction.

[0504] 1. Incubate at 60 °C for an additional 24 hours to allow gap filling and ligation to circularize the capture region.

[0505] Exonucleases Selected for the Circularized Products: (1 hour)

[0506] 1. Prepare an exonuclease mixture to be added to each capture reaction to remove uncaptured gDNA, excess probes, and blocking oligonucleotides:

[0507] Reagents Volume per Sample (μl) Final Concentration in the Reaction Exo I 20 U / μl 2 1.7 U / μl Exo III 100 U / μl 2 8.3 U / μl

[0508] 2. Lower the temperature of the capture reaction to 37 °C and incubate for at least 1 minute before adding 4 μl of the exonuclease mixture.

[0509] 3. Incubate at 37 °C for 15 minutes.

[0510] 4. Inactivate the exonuclease by heating the reaction at 95 °C for 2 minutes.

[0511] 5. Use 100 ng of the reaction product as a template for rolling circle amplification.

[0512] 4.3 Rolling circle amplification

[0513] 1. Vacuum dry 20 μl of 2X annealing buffer to 1 μl.

[0514] 2. Add 40 μl of circular DNA (about 10 ng).

[0515] 3. Add 4 μl of 50 μM random primers.

[0516] 4. Incubate at 90 °C for 5 minutes, then slowly cool to room temperature.

[0517] 5. Add the following reagents: 10 μl of 10x Phi29 buffer, 2 μl of 100x BSA, 2 μl of 10 mM dNTPs, 2 μl of Phi29 polymerase, 4 μl of pyrophosphatase at 0.1 U / μl, and 40 μl of water.

[0518] 6. Incubate at 30 °C for 19 hours and then at 65 °C for 10 minutes.

[0519] 7. Wash the reaction product with Ampure XP beads (0.4 V).

[0520] The washed reaction product is used for any long-read sequencing protocol.

[0521] 5. List of Designed and Tested Backbones

[0522] >BB1(199bp)

[0523] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0524] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0525] GTAGAGTACGACTGCGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCA

[0526] CGATCGATGACTAGATGCTGCATGACATAGCCC

[0527] >BB2(259bp)

[0528] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0529] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0530] GTCACCGGGTCTTCGAGAAGACCTGTTTAGAGTACGACTGCAAATGGCTCTAGA

[0531] GGTACCCGTTACATAACTTACGCAGATGTGATCAGTGACTACGTGACACTGTACA

[0532] TCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC

[0533] >BB2_100 (341)

[0534] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTC

[0535] ATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTAC

[0536] GTCACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATAT

[0537] ATATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTTAGAGTACGA

[0538] CTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGTGATCAGTGA

[0539] CTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCC

[0540] C

[0541] >BB3(514bp)

[0542] AACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTC

[0543] ACATGTGAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCT

[0544] GTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAA

[0545] ATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGT

[0546] TTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCT

[0547] TTATATATCTTGTGGAAAGGACGAAACACCGGGTCTTCGAGAAGACCTGTTTTAG

[0548] AGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTG

[0549] GCACCGAGTCGGTGCTTTTTTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGG

[0550] CTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGCAGACAAATGGCTCTAGAGGTAC

[0551] CCGTTACATAACTTA

[0552] >BBpX2(557bp)

[0553] GGGCATGCACAGATGTACACGAACGCCAGCAACGCGGCCTTTTTACGGTTCCTG

[0554] GCCTTTTGCTGGCCTTTTGCTCACATGTGAGGGCCTATTTCCCATGATTCCTTCAT

[0555] ATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTA

[0556] AACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGT

[0557] AGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTT

[0558] GAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGG

[0559] GTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGT

[0560] CCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTGTTTTAGAGCTA

[0561] GAAATAGCAAGTTAAAATAAGGCTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGC

[0562] AGACAAATGGCTCTAGAGGTACCCGTTACATAACTTATAGATGCTGCATGACATAG

[0563] CCC

[0564] 6. Additional Details on the Generation of BB2_100 and BBpX2

[0565] BB2 was optimized by inserting a flexible sequence using the BbsI cloning site present in BB2. An insertion consisting of a 100 bp long DNA extension obtained in silico design (see Section 7) was added, and for cloning purposes, a BbsI restriction site (bold in the sequence below) was added at the very end. The complete insertion (sense and antisense) was obtained by annealing two shorter oligonucleotides. The oligonucleotides were ordered as single-stranded oligonucleotides from IDT DNA Technologies. The forward and reverse strands were annealed, and the annealed product, now dsDNA with sticky ends, was resolved on an agar gel. Subsequently, the insertion was cloned into BB2 by a Golden-Gate cloning reaction similar to the reaction described by (Ran et al. 2013).

[0566] The complete insertion sequence and oligonucleotides used to generate it are as follows:

[0567] >Insert BB2_100

[0568] CACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATATAT

[0569] ATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTT

[0570] >Sense oligonucleotide

[0571] CACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATATAT

[0572] ATATATATGGATATGTATGGATATATATATATATGGATATGGAT

[0573] >Antisense oligonucleotide

[0574] AAACATCCATATCCATATATATATATATCCATACATATCCATATATATATATATATATCC

[0575] ATATATCCATATATATATATATCCATATATATATCCATATATAT

[0576] BBpX2 was obtained by adding an SrfI-half-site (GGGC) to the extreme end of the PCR amplicon and the remaining universal primer sequences (underlined in the sequences below). BB3 was used as the template for the PCR reaction.

[0577] >BBpX to BBpX2-F

[0578]

[0579] >BBpX to BBpX2-R

[0580]

[0581] The above sequences have a template annealing part in lowercase and a flanking region in uppercase. The SrfI-half-site is highlighted in orange, and the remaining uppercase sequences are part of the constant sequences present at the ends of all backbones. The constant sequences are not essential for the backbones, but are useful for standardizing their amplification during the production steps. The following primers are capable of amplifying any backbone manufactured so far:

[0582] >Universal SrfI-BB-F

[0583] GGGCATGCACAGATGTACACG

[0584] >Universal SrfI-BB-R

[0585] GGGCTATGTCATGCAGCATCTA

[0586] 7. List of Insertion Sequences

[0587] >Insertion 17.1 (TP53, chr17:7576971-7577132)

[0588] TAACTGCACCCTTGGTCTCCTCCACCGCTTCTTGTCCTGCTTGCTTACCTCGCTTA

[0589] GTGCTCCCTGGGGGCAGCTCGTGGTGAGGCTCCCCTTTCTTGCGGAGATTCTCTT

[0590] CCTCTGTGCGCCGGTCTCTCCCAGGACAGGCACAAACACGCACCTCAAAG

[0591] >Insertion 17.2 (TP53, chr17:7578161-7578394)

[0592] CAGTTGCAAACCAGACCTCAGGCGGCTCATAGGGCACCACCACACTATGTCGAA

[0593] AAGTGTTTCTGTCATCCAAATACTCCACACGCAAATTTCCTTCCACTCGGATAAG

[0594] ATGCTGAGGAGGGGCCAGACCTAAGAGCAATCAGTGAGGAATCAGAGGCCTGG

[0595] GGACCCTGGGCAACCAGCCCTGTCGTCTCTCCAGCCCCAGCTGCTCACCATCGC

[0596] TATCTGAGCAGCGCTCAT

[0597] 8. In Silico Design of Flexible DNA Sequences

[0598] To improve the flexibility of BB2, a 100 bp sequence was added to the flexible DNA sequence, which was specifically designed by a simple genetic algorithm. The same method was used to de novo design the entire backbone core sequence. Restriction sites, barcodes, or primer sites of these backbone core sequences can be added as described elsewhere in this article.

[0599] The backbone core sequence was optimized based on an evolutionary selection algorithm, which optimized the following components of the sequence:

[0600] 1) High molecular flexibility

[0601] 2) High sequence entropy

[0602] 3) GC content between 30% and 60%, ideally close to 50%

[0603] 4) Lack of long self-complementary extensions

[0604] 5) Lack of long oligonucleotide polymers (NNNNNNN)

[0605] 6) Lack of repeated motifs (kmer)

[0606] Flexibility calculation

[0607] Use the Python implementation of the TwistFlex algorithm (http: / / margalit.huji.ac.il / TwistFlex / ) (Menconi et al. 2015) to calculate the DNA flexibility at the twist angles of the input sequence. The flexibility of each single dinucleotide is calculated according to the following angle table:

[0608]

[0609] Subsequently, the evolutionary algorithm for backbone optimization takes into account the average flexibility of the entire sequence when selecting. The average flexibility of the DNA sequence is calculated as the sum of all dinucleotide angles divided by the total number of dinucleotides. The flexibility threshold for our backbone is an average of 12.5 degrees. Discard any sequence with an average flexibility lower than 12.5 degrees.

[0610] Entropy calculation for determining sequence complexity

[0611] The Shannon entropy of a string is defined as the minimum average number of bits per symbol required to encode that string. The formula for calculating Shannon entropy is:

[0612]

[0613] where p i is the probability that the character number i appears in the sequence. This calculation can also be performed through http: / / www.shannonentropy.netmark.pl /

[0614] Implement the above formula using the following Python code:

[0615]

[0616]

[0617] The minimum entropy value expected for our backbone core sequence is 1.5 Sh. Discard each sequence with a smaller entropy value.

[0618] Self-complementarity

[0619] Filter the selected backbone core sequences by the presence of 8-base stretches of self-complementarity. Backbones having 8 or more consecutive bases of self-complementarity on the same strand are discarded.

[0620] Absence of repetitive motifs (kmers)

[0621] Filter out backbone core sequences containing motifs with 6 bases that are repeated more than twice in the sequence.

[0622] Design an evolutionary algorithm for newer backbones (past BB2)

[0623] The newer backbones consist of flexible DNA and a pair of fixed sequences at the ends (the universal SrfI-BB F / R described in paragraph 6), which serve as primer annealing sites for backbone PCR amplification and for adding semi-restriction enzyme sites.

[0624] As with any genetic algorithm (GA), Cyclomics' GA consists of a main loop in which sequence libraries are scored and selected. The selected sequences are then used as input (parent sequences) to generate new sequences (child sequences). The parent and child sequences are grouped in a new library, prepared for the next iteration. The pseudocode for this loop is as follows:

[0625] For each iteration:

[0626]

[0627] The algorithm is implemented entirely in Python, with the pairing operator, mutation operator, and main loop all implemented from scratch according to general guidelines in the literature (Hwang and Jang 2008; 1996; Coello Coello and Lamont 2004; Lobo, Lima, and Michalewicz 2007). The pairing operator acts on strings, sequences, and makes a single swap at a random position. The mutation operator adds random mutations to the parent or child sequences, which may include small deletions and duplications. The filtering step is used to prune sequence libraries with CG content imbalance, low sequence entropy, and unwanted repeated kmers before the selection step. Selection itself simply collects the best sequences according to a flexibility score. The selected sequences are used as parent sequences to generate child sequences using the pairing operator. To calculate sequence flexibility, existing code (Menconi et al. 2015) was adapted to suit our purposes.

[0628] 9. Results

[0629] PCR was performed using primers 17.2-F (CAGTTGCAAACCAGACCTCA) and 17.2-R (ATGAGCGCTGCTCAGATAG) to obtain a PCR product of 234 bp in length that covered the coding exon TP53 (chr17:7578161-7578394, GRCh37). According to the standard procedure, the DEPCR product designated 17.2 was ligated to pJET (Thermo Fisher). The ligation product was transformed into E. coli Top10, and a single colony was picked for collection of the (clonally propagated) pJET-17.2 plasmid. The sequence of 17.2 was verified by Sanger sequencing and found to be identical to the reference genome (GRCh37). The phosphorothioate (PTO)-modified primer 17.2-R (ATGAGCGCTGCTCAGATA*G*, where * indicates PTO modification) (5 μM) was annealed to 50 ng of pJET-17.2 in a final volume of 20 μL in the presence of 5 mM EDTA. The reaction mixture was heated to 95 °C for 5 minutes and then cooled to 4 °C.

[0630] The 20-μL annealing reaction was supplemented with 0.2 μ inorganic pyrophosphatase (Thermo Fisher), 10 μL Phi29 (NEB), 1 μL of 10 mM dNTPs, 1 μL of 100X BSA solution (NEB, 20 mg / mL), and 5 μL of Phi29 10X reaction buffer (NEB). The resulting reaction mixture was incubated at 30 °C for 3 hours and then at 65 °C for 10 minutes. The amplified high-molecular-weight DNA was purified using Ampure beads (Agencourt), and then a 1D nanopore library (Oxford nanopore Technologies, SQK-LSK108) was prepared. The resulting library was run on a MinION flow cell (FLOI-MIN 106, R9.4 chemistry) for 48 hours.

[0631] 10. DNA Quantification by Gel Densitometry

[0632] Agarose gel densitometry is a method for quantifying DNA by image analysis of gel bands that is performed by comparing the pixel brightness between 1) a sequence ladder and a band of interest or 2) an input band and a product band.

[0633] 10.1 Using Sequence Ladders as References

[0634] Given a picture of an agar gel containing a DNA sequence ladder of known amounts, we can use ImageJ software to estimate the brightness intensity of the bands. The correct circularized products are quantified and compared to the input to calculate the circularization efficiency. The area and average intensity of the bands on each gel are determined using the measurement function of ImageJ. The average intensity of the determined image background is as close as possible to the band in question and subtracted from the band intensity. The resulting intensity is multiplied by the area of the band (referred to as the level). To create a reference level, the ratio of the band corresponding to 400 base pairs in each image in the 50 bp DNA sequence ladder (Thermo Fisher Scientific), which shows the intensity of 15 ng of DNA, was also calculated. To calculate the DNA content (in ng) of each band, the calculated level was divided by the reference level and multiplied by 15. The molar content of DNA: μg to pmol was determined using the Promega DNA conversion tool dsDNA. The circularization efficiency is calculated by dividing the correctly formed product in moles by the molar input of the insert and multiplying by 100%. To validate this method, the DNA content of other bands in the DNA sequence ladder was calculated and compared to the predicted DNA content. The DNA content of each band was plotted against the predicted and calculated values ( Figure 6 ).

[0635] 10.2 Direct Comparison of Input Bands and Product Bands

[0636] The optional process of estimating circularization efficiency by gel image analysis requires at least two lanes on the gel, one for the input DNA and the other for the product. Using ImageJ, we can estimate the ratio between the insert (input) before the reaction and the remaining insert (unreacted) after the reaction.

[0637] Measure and compare the brightness of the bands within the yellow rectangles (input DNA on the left band and unreacted DNA on the right band). In this case, the ratio between the input and the unreacted is 66:33. We can conclude that 50% of the initial DNA has reacted ( Figure 7 A).

[0638] Next, we compare the product bands ( Figure 7 B).

[0639] We know that the top band represents the desired product, and the lower band represents the uncircularized product. From the ratio of these two bands, we can establish the circularization efficiency, which is defined as the amount of input DNA that is correctly circularized into the final product. If A is the ratio of the input that reacted, and B is the ratio of the correct product, then the efficiency is A * B. In this case, 50% * 50% = 25%.

[0640] References

[0641] - Thomas. 1996. Evolutionary Algorithms in Theory and Practice: Evolution Strategies, Evolutionary Programming, Genetic Algorithms. Oxford University Press on Demand.

[0642] - Coello Coello, Carlos A., and Gary B. Lamont. 2004. Applications of Multi-Objective Evolutionary Algorithms. World Scientific.

[0643] - Hwang, Gi-Hyun, and Won-Tae Jang. 2008. “An Adaptive Evolutionary Algorithm Combining Evolution Strategy and Genetic Algorithm (Application of Fuzzy Power System Stabilizer).” In Advances in Evolutionary Algorithms.

[0644] - Lobo, F.J., Claudio F. Lima, and Zbigniew Michalewicz. 2007. Parameter Setting in Evolutionary Algorithms. Springer Science & Business Media.

[0645] - Menconi, Giulia, Andrea Bedini, Roberto Barale, and Isabella Sbrana. 2015. “Global Mapping of DNA Conformational Flexibility on Saccharomyces Cerevisiae.” PLoS Computational Biology 11(4): e1004136.

[0646] -Ran, F. Ann, Patrick D. Hsu, Jason Wright, Vineeta Agarwala, David A. Scott, and Feng Zhang. 2013. “Genome Engineering Using the CRISPR-Cas9 System.” Nature Protocols 8(11): 2281-2308.

[0647] Example 2

[0648] Sequencing of tandem DNA molecules

[0649] We cloned the PCR product covering locus 17:7578265 of the TP53 gene into pJET (Materials and Methods, Section 9). After bacterial transformation, single colonies were picked and plasmid DNA was isolated to confirm the presence of the TP53 insert (data not shown). Next, we performed rolling circle amplification (RCA) on the isolated plasmid using phi29 polymerase and random hexamer primers (Materials and Methods, Section 9). We obtained high molecular weight RCA products larger than 20 kb as estimated by gel electrophoresis ( Figure 8 A). This product was used as the input for a 1D library preparation, which was sequenced on an Oxford Nanopore Technologies (ONT) MinION instrument, and the library was sequenced for 48 hours according to the manufacturer's instructions. A total of 16,248 sequencing reads were generated from this sample, with an average read length of 5.7 kb ( Figure 8 B), and 2,083 reads longer than 10 kb. Using LAST ( et al. 2011), the nanopore / MinION sequencing reads were mapped to the human reference genome (GRCh37, amplified with the pJET sequence). A subset of 2,083 reads longer than 10 kb was examined to determine the configuration of the replacement backbone (pJET) and the insert (TP53 fragment), which is referred to as BI and is likely to be expected from the circularized template used as the input. We observed that all 2,083 reads had multiple copies of BI. We calculated a “mode score” for all reads longer than 10 kb, which is a number between 1 and 100 representing the regularity of the BI repeats, calculated as BI / ((B + I) / 2)*100, where BI is the amount of the pJET-17.2 fragment in a nanopore read, B is the amount of the pJET fragment in a nanopore read, and I is the amount of the 17.2 fragment in a nanopore read. The mode scores for most of the long reads were 100, indicating the correctness of the RCA product, i.e., the repetition of the BI units. Figure 8C). We used these reads to extract the inserted sequences (17.2-TP53 fragments), and subsequently aligned the insertions of each read using Muscle (Edgar 2004a, [b]2004)( Figure 8 D). For each read, we applied the aligned 17.2 fragments to a majority voting scheme to obtain a consensus sequence. The consensus sequence was compared with the reference sequence to determine the accuracy as a function of the BI copy number in the reads( Figure 8 E). This experiment demonstrated the concept of obtaining accurate consensus sequence reads by nanopore sequencing of multiple copies of DNA molecules (Li et al. 2016). In the next step, the backbone design was improved to reduce the sequencing throughput of the backbone (pJET-3kb in the above experiment).

[0650] Optimization of the linear dsDNA backbone

[0651] Backbone Design Principle

[0652] As a first step in optimizing the capture and circularization of short DNA molecules, we tested different designs of the backbone sequences that mediate this process. We compared three parameters: the length of the backbone (longer DNA molecules (backbones) are thought to be more easily circularized, but this results in a waste of sequencing information as most of each read will consist of the backbone sequence). Second, DNA molecules shorter than about 200 bp are thought to be difficult to circularize due to their relative stiffness (Shore, Langowski, and Baldwin 1981). Third, the flexibility of a given DNA molecule also depends on its base composition and sequence.

[0653] The first-generation backbones were BB1 and BB2 (sequences in Section 5 of Materials and Methods), which were designed according to the general principles highlighted in Section 2 of Materials and Methods. These backbones were designed to serve as basic building blocks that we could improve. These backbones included a combination of several elements that could help capture DNA molecules and subsequent amplification and sequencing, such as restriction enzyme sites, barcode sequences, and / or nicking enzyme sites. We used BB1 to detect circularization, but it was not very effective and was not carried forward for further testing.

[0654] Shore, Langowski, and Baldwin (1981) proposed that the cyclization of short DNA molecules with blunt ends might be suboptimal if efficiency was required. Next, we generated longer backbones that would allow for more efficient cyclization within a short time period (∼1 h). The resulting backbone was named BB3, a 514-bp long dsDNA fragment generated by PCR amplification of a portion of plasmid pX330 (Materials and Methods, Sections 5 and 6). BB3 does not contain restriction enzyme sites at its ends, and it was only used to test different ligation conditions (see below). We also generated a modified version of BB3 by adding SrfI sites, resulting in BBpX2.

[0655] We used the free energy values per base pair (Breslauer et al. 1986) and the deviation of the twist angle (degrees) (Sarai et al. 1989) to calculate the flexibility of any given DNA sequence. We used a genetic algorithm to generate a set of sequences that were selected for high flexibility, short length, and optimized for other parameters such as GC content, the presence of repetitive motifs, and the self-complementarity of the sequence. A more detailed description of the genetic optimization algorithm and more details of the backbone structure are given in Materials and Methods, Section 8. To increase the cyclization efficiency of the backbone while keeping the sequence short, we computationally designed several extensions of flexible DNA. We used such extensions to increase the flexibility of the backbone. A 100-bp long flexible sequence was added in the middle of the sequence of backbone 2 (BB2) to modify backbone 2. The resulting backbone (BB2_100) is 341 bp long and includes SrfI half-restriction enzyme sites at its very ends.

[0656] Measure the Reaction Products of BB2 Backbone and BB3 Backbone in the Circularization Reaction

[0657] We tested BB2 and BB3 as inserts together with PCR products in the cyclization reaction, 17.2 (Materials and Methods, Sections 2, 5, and 7). The first experiment was conducted to determine the best reaction conditions for obtaining optimal cyclization products. The reactions were carried out with and without the addition of Plasmid Safe deoxyribonuclease to obtain a clear view of the linear and cyclized reaction products ( Figure 9 A). The results of the cyclization efficiency of BB2 were consistent, but the cyclized products were not very abundant. The cyclization reaction with BB3 resulted in visible circular products ( Figure 9 A), while no clear reaction products were visible with BB2. However, running the entire reaction mixture on a gel allowed for the observation of weak bands after Plasmid Safe digestion, indicating that the correct cyclized products consisted of BB2 and insert 17.2 ( Figure 9 B).

[0658] Effect of Different Backbone-Insert Ratios on Circularization Efficiency

[0659] Next, we evaluated the effect of different backbone - insert ratios on the specificity and efficiency of circular backbone - insert product formation, which was also our goal. Therefore, we performed cyclization experiments using BB3, along with a 234 bp PCR product (17.2, Materials and Methods, section 7)( Figure 10 )

[0660] The best cyclization efficiency was obtained with a backbone - to - insert molar ratio of 3:1. This setting was maintained for further characterization. Note that the strategy used here is very different from that used in standard plasmid - based cloning. In standard cloning, plasmids are usually dephosphorylated to avoid self - cyclization and an excess of phosphorylated insert is added to the reaction. For the Cyclomics technology, the backbone is phosphorylated and over - phosphorylated, while the insert is dephosphorylated. We found that this avoided ligation - dependent target concatemerization and increased the ligation efficiency of backbone targets.

[0661] Effect of Flexible Extension on Backbone Circularization

[0662] In subsequent experiments, we tested the effect of adding flexible DNA extensions on backbone cyclization. Therefore, we compared the cyclization efficiency of BB2 with BB2_100 (see above) and BB3 in the cyclization reaction of insert 17.2( Figure 11 ) The flexible region in BB2_100 has a rich TA repeat but is still complex enough to be clearly mapped to the reference sequence. In this test, we also evaluated the effect of HMGB1 and SrfI on the cyclization reaction of BB2_100. HMGB1 is a known DNA - bending protein that can potentially improve cyclization (Belgrano et al. 2013). We observed an increase in the cyclization efficiency of BB2_100 compared to BB2, especially when considering a backbone:insert ratio of 3:1. Therefore, we concluded that the backbone design can be optimized by adding flexible DNA extensions to improve cyclization efficiency. The overnight cyclization of BB2_100 and 17.2 gave a higher cyclization efficiency, estimated to be approximately 26%( Figure 12 ), indicating that better reaction performance can be obtained by adjusting both the backbone design and the reaction conditions.

[0663] Effect of SrfI on the Formation of Backbone Circularization Products

[0664] A necessary part of our backbone is the presence of separate restriction enzyme cleavage sites at the very ends. If the backbone self-circularizes without an insert, the complete restriction enzyme cleavage sites are reconstituted, making the backbone vulnerable to specific nucleases. In the following example, a half-restriction enzyme cleavage site of SrfI (GCCC|GGGC) is added at the end of BB3 to generate a new backbone we call BBpX2, and SrfI nuclease and T4 ligase are added to the reaction mixture. The advantage of SrfI is that it can recognize an 8-base-long site, while most commercially available alternatives can only recognize 6-base-long sites. Other sequences we are evaluating are PmeI (GTTT|AAAC) and SweI (ATTT|AAAT). If the ligation reaction is carried out in the presence of Srf1, any self-circularized backbone is vulnerable to restriction enzyme cleavage and thus returns to its original linear form.

[0665] By comparing the first two lanes, the effect of the restriction endonuclease can be clearly seen ( Figure 13 ). When SrfI is present, the linear backbone is maintained (thick bold band), and the whole reaction results in few by-products. In the absence of SrfI (first lane), most of the backbone is wasted in the formation of several by-products. The effect of SrfI can be further evaluated by the treatment effect of Plasmid Safe deoxyribonuclease (last two lanes), which results in the degradation of linear DNA. If SrfI is added to the reaction (last lane), only the expected product is generated, whereas if SrfI is not added (third lane), many unwanted circular by-products are produced.

[0666] Dephosphorylation of Insert

[0667] To avoid self-polymerization of the insert, we use Antarctic phosphatase for enzymatic dephosphorylation, which ensures high reactivity at low temperatures and is completely inactivated in only 5 minutes at 65 °C.

[0668] Barcoding strategy

[0669] Molecular barcoding is a strategy for tagging individual DNA molecules in order to classify the sequencing results of DNA molecules. Barcodes can be used to classify sequencing reads by sample (bioinformatically), thus allowing multiple samples to be pooled in a single sequencing experiment (Wong, Jin, and Moqtaderi 2013). In this case, only a limited number of unique barcodes are used, one for each sample.

[0670] In addition, barcodes can be used to individually label each DNA molecule, and such barcodes are commonly referred to as unique molecular identifiers (UMIs). In this case, a large number of unique barcodes / UMIs (random sequences) are used to minimize the chance that any two unrelated sequences will obtain the same barcode. UMIs can be used to obtain absolute quantification of individual sequences (Kivioja et al. 2011).

[0671] Another application of UMIs is the detection and quantification of low-frequency mutations (Kou et al. 2016), such as in cancer samples. This involves labeling individual DNA molecules, followed by PCR amplification and deep sequencing. Subsequently, sequence reads can be grouped by UMI sequence, and possible mutations can be detected and identified from sequencing errors. Newman et al. (Newman et al. 2016) outlined the superior application of UMIs in ctDNA mutation detection.

[0672] We envision designing the backbone sequence with sample-specific barcodes and UMIs ( Figure 14 ). This strategy enables the pooled sequencing of multiple independent samples and enhanced mutation detection capabilities. The sample-specific barcodes are 5-20 nucleotides in length and can be placed at any position in the backbone sequence, provided that they do not affect the flexibility of the backbone (and thus ligation efficiency). A random strand of 5-20 nucleotides representing the UMI will also be added to the backbone to label individual DNA molecules. By requiring at least two or more different molecules with mutations, i.e., both molecules should have unique UMIs, UMIs can be used to improve mutation detection.

[0673] Rolling circle amplification of circular DNA molecules

[0674] The circular DNA product obtained by the backbone and insert circularization reaction can serve as a template for generating concatemers via rolling circle amplification (RCA). We tested RCA using DNA (inserts) from different sources, including cfDNA, PCR amplicons, plasmids, and cDNA ( Figure 15 ), and using random hexamer primers.

[0675] Site-directed RCA

[0676] In addition to the typical RCA reaction, which involves random hexamers used to initiate amplification, we designed a method of using specific primers to direct amplification to regions of interest, a method known as site-directed RCA. This method may be useful in cases where only specific genes rather than the entire genome need to be sequenced. Currently, the way to implement this method is through PCR enrichment of the gene of interest (Dowthwaite and Pickford 2015). However, it is well known that PCR amplification increases errors in amplicons (Shuldiner, Nirula, and Roth 1989); (Diaz-Cano 2001), and even a single amplification error that occurs early in the PCR reaction can lead to bias in the final result (Diaz-Cano 2001; Quach, Goodman, and Shibata 2004); (Arbeithuber, Makova, and Tiemann-Boege 2016).

[0677] To test whether we could achieve site-specific enrichment of a target region without using PCR, we combined the circularomics assay with site-directed RCA. Briefly, we cloned two different regions 17.1 and 17.2 of the TP53 gene (Materials and Methods, Section 7) into the pJET vector. The modified pJET vector was used as the template for the RCA reaction at a 1:1 molar ratio, where a specific primer targeting 17.2 rather than 17.1 (17.2R, Materials and Methods, Section 9) was used instead of random hexamers. The reaction products were sequenced using the nanopore MinION instrument, and the number of reads containing 17.1 and 17.2 was compared ( Figure 16 ). We observed that sequencing reads containing 17.2 appeared at 14x coverage compared to reads containing 17.1, indicating that target-selective RCA can be achieved with specific primers.

[0678] One-pot reaction design

[0679] To enhance the use of circularomics technology, we focused on developing a streamlined experimental procedure. Therefore, we restricted time-consuming and laborious steps such as DNA purification, concentration, and gel electrophoresis as much as possible. For this purpose, we designed a protocol consisting of three simple consecutive steps that can be carried out in a single tube to limit the need for purification or buffer exchange.

[0680] The steps are: 1) circularization, 2) removal of linear DNA, and 3) rolling circle amplification ( Figure 17 ).

[0681] The first reaction of the circleomics protocol involves inserting DNA (I) and backbone (BB), which are mixed together in the presence of T4 DNA ligase and the restriction enzyme SrfI. The mixture is left at room temperature for 1 to 4 hours and then the enzymes are heat inactivated at 70 °C for 30 minutes. The second step of the circleomics protocol is to add Plasmid Safe enzyme, its buffer, and 1 mM ATP to the reaction mixture. The mixture is incubated at 37 °C for 30 minutes and then inactivated. Before proceeding with rolling circle amplification (Reaction 3), RCA-primers are added to the mixture and a rapid annealing step is performed by heating the reaction to 98 °C for 5 minutes. After the mixture has cooled to room temperature, Phi29, pyrophosphatase, and the other components of the RCA reaction are added. The reaction is then incubated at 30 °C for at least 3 hours.

[0682] Consensus sequence calling

[0683] To detect mutations from the long reads of concatemers, a consensus sequence ( Figure 18 ) of the target sequence was generated. For this purpose, based on the LAST split-reads mapped to the reference genome, the long reads were separated into backbone sequences and target sequences ( et al. 2011). The target sequences were passed to the GATK UnifiedGenotyper for variant calling (DePristo et al. 2011). Causal filtering was employed to optimize sensitivity and specificity based on the variant confidence scores.

[0684] Application examples of circleomics technology

[0685] Targeted sequencing of TP53 mutations in ovarian cancer genomic DNA

[0686] We tested the circleomics method in three tumor biopsies with known TP53 mutations (chr17:7578265, A->T, hg19) with variable frequencies of the mutation (1%, 9%, 14%), as previously evaluated using short-read targeted IonTorrent sequencing (Hoogstraat et al. 2014). Briefly, we performed PCR on the target site and ligated the resulting products to a backbone that was specifically designed and optimized to facilitate efficient capture of short DNA products. Subsequently, the ligation products were amplified and concatemerized to form long DNA molecules with repeated copies of the target / insert and the backbone. The long DNA molecules were sequenced for several hours using a nanopore MinION instrument (1D ligation-based library preparation). We obtained a total of 206,048 sequence reads for all three samples, which were processed by mapping to LAST ( et al. 2011) and a custom algorithm for consensus sequence calling ( Figure 18)。Next, we estimated the mutation frequencies from the consensus sequences and observed TP53 mutation frequencies of 0.5%, 7.6%, and 14%, respectively, providing proof of concept for the detection of low-frequency somatic mutations in cancer DNA using circleomics technology( Figure 19 )。

[0687] References

[0688] -Arbeithuber, Barbara, Kateryna D. Makova, and Irene Tiemann-Boege. 2016. “Artifactual Mutations Resulting from DNA Lesions Limit Detection Levels in Ultrasensitive Sequencing Applications.” DNA Research: An International Journal for Rapid Publication of Reports on Genes and Genomes 23(6): 547-59.

[0689] -Belgrano, Fabricio S., Isabel C. de Abreu da Silva, Francisco M. Bastos de Oliveira, Marcelo R. Fantappié, and Ronaldo Mohana-Borges. 2013. “Role of the Acidic Tail of High Mobility Group Protein B1 (HMGB1) in Protein Stability and DNA Bending.” PloS One 8(11): e79572.

[0690] -Breslauer, K. J., R. Frank, H. Blocker, and L. A. Marky. 1986. “Predicting DNA Duplex Stability from the Base Sequence.” Proceedings of the National Academy of Sciences 83(11): 3746-50.

[0691] - DePristo, Mark A., Eric Banks, Ryan Poplin, Kiran V. Garimella, Jared R. Maguire, Christopher Hartl, Anthony A. Philippakis, et al. 2011. “A Framework for Variation Discovery and Genotyping Using next-Generation DNA Sequencing Data.” Nature Genetics 43(5): 491-98.

[0692] - Diaz-Cano, Salvador J. 2001. “Are PCR Artifacts in Microdissected Samples Preventable?” Human Pathology 32(12): 1415.

[0693] - Dowthwaite, Gary, and Jo Pickford. 2015. “PCR-Based DNA Enrichment Enhances Detection of Mutations in Oncology.” MLO: Medical Laboratory Observer 47(11): 18, 20.

[0694] - Edgar, Robert C. 2004a. “MUSCLE: Multiple Sequence Alignment with High Accuracy and High Throughput.” Nucleic Acids Research 32(5): 1792-97.

[0695] - ———. 2004b. “MUSCLE: A Multiple Sequence Alignment Method with Reduced Time and Space Complexity.” BMC Bioinformatics 5(August): 113.

[0696] - Hoogstraat, Marlous, Mirjam S. de Pagter, Geert A. Cirkel, Markus J. van Roosmalen, Timothy T. Harkins, Karen Duran, Jennifer Kreeftmeijer, et al. 2014. “Genomic and Transcriptomic Plasticity in Treatment-Naive Ovarian Cancer.” Genome Research 24(2): 200-211.

[0697] - Szymon M., Raymond Wan, Kengo Sato, Paul Horton, and Martin C. Frith. 2011. “Adaptive Seeds Tame Genomic Sequence Comparison.” Genome Research 21(3): 487-93.

[0698] - Kivioja, Teemu, Anna Kasper Karlsson, Martin Bonke, Martin Enge, Sten Linnarsson, and Jussi Taipale. 2011. “Counting Absolute Numbers of Molecules Using Unique Molecular Identifiers.” Nature Methods 9(1): 72-74.

[0699] - Kou, Ruqin, Ham Lam, Hairong Duan, Li Ye, Narisra Jongkam, Weizhi Chen, Shifang Zhang, and Shihong Li. 2016. “Benefits and Challenges with Applying Unique Molecular Identifiers in Next Generation Sequencing to Detect Low Frequency Mutations.” PloS One 11(1): e0146638.

[0700] -Li, Chenhao, Kern Rei Chng, Esther Jia Hui Boey, Amanda Hui Qi Ng, Andreas Wilm, and Niranjan Nagarajan. 2016. “INC-Seq: Accurate Single Molecule Reads Using Nanopore Sequencing.” GigaScience 5(1): 34.

[0701] -Newman, Aaron M., Alexander F. Lovejoy, Daniel M. Klass, David M. Kurtz, Jacob J. Chabon, Florian Scherer, Henning Stehr, et al. 2016. “Integrated Digital Error Suppression for Improved Detection of Circulating Tumor DNA.” Nature Biotechnology 34(5): 547-55.

[0702] -Quach, Nancy, Myron F. Goodman, and Darryl Shibata. 2004. “In Vitro Mutation Artifacts after Formalin Fixation and Error Prone Translesion Synthesis during PCR.” BMC Clinical Pathology 4(1). doi:10.1186 / 1472-6890-4-1.

[0703] -Sarai, A., J. Mazur, R. Nussinov, and R. L. Jernigan. 1989. “Sequence Dependence of DNA Conformational Flexibility.” Biochemistry 28(19): 7842-49.

[0704] -Shore, D., J. Langowski, and R. L. Baldwin. 1981. “DNA Flexibility Studied by Covalent Closure of Short Fragments into Circles.” Proceedings of the National Academy of Sciences of the United States of America 78(8): 4833-37.

[0705] -Shuldiner, Alan R., Ajay Nirula, and Jesse Roth. 1989. “Hybrid DNA Artifact from PCR of Closely Related Target Sequences.” Nucleic Acids Research 17(11): 4409-4409.

[0706] -Wong, Koon Ho, Yi Jin, and Zarmik Moqtaderi. 2013. “Multiplex Illumina Sequencing Using DNA Barcoding.” Current Protocols in Molecular Biology / Edited by Frederick M. Ausubel... [et Al.] Chapter 7: Unit 7.11.

[0707] Example 3

[0708] Materials and Methods

[0709] Circularization and RCA Amplification of Short PCR Oligonucleotides

[0710] Materials

[0711] Backbone (BB) BB2.4 with barcode 10-50 ng / μl 243-244 bp

[0712] Insert (I) Blunt-ended PCR amplicons 10-50 ng / μl 100-250 bp

[0713] CutSmart Buffer 10X (supplied by NEB#R0629)

[0714] ATP 10 mM (NEB#P0756)

[0715] dNTPs 10 mM (ThermoFisher#R0192)

[0716] T4 Ligase 400 U / μl (NEB#M0202S)

[0717] SrfI (Restr.Enz.) 20 U / μl (NEB#R0629)

[0718] Plasmid-Safe Buff. 10X (Lucigen#E3101K)

[0719] Plasmid-Safe Enz. 10 U / ul (Lucigen#E3101K)

[0720] Annealing Buffer 5X (50 mM Tris@pH 7.5 - 8.0, 250 mM NaCl, 5 mM EDTA)

[0721] Phi29 Buffer 10X (supplied by ThermoFisher#EP0091)

[0722] BSA 10 mg / ml (NEB#B9001)

[0723] Pyrophosphatase 0.1 U / μl (ThermoFisher#EF0221)

[0724] Phi29 DNA Polym. 10 U / μl (ThermoFisher#EP0091)

[0725] Exo-Res.RND Primer 500 μM (ThermoFisher#SO181)

[0726] Wizard SV Gel and PCR Clean-Up System (Promega#A9282)

[0727] The backbone must be phosphorylated, either by using phosphorylated primers via PCR production, or by using PNK of non-phosphorylated PCR products or synthetic DNA duplexes (T4 polynucleotide kinase).

[0728] The insert must be amplified via PCR with non-phosphorylated primers or dephosphorylated using Antarctic phosphatase.

[0729] Both the insert and the backbone must be blunt-ended. The preferred method is to use Phusion polymerase (amplicon leaving blunt-ended termini).

[0730] Both the insert and the backbone must be bufferless when purified by column or beads.

[0731] If, in the PCR reaction for the production of BB or I, more than one product is generated, then gel purification of the expected product is necessary.

[0732] If the template for I or BB amplification is circular (e.g., plasmid), then gel purification of the PCR product is necessary.

[0733] Method

[0734] Circularization

[0735] Reaction mixture (1X): (The molar ratio of BB:I should be 3:1)

[0736] BB X μl

[0737] I X μl

[0738] CutSmart buffer (10X) 5 μl

[0739] ATP (10 mM) 10 μl (2 mM final concentration)

[0740] H2O to 46 μl

[0741] T4 ligase 2 μl

[0742] SrfI (Restr.Enz.) 2 μl

[0743] Total 50 μl

[0744] Prepare the above reaction mixture on ice and in a PCR tube.

[0745] Vortex and spin.

[0746] Place in a thermal cycler and run the following program: (16 °C x 10’ >> 37 °C x 10’) x 8 >> 70 °C x 20’.

[0747] Add 1 μl of SrfI and run the following program: (This step is to digest any remaining BB - BB) 37 °C x 15’ >> 70 °C x 20’.

[0748] In a 50 mol / l reaction, it is recommended that the maximum amount of DNA used in the reaction (taking into account both I and BB) is 400 ng. The ratio of BB:I should not change.

[0749] Example ratio calculation: (len(X) = length of X, in base pairs), where: len(I) = 130 bp; len(BB) = 245 bp; len(BB) / len(I) = 245 / 130 = 1.88; starting from 50 ng of I, then 50 * 1.88 * 3 = 282 ng of BB is needed to achieve a 3:1 ratio.

[0750] Removal of linear DNA

[0751] Take out 4 μl of the circularization reaction as the negative control for the gel to be run later.

[0752] Add to the remaining cycling reaction (46 μl):

[0753] - ATP 10 mM 6 μl

[0754] - Plasmid-Safe buffer 10X 6 μl

[0755] - Plasmid-Safe enzyme 2 μl

[0756] Incubate at 37 °C for 30'.

[0757] Inactivate at 70 °C for 30'.

[0758] Run the entire reaction (S) and negative control (C-) on a 1.7% agar gel.

[0759] Agar purification of the band corresponding to circular BB-I ( Figure 25 )

[0760] Elute twice with 30 μl of water.

[0761] Rolling circle amplification

[0762] Add to the purified circular BB-I (about 50 μl at this time):

[0763] - Annealing buffer (5X) 12 μl

[0764] - Exo-Res.RND primer (500 μM) 1 μl

[0765] Heat the solution at 98 °C for 5', then cool slowly at room temperature.

[0766] Add:

[0767] - Phi29 buffer (10X) 10 μl

[0768] - BSA 2 μl

[0769] - dNTPs 10 μl

[0770] - 4 μl of pyrophosphatase

[0771] - 2 μl of Phi29 polymerase

[0772] - H2O to 100 μl

[0773] Incubate the reaction at 30 °C for at least 3 h.

[0774] Inactivate at 70 °C for 10’.

[0775] Run 5 μl on a 0.5% agar gel.

[0776] Running the RCA reaction overnight will produce more products. However, it is not clear whether the quality of the concatemers will be affected.

[0777] Quality check

[0778] The following procedure allows a rough estimate of the amounts of BB-I and BB-only monomers in the RCA product. Using the restriction enzyme cleavage sites present in the backbone (BglII in the following example), the RCA product can be digested, and the resulting band pattern can be used to infer the exact content of the RCA product.

[0779] As Figure 27 shown, BB200_4 (243 bp) and S1_WT (158 bp) were circularized and amplified by RCA. When digesting concatemers made of BB-I, we expect bands around 400 bp, while if the concatemers only have BB, the bands produced should be around 250 bp. Concateners formed only by I will not be digested, making the RCA bands visible.

[0780] Library preparation

[0781] DNA purification:

[0782] - Add an equal volume of Dyna beads, mix gently, and incubate for 5 minutes at room temperature.

[0783] - Insert the tube into the magnetic rack and wait for 5 minutes to allow the beads to aggregate on the wall.

[0784] - Remove the buffer.

[0785] - Rinse gently with 700 μl of 70% ethanol.

[0786] - Remove the ethanol and repeat the rinse step once more.

[0787] - Evaporate the residual ethanol.

[0788] - Remove the tube from the magnetic rack.

[0789] - Elute the DNA from the beads with 100 μl of ultrapure water.

[0790] Dissolved branched DNA:

[0791] - Add 4 μl of T7 endonuclease (NEB#M0302S).

[0792] - Incubate at 37 °C for 1 h.

[0793] Library preparation:

[0794] - Proceed with nanopore library preparation, which can be 1D ligation preparation or rapid preparation.

[0795] List of backbone and insert sequences

[0796] - len = length of the backbone in base pairs

[0797] - mean_flex = average of the calculated DNA flexibilities of all consecutive segments of 50 base pairs in the sequence

[0798] - max_flex = maximum DNA flexibility calculated for segments of 50 base pairs in the sequence

[0799] - Entropy = Shannon entropy of the DNA sequence

[0800] - GC% = percentage of GC bases in the backbone

[0801] >BB100_1(len:143 mean_flex:12.89 max_flex:14.71 Entropy:2.0 GC%:48.25)

[0802] GGGCATGCACAGATGTACACGATTCCCAACACACCGTGCGGGCCATCGACCTATG

[0803] CATACCGTACATATCATATATAAATCACATAATTTATTATACGTATGTCGCGCGGGTG

[0804] GCTGTGGGTAGATGCTGCATGACATAGCCC

[0805] >BB100_2(len:143 mean_flex:13.29 max_flex:14.95 Entropy:1.96 GC%:37.76)

[0806] GGGCATGCACAGATGTACACGCACTACATGCCAATGCCCAAGCAGTGCGCATATC

[0807] ACGTATCATATCTAATATATTATAATATTATGATAATGAGTATTTATTTAATTTGTTTG

[0808] TGTGAGGTAGATGCTGCATGACATAGCCC

[0809] >BB100_3(len:143 mean_flex:12.78 max_flex:14.1 entropy:1.95 GC%:44.06)

[0810] GGGCATGCACAGATGTACACGCATTGGCCGTCTGTGCTGTCCATGGATCGTCTGA

[0811] TTGATATGATATCATATATTATAATTATACAGTAAGGTGATTGGGTATTGAGGGTTGT

[0812] GTGGTTGGTAGATGCTGCATGACATAGCCC

[0813] >BB100_4(len:145 mean_flex:12.89 max_flex:14.06 entropy:1.95 GC%:44.14)

[0814] GGGCATGCACAGATGTACACGGTAGACATGCGAAGCGTGCGATGACAATCGATG

[0815] TGGACATCATGCATATATATGTTGTATAATTAAACAAATATGTGTAGTGTGTGAGGT

[0816] GGGTGTAGGAAGTAGATGCTGCATGACATAGCCC

[0817] >BB100_5(len:143 mean_flex:13.27 max_flex:14.34 entropy:1.9 GC%:37.76)

[0818] GGGCATGCACAGATGTACACGTTGTCATGGGAATTTGTGGTTATGAAATGAGTAT

[0819] GCGACGAATATGTATACATATATATTAAATTATAGAGTGATGTATGAGTTTGTGATG

[0820] TGTGGTGTATAGATGCTGCATGACATAGCCC

[0821] >BB200_1(len:243 mean_flex:13.0 max_flex:14.9 entropy:1.99 GC%:44.86)

[0822] GGGCATGCACAGATGTACACGGCGGCGCAAGATGATGTGCCGAACCTGACATGG

[0823] CATCGACTGGTATGGATCAATACTGATGCGATATCGATACCGGATAAATCATATATG

[0824] CATAATATCACATTATATTAATTATAATACATCGGCGTACATATACACGTACGCATCA

[0825] TTTCACTATCTATCGGTACTATACGTAGTGCCGGTCTGTTGGCCGGGCGACATAGA

[0826] TGCTGCATGACATAGCCC

[0827] >BB200_2(len:244 mean_flex:13.15 max_flex:14.69 entropy:1.96 GC%:38.52)

[0828] GGGCATGCACAGATGTACACGTGACGCAACGATGATGTTAGCTATTTGTTCAATG

[0829] ACAAATCTGGTATGATCAATACCGATGCGATATTGATATCTGATAACTCATATATGT

[0830] AGAATATCACATTATATTTATTATAATACATCGTCGAACATATACACAATGCATCTTA

[0831] TCTATACGTATCGGGATAGCGTTGGCATAGCACTGGATGGCATGACCCTCATTAGA

[0832] TGCTGCATGACATAGCCC

[0833] >BB200_3(len:244 mean_flex:13.06 max_flex:14.9 entropy:1.96 GC%:39.75)

[0834] GGGCATGCACAGATGTACACGAGACCGCAAGATGATGTTCATTCTTGAACATGAG

[0835] ATCGGATGGGTATGGATCAATACCGATGCGATATGATAACTGATAAATCATATATCT

[0836] ATAATATCACATTATATTAATTATAATACAGGATCGTTACATGCATACACAATGTATA

[0837] CTATACGTATTCGGTAGTTAGTGTACGGTCGGAATGGAGGTGGTGGCGGTGATAG

[0838] ATGCTGCATGACATAGCCC

[0839] >BB200_4(len:243 mean_flex:13.29 max_flex:14.44 entropy:1.93 GC%:34.57)

[0840] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGA

[0841] GATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGC

[0842] ATAATCTCACATTATATTTATTATAATAAATCATCGTAGATATACACAATGTGAATTG

[0843] TATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATG

[0844] CTGCATGACATAGCCC

[0845] >BB200_5(len:243 mean_flex:13.37 max_flex:14.52 entropy:1.94 GC%:35.8)

[0846] GGGCATGCACAGATGTACACGAATCCGTGAGATGACTATCTTATTTGTGACATTCA

[0847] TCGATCTGGATATGATCAATACCATGCGATATTGATTACTGATAAATCATATATGTAG

[0848] AATATCACATTATATTAATTATAATAAATCGTCGTACATATACATCCACAATTAGCTA

[0849] TGTATACTATCTATAGAGATGGTGCATCATCGTACTCCACCATTCCCACTAGATGCT

[0850] GCATGACATAGCCC

[0851] >BB300_1(len:348 mean_flex:13.12 max_flex:14.77 entropy:1.98 GC%:41.67)

[0852] GGGCATGCACAGATGTACACGCATAAGACCACAGGGTGCAAATCTGGATTGCGG

[0853] CATGGATGATTCATCATCGTGGCATATTCGCTATGGATATATCCATCATAATACATTG

[0854] ATACGTCATGCGTATAATCGCATTATATGTCGATATTGGTCATAGGGATACATCCGT

[0855] GTATACTATCGTATATGCGTGCAATGTAGCCATGTTAATCATGCTATAACCATAACA

[0856] TAAATATAATATATACAGATGGTGTATCTCTACTTATGTATGCTTGTATAGTAATGTC

[0857] GATACTGATGGGTCTCCGGCCCACTACACCACCTGGCCGCTCTAGATGCTGCATG

[0858] ACATAGCCC

[0859] >BB300_2(len:343 mean_flex:13.26 max_flex:14.34 entropy:1.98 GC%:40.82)

[0860] GGGCATGCACAGATGTACACGGGCAATCCGCCAGGGTTCAAATATGGATATGTGA

[0861] TGATCGATTCAACATGCACATATGCACGATATCATATATTACTCCAGATGTCATCAT

[0862] CGTCGTGCGTATATGAGATATGTATTTATGCATATAATCCACCATACATGGTAGCGA

[0863] TATTATAGTGCGATTATGTGTATATGACTATCATGGCTATTGTTAATATATAAATCATA

[0864] ACCATACCACTTCCACGCCTGGTATGGCGTATAGTATAGAGATATTGTGTGATGCC

[0865] CTATGTCGACCATGATGTGCCGTTGTACTGCCAATCCTAGATGCTGCATGACATAG

[0866] CCC

[0867] >BB300_3(len:344 mean_flex:13.47 max_flex:14.8 entropy:1.95 GC%:36.34)

[0868] GGGCATGCACAGATGTACACGTATCCATGCAGCTTATTGTAACTAGCGCATGCAC

[0869] GTGGTGATTCATCACATCTATATATACGATATGATATATTACACATATTTGCATAGTAT

[0870] CATCCGGTGTGATATCATCCGATATGCTCATACTTATTCATTGGTAGCATTGCATTG

[0871] ATGGATCAATAGTTATTATGACATCATGGCATGTACAATTATAAATAATACAACATA

[0872] CATAAATATACTATACACATCGTGTATGTGTTATACAGATCTGTGTGATGTATGATA

[0873] ATGTAATGGCGTCGAACACCACAAGGCAGTCCTATAATAGATGCTGCATGACATA

[0874] GCCC

[0875] >BB300_4(len:344 mean_flex:13.37 max_flex:14.57 entropy:1.94 GC%:37.5)

[0876] GGGCATGCACAGATGTACACGGTCCATTACAATCGAATCTATATCCCAATGTGTAT

[0877] CGATTATCACCACAATGACATAATACGATATCATATATTACTCCATATGCCTTACGTC

[0878] AGATCGTTATATGAGATATGTATTCATGCATATGATATCCCACAGTACACGTCGTCT

[0879] AATGCCATCATGAATGTATGACATATCTAGTCGATTATACATAATATAACATACCAAT

[0880] ATAACAATATCTATACACATTTGATGGCGTATAGTATAAAGATATTGTGGCAATGCC

[0881] CATACACCACTGACTGTCGCCGATCATTCCTACCACTAGATGCTGCATGACATAGC

[0882] CC

[0883] >BB300_5(len:344 mean_flex:13.51 max_flex:14.89 entropy:1.91 GC%:33.43)

[0884] GGGCATGCACAGATGTACACGACCGACCGTGAAAGTGATTCAGAATGATGTGCA

[0885] TGAATGTTATCATGACATGATTTATGATGCACTGATATATGCATATTATAATATTGTA

[0886] CAATGTCGTATATACGACATATCTATACTATGAATTATGGCATCATGGACAATAGAT

[0887] GGTAAGGTATAGTACGATCTATATAGCATGTTGAAATGGGATATAAATTATCATAAA

[0888] CATACATACTTAACTAATATCAAGATGATATGTGTATGACATCAGAATGATAGTAGT

[0889] AATGAGTATTGTCAGATGTATGTACGAATATCACACGATTAGATGCTGCATGACAT

[0890] AGCCC

[0891] >Insert S1 WT (TP53, chr17:7577450-7577649)

[0892] AGGCTGGGGCACAGCAGGCCAGTGTGCAGGGTGGCAAGTGGCTCCTGACCTGG

[0893] AGTCTTCCAGTGTGATGATGGTGAGGATGGGCCTCCGGTTCATGCCGCCCATGCA

[0894] GGAACTGTTACACATGTAGTTGTAGTGGATGGTGGTACAGTCAGAGCCAACCTAG

[0895] GAGATAACACAGGCCCAAGATGAGGCCAGTGCGCCTT

[0896] >Insert 17.2 (TP53, chr17:7578161-7578394)

[0897] CAGTTGCAAACCAGACCTCAGGCGGCTCATAGGGCACCACCACACTATGTCGAA

[0898] AAGTGTTTCTGTCATCCAAATACTCCACACGCAAATTTCCTTCCACTCGGATAAG

[0899] ATGCTGAGGAGGGGCCAGACCTAAGAGCAATCAGTGAGGAATCAGAGGCCTGG

[0900] GGACCCTGGGCAACCAGCCCTGTCGTCTCTCCAGCCCCAGCTGCTCACCATCGC

[0901] TATCTGAGCAGCGCTCAT

[0902] Related to Figure 24 Bioinformatics

[0903] Using Tombo's DNA model (Fasta->raw), both forward and reverse (https: / / github.com / nanoporetech / tombo), will generate the expected reference signal for every possible insertion (for every possible base pair at the target position). Also create forward and reverse expected signals for the backbone.

[0904] Use dynamic time warping (DTW) to map the expected backbone signal to the read. If the expected backbone signal overlaps when aligned with the read, pick the best result and discard the sub-optimal results. Then fragment the read based on the orientation of the appropriate backbone.

[0905] Subsequently, use DTW to map all possible expected insertion signals to the read. Again, overlapping results are discarded and only the best results are retained. The optimal matching result (minimum DTW error) is retained for each read. The number of specific insertions (representing the specific base at the target position) determines the most likely base at the target position in the read.

[0906] Results

[0907] Circularization efficiency of different backbones

[0908] To be able to experimentally evaluate the efficiency of circularizing short DNA amplicons with different backbones, a 234bp PCR amplicon (insert 17.2) was ligated to backbones from three different backbone series (BB100_1 / 2 / 3 / 4 / 5, BB200_2 / 4 / 5, and BB300).

[0909] The main chain sequences and physical properties are reported below. The detailed protocols are disclosed in Materials and Methods.

[0910] Figure 21 (Left) shows the product of the cyclization reaction.

[0911] After cyclization, the reaction was supplemented with an enzyme mixture (PlasmidSafeLucigen #E3101K) to digest linear DNA. Figure 21 (Right) shows the residual product (circular DNA).

[0912] The BB200 series showed the best efficiency to date. To further characterize the efficiency of BB200_2 / 4 / 5, these 3 main chains were ligated with the same amplicons in the absence of the restriction enzyme SrfI. The rationale behind this experiment is that the ligation efficiency of the main chains can be estimated by the amount of multimers formed in the reaction. As Figure 22 can be seen, the ligation efficiency of BB200_4 was significantly higher than that of BB200_2 and BB200_5.

[0913] The higher ligation efficiency of BB200_4 was reflected in higher cyclization and RCA product formation efficiency. In Figure 23 , the sequencing read counts from two independent experiments (blue and red) were plotted, in which an equimolar mixture of BB200_2, BB200_4, and BB200_5 was used to generate concatemers. The sequencing results were consistent with the previous experimental results, indicating that the vast majority of the sequenced beads contained BB200_4.

[0914] New (optimized) barcode sequences with better ligation efficiency

[0915] BB200_4 is the most efficient main chain tested to date in the cyclization reaction.

[0916] Coupling with the possibility of strand-specific rolling circle amplification Strand-specific mutation calling

[0917] The Cyclomics method generates double-stranded DNA circles. One advantage of circles with double strands is that one of the strands can be preferably used as a template for RCA, for example, by using strand-specific primers to initiate the reaction and then following a known procedure (https: / / www.sciencedirect.com / science / article / pii / S0042682212002814). In this way, the Cyclomics method enables the selective amplification of the sense or antisense sequences of a given DNA sequence. Such strand-specific amplification is not possible using the smartbell method, but it is highly beneficial for obtaining accurate variant calls from nanopore sequencing data in an efficient manner.

[0918] In Figure 24 this paper, we presented an example case where the detection rate of the correct bases was different when analyzing data from two different strands of a DNA molecule. The data was from an experiment where amplicons 200bp (insert S1WT) long were circularized with BB200_4 and amplified as per the protocol reported.

[0919] Data analysis of the sequencing results can determine the accuracy of base calling for each strand. Specifically, we noted that it is often difficult to distinguish between C and A bases due to their similar raw signal intensities. However, the signal from T is very different from the signals from all other bases and is easily correctly classified. For example, if an A mutation in the forward strand is expected, sequencing of the reverse strand will result in a clearer result because an A in the forward strand will be mis-called as a G. Thus, specific enrichment of the reverse strand is advantageous in this scenario.

[0920] Figure 24 The highlighted data in this paper shows an example of the difference in identifying bases on the forward or reverse strand. It emphasizes how the correct bases on the reverse strand data can be inferred by using a simple cut-off point (Y < 0.3) on the Y-axis. The same method does not work for the forward strand. Thus, in this case, amplifying and sequencing both strands would result in a waste of data and, more problematically, would lead to misleading mutation detection at specific positions with a high false positive rate. Instead, strand-specific enrichment would result in higher sensitivity (most reads will come from the best strand) and no false positive calls.

Claims

1. A method for preparing a double-stranded target DNA molecule for sequencing, comprising: - providing a double-stranded backbone DNA molecule having a 5'-end and a 3'-end, wherein the 5'-end and the 3'-end: - are ligation-compatible with the 5'-end and the 3'-end of the target DNA; - form a first restriction enzyme recognition site when self-ligated; - are in a form capable of self-ligation; and - providing, if not already present, 5'-ends and 3'-ends for the target DNA that are in a form that prevents self-ligation and are ligation-compatible with the 5'-end and the 3'-end of the backbone DNA; The method further comprises: - ligating the target DNA to the backbone DNA in the presence of a ligase and a first restriction enzyme that cleaves the first restriction enzyme recognition site, thereby generating at least one DNA circle comprising the backbone DNA molecule and the target DNA molecule; - optionally removing linear DNA; - generating concatemeric DNA molecules comprising an ordered array of copies of the at least one DNA circle by rolling circle amplification; and - sequencing the at least one concatemer, wherein, - the target DNA is 20 to 300 base pairs; - the form allowing self-ligation is a 5'-phosphate group at one DNA end and a 3'-hydroxyl group at the other DNA end, and the form preventing self-ligation is a 5'-hydroxyl group at one DNA end and a 3'-hydroxyl group at the other DNA end; - the ligation of the ends of the target DNA to the ends of the backbone produces a target-backbone linker having a sequence that is not recognized / cut by the restriction enzyme that cuts the restriction enzyme site formed by self-ligation of the backbone; - the backbone comprises a linker, and the linker comprises a sequence of 20 to 900 nucleotides; - the backbone has a length of 20 to 1000 nucleotides and has a flexibility score of 10 or higher.

2. A method for evaluating the target DNA capture efficiency of a backbone, comprising the steps of claim 1 and further comprising comparing the target DNA capture efficiencies between different backbones.

3. A collection of linear DNA molecules (backbones) having a length of 20 to 1000 nucleotides, the DNA molecules comprising a 5'-end and a 3'-end, the 5'-end comprising a part of a first restriction enzyme recognition site at the outermost end, and the 3'-end comprising another part of the first restriction enzyme recognition site at the outermost end, and the 5'-end and the 3'-end being ligation-compatible with each other and forming a restriction enzyme (first restriction enzyme) recognition site when self-ligated, and wherein each backbone comprises: a linker; an identifier sequence (barcode), the identifier sequence being different from the sequences of the identifiers of other backbones in the collection; and optionally a restriction site for a nicking enzyme, wherein, - the collection of linear DNA molecules (backbones) has a flexibility score of 10 or higher, and; - the linker comprises a sequence of 20 to 900 nucleotides.

4. A method for determining the sequences of a collection of nucleic acid molecules, comprising: - providing a double-stranded target DNA molecule having recombinase recognition sites specific for a target site-specific recombinase at the 5'-end and the 3'-end; - Provide a backbone comprising the recognition site, the recognition site being separated by DNA comprising a linker; - Incubate the target DNA molecule with the backbone in the presence of the target site-specific recombinase, which is preferably Cre recombinase, FLP recombinase or phage λ integrase, thereby generating a DNA loop comprising the backbone and the target DNA molecule; - Optionally remove the linear DNA; - Generate a concatemer comprising an ordered array of copies of at least two of the DNA loops by rolling circle amplification; and - Sequence the concatemer, wherein, - The target DNA is 20 to 400 base pairs; - The backbone has a length of 20 to 1000 nucleotides and has a flexibility score of 10 or higher, and - The linker comprises a sequence of 20 to 900 nucleotides.

5. A kit comprising the set of linear DNA molecules according to claim 3.

Citation Information

Patent Citations

  • Window

    EP0910722A1

  • Wellbores utilizing fiber optic-based sensors and operating devices

    EP0910725A1