Method for preparing nucleic acid molecules for sequencing
By connecting the target DNA to the double-stranded main chain DNA to form a DNA loop and generating multi-conjunctive DNA molecules through rolling loop amplification, the problems of high error rates and difficulty in analyzing long repeat elements in existing sequencing technologies are solved, and high accuracy and long-length sequencing are achieved.
Patent Information
- Application Number
- CN201880089087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-11
- Filing Date
- 2018-12-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2038-12-11
AI Technical Summary
Existing sequencing technologies have problems with high error rates and difficulty in analyzing long repeat elements and structural variations in the genome.
By providing double-stranded backbone DNA molecules with 5' and 3' ends, the target DNA is linked to the backbone DNA to form a DNA loop and multi-conjunctive DNA molecules are sequenced by rolling loop amplification.
It improves the accuracy and length of sequencing, can effectively analyze long repeat elements and structural variations in the genome, and reduces the error rate.
Smart Images

Figure BDA0002627214160000161 
Figure BDA0002627214160000162 
Figure BDA0002627214160000163
Abstract
Description
Technical Field
[0001] The present invention relates to means and methods for determining the sequence of nucleic acid molecules. In particular, the present invention relates to methods utilizing rolling circle amplification of nucleic acid molecules whose sequences are to be determined. Background Art
[0002] Sequencing methods have evolved over time. The old Sanger sequencing method has been replaced by the currently commonly used next-generation sequencing (NGS) methods. Recently, these methods have been reviewed in the literature (Goodwin et al. 2016; Nature Reviews | Genetics Volume 17: pp 333-351: doi: 10.1038 / nrg.2016.49). The most commonly used NGS methods rely on the sequencing of short stretches of DNA. Sequencing techniques for short stretches of DNA have inherent errors. Errors can be reduced by independently sequencing multiple copies of the same target sequence. However, for each single sequence read, it is impossible to determine whether the change represents an error or a true mutation. The accumulated evidence across several independent sequence reads allows filtering mutations introduced during amplification and errors in sequencing longer target DNA (which can also be sequenced using short read methods). This is usually done by sequencing overlapping fragments that can be aligned to produce assembled longer sequences. This so-called short-read paired-end technology has been very successful in sequencing large target nucleic acids and is very useful for various genome projects. Genome projects have revealed that genomes are highly complex, with many long repetitive elements, copy number changes, and structural variations. Many of these elements are so long that short-read paired-end technologies are insufficient to resolve them. Long-read sequencing provides read lengths exceeding several kilobases and allows the resolution of these large structural features throughout the genome. Two popular platforms for long-read sequencing are Pacific Biosciences systems (RSII and Sequel) and Oxford Nanopore systems (MK1 MinION and PromethION). Both are single-molecule sequencers. Both platforms allow read lengths exceeding 55 kb or even longer. However, these systems have even higher error rates than next-generation (second-generation) sequencers. These errors can be reduced by increasing the number of sequencing times for the same target nucleic acid (Goodwin et al 2016; doi: 10.1038 / nrg.2016.49).
[0003] The present invention provides new protocols for preparing nucleic acid molecules for sequencing. Summary of the invention
[0004] An embodiment of the present invention provides a method for preparing a double-stranded target DNA molecule for sequencing, comprising:
[0005] - providing a double-stranded backbone DNA molecule comprising a 5' end and a 3' end, wherein the double-stranded backbone DNA molecule:
[0006] - is ligation-compatible with the 5' and 3' ends of the target DNA;
[0007] - forms a first restriction enzyme recognition site upon self-ligation;
[0008] - in a form capable of self-connection; and
[0009] - providing the target DNA with 5' and 3' ends, if not already present, in a form that prevents self-ligation and is ligation-compatible with the 5' and 3' ends of the backbone DNA;
[0010] The method further comprises:
[0011] - In the presence of a ligase and a first restriction enzyme that cuts the first restriction enzyme recognition site, ligating the target DNA to the backbone DNA, thereby generating a matrix comprising a backbone DNA molecule and a target DNA molecule
[0012] at least one DNA loop;
[0013] -optionally remove linear DNA;
[0014] - generating concatemeric DNA molecules comprising an ordered array of copies of the at least one DNA circle by rolling circle amplification; and
[0015] - sequencing said at least one concatemer.
[0016] Also provided is a collection of DNA molecules (backbones) with a length of 50 to 1000 nucleotides, the DNA molecules comprising a 5' end and a 3' end, the 5' end comprising a portion of a first restriction enzyme recognition site at the most terminal end, the 3' end comprising another portion of the first restriction enzyme recognition site at the most terminal end, and the 5' end and the 3' end are compatible with each other for ligation and can form a restriction enzyme (first restriction enzyme) recognition site when self-ligated, and wherein each of the backbones comprises:
[0017] Connector;
[0018] an optional identifier sequence (barcode) that is different from the sequence of the identifiers of the other backbones in the set;
[0019] a second identifier, optionally unique to the set of backbone molecules; and
[0020] Optionally a restriction site for a nicking enzyme.
[0021] Further provided is a method for determining the sequence of a collection of nucleic acid molecules comprising:
[0022] - providing a double-stranded target DNA molecule having a 5' end and a 3' end, wherein there are overhanging adenine residues at the 3' end of both strands of the DNA molecule;
[0023] - providing a collection of double-stranded backbone DNA molecules, the double-stranded backbone DNA molecules comprising 5' and 3' ends that are ligation-compatible with the 5' and 3' ends of the target DNA;
[0024] The method further comprises:
[0025] - ligating the target DNA to the backbone in the presence of a ligase, thereby generating a DNA circle comprising the backbone and the target DNA molecule;
[0026] -optionally remove linear DNA;
[0027] - generating concatemers of ordered arrays comprising at least two copies of said DNA circles by rolling circle amplification; and
[0028] - sequencing the concatemers.
[0029] Further provided is a method for determining the sequence of a collection of nucleic acid molecules comprising:
[0030] - providing a double-stranded target DNA molecule, wherein the double-stranded target DNA molecule has a recombinase recognition site specific for a target site-specific recombinase at the 5' end and the 3' end;
[0031] - providing a backbone comprising said recognition site separated by a DNA comprising a linker;
[0032] - incubating the target DNA molecule with the backbone in the presence of the target site-specific recombinase, preferably Cre recombinase, FLP recombinase or bacteriophage lambda integrase, thereby generating a DNA circle including the backbone and the target DNA molecule;
[0033] -optionally removing linear DNA; and
[0034] - generating concatemers of ordered arrays comprising at least two copies of said DNA circles by rolling circle amplification; and
[0035] - sequencing the concatemers.
[0036] In a preferred embodiment, the backbone is a loop comprising two recombinase recognition sites, separated on one side by a DNA comprising a connector and on the other side by a DNA encoding a restriction enzyme recognition site, and wherein the restriction enzyme site is the only recognition site in the backbone for the restriction enzyme. In this embodiment, the method preferably further comprises digesting the DNA with the restriction enzyme after the recombination, and subsequently removing the linear DNA, followed by producing the concatemer.
[0037] Further provided is a method for determining the sequence of a collection of nucleic acid molecules comprising:
[0038] - providing a double-stranded target DNA molecule having a recombinase recognition site for a target site-specific recombinase at the 5' end and the 3' end;
[0039] - providing a collection of double-stranded circular backbone DNA molecules comprising the recombinase recognition site and a linker;
[0040] The method further comprises:
[0041] - incubating the target DNA molecule with the backbone in the presence of a target site-specific recombinase for the recognition site, thereby generating a DNA circle comprising the backbone and the target DNA molecule;
[0042] -optionally removing linear DNA; and
[0043] - generating concatemers of ordered arrays comprising at least two copies of said DNA circles by rolling circle amplification; and
[0044] - sequencing the concatemers.
[0045] Further provided are kits comprising one or more backbones. DETAILED DESCRIPTION
[0046] Means and methods as described herein can repeatedly determine the sequence of the same target DNA molecule. This can be used as a means of correcting errors. This is different from the classical second generation sequencing method of correcting errors by sequencing a plurality of independent molecules covering the same genomic locus. In this case, each reading generally represents a sequencing event of a molecule. With the method of the present invention, a single (target) molecule is repeatedly replicated, so a plurality of sequencing events representing the same molecule are read once.
[0047] The target nucleic acid is usually double-stranded DNA. Single-stranded DNA or RNA that needs to determine the sequence can be easily converted into double-stranded DNA by methods known in the art. Such methods include but are not limited to cDNA synthesis, reverse transcriptase (RT) polymerase chain reaction (PCR), PCR, random primer extension, etc. Before performing the method, the target DNA is linear or is made into linear.
[0048] The main chain is usually double-stranded DNA. In the method of using restriction enzyme to connect the target DNA to the main chain, the main chain is usually linear or made linear before or during the method. In the method of using target site specific recombinase to insert the target DNA into the main chain, the main chain can be linear or made circular before or during the method.
[0049] Self-ligation is defined herein as the ligation of the 5' end of one and the same nucleic acid molecule to the 3' end.
[0050] The 5' and 3' ends of the target DNA are selected so that the 5' and 3' ends of the target DNA are compatible with the 5' and 3' ends of the main chain used in the reaction. Compatible connection refers to the connection of the ends to each other to produce a double-stranded DNA with correctly paired nucleotides without a gap at the connection joint. Of course, a gap can be introduced later to allow the RCA reaction to be initiated. The flat end is compatible with other flat ends. If the protruding DNA chains can be annealed together without unpaired bases, the DNA with sticky (also referred to as "sticking") ends is compatible with other sticky ends. This is usually the case when the ends have complementary sequences. "Connection-compatible ends" are also referred to as "compatible ends" or "compatible sticking ends" or "compatible sticky ends" in the art.
[0051] Double-stranded target DNA molecules include sequences of nucleic acid molecules to be determined. Nucleic acid molecules to be determined may already be double-stranded DNA with 5' and 3' ends that are compatible with the main chain to be used. Sometimes nucleic acid needs to be made into double-stranded DNA, for example, in the case of cDNA or mRNA. Target DNA may already have suitable 5' and 3' ends, for example, various polymerases produce blunt-ended fragments. This blunt-ended fragment is compatible with the main chain with blunt 5' and 3' ends. Target nucleic acid may also provide suitable 5' and 3' ends, for example, by digesting with suitable one or more restriction enzymes, or by adding deoxynucleotides by terminal transferase. Suitable 5' and 3' ends may also be introduced by the insertion of restriction enzyme sites, recombinase recognition sites and / or homology. For example, by connecting the adaptor containing the site to the target DNA, or by amplifying the target DNA with a primer containing restriction enzyme sites, recombinase recognition sites and / or homology.
[0052] The enzyme that produces the end that is compatible with the end of main chain for connection, but is different from the nucleotide in the region of protruding end immediately adjacent is available.In this embodiment, preferably, the recognition site of enzyme is different from the restriction enzyme site of described first restriction enzyme.Like this, the connection of compatible end does not produce the site that can be cut by described first restriction enzyme.If restriction enzyme is used to provide target nucleic acid with suitable end, preferably, enzyme is the enzyme that produces blunt end.In one embodiment, by digestion with one or more restriction enzymes, provide the target DNA molecule with 5' end and 3' end that are compatible with the 5' end and 3' end of main chain to be used for connection.
[0053] In one embodiment, ligating the end of the target DNA to the end of the backbone generates a target-backbone linker having a sequence recognized / cleaved by a restriction enzyme that is not cleaved by the (first) restriction enzyme site formed by self-ligation of the backbone.
[0054] In a preferred embodiment, the form that prevents self-connection is 5'-hydroxyl of one DNA end and 3'-hydroxyl of another DNA end, and the form that allows self-connection is 5'-phosphate group of one DNA end and 3'-hydroxyl of another DNA end.Connection needs the existence of 5'-phosphate group.Remove by suitable phosphatase on the 5' end of nucleic acid molecule, can prevent self-connection and be connected to other DNA molecules that are similarly processed.Even if the end has a connection compatible end, connection is also prevented.
[0055] In one embodiment, the backbone includes a recognition site for a nicking enzyme.
[0056] The target DNA molecule has a 5' end and a 3' end that prevent self-ligation. Preferably, the target DNA is in a form that prevents ligation to other target DNA molecules. These two requirements can be met by providing a dephosphorylated end or adding nucleotides (3' overhang) to the 3' end of the target DNA molecule.
[0057] When the 5' and 3' ends of the target DNA are incompatible for connection, self-connection is inherently prevented. However, in these cases, preferably, connection to other target molecules is also prevented. Therefore, in this environment, preferably, the ends are also provided in a dephosphorylated form. Incompatible ends are, for example, but not limited to, blunt ends and protruding ends or protruding ends in which the protruding nucleotides (protruding ends) are incompatible.
[0058] Preventing self-ligation and / or preventing ligation to other target DNA molecules is not necessarily absolute. The process can / will occur to some extent. This is permissible in the method of the present invention. Even if the ligation efficiency is low, good reading can be obtained.
[0059] The 5' end and 3' end of the backbone DNA may be compatible for connection with each other. In this embodiment, preferably, the 5' end and 3' end of the target DNA are also compatible for connection with each other. Preferably, self-connection of the ends of the backbone is not prevented. Preferably, the 5' end of the target DNA is dephosphorylated. Preferably, connection is performed in the presence of a restriction enzyme that recognizes and cuts the first restriction enzyme site.
[0060] In an embodiment of capturing double-stranded target DNA, the backbone is a double-stranded nucleic acid molecule. Such backbone includes a 5' end and a 3' end that are ligation-compatible with the 5' end and the 3' end of the target DNA. The 5' end and the 3' end of the backbone may also be ligation-compatible with each other.
[0061] In an embodiment, the backbone comprises one or more of the following moieties:
[0062] - encoding the first part of the first restriction enzyme cleavage site, preferably the 5' end of the first half of the first restriction enzyme cleavage site (eg, see 1 in the schematic example below),
[0063] - one or more sites that allow nicking of the double-stranded backbone sequence (e.g., see 2 below),
[0064] - one or more type 1 or type 2 restriction enzyme sites (e.g., see 3 below),
[0065] - a secondary cloning site (see, for example, 4 below),
[0066] - Flexible DNA extension that allows efficient circularization (bending) of the backbone molecule, 5) below
[0067] - A unique molecular barcode (identifier) sequence used to label each individual backbone molecule (e.g., see 6 below)
[0068] - encoding another part, preferably the 3' end of the other half of the first mentioned restriction site.
[0069] - Phosphorylation at the 5' end of the backbone molecule and a hydroxyl group at the 3' end of the backbone molecule.
[0070] - Secondary barcode sequences that can be used to identify individual samples.
[0071] Schematic example of a double-stranded backbone sequence: (1)(2)(3)(4)(5)(6)(1)
[0073] 5'-GGGC..CCTCAGC..ATTTAAAT..GTCTTCGAGAAGAC..CATACTATCATG..(N)..GCCC-3'
[0074] 3'-CCCG..GGAGTCG..TAAATTTA..CAGAAGCTCTTCTG..GTATGATAGTAC..(N)..CGGG-5'
[0075] Dots represent 0 nucleotides; 1 nucleotide; 2 nucleotides or more.
[0076] The sequences GGGC and GCCC represent half of a restriction enzyme site. The sequence constitutes a SrfI site, but another restriction enzyme site will also work. In the case of SrfI (GCCC | GGGC), the advantage is that it is a blunt-ended site. Another advantage is that it identifies a site of 8 bases long, while most commercially available substitutes identify sites of 6 bases long.
[0077] Preferably, the first restriction enzyme site does not appear elsewhere in the backbone sequence.
[0078] Ligation of ligation-compatible ends can generate a restriction enzyme site. This occurs if the end and the flanking sequence (if any) encode a restriction enzyme site when ligated to each other. As an example, an end of a double-stranded DNA molecule having a single strand of the sequence 5'-AATT... is ligation-compatible with a double-stranded DNA molecule having a single strand of the sequence ...TTAA-5', where the dots indicate the double-stranded portion and 5' or 3' indicates the free end of the respective molecule. Ligation of the two ends results in a molecule having the following double-stranded sequence:
[0079] …AATT…
[0080] …TTAA…
[0081] The overhangs are identical to those produced by the EcoRI restriction enzyme. Ligation only produced the restriction site EcoRI in some cases, i.e., where the bold nucleotides have the indicated bases:
[0082] …GAATTC…
[0083] …CTTAAG…
[0084] EcoRI cannot cut when the bold nucleotides have different bases. For example, the following sequence is not cut by EcoRI:
[0085] …CAATTC…; or …AAATTC…; or …GAATTA…
[0086] …GTTAAG…;…TTTAAG…;…CTTAAT…
[0087] The sequence of the ends of the target DNA thus determines whether the ligation adapter formed by ligation of compatible ends can be digested by the enzyme that cuts the first restriction enzyme site.
[0088] In embodiments, the insertion capture efficiency of the backbone can be optimized, wherein higher efficiency is reflected in more efficient circularization and rolling circle amplification (RCA) product formation. The insertion capture efficiency of the backbone can be estimated by the number of multimers that can be formed.
[0089] In the method for sequencing target DNA described herein, preferably, the target DNA is connected to the main chain DNA without generating the first restriction enzyme recognition site at the target / main chain DNA joint. In the present invention, preferably, the self-connection of the main chain generates the first restriction enzyme site, and the main chain is connected to the target DNA without generating the site. The preferred first restriction enzyme site is the enzyme that allows the most sequence variation in the joint joint. Since the sequence of the main chain has a part of the recognition sequence of the first restriction enzyme site, and preferably half, the variation comes from the sequence of the target end. In the case where the first restriction enzyme site is an EcoRI site, the main chain sequence encoding the first restriction enzyme site has a 5' end with a sequence of 5'-AATTC. Depending on the base of the nucleotides flanking the overhang in the target DNA, the joint with the target DNA can have one of four different sequences. Only when the target sequence has an end with a sequence of 5'-AATTC.., the joint with the main chain is EcoRI-digestible. The joint with other sequences is not EcoRI-digestible. By selecting an enzyme that generates little or no overhang and by selecting an enzyme that requires more specific bases at the recognition site, the variation in the joint is improved. The first restriction enzyme site preferably includes 6 and more preferably 8 and more preferably bases. Therefore, the enzyme that cuts the first restriction enzyme site is preferably at least 6 cutters, more preferably at least 7 cutters, and more preferably 8 cutters. The number indicates the number of bases in the recognition site of the enzyme. For example, EcoRI is a 6-cutter; AluI that recognizes AGCT is a 4-cutter. There are also 5-cutters (e.g., AvaII), 7-cutters (e.g., BbvCI), 8-cutters (e.g., NotI) and even other restriction enzymes. In addition, preferably less or no overhangs, this ensures a high probability of sequence variation in the connecting joint, and this reduces the possibility that the joint of the target sequence and the main chain sequence is the first restriction enzyme site. The first restriction enzyme with more nucleotides in the recognition site is also preferred because this enzyme allows larger target nucleic acid insertion. The method is suitable for a variety of target nucleic acid sources. The method of the present invention can be performed with two or more main chains with different first restriction enzyme sites. In this way, more target molecules can be captured in the DNA ring. In the case where the target DNA has two very close first restriction enzyme sites, the intermediate sequence can be effectively sequenced, for example, by capturing the intermediate sequence with a backbone having other first restriction enzyme sites. Mentioning "first" in the context of restriction enzyme sites refers to the position of the site on the backbone (half of the site). The restriction enzyme recognition sites at other positions in the backbone will be referred to as second, third restriction enzyme recognition sites, etc.
[0090] Preferred first restriction enzyme recognition sites are sites for restriction enzymes SrfI (GGGC|GCCC), PmeI (GTTT|AAAC) and SweI (ATTT|AAAT). Particularly preferred first restriction enzyme sites are sites for restriction enzyme SrfI.
[0091] The 5' end of the main chain is included in the most terminal part of the first restriction enzyme recognition site. It can but does not necessarily contain other nucleotides inside. The number of nucleotides at the end can be changed. The 5' end usually has between 2 and 15, preferably between 2 and 10, preferably between 2 and 8 nucleotides, more preferably 2, 3, 4, 5, 6, 7 or 8 nucleotides. In some embodiments, the 5' end is 3 or 4 nucleotides.
[0092] The 3' end of the main chain is included in the most terminal part of the first restriction enzyme recognition site. It may but not necessarily contain other nucleotides inside. The number of nucleotides at the end can be changed. The 3' end usually has between 2 and 15, preferably between 2 and 10, preferably between 2 and 8 nucleotides, more preferably 2, 3, 4, 5, 6, 7 or 8 nucleotides. In some embodiments, the 3' end is 3 or 4 nucleotides.
[0093] The 5' and 3' ends of the target DNA are preferably flat ends. If self-ligation is not prevented, the 5' and 3' ends may also be sticky ends that can be connected together. The 5' and 3' ends of the target DNA are preferably provided in a dephosphorylated form to prevent self-ligation. The 5' and 3' ends of the target DNA may also be sticky ends that cannot be connected together, such as adenine overhangs added by terminal transferase.
[0094] Preferably, connection is carried out in the presence of a ligase and a restriction enzyme (first restriction enzyme) that cuts the first restriction enzyme site. The end of the main chain is connected to the end of the target DNA to produce a double-stranded DNA ring. In the method of the present invention, the self-connection of the main chain is often not prevented. In the presence of a ligase, the ends of the two ends of the main chain that are connected to each other or connected to other main chains can hinder the capture of target nucleic acids by the main chain. The connection of the main chain end can be resisted by the presence of the first restriction enzyme. Since this connection usually produces (reproduces) the first restriction enzyme site, the main chain is linearized and / or de-multiplexed. Ligation reaction is carried out under buffer conditions that support effective connection and effective cutting using the first restriction enzyme at the same time. The method of the present invention is particularly suitable for producing a DNA ring with a main chain and a target nucleic acid.
[0095] In embodiments of the invention, linear DNA, if present, is preferably removed prior to rolling circle amplification. Performing rolling circle amplification after removal of linear DNA generally produces higher molecular weight concatemers of the backbone and target DNA.
[0096] The method comprises: subjecting the DNA ring produced in the ligation reaction to rolling circle amplification (RCA). The rolling circle amplification produces an ordered array of at least two copies of the DNA ring. The rolling circle amplification produces a high molecular weight DNA molecule. The high molecular weight is suitable for sequencing, especially for long read sequencing.
[0097] Rolling circle amplification has been reviewed in the literature (Mohsen and Kool (2016) Acc Chem Res. Vol 49 (11): pp2540-2550; published online on Oct 24, 2016. doi: 10.1021 / acs.accounts.6b00417). The terms rolling circle amplification and rolling circle replication are sometimes used interchangeably in the art. In other cases, rolling circle amplification is used to refer to the replication of naturally occurring plasmids and viral genomes. These terms refer to similar basic principles, that is, the same circular DNA that produces longer nucleic acid molecules is repeatedly copied with an ordered array of backbone-target nucleic acid copies. Existing techniques for rolling circle amplification allow the production of large arrays containing many copies of the generated DNA ring. Multiple concatemers can have 2 or more copies, preferably 4 or more copies of the generated ring.
[0098] Rolling circle amplification is performed by polymerase, and a commonly used introduction sequence is required to generate a promoter. Specific polymerases with high processivity can be used to produce relatively long concatemers. Polymerases with high processivity are polymerases that can polymerize one thousand or more nucleotides without separating from the DNA template. They can preferably polymerize two thousand, three thousand, four thousand or more nucleotides without separating from the DNA template. Polymerases with high processivity are discussed in the literature (Kelman et al.; 1998: Structure Vol 6; pp 121-125). Rolling circle amplification uses polymerases with high processivity and strand displacement ability, such as phi29 polymerase, to produce very high molecular weight concatemers. The polymerase can polymerize 10 kb or more. Therefore, the preferred polymerase with high processivity is a polymerase that polymerizes 10 kb or more without separating from the DNA template (Blanco et al.; 1999. J. Biol. Chem. 264 (15): 8935-40). In the presence of one or more suitable primers, polymerization can start on the nick of double-stranded DNA or DNA can be unzipped and annealed. Examples of suitable primers are hexamer random primers, one or more main chain specific primers, one or more target nucleic acid specific primers or their combinations. When the target nucleic acid sequence is unknown or when a variety of target nucleic acid sequences are to be sequenced, random primers are usually preferred. One or more specific primers can be used for the sequencing of specific target nucleic acids of known basic sequences. In one embodiment, one or more primers are peculiar to the main chain. Such primers can be used in different scenarios, such as, but not limited to, high-throughput systems with optimized main chains.
[0099] The advantage of having double-stranded circular DNA is that one of the chains can be used as a template for rolling circle amplification. For example, by using chain-specific primers to start the RCA reaction. Data analysis of Oxford nanopore sequencing results allows the accuracy of base-calling and variation calling to be determined for each chain respectively. Specifically, we found that due to the similarity of the original current intensity of C bases and A bases, it is often difficult to distinguish them. However, the current signal from T is substantially different from all other bases and is easily correctly classified. For example, if A is expected to mutate on the positive chain, the sequencing of the reverse chain will obtain clearer results because the A in the positive chain may be misjudged as G. Therefore, in this scenario, the specific enrichment of the reverse chain will be advantageous. Therefore, in a preferred embodiment, the rolling circle initiator primer is a chain selective primer.
[0100] Further optimization for obtaining strand-specific sequences may involve using (or additionally using) real-time selective sequencing methods, such as those described in the prior art (Loose et al. 2016. Nature methods. Real-time selective sequencing using nanopore technology).
[0101] The backbone is preferably 20 to 1000 nucleotides long, preferably 20 to 800 nucleotides long, preferably 50 to 800 nucleotides long, more preferably 100 to 600 nucleotides long, preferably 200 to 600 nucleotides long. Depending on the application, the target nucleic acid is preferably 40 to 15000 nucleotides long.
[0102] DNA that circulates freely or is bound to cell particles in blood or other body fluid samples is usually less than 400 nucleotides. Target nucleic acid molecules of this length are particularly suitable for the method of the present invention. Other samples with relatively small nucleic acid molecules are some types of forensic samples, fossil samples, nucleic acid samples isolated from environments that are naturally unfriendly to nucleic acid molecule integrity, such as fecal samples, surface water samples, and other microbial-rich samples. For small target DNA (less than 100 nucleotides), it is preferred to use a larger backbone as described herein. The target nucleic acid may also be double-stranded circulating tumor DNA (ctDNA) or cell-free DNA (cfDNA) present in liquid biopsies (including but not limited to blood, saliva, pleural fluid, or ascites). The target nucleic acid may also be a double-stranded or single-stranded cDNA derived from messenger RNA, microRNA, CRISPR RNA, non-coding RNA, viral RNA, or RNA from other sources. The target nucleic acid may also be a double-stranded DNA derived from double-stranded DNA from genomic DNA, PCR products, plasmid DNA, viral DNA, or other sources. The means and methods of the present invention are particularly suitable for capturing small DNA. Preferably 400 or less base pairs. The target DNA is captured in the backbone of the present invention. The captured DNA is also referred to as inserted DNA or target DNA. The target DNA is preferably 400 base pairs or less, more preferably 300 base pairs or less, more preferably 200 base pairs or less, more preferably 150 base pairs or less. The lower limit of the target DNA is preferably 20 base pairs, more preferably 30 base pairs, more preferably 40 base pairs, and more preferably 50 base pairs. Any lower limit can be combined with any upper limit.
[0103] Here, the size of the DNA fragment is given in nucleotides. This of course refers to the number in one strand. It is also possible to give the size of double-stranded DNA in base pairs. So, a DNA with 400 nucleotides is 400 base pairs long.
[0104] The generated concatemers can be sequenced by a variety of different methods. Among them, long-read sequencing methods are preferred. Technicians can use a variety of long-read sequencing methods. Their common feature is that molecules greater than 200 nucleotides are produced in the sequencing reaction. Usually greater than 500 nucleotides and even thousands of nucleotides in length. Two existing platforms for long-read, real-time, single-molecule sequencing are Pacific Biosciences systems (RSII and Sequel) and Oxford Nanopore systems (MK1 MinION, GridION and PromethION). These allow read lengths of more than 55kb or even longer (Goodwin et al. 2016; doi: 10.1038 / nrg.2016.49). The long-read system is preferably a single-molecule real-time sequencing system. The single-molecule system does not rely on a clonal population of amplified DNA fragments to produce a detectable signal. These systems fix the sequence-determining protein at a specific position and allow the nucleic acid chain to advance through the protein. The existing Pacific Bioscience system uses polymerase, while the Oxford Nanopore system currently uses membrane channel proteins. In a preferred embodiment, the sequencing method is a single molecule real-time (SMRT) sequencing method. The generated concatemers have an ordered array of at least one copy of said DNA circle, preferably at least 2, 3, 4 or preferably 5 copies of said DNA circle.
[0105] In some embodiments of the present invention, the backbone has an identifier. Such an identifier is also referred to as a barcode. A barcode or identifier is an extension of a nucleic acid whose sequence may be different between backbones. The barcode may be used to group the sequencing results of a specific DNA ring. The barcode may identify a DNA ring. The barcode may be used to group the sequencing results of fragments of an ordered array of concatemers produced by RCA of a DNA ring. The method using a backbone with a barcode generally has one or more backbone sets, wherein the backbones in the set have unique barcodes in other similar or identical backbones. Two or more backbone sets may be used, for example, to accommodate the different first restriction enzyme sites mentioned above, or to identify the sequencing results of different samples. The barcodes between the sets may be the same, because the sequence differences in another part of the backbone identify the set. The backbone set may include more than one copy of a specific barcode containing the backbone. The combination of a barcode with a specific overall target sequence may also clearly identify the nucleic acid as being derived from a specific DNA ring, for example, when the target nucleic acid is complex and / or the number of identical barcodes in the backbone set is low. The sequencing results of a set of DNA circle sequences can be used to filter out errors such as amplification errors or polymerase errors. This is schematically illustrated in Figure 1 A and Figure 1 B. The backbone in the method as disclosed herein preferably comprises at least two backbones having unique identifiers.
[0106] In the connection step, a DNA ring is generated. Generally, the longer the molecule is, the higher the efficiency of the cyclization is. Flexible molecules are more easily cyclized than rigid molecules. Small target nucleic acids (20 to 200 nucleotides) are captured by larger main chains, and the larger the main chain is, the higher the capture efficiency is. For small target nucleic acid main chains, preferably 200 or more nucleotides, preferably 300 or more nucleotides, more preferably 300 or more nucleotides, more preferably 400 or more nucleotides, preferably between 450 and 650 nucleotides. Smaller main chains generally allow each DNA ring to have more concatemers. The average length of the target nucleic acid and the length of the main chain in the DNA ring are preferably 90 to 16000 nucleotides, preferably 200 to 12000 nucleotides, preferably 300 to 8000 nucleotides, preferably 400 to 4000 nucleotides, preferably 500 to 2000 nucleotides. The average length of the target nucleic acid plus the main chain nucleic acid is preferably about 1000 nucleotides.
[0107] The backbone DNA molecule preferably includes the following sequence:
[0108] >BB1(199bp)
[0109] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTAGAGTACGACTGCGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC;
[0110] >BB2(259bp)
[0111] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTCACCGGGTCTTCGAGAA GACCTGTTTAGAGTACGACTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC;
[0112] >BB2_100(341)
[0113] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTCACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATATATATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTTAGAGTACGACTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC; or
[0114] >BBpX2(557bp)
[0115] GGGCATGCACAGATGTACACGAACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTGAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGGTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGCAGACAAATGGCTCTAGAGGTACCCGTTACATAACTTATAGATGCTGCATGACATAGCCC。
[0116] >BB100_1(143bp)
[0117] GGGCATGCACAGATGTACACGATTCCCAACACACCGTGCGGGCCATCGACCTATGCATACCGTACATATCATATATAAATCACATAATTTATTATACGTATGTCGCGCGGGTGGCTGTGGGTAGATGCTGCATGACATAGCCC
[0118] >BB100_2(143bp)
[0119] GGGCATGCACAGATGTACACGCACTACATGCCAATGCCCAAGCAGTGCGCATATCACGTATCATATCTAATATATTATAATATTATGATAATGAGTATTTATTTAATTTGTTTGTGTGAGGTAGATGCTGCATGACATAGCCC
[0120] >BB100_3(143bp)
[0121] GGGCATGCACAGATGTACACGCATTGGCCGTCTGTGCTGTCCATGGATCGTCTGATTGATATGATATCATATATTATAATTATACAGTAAGGTGATTGGGTATTGAGGGTTGTGTGGTTGGTAGATGCTGCATGACATAGCCC
[0122] >BB100_4(145 bp)
[0123] GGGCATGCACAGATGTACACGGTAGACATGCGAAGCGTGCGATGACAATCGATGTGGACATCATGCATATATATGTTGTATAATTAAACAAATATGTGTAGTGTGTGAGGTGGGTGTAGGAAGTAGATGCTGCATGACATAGCCC
[0124] >BB100_5(143 bp)
[0125] GGGCATGCACAGATGTACACGTTGTCATGGGAATTTGTGGTTATGAAATGAGTATGCGACGAATATGTATACATATATATTAAATTATAGAGTGATGTATGAGTTTGTGATGTGTGGTGTATAGATGCTGCATGACATAGCCC
[0126] >BB200_1(243 bp)
[0127] GGGCATGCACAGATGTACACGGCGGCGCAAGATGATGTGCCGAACCTGACATGGCATCGACTGGTATGGATCAATACTGATGCGATATCGATACCGGATAAATCATATATGCATAATATCACATTATATTAATTATAATACATCGGCGTACATATACACGTACGCATCATTTCACTATCTATCGGTACTATACGTAGTGCCGGTCTGTTGGCCGGGCGACATAGATGCTGCATGACATAGCCC
[0128] >BB200_2(244 bp)
[0129] GGGCATGCACAGATGTACACGTGACGCAACGATGATGTTAGCTATTTGTCAATGACAAATCTGGTATGATCAATACCGATGCGATATTGATATCTGATAACTCATATATGTAGAATATCACATTATATTTATTATAATACATCGTCGAACATATACACAATGCATCTTATCTATACGTATCGGGATAGCGTTGGCATAGCACTGGATGGCATGACCCTCATTAGATGCTGCATGACATAGCCC
[0130] >BB200_3(244 bp)
[0131] GGGCATGCACAGATGTACACGAGACCGCAAGATGATGTTCATTCTTGAACATGAGATCGGATGGGTATGGATCAATACCGATGCGATATGATAACTGATAAATCATATATCTATAATATCACATTATATTAATTATAATACAGGATCGTTACATGCATACACAATGTATACTATACGTATTCGGTAGTTAGTGTACGGTCGGAATGGAGGTGGTGGCGGTGATAGATGCTGCATGACATAGCCC
[0132] >BB200_4(243 bp)
[0133] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGAGATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGCATAATCTCACATTATATTTATTATAATAAATCATCGTAGATATACACAATGTGAATTGTATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATGCTGCATGACATAGCCC
[0134] >BB200_5(243 bp)
[0135] GGGCATGCACAGATGTACACGAATCCGTGAGATGACTATCTTATTTGTGACATTCATCGATCTGGATATGATCAATACCATGCGATATTGATTACTGATAAATCATATATGTAGAATATCACATTATATTAATTATAATAAATCGTCGTACATATACATCCACAATTAGCTATGTATACTATCTATAGAGATGGTGCATCATCGTACTCCACCATTCCCACTAGATGCTGCATGACATAGCCC
[0136] >BB300_1(348 bp)
[0137] GGGCATGCACAGATGTACACGCATAAGACCACAGGGTGCAAATCTGGATTGCGGCATGGATGATTCATCATCGTGGCATATTCGCTATGGATATATCCATCATAATACATTGATACGTCATGCGTATAATCGCATTATATGTCGATATTGGTCATAGGGATACATCCGTGTATACTATCGTATATGCGTGCAATGTAGCCATGTTAATCATGCTATAACCATAACATAAATATAATATATACAGATGGTGTATCTCTACTTATGTATGCTTGTATAGTAATGTCGATACTGATGGGTCTCCGGCCCACTACACCACCTGGCCGCTCTAGATGCTGCATGACATAGCCC
[0138] >BB300_2(343 bp)
[0139] GGGCATGCACAGATGTACACGGGCAATCCGCCAGGGTTCAAATATGGATATGTGATGATCGATTCAACATGCACATATGCACGATATCATATATTACTCCAGATGTCATCATCGTCGTGCGTATATGAGATATGTATTTATGCATATAATCCACCATACATGGTAGCGATATTATAGTGCGATTATGTGTATATGACTATCATGGCTATTGTTAATATATAAATCATAACCATACCACTTCCACGCCTGGTATGGCGTATAGTATAGAGATATTGTGTGATGCCCTATGTCGACCATGATGTGCCGTTGTACTGCCAATCCTAGATGCTGCATGACATAGCCC
[0140] >BB300_3(344 bp)
[0141] GGGCATGCACAGATGTACACGTATCCATGCAGCTTATTGTAACTAGCGCATGCACGTGGTGATTCATCACATCTATATATACGATATGATATATTACACATATTTGCATAGTATCATCCGGTGTGATATCATCCGATATGCTCATACTTATTCATTGGTAGCATTGCATTGATGGATCAATAGTTATTATGACATCATGGCATGTACAATTATAAATAATACAACATACATAAATATACTATACACATCGTGTATGTGTTATACAGATCTGTGTGATGTATGATAATGTAATGGCGTCGAACACCACAAGGCAGTCCTATAATAGATGCTGCATGACATAGCCC
[0142] >BB300_4(344 bp)
[0143] GGGCATGCACAGATGTACACGGTCCATTACAATCGAATCTATATCCCAATGTGTATCGATTATCACCACAATGACATAATACGATATCATATATTACTCCATATGCCTTACGTCAGATCGTTATATGAGATATGTATTCATGCATATGATATCCCACAGTACACGTCGTCTAATGCCATCATGAATGTATGACATATCTAGTCGATTATACATAATATAACATACCAATATAACAATATCTATACACATTTGATGGCGTATAGTATAAAGATATTGTGGCAATGCCCATACACCACTGACTGTCGCCGATCATTCCTACCACTAGATGCTGCATGACATAGCCC
[0144] >BB300_5(344bp)
[0145] GGGCATGCACAGATGTACACGACCGACCGTGAAAGTGATTCAGAATGATGTGCATGAATGTTATCATGACATGATTTATGATGCACTGATATATGCATATTATAATATTGTACAATGTCGTATATACGACATATCTATACTATGAATTATGGCATCATGGACAATAGATGGTAAGGTATAGTACGATCTATATAGCATGTTGAAATGGGATATAAATTATCATAAACATACATACTTAACTAATATCAAGATGATATGTGTATGACATCAGAATGATAGTAGTAATGAGTATTGTCAGATGTATGTACGAATATCACACGATTAGATGCTGCATGACATAGCCC
[0146] The backbone DNA molecule most preferably comprises the following sequence:
[0147] >BB200_4(243bp)
[0148] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGAGATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGCATAATCTCACATTATATTTTATTATAATAAATCATCGTAGATATACACAATGTGAATTGTATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATGCTGCATGACATAGCCC
[0149] The flexibility of the main chain of fixed length can be adjusted by the sequence of customizing the main chain. According to the specific sequence of the molecule, different DNA molecules have different flexibility. Different sequences are provided by selecting different first restriction enzyme sites, different barcode sequences and different sequences for other elements in the main chain. Preferably, flexibility is adjusted by the sequence of the dedicated part of the custom main chain sequence. This dedicated part is further referred to as "connector". Connector preferably includes 20 to 900 nucleotides, preferably 25 to 900 nucleotides, preferably 30 to 900 nucleotides, preferably 30 to 800 nucleotides, preferably 50 to 700 nucleotides, preferably 100 to 600 nucleotides, preferably 150 to 500 nucleotides. Connector can be a continuous sequence on the main chain or be divided into two, three, four or more continuous sequences. Connector is preferably a continuous sequence on the main chain or is divided into two, three, four continuous sequences, preferably one, two or three, preferably one or two, and more preferably a continuous sequence.
[0150] The values of the dissociation energy per base pair (Breslauer et al. 1986) and the deviations in torsion angles (degrees) (Sarai et al. 1989) can be used to calculate the flexibility of any given DNA sequence. An example of such a calculation is:
[0151] Flexible computing
[0152] The python implementation of the TwistFlex algorithm (http: / / margalit.huji.ac.il / TwistFlex / ) (Menconi et al. 2015) can be used to calculate DNA flexibility at the torsion angles of the input sequence. The flexibility of each single dinucleotide is calculated based on the following angle table:
[0153]
[0154] Subsequently, the evolutionary algorithm used for main chain optimization takes into account the average flexibility of the entire sequence when selecting. The average flexibility of the DNA sequence is calculated as the summation of all dinucleotide angles divided by the total number of dinucleotides. The flexibility score of a suitable main chain is 10 or greater, preferably 11 or greater, preferably 12 or greater, preferably 12.5 or greater (dinucleotide angles / dinucleotides in the main chain). It is generally not desirable to be greater than 14 flexibility.
[0155] Entropy calculation for determining sequence complexity
[0156] The Shannon entropy of a string is defined as the minimum average number of bits per symbol required to encode the string. The formula used to calculate the Shannon entropy is:
[0157]
[0158] Among them, p i is the probability that character number i appears in the sequence.
[0159] This calculation can also be performed via http: / / www.shannonentropy.netmark.pl / .
[0160] Use the following python code to implement the above formula:
[0161]
[0162]
[0163] Preferred backbones have a Shannon entropy value of 1.5 Sh or higher, preferably 1.5 or higher, preferably 2.5 or higher, more preferably 3.5 or higher.
[0164] Self-complementarity
[0165] The backbone core sequence preferably does not have 8 or more continuous bases that are self-complementary in the same chain. The exception is the intentional insertion of one or more restriction enzyme sites or one or more other functional sequences. The sequence occasionally introduces self-complementary bases in the same chain. If possible, this base greater than 8 should be avoided, but this is permissible in the functional backbone. However, when designing a new backbone, this sequence is preferably avoided if possible. The same is true for kmer discussed below.
[0166] Lack of repetitive motifs (kmer)
[0167] The backbone core sequence preferably does not have a motif of 6 bases that is repeated more than twice in the sequence.
[0168] The flexibility of the backbone and the Shannon index of the backbone can be adjusted by adding linkers to the backbone. The effect of the specific sequence of the linker on the flexibility and complexity score of the backbone can be easily calculated.
[0169] The linker preferably has one or more of the following characteristics: (i) the overall complexity of the linker sequence is preferably high. The above-mentioned Shannon entropy formula is a method for determining the value of the complexity of a given sequence; (ii) copies of DNA motifs longer than 5 bases preferably appear no more than twice, preferably no more than once, in the linker sequence. In a preferred embodiment, the linker does not include copies of DNA motifs longer than 5 bases (i.e., the linker sequence does not contain repeated motifs, wherein the motif is more than 6 consecutive bases); (iii) the linker preferably does not include more than two, preferably does not include more than one, preferably does not include more than 6, self-complementary sequences of nucleotides (inverted repeats) separated by less than 10 nucleotides. The mentioned criteria help to avoid, in general, the appearance of complex secondary structures in the single-stranded version of the linker. The possibility and strength of secondary structures can also be calculated by other means.
[0170] The backbone preferably comprises a GC content of 30%-60%; preferably 40%-60%; preferably 40%-50%; preferably 45%-55%.
[0171] The main chain preferably has one or more of the following features. The first restriction enzyme site is preferably used to produce a blunt-ended restriction enzyme. It has been observed that this improves the capture of the target nucleic acid. The main chain preferably includes a recognition site for a DNA nickase that is used to produce a primer site for rolling circle amplification. Additional restriction enzyme recognition sites can be used to carry out the sequential connection of multiple short DNA molecules to a circular DNA. The main chain preferably includes a molecular identifier that distinguishes the original captured nucleic acid and its subsequent sequencing read length.
[0172] The method of the present invention can be used to capture two or more target nucleic acids in order on each backbone. A single capture step can sometimes capture two target nucleic acid molecules at the same time. The probability of this happening is deliberately low because measures are taken to prevent self-connection. The orderly capture of two or more target nucleic acids can be a desired feature. Additional restriction enzyme sites can be incorporated into the backbone. Once the first target nucleic acid is captured, the method can be repeated by adding restriction enzymes that cut additional restriction enzyme sites. The DNA loop is cut and linearized by the second restriction enzyme and is prepared to be connected to the target nucleic acid. If the second restriction enzyme produces the same type of end (e.g., blunt end) as the first restriction enzyme, the reaction can continue to capture the target nucleic acid that was not captured in the first iteration of the method. Alternatively, the DNA loop can be purified (e.g., by removing linear DNA), and a new target nucleic acid can be added, the end of the target nucleic acid is compatible with the backbone end produced by the second restriction enzyme. By adding further restriction enzyme sites to the backbone, this step can of course be repeated for orderly capture of the third, fourth, etc. target nucleic acids. When more than one target nucleic acid is captured, preferably, the second and first restriction enzyme sites are enzyme sites that are not often cut. This enzyme is preferably 8 or more cutters. This enzyme is preferably a flat end that produces an enzyme. This orderly capture allows simultaneous sequencing of more than one nucleic acid. Different target nucleic acids can be identified according to their position in the backbone, that is, according to the flanking backbone sequence they insert.
[0173] In addition, the backbone serves as a control sequence during data analysis. Based on the backbone sequence read length, the error rate of each sequencing read can be inferred, allowing accurate estimation of the likelihood of genetic variation within the captured nucleic acid sequence.
[0174] Byproducts may be generated in the methods of the present invention. The number of single backbone, single target DNA containing DNA circles is influenced by, for example, the molar ratio of backbone / sample, which should promote the formation and insertion of molecules with backbones rather than unwanted byproducts ( Figure 2 ), for example: (i) linear DNA formed by random concatenation of backbone and sample DNA, (ii) circular DNA containing only backbone or only sample DNA, (iii) circular DNA containing excess backbone or sample DNA.
[0175] In an embodiment of the present invention, preferably, the molar ratio of the backbone molecule to the target nucleic acid molecule is in the range of 1:10 to 10:1. Preferably, the ratio is in the range of 1:5 to 5:1, preferably in the range of 1:2 to 2:1. Preferably, the average ratio is 1:1.
[0176] The methods described herein involving rolling circle amplification are preferably performed without the use of switching containers. The generated concatemers can be sequenced in the same container or in different containers.
[0177] The method of the present invention preferably produces concatemers as long as the linear dsDNA formed by multiple units consisting of target nucleic acid-backbone copies (>10Kb). The concatemerization / polymerization of such units is conducive to distinguishing the detection of true genetic variations from sequencing errors. In fact, in the case of rare genetic variations with a frequency of less than 1% in the DNA molecule library, direct sequencing such as short-read sequencing may no longer be applicable because the sequencing error rate is higher than the mutation frequency. Using the method described herein, the same rare sequence (genetic variation) is represented multiple times in long concatemers, which provides high confidence in the presence of mutations, even if the mutation frequency is low in the original library of nucleic acid molecules.
[0178] The backbone includes a 3' sequence encoding a portion of a first restriction enzyme recognition site (restriction enzyme site) and a 5' sequence encoding another portion of a first restriction enzyme cleavage site. The backbone may contain further elements, such as one or more of the following: (i) one or more sites that allow nicking of the double-stranded backbone sequence; (ii) one or more type 1 or type 2 restriction enzyme cleavage sites; (iii) a secondary cloning site; (iv) a flexible DNA extension (linker) that allows efficient circularization (bending) of the backbone molecule; and (v) a unique molecular barcode sequence used to label each single backbone molecule.
[0179] The backbone preferably has 5'-phosphorylation at both ends of the backbone molecule.
[0180] The present invention also provides a collection of linear DNA molecules (backbone) with a length of 20 to 1000 nucleotides, the linear DNA molecules comprising a 5' end and a 3' end, the 5' end comprising a portion of a first restriction enzyme recognition site at the most terminal end, the 3' end comprising another portion of the first restriction enzyme recognition site at the most terminal end, and the 5' end and the 3' end are compatible with each other for ligation and form a restriction enzyme (first restriction enzyme) recognition site when self-ligated, and wherein each chain of the backbone comprises:
[0181] Connector;
[0182] an identifier sequence (barcode) that is different from the sequence of identifiers of other backbones in the set; and
[0183] Optional restriction enzyme site for nicking enzyme.
[0184] The main chain is a preferred main chain in the method described herein. The first part and the second part of the first restriction enzyme cutting site together form a complete recognition site of the first restriction enzyme cutting site, and are located at a position on the molecule that allows the two parts to form an operably connected position of the first restriction enzyme cutting site. Operably connected in the context refers to the effectiveness of being cut by the first restriction enzyme. The main chain preferably further includes a second restriction enzyme cutting site that is a type I or type II restriction enzyme site. The main chain preferably further includes a restriction enzyme site (Golden-Gate cloning site) for a type II restriction enzyme that can produce a non-palindromic overhang. The connector is preferably a connector as described above. The main chain preferably includes a nucleic acid molecule (captured nucleic acid molecule) in the first restriction enzyme cutting site. The main chain preferably includes a captured nucleic acid molecule library.
[0185] A kit comprising a backbone as described herein is further provided. The kit preferably comprises a collection of backbone molecules as described herein. The kit preferably further comprises a polymerase with high processivity and optionally one or more polymerizing primers. The kit preferably further comprises: a ligase and the first restriction enzyme; and / or the target site-specific recombinase. The kit preferably further comprises a DNA exonuclease. The latter enzyme is suitable for removing linear DNA before generating a concatemer of the DNA ring.
[0186] In one aspect, the present invention provides a method for determining the sequence of a collection of nucleic acid molecules, the method comprising:
[0187] - providing a double-stranded target DNA molecule having a 5' end and a 3' end, the double-stranded target DNA molecule having overhanging adenine residues at the 3' ends of both strands of the DNA molecule;
[0188] - Providing a collection of double-stranded backbone DNA molecules comprising 5' and 3' ends that are ligation-compatible with the 5' and 3' ends of the target DNA.
[0189] The method further comprises:
[0190] - ligating the target DNA to the backbone in the presence of a ligase, thereby generating a DNA circle comprising the backbone and the target DNA molecule;
[0191] -optionally remove linear DNA;
[0192] - generating concatemers of ordered arrays comprising at least two copies of said DNA circles by rolling circle amplification; and
[0193] - sequencing the concatemers.
[0194] The ends compatible with the prominent 3' adenine connection are the ends with 5' prominent thymidine bases or their analogs. The difference between this method and the method described above is that, since all ends have a'-protruding adenine bases, the molecular connection between or within the target is inherently suppressed. Therefore, the self-connection of the target nucleic acid or the connection of one end to another target nucleic acid molecule is inherently impossible. The prominent bases are the nucleotides at the ends of the nucleic acid molecules, and they do not base pair with the bases on the opposite strand. There is no opposite base for the prominent bases. This protrusion is also referred to as a sticky end or a sticking end. The same is true for the main chain. They inherently prevent self-connection. In this embodiment, the main chain does not have to have a portion of the first restriction enzyme site at the end. The end therefore cannot be connected to produce the first restriction enzyme site. Therefore, the connection does not have to be carried out in the presence of the first restriction enzyme. The remaining steps and definitions can be the same as described elsewhere in this article.
[0195] Further provided is a method for determining the sequence of a collection of nucleic acid molecules comprising:
[0196] Providing a double-stranded target DNA molecule having recombinase recognition sites at the 5' end and the 3' end, wherein the recombinase recognition sites are specific for a target site-specific recombinase;
[0197] providing a backbone comprising said recognition site separated by a DNA comprising a linker;
[0198] Incubating the target DNA molecule with the backbone in the presence of the target site-specific recombinase, wherein the target site-specific recombinase is preferably Cre recombinase, FLP recombinase or bacteriophage lambda integrase, thereby generating a DNA circle including the backbone and the target DNA molecule;
[0199] optionally removing linear DNA; and
[0200] generating concatemers of an ordered array comprising at least two copies of said DNA circles by rolling circle amplification; and
[0201] The concatemers are sequenced.
[0202] In a preferred embodiment, the main chain is a ring comprising two recombinase recognition sites, which are separated by the DNA comprising a connector on one side and by the DNA encoding a further restriction enzyme recognition site on the other side, and wherein the further restriction enzyme site is the unique recognition site of the restriction enzyme in the main chain. In this embodiment, the method preferably further comprises: before producing the concatemer, digesting the DNA and subsequently removing the linear DNA after recombination with the restriction enzyme. The further restriction enzyme site is preferably 6 or more cutters, preferably 7 or more cutters, preferably 8 cutters. The end produced by enzyme cutting is not necessarily a blunt end. In a preferred embodiment, the further restriction enzyme is not a blunt end cutter.
[0203] Target site specific recombinase
[0204] Target site specific recombinase is a gene recombinase. Target site specific DNA recombinase is widely used in multicellular organisms to manipulate genome structure and control gene expression. These enzymes, derived from bacteria and fungi, catalyze directional sensitive DNA exchange reactions between short (30-40 nucleotides) target sequences for each recombinase. These reactions can carry out four basic functional modules, excision / insertion, flipping, translocation and box exchange. Non-limiting examples of recombinases are Cre recombinase, Hin recombinase, Tre recombinase and FLP recombinase. Cre recombinase is the first widely used recombinase. It is a tyrosine recombinase derived from P1 phage. The enzyme uses a topoisomerase I-like mechanism to carry out site-specific recombination events. The enzyme (38kDa) is a member of the integrase family of site-specific recombinases, and it is known that it catalyzes site-specific recombination events between two DNA recognition sites (LoxP sites). The 34 base pairs (bp) loxP recognition site consists of two 13bp palindromic sequences, which are located on the side of the 8bp separated region. The product of Cre-mediated recombinase at loxP site depends on the position and relative orientation of loxP site. Two separate DNA species, both including loxP site, can be fused by Cre-mediated recombination. The DNA sequence found between two loxP sites is called "knock-in (floxed)".
[0205] Red / ET recombinase
[0206] Recombineering utilizes phage-derived protein pairs, either RecE / RecT from Rac phage or Redα / Redβ from λ phage, to facilitate the cloning or subcloning of DNA fragments into vectors without the need for restriction enzyme sites or ligases. RecE / RecT, Redα / Redβ, and other similar protein pairs are further referred to herein as RecE / RecT protein pairs. A limitation of the original homologous recombination technique was due to the fact that the bacterial RecBCD nuclease degrades linear DNA and the initial events had to be studied in RecBCD-deficient strains (7). This problem was overcome with the discovery that Redγ can assist Redα and Redβ to inhibit RecBCD nuclease activity, allowing the technique to be applied to Escherichia coli (E. coli) and other commonly used strains. In addition, recombination efficiency was increased 10-100 fold. The combination of these three enzymes (α, β and γ, or E, T and γ) in one vector is named Red / ET recombination, and the basic principle of the method is that in order for recombination to occur in a linear fragment, at both ends of a double-strand break (DSBs), and in another linear or circular plasmid, two homology regions greater than 15 bp, preferably greater than 20 bp, preferably greater than 30 bp, and preferably greater than 42 bp are required. Directed insertion can be achieved by flanking the target DNA and the insertion site with two different homology regions. DSBs are essential so that RecE or Redα can bind and degrade one strand of the DNA (5' to 3') and simultaneously load RecT or Redβ to the exposed single-stranded strand. The single-stranded DNA loaded with RecT or Redβ recombinase finds a perfect match sequence and joins the two sequences by strand invasion or annealing.
[0207] Insertion of homology regions (HRs) is usually achieved by including them in the oligonucleotides used for amplification of the product, which serves as a linear substrate for the recombination event. If the process requires longer DNA fragments, HRs can be inserted by traditional restriction / ligation techniques using plasmids or adapters.
[0208] Restriction enzyme recognition site
[0209] Restriction enzyme recognition sites are also often referred to simply as restriction enzyme sites, restriction enzyme cleavage sites or restriction recognition sites. They are locations on a DNA molecule containing a specific sequence of nucleotides that can be recognized by a restriction enzyme. These are usually palindromic sequences. A specific restriction enzyme can cleave a sequence between two nucleotides at or near its recognition site. The enzyme usually cleaves both strands of the DNA molecule, usually followed by separation of the ends. So-called nicking enzymes also recognize restriction enzyme cleavage sites, but only cleave one of the two chains. The resulting DNA molecule is still bound, but one of the two chains has a nick.
[0210] Restriction enzyme type
[0211] Naturally occurring restriction endonucleases (restriction enzymes) can be divided into four groups (type I, type II, type III and type IV) based on their composition and enzyme cofactor requirements, the nature of their target sequence and the position of their DNA cleavage site relative to the target sequence. However, DNA sequence analysis of restriction enzymes shows great diversity, indicating that there are more than four types. All types of enzymes recognize specific short DNA sequences and perform endonucleolytic cleavage of the DNA to produce specific fragments with 5'-phosphate termini.
[0212] Type I enzymes (EC 3.1.21.3) cleave at a site distal to the recognition site and require ATP and S-adenosyl-L-methionine to function. They are multifunctional in that they have both restriction enzyme and methylase (EC 2.1.1.72) activities.
[0213] Type II enzymes (EC 3.1.21.4) cleave within the recognition site or at a short specific distance from the recognition site. Most type II enzymes require magnesium ions. They usually have a single function (restriction).
[0214] DNA phosphorylation
[0215] Single-stranded or double-stranded DNA with a 5'-hydroxyl terminus must have a 5' phosphate group for efficient ligation. 5' ends without such a phosphate group can be phosphorylated prior to ligation. Several polynucleotide kinases, including T4 PNK (NEB #M0201) and T4 PNK (minus 3' phosphatase) (NEB #M0236), can be used to transfer the γ-phosphate of ATP to the 5' end of DNA.
[0216] DNA dephosphorylation
[0217] Digested DNA usually has a 5' phosphate group required for ligation. To prevent self-ligation, the 5' phosphate can be removed prior to ligation. Dephosphorylation of the 5' end inhibits self-ligation, allowing the technician to manipulate the DNA as needed before re-ligation. Dephosphorylation can be accomplished using any of several phosphatases, including the Rapid Dephosphorylation Kit (NEB #M0508), Shrimp Alkaline Phosphatase (rSAP) (NEB #M0371), Calf Intestinal Alkaline Phosphatase (CIP) (NEB #M0290), and Antarctic Phosphatase (NEB #M0289).
[0218] DNA ligation
[0219] DNA ligation is a central step in many modern molecular biology workflows. DNA ligase catalyzes the formation of a phosphodiester bond between the 3' hydroxyl and 5' phosphate of adjacent DNA residues. In the laboratory, this reaction is used to join dsDNA fragments with blunt or sticky ends to form recombinant DNA plasmids, to add barcoded adapters to fragmented DNA in next-generation sequencing, and many other applications. DNA ligase from T4 bacteriophage is the most commonly used ligase. It can join sticky or cohesive ends of DNA, oligonucleotides, and RNA and RNA-DNA hybrids. It can also ligate blunt-ended DNA with high efficiency. Using Circ Ligase TM II ssDNA Ligase* (Epicentre) efficiently ligates single-stranded DNA. It is a thermostable enzyme that catalyzes the intramolecular ligation (i.e., circularization) of ssDNA templates containing a 5'-phosphate and a 3'-hydroxyl group. Circ Ligase TM II ssDNA ligase* ligates the ends of ssDNA in the absence of complementary sequences. Therefore, this enzyme is very useful in synthesizing circular ssDNA molecules from linear ssDNA. Circular ssDNA molecules can be used as substrates for rolling circle replication or circular transcription.
[0220] For the sake of clarity and concise description, when referring to a step as being carried out on or with one or more substrates and which step is catalyzed by one or more enzymes, the step is carried out by contacting the substrate with the enzyme. This is usually accomplished by adding the enzyme to the substrate in an appropriate buffer.
[0221] For clarity and concise description, features are described herein as part of the same or separate embodiments, however, it is to be understood that the scope of the invention may include embodiments having combinations of all or some of the described features. BRIEF DESCRIPTION OF THE DRAWINGS
[0222] Figure 1 A) Schematic diagram of a method for capturing small nucleic acid molecules and generating concatemers by using a backbone and rolling circle amplification. B) Schematic diagram of a sequencing reaction using a short read length and no backbone and a long read sequencing using a backbone.
[0223] Figure 2 , Figure 1 Examples of possible linear and circular byproducts of the cyclization reaction indicated in . The large circle shading in the upper left corner of the circular byproduct diagram represents the backbone sequence. The other shading is the target sequence.
[0224] Figure 3 , schematic example of a double-stranded backbone sequence.
[0225] (1) indicates the 5' and 3' sequences that together encode the first restriction enzyme recognition site.
[0226] (2) is a restriction enzyme site for the nicking enzyme BbvCI. Any other nicking site will do, however the advantage of using BbvCI is that two forms of this enzyme are commercially available, one that cleaves DNA on the plus strand and another that cleaves DNA on the minus strand. The nicked DNA is an effective introduction site for rolling circle amplification (RCA) reactions. Depending on the circumstances, we may want to use the nicked DNA instead of a DNA primer to initiate the polymerization reaction.
[0227] (3) is an additional blunt-ended restriction site, in this example, a recognition site for SweI. The second blunt-ended restriction site allows the capture of a second DNA fragment in a further circularization reaction.
[0228] (4) is a cloning site, in this example a double flipped BbsI site, which can be used for simple extension of the backbone via Golden-Gate or other types of cloning.
[0229] (5) represents a flexible DNA extension (linker), which has variable length and assists in efficient circularization.
[0230] (6) Capital N indicates a stretch of nucleic acid encoding a unique identifier. It is a barcode-like sequence. It can encode one or more (random) barcodes of any suitable size.
[0231] Element (1) is positioned at the very end, and elements 2-6 may have any order and may be present or absent depending on the situation.
[0232] Figure 4 We have developed methods to detect gene fusions based on targeted cDNA synthesis, single-stranded DNA circularization (ssDNA), and targeted rolling circle amplification. The cDNA produced by the reverse transcriptase step is shaded gray in the left figure. The cDNA at the bottom is the fusion gene DNA, and there are two shades indicating a portion from one gene and a portion from another gene. It is clear that the RCA assay produces multiple concatemers of fusions. Of course, this method can also be used to determine the sequence of one or more cDNAs that are not the result of a gene fusion.
[0233] Figure 5 , Schematic diagram of MIP probes and methods.
[0234] Figure 6 , plot the DNA content and predicted value of each band.
[0235] Figure 7 , Determine the cyclization efficiency: A) Comparison of insertion before and after the reaction. B) Comparison of the cyclized product and the unreacted product.
[0236] Figure 8 Results of proof-of-concept experiments. (A) Gel image indicating RCA products used for subsequent sequencing on the nanopore MinION. (B) Distribution of nanopore read lengths generated by a MinION R9.4 experiment using the samples indicated in (A) as input. (C) Distribution of pattern scores for 2083 read lengths greater than 10 kb. (D) Schematic representation of nanopore sequence reads, replacement inserts (green) and backbones (red), alignment of insert sequences, and generation of consensus sequences from aligned inserts. (E) Accuracy of consensus sequences.
[0237] Fig. 9 , Circularization of Backbone 2 (BB2) and Backbone 3 (BB3) with Insert 17.2 in a 3:1 ratio. A) Comparison of BB2 and BB3. Red asterisk: correct circularization product. Multiple bands in condition 3: linear product vs. circular product, connection of multiple backbones. Additional bands in condition 4: circularized backbone. B) Successful circularization using Backbone 2 (BB2). Yellow asterisk: correct circularization product. In this gel image, the entire reaction was loaded into each lane. Lanes 1 and 2 represent the circularization of BB2 with Insert 17.2 in a 1:1 ratio before and after PlasmidSafe digestion. Lanes 3 and 4 represent the same circularization reaction using BB2 with Insert 17.2 in a 3:1 ratio.
[0238] Fig.10 , backbone circularization efficiency at different backbone insertion ratios. The ligation products of different backbone (BB3) to insert (17.2) ratios were qualitatively and quantitatively detected. (A) Agarose gel showing the circularization input and reaction products. PlasmidSafe digestion was used to remove linear products after the circularization reaction. Red asterisk: correct circularization product. Yellow asterisk: remainder of the insertion input. (B) Quantification of circularization efficiency. The quantification of circularization efficiency is defined as P / I*100, where P is the amount of correct backbone insertion product (in moles, red asterisk), and I is the amount of input insertion (in moles). ImageJ software ( https: / / en.wikipedia.org / wiki / ImageJ ) The intensity and surface area of the bands were measured. The data were normalized using the GeneRuler 50 bp DNA ladder as a reference. See also Section 10 of Materials and Methods.
[0239] Fig.11 , BB2_100 (orange column) with and without SrfI and HMGB1. Ligation was performed using the backbone to insert ratio. The blue column represents a control experiment in which BB2 and BB3 were ligated to the same insert without SrfI and HMGB1. As described above ( Fig.10 Figure legend) Quantification of cyclization efficiency.
[0240] Fig.12 , Visual display of the reaction products of the cyclization with backbone BB2_100 and insert 17.2. Red asterisk: correct product. Orange box: expected position of the remaining insert after cyclization. The complete disappearance of the insert band after ligation shows that the insert is completely ligated. Circularized 1: before Plasmid Safe DNase treatment. Circularized 2: after Plasmid Safe DNase treatment.
[0241] Fig.13 Effect of adding restriction enzyme SrfI to the cyclization reaction. Circularization reaction was performed using BB3 together with insert 17.2. Reactions were performed in the presence and absence of SrfI and Plasmid Safe DNase.
[0242] Fig.14 Figure 2. Barcoding strategies useful for the described technologies. (A) Tagging of single DNA molecules using unique molecular identifiers to improve mutation discovery. (B) Tagging of single samples using sample-specific barcodes for pooling of sequencing experiments.
[0243] Fig.15 , RCA products using various DNA templates. RCA was performed using circular DNA templates derived from various sources. (A) Cell-free DNA circularized using backbone BB2; (B) plasmid pX_Zeo; (C) ss-cDNA self-circularized using Circ Ligase II (Epicentre #CL9021K); (D) PCR product 17.2 cloned into plasmid pJET. As a reference, a long range 1Kb ladder was used. The taller band of the ladder was 10Kb long, and the RCA products were estimated to be between 20kb and 100kb long.
[0244] Fig.16 , the number of read lengths containing 17.1 and 17.2. The ratio between reads containing 17.1 and reads containing 17.2 is 1:14, indicating a clear enrichment of the target region due to site-directed RCA.
[0245] Fig.17 , one Overview of the reaction products of the steps in the pot reaction design. The insert (17.2) and backbone (BB2_100) are circularized to produce the products indicated at (1). The linear DNA product is digested with Plasmid Safe DNase as indicated at (2). RCA reaction products are formed based on (2) as input as indicated at (3).
[0246] Fig.18, Overview of consensus sequence calling methods for short-read sequencing (left) and long-read sequencing (right).
[0247] Fig.19 (A) Example of a mapped insertion-derived Cyclomics sequencing read with a TP53 mutation (red box). (B) Graph showing the fraction of insertions supporting non-reference alleles among 558 read lengths with >4 insertions. Four reads showed a high fraction of non-reference alleles, and these contained insertions with the expected chr17:7578265, A->T mutation.
[0248] Fig. 20 , capture DNA with target site specific recombinase. The recognition sites for target site specific recombinase are indicated by letters A and B. The target DNA is indicated by the wording "insert". Sites A and B can be introduced in various ways, such as by connecting an adapter with sites to the insert DNA or by amplifying the insert with primers, which include sequences encoding the sites A and B. The main chain is indicated by the term "main chain". In the figure, the main chain is a circular molecule including DNA between two sites A and B. The intermediate DNA includes restriction enzyme sites unique to the entire main chain. The arrow indicates that the insertion and main chain are first recombined by adding a recombinase and then adding a restriction enzyme. The restriction enzyme will only cut the unreacted main chain and the main chain in which the connector is inserted and replaced. The linearized DNA can be removed by adding a suitable exonuclease.
[0249] Fig.21 , Comparison between the efficiency of ligation reaction products and different backbone designs. Left: Ligation of different backbones with a 250bp PCR amplicon. Right: Remaining circular product after digestion of linear DNA with Plasmid Safe DNase. Ligation to all members of the BB200 series showed high circularization efficiency, as indicated by the formation of a circular product including the PCR product and the backbone.
[0250] Fig. 22 , Comparison of the ligation efficiency of backbones from the BB200 series. Left: Gel showing the ligation products of the 3 backbones. BB200_4 shows a brighter band indicating more product was formed during the reaction. Right: Measurement of the brightness of the bands. From top to bottom: BB200_2, BB200_4 and BB200_5.
[0251] Fig.23 , the difference in the number of sequencing reads derived from the RCA products formed by the ligation of the backbones of the BB200 series. The percentage of reads from different backbones in two independent experiments (red and blue). The backbones were initially mixed in a 1:1:1 ratio. BB200_4 with a higher number of reads and Figure 2This is consistent with the higher ligation efficiency shown in .
[0252] Fig.24 , gene TP53 (GRCh37 17: 7577518). The Y-axis represents the distance (average match score) between the patterned nanopore signal corresponding to the reference sequence and the signal derived from the experimental sequence. The larger the distance, the more difficult it is to infer the correct base. On the X-axis is the number of insertions that appear in the read fragment. The inferred bases are indicated in different colors. The signal from the positive strand is not as clear as the signal measured on the reverse strand. This makes it difficult to distinguish the correct base from other possible bases even when the calculated distance is low (A, blue).
[0253] Fig.25 , Agarose gel depicting the products of the cyclization reaction between the backbone and the insert (S). Negative control is indicated as C-. The band corresponding to the circular BB-I product was isolated from the gel.
[0254] Fig.26 , Agarose gel showing example products after rolling circle amplification.
[0255] Fig. 27 , BB200_4 (243 bp, as indicated by BB in the figure) and S1_WT (158 bp, as indicated by I in the figure) were circularized and amplified by RCA. When the concatemers made by BB-I are digested, we expect a band of about 400 bp, whereas if the concatemers consist only of BB, a band of about 250 bp is produced. The concatemers formed only by I will not be digested, making the RCA bands visible.
[0256] Example
[0257] Example 1
[0258] Materials and methods
[0259] The method described herein is to sequence the nucleic acid sequence in the nucleic acid molecule (DNA) library to allow the detection of gene mutations, such as ctDNA (circulating tumor DNA), cfDNA (cell-free DNA), genomic DNA, RNA, polymerase chain reaction (PCR) products or other products. In this process, the "product" obtained is a long linear dsDNA (> 10Kb), formed by multiple units consisting of nucleic acid-main chain copies. The multimerization / polymerization of this unit is necessary to distinguish the detection of true gene mutations from sequencing errors. In fact, in the case of rare gene mutations with a frequency of less than 1% in the DNA molecule library, direct sequencing, for example, short read sequencing, is no longer applicable because the sequencing error rate is higher than the mutation frequency. Using the above method, the same rare sequence (gene mutation) is represented multiple times in the long multimer, which provides a very high confidence in the presence of mutations, even if the mutation frequency in the original nucleic acid molecule library is low.
[0260] The design of the backbone of the capture nucleic acid provides many advantages for efficient and specific acquisition of nucleic acids and is critical for computational analysis of sequencing data.
[0261] Preferred backbone molecules are characterized by:
[0262] 1) Blunt-ended restriction sites encoded at the extreme ends of the backbone are used to improve the ligation efficiency of short DNAs.
[0263] 2) A recognition site for a DNA nickase, which is used to generate a "single-stranded template" for rolling circle amplification.
[0264] 3) Restriction enzyme target sites can be used to continuously connect multiple short DNA molecules into a circular DNA.
[0265] 4) Molecular identifiers that allow differentiation between the original captured nucleic acid and its subsequent sequencing read.
[0266] In addition, the backbone serves as a control sequence during data analysis. Based on the backbone sequence reads, the error rate of each sequencing read can be inferred, allowing accurate estimation of the probability of genetic variation within the captured nucleic acid sequence.
[0267] 2. Materials and methods used for cyclomics technology
[0268] 2.1 Main chain design
[0269] We went through an iterative approach of replacing backbone designs, followed by experimental testing to find the best backbone that enables the most efficient circularization (i.e., capture of double-stranded nucleic acid molecules). Our basic backbone design can include one or more of the following parts:
[0270] 1) 3' sequence encoding half of the restriction enzyme cleavage site.
[0271] 2) One or more sites that allow nicking of the double-stranded backbone sequence
[0272] 3) One or more type 1 or type 2 restriction enzyme sites
[0273] 4) Secondary cloning site
[0274] 5) Flexible DNA extension that allows efficient circularization (bending) of the backbone molecule
[0275] 6) Unique molecular barcode sequence used to label each single backbone molecule
[0276] 7) A 5' sequence encoding the other half of the same blunt-ended restriction site used in 1
[0277] 8) Phosphorylation of the 3' and 5' ends of the main chain molecule
[0278] Flexible DNA stretches were designed with the help of a custom-made evolutionary algorithm that utilized various selection criteria among the following:
[0279] 1) The high overall complexity of the sequence
[0280] 2) Lack of repetitive DNA motifs longer than 5 bases
[0281] 3) Lack of self-complementary sequences of more than 5 nucleotides
[0282] 4) Select the most flexible sequence in each design iteration cycle
[0283] According to the above design, use the mFold server ( http: / / unafold.rna.albany.edu / ) manually checked each sequence and modified it to minimize the formation of hairpins and generally complex secondary structures.
[0284] 2.2 Preparation of DNA template for rolling circle amplification
[0285] The following protocol was prepared for the preparation of a template suitable for the RCA reaction. Any circular DNA is a suitable template and different protocols are available and known in the art to treat dsDNA or ssDNA.
[0286] Examples of dsDNA we can circularize are: cfDNA, ctDNA, sheared genomic DNA, PCR amplicons.
[0287] Examples of ssDNA include: cDNA, viral DNA.
[0288] 2.3 dsDNA cyclization reaction
[0289] Here, the dephosphorylated dsDNA molecule is called an “insert”, which is linked to the two ends of the phosphorylated backbone to form a circular dsDNA product.
[0290] The reaction is carried out using both DNA ligase and restriction enzymes under appropriate buffer conditions.
[0291] Buffer conditions have been optimized to allow ligation, restriction digestion and PlasmiSafe treatment in a one-pot reaction without the need for intermediate DNA purification steps.
[0292] Considering that the backbone in the example has a SrfI half-site at the very end, the components of the reaction are as follows:
[0293] Buffer 1X
[0294] 50 mM potassium acetate
[0295] 20 mM Tris-acetate
[0296] 10 mM magnesium acetate
[0297] 100 μg / ml BSA
[0298] 1mM ATP
[0299] 10mM DTT
[0300] DNA and enzymes
[0301] Main chain + insert at a molar ratio of 3:1
[0302] 1 unit T4 DNA ligase
[0303] 1 unit Srf1
[0304] 1 unit HMGB1
[0305] H2O was added to a final volume of 20 to 50 μl (depending on the amount of DNA loaded), followed by incubation at 22°C for 1 h and subsequent heat inactivation at 65°C for 15 min.
[0306] The presence of restriction enzymes increases the overall yield of the reaction, avoids the accumulation of backbone concatemers, and prevents insert concatemerization by preventive dephosphorylation. HMGB1 (high mobility group protein 1) is used to promote the bending of short DNA, thereby improving circularization efficiency.
[0307] The most abundant product of the above reaction is a circular dsDNA containing a backbone and an insert.
[0308] Removal of linear DNA
[0309] To remove residual linear dsDNA, our templates were treated with 1 μl of Plasmid-Safe DNase at 37 °C for 15 min and then heat-inactivated at 70 °C for 30 min.
[0310] 3. Materials and methods for detection of fusion genes
[0311] Based on RNA extracted from human cells, we developed a protocol using circularization and rolling circle amplification to detect fusion genes. In this case, ssDNA (as opposed to dsDNA) was used as input for the circularization and amplification reactions. This protocol can be extended to sequence any RNA of interest.
[0312] The first part of the protocol involves standard procedures of "RNA extraction" and "cDNA amplification", e.g., Trizol-based RNA isolation followed by polyT-primed cDNA synthesis using reverse transcriptase.
[0313] After digestion of the RNA from the RNA-DNA hybrid, we obtain linear ssDNA. At this point, we use ssDNA ligase to self-circularize the input DNA.
[0314] Different ssDNA ligases are available on the market. We used Circ Ligase II (Epicentre) for proof-of-principle experiments. Circular ssDNA obtained following the supplier's protocol has been successfully used as a template for RCA reactions using specific primers to direct the amplification of the fusion gene of interest.
[0315] The following protocol details all passaging just after RNA isolation.
[0316] 3.1 Removal of residual DNA
[0317] Buffers, enzymes and inactivators were purchased from Thermo Fisher (TURBO DNase Kit).
[0318] In a 0.5 ml tube mix:
[0319] 10x reaction buffer 1 μl
[0320] Extracted RNA 1 μg
[0321] TURBO DNase 0.5 μl
[0322] HO to 10 μl final volume
[0323] Next the solution was mixed and incubated at 37°C for 30 min. 2 μl of inactivator was added to inactivate the solution. Mixing was continued for 5 min.
[0324] 3.2 cDNA synthesis
[0325] We used Invitrogen's Super Script II kit, but any other cDNA transcription kit can be used instead.
[0326] Add to the previous reaction:
[0327] Primer 2mM (random hexamer or specific) 1μl primer is phosphorylated
[0328] dNTP (10 mM each) 1 μl
[0329] Incubate at 65°C for 5 min and then place on ice for 5 min. During this step the primers are annealed to the template.
[0330] Next, add:
[0331] 5x First Strand Buffer 4 μl
[0332] 100mM DTT 1μl
[0333] After incubation at 42° C. for 2 min, 1 μl of SSII enzyme was added and incubated for 45 min at 42° C. Finally, we inactivated the reaction at 70° C. for 15 min.
[0334] 3.3 RNA Removal
[0335] RNaseH was purchased from Thermo Fisher.
[0336] To the previous reaction, 1 μl of RNase H enzyme was added and incubated at 37°C for 20 min, followed by heat inactivation at 70°C for 10 min.
[0337] 3.4 Chelation of divalent metal ions
[0338] We added this step to reduce the free Mg 2+ concentrations that would inhibit ssCirc ligase II used in the next reaction.
[0339] To complex all free Mg2+, add to the previous reaction:
[0340] 50mM EDTA stock (0,9EDTA + 50ml H2O) 2μl
[0341] 3.5 ssDNA cyclization
[0342] For this reaction we used the ssDNACirc Ligase II Kit from Epicentre.
[0343] 10x reaction buffer 2 μl
[0344] MnCl2 (Manganese (Mn) should not be confused with magnesium (Mg)) 1μl
[0345] Betaine 4μl
[0346] Circ Ligase II 1 μl
[0347] ss-cDNA 10p mole
[0348] HO to a final volume of 20 μl
[0349] Incubate at 60°C for 1-2 hours, then heat inactivate the reaction at 70°C for 10 min.
[0350] At this point, the reaction is treated with Plasmid Safe (optional) and used as a template for the RCA reaction (step below).
[0351] 3.6 Primer annealing
[0352] Depending on the situation, we use random primers or backbone-specific primers or target-specific primers.
[0353] In this step, the template DNA may also be single-stranded circular DNA, as in the case of self-circularized cDNA.
[0354] The primers have two 3' terminal phosphorothioate (PTO) modified nucleotides that are resistant to the 3'→5' exonuclease activity of proofreading DNA polymerases such as phi29 DNA polymerase. They also have a 5'-hydroxyl terminus and a 3'-hydroxyl terminus.
[0355] If the circular DNA is in water, such as in the case of a previous purification, add 11% of the volume of 10X Annealing Buffer.
[0356] 10X Annealing Buffer
[0357] 100 mM Tris, pH 7.5-8.0
[0358] 500mM NaCl
[0359] 10mM EDTA
[0360] Concentrated primers (50-100 μM) were added to the reaction to a final concentration of 5 μM.
[0361] The reaction temperature was adjusted to 98°C and then slowly cooled to room temperature.
[0362] Rolling circle amplification
[0363] For a 50 μl reaction, calculate the following volumes.
[0364] While the template is in annealing buffer, take 20 μl of template and add to it:
[0365] 5 μl Phi29 buffer (10x)
[0366] 1 μl BSA
[0367] 1μl dNTPs (10μM)
[0368] 2 μl pyrophosphatase
[0369] 1 μl Phi29 DNA polymerase
[0370] H2O to 50 μl
[0371] While the template is in circularization buffer, take 46 μl of template and add to it:
[0372] 0.5 μl Uracil-DNA glycosidase (for removing any deaminated cytosine from DNA)
[0373] 0.5 μl formamidopyrimidine-DNA glycosidase (for removal of 8-oxo-guanine products)
[0374] Reaction conditions
[0375] Depending on the amount of input DNA (template), the reaction is run at:
[0376] >>3h@30℃, if the template is 10-50ng
[0377] >>6h@30℃, if the template is 5-10ng
[0378] >>12h@30℃, if the template is 0.5-5ng
[0379] 4. Materials and Methods for Targeted Cyclomics
[0380] To enable ultra-precise targeted sequencing of any double-stranded DNA molecule, we designed a workflow based on the existing molecular inversion probe (MIP) technology. The uniqueness of this design lies in the MIP capture backbone (minimizing backbone size, adding unique molecular barcodes, probe specificity and distance), and the combination of the assay and rolling circle amplification.
[0381] 4.1 Probe Generation
[0382] Amplify oligonucleotides away from the array using PCR (MIP precursor): (2.5 hours)
[0383] 1. Array-extracted MIP precursor oligonucleotides (mixture of 100-mers obtained from Agilent) were dissolved in Tris-EDTA buffer between pH 8 and 0.1% to a final concentration of 100 nM.
[0384] 2. Prepare 400 μl of the following PCR mixture in a 1.5 ml centrifuge tube.
[0385] Reagents Volume (μl) Final concentration 2x iProof HF PCR Masterbatch Mix (Biorad) 200 1x Oligo_Fwd_Amp primer (100 μM) 2 500nM Oligo_Rev_Amp primer (100 μM) 2 500nM SYBRGreen I 100x(Invitrogen)** 1 0.2X Template (100 nM in 0.1% Tween) 1 250pM <![CDATA[H2O]]> 194
[0386] It was divided into 8 x 50 μl reactions in 0.2 ml PCR tubes. One PCR preparation yielded approximately 1.5 μg of amplified DNA.
[0387] 3. Use the following PCR cycling program on a real-time thermal cycler such as the Biorad MJ Mini.
[0388] 1) 98℃ for 30 seconds
[0389] 2) 98℃ for 10 seconds
[0390] 3) 60℃ for 30 seconds
[0391] 4) 72℃ for 30 seconds (plate reading)
[0392] 5) Repeat steps 2 to 4 for 25 cycles
[0393] 6) 4℃ unlimited time
[0394] 4. Combine and clean the PCR reactions on a column using the QIAquick PCR Purification Kit according to the manufacturer's instructions. Elute with 90 μl of elution buffer.
[0395] 5. Quantify 1 μl of amplified DNA using the Qubit High Sensitivity dsDNA Assay Kit. Capture exons using a molecular inversion probe.
[0396] 6. Analyze 1 μl of amplified DNA on a 6% TBE PAGE gel (Invitrogen) to verify amplification. When the primer was increased by 10 bp, the product showed a single band at 110 bp.
[0397] The PCR product was digested with a nicking restriction endonuclease to generate a 70-mer MIP (7.5 hours):
[0398] 1. Add 10 μl of NEB-2 (10x) and 5 μl of Nt.AlwI (10 U / μl; NEB) to 85 μl of PCR product (total volume 100 μl).
[0399] 2. Mix and divide into two test tubes, 50 μl in each test tube. Incubate at 37℃ for 3h and then at 80℃ for 20min in a thermal cycler.
[0400] 3. The temperature was lowered to 65°C and maintained for at least 1 min. 2.5 μl of Nb.BsrDI (2 U / μl; NEB) was added to each 50 μl reaction.
[0401] 4. Place at 65°C for 3 hours, then at 80°C for 20 minutes.
[0402] 5. Purify two 50 μl digestion reactions on one column using the QIAquick Nucleotide Removal Kit. Elute each column with 30 μl of elution buffer. We observed a yield of 80-90% for this step.
[0403] Quantify available probe using denaturing gel (2 hours):
[0404] 1. Accurate quantification of the available MIPs in the cleavage probe mix is important because it determines how much probe mix to add to the capture reaction.
[0405] 2. Prepare two-fold dilutions of NEB 100 bp DNA ladder (the dilutions we used ranged from 500 ng to 62 ng).
[0406] 3. Mix 2x TBE-Urea sample buffer (Invitrogen) with 1 μl of the enzyme-cleaved probe and the above dilution solution.
[0407] 4. We denatured the DNA by heating it to 95°C for 5 minutes and then immediately transferring it to ice.
[0408] 5. The samples were run on precast 6% TBE-urea denaturing PAGE gels (Invitrogen) at 160 V for 1 h.
[0409] 6. Quantify the amount of available MIP in the digestion mix by comparing the ladder dilution and the intensity of the 70 bp band. We use this MIP concentration when determining the volume of probe mix to add to the capture reaction.
[0410] 4.2 Exon capture using molecular inversion probes
[0411] Hybridization of probe to genomic DNA (37 hours):
[0412] 1. For each sample to be captured, we added the following reagents in a 0.2 ml PCR tube. The final capture reaction volume was 25 μl. Since there was no size selection for the 70 bp MIP, the volume of probe mix to be added was based on the concentration of the available MIP.
[0413]
[0414] 2. Denature at 95°C for 10 minutes.
[0415] 3. Incubate at 60°C for at least 36 h to hybridize the MIP to the gDNA.
[0416] Circularization of captured exons (1 day):
[0417] 1. Prepare the mixture of ligase and polymerase to be added to each capture reaction:
[0418]
[0419] This mixture was placed on ice and kept cool before adding 4.7 μl to the capture reaction.
[0420] 1. Incubate at 60°C for an additional 24 hours to allow gap filling and ligation to circularize the capture region.
[0421] Exonuclease selected for circularization of the product: (1 hour)
[0422] 1. Prepare the exonuclease mix to be added to each capture reaction to remove uncaptured gDNA, excess probe, and blocking oligonucleotide:
[0423] Reagents Volume of each sample (μl) Final concentration in the reaction Exo I 20U / μl 2 1.7U / μl Exo III 100U / μl 2 8.3U / μl
[0424] 2. Reduce the temperature of the capture reaction to 37°C and incubate for at least 1 minute before adding 4 μl of the exonuclease mix.
[0425] 3. Incubate at 37°C for 15 minutes.
[0426] 4. Inactivate the exonuclease by heating the reaction at 95°C for 2 minutes.
[0427] 5. Use 100 ng of the reaction product as a template for rolling circle amplification.
[0428] 4.3 Rolling circle amplification
[0429] 1. Vacuum dry 20 μl 2X Annealing Buffer to 1 μl.
[0430] 2. Add 40 μl circular DNA (approximately 10 ng).
[0431] 3. Add 4 μl of 50 μM random primers.
[0432] 4. Incubate at 90°C for 5 minutes and then slowly cool to room temperature.
[0433] 5. Add the following reagents: 10 μl 10x Phi29 buffer, 2 μl 100x BSA, 2 μl 10 mM dNTPs, 2 μl Phi29 polymerase, 4 μl pyrophosphatase 0.1 U / μl, 40 μl water.
[0434] 6. Incubate at 30°C for 19 hours, then at 65°C for 10 minutes.
[0435] 7. Wash the reaction product with Ampure XP beads (0.4V).
[0436] The cleaned reaction products are used in any long-read sequencing protocol.
[0437] 5. List of designed and tested backbones
[0438] >BB1(199bp)
[0439] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTAGAGTACGACTGCGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC
[0440] >BB2(259bp)
[0441] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTCACCGGGTCTTCGAGA AGACCTGTTTAGAGTACGACTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC
[0442] >BB2_100(341)
[0443] GGGCATGCACAGATGTACACGTACGATCATGTACGTCACGCGAGTGCACGTCGTCATAGCTGTCGAGTACTGTACTGACTGTCTCGAGCCTCAGCGAGTATTTAAATCTACGTCACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATA TATATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTTAGAGTACGACTGCAAATGGCTCTAGAGGTACCCGTTACATAACTTACGCAGATGTGATCAGTGACTACGTGACACTGTACATCAGCACGATCGATGACTAGATGCTGCATGACATAGCCC
[0444] >BB3(514bp)
[0445] AACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTGAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGGTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGCAGACAAATGGCTCTAGAGGTACCCGTTACATAACTTA
[0446] >BBpX2(557bp)
[0447] GGGCATGCACAGATGTACACGAACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTGAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGA GAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTT GAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGGTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGA GTCGGTGCTTTTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTTTTAGCGCGTGCGCCAATTCTGCAGACAAATGGCTCTAGAGGTACCCGTTACATAACTTATAGATGCTGCATGACATAGCCC
[0448] 6. Additional details on the generation of BB2_100 and BBpX2
[0449] BB2 was optimized by inserting a flexible sequence using the BbsI cloning site present in BB2. An insert consisting of a 100 bp long DNA stretch obtained in silico design (see Section 7) was added and a BbsI restriction site was added at the very end for cloning purposes (bold in the sequence below). The complete insert (sense and antisense) was obtained by annealing two shorter oligonucleotides. The oligonucleotides were ordered as single stranded oligonucleotides from IDT DNA technologies. The positive and reverse strands were annealed and the annealed products, now dsDNA with sticky ends, were resolved on an agarose gel. Subsequently, the insert was cloned into BB2 by a Golden-Gate cloning reaction similar to the one described by (Ran et al. 2013).
[0450] The complete insert sequence and oligonucleotides used to generate it are as follows:
[0451] >Insert BB2_100
[0452] CACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATATATATATATGGATATGTATGGATATATATATATATGGATATGGATGTTT
[0453] >Sense Oligonucleotides
[0454] CACCATATATATGGATATATATATGGATATATATATATATGGATATATGGATATATATATATATATATGGATATGTATGGATATATATATATATGGATATGGAT
[0455] >Antisense Oligonucleotides
[0456] AAACATCCATATCCATATATATATATATCCATACATATCCATATATATATATATATATCCATATATCCATATATATATATATCCATATATATATCCATATATAT
[0457] BBpX2 was obtained by adding a SrfI-half-site (GGGC) and the rest of the universal primer sequence (underlined in the sequence below) to the very end of the PCR amplicon. BB3 was used as template for the PCR reaction.
[0458] >BBpX to BBpX2-F
[0459] GGGCATGCACAGATGTACACG aacgccagcaacgcggc
[0460] >BBpX to BBpX2-R
[0461] GGGCTATGTCATGCAGCATCTA taagttatgtaacgggtacctct
[0462] The above sequence has the template annealing part in lowercase and the flanking regions in uppercase. The SrfI-half-site is highlighted in orange, and the remaining uppercase sequence is part of a constant sequence present at the end of all backbones. The constant sequence is not essential for the backbones, but is useful to standardize their amplification during the production steps. The following primers are indeed able to amplify any backbone made so far:
[0463] >Universal SrfI-BB-F
[0464] GGGCATGCACAGATGTACACG
[0465] >Universal SrfI-BB-R
[0466] GGGCTATGTCATGCAGCATCTA
[0467] 7. List of Insertion Sequences
[0468] >Insert 17.1 (TP53, chr17:7576971-7577132)
[0469] TAACTGCACCCTTGGTCTCCTCCACCGCTTCTTGTCCTGCTTGCTTACCTCGCTTAGTGCTCCCTGGGGGCAGCTCGTGGTGAGGCTCCCCTTTCTTGCGGAGATTCTCTTCCTCTGTGCGCCGGTCTCCCAGGACAGGCACAAACACGCACCTCAAAG
[0470] >Insert 17.2 (TP53, chr17:7578161-7578394)
[0471] CAGTTGCAAACCAGACCTCAGGCGGCTCATAGGGCACCACCACACTATGTCGAAAAGTGTTTCTGTCATCCAAATACTCCACACGCAAATTTCCTTCCACTCGGATAAGATGCTGAGGAGGGGCCAGACCTAAGAGCAATCAGTGAGGAATCAGAGGCCTGGGGACCCTGGGCAACCAGCCCTGTCGTCTCCAGCCCCAGCTGCTCACCATCGCTATCTGAGCAGCGCTCAT
[0472] 8. Computer design of flexible DNA sequences
[0473] To improve the flexibility of BB2, a 100 bp sequence was added to the flexible DNA sequence, which was specially designed by a simple genetic algorithm. The same method was used to design the entire backbone core sequence from scratch. Restriction sites, barcodes or primer sites of these backbone core sequences can be added as described elsewhere in this article.
[0474] The backbone core sequence was optimized based on an evolutionary selection algorithm, which optimized the following components of the sequence:
[0475] 1) High molecular flexibility
[0476] 2) High sequence entropy
[0477] 3) GC content between 30% and 60%, ideally close to 50%
[0478] 4) Lack of long self-complementary stretches
[0479] 5) Lack of long oligonucleotide polymers (NNNNNNN)
[0480] 6) Lack of repeated motifs (kmer)
[0481] Flexible computing
[0482] The flexibility of DNA at the torsion angle of the input sequence was calculated using the python implementation of the TwistFlex algorithm (http: / / margalit.huji.ac.il / TwistFlex / ) (Menconi et al. 2015). The flexibility of each single dinucleotide was calculated according to the following angle table:
[0483]
[0484] Subsequently, the evolutionary algorithm used for backbone optimization took into account the average flexibility of the entire sequence when selecting. The average flexibility of a DNA sequence is calculated as the sum of all dinucleotide angles divided by the total number of dinucleotides. Our backbone flexibility threshold is an average of 12.5 angles. Any sequence with an average flexibility below 12.5 angles is discarded.
[0485] Entropy calculation for determining sequence complexity
[0486] The Shannon entropy of a string is defined as the minimum average number of bits per symbol required to encode the string. The formula used to calculate the Shannon entropy is:
[0487]
[0488] Among them, p i is the probability of character number i appearing in the sequence. This calculation can also be performed at http: / / www.shannonentropy.netmark.pl / .
[0489] Use the following python code to implement the above formula:
[0490]
[0491] The minimum entropy value expected for our backbone core sequence is 1.5Sh. Each sequence with a smaller entropy value is discarded.
[0492] Self-complementarity
[0493] The selected backbone core sequences were filtered by the presence of self-complementary 8 base stretches. Backbones with 8 or more self-complementary consecutive bases on the same strand were discarded.
[0494] Lack of repetitive motifs (kmer)
[0495] Backbone core sequences containing motifs of 6 bases that were repeated more than twice in the sequence were filtered out.
[0496] Evolutionary algorithm for designing newer main chains (beyond BB2)
[0497] The newer backbone consists of flexible DNA and a pair of fixed sequences at the ends (the universal SrfI-BB F / R described in paragraph 6) that serve as primer annealing sites for backbone PCR amplification and for the addition of semi-restriction sites.
[0498] Like any genetic algorithm (GA), Cyclomics' GA consists of a main loop in which a library of sequences is scored and selected. The selected sequences are then used as input (parents) to generate new sequences (children). Both parents and children are grouped in a new library, ready for the next iteration. The pseudocode for this loop is as follows:
[0499] For each iteration:
[0500]
[0501]
[0502] The algorithm is implemented entirely in Python, and the pairing and mutation operators as well as the main loop are based on the literature (Hwang and Jang 2008; 1996; CoelloCoello and Lamont 2004; Lobo, Lima, and Michalewicz 2007). The pairwise operator acts on strings, sequences, and performs single swaps at random positions. The mutation operator adds random mutations to the parent or daughter sequence, which may include small deletions and duplications. The filtering step is used to trim the library of sequences with CG content imbalance, low sequence entropy, and unwanted kmer repeats before the selection step. The selection itself, simply collects the best sequences according to the flexibility score. The selected sequences are used as parent sequences to generate daughter sequences using the pairwise operator. To calculate sequence flexibility, an existing code (Menconi et al. 2015) was adapted to our purpose.
[0503] 9. Results
[0504] Primers 17.2-F (CAGTTGCAAACCAGACCTCA) and 17.2-R (ATGAGCGCTGCTCAGATAG) were used to perform PCR to obtain a PCR product of 234 bp in length covering the coding exon TP53 (chr17:7578161-7578394, GRCh37). According to standard procedures, the DEPCR product referred to as 17.2 was connected to pJET (Thermo Fisher). The connection product was transformed into E.coli Top10, and a strain of bacterium selected for collecting (clonal propagation) pJET-17.2 plasmid. The sequence of 17.2 was verified by Sanger sequencing, and it was found to be identical to the reference genome (GRCh37). Phosphorothioate (PTO)-modified primer 17.2-R (ATGAGCGCTGCTCAGATA*G*, where * indicates PTO modification) (5 μM) was annealed to 50 ng pJET-17.2 in the presence of 5 mM EDTA in a final volume of 20 μ L. The reaction mixture was heated to 95° C. for 5 minutes and then cooled to 4° C.
[0505] The 20 μL annealing reaction was supplemented with 0.2 μ inorganic pyrophosphatase (ThermoFisher), 10 μ Phi29 (NEB), 1 μ L 10 mM dNTPs, 1 μ L 100X BSA solution (NEB, 20 mg / mL) and 5 μ L Phi29 10X reaction buffer (NEB). The resulting reaction mixture was incubated at 30 ° C for 3 hours and then incubated at 65 ° C for 10 minutes. The amplified high molecular weight DNA was purified using Ampure beads (Agencourt) and then 1D nanopore libraries (Oxford nanopore Technologies, SQK-LSK108) were prepared. The resulting library was run on a MinION flow cell (FLOI-MIN 106, R9.4 chemistry) for 48 hours.
[0506] 10. DNA Quantification by Gel Densitometry
[0507] Agarose gel densitometry is a method for quantifying DNA by image analysis of gel bands by comparing pixel intensity 1) between the sequence ladder and the band of interest or 2) between the input band and the product band.
[0508] 10.1 Using the Ladder as a Reference
[0509] Given an image of an agarose gel containing a known amount of DNA ladder, we can use ImageJ software to estimate the brightness intensity of the bands. The correct circularization products are quantified and compared to the input to calculate the circularization efficiency. The area and average intensity of the bands on each gel are determined using the measurement function of ImageJ. The average intensity of the image background is determined to be as close as possible to the band in question and is subtracted from the band intensity. The resulting intensity is multiplied by the area of the band (called the level). To create a reference level, the ratio of the band corresponding to 400 base pairs in each image to the 50 bp DNA ladder (Thermo Fisher Scientific) is also calculated, which shows the intensity of 15 nanograms of DNA. To calculate the DNA content of each band in nanograms, the calculated level is divided by the reference level and multiplied by 15. The molar content of DNA is determined using the Promega DNA Conversion Tool dsDNA: μg to pmol. The circularization efficiency is calculated by dividing the correct product measured in moles by the inserted molar input and multiplying by 100%. To validate this approach, the DNA content of the other bands in the DNA ladder was calculated and compared to the predicted DNA content. The DNA content of each band was plotted as the predicted and calculated values ( Figure 6 ).
[0510] 10.2 Direct comparison of input and product bands
[0511] The optional process of estimating circularization efficiency by gel image analysis requires at least two lines on the gel, one for the input DNA and one for the product. Using ImageJ, we can estimate the ratio between the insert before the reaction (input) and the insert remaining after the reaction (unreacted).
[0512] Measure and compare the brightness of the bands within the yellow rectangle (input DNA on the left band and unreacted DNA on the right band). In this case, the ratio between input and unreacted is 66:33. We can conclude that 50% of the initial DNA has reacted ( Figure 7 A).
[0513] Next we compare the product bands ( Figure 7 B).
[0514] We know that the top band represents the desired product, while the bottom band represents the uncircularized product. From the ratio of these two bands, we can establish the circularization efficiency, which is defined as the amount of input DNA that is correctly circularized to the final product. If A is the ratio of the inputs to the reaction, and B is the ratio of the correct products, then the efficiency is A*B. In this case, 50%*50%=25%.
[0515] References
[0516] Thomas.1996.Evolutionary Algorithms in Theory and Practice:Evolution Strategies,Evolutionary Programming,Genetic Algorithms.OxfordUniversity Press on Demand.
[0517] CoelloCoello,Carlos A.,and Gary B.Lamont.2004.Applications ofMulti-Objective Evolutionary Algorithms.World Scientific.
[0518] Hwang,Gi-Hyun,and Won-Tae Jang.2008.“An Adaptive EvolutionaryAlgorithm Combining Evolution Strategy and Genetic Algorithm(Application ofFuzzy Power System Stabilizer).”In Advances in Evolutionary Algorithms.
[0519] Lobo,F.J.,Cláudio F.Lima,and Zbigniew Michalewicz.2007.ParameterSetting in Evolutionary Algorithms.Springer Science&Business Media.
[0520] Menconi,Giulia,Andrea Bedini,Roberto Barale,and IsabellaSbrana.2015.“Global Mapping of DNA Conformational Flexibility onSaccharomyces Cerevisiae.”PLoS Computational Biology 11(4):e1004136.
[0521] Ran, F. Ann, Patrick D. Hsu, Jason Wright, Vineeta Agarwala, David A. Scott, and Feng Zhang. 2013. “Genome Engineering Using the CRISPR-Cas9 System.” Nature Protocols 8(11): 2281-2308.
[0522] Example 2
[0523] Sequencing of tandem DNA molecules
[0524] We cloned the PCR product covering site 17:7578265 of the TP53 gene into pJET (Materials and Methods, Section 9). After bacterial transformation, individual colonies were picked and plasmid DNA was isolated to confirm the presence of the TP53 insertion (data not shown). Next, we performed rolling circle amplification (RCA) of the isolated plasmid using phi29 polymerase and random hexamer primers (Materials and Methods, Section 9). We obtained a high molecular weight RCA product of >20 kb as estimated by gel electrophoresis ( Figure 8 A). This product was used as input for a 1D library preparation, which was used for sequencing on an Oxford Nanopore Technologies (ONT) MinION instrument, and the library was sequenced for 48 hours according to the manufacturer's instructions. This sample generated a total of 16,248 sequencing reads with an average read length of 5.7 kb ( Figure 8 B), and 2083 reads longer than 10 kb. Using LAST( et al., 2011), Nanopore / MinION sequencing reads were mapped to the human reference genome (GRCh37, amplified using pJET sequences). A subset of 2083 reads larger than 10 kb were examined to determine the configuration of replacement backbone (pJET) and insertion (TP53 fragment), referred to as BI, that might be expected from the circularized template used as input. We observed that all 2083 reads had multiple copies of BI. For all reads longer than 10 kb we calculated a “pattern score”, a number between 1 and 100 representing the regularity of BI duplication, calculated as BI / ((B+I) / 2)*100, where BI is the amount of pJET-17.2 fragment in one Nanopore read, B is the amount of pJET fragment in one Nanopore read, and I is the amount of 17.2 fragment in one Nanopore read. Most long reads had a pattern score of 100, indicating correctness of the RCA product, i.e. duplication of BI units ( Figure 8 C). We used these reads to extract the insert sequence (17.2-TP53 fragment) and subsequently aligned each read to the insertion using Muscle (Edgar 2004a, [b] 2004) ( Figure 8 D). For each read, we applied the aligned 17.2 fragments to a majority voting scheme to obtain a consensus sequence. The consensus sequence was compared to the reference sequence to determine accuracy as a function of BI copy number in the read ( Figure 8 E). This experiment proved the concept of obtaining accurate consensus sequence reads based on nanopore sequencing of multiple copies of DNA molecules (Li et al. 2016). In the next step, the backbone design was improved to reduce the sequencing throughput of the backbone (pJET-3kb in the above experiment).
[0525] Optimization of linear dsDNA backbone
[0526] Main chain design principle
[0527] As a first step to optimize the capture and circularization of short DNA molecules, we tested different designs of the backbone sequence that mediates the process. We compared three parameters: the length of the backbone (longer DNA molecules (backbones) are thought to be easier to circularize, but this leads to a waste of sequencing information because the majority of each read will consist of the backbone sequence). Second, DNA molecules shorter than about 200 bp are thought to be difficult to circularize due to their relative rigidity (Shore, Langowski, and Baldwin 1981). Third, the flexibility of a given DNA molecule also depends on its base composition and sequence.
[0528] The first generation backbones were BB1 and BB2 (sequences in Materials and Methods, Section 5), which were designed following the general principles highlighted in Materials and Methods, Section 2. These backbones were designed with the intent to serve as basic building blocks that we could improve upon. These backbones include a combination of several elements that aid in the capture of DNA molecules and subsequent amplification and sequencing, such as restriction enzyme sites, barcode sequences, and / or nickase sites. We used BB1 to detect circularization, but it was not very efficient and was not taken forward for further testing.
[0529] Shore, Langowski and Baldwin 1981 suggested that circularization of short DNA molecules with blunt ends may be suboptimal if efficiency is required. Next, we generated a longer backbone that would allow more efficient circularization in a short period of time (~1 h). The resulting backbone was named BB3 and was a 514 bp long dsDNA fragment generated by PCR amplification of a portion of plasmid pX330 (Materials and Methods, Sections 5 and 6). BB3 does not contain restriction enzyme sites at the ends, and it was only used to test different ligation conditions (see below). We also generated a modified version of BB3 by adding an SrfI site, resulting in BBpX2.
[0530] We used the values of the free energy per base pair (Breslauer et al. 1986) and the deviations in torsion angles (degrees) (Sarai et al. 1989) to calculate the flexibility of any given DNA sequence. We used a genetic algorithm to generate a set of sequences that were selected for high flexibility, short length, and to optimize other parameters (such as GC content, presence of repeating motifs, and self-complementarity of the sequences). A more detailed description of the genetic optimization algorithm and further details of the backbone structure are given in Materials and Methods, Section 8. In order to improve the efficiency of backbone circularization while keeping the sequences short, we designed several stretches of flexible DNA by computer. We used this stretch to increase the flexibility of the backbone. Backbone 2 (BB2) was modified by adding a 100 bp long flexible sequence in the middle of its sequence. The resulting backbone (BB2_100) was 341 bp long and included a SrfI semi-restriction site at its extreme end.
[0531] Measuring the reaction products of the BB2 backbone and the BB3 backbone in the cyclization reaction
[0532] We tested BB2 and BB3 as inserts in the circularization reaction together with the PCR product, 17.2 (Materials and Methods, Sections 2, 5, and 7). The first experiment was performed to determine the best reaction conditions to obtain the best circularized product. Reactions were performed with and without the addition of Plasmid Safe DNase to obtain a clear view of the linear and circularized reaction products ( Fig. 9 A). The results of BB2's cyclization efficiency are consistent, but the cyclization product is not very abundant. The cyclization reaction with BB3 gave a visible cyclic product ( Fig. 9 A), whereas no clear reaction product was visible with BB2. However, running the entire reaction mixture on a gel allowed for the observation of a weak band after Plasmid Safe digestion, indicating that the correct cyclization product consisted of BB2 and insert 17.2 ( Fig. 9 B).
[0533] Effects of different backbone-insert ratios on cyclization efficiency
[0534] Next, we evaluated the effect of different backbone-insert ratios on the specificity and efficiency of circular backbone-insert product formation, which was our goal. Therefore, we performed circularization experiments using BB3 and simultaneously used the 234 bp PCR product (17.2, Materials and Methods, Section 7) ( Fig.10 ).
[0535] The best circularization efficiency was obtained with a 3:1 molar ratio of backbone to insert. This setting was maintained for further characterization. Note that the strategy used here is very different from that used for standard plasmid-based cloning. In standard cloning, plasmids are often dephosphorylated to avoid self-circularization and excess phosphorylated insert is added to the reaction. For the Cyclomics technique, the backbone is phosphorylated and hyperphosphorylated, while the insert is dephosphorylated. We found that this avoids multimerization of ligation-dependent targets and improves backbone-target ligation efficiency.
[0536] Effect of flexible extension on main chain cyclization
[0537] In subsequent experiments, we tested the effect of adding flexible DNA extensions on backbone circularization. We therefore compared the effects of BB2 with BB2_100 (see above) and BB3 ( Fig.11 ) circularization efficiency. The flexible region in BB2_100 is rich in TA repeats but is still complex enough to be clearly mapped to the reference sequence. In this test, we also evaluated the effects of HMGB1 and SrfI on the circularization reaction of BB2_100. HMGB1 is a DNA bending protein known to potentially improve circularization (Belgrano et al. 2013). We observed that the circularization efficiency of BB2_100 was improved compared to BB2, especially when considering a backbone:insert ratio of 3:1. Therefore, we conclude that the backbone design can be optimized by adding flexible DNA extensions to improve circularization efficiency. Circularization of BB2_100 and 17.2 overnight resulted in a higher circularization efficiency, estimated to be approximately 26% ( Fig.12 ), which indicates that better reaction performance can be obtained by adjusting both the backbone design and the reaction conditions.
[0538] Effect of SrfI on the formation of main-chain cyclization products
[0539] A necessary part of our backbones is the presence of separate restriction sites at the very ends. If the backbone self-circularizes without an insertion, the intact restriction site is reconstituted, making the backbone susceptible to specific nucleases. In the example below, a semi-restriction site for SrfI (GCCC|GGGC) was added to the end of BB3 to generate a new backbone we call BBpX2, and SrfI nuclease and T4 ligase were added to the reaction mixture. SrfI has the advantage of being able to recognize sites that are 8 bases long, while most commercially available alternatives only recognize sites that are 6 bases long. Other sequences we are evaluating are PmeI (GTTT|AAAC) and SweI (ATTT|AAAT). If the ligation reaction is performed in the presence of Srf1, any backbone that self-circularizes will be susceptible to restriction enzyme cleavage, returning to the original linear form.
[0540] By comparing the first two lanes, the effect of the restriction enzyme can be clearly seen ( Fig.13 ). When SrfI is present, the linear backbone is maintained (thick bands) and the entire reaction results in very few side products. In the absence of SrfI (first lane), most of the backbone is wasted in the formation of several side products. The impact of SrfI can be further evaluated by the effect of Plasmid Safe DNase treatment (last two lanes), which results in the degradation of linear DNA. If SrfI is added to the reaction (last lane), only the expected products are generated, whereas, if SrfI is not added (third lane), many unwanted circular side products are generated.
[0541] Dephosphorylation of insertion
[0542] To avoid self-polymerization of the inserts, we performed enzymatic dephosphorylation using Antarctic phosphatase, an enzyme that ensures high reactivity at low temperatures and complete inactivation at 65 °C in just 5 min.
[0543] Barcode strategy
[0544] Molecular barcoding is a strategy to label individual DNA molecules in order to classify sequencing results of DNA molecules. Barcoding can be used to classify sequencing reads by sample (bioinformatically), allowing multiple samples to be combined in a single sequencing experiment (Wong, Jin, and Moqtaderi 2013). In this case, only a limited number of unique barcodes are used, one per sample.
[0545] Additionally, barcodes can be used to individually label each DNA molecule, and such barcodes are often referred to as unique molecular identifiers (UMIs). In this case, a large number of unique barcodes / UMIs (random sequences) are used to minimize the chance that any two unrelated sequences will have the same barcode. UMIs can be used to obtain absolute quantification of individual sequences (Kivioja et al. 2011).
[0546] Another application of UMIs is the detection and quantification of low-frequency mutations (Kou et al. 2016), for example in cancer samples. This involves labeling individual DNA molecules followed by PCR amplification and deep sequencing. Sequence reads can then be grouped by UMI sequence, and possible mutations can be detected and identified from sequencing errors. Newman et al. (Newman et al. 2016) outlined the superior application of UMIs in ctDNA mutation detection.
[0547] We envision designing backbone sequences using sample-specific barcodes and UMIs ( Fig.14 ). This strategy enables the combined sequencing of multiple independent samples, and enhanced mutation detection capabilities. Sample-specific barcodes are 5-20 nucleotides in length and can be placed anywhere in the backbone sequence, provided that they do not affect the flexibility of the backbone (and thus the ligation efficiency). A random chain of 5-20 nucleotides representing the UMI will also be added to the backbone to label individual DNA molecules. UMIs can be used to improve mutation detection by requiring at least two or more different molecules with a mutation, i.e., both molecules should have a unique UMI.
[0548] Rolling circle amplification of circularized DNA molecules
[0549] The circular DNA product obtained by the circularization reaction of the backbone and insert can be used as a template for the generation of concatemers via rolling circle amplification (RCA). We used different sources (including cfDNA, PCR amplicons, plasmids, and cDNA ( Fig.15 )) of DNA (insert) and used random hexamer primers to test RCA.
[0550] Site-directed RCA
[0551] In addition to the typical RCA reaction, which involves random hexamers used to initiate amplification, we designed the use of specific primers to direct amplification to the region of interest, an approach known as site-directed RCA. This approach may be useful in situations where only specific genes need to be sequenced rather than the entire genome. The current way to implement this approach is to enrich the gene of interest by PCR (Dowthwaite and Pickford 2015). However, it is well known that PCR amplification increases errors in the amplicon (Shuldiner, Nirula, and Roth 1989); (Diaz-Cano 2001), and even a single amplification error that occurs early in the PCR reaction can lead to biased final results (Diaz-Cano 2001; Quach, Goodman, and Shibata 2004); (Arbeithuber, Makova, and Tiemann-Boege 2016).
[0552] To test whether we could obtain site-specific enrichment of a target region without using PCR, we combined the cycloomics assay with site-directed RCA. Briefly, we cloned two different regions of the TP53 gene, 17.1 and 17.2 (Materials and Methods, Section 7) into the pJET vector. A modified pJET vector was used as template for the RCA reaction at a 1:1 molar ratio, where specific primers targeting 17.2 instead of 17.1 (17.2R, Materials and Methods, Section 9) were used instead of random hexamers. The reaction products were sequenced using a nanopore MinION instrument, and the number of reads containing 17.1 and 17.2 were compared ( Fig.16 ). We observed that sequencing reads containing 17.2 were more abundant at 14x accession compared to reads containing 17.1, suggesting that target-selective RCA can be achieved by specific primers.
[0553] One-pot reaction design
[0554] To enhance the use of cycling omics technologies, we focused on developing a streamlined experimental process. Therefore, we limited as much as possible the time-consuming and labor-intensive steps of DNA purification, concentration, and gel electrophoresis. To this end, we designed a protocol consisting of three simple sequential steps that can be performed in a single tube to limit the need for purification or buffer exchange.
[0555] The steps are: 1) circularization, 2) removal of linear DNA and 3) rolling circle amplification ( Fig.17 ).
[0556] The first reaction of the cyclomics protocol involves the insert DNA (I) and the backbone (BB) being mixed together in the presence of T4 DNA ligase and the restriction enzyme SrfI. The mixture is left at room temperature for 1 to 4 hours and then the enzyme is heat-inactivated at 70°C for 30 minutes. The second step of the cyclomics protocol is the addition of Plasmid Safe enzyme and its buffer and 1 mM ATP to the reaction mixture. The mixture is incubated at 37°C for 30 minutes and then inactivated. Before continuing with rolling circle amplification (reaction 3), RCA-primers are added to the mixture and a rapid annealing step is performed by heating the reaction to 98°C for 5 minutes. After the mixture has cooled at room temperature, Phi29, pyrophosphatase and the other components of the RCA reaction are added. The reaction is then incubated at 30°C for at least 3 hours.
[0557] Consensus sequence calling
[0558] To detect mutations from long concatenated reads, a consensus sequence of the target sequence was generated ( Fig.18 To this end, long reads were divided into backbone and target sequences based on the LAST decomposition-reads mapped to the reference genome ( et al. 2011). Target sequences were passed to GATK UnifiedGenotyper for variant calling (DePristo et al. 2011). Causal filtering was applied to optimize sensitivity and specificity based on variant confidence scores.
[0559] Examples of cyclomics applications
[0560] Targeted sequencing of TP53 mutations in genomic DNA from ovarian cancer
[0561] We tested the cyclomics approach in tumor biopsies with three known TP53 mutations (chr17:7578265, A->T, hg19) with variable frequencies (1%, 9%, 14%) as previously evaluated using short-read targeted IonTorrent sequencing (Hoogstraat et al. 2014). Briefly, we performed PCR on the target loci and ligated the resulting products to a backbone specifically designed and optimized to promote efficient capture of short DNA products. Subsequently, the ligated products were amplified and concatenated to form long DNA molecules with duplicate copies of the target / insert and the backbone. The long DNA molecules were sequenced in a few hours using the nanopore MinION instrument (1D ligation-based library preparation). We obtained a total of 206,048 sequence reads for all three samples, which was compared with LAST ( et al. 2011) and a custom algorithm for consensus calling to process ( Fig.18Next, we estimated mutation frequencies from consensus sequence reads and observed TP53 mutation frequencies of 0.5%, 7.6%, and 14%, respectively, providing proof of concept for using cyclo-omics technology to detect low-frequency somatic mutations in cancer DNA ( Fig.19 ).
[0562] References
[0563] -Arbeithuber, Barbara, Kateryna D. Makova, and Irene Tiemann-Boege. 2016. “Artifactual Mutations Resulting from DNA Lesions Limit Detection Levels in Ultrasensitive Sequencing Applications.” DNA Research: An International Journal for Rapid Publication of Reports on Genes and Genomes 23(6): 547-59.
[0564] -Belgrano, Fabricio S., Isabel C. de Abreu da Silva, Francisco M. Bastosde Oliveira, Marcelo R. Fantappié, and Ronaldo Mohana-Borges. 2013. "Role of the Acidic Tail of High Mobility Group Protein B1(HMGB1) in Protein Stability and DNA Bending." PloS One 8(11):e79572.
[0565] -Breslauer, KJ, R. Frank, H. Blocker, and LA Marky. 1986. "Predicting DNADuplex Stability from the Base Sequence." Proceedings of the National Academy of Sciences 83(11): 3746-50.
[0566] - DePristo, Mark A., Eric Banks, Ryan Poplin, Kiran V. Garimella, Jared R. Maguire, Christopher Hartl, Anthony A. Philippakis, et al. 2011. “A Framework for Variation Discovery and Genotyping Using next-Generation DNA Sequencing Data.” Nature Genetics 43(5): 491-98.
[0567] - Diaz-Cano, Salvador J. 2001. “Are PCR Artifacts in Microdissected Samples Preventable?” Human Pathology 32(12): 1415.
[0568] - Dowthwaite, Gary, and Jo Pickford. 2015. “PCR-Based DNA Enrichment Enhances Detection of Mutations in Oncology.” MLO: Medical Laboratory Observer 47(11): 18, 20.
[0569] - Edgar, Robert C. 2004a. “MUSCLE: Multiple Sequence Alignment with High Accuracy and High Throughput.” Nucleic Acids Research 32(5): 1792-97.
[0570] - ———. 2004b. “MUSCLE: A Multiple Sequence Alignment Method with Reduced Time and Space Complexity.” BMC Bioinformatics 5(August): 113.
[0571] -Hoogstraat, Marlous, Mirjam S. de Pagter, Geert A. Cirkel, Markus J. van Roosmalen, Timothy T. Harkins, Karen Duran, Jennifer Kreeftmeijer, et al. 2014. “Genomic and Transcriptomic Plasticity in Treatment-Naive Ovarian Cancer.” Genome Research 24(2): 200-211.
[0572] Szymon M., Raymond Wan, Kengo Sato, Paul Horton, and Martin C. Frith. 2011. “Adaptive Seeds Tame Genomic Sequence Comparison.” Genome Research 21(3): 487-93.
[0573] -Kivioja, Teemu, Anna Kasper Karlsson, Martin Bonke, Martin Enge, Sten Linnarsson, and Jussi Taipale. 2011. “Counting Absolute Numbers of Molecules Using Unique Molecular Identifiers.” Nature Methods 9(1): 72-74.
[0574] -Kou, Ruqin, Ham Lam, Hairong Duan, Li Ye, Narisra Jongkam, Weizhi Chen, Shifang Zhang, and Shihong Li. 2016. “Benefits and Challenges with Applying Unique Molecular Identifiers in Next Generation Sequencing to Detect Low Frequency Mutations.” PloS One 11(1): e0146638.
[0575] -Li, Chenhao, Kern Rei Chng, Esther Jia Hui Boey, Amanda Hui Qi Ng, Andreas Wilm, and Niranjan Nagarajan. 2016. “INC-Seq: Accurate Single Molecule Reads Using Nanopore Sequencing.” GigaScience 5(1): 34.
[0576] -Newman, Aaron M., Alexander F. Lovejoy, Daniel M. Klass, David M. Kurtz, Jacob J. Chabon, Florian Scherer, Henning Stehr, et al. 2016. “Integrated Digital Error Suppression for Improved Detection of Circulating Tumor DNA.” Nature Biotechnology 34(5): 547-55.
[0577] -Quach, Nancy, Myron F. Goodman, and Darryl Shibata. 2004. “In Vitro Mutation Artifacts after Formalin Fixation and Error Prone Translesion Synthesis during PCR.” BMC Clinical Pathology 4(1). doi:10.1186 / 1472-6890-4-1.
[0578] -Sarai, A., J. Mazur, R. Nussinov, and R. L. Jernigan. 1989. “Sequence Dependence of DNA Conformational Flexibility.” Biochemistry 28(19): 7842-49.
[0579] -Shore, D., J. Langowski, and RLBaldwin. 1981. "DNA Flexibility Studied by Covalent Closure of Short Fragments into Circles." Proceedings of the National Academy of Sciences of the United States of America 78(8): 4833-37.
[0580] -Shuldiner, Alan R., Ajay Nirula, and Jesse Roth. 1989. "Hybrid DNAArtifact from PCR of Closely Related Target Sequences." Nucleic Acids Research 17(11): 4409-4409.
[0581] -Wong, Koon Ho, Yi Jin, and Zarmik Moqtaderi. 2013. "Multiplex IlluminaSequencing Using DNA Barcoding." Current Protocols in Molecular Biology / Edited by Frederick M.Ausubel...[et Al.]Chapter 7: Unit 7.11.
[0582] Example 3
[0583] Materials and methods
[0584] Circularization and RCA amplification of short PCR oligonucleotides
[0585] Material
[0586]
[0587] Wizard SV Gel and PCR Cleanup System (Promega #A9282)
[0588] The backbone must be phosphorylated, either by PCR production using phosphorylated primers or by PNK using non-phosphorylated PCR products or synthetic DNA duplexes (T4 polynucleotide kinase).
[0589] Inserts must be amplified via PCR using non-phosphorylated primers or dephosphorylated using Antarctic phosphatase.
[0590] Both the insert and the backbone must be blunt ended. The preferred method is to use Phusion polymerase (which leaves amplicons with blunt ends).
[0591] Both the insert and the backbone must be buffer-free when purified by column or beads.
[0592] In cases where the PCR reaction used to produce BB or I produces more than one product, then gel purification of the desired product is necessary.
[0593] If the template used for I or BB amplification is circular (e.g., a plasmid), gel purification of the PCR product is necessary.
[0594] method
[0595] Cyclization
[0596] Reaction mixture (1X): (the molar ratio of BB:I should be 3:1)
[0597]
[0598] Prepare the above reaction mixture on ice and in a PCR tube.
[0599] Vortex and swirl.
[0600] Place in thermocycler and run the following program: (16°C x 10' >> 37°C x 10') x 8 >> 70°C x 20'.
[0601] Add 1 μl of SrfI and run the following program: (This step is to digest any residual BB-BB) 37°C x 15' >> 70°C x 20'.
[0602] The maximum amount of DNA recommended to be used in this reaction (taking both I and BB into account) is 400 ng in a 50 mol / l reaction. The BB:I ratio should not be altered.
[0603] Example ratio calculation: (len(X) = length of X in base pairs), where: len(I) = 130 bp; len(BB) = 245 bp; len(BB) / len(I) = 245 / 130 = 1.88; starting from 50 ng of I, then 50*1.88*3 = 282 ng of BB are needed to achieve a 3:1 ratio.
[0604] Linear DNA Removal
[0605] Remove 4 μl of the cyclization reaction as a negative control for the gel run later.
[0606] To the remaining cycle reaction (46 μl) add:
[0607] -ATP 10mM 6μl
[0608] -Plasmid-Safe Buffer 10X 6μl
[0609] - Plasmid-Safe enzyme 2 μl
[0610] Incubate at 37 °C for 30'.
[0611] Inactivate at 70°C for 30'.
[0612] The entire reaction (S) and negative control (C-) were run on a 1.7% agarose gel.
[0613] Agar purification of the band corresponding to circular BB-I ( Fig.25 ).
[0614] Elute twice with 30 μl of water.
[0615] Rolling circle amplification
[0616] To the purified cyclic BB-I (approximately 50 μl at this time) add:
[0617] - Annealing buffer (5X) 12μl
[0618] -Exo-Res.RND primer (500μM) 1μl
[0619] The solution was heated at 98 °C for 5' and then slowly cooled at room temperature.
[0620] Add to:
[0621]
[0622] Incubate the reaction at 30 °C for at least 3 h.
[0623] Inactivate at 70°C for 10'.
[0624] Run 5 μl on a 0.5% agarose gel.
[0625] Running the RCA reaction overnight will yield more product. However, it is unclear whether the quality of the concatemer will be affected.
[0626] Quality Check
[0627] The following procedure allows a rough estimate of the amount of BB-I and BB-only monomers in the RCA product. Using the restriction sites present in the backbone (BglII in the example below), the RCA product can be digested and the resulting banding pattern can be used to infer the exact content of the RCA product.
[0628] like Fig. 27 As shown in , BB200_4 (243 bp) and S1_WT (158 bp) were circularized and amplified by RCA. When digesting a concatemer made of BB-I, we expect a band around 400 bp, while if the concatemer is only BB, the resulting band should be around 250 bp. Concatemers formed by only I will not be digested, making the RCA band visible.
[0629] Library preparation
[0630] DNA purification:
[0631] - Add an equal volume of Dyna beads, mix gently, and incubate at room temperature for 5 minutes.
[0632] - Insert the tube into the magnetic rack and wait 5 minutes to allow the beads to aggregate on the walls.
[0633] - Remove buffer.
[0634] - Rinse gently with 700 μl of 70% ethanol.
[0635] - Remove the ethanol and repeat the rinse step one more time.
[0636] - Evaporate the residual ethanol.
[0637] - Remove the test tubes from the magnetic rack.
[0638] - Elute the DNA from the beads with 100 μl of ultrapure water.
[0639] Dissolving branched DNA:
[0640] - Add 4 μl of T7 endonuclease (NEB #M0302S).
[0641] - Incubate at 37°C for 1 h.
[0642] Library preparation:
[0643] - Continue with nanopore library preparation, which can be 1D ligation preparation or rapid preparation.
[0644] List of backbone and insert sequences used
[0645] -len = backbone length in base pairs
[0646] -mean_flex = the average value of the calculated DNA flexibility for all consecutive segments of 50 base pairs in the sequence
[0647] - max_flex = Maximum DNA flexibility calculated for a 50 base pair segment in the sequence
[0648] - Entropy = Shannon entropy of the DNA sequence
[0649] - GC% = percentage of GC bases in the backbone
[0650] >BB100_1(len: 143 mean_flex: 12.89 max_flex: 14.71 Entropy: 2.0 GC%: 48.25)
[0651] GGGCATGCACAGATGTACACGATTCCCAACACACCGTGCGGGCCATCGACCTATGCATACCGTACATATCATATATAAATCACATAATTTATTATACGTATGTCGCGCGGGTGGCTGTGGGTAGATGCTGCATGACATAGCCC
[0652] >BB100_2(len: 143 mean_flex: 13.29 max_flex: 14.95 Entropy: 1.96 GC%: 37.76)
[0653] GGGCATGCACAGATGTACACGCACTACATGCCAATGCCCAAGCAGTGCGCATATCACGTATCATATCTAATATATTATAATATTATGATAATGAGTATTTATTTAATTTGTTTGTGTGAGGTAGATGCTGCATGACATAGCCC
[0654] >BB100_3(len: 143 mean_flex: 12.78 max_flex: 14.1 Entropy: 1.95 GC%: 44.06)
[0655] GGGCATGCACAGATGTACACGCATTGGCCGTCTGTGCTGTCCATGGATCGTCTGATTGATATGATATCATATATTATAATTATACAGTAAGGTGATTGGGTATTGAGGGTTGTGTGGTTGGTAGATGCTGCATGACATAGCCC
[0656] >BB100_4(len: 145 mean_flex: 12.89 max_flex: 14.06 Entropy: 1.95 GC%: 44.14)
[0657] GGGCATGCACAGATGTACACGGTAGACATGCGAAGCGTGCGATGACAATCGATGTGGACATCATGCATATATATGTTGTATAATTAAACAAATATGTGTAGTGTGTGAGGTGGGTGTAGGAAGTAGATGCTGCATGACATAGCCC
[0658] >BB100_5(len:143 mean_flex:13.27 max_flex:14.34 entropy:1.9 GC%:37.76)
[0659] GGGCATGCACAGATGTACACGTTGTCATGGGAATTTGTGGTTATGAAATGAGTATGCGACGAATATGTATACATATATATTAAATTATAGAGTGATGTATGAGTTTGTGATGTGTGGTGTATAGATGCTGCATGACATAGCCC
[0660] >BB200_1(len:243 mean_flex:13.0 max_flex:14.9 entropy:1.99 GC%:44.86)
[0661] GGGCATGCACAGATGTACACGGCGGCGCAAGATGATGTGCCGAACCTGACATGGCATCGACTGGTATGGATCAATACTGATGCGATATCGATACCGGATAAATCATATATGCATAATATCACATTATATTAATTATAATACATCGGCGTACATATACACGTACGCATCATTTCACTATCTATCGGTACTATACGTAGTGCCGGTCTGTTGGCCGGGCGACATAGATGCTGCATGACATAGCCC
[0662] >BB200_2(len:244 mean_flex:13.15 max_flex:14.69 entropy:1.96 GC%:38.52)
[0663] GGGCATGCACAGATGTACACGTGACGCAACGATGATGTTAGCTATTTGTTCAATGACAAATCTGGTATGATCAATACCGATGCGATATTGATATCTGATAACTCATATATGTAGAATATCACATTATATTTATTATAATACATCGTCGAACATATACACAATGCATCTTATCTATACGTATCGGGATAGCGTTGGCATAGCACTGGATGGCATGACCCTCATTAGATGCTGCATGACATAGCCC
[0664] >BB200_3(len:244mean_flex:13.06max_flex:14.9 entropy:1.96GC%:39.75)
[0665] GGGCATGCACAGATGTACACGAGACCGCAAGATGATGTTCATTCTTGAACATGAGATCGGATGGGTATGGATCAATACCGATGCGATATGATAACTGATAAATCATATATCTATAATATCACATTATATTAATTATAATACAGGATCGTTACATGCATACACAATGTATACTATACGTATTCGGTAGTTAGTGTACGGTCGGAATGGAGGTGGTGGCGGTGATAGATGCTGCATGACATAGCCC
[0666] >BB200_4(len:243mean_flex:13.29max_flex:14.44 entropy:1.93GC%:34.57)
[0667] GGGCATGCACAGATGTACACGAATCCCGAAGATGTTGTCCATTCATTGAATATGAGATCTCATGGTATGATCAATATCGGATGCGATATTGATACTGATAAATCATATATGCATAATCTCACATTATATTTATTATAATAAATCATCGTAGATATACACAATGTGAATTGTATACAATGGATAGTATAACTATCCAATTTCTTTGAGCATTGGCCTTGGTGTAGATGCTGCATGACATAGCCC
[0668] >BB200_5(len:243 mean_flex:13.37 max_flex:14.52 entropy:1.94 GC%:35.8)
[0669] GGGCATGCACAGATGTACACGAATCCGTGAGATGACTATCTTATTTGTGACATTCATCGATCTGGATATGATCAATACCATGCGATATTGATTACTGATAAATCATATATGTAGAATATCACATTATATTAATTATAATAAATCGTCGTACATATACATCCACAATTAGCTATGTATACTATCTATAGAGATGGTGCATCATCGTACTCCACCATTCCCACTAGATGCTGCATGACATAGCCC
[0670] >BB300_1(len:348 mean_flex:13.12 max_flex:14.77 entropy:1.98 GC%:41.67)
[0671] GGGCATGCACAGATGTACACGCATAAGACCACAGGGTGCAAATCTGGATTGCGGCATGGATGATTCATCATCGTGGCATATTCGCTATGGATATATCCATCATAATACATTGATACGTCATGCGTATAATCGCATTATATGTCGATATTGGTCATAGGGATACATCCGTGTATACTATCGTATATGCGTGCAATGTAGCCATGTTAATCATGCTATAACCATAACATAAATATAATATATACAGATGGTGTATCTCTACTTATGTATGCTTGTATAGTAATGTCGATACTGATGGGTCTCCGGCCCACTACACCACCTGGCCGCTCTAGATGCTGCATGACATAGCCC
[0672] >BB300_2(len:343 mean_flex:13.26 max_flex:14.34 entropy:1.98 GC%:40.82)
[0673] GGGCATGCACAGATGTACACGGGCAATCCGCCAGGGTTCAAATATGGATATGTGATGATCGATTCAACATGCACATATGCACGATATCATATATTACTCCAGATGTCATCATCGTCGTGCGTATATGAGATATGTATTTATGCATATAATCCACCATACATGGTAGCGATATTATAGTGCGATTATGTGTATATGACTATCATGGCTATTGTTAATATATAAATCATAACCATACCACTTCCACGCCTGGTATGGCGTATAGTATAGAGATATTGTGTGATGCCCTATGTCGACCATGATGTGCCGTTGTACTGCCAATCCTAGATGCTGCATGACATAGCCC
[0674] >BB300_3(len:344 mean_flex:13.47 max_flex:14.8 entropy:1.95 GC%:36.34)
[0675] GGGCATGCACAGATGTACACGTATCCATGCAGCTTATTGTAACTAGCGCATGCACGTGGTGATTCATCACATCTATATATACGATATGATATATTACACATATTTGCATAGTATCATCCGGTGTGATATCATCCGATATGCTCATACTTATTCATTGGTAGCATTGCATTGATGGATCAATAGTTATTATGACATCATGGCATGTACAATTATAAATAATACAACATACATAAATATACTATACACATCGTGTATGTGTTATACAGATCTGTGTGATGTATGATAATGTAATGGCGTCGAACACCACAAGGCAGTCCTATAATAGATGCTGCATGACATAGCCC
[0676] >BB300_4(len:344 mean_flex:13.37 max_flex:14.57 entropy:1.94 GC%:37.5)
[0677] GGGCATGCACAGATGTACACGGTCCATTACAATCGAATCTATATCCCAATGTGTATCGATTATCACCACAATGACATAATACGATATCATATATTACTCCATATGCCTTACGTCAGATCGTTATATGAGATATGTATTCATGCATATGATATCCCACAGTACACGTCGTCTAATGCCATCATGAATGTATGACATATCTAGTCGATTATACATAATATAACATACCAATATAACAATATCTATACACATTTGATGGCGTATAGTATAAAGATATTGTGGCAATGCCCATACACCACTGACTGTCGCCGATCATTCCTACCACTAGATGCTGCATGACATAGCCC
[0678] >BB300_5(len:344 mean_flex:13.51 max_flex:14.89 entropy:1.91 GC%:33.43)
[0679] GGGCATGCACAGATGTACACGACCGACCGTGAAAGTGATTCAGAATGATGTGCATGAATGTTATCATGACATGATTTATGATGCACTGATATATGCATATTATAATATTGTACAATGTCGTATATACGACATATCTATACTATGAATTATGGCATCATGGACAATAGATGGTAAGGTATAGTACGATCTATATAGCATGTTGAAATGGGATATAAATTATCATAAACATACATACTTAACTAATATCAAGATGATATGTGTATGACATCAGAATGATAGTAGTAATGAGTATTGTCAGATGTATGTACGAATATCACACGATTAGATGCTGCATGACATAGCCC
[0680] >Insert S1 WT (TP53, chr17:7577450 - 7577649)
[0681] AGGCTGGGGCACAGCAGGCCAGTGTGCAGGGTGGCAAGTGGCTCCTGACCTGGAGTCTTCCAGTGTGATGATGGTGAGGATGGGCCTCCGGTTCATGCCGCCCATGCAGGAACTGTTACACATGTAGTTGTAGTGGATGGTGGTACAGTCAGAGCCAACCTAGGAGATAACACAGGCCCAAGATGAGGCCAGTGCGCCTT
[0682] >Insert 17.2 (TP53, chr17:7578161-7578394)
[0683] CAGTTGCAAACCAGACCTCAGGCGGCTCATAGGGCACCACCACACTATGTCGAAAAGTGTTTCTGTCATCCAAATACTCCACACGCAAATTTCCTTCCACTCGGATAAGATGCTGAGGAGGGGCCAGACCTAAGAGCAATCAGTGAGGAATCAGAGGCCTGGGGACCCTGGGCAACCAGCCCTGTCGTCTCCAGCCCCAGCTGCTCACCATCGCTATCTGAGCAGCGCTCAT
[0684] and Fig.24 Bioinformatics
[0685] Using Tombo's DNA model (Fasta->raw), both forward and reverse (https: / / github.com / nanoporetech / tombo), an expected reference signal is generated for every possible insertion (for every possible base pair at the target position). Expected signals for the backbone are also created, forward and reverse.
[0686] The expected backbone signal is mapped to the read using dynamic time warping (DTW). If the expected backbone signal overlaps when aligned with the read, the best result is picked and the suboptimal result is removed. The read is then cut into fragments based on the orientation of the appropriate backbone.
[0687] Subsequently, all possible expected insertion signals are mapped to the reads using DTW. Again, overlapping results are removed and only the best results are retained. The best match result (smallest DTW error) is retained for each read. The number of specific insertions (indicating a specific base at the target position) determines the most likely base read at the target position.
[0688] result
[0689] Cyclization efficiency of different main chains
[0690] In order to be able to experimentally evaluate the efficiency of different backbones in circularizing short DNA amplicons, a 234 bp PCR amplicon (insert 17.2) was ligated with backbones derived from three different backbone series (BB100_1 / 2 / 3 / 4 / 5, BB200_2 / 4 / 5 and BB300).
[0691] The backbone sequences and physical properties are reported below. The detailed protocols are disclosed in Materials and Methods.
[0692] Fig.21 (Left) The product of the cyclization reaction is shown.
[0693] After circularization, the reaction was supplemented with an enzyme mix (PlasmidSafeLucigen #E3101K) to digest the linear DNA. Fig.21 The residual product (circular DNA) is shown in (right).
[0694] The BB200 series showed the best efficiency so far. To further characterize the efficiency of BB200_2 / 4 / 5, the 3 backbones were ligated with the same amplicon in the absence of the restriction enzyme SrfI. The rationale behind this experiment is that the backbone ligation efficiency can be estimated by the amount of multimers formed in the reaction. Fig. 22 It can be seen that the connection efficiency of BB200_4 is significantly higher than that of BB200_2 and BB200_5.
[0695] The higher efficiency of BB200_4 ligation is reflected in more efficient cyclization and RCA product formation. Fig.23 In Figure 2, sequencing read counts from two independent experiments (blue and red) in which equimolar mixtures of BB200_2, BB200_4, and BB200_5 were used to generate concatemers are plotted. The sequencing results are consistent with previous experimental results, indicating that the vast majority of sequenced beads contained BB200_4.
[0696] New (optimized) barcode sequences that are better in terms of ligation efficiency
[0697] BB200_4 is the most efficient backbone tested so far in the cyclization reaction.
[0698] Coupling strand-specific mutation calling with the possibility of strand-specific rolling circle amplification
[0699] The cyclomics approach generates double-stranded DNA circles. One advantage of having double-stranded circles is that one of the strands can be preferentially used as a template for RCA, for example by using strand-specific primers to initiate the reaction, followed by a known process (https: / / www.sciencedirect.com / science / article / pii / S0042682212002814). In this way, the cyclomics approach enables selective amplification of the sense or antisense sequence of a given DNA sequence. This strand-specific amplification is not possible using smartbell methods, but it is of great benefit in obtaining accurate variant calls from nanopore sequencing data in an efficient manner.
[0700] exist Fig.24 In , we show example situations where the detection rates of correct bases are different when analyzing data from two different strands of a DNA molecule. The data are from an experiment where a 200 bp (insert S1 WT) long amplicon was circularized with BB200_4 and amplified as specified in the reported protocol.
[0701] Data analysis of sequencing results can determine the accuracy of base calls for each strand. Specifically, we noticed that C bases and A bases are often difficult to distinguish due to their similar raw signal intensities. However, the signal from T is very different from the signals from all other bases and is easily classified correctly. For example, if an A mutation in the forward strand is expected, sequencing the reverse strand will lead to clearer results because the A in the forward strand will be miscalled as a G. Therefore, specific enrichment of the reverse strand is advantageous in this scenario.
[0702] Fig.24 The data highlighted in the middle shows an example of the difference in identifying bases on the forward or reverse strand. Emphasize how the correct bases on the reverse strand data can be inferred by using a simple cutoff on the Y axis (Y < 0.3). The same approach does not work for the forward strand. Therefore, in this case, amplification and sequencing of both strands results in a waste of data and, more problematically, can lead to misleading mutation detection at specific positions with a high false positive rate. In contrast, strand-specific enrichment results in higher sensitivity (most of the reads will come from the best strand) and no false positive calls.
Claims
1. A method for preparing a double-stranded target DNA molecule for sequencing, comprising: - providing a double-stranded backbone DNA molecule comprising a 5' end and a 3' end, wherein the molar ratio of the backbone DNA molecule to the target DNA molecule ranges from 1:5 to 5:1, and the 5' end and the 3' end of the backbone DNA molecule: - is ligation-compatible with the 5' and 3' ends of the target DNA; - forms a first restriction enzyme recognition site upon self-ligation; - in a form capable of self-connection; and - the target DNA is in a form that prevents self-ligation and its 5' end and 3' end are ligation-compatible with the 5' end and 3' end of the backbone DNA molecule; The method further comprises: - ligating the target DNA molecule to the backbone DNA molecule in the presence of a ligase and a first restriction enzyme that cuts the first restriction enzyme recognition site, thereby generating at least one DNA circle including the backbone DNA molecule and the target DNA molecule; -optionally remove linear DNA; - generating concatemeric DNA molecules comprising an ordered array of copies of said at least one DNA circle by rolling circle amplification; and - sequencing said at least one concatemer, in, - The target DNA molecule contains 20 to 300 base pairs; - the form that allows self-ligation is the 5'-phosphate group of one DNA end and the 3'-hydroxyl group of the other DNA end, and the form that prevents self-ligation is the 5'-hydroxyl group of one DNA end and the 3'-hydroxyl group of the other DNA end; - ligating the end of the target DNA molecule to the end of the backbone DNA molecule to produce a target-backbone linker having a sequence that is not recognized or cleaved by the restriction enzyme that cleaves the restriction enzyme site formed by self-ligation of the backbone; - the backbone DNA molecule comprises a linker, the linker comprising a sequence of 20 to 500 nucleotides; - the backbone DNA molecule has a length of 200 to 600 nucleotides and has a flexibility score of 10 or higher, wherein the backbone DNA molecule has a GC content of 30-60% and a Shannon entropy value of 1.5 Sh or higher, and wherein the sequence of the linker does not have a repetitive DNA motif of more than 5 nucleotides and does not have a self-complementary motif of more than 6 nucleotides separated by less than 10 nucleotides.
2. The method of claim 1, wherein the concatemers are sequenced by long read sequencing.
3. The method of claim 1, wherein two or more backbone DNA molecules are provided.
4. The method of claim 2, wherein two or more backbone DNA molecules are provided.
5. The method of claim 3, wherein at least two backbone DNA molecules comprise a unique identifier sequence.
6. The method of claim 4, wherein at least two backbone DNA molecules comprise a unique identifier sequence.
7. The method according to any one of claims 1 to 6, wherein the specific ligation-compatible 5' end and 3' end are blunt ends.
8. A method for evaluating the target DNA capture efficiency of a backbone DNA molecule, comprising the steps of claim 1 and further comprising comparing the target DNA molecule capture efficiency between different backbone DNA molecules.
9. A collection of linear double-stranded backbone DNA molecules of 200 to 600 nucleotides in length, said DNA molecules comprising a 5' end and a 3' end, said 5' end comprising a portion of a first restriction enzyme recognition site at the very end, and said 3' end comprising another portion of the first restriction enzyme recognition site at the very end, and said 5' end and said 3' end are ligation compatible with each other and form a first restriction enzyme recognition site when self-ligated, and wherein each of said backbone DNA molecules comprises: Connector; an identifier sequence that is different from the sequences of identifiers of other backbone DNA molecules in the set; and optionally a restriction site for nicking enzyme, in, - said collection of linear double-stranded backbone DNA molecules has -Flexibility score of 10 or higher, -30-60% GC content, -1.5 Sh or higher Shannon entropy, and - said linker comprises a sequence of 20 to 500 nucleotides, - said sequence of said linker does not have a repetitive DNA motif of more than 5 nucleotides, and - said sequence of said linker does not have self-complementary motifs of more than 6 nucleotides separated by less than 10 nucleotides, or a combination thereof.
10. The collection of backbone DNA molecules according to claim 9, wherein the backbone DNA molecules further comprise restriction enzyme sites for type II restriction enzymes capable of generating non-palindromic overhangs, thereby providing Golden-Gate cloning sites.
11. A collection of backbone DNA molecules according to claim 9, wherein the linkers comprise sequences of 30 to 900 nucleotides; the linkers have a high overall complexity; the linkers do not have repetitive DNA motifs of more than 5 nucleotides, or do not have self-complementary motifs of more than 3 nucleotides separated by less than 10 nucleotides, or a combination thereof.
12. A collection of backbone DNA molecules according to claim 10, wherein the linkers comprise sequences of 30 to 900 nucleotides; the linkers have a high overall complexity; the linkers do not have repetitive DNA motifs of more than 5 nucleotides, or do not have self-complementary motifs of more than 3 nucleotides separated by less than 10 nucleotides, or a combination thereof.
13. The collection of backbone DNA molecules according to any one of claims 9 to 12, further comprising a nucleic acid molecule in the first restriction enzyme recognition site, wherein the nucleic acid molecule in the first restriction enzyme recognition site is a captured nucleic acid molecule.
14. A collection of backbone DNA molecules according to any one of claims 9 to 12, comprising a library of captured nucleic acid molecules.
15. A method for determining the sequence of a collection of nucleic acid molecules, comprising: - providing a double-stranded target DNA molecule having a recombinase recognition site specific for a target site-specific recombinase at the 5' end and the 3' end; - providing a backbone DNA molecule comprising said recognition site, said recognition site being separated by a DNA comprising a linker, wherein the target DNA molecule and the backbone DNA molecule are provided in a molar ratio ranging from 1:5 to 5:1; - incubating the target DNA molecule with the backbone DNA molecule in the presence of the target site-specific recombinase, thereby generating a DNA circle including the backbone DNA molecule and the target DNA molecule; -optionally remove linear DNA; - generating concatemers of an ordered array comprising at least two copies of said DNA circles by rolling circle amplification; and - sequencing said concatemer, in, - the target DNA molecule comprises 20 to 400 base pairs; - the backbone DNA molecule has a length of 200 to 600 nucleotides and has -Flexibility score of 10 or higher, -30-60% GC content, -1.5 Sh or higher Shannon entropy value, - the linker comprises a sequence of 20 to 500 nucleotides, and - said sequence of said linker does not have a repetitive DNA motif of more than 5 nucleotides, and - said sequence of said linker does not have self-complementary motifs of more than 6 nucleotides separated by less than 10 nucleotides, or a combination thereof. The method according to claim 15 , wherein the target site-specific recombinase is Cre recombinase, FLP recombinase or bacteriophage λ integrase.
17. A kit comprising a collection of linear double-stranded backbone DNA molecules according to any one of claims 11 to 14.
18. The kit according to claim 17, further comprising a high processivity polymerase and optionally one or more polymerase primers.
19. The kit according to claim 17, further comprising a ligase and the first restriction enzyme; and / or a target site-specific recombinase.
20. The kit according to claim 18, further comprising a ligase and the first restriction enzyme; and / or a target site-specific recombinase.
21. The kit according to any one of claims 17 to 20, further comprising a DNA exonuclease.
Citation Information
Patent Citations
Sequence tag directed subassembly of short sequencing reads into long sequencing reads
US20100069263A1
Sequencing using concatemers of copies of sense and antisense strands
US20160376647A1