method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- OXFORD NANOPORE TECH LTD
- Filing Date
- 2023-04-04
- Publication Date
- 2026-04-13
AI Technical Summary
The prior art is inefficient and costly when binding a nucleic acid adapter to a target polynucleotide and building a nucleic acid library, especially in polynucleotide serialization and recognition techniques.
The topoisomerase activation adapter is used to bind the topoisomerase to the double-stranded urea nucleotide, and bind to the 5' end of the target polynucleotide through specific trinucleotide sequences (such as GGT, GGGG, etc.), thereby improving the binding efficiency of the adapter to the target polynucleotide.
It improves the efficiency and quality of nucleic acid library construction, reduces dependence on high fluorescent chemicals, reduces experimental costs, and improves the efficiency of polynucleotide serialization and recognition.
Smart Images

Figure 00000041_0000 
Figure 00000041_0001 
Figure 00000041_0002
Abstract
Description
[Technical field]
[0001] The present disclosure relates to methods for attaching nucleic acid adaptors to target polynucleotides and methods for preparing nucleic acid libraries using topoisomerase-activated adaptors. The present disclosure also relates to methods for generating adaptive PCR amplification products and kits for carrying out the methods of the present disclosure. [Background technology]
[0002] Currently, there is a need for rapid and inexpensive polynucleotide (e.g., DNA or RNA) sequencing and identification techniques across a wide range of applications. Conventional techniques are time-consuming and expensive, mainly relying on amplification techniques to generate large amounts of polynucleotides and requiring large amounts of special fluorescent chemicals for signal detection.
[0003] Transmembrane pores (nanopores) have great potential as direct electrical biosensors for polymers and a variety of small molecules. In particular, nanopores have recently attracted attention as a potential DNA sequencing technology.
[0004] A potential is applied across the nanopore, causing a change in current flow when an analyte, such as a nucleotide, is transiently present within the barrel for a period of time. Detection of the nucleotide by the nanopore results in a current change of known signature and duration. In strand sequencing methods, a single polynucleotide strand is passed through the pore and an identifier for the nucleotide is derived. Strand sequencing may involve the use of molecular brakes to control the movement of the polynucleotide through the pore.
[0005] There are many commercial situations, such as polynucleotide sequencing and identification technologies, that require the attachment of adapters to target polynucleotides and the preparation of nucleic acid libraries, which can be achieved using topoisomerase-activated adapters.
[0006] Cheng and Shuman (Nucleic Acids Research, 2000, Vol. 28, No. 9, 1893-1898) describe DNA strand transfer catalyzed by vaccinia topoisomerase and the ligation of DNA containing 3' mononucleotide overhangs.
[0007] US 2002 / 0068290 describes topoisomerase-activated nucleotide adaptors and methods and compositions for the rapid conjugation of target nucleic acid sequences with topoisomerase-activated adaptor sequences that provide a specific function to the target.
[0008] US 2016 / 0348152 describes compositions comprising activated topoisomerase adaptors and methods of using the activated topoisomerase adaptors to prepare libraries of target DNA duplexes derived from sample polynucleotides for use in next-generation sequencing methods.
[0009] WO 2016 / 059436 describes a method of using topoisomerase to attach double-stranded DNA to an RNA strand for use in nanopore sequencing. Summary of the Invention
[0010] The inventors have identified novel methods of attaching nucleic acid adaptors to double-stranded target polynucleotides and novel methods of preparing nucleic acid libraries. In particular, the inventors have determined that the disclosed methods utilizing topoisomerase-activated sequencing adaptors can be more easily covalently attached to the 5' ends of some amplification products than others, thus advantageously enabling improved methods for generating nucleic acid libraries.
[0011] The method includes using a topoisomerase-activated adaptor to covalently attach a nucleic acid adaptor to a 5' end of a double-stranded target polynucleotide that includes a triplet nucleotide sequence at the 5' end of one or both strands, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. The inventors have shown that a topoisomerase-activated adaptor can be more easily, e.g., more efficiently, covalently attached to the 5' end of a double-stranded polynucleotide target sequence having a 5' terminal triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG.
[0012] Polynucleotide sequencing applications typically require that a specific sequencing adaptor be attached to a target polynucleotide in order to sequence the target polynucleotide. Maximizing the percentage of polynucleotides in a sample that contain attached sequencing adaptors also maximizes the percentage of polynucleotides that are sequenced, improving overall sequencing efficiency. Furthermore, by ensuring that the target double-stranded polynucleotide contains a 5'-terminal triplet nucleotide sequence to which topoisomerase more readily binds a topoisomerase-activated adaptor (e.g., a 5'-terminal triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG), more consistent adaptor attachment can be achieved across multiple different target polynucleotides (e.g., target amplicons), resulting in improved subsequent sequencing performance and reduced sampling bias in library preparation. Thus, the disclosed method allows for increased and / or improved sequencing information obtained from a polynucleotide sample.
[0013] In a first aspect, the present disclosure provides a method for preparing a nucleic acid library, the method comprising:
[0014] (a) providing a plurality of topoisomerase-activated adaptors, each topoisomerase-activated adaptor comprising a topoisomerase bound to a double-stranded oligonucleotide;
[0015] (b) providing a plurality of double-stranded target polynucleotides, each double-stranded target polynucleotide comprising a triplet nucleotide sequence at a 5' end of one or both strands, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG;
[0016] (c) contacting a topoisomerase-activated adaptor with a plurality of double-stranded target polynucleotides such that the topoisomerase covalently attaches the activated adaptor to the 5' ends of the double-stranded target polynucleotides;
[0017] Thereby, a nucleic acid library is prepared.
[0018] The present disclosure also provides a method of attaching a polynucleotide adaptor to a double stranded target polynucleotide, the method comprising:
[0019] (a) providing a topoisomerase-activatable adaptor comprising a topoisomerase bound to a double-stranded oligonucleotide;
[0020] (b) providing a double-stranded target polynucleotide comprising a triplet nucleotide sequence at a 5' end of one or both strands, wherein the triplet nucleotide sequence is selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG;
[0021] (c) contacting a topoisomerase-activated adaptor with the double-stranded target polynucleotide such that the topoisomerase covalently attaches the activated adaptor to the 5' end of the double-stranded target polynucleotide;
[0022] Thereby, the nucleic acid adaptor is attached to the double-stranded target polynucleotide.
[0023] The present disclosure also provides a method for generating an adaptive PCR amplification product, the method comprising:
[0024] (i) amplifying a first nucleic acid using a first oligonucleotide primer and a second oligonucleotide primer to generate a first amplification product, where at least one of the first primer and the second primer comprises a 5' tail that terminates in a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG to generate the first amplification product, and amplifying a second nucleic acid using a third primer and a fourth primer to generate a second amplification product, where at least one of the third primer and the fourth primer comprises a 5' tail that terminates in a triplet nucleotide sequence that is not GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, or GCG to generate the second amplification product;
[0025] (ii) contacting the first amplification product and the second amplification product with a topoisomerase-activated adaptor, such that the topoisomerase covalently attaches the adaptor to the 5' end of the first amplification product and to the 5' end of the second amplification product.
[0026] The present disclosure also provides a kit, the kit comprising:
[0027] - topoisomerase,
[0028] a double-stranded oligonucleotide adaptor comprising a topoisomerase target site;
[0029] - a primer pair consisting of a first oligonucleotide primer and a second oligonucleotide primer, wherein at least one of the first primer and the second primer is a 5' tailed primer having a triplet nucleotide sequence at the 5' end of the tail selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. [Brief description of the drawings]
[0030] [Figure 1] Figure 2 shows the total number of reads obtained after sequencing the amplification products obtained in Example 1 using control primers and NNN-tailed primers on an Oxford Nanopore Technologies MinION sequencer (compare all control primers with all NNN-tailed primers). [Diagram 2] Total number of reads obtained after sequencing the amplification products obtained in Example 1 using the control primer and the NNN-tailed primer on an Oxford Nanopore Technologies MinION sequencer (comparing control primer to NNN-tailed primer for each target gene). [Diagram 3] Graph of the data presented in Table A showing the total number of reads obtained at each NNN motif for all six targets. [Figure 4] Percentage of template modification for control and three different targets, GGK tailed, 0x = no adapter ligated, 1x = adapter ligated to one end of the amplicon, 2x = adapter ligated to both ends of the amplicon, with each target shown from left to right. [Diagram 5] Number of reads obtained for control and GGK-tailed amplicons. [Figure 6] Resulting read counts showing control and GGK-tailed amplification products of the three target genes. [Figure 7]Percentage of template modification for 12 different targets, control and GGK-tailed, where 0x = no adapter ligated, 1x = adapter ligated to one end of the amplicon, and 2x = adapter ligated to both ends of the amplicon, with each target shown from left to right. [Figure 8] Read counts for E. coli target and control amplicons. [Figure 9] E. coli target, number of reads of GGK-tailed amplicons. [Figure 10] Percentage of template modification of control, GGK and ACA tails. 0x = no adapter ligated, 1x = adapter ligated to one end of the amplicon, 2x = adapter ligated to both ends of the amplicon, with each target shown from left to right. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] The present invention will be described with respect to certain embodiments and with reference to certain drawings, but the disclosure is not limited thereto, but only by the claims. Any reference signs in the claims should not be interpreted as limiting the scope. Of course, it should be understood that not necessarily all aspects or advantages are achieved in accordance with a particular embodiment. Thus, for example, a person skilled in the art will recognize that the disclosed embodiments can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other aspects or advantages taught or suggested herein.
[0032] The disclosure, both as to its configuration and method of operation, as well as its features and advantages, will be best understood by reference to the following detailed description in conjunction with the accompanying drawings. Aspects and advantages of the present invention will be made apparent and elucidated with reference to the embodiment(s) described below. When "one embodiment" or "one embodiment" is referred to throughout this specification, it means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one disclosed embodiment. Thus, the appearances of the phrases "in one embodiment" or "in one embodiment" in various places in this specification do not necessarily all refer to the same embodiment, although they may. Similarly, in describing the disclosed exemplary embodiments, it should be understood that various features may be grouped together in one embodiment, drawing, or description thereof in order to streamline the disclosure and facilitate understanding of one or more of the various aspects of the invention. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects reside in fewer than all features of a single foregoing disclosed embodiment.
[0033] Unless the context dictates otherwise, it is to be understood that the "embodiments" of the present disclosure can be specifically combined, and that specific combinations of all disclosed embodiments are further disclosed embodiments of the claimed invention (unless otherwise implied by the context).
[0034] Furthermore, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, includes reference to a "topoisomerase," etc.
[0035] Unless otherwise indicated, nucleic acid sequences herein are written left to right in 5' to 3' orientation.
[0036] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0037] Methods for preparing a nucleic acid library The present disclosure provides a method for preparing a nucleic acid library, the method comprising:
[0038] (a) providing a topoisomerase-activated adaptor comprising a topoisomerase bound to a double-stranded oligonucleotide, or a plurality of topoisomerase-activated adaptors, each comprising a topoisomerase bound to a double-stranded oligonucleotide;
[0039] (b) providing a plurality of double-stranded target polynucleotides, each double-stranded target polynucleotide comprising a triplet nucleotide sequence at a 5' end of one or both strands, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG;
[0040] (c) contacting the topoisomerase-activated adaptors with a plurality of double-stranded target polynucleotides such that the topoisomerase covalently attaches the activated adaptors to the 5' ends of the double-stranded target polynucleotides;
[0041] Thereby, a nucleic acid library is prepared.
[0042] For example, a method for preparing a nucleic acid library, the method comprising:
[0043] (a) providing a plurality of topoisomerase-activated adaptors, each topoisomerase-activated adaptor comprising a topoisomerase bound to a double-stranded oligonucleotide;
[0044] (b) providing a plurality of double-stranded target polynucleotides, each double-stranded target polynucleotide comprising a triplet nucleotide sequence at a 5' end of one or both strands, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, and CGT, or GGT, GGG, TGT, GTT, and TGG;
[0045] (c) contacting the topoisomerase-activated adaptors with a plurality of double-stranded target polynucleotides such that the topoisomerase covalently attaches the activated adaptors to the 5' ends of the double-stranded target polynucleotides;
[0046] Thereby, a nucleic acid library is prepared.
[0047] The nucleic acid libraries generated by the methods of the present invention can include a plurality of double-stranded target polynucleotides having topoisomerase-activating adaptors attached to one or both ends.
[0048] The nucleic acid library can be a sequence library. The term "sequencing library" can refer to nucleic acids prepared for sequencing, for example, using next-generation sequencing methods, such as nanopore sequencing methods. The nucleic acids in the sequencing library can be optionally amplified nucleic acids, for example, in the form of amplification products, for example, using PCR.
[0049] Double-stranded target polynucleotide The double-stranded target polynucleotide may comprise a nucleic acid. The nucleic acid may comprise one or more natural nucleic acids, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). The double-stranded target polynucleotide may comprise a double-stranded DNA or a double-stranded RNA. The double-stranded target polynucleotide may comprise a DNA / RNA duplex, for example, a single RNA strand hybridized to a single DNA strand.
[0050] The nucleic acid may comprise one or more synthetic nucleic acids. Synthetic nucleic acids are known in the art. For example, the nucleic acid may comprise peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) and / or other synthetic polymers with nucleotide side chains. If the polynucleotide is a PNA, the PNA backbone may be composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. If the polynucleotide is a GNA, the GNA backbone may be composed of repeating glycol units linked by phosphodiester bonds. If the polynucleotide is a TNA, the TNA backbone may be composed of repeating threose sugars linked by phosphodiester bonds. If the polynucleotide is an LNA, the LNA backbone may be formed from ribonucleotides as described above, with an additional bridge connecting the 2' oxygen and 4' carbon of the ribose moiety.
[0051] The double-stranded target polynucleotide may be of any length. For example, the double-stranded target polynucleotide may be at least about 10, at least about 50, at least about 70, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 400 or at least about 500 nucleotides in length. The double-stranded target polynucleotide may be at least about 1,000, at least about 5,000, at least about 10,000, at least about 100,000, at least about 500,000, at least about 1,000,000 or at least about 10,000,000 nucleotides in length or more. The double-stranded target polynucleotide is preferably about 30 to about 10,000 nucleotides in length, for example, about 50 to about 5,000 nucleotides, about 100 to about 2,000 nucleotides or about 500 to about 1,000 nucleotides in length. The double-stranded target polynucleotide itself may be a fragment of a longer polynucleotide.
[0052] The double-stranded target polynucleotide may be linear. The double-stranded target polynucleotide may be circular. The double-stranded target polynucleotide of interest may be an end-to-end RNA or DNA molecule.
[0053] The double-stranded target polynucleotide comprises a triplet nucleotide sequence at the 5'-end of one or both strands of the target polynucleotide, preferably at the 5'-end of both strands of the target polynucleotide, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. In one embodiment, the triplet nucleotide sequence is selected from GKK and KGK, where K is G or T. Thus, in one embodiment, the triplet nucleotide sequence is selected from GGG, GTG, GGT, GTT, TGG and TGT. In one embodiment, the triplet nucleotide sequence is GGK. Thus, in one embodiment, the triplet nucleotide sequence is GGG or GGT. Optionally, the triplet nucleotide sequence is GGG. Optionally, the triplet sequence is GGT.
[0054] The triplet nucleotide sequence at the 5' end of one or both strands of the multiple double-stranded target polynucleotides can be a first triplet sequence immediately 5' to a second triplet nucleotide sequence, i.e., the 3' end of the first triplet sequence is joined to the 5' end of the second triplet sequence, and the second triplet nucleotide sequence is selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. In one embodiment, the second triplet nucleotide sequence is selected from GKK and KGK, where K is G or T. Thus, in one embodiment, the second triplet nucleotide sequence is selected from GGG, GTG, GGT, GTT, TGG and TGT. Optionally, the second triplet nucleotide sequence is GGK, i.e., GGT or GGG. Optionally, the second triplet nucleotide sequence is GGG. Optionally, the second triplet nucleotide sequence is GGT. Thus, the double-stranded target polynucleotide may comprise a sextuplet nucleotide sequence at the 5'-end of one or both strands of the target polynucleotide, preferably at the 5'-end of both strands of the target polynucleotide, the sextuplet nucleotide sequence consisting of a first triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, and a second triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. In one embodiment, the first triplet nucleotide sequence is selected from GKK and KGK, where K is G or T. Thus, in one embodiment, the first triplet nucleotide sequence is selected from GGG, GTG, GGT, GTT, TGG and TGT. Optionally, the first triplet nucleotide sequence is GGK, i.e., GGT or GGG. Optionally, the first triplet nucleotide sequence is GGT. Optionally, the first triplet nucleotide sequence is GGG. Thus, exemplary sextuplet sequences include GGKGGK, such as GGTGGT, GGGGGG, GGTGGG, or GGGGGT.
[0055] The second triplet sequence may be immediately 5' to a third triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. In one embodiment, the third triplet nucleotide sequence is selected from GKK and KGK, where K is G or T. Thus, in one embodiment, the third triplet nucleotide sequence is selected from GGG, GTG, GGT, GTT, TGG and TGT. Optionally, the third triplet nucleotide sequence is GGK, i.e., GGT or GGG. Optionally, the third triplet nucleotide sequence is GGG. Optionally, the third triplet nucleotide sequence is GGT. Thus, the double-stranded target polynucleotide comprises a 9mer nucleotide sequence at the 5'-end of one or both strands of the target polynucleotide, preferably at the 5'-end of both strands of the target polynucleotide, the 9mer nucleotide sequence being comprised of a first triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, a second triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, and a third triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. Exemplary 9mer sequences include GGKGGKGGK, such as GGTGGTGGT, GGTGGGTGGGG, GGGGGGGGG, GGGGGGGGT, GGTGGGGGT, GGTGGGGGG, GGGGGGTGGT or GGGGTGGGG.
[0056] The double-stranded target polynucleotide may be blunt-ended. The double-stranded target polynucleotide may be generated using a blunted polymerase. Such blunted polymerases as NEB's Q5 are well known in the art.
[0057] The double-stranded target polynucleotide may optionally include a nucleotide overhang at the 3' overhang. The overhang may be generated by a polymerase known in the art. The 3' overhang may be, for example, an "A-tail" that includes a single dA nucleotide. The A-tailed overhang may be generated using any method known in the art.
[0058] The ends of a double-stranded target polynucleotide are preferably complementary to the ends of an adaptor contained within a topoisomerase-activated adaptor described herein; for example, a blunt-ended adaptor is typically used with a blunt-ended target polynucleotide, and a dT-tailed adaptor is typically used with a dA-tailed double-stranded target polynucleotide.
[0059] Multiple double-stranded target polynucleotides The method for preparing nucleic acid library comprises providing a plurality of double-stranded target polynucleotides.Therefore, in this method, at least two double-stranded target polynucleotides are provided.For example, at least about 5, at least about 10, at least about 50 or at least about 100 double-stranded target polynucleotides can be provided.
[0060] Multiple double-stranded target polynucleotides may be present in a sample. The present invention is typically performed on a sample known to contain or suspected to contain multiple double-stranded target polynucleotides.
[0061] The sample may be a biological sample. These methods may be performed in vitro using samples taken or extracted from any organism or microorganism. The organism or microorganism is typically an archaea, prokaryote or eukaryote, typically belonging to one of the five kingdoms: plantae, animalia, fungi, monera, protista. The invention may be performed in vitro on samples taken or extracted from any virus. The sample is preferably a liquid sample. The sample typically comprises a body fluid of the patient. The sample may be urine, lymph, saliva, mucus, amniotic fluid, etc., but is preferably blood, plasma, serum.
[0062] Typically the samples are from humans, but may also be from other mammals, such as commercially farmed animals such as horses, cows, sheep, fish, chickens, pigs, or pets such as cats and dogs. Alternatively, the samples may be from plants, e.g. samples obtained from commercial crops such as grains, legumes, fruits, vegetables, e.g. wheat, barley, oats, canola, corn, soybean, rice, rhubarb, bananas, apples, tomatoes, potatoes, grapes, tobacco, beans, lentils, sugar cane, cocoa, cotton.
[0063] The sample may be a non-biological sample. The non-biological sample may preferably be a liquid sample. Examples of non-biological samples include surgical fluids, drinking water, water such as sea water or river water, and reagents for clinical testing.
[0064] Samples are typically processed before use in the present invention, for example by centrifugation or by passing through a membrane that filters out unwanted molecules or cells, such as red blood cells. Samples may be subjected to the methods disclosed herein immediately after collection. Samples may also typically be stored, preferably below -70°C, prior to analysis.
[0065] The double-stranded target polynucleotide may comprise a triplet nucleotide sequence at the 5'-end of one or both strands of the target polynucleotide, preferably at the 5'-end of both strands of the target polynucleotide, the triplet nucleotide sequence being selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, as described above, e.g., selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or from GGT, GGG, TGT, GTT and TGG.
[0066] In some embodiments, the multiple double-stranded target polynucleotides may include one or more, e.g., two, five, ten or more, double-stranded polynucleotides comprising a triplet nucleotide sequence at the 5'-end of one or both strands of the target polynucleotide, preferably at the 5'-end of both strands of the target polynucleotide, wherein the triplet nucleotide sequence is selected from CAA, ATG, AAT, TAC, GCG, TAG, AAC, ATC, GCA, GCC, TAA, TTT, ATA, AAG, CCT, AAA, CCG, TTG, TTC, TCG, TCA, CCC, CCA, TCT, ACT, ACC, TCC, ACG, TTA and ACA, e.g., ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA, and optionally the triplet sequence is ACA.
[0067] Generation of double-stranded polynucleotides The method may further include generating a double-stranded target polynucleotide, such as a plurality of double-stranded polynucleotides comprising a triplet nucleotide sequence at the 5' end of one or both strands, by PCR (polymerase chain reaction) using a first oligonucleotide primer and a second oligonucleotide primer, wherein at least one of the first primer and the second primer comprises a 5' tail that terminates in the triplet nucleotide sequence.
[0068] A plurality of double-stranded target polynucleotides comprising a triplet nucleotide sequence at the 5' end of both strands can be generated by using a first oligonucleotide primer and a second oligonucleotide primer, each of which comprises a 5' tail that terminates in a triplet nucleotide sequence.
[0069] An oligonucleotide primer having a 5' tail is known in the art as a "tailed" primer. A tailed primer comprises a 5' nucleotide sequence ("tail") that is non-complementary to a nucleotide sequence in a predetermined polynucleotide target and a 3' nucleotide sequence (primer sequence) that is complementary to a specific sequence in a predetermined polynucleotide target. Thus, the 5' tail of the primer can be non-complementary to a nucleotide sequence present at the primer binding site of the target polynucleotide. In an embodiment of the present disclosure that includes generating a plurality of double-stranded target polynucleotides that include a triplet nucleotide sequence at the 5' end of one or both strands, the 5' tail terminates with a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, e.g., GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or GGT, GGG, TGT, GTT and TGG, and thus the triplet nucleotide sequence is present at the 5' end of the 5' tail. In some embodiments, the 5' tail ends with GKK or KGK, where K is G or T. Thus, in some embodiments, the 5' tail ends with GGG, GTG, GGT, GTT, TGG or TGT. In some embodiments, the 5' tail ends with GGK. Thus, in some embodiments, the 5' tail ends with GGG or GGT. In some embodiments, the 5' tail ends with GGG. In some embodiments, the 5' tail ends with GGT. The use of such tailed primers incorporates a 5' tail triplet sequence at the 5' end of the amplification product generated from the PCR step. Such amplification products constitute double stranded target polynucleotides. Advantageously, by using the 5' tailed primers described herein, the triplet nucleotide sequence of the invention can be added to the 5' end of a known primer sequence of the target nucleic acid sequence to be amplified, thus providing the advantages of the invention when such known primer sequences are used in PCR.
[0070] The 5' tail can consist of or include a triplet nucleotide sequence at the 5' end. The 5' tail can be of any suitable length, for example, from 3 to about 15 nucleotides, for example, from about 4, 5 or 6 nucleotides to about 8, 9, 10 or 12 nucleotides. In some embodiments, the triplet nucleotide sequence at the 5' end of the tail is followed by up to about 3, about 6, about 9, or about 12 additional N nucleotides, where N is A, C, G, or T, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 additional N nucleotides, for example, NNN, NNNNNN, NNNNNNNNN, or NNNNNNNNNNNN. Thus, in some embodiments, the 5' tail comprises a first triplet nucleotide sequence selected from (from the 5' end) GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, or selected from GKK and KGK, or GGK, followed by the nucleotide sequence NNN, or NNNNNN, or NNNNNNNNN, or NNNNNNNNNNNN, where N is A, C, G or T.
[0071] The 5' tail may comprise or consist of a sextuplet nucleotide sequence, the sextuplet nucleotide sequence consisting of a first triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, and a second triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. The first triplet sequence and the second triplet sequence may be independently selected from, for example, GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or from GGT, GGG, TGT, GTT and TGG, respectively. In one embodiment, the first triplet nucleotide sequence is selected from GKK and KGK, where K is G or T. Thus, in one embodiment, the first triplet nucleotide sequence is selected from GGG, GTG, GGT, GTT, TGG and TGT. In one embodiment, the first triplet nucleotide sequence is GGK, i.e., GGT or GGG. Optionally, the first triplet nucleotide sequence is GGT. Thus, exemplary sextuplet sequences include GGKKKK and GGKGGK, for example, GGTGGT, GGGGGG, GGTGGG or GGGGGT.
[0072] The 5' tail may comprise or consist of a 9mer nucleotide sequence, the 9mer nucleotide sequence consisting of a first triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGT and GCG, a second triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, and a third triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. The first triplet sequence, the second triplet sequence and the third triplet sequence may be independently selected from, for example, GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or from GGT, GGG, TGT, GTT and TGG, respectively. Exemplary 9mer sequences include GGKKKKKKK and GGKGGKKKK, GGKGGKGGK, e.g., GGTGGTGGT, GGTGGTGGG, GGGGGGGGG, GGGGGGGGT, GGTGGGGGT, GGTGGGGGG, GGGGGGGGG, GGGGGTGGT or GGGGGTGGG. In an embodiment of the present disclosure that includes the generation of a plurality of double-stranded target polynucleotides comprising a triplet nucleotide sequence at the 5'-end of one or both strands by PCR, the PCR step can be carried out using any suitable polymerase known in the art. In particular, the suitable polymerase can be determined by the desire of the user to generate a double-stranded target polynucleotide amplification product having blunt ends or nucleotide overhang ends. Such polymerases are well known in the art.
[0073] Topoisomerase-activatable adaptors Topoisomerases, such as Vaccinia virus topoisomerase I, act in vivo to help regulate the positive and negative supercoiling of DNA. Vaccinia virus topoisomerase I binds to double-stranded DNA and cleaves the phosphodiester backbone of one strand at the 3' end of the target sequence (C / T)CCTT. The cleavage reaction conserves binding energy by forming a covalent adduct between the 3' phosphate of the incised strand and Tyr-274 of the topoisomerase protein. Topoisomerases can either religate the covalently bound strand through the same bond that was originally cleaved (occurring during relaxation of supercoiled DNA) or religate to a heterologous acceptor DNA to create a recombinant molecule. Dissociation of the nucleic acid associated with the free 5' end generated by the topoisomerase cleavage event allows another nucleic acid fragment with a compatible end containing a free 5'-OH to bind to the activated topoisomerase-DNA complex. Vaccinia DNA topoisomerase is described in Cheng and Shuman (Nucleic Acids Research, 2000, Vol. 28, No. 9, 1893-1898).
[0074] The methods described herein may include providing a topoisomerase-activated adaptor that includes a topoisomerase bound to a double-stranded oligonucleotide. Thus, the adaptor may be a topoisomerase-activated polynucleotide adaptor. The term "topoisomerase-activated adaptor" may refer to a polynucleotide structure that includes a double-stranded oligonucleotide region at or near the 3' end of a first end to which a topoisomerase is covalently attached. In one embodiment, the topoisomerase-activated adaptor is not a cloning vector.
[0075] Any suitable topoisomerase known in the art can be used in this method. The topoisomerase can be topoisomerase I. The topoisomerase can be vaccinia virus topoisomerase I.
[0076] The double-stranded oligonucleotide may comprise nucleic acid. The nucleic acid may comprise one or more natural nucleic acids, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). The double-stranded target polynucleotide may comprise double-stranded DNA or double-stranded RNA. The double-stranded target polynucleotide may comprise a DNA / RNA duplex, for example, a single RNA strand hybridized to a single DNA strand.
[0077] The nucleic acid may comprise one or more synthetic nucleic acids. Synthetic nucleic acids are known in the art. For example, the nucleic acid may comprise peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) and / or other synthetic polymers with nucleotide side chains. If the polynucleotide is a PNA, the PNA backbone may be composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. If the polynucleotide is a GNA, the GNA backbone may be composed of repeating glycol units linked by phosphodiester bonds. If the polynucleotide is a TNA, the TNA backbone may be composed of repeating threose sugars linked by phosphodiester bonds. If the polynucleotide is an LNA, the LNA backbone may be formed from ribonucleotides as described above, with an additional bridge connecting the 2' oxygen and 4' carbon of the ribose moiety.
[0078] The double-stranded oligonucleotide can be of any suitable length, such as about 10 to about 100 nucleotides, such as about 40, 50, 60, 70, 80 or 90 nucleotides. The double-stranded oligonucleotide can include a single-stranded overhang at one or both ends. The single-stranded overhang can be of any suitable length, such as 1 to about 30 nucleotides, such as at least about 5, about 10, about 15 or about 20 nucleotides.
[0079] The adaptor typically comprises a topoisomerase target site. Preferably, the topoisomerase target site comprises a [C / T]CCTT nucleotide sequence, such as CCCTT or TCCTT. Topoisomerase is known to recognize the [C / T]CCTT site in double-stranded DNA, cleave the backbone of the double-stranded DNA, and covalently bind to the 3' hydroxyl group of the cleaved backbone to generate a topoisomerase-activated adaptor and a leaving group. The topoisomerase then attaches the adaptor to the 5' end of the double-stranded target polynucleotide. The adaptor covalently attached to the topoisomerase is preferably blunt-ended or can have a nucleotide overhang, such as a 5' dT tail. The type of ends of the topoisomerase adaptor are well known in the art and can be selected based on the corresponding ends of the double-stranded target polynucleotide, so that the ends of the adaptor and the ends of the double-stranded target polynucleotide are complementary. For example, if the double-stranded target polynucleotide is blunt-ended, the adaptors are also preferably blunt-ended, or if the double-stranded target polynucleotide is dA-tailed, the adaptors are preferably dT-tailed.
[0080] Preferably, the adaptor comprises a [C / T]CCTT sequence at the 3' end, such a sequence allows for the generation of a topoisomerase-activated adaptor by initially including the adaptor within a double-stranded DNA sequence, whereby a topoisomerase cleaves the DNA backbone and covalently binds to the 3' hydroxyl group of the [C / T]CCTT sequence, thereby generating a topoisomerase-activated adaptor.
[0081] The adaptor may further include any components necessary to further functionalize the adaptor. Such components are known in the art. The adaptor may include a click-reactive group, a barcode, a fluorophore, a binder, a pull-down group, a tethering moiety, a marker, a modified base, an abasic residue, a sequencing adaptor, an intermediate adaptor, an amplification adaptor, a hairpin adaptor, a unique molecular identifier, a helicase binding site, and / or a spacer. Examples of polynucleotide sequencing adaptors suitable for use in nanopore sequencing are described in WO 2015 / 110813 and WO 2020 / 234612, both of which are incorporated herein by reference. The adaptor may include components as described above, such as a click group, that allow for subsequent attachment of additional adaptors, such as polynucleotide sequencing adaptors, to the adaptor. Thus, by way of example, attachment of a topoisomerase-activated adaptor to a target double-stranded polynucleotide, as described herein, may facilitate subsequent attachment of a polynucleotide sequencing adaptor.
[0082] The adaptor may include a marker. The marker may be any suitable marker that allows a person skilled in the art to identify where the adaptor is attached to the double-stranded target polynucleotide or whether one or more adaptors are attached. Exemplary markers include one or more biotin molecules, one or more modified bases, one or more abasic residues, one or more base-base complexes, or one or more protein-base complexes.
[0083] Preferably, the marker is detectable by one or more sequencing techniques known in the art, such as nanopore-based sequencing. The marker may be optically detectable. For example, the marker may fluoresce under excitation with light of an appropriate wavelength. For example, the marker may include one or more fluorescent bases, such as Cy3 or Cy5. Any optical and / or fluorescent marker that is determined to be suitable by a person skilled in the art may be used. Other detectable markers include non-standard bases, abasics, and spacers. Any abasic and / or spacer that is determined to be suitable by a person skilled in the art may be used. In particular, exemplary spacers may include C3, PC spacer, hexanediol, spacer 9, spacer 18, or 1',2'-dideoxyribose (dSpacer).
[0084] The adaptor may include one or more modified bases. The modified base may be any suitable modified base. The modified base may be, for example, a nucleotide labeled with biotin (i.e., a biotinylated nucleotide) or a nucleotide labeled with digoxigenin (i.e., a digoxigenin-labeled nucleotide). The modified base, such as a biotin-labeled or digoxigenin-labeled base, may allow the double-stranded target polynucleotide to be bound to a solid surface, for example, a surface coated with streptavidin or anti-digoxigenin, respectively.
[0085] The adaptor may include a pull-down group. The pull-down group may be any suitable pull-down group that allows a person skilled in the art to purify or isolate a double-stranded target polynucleotide, or to immobilize a double-stranded target polynucleotide by binding it to another substance. The other substance may be, for example, a nucleic acid construct, a nucleic acid molecule, a polypeptide, a protein, a membrane, or a solid surface. An exemplary pull-down group includes one or more polypeptides and one or more hydrophobic anchors.
[0086] The pull-down group may, for example, comprise one or more modified nucleotide bases. The modified base may be a nucleotide labeled with biotin (i.e., a biotinylated nucleotide) or a nucleotide labeled with digoxigenin (i.e., a digoxigenin-labeled nucleotide). The modified base in the pull-down group, such as a biotin or digoxigenin-labeled base, allows the double-stranded target polynucleotide to be tethered to a solid surface, for example, a surface coated with streptavidin or anti-digoxigenin. Any suitable tether allows for binding to a solid surface. The solid surface that may be bound to the double-stranded target polynucleotide prepared by the method disclosed herein includes, for example, nanogold, polystyrene beads, and Qdots.
[0087] The pull-down group can include, for example, a tethering moiety that includes a hydrophobic anchor, and optionally, the hydrophobic anchor includes a hydrophobic nucleotide base. The tethering moiety can be a lipid, a fatty acid, a sterol, a carbon nanotube, a protein or an amino acid, cholesterol, palmitic acid, or octyl tocopherol.
[0088] The adaptor allows for further manipulation of the double-stranded target polynucleotide or allows for direct sequencing of the double-stranded target polynucleotide (e.g., sequencing adaptor). The adaptor can be any adaptor that is deemed appropriate by one of skill in the art.
[0089] Adapters are known in the art. Adapters can include, for example, nucleotide sequences that allow protein binding, sequencing adaptors, PCR adaptors, hairpin adaptors, adaptors that allow circularization and / or rolling circle amplification of target polynucleotides, unique molecular identifiers (UMIs), oligonucleotide splints, click chemistry moieties, exonuclease-resistant bases and / or phosphorothioate bonds. Adapters can be simultaneously bound to desired proteins, such as motor enzymes.
[0090] The adaptor may comprise an RNA and / or DNA sequence that can be recognized and bound by a DNA and / or RNA binding protein. For example, in the methods described herein, the adaptor may be bound by a motor enzyme, such as a helicase or translocase.
[0091] The adapters can be sequence motifs that can be specifically recognized by certain DNA and / or RNA binding proteins. The adapters can be RNA / DNA hybrid sequences that can be specifically recognized and bound by DNA and / or RNA binding proteins. For example, the sequence motifs can be recognized by DNA binding proteins characterized by structural domains such as helix-turn-helix, zinc finger, leucine zipper, winged helix, winged helix-turn-helix, helix-loop-helix, HMG box, Wor3, OB fold, etc. The adapters can be lacO or tetO arrays that can be bound by lac or tet repressor proteins. The adapters can be RNA-DNA hybrid sequences that can be bound by antibodies. The RNA-DNA hybrid markers can be bound by the S9.6 antibody.
[0092] The adaptor may comprise a nucleotide sequence suitable for hybridization of an oligonucleotide. In particular, the oligonucleotide may comprise complementary bases for hybridizing with the adaptor, thus allowing for extension (linear amplification) of a complementary polynucleotide sequence or priming of a polymerase chain reaction.
[0093] The adaptor may include a nucleotide sequence that functions as a unique molecular identifier (UMI). The UMI can be detected by any sequencing method that the skilled artisan considers appropriate. In particular, the UMI can be used to detect and quantify double-stranded target polynucleotides. The double-stranded polynucleotide sequence reads of the target can be clustered according to the presence of the UMI, thus improving the accuracy of single molecule sequencing.
[0094] The adaptor can be any adaptor that can form a covalent bond with another molecule, for example, via click chemistry. The adaptor can include a group that allows copper-free click chemistry. An exemplary group that can be applied to the adaptor of the present method is a 5'DBCO group.
[0095] Methods for determining the presence or absence or one or more properties of a double-stranded target polynucleotide covalently bound to an activated adaptor - Patent Application 20070123333 The product of any of the methods described herein, such as a double-stranded target polynucleotide with a covalently attached adaptor, or a double-stranded polynucleotide of a nucleic acid library containing double-stranded target polynucleotides, can be subjected to a nanopore-based method to determine the presence, absence, or one or more properties of said product. Thus, the method of the present disclosure may further include determining the presence, absence, or one or more properties of the target double-stranded target polynucleotide covalently attached to the adaptor by a topoisomerase as described above. The presence, absence, or one or more properties of such a target double-stranded target polynucleotide may be determined by:
[0096] (a) contacting a target polynucleotide with a nanopore or a variant thereof, such that the target polynucleotide is translocated relative to the pore;
[0097] (b) performing one or more measurements as the polynucleotide translocates relative to the pore;
[0098] The presence, absence, or one or more characteristics of the polynucleotide are thereby determined.
[0099] Nanopore-based methods for detecting target polynucleotides in a sample have been previously described (WO 2018 / 060740). Nanopore-based methods for characterizing target polynucleotides in a sample have been previously described (WO 2015 / 124935). Thus, any suitable nanopore-based method of polynucleotide characterization would be suitable for application in the methods of determining the presence, absence, or one or more properties of a double-stranded target polynucleotide covalently attached to an activating adaptor described herein.
[0100] The method may further comprise monitoring for the presence or absence of an effect on the potential difference applied across the membrane as a result of the interaction of the target polynucleotide with the transmembrane pore, thereby determining the presence or absence of the target polynucleotide. The effect is indicative of the double-stranded target polynucleotide interacting with the transmembrane pore. The effect may be caused by the translocation of one or both strands of the double-stranded target polynucleotide through the pore of the adaptor coupled to one of the components of the double-stranded target polynucleotide. The effect may be monitored using electrical and / or optical measurements. In this case, the effect is a measured change in an electrical or optical quantity. The electrical measurement may be a current measurement, an impedance measurement, a tunneling measurement or a field effect transistor (FET) measurement. The effect may be a change in the flow of ions through the transmembrane pore, resulting in a current, a change in resistance, or a change in optical properties. The effect may be an electron tunneling effect across the transmembrane pore. The effect may be a change in the potential due to the interaction of the double-stranded target polynucleotide with the transmembrane pore, which is monitored using a local potential sensor in a FET measurement.
[0101] In the methods of characterizing a double-stranded target polynucleotide described herein, contacting the double-stranded target polynucleotide with the pore results in at least one nucleic acid strand of the double-stranded target polynucleotide translocating through the pore.
[0102] In the methods of characterizing a double-stranded target polynucleotide described herein, one or more measurements are performed to indicate one or more properties of the double-stranded target polynucleotide, the properties being selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0103] Step (b) comprises contacting the target polynucleotide with a transmembrane pore and allowing the modified polynucleotide to translocate through the pore. The target polynucleotide is in contact with the transmembrane pore and both are capable of translocating through the pore.
[0104] The method is preferably carried out by applying a potential across the pore. The applied potential can be a voltage potential. Alternatively, the applied potential can be a chemical potential. One example is the use of a salt gradient across the amphiphilic layer. Salt gradients are disclosed in Holden et al., J Am Chem Soc. 2007 Jul 11; 129(27): 8650-5. In some cases, the current passing through the pore as the polynucleotide moves relative to the pore is used to determine the sequence of the double-stranded target polynucleotide. This is strand sequencing.
[0105] The methods for characterizing a target polynucleotide can be used to characterize, e.g., sequence, the entirety or only a portion of a double-stranded target polynucleotide.
[0106] A transmembrane pore is a structure that traverses a membrane to some extent. It allows hydrated ions driven by an applied potential to flow across or within the membrane. A transmembrane pore typically traverses the entire membrane, so that hydrated ions flow from one side of the membrane to the other side of the membrane. However, a transmembrane pore does not have to traverse the membrane. It may be closed on one side. For example, a pore can be a well, gap, channel, trench or slit in the membrane through which hydrated ions flow (or into).
[0107] Any suitable transmembrane pore may be used. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, solid pores. In any of the methods described herein, the pore may allow double-stranded polynucleotides and bound polynucleotides to translocate through the pore. In any of the methods described herein, the pore may allow double-stranded polynucleotides and bound polynucleotides to translocate through the pore. In any of the methods described herein, the pore may allow double-stranded polynucleotides to translocate. In any of the methods described herein, the pore may allow single-stranded polynucleotides to translocate.
[0108] The pores may be present in any suitable membrane. Suitable membranes are well known in the art. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, that have both at least one hydrophilic portion and at least one lipophilic or hydrophobic portion. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic molecules may be synthetic or natural. Non-natural amphiphiles and amphiphiles that form monolayers are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer subunits are polymerized to form a single polymer chain. Block copolymers typically have properties provided by each monomer subunit. However, block copolymers may have unique properties that polymers formed from individual subunits do not have. Block copolymers can be designed such that one of the monomer subunits is hydrophobic (i.e., lipophilic) and the other subunit(s) is hydrophilic in aqueous media. In this case, the block copolymer can have amphiphilic properties and form structures that mimic biological membranes. Block copolymers can be diblock (composed of two monomer subunits), but can also be composed of two or more monomer subunits to form more complex arrangements that function as amphiphiles. Copolymers can be triblock, tetrablock, or pentablock copolymers.
[0109] The amphiphilic layer can be a planar lipid bilayer or a supported bilayer.
[0110] The amphiphilic layer can be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as an excellent platform for various experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of various substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers or liposomes. The lipid bilayer can be a planar lipid bilayer. Suitable lipid bilayers are disclosed in International Application No. PCT / GB08 / 000563 (published as WO 2008 / 102121), International Application No. PCT / GB08 / 004127 (published as WO 2009 / 077734), and International Application No. PCT / GB2006 / 001057 (published as WO 2006 / 100484).
[0111] Methods for forming lipid bilayers are known in the art. Suitable methods are disclosed in the Examples. Lipid bilayers are generally formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566). Other common methods for forming bilayers include tip dipping, painting of bilayers and patch clamping of liposome bilayers.
[0112] The lipid bilayer may be formed as described in International Application No. PCT / GB08 / 004127 (published as WO 2009 / 077734).
[0113] The membrane may be a solid layer. The solid layer is not of biological origin. In other words, the solid layer is not derived from or isolated from a biological environment such as a living organism or cell, or from a synthetic version of a biologically available structure. The solid layer may be formed from both organic and inorganic materials, including but not limited to microelectronic materials, insulating materials such as Si3N4, A12O3, SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon, elastomers such as two-component addition-cured silicone rubber, glass, and the like. The solid layer may be formed of a monolayer, such as graphene, or a layer only a few atoms thick. A suitable graphene layer is disclosed in International Application No. PCT / US2008 / 010637 (published as WO 2009 / 035647).
[0114] The method is typically carried out using (i) an artificial amphiphilic layer containing a pore, (ii) an isolated natural lipid bilayer containing a pore, or (iii) a cell into which a pore has been inserted. The method is typically carried out using an artificial amphiphilic layer, such as an artificial lipid bilayer, which may contain, in addition to the pore, other transmembrane and / or intramembrane proteins, as well as other molecules. Suitable equipment and conditions are described below. The method of the present disclosure is carried out in vitro.
[0115] The double-stranded target polynucleotide may be bound to the membrane. This may be carried out using any known method. In particular, the double-stranded target polynucleotide may be bound to the membrane via a suitable modified polynucleotide linked to the fragmented target polynucleotide. Alternatively, the modified polynucleotide linked to the fragmented target polynucleotide may be modified to introduce a binding element or anchor element for binding the double-stranded target polynucleotide to the membrane. If the membrane is an amphiphilic layer such as a lipid bilayer (as described in detail above), the double-stranded target polynucleotide is preferably bound to the membrane via a polypeptide present in the membrane or a hydrophobic anchor present in the membrane. The hydrophobic anchor is preferably a lipid, a fatty acid, a sterol, a carbon nanotube or an amino acid.
[0116] The double-stranded target polynucleotide can be directly attached to the membrane. The polynucleotide is preferably attached to the membrane via a linker. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polypeptides, etc. If the polynucleotide is directly attached to the membrane, the characterization run cannot continue to the end of the double-stranded target polynucleotide due to the distance between the membrane and the pore, and some data will be lost. If a linker is used, the polynucleotide can be processed to completion. If a linker is used, the linker can be attached to any position of the polynucleotide. The linker is preferably attached to the double-stranded target polynucleotide at the tail polymer.
[0117] This bond can be stable or temporary. In certain applications, the temporary nature of the bond is preferred. If a stable binding molecule is attached directly to the 5' or 3' end of the polynucleotide, the characterization run cannot continue to the end of the polynucleotide due to the distance between the bilayer and the pore, and some data will be lost. If the bond is temporary, the polynucleotide can be processed to completion when the binding end is randomly released from the bilayer. Chemical groups that form stable or temporary bonds with the membrane are described in detail below. Polynucleotides can be temporarily attached to amphiphilic layers such as cholesterol or lipid bilayers using fatty acyl chains. Any fatty acyl chain with a length of 6 to 30 carbon atoms, such as hexadecanoic acid, can be used.
[0118] Suitable conjugation methods are disclosed in WO 2012 / 164270, WO 2015 / 110813 and WO 2015 / 150787.
[0119] A common technique for amplifying a section of genomic DNA is to use the polymerase chain reaction (PCR), where two synthetic oligonucleotide primers can be used to generate multiple copies of the same section of DNA, where each copy will have a synthetic polynucleotide 5' of each strand of the duplex. By using an antisense primer that has a reactive group such as cholesterol, thiol, biotin, lipid, etc., each copy of the amplified target DNA will contain a reactive group for binding.
[0120] The transmembrane pore is preferably a transmembrane protein pore. A transmembrane protein pore is a polypeptide or an assembly of polypeptides that allows hydrated ions, such as an analyte, to flow from one side of a membrane to the other side of the membrane. In the present disclosure, the transmembrane protein pore can form a pore that allows hydrated ions driven by an applied electric potential to flow from one side of a membrane to the other side. The transmembrane protein pore preferably allows analytes, such as nucleotides, to flow from one side of a membrane, such as a lipid bilayer, to the other side. The transmembrane protein pore allows polynucleotides, such as DNA or RNA, to move through the pore.
[0121] The transmembrane protein pore may be monomeric or oligomeric. The pore may be composed of multiple repeating subunits, such as 6, 7, 8 or 9 subunits. The pore may be a hexameric, heptameric, octameric or nonameric pore.
[0122] A transmembrane protein pore typically comprises a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and provide the strands for a transmembrane beta barrel or channel or a transmembrane alpha helical barrel or channel.
[0123] The barrel or channel of a transmembrane protein pore typically comprises amino acids that facilitate interaction with an analyte, such as a nucleotide, polynucleotide or nucleic acid. These amino acids are preferably located near the constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine, histidine, or aromatic amino acids, such as tyrosine, tryptophan. These amino acids typically facilitate interaction of the pore with a nucleotide, polynucleotide or nucleic acid.
[0124] Methods for Attaching Polynucleotide Adaptors to Double-Stranded Target Polynucleotides The present disclosure provides a method for attaching a polynucleotide adaptor to a double stranded target polynucleotide, the method comprising:
[0125] (a) providing a topoisomerase-activatable adaptor comprising a topoisomerase bound to a double-stranded oligonucleotide;
[0126] (b) providing a double stranded target polynucleotide comprising a triplet nucleotide sequence at a 5' end of one or both strands, wherein the triplet nucleotide sequence is selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, e.g., GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or GGT, GGG, TGT, GTT and TGG;
[0127] (c) contacting a topoisomerase-activated adaptor with the double-stranded target polynucleotide such that the topoisomerase covalently attaches the activated adaptor to the 5' end of the double-stranded target polynucleotide;
[0128] Thereby, the polynucleotide adaptor is bound to the double-stranded target.
[0129] The features of the method are as described in detail above with respect to the method of preparing a nucleic acid library. Any of the features described with respect to the method of preparing a nucleic acid library can be similarly applied to the method of attaching polynucleotide adaptors to double-stranded target polynucleotides.
[0130] The requirements for binding adaptors to double-stranded target polynucleotides using topoisomerase-activated adaptors are well known in the art, and therefore any suitable reaction conditions may be used, such as those described in the Examples herein.
[0131] Methods for generating adaptive PCR amplicons The present invention provides a method for generating an adaptive PCR amplification product, the method comprising:
[0132] (i) amplifying a first nucleic acid using a first oligonucleotide primer and a second oligonucleotide primer to generate a first amplification product, wherein at least one of the first primer and the second primer comprises a 5' tail terminating with a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG, e.g., a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or from GGT, GGG, TGT, GTT and TGG. and amplifying the second nucleic acid using a third primer and a fourth primer to generate a second amplification product, wherein at least one of the third primer and the fourth primer comprises a 5' tail terminating with a triplet nucleotide sequence that is not GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, or GCG, e.g., a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, and CGT, or from GGT, GGG, TGT, GTT, and TGG;
[0133] (ii) contacting the first amplification product and the second amplification product with a topoisomerase-activated adaptor, such that the topoisomerase covalently attaches the adaptor to the 5' end of the first amplification product and to the 5' end of the second amplification product.
[0134] As mentioned above, the present disclosure relates to a method for covalently attaching an adaptor to a double-stranded target polynucleotide comprising a 5' tail terminating in a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. The inventors have shown that a topoisomerase-activated sequencing adaptor can more easily, i.e., more efficiently, covalently attach an adaptor to the 5' end of a double-stranded polynucleotide target sequence having a 5' terminal triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG (see Example 1, Table A and FIG. 3). Polynucleotide sequencing applications tend to rely on target polynucleotides having specific sequencing adaptors to sequence the target polynucleotide. Maximizing the proportion of polynucleotides in a sample that contain sequencing adaptors also maximizes the proportion of polynucleotides that are sequenced. Thus, the method of the present disclosure allows for increased and / or improved sequencing information obtained from a polynucleotide sample.
[0135] Thus, the method of generating adaptive amplification products described herein can be utilized to covalently attach sequencing adaptors to certain nucleic acid targets with higher efficiency and to other targets with lower efficiency. For example, amplification primers used in a multiplex PCR assay to generate a sequencing library may vary in the level of amplification products in the final library due to variable primer hybridization and the resulting PCR amplification efficiency. In this regard, once a particular nucleic acid target is identified that has reduced amplification compared to other targets in a sample (i.e., due to reduced primer efficiency), such target may be amplified using a first primer and a second primer to generate a first amplification product, at least one of the first primer and the second primer comprising a 5' tail that terminates with a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG. Having such a tail means that when the first amplification product is subjected to step (ii), i.e., when the first amplification product is contacted with a topoisomerase-activated adapter and the topoisomerase covalently attaches the adapter to the 5' end of the first amplification product, the triplet will improve the efficiency of adapter binding in the amplification product. As discussed above, improved adapter binding efficiency will result in a greater proportion of a particular target sequence being sequenced, biasing the observed amplicon representation during a sequencing run, thereby compensating for the observed reduction in primer / amplification efficiency.
[0136] The reverse principle can be applied to certain nucleic acid targets that undergo increased amplification relative to other targets in a sample. Such targets can be amplified using a third primer and a fourth primer to generate an amplification product, where at least one of the third primer and the fourth primer is selected from the group consisting of GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GCG. do not haveIt comprises a 5' tail that terminates in a triplet nucleotide sequence. Such triplet nucleotides include those other than the top 10 in Table A and Figure 3, such as CAA, ATG, AAT, TAC, GCG, TAG, AAC, ATC, GCA, GCC, TAA, TTT, ATA, AAG, CCT, AAA, CCG, TTG, TTC, TCG, TCA, CCC, CCA, TCT, ACT, ACC, TCC, ACG, TTA or ACA. Having such a tail means that when the second amplification product is subjected to step (ii), i.e., when the second amplification product is contacted with a topoisomerase-activated adapter and the topoisomerase covalently attaches the adapter to the 5' end of the second amplification product, the triplet reduces the efficiency of adapter attachment on the amplification product. As discussed above, reduced adapter binding efficiency will result in a smaller proportion of a particular target sequence being sequenced, biasing the amplicon representation observed during a sequencing run, thereby compensating for increased primer / amplification efficiency.
[0137] The amplification efficiency of the target may be determined by any suitable method in the art. For large multiplex assays, it is preferable to use sequencing analysis to determine amplification bias and determine whether the method of generating PCR adaptive amplification products described herein can successfully adjust the amount of sequencing information obtained from a specific target. Thus, the method may further include a step, prior to (i), of determining, optionally by sequencing, whether a specific nucleic acid target has been disproportionately amplified during the preparation of the sequencing library. As described above, a nucleic acid target identified as having increased or decreased amplification efficiency, for example, as evidenced by disproportionately high or low sequence reads, prompts the user to deploy steps (i) and (ii) to generate adaptive PCR amplification products with variable adaptive efficiency to compensate for the previously determined disproportionate amplification efficiency. The method may further include a further step, after (ii), of determining, optionally by sequencing, whether the method has successfully compensated for the observed increase or decrease in amplification efficiency of a specific nucleic acid target. The method of generating adaptive PCR amplification products, optionally including the above steps before (i) and after (ii), can be repeated any number of times as deemed appropriate to achieve a sufficient level of compensation for increases or decreases in amplification efficiency.
[0138] At least one of the first and second primers may include a 5' tail that terminates in the triplet nucleotide sequence GGT or GGG.
[0139] The 5' tail of at least one of the third and fourth primers may terminate in a triplet nucleotide sequence selected from CAA, ATG, AAT, TAC, GCG, TAG, AAC, ATC, GCA, GCC, TAA, TTT, ATA, AAG, CCT, AAA, CCG, TTG, TTC, TCG, TCA, CCC, CCA, TCT, ACT, ACC, TCC, ACG, TTA or ACA.
[0140] The 5' tail of at least one of the third and fourth primers may terminate in a triplet nucleotide sequence selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA. Optionally, the third and fourth primers comprise a 5' tail that terminates in the triplet nucleotide sequence ACA.
[0141] In one embodiment, the 5' tail of at least one of the first and second primers may terminate in a triplet nucleotide sequence selected from GGT or GGG, and the 5' tail of at least one of the third and fourth primers may terminate in a triplet nucleotide sequence selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA.
[0142] In one embodiment, the 5' tail of at least one of the first and second primers may terminate with a triplet nucleotide sequence selected from GGT or GGG, and the 5' tail of at least one of the third and fourth primers may terminate with the triplet nucleotide sequence ACA.
[0143] kit The kit provided is:
[0144] Topoisomerase and
[0145] a double-stranded oligonucleotide adaptor comprising a topoisomerase target site;
[0146] a primer pair consisting of a first oligonucleotide primer and a second oligonucleotide primer, wherein at least one of the first primer and the second primer is a 5' tailed primer having a triplet nucleotide sequence at the 5' end of the tail selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG.
[0147] Any of the topoisomerases described above with respect to other aspects of the disclosure can be used in the kit.
[0148] Any of the double-stranded oligonucleotide adaptors described above with respect to other aspects of the present disclosure can be used in the kit. Preferably, the topoisomerase target site comprises a [C / T]CCTT nucleotide sequence. The adaptor may further comprise a click reactive group, a fluorophore, a binder, a pull-down group, a tethering moiety, a marker, a modified base, an abasic residue, a sequencing adaptor, an intermediate adaptor, an amplification adaptor, a hairpin adaptor, a unique molecular identifier, a helicase binding site, and / or a spacer.
[0149] The topoisomerase can be loaded onto a double-stranded oligonucleotide adaptor, for example, the adaptor can be a topoisomerase-activated adaptor as described herein.
[0150] At least one of the first and second primers is a 5' tailed primer having a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG at the 5' end of the tail, and the pair of primers may include any of the features of the primers described above with respect to other aspects of the disclosure.
[0151] In the primer pair, at least one of the first and second primers is a 5' tailed primer having a triplet nucleotide sequence that may be selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG and CGT, or GGT, GGG, TGT, GTT and TGG.
[0152] At least one of the first and second primers may include a 5' tail that terminates in the triplet nucleotide sequence GGT or GGG.
[0153] The kit may further comprise a pair of primers consisting of a third oligonucleotide primer and a fourth oligonucleotide primer, wherein at least one of the third and fourth oligonucleotide primers is a 5' tailed primer that terminates with a triplet nucleotide sequence that is not GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT or GCG. Optionally, the 5' tail of at least one of the third and fourth primers terminates with a triplet nucleotide sequence selected from CAA, ATG, AAT, TAC, GCG, TAG, AAC, ATC, GCA, GCC, TAA, TTT, ATA, AAG, CCT, AAA, CCG, TTG, TTC, TCG, TCA, CCC, CCA, TCT, ACT, ACC, TCC, ACG, TTA or ACA. Optionally, the 5' tail of at least one of the third and fourth primers terminates in a triplet nucleotide sequence selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA. Optionally, at least one of the third and fourth primers comprises a 5' tail that terminates in the triplet nucleotide sequence ACA.
[0154] Although specific embodiments, specific configurations, and materials and / or molecules have been described herein for methods according to the present disclosure, it should be understood that various changes or modifications in form and details can be made without departing from the scope and spirit of the present invention. The foregoing embodiments and the following examples are provided for illustrative purposes only and should not be considered as limiting this application. This application is limited only by the claims. EXAMPLES
[0155] Experiments detailed in the Examples below were performed to investigate whether topoisomerase-activated sequencing adaptors covalently attach to the 5' ends of some amplicons more readily than to other ends, thereby biasing the amplicons representation observed when sequencing was performed.
[0156] Example 1 This experiment was performed to determine whether sequences at the 5' end of target amplicons influence downstream representation in sequencing of the amplicons after ligation of topoisomerase-activated sequencing adaptors.PCR primers for three different target genes in Mycobacterium tuberculosis (TB) were generated, in each case both with and without an additional three-base sequence (NNN) appended to the 5' end of the primer (referred to as an NNN tail).
[0157] Preparation of adapters: Two different types of topoisomerase adaptors (topo adaptors) were prepared: dT-tailed and blunt.
[0158] dT Tail Type: [ka] [ka] [ka]
[0159] Brand Type: [ka] [ka] [ka]
[0160] DBCO-TEG=a click group that facilitates downstream attachment to sequencing adaptors.
[0161] mU = 2'-O-methyluridine.
[0162] The strand portion has the following functions, which are underlined.
[0163] Not underlined: overhang for binding to sequencing adapters.
[0164] Dotted underline: Barcode.
[0165] Single underline: topoisomerase binding site containing the CCCTT motif and adjacent regions.
[0166] Double underline: leaving group.
[0167] The adapters were prepared in the following manner.
[0168] DNA adapters were generated by annealing 10 μM strands in 10 mM Tris (pH 7.5), 50 mM NaCl from 95°C at 2°C / min.
[0169] Topo-DNA adapters were formed by mixing the following and incubating at 37° C. for 1 hour:
[0170] 4 μl nuclease-free HO
[0171] 1 μl 1M Tris (pH 7.5)
[0172] 5 μl DNA adapter
[0173] 10 μl Vaccinia topoisomerase
[0174] After 1 hour incubation, 20 μl of 50 mM Tris pH 7.5, 100 mM NaCl, 5% glycerol is added.
[0175] Primer: The following primers were mixed to generate forward and reverse primer pairs for the three TB targets (rpoB, gyrA, fabG1) with or without NNN tails: [Table 1] [Table 2]
[0176] PCR: For each of the six primer mixes, the following mixes were assembled: [Table 3]
[0177] Amplification was carried out according to the following conditions: [Table 4]
[0178] Samples were quantified using a Qubit system (ThermoFisher Scientific) and diluted to 25 ng / μL in water.
[0179] Sequencing analysis: The following reactions were assembled for control and NNN-tailed templates and incubated at 65°C for 10 min and 80°C for 2 min. [Table 5]
[0180] The following steps were carried out:
[0181] -For SPRI purification (0.7x), wash twice with 70% EtOH and elute in 14 µL of EB.
[0182] -For the addition of sequencing adaptors, add 1 μL of RAP-F fast sequencing adaptors to 11 μL of eluted library and incubate at room temperature for 10 min before running on a MinION sequencer (Oxford Nanopore Technologies) for 12 h.
[0183] The total number of reads was measured. The reads were then downsampled and aligned to the TB genome, and the number of reads (forward and reverse) for each target was recorded.
[0184] It was noted that altering the 5' tail sequence of the primers altered sequencing performance, resulting in a decrease in overall sequencing performance. This is shown in Figure 1 by the overall decrease in the total reads of amplicons obtained with NNN-tailed primers compared to amplicons obtained with control primers, indicating that the use of NNN-tailed primers reduced topoisomerase-mediated binding of sequencing adapters to these amplicons.
[0185] FIG. 2 shows the total number of reads obtained after sequencing the amplification products obtained in Example 1 using control and NNN-tailed primers on an Oxford Nanopore Technologies MinION sequencer (comparing control and NNN-tailed primers for individual target genes).
[0186] Additionally, it was noted that the performance of the individual amplification products was also variable, in particular the increased representation of rpoB when NNN-tailed primers were used compared to the control.
[0187] Since multiple different NNN sequences were tested, the frequency with which each trimer sequence within the NNN motif was sequenced was determined. The number of reads for each different trimer sequence is shown in Table A, along with results for each individual target gene (forward (F) and reverse (R)) and all results combined. [Table 6-1] [Table 6-2] [Table 6-3]
[0188] FIG. 3 is a graph of the data presented in Table A, showing the total number of reads obtained at each NNN motif for all six targets.
[0189] It has been observed that certain motifs are more frequently observed than others and therefore associated with higher read counts than others. The 10 motifs shown in Table A to be associated with the highest read counts are triplet nucleotide sequences useful in the methods provided herein. In particular, the motifs GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, GCG were associated with the highest total read counts (shown on the left side of the graph in Figure 3).
[0190] The most commonly observed motifs were GGT and GGG (GGK, where K represents G or T). The least common was ACA. The 10 motifs shown in Table A to be associated with the highest number of reads are the most susceptible to modification by topo adaptors when present at the 5' end of one or both strands of multiple double-stranded target polynucleotides in the methods herein. Thus, targets with motifs selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT and GCG at the 5' end, particularly GGT or GGG (GGK), are susceptible to modification by topo adaptors.
[0191] The motifs shown in Table A to be associated with the lowest number of reads are the least susceptible to topo adaptor modification when present at the 5' end of one or both strands of a double-stranded target polynucleotide. Thus, targets with motifs selected from CAA, ATG, AAT, TAC, GCG, TAG, AAC, ATC, GCA, GCC, TAA, TTT, ATA, AAG, CCT, AAA, CCG, TTG, TTC, TCG, TCA, CCC, CCA, TCT, ACT, ACC, TCC, ACG, TTA and ACA at the 5' end are the least susceptible to topo adaptor modification. Targets with motifs selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA are even less susceptible. The triplet sequence ACA is the least susceptible.
[0192] Example 2 This experiment was performed using the same TB target gene as in Example 1 to determine whether the presence of a GGK tail would improve amplicon representation (the topo adapters generate more uniform modifications across the amplicon) and improve sequencing output.
[0193] Primer: The following primers were mixed to generate forward and reverse primer pairs for three TB targets (rpoB, gyrA, fabG1) with and without GGK tails. (In this and the following examples, primers designated as "GGK" primers with GGK tails contain a mixture of GGG and GGT tails obtained during primer production.) [Table 7] [Table 8]
[0194] PCR: For each of the six primer mixes, the following mixes were assembled: [Table 9]
[0195] Amplification was carried out according to the following conditions: [Table 10]
[0196] Samples were quantified using a Qubit system (ThermoFisher Scientific) and diluted to 25 ng / μL in water.
[0197] Agilent Analysis: The following reactions were set up for each of the six amplification products: Incubation was at 65° C. for 10 min and 80° C. for 2 min. Topo adapters were prepared according to the method described in Example 1. [Table 11]
[0198] A 2 μL sample of each was run on an Agilent 2100 Bioanalyzer 12000 chip.
[0199] Addition of the GGK tail significantly improved the modification of rpoB by the topo adapter, as well as the modification of gyrA (Fig. 4).
[0200] Figure 4 shows the percentage template modification of three different targets, control and GGK tailed, with 0x = no adapter attached, 1x = adapter attached to one end of the amplicon, and 2x = adapter attached to both ends of the amplicon, with each target shown from left to right.
[0201] Sequencing analysis: The following reactions were set up for the control and NNN-tailed templates: Incubations were at 65°C for 10 min and 80°C for 2 min. [Table 12]
[0202] The following steps were carried out:
[0203] -For SPRI purification (0.7x), wash twice with 70% EtOH and elute in 14 µL of EB.
[0204] -For the addition of sequencing adaptors, add 1 μL of RAP-F fast sequencing adaptors to 11 μL of eluted library and incubate at room temperature for 10 min before running on a MinION sequencer (Oxford Nanopore Technologies) for 12 h.
[0205] The total number of reads was measured. The reads were then downsampled and aligned to the TB genome, and the number of reads (forward and reverse) for each target was recorded.
[0206] Changing the primer ends to GGK improved sequencing performance (increased read numbers). The performance of the individual amplicons also changed, resulting in more uniform representation across the three different targets. Figure 5 shows the number of reads obtained for the control and GGK-tailed amplicons. Figure 6 shows the number of reads obtained, showing the control and GGK-tailed amplicons for the three target genes.
[0207] Example 3 This experiment was performed using a target sequence from Escherichia coli to determine whether the presence of a GGK tail would have the added benefit of improving modification by topo adapters, flattening amplicon representation, and improving sequencing output.
[0208] Primer: The following primers were mixed to generate forward and reverse primer pairs for the 12 E. coli targets, with or without a GGK tail. [Table 13] [Table 14]
[0209] PCR: Each of the 24 targets was amplified using NEB (New England Biolabs) Hot Start Master Mix according to the manufacturer's instructions.
[0210] -Samples were quantified using a Qubit system (ThermoFisher Scientific) and diluted to 25ng / μL in water.
[0211] Agilent Analysis: For each of the 24 amplification products, the following reactions were assembled: Incubation was at 65° C. for 10 min and 80° C. for 2 min. Topo adapters were prepared according to the method described in Example 1. [Table 15]
[0212] A 2 μL sample of each was run on an Agilent 2100 Bioanalyzer 12000 chip.
[0213] For the majority of targets containing GGK tails, modification by the topo adapters was significantly improved, especially for amplicons 2, 6, and 9 (Figure 7). Figure 7 shows the percentage template modification for 12 different targets, both control and GGK tailed, with 0x = no adapter attached, 1x = adapter attached to one end of the amplicon, and 2x = adapter attached to both ends of the amplicon, with each target shown from left to right.
[0214] Sequencing analysis: -Tailed or non-tailed amplification products were mixed and then modified with blunted topo adapters by incubating at 65°C for 10 min and 80°C for 2 min.
[0215] The -Topo-modified pool was SPRI purified (0.7x), washed twice with 70% EtOH, and eluted in 14 μL EB.
[0216] -For the addition of sequencing adaptors, add 1 µL of RAP-F to 11 µL of eluted library and incubate for 10 min at room temperature before running on a MinION sequencing instrument.
[0217] Reads were downsampled and aligned to the E. coli genome, and the number of reads (forward and reverse) for each target was recorded. In untailed samples, amplicon representation was highly variable, with particularly low representation of amplicons 2, 6, and 9 (consistent with low levels of topo-modification), and forward / reverse bias was also observed (Figure 8). Figure 8 shows the read counts for E. coli target and control amplicons.
[0218] The addition of the GGK tail resulted in a flatter overall representation of the target, with a more even split between forward and reverse reads (Figure 9). Figure 9 shows the read counts for the E. coli target, GGK-tailed amplicon.
[0219] Example 4 This experiment was performed to determine whether the presence of an ACA tail reduces modification by topoadapters.
[0220] Primer: The following primers were mixed to generate forward and reverse primer pairs for E. coli target 1 (from Example 3) with a GGK tail, an ACA tail or no tail. [Table 16]
[0221] PCR: -Each of the three targets was amplified using NEB Q5 Hot Start Master Mix according to the manufacturer's instructions.
[0222] -Samples were quantified using a Qubit system (ThermoFisher Scientific) and diluted to 25ng / μL in water.
[0223] Agilent Analysis: The following reactions were set up for each of the three amplification products: Incubation was at 65° C. for 10 min and 80° C. for 2 min. Topo adapters were prepared according to the method described in Example 1. [Table 17]
[0224] A 2 μL sample of each was run on an Agilent 2100 Bioanalyzer 12000 chip.
[0225] Topoisomerase modification was improved by the addition of a GGK tail, but was significantly inhibited by the addition of an ACA tail (Figure 10). Figure 10 shows the percentage of template modification for control, GGK and ACA tails. 0x = no adapter ligated, 1x = adapter ligated to one end of the amplicon, 2x = adapter ligated to both ends of the amplicon, with each target shown from left to right.
Claims
1. A method for preparing a nucleic acid library, (a) To provide a plurality of topoisomerase activation adapters, wherein each topoisomerase activation adapter contains vaccinia virus topoisomerase I bound to a double-stranded oligonucleotide, (b) To provide a plurality of double-stranded target polynucleotides, wherein each double-stranded target polynucleotide includes a triplet nucleotide sequence at the 5' end of one or both strands, and the triplet nucleotide sequence is selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC. (c) The topoisomerase activation adapter is brought into contact with the plurality of double-stranded target polynucleotides, and as a result, the topoisomerase covalently bonds the activation adapter to the 5' end of the double-stranded target polynucleotides, A method for preparing a nucleic acid library.
2. The method according to claim 1, further comprising generating the plurality of double-stranded target polynucleotides, each having the triplet nucleotide sequence at the 5' end of one or both strands, by PCR using a first oligonucleotide primer and a second oligonucleotide primer, wherein at least one of the first primer and the second primer has a 5' tail ending with the triplet nucleotide sequence.
3. The method according to claim 2, wherein both the first oligonucleotide primer and the second oligonucleotide primer include a 5' tail ending in the triplet nucleotide sequence.
4. The method according to claim 1, wherein the adapter is a sequence determination adapter.
5. The method according to claim 1, wherein the triplet nucleotide sequence is GGT or GGG.
6. The triplet nucleotide sequence at the 5' end of one or both strands of the plurality of double-stranded target polynucleotides is located immediately 5' to a second triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC. The method according to claim 1, wherein the second triplet nucleotide sequence at the 5' end of one or both strands of the plurality of double-stranded target polynucleotides is located immediately 5' to a third triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC.
7. (i) The double-stranded target polynucleotide is a blunt end, or (ii) The double-stranded target polynucleotide includes a nucleotide overhang, The method according to claim 1, wherein the double-stranded target polynucleotide is optionally A-tailed.
8. The method according to claim 1, wherein the adapter further comprises a click reaction group, a phosphor, a binder, a pull-down group, a tethering moiety, a marker, a modified base, a debased residue, an intermediate adapter, an amplification adapter, a hairpin adapter, a unique molecular identifier, a helicase binding site and / or a spacer.
9. The method further includes determining the presence or absence of the double-stranded target polynucleotide covalently bonded to the adapter, or one or more of its properties, (a) The target polynucleotide is brought into contact with a nanopore or a variant thereof, and as a result the target polynucleotide moves relative to the pore, (b) The measurement is performed once or more when the polynucleotide moves toward the pore, The method according to claim 1, wherein the presence or absence of the polynucleotide or one or more properties are determined accordingly.
10. A method for attaching a polynucleotide adapter to a double-stranded target polynucleotide, (a) To provide a topoisomerase activation adapter containing vaccinia virus topoisomerase I bound to a double-stranded oligonucleotide, (b) To provide a double-stranded targeted polynucleotide comprising a triplet nucleotide sequence at the 5' end of one or both strands, wherein the triplet nucleotide sequence is selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC, (c) The topoisomerase activation adapter is brought into contact with the double-stranded target polynucleotide, and as a result, the topoisomerase covalently bonds the activation adapter to the 5' end of the double-stranded target polynucleotide, A method for attaching a polynucleotide adapter to a double-stranded target polynucleotide.
11. A method for generating adaptive PCR amplification products, (i) Amplifying a first nucleic acid using a first oligonucleotide primer and a second oligonucleotide primer to produce a first amplification product, wherein at least one of the first primer and the second primer comprises a 5' tail ending in a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC; and amplifying a second nucleic acid using a third primer and a fourth primer to produce a second amplification product, wherein at least one of the third primer and the fourth primer comprises a 5' tail ending in a triplet nucleotide sequence that is not GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, or GGC; and (ii) A method comprising contacting the first amplification product and the second amplification product with a topoisomerase-activated adapter containing vaccinia virus topoisomerase I bound to the adapter, thereby causing the topoisomerase to covalently bond the adapter to the 5' end of the first amplification product and the 5' end of the second amplification product.
12. (i) At least one of the first primer and the second primer includes a 5' tail ending in the triplet nucleotide sequence GGT or GGG, and / or (ii) The 5' tail of at least one of the third primer and the fourth primer is terminated with a triplet nucleotide sequence selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG and AAA, The method according to claim 11, wherein at least one of the third primer and the fourth primer optionally includes a 5' tail ending in the triplet nucleotide sequence ACA.
13. It's a kit, -Vaccinia virus topoisomerase I, - A double-stranded oligonucleotide adapter containing a topoisomerase target site, A kit comprising a primer pair consisting of a first oligonucleotide primer and a second oligonucleotide primer, wherein at least one of the first and second primers is a 5'-tailed primer having a triplet nucleotide sequence selected from GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, and GGC at the 5' end of the tail.
14. (i) At least one of the first primer and the second primer includes a 5' tail ending in the triplet nucleotide sequence GGT or GGG, (ii) The kit further comprises a pair of primers comprising a third oligonucleotide primer and a fourth oligonucleotide primer, wherein at least one of the third oligonucleotide primer and the fourth oligonucleotide primer is a 5' tailed primer ending in a triplet nucleotide sequence that is not GGT, GGG, TGT, GTT, TGG, GGA, GTG, CGG, CGT, or GGC. Optionally, the 5' tail of at least one of the third and fourth primers is terminated with a triplet nucleotide sequence selected from ACA, TTA, ACG, TCC, ACC, ACT, TCT, CCA, CCC, TCA, TCG, TTC, TTG, CCG, and AAA. The kit according to claim 13, wherein, optionally, at least one of the third primer and the fourth primer comprises a 5' tail ending in the triplet nucleotide sequence ACA.
15. The kit according to claim 13 or 14, wherein the topoisomerase is loaded onto the adapter, and / or the adapter further comprises a click reaction group, a phosphor, a binder, a pull-down group, a tethering moiety, a marker, a modified base, a debasing residue, an intermediate adapter, an amplification adapter, a hairpin adapter, a unique molecular identifier, a helicase binding site, and / or a spacer.