Preparation of a single-stranded circular DNA template for single-molecule sequencing
The method of ligation with stem-loop adapters and exonuclease treatment enriches the sequencing library, addressing inefficiencies in nucleic acid sequencing by enabling separate sequencing of each strand and improving read lengths and yield.
Patent Information
- Application Number
- JP2024083207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-06-15
- Filing Date
- 2024-05-22
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2039-02-28
AI Technical Summary
Existing nucleic acid sequencing methods face inefficiencies due to the presence of linear nucleic acid by-products, which reduce sequencing performance, particularly when long target molecules are sequenced with separate single-pass reads or when each strand needs to be read separately.
A method involving ligation of double-stranded target nucleic acids with adapters forming a stem-loop structure, followed by exonuclease treatment to enrich the doubly-adapted nucleic acids, and subsequent cleavage at specific sites to generate extendable ends for separate sequencing of each strand, using exonucleases like exonuclease III and uracil-DNA-N-glycosylase, endonucleases, and adapters with cleavage sites or self-priming capabilities.
Enriches the sequencing library with long single molecules, eliminates the need for separate primers, reduces strand orientation bias, and allows for efficient sequencing of long target nucleic acids with high yield and accurate read lengths.
Smart Images

Figure 0007717903000003 
Figure 0007717903000004 
Figure 0007717903000005
Abstract
Description
Technical Field
[0001] The present invention relates to the field of nucleic acid analysis, and more particularly, to the preparation of templates for nucleic acid sequencing.
Background Art
[0002] Determination of the sequence of single molecule nucleic acids involves preparing a library of target molecules for the sequencing step. Linear nucleic acid libraries often coexist with linear nucleic acid by-products that reduce the performance of the sequencing method. There are library preparation methods that generate circular double-stranded templates that allow both strands of the target sequence to be read multiple times by successive polymerase reads. See U.S. Pat. Nos. 7,302,146 and 8,153,375. In some applications, it is necessary to read long target molecules with separate single-pass reads or to make each strand more desirable. The present invention is a method for efficiently generating a library and separately sequencing each strand of a target nucleic acid. The method has a number of advantages that are described in detail below.
Summary of the Invention
[0003] In some embodiments, the present invention is a method for separately sequencing each strand of a target nucleic acid, comprising: in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly-adapted target nucleic acid, wherein the adapter is a single strand that forms a double-stranded stem and a single-stranded loop, and the stem contains at least one strand cleavage site; contacting the reaction mixture with an exonuclease to thereby enrich the doubly-adapted target nucleic acid; contacting the reaction mixture with a cleavage agent to cleave the doubly-adapted target nucleic acid at the cleavage site to form extendable ends on each strand; and extending the extendable ends to thereby separately sequence each strand of the target nucleic acid, wherein the extension terminates at the cleavage site and does not proceed to the complementary strand. The ligation can be, for example, by ligation of sticky ends of the target nucleic acid and the adapter. The exonuclease can be selected from one or both of exonuclease III and exonuclease VII. The adapter can contain at least one barcode. In some embodiments, the cleavage site contains one or more deoxyuracils, and the cleavage agent contains uracil-DNA-N-glycosylase (UNG) and an endonuclease, such as endonuclease III, endonuclease IV, or endonuclease VIII. In some embodiments, the cleavage site contains one or more ribonucleotides, and the cleavage agent contains RNaseH. In some embodiments, the cleavage site contains one or more abasic sites, and the cleavage agent contains an endonuclease selected from endonuclease III, endonuclease IV, and endonuclease VIII. The adapter can contain exonuclease-protecting nucleotides, such as those containing phosphorothioate groups, etc.
[0004] In some embodiments, the adapter comprises a ligand for the capture moiety. For example, the ligand can be biotin or modified biotin, and the capture moiety comprises avidin or streptavidin. The ligand can be an adenosine sequence (oligo-dA), and the capture moiety comprises a thymidine sequence (oligo-dT).
[0005] In some embodiments, the method further comprises a target enrichment step, e.g., using a target-specific probe, prior to sequencing.
[0006] In some embodiments, the invention is a method of creating a library of target nucleic acids for separately sequencing each strand of a target nucleic acid, the method comprising, in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly adapter-ligated target nucleic acid, wherein the adapter is single-stranded and forms a double-stranded stem and a single-stranded loop, and the stem comprises a strand cleavage site; contacting the reaction mixture with an exonuclease to thereby enrich the doubly adapter-ligated target nucleic acid; and contacting the reaction mixture with a cleavage agent to cleave the doubly adapter-ligated target nucleic acid at the cleavage site to form extendable ends on each strand. In some embodiments, the invention is a method of determining the sequence of a library of target nucleic acids in a sample, the method comprising forming the library of target nucleic acids as described above; and extending the extendable ends to thereby separately sequence each strand of the target nucleic acids of the library, wherein the extension terminates at the cleavage site and does not proceed to the complementary strand.
[0007] In some embodiments, the present invention is a method for separately sequencing each strand of a target nucleic acid, comprising the steps of: in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly adapter-ligated target nucleic acid, wherein the adapter is a single strand that forms a double-stranded stem and a single-stranded loop, the loop contains a primer binding site, and the adapter further contains a strand synthesis termination site; contacting the reaction mixture with an exonuclease to thereby concentrate the doubly adapter-ligated target nucleic acid; contacting the reaction mixture with a primer capable of hybridizing to the primer binding site; and extending the primer to thereby separately sequence each strand of the target nucleic acid, wherein the extension terminates at the strand synthesis termination site and does not proceed to the complementary strand. The strand synthesis termination site can be selected from a nick, a gap, a depurinated nucleotide, and a non-nucleotide linker.
[0008] In some embodiments, the present invention is a method for creating a library for separately sequencing each strand of a target nucleic acid, comprising the steps of: in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly adapter-ligated target nucleic acid, wherein the adapter is a single strand that forms a double-stranded stem and a single-stranded loop, the loop contains a primer binding site, and the adapter further contains a strand synthesis termination site; and contacting the reaction mixture with an exonuclease to thereby concentrate the doubly adapter-ligated target nucleic acid and form a library of doubly adapter-ligated target nucleic acids. In some embodiments, the present invention is a method for determining the sequences of a library of target nucleic acids in a sample, comprising the steps of: forming the library of target nucleic acids as described above; contacting the library with a primer capable of hybridizing to the primer binding site; and extending the primer to thereby separately sequence each strand of the target nucleic acids in the library, wherein the extension terminates at the strand synthesis termination site and does not proceed to the complementary strand. BRIEF DESCRIPTION OF THE DRAWINGS
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0010] Detailed Description of the Invention Definitions The following definitions assist in the understanding of the present disclosure.
[0011] The term "sample" means any composition that contains or is presumed to contain a target nucleic acid. This includes tissue or liquid samples isolated from an individual, such as skin, plasma, serum, cerebrospinal fluid, lymph fluid, synovial fluid, urine, tears, blood cells, organs, and tumors, etc., and also in vitro culture samples established from cells taken from an individual, including formalin - fixed paraffin - embedded tissue (FFPET) and nucleic acids isolated therefrom. The sample also includes cell - free materials, such as cell - free DNA (cfDNA) or cell - free blood fractions containing circulating tumor DNA (ctDNA).
[0012] The term "nucleic acid" refers to a polymer of nucleotides (e.g., both natural and unnatural ribonucleotides and deoxyribonucleotides) including DNA, RNA, and their subcategories such as cDNA, mRNA, etc. Nucleic acids can be single-stranded or double-stranded and generally contain 5'-3' phosphodiester bonds, although in some cases nucleotide analogs may have other bonds. Nucleic acids may contain bases of natural origin (adenosine, guanosine, cytosine, uracil, and thymidine), as well as unnatural bases. Some examples of unnatural bases include, for example, those described in Seela et al., (1999) Helv. Chim. Acta 82:1640. Unnatural bases may have certain functions, such as increasing the stability of nucleic acid duplexes, inhibiting nuclease digestion, or preventing primer extension or strand polymerization.
[0013] The terms "polynucleotide" and "oligonucleotide" are used synonymously. A polynucleotide is a single-stranded or double-stranded nucleic acid. Oligonucleotide is a term sometimes used to describe shorter polynucleotides. Oligonucleotides are prepared by any suitable method known in the art, for example, by methods related to direct chemical synthesis as described in Narang et al. (1979) Meth. Enzymol. 68:90-99; Brown et al. (1979) Meth. Enzymol. 68:109-151; Beaucage et al. (1981) Tetrahedron Lett. 22:1859-1862; Matteucci et al. (1981) J. Am. Chem. Soc. 103:3185-3191.
[0014] The term "modified nucleotide" is used herein to describe a nucleotide in DNA that has a base other than the four conventional DNA bases consisting of adenosine, guanosine, thymidine, and cytosine. dA, dG, dC, and dT are conventional nucleotides. However, deoxyuracil (dU) and deoxyinosine (dI) are modified nucleotides in DNA. Also, ribonucleosides (rA, rC, rU, and rG) inserted into DNA are also considered "modified nucleotides" in the context of the present invention. Finally, non-nucleotide moieties (e.g., PEG) inserted in place of nucleotides into a nucleic acid strand are also considered "modified nucleotides" in the context of the present invention.
[0015] The term "primer" means a single-stranded oligonucleotide that hybridizes to the sequence of a target nucleic acid ("primer binding site") and acts as a starting point for synthesis along the complementary strand of the nucleic acid under conditions suitable for such synthesis. and is capable of doing so.
[0016] The term "adapter" means a nucleotide sequence that can be added to a sequence to incorporate additional properties into that sequence. An adapter is typically an oligonucleotide that can be single-stranded or double-stranded, or an oligonucleotide that can have both single-stranded and double-stranded portions. The term "adapter-ligated target nucleic acid" means a nucleic acid to which an adapter is conjugated at one or both ends.
[0017] The term "ligation" refers to a condensation reaction that links two nucleic acid strands in which the 5'-phosphate group of one molecule reacts with the 3'-hydroxyl group of another molecule. Ligation is typically an enzymatic reaction catalyzed by ligase or topoisomerase. Ligation can link two single strands to generate a single-stranded molecule. Ligation can also link the two strands belonging to a double-stranded molecule, thereby linking two double-stranded molecules. Ligation can also link both strands of one double-stranded molecule to both strands of another double-stranded molecule, thereby linking two double-stranded molecules. Ligation can also link the two ends of a strand within a double-stranded molecule, thereby repairing a nick in the double-stranded molecule.
[0018] The term "barcode" refers to a nucleic acid sequence that can be detected and identified. Barcodes can be incorporated into various nucleic acids. Barcodes are long enough, for example, 2, 5, 20 nucleotides, such that as a result, nucleic acids incorporating the barcode in a sample can be identified or grouped by the barcode.
[0019] The term "multiplex identifier" or "MID" refers to a barcode that identifies the source of a target nucleic acid (e.g., the sample from which the nucleic acid is derived). All or substantially all target nucleic acids from the same sample share the same MID. Target nucleic acids from different sources or samples can be mixed and sequenced simultaneously. By using MIDs, sequence reads can be assigned to the individual samples from which the target nucleic acids originated.
[0020] The term "unique molecular identifier" or "UID" refers to a barcode that identifies the nucleic acid to which it is attached. All or substantially all target nucleic acids from the same sample have different UIDs. All or substantially all progeny (e.g., amplicons) derived from the same original target nucleic acid share the same UID.
[0021] The terms "universal primer" and "universal priming binding site" or "universal priming site" mean a primer and a primer binding site that are present in different target nucleic acids (typically, via addition in vitro). The universal priming site is added to a plurality of target nucleic acids using an adapter or using a target-specific (non-universal) primer having a universal priming site at its 5' portion. A universal primer can bind to the universal priming site and induce primer extension therefrom.
[0022] More generally, the term "universal" means a nucleic acid molecule (e.g., a primer or other oligonucleotide) that can be added to any target nucleic acid and can perform its function independently of the target nucleic acid sequence. A universal molecule can perform its function by hybridizing to a complement, e.g., by hybridizing a universal primer to a universal primer binding site or by hybridizing a universal circularizing oligonucleotide to a universal primer sequence.
[0023] As used herein, the terms "target sequence", "target nucleic acid" or "target" mean a portion of a nucleic acid sequence in a sample that is detected or analyzed. The term target includes all variants of the target sequence, e.g., one or more mutant variants and wild-type variants.
[0024] The term "amplification" means a process of creating additional copies of a target nucleic acid. Amplification can have two or more cycles, e.g., multiple cycles of exponential amplification. Amplification may also have only one cycle (creation of a single copy of the target nucleic acid). The copies may have additional sequences, e.g., sequences present in the primers used for amplification. Amplification can also produce replication of only one strand (linear amplification), preferentially one strand (asymmetric PCR).
[0025] The term "sequence determination" means any method for determining the sequence of nucleotides in a target nucleic acid.
[0026] The term "self-priming adapter" means an adapter capable of initiating strand extension (strand replication) from the adapter itself. A self-priming adapter is contrasted with a conventional adapter that includes a primer binding site where separate primer molecules bind to the adapter to initiate strand extension from the primer.
[0027] Single molecule sequencing methods include the step of generating a library of adapter-ligated target nucleic acids. In some methods, the library is created from linear adapter-ligated target nucleic acids. During the workflow for preparing a linear library, a sequencing adapter (Y adapter or h adapter) is ligated to double-stranded DNA and then loaded onto a sequencer. Unfortunately, the ligation step is not 100% efficient, and partially ligated products or unligated products are generated. These by-products reduce the yield of effective sequencing and the performance of the instrument, for example, by competing for binding to the sequencing polymerase. To enrich for fully ligated products, modified adapters have been designed. This new adapter enables an exonuclease step that removes partially ligated products or unligated products.
[0028] The method of the present invention has a number of advantages. The method enables sequencing of long single molecules of a highly enriched library (due to the effect of exonuclease-mediated enrichment). Furthermore, the use of self-priming adapters streamlines the sequencing workflow by eliminating the need for another primer and primer annealing step. Still further, since the same adapter is ligated to both ends of each target nucleic acid and the same priming mechanism is used, there is no strand orientation bias. Still further, different types of target nucleic acids are compatible with these adapters, for example genomic DNA (gDNA) or amplification products. The size of the target nucleic acid and the final read length are limited only by the sequencing platform used and not by the library design.
[0029] The present invention includes the step of detecting a target nucleic acid in a sample. In some embodiments, the sample is derived from a subject or patient. In some embodiments, the sample may include a fragment of solid tissue or solid tumor derived from a subject or patient, for example by biopsy. The sample may also include a body fluid (e.g., urine, sputum, serum, plasma or lymph fluid, saliva, phlegm, sweat, tear fluid, cerebrospinal fluid, amniotic fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, cyst fluid, bile, gastric juice, intestinal juice, and / or fecal sample). The sample may include whole blood or a blood fraction in which tumor cells may be present. In some embodiments, the sample, particularly a liquid sample, may include cell-free material, e.g., cell-free DNA or RNA including cell-free tumor DNA or tumor RNA. The present invention is particularly suitable for analyzing rare and small amounts of targets. In some embodiments, the sample is a cell-free sample, e.g., a sample derived from cell-free blood in which cell-free tumor DNA or tumor RNA is present. In other embodiments, the sample is a cultured sample, e.g., a culture or culture supernatant containing or suspected of containing nucleic acids derived from an infectious agent or infectious agents. In some embodiments, the infectious agent is a bacterium, protozoan, virus or mycoplasma.
[0030] The target nucleic acid is the nucleic acid of interest that may be present in the sample. In some embodiments, the target nucleic acid is a gene or a gene fragment. In other embodiments, the target nucleic acid contains a genetic variant, such as a polymorphism including a single nucleotide polymorphism or variant (SNP or SNV), or a gene rearrangement that occurs, for example, in a gene fusion. In some embodiments, the target nucleic acid contains a biomarker. In other embodiments, the target nucleic acid is a characteristic of a particular organism, for example, useful in the identification of a pathogenic organism or a characteristic of a pathogenic organism, such as drug sensitivity or drug resistance. In still other embodiments, the target nucleic acid is a characteristic of a human subject, for example, an HLA or KIR sequence that defines the HLA or KIR genotype unique to the subject. In still other embodiments, all sequences in the sample are target nucleic acids, for example, in shotgun genome sequencing.
[0031] In embodiments of the present invention, the double-stranded target nucleic acid is converted to the template configuration of the present invention. In some embodiments, the target nucleic acid naturally exists in a single-stranded form (for example, RNA including mRNA, microRNA, viral RNA; or single-stranded viral DNA). The single-stranded target nucleic acid is converted to a double-stranded form to enable further steps of the claimed method.
[0032] Long target nucleic acids can be fragmented, but in some applications, long target nucleic acids may be desired to achieve long reads. In some embodiments, the target nucleic acid is naturally fragmented, for example, circulating cell-free DNA (cfDNA) or chemically degraded DNA, such as that identified in preserved samples. In other embodiments, the target nucleic acid is fragmented in vitro, for example, by physical means, such as sonication, or by endonuclease digestion, such as restriction digestion.
[0033] In some embodiments, the present invention includes a target enrichment step. The enrichment may be by capturing the target sequence via one or more target-specific probes. Nucleic acids in the sample can be denatured and contacted with single-stranded target-specific probes. The probes may include a ligand for an affinity capture moiety, such that after hybridization complexes are formed, they are captured by providing the affinity capture moiety. In some embodiments, the affinity capture moiety is avidin or streptavidin and the ligand is biotin or desthiobiotin. In some embodiments, this moiety is bound to a solid support. As will be described in more detail below, the solid support may include superparamagnetic spherical polymer particles, such as DYNABEADS™ magnetic beads or magnetic glass particles.
[0034] In some embodiments of the present invention, an adapter molecule is ligated to the target nucleic acid. The ligation can be blunt-end ligation or more efficient sticky-end ligation. The target nucleic acid or the adapter can be made blunt-ended by "end repair" including strand filling, i.e., by extending the 3'-end with a DNA polymerase to remove the 5'-overhang. In some embodiments, blunt-end adapters and target nucleic acids can be made sticky by adding a single nucleotide to the 3'-end of the adapter, and also by adding a single complementary nucleotide to the 3'-end of the target nucleic acid, for example, by a DNA polymerase or a terminal transferase. In still other embodiments, the adapter and the target nucleic acid can obtain sticky ends (overhangs) by digestion with a restriction endonuclease. The latter option is more advantageous for known target sequences that are known to contain a restriction enzyme recognition site. In some embodiments, other enzymatic steps may be required to achieve ligation. In some embodiments, a polynucleotide kinase may be used to add a 5'-phosphate to the target nucleic acid molecule and the adapter molecule.
[0035] In some embodiments, the adapter comprises a double-stranded portion (stem) and a single-stranded portion (loop) distal to the stem. The stem comprises a strand cleavage site (FIG. 3). In some embodiments, the loop comprises a primer binding site (FIG. 4). In such embodiments, the annealed primer may initiate replication of the strand. In some embodiments, the adapter comprises a double-stranded portion (stem) and a single-stranded portion (loop) distal to the stem that comprises a very small loop. Such an adapter can be a stem-loop adapter or a hairpin adapter (FIG. 7). The stem comprises a strand cleavage site that enables the cleaved adapter to self-prime, i.e., to initiate replication of the strand without a separate primer.
[0036] In some embodiments, the adapter molecule is an artificially synthesized sequence in vitro. In other embodiments, the adapter molecule is a sequence of natural origin synthesized in vitro. In yet other embodiments, the adapter molecule is an isolated molecule of natural origin.
[0037] In some embodiments, the invention includes introduction of a barcode into a target nucleic acid by ligation of a barcode-containing adapter. Sequencing of individual molecules typically requires molecular barcodes as described, for example, in U.S. Patent Nos. 7,393,665; 8,168,385; 8,481,292; 8,685,678; and 8,722,368. Unique molecular barcodes are short artificial sequences that are added to each molecule in a sample, e.g., a patient sample, typically during an early step of in vitro manipulation. The barcode marks the molecule and its progeny. Unique molecular barcodes (UIDs) have multiple uses. The barcode can track each individual nucleic acid molecule in a sample to assess, for example, the presence and amount of circulating tumor DNA (ctDNA) molecules in a patient's blood and detect and monitor cancer without a biopsy. See U.S. Patent Application Nos. 14 / 209,807 and 14 / 774,518. Unique molecular barcodes may also be used to correct sequencing errors. The progeny of a single target molecule are all marked with the same barcode, forming a barcoded family. Variations in sequences not shared by all members of a barcoded family are discarded as artifacts and not true mutations. Since the entire family represents a single molecule in the original sample, the barcode may also be used for positional de-duplication and target quantification. See the same reference.
[0038] In some embodiments of the invention, the adapter includes one or more barcodes. The barcode can be a multiplex sample ID (MID) used to identify the source of a sample in which samples are pooled (multiplexed). The barcode can also function as a UID used to identify each original molecule and its progeny. The barcode can also be a combination of a UID and an MID. In some embodiments, a single barcode is used as both a UID and an MID.
[0039] In some embodiments, each barcode comprises a predefined array. In other embodiments, the barcode comprises a random array. The barcode can be 1 to 40 nucleotides in length.
[0040] In the method of the present invention, the adapter comprises a cleavage site. The cleavage site is selected from modified nucleotides for which a specific endonuclease is available. A non-limiting list of examples of modified nucleotide - endonuclease pairs includes deoxyuracil - Uracil - DNA - N - glycosylase (UNG) and endonuclease; abasic site - AP endonuclease; 8 - oxoguanine - 8 - oxoguanine DNA glycosylase (also known as Fpg [formamidopyrimidine (fapy) - DNA glycosylase]); deoxyinosine - alkyladenine glycosylase (AAG) and endonuclease, and ribonucleotide - RNaseH. In some embodiments of the present invention, the cleavage site is used to generate an extendable 3'-end. In such embodiments, the cleavage site is located in the stem portion of the adapter, whereby replication of the complementary strand can be initiated.
[0041]
[0042] Different cleavage agents produce different products. In some embodiments, endonuclease VIII (Endo VIII), which produces a mixture of products containing 3'-P, is used. In other embodiments, endonuclease III (Endo III), which produces 3'-phospho-α,β-unsaturated aldehyde, is used. In still other embodiments, endonuclease IV (Endo IV), which produces a 3'-OH terminus, is used. A non-extendable terminus is advantageous in embodiments where separate sequencing primers are used. An extendable 3'-end (3'-OH) is advantageous when there are no separate sequencing primers and the sequencing reaction is self-primed by the extendable 3'-end.
[0043] In other embodiments, the cleavage site is used as a chain synthesis (elongation) terminator. In such embodiments, the cleavage site is placed anywhere in the adapter upstream of the primer binding site.
[0044] In some embodiments, the method includes contacting the reaction mixture with an endonuclease capable of cleaving the cleavage site under conditions under which such cleavage can occur.
[0045] In some embodiments, the sequencing reaction includes a self-priming step. Endonuclease strand cleavage at the cleavage site generates a free 3'-end that can be extended by sequencing by synthesis without using separate sequencing primers (Figure 7).
[0046] In some embodiments, the cleavage site acts as a chain termination step. By extending a strand (or primer) from one of the two adapters in the adapter-ligated target nucleic acid, the sequencing DNA polymerase reaches the cleavage site of the second adapter in the target nucleic acid. The strand break acts as a chain elongation terminator. In other embodiments, the modified nucleotide is a non-nucleotide polymer, such as polyethylene glycol (PEG), such as hexaethylene glycol (HEG). These moieties are not cleaved by endonucleases but act as chain synthesis terminators of the replicating nucleic acid strand.
[0047] In some embodiments, the method involves affinity capture of an adapter-ligated target nucleic acid or any other sequencing intermediate (e.g., the ternary complex of a pore protein, DNA polymerase, and template used in nanopore sequencing). For that purpose, the adapter may incorporate an affinity ligand (e.g., biotin) that enables the target to be captured by an affinity capture moiety (e.g., via streptavidin). In some embodiments, desthiobiotin is used. In some embodiments, the affinity capture utilizes an affinity molecule (e.g., streptavidin) bound to a solid support. The solid support can be a suspension in solution (e.g., glass beads, magnetic beads, polymer beads or other similar particles), or a solid-phase support (e.g., a silicon wafer, a glass slide, etc.). Examples of liquid-phase supports include superparamagnetic spherical polymer particles, such as DYNABEADS™ magnetic beads, or magnetic glass particles as described in U.S. Pat. Nos. 656,568; 6,274,386; 7,371,830; 6,870,047; 6,255,477; 6,746,874; and 6,258,531. In some embodiments, the affinity ligand is a nucleic acid sequence and the affinity molecule is a complementary sequence. In some embodiments, the solid substrate contains a poly-T oligonucleotide while the adapter contains at least a partially single-stranded poly-A portion.
[0048] In some embodiments, strand isolation is enhanced by various agents selected from single-stranded binding proteins, such as bacterial SSB, low complexity DNA C0t DNA (DNA enriched for repetitive sequences), or chemicals, such as alkalis, glycerol, urea, DMSO, or formamide.
[0049] In some embodiments, the invention includes an exonuclease digestion step after the adapter ligation step. The exonuclease removes any nucleic acid containing free ends from the reaction mixture. The exonuclease digestion enriches for the double adapter-ligated target nucleic acid that is topologically circular, i.e., does not contain free ends. Unligated target nucleic acids, nucleic acids ligated to only one adapter, and excess adapters are removed from the reaction mixture.
[0050] The exonuclease may be a single-strand specific exonuclease, a double-strand specific exonuclease, or a combination thereof. The exonuclease can be one or more of exonuclease I, exonuclease III, and exonuclease VII.
[0051] In some embodiments, the invention includes a method of creating a library of adapter-ligated target nucleic acids ready for sequencing as described herein, as well as a library generated by that method. Specifically, the library includes a collection of adapter-ligated target nucleic acids derived from the nucleic acids present in the sample. The adapter-ligated target nucleic acid molecules of the library each contain a target sequence ligated to an adapter sequence at each end and are topologically circular molecules that include cleavage sites and optionally primer binding sites.
[0052] In some embodiments, the invention includes detecting a target nucleic acid in a sample by nucleic acid sequencing. The plurality of nucleic acids, including all nucleic acids in the sample, can be converted into the library of the invention and sequenced.
[0053] In some embodiments, the method further includes removing damaged or fragmented targets from the library to improve the quality and length of sequencing reads. This step further includes contacting the library with one or more of uracil DNA N-glycosylase (UNG or UDG), AP nuclease, and Fpg (formamidopyrimidine [fapy]-DNA glycosylase), also known as 8-oxoguanine DNA glycosylase, to degrade such damaged target nucleic acids.
[0054] Sequencing can be performed by any method known in the art. Particularly advantageous is high-throughput single molecule sequencing capable of reading long target nucleic acids. Examples of such technologies include the Pacific Biosciences platform using SMRT (Pacific Biosciences, Menlo Park, Cal.), or platforms using nanopore technology, such as those manufactured by Oxford Nanopore Technologies (Oxford, UK), or Roche Sequencing Solutions (Roche Genia, Santa Clara, Cal.), and other existing or future DNA base sequencing technologies, regardless of the involvement of synthetic sequencing. The sequencing step can utilize sequencing primers specific to the platform.
[0055] Analysis and error correction In some embodiments, the sequencing step includes sequence analysis including a sequence alignment step. In some embodiments, the alignment is used to determine a consensus sequence from a plurality of sequences, for example, those having the same barcode (UID). In some embodiments, the barcode (UID) is used to determine a consensus from a plurality of sequences all having the same barcode (UID). In other embodiments, the barcode (UID) is used to remove artifacts, i.e., variations that are present in some sequences having the same barcode (UID) but not in all sequences. Such artifacts resulting from sample pretreatment or sequencing errors can be removed.
[0056] In some embodiments, the number of each sequence in a sample can be quantified by quantifying the relative number of sequences with each barcode (UID) in the sample. Each UID represents a single molecule in the original sample, and by counting different UIDs associated with each sequence variant, the proportion of each sequence in the original sample can be determined. One skilled in the art can determine the number of sequence reads necessary to determine a consensus sequence. In some embodiments, the relative number is the number of reads per UID ( "sequence depth") necessary for accurate quantitative results. In some embodiments, the desired depth is 5 to 50 reads per UID.
[0057] For example, a prior art method of forming a library of adapter-ligated target nucleic acids for amplification and sequencing is shown in FIG. 1. The library contains linear adapter-ligated target nucleic acids. In contrast, the method of the present invention (shown in FIG. 2) forms topologically closed nucleic acids that are resistant to exonuclease digestion and can be enriched using exonuclease. An embodiment of an adapter having one deoxyuracil (dU) is shown in FIG. 3. Other embodiments of adapters having multiple dUs are described in FIGS. 4 and 7. Yet another embodiment of the adapter is a non-nucleoside instead of dU ( a polymer of nucleotides, such as polyethylene glycol (PEG), such as having hexaethylene glycol (HEG) (Figure 5).
[0058] The method begins by ligating a stem-loop adapter to the ends of double-stranded nucleic acids in a reaction mixture (Figure 2). The resulting structure is a topologically circular (closed) nucleic acid lacking free 5' and 3' ends. Any unligated target nucleic acids and unused adapter molecules can be removed from the reaction mixture, thereby concentrating the adapter-ligated target nucleic acids (Figure 2). The adapter contains at least one deoxyuracil (dU) in the stem portion that serves as the site of strand cleavage (Figure 3). Next, uracil is cleaved using uracil-DNA glycosylase (e.g., UNG) to expose an abasic site in the DNA. Further, strand breaks are generated using an AP lyase, an endonuclease (e.g., endonuclease IV), or non-enzymatic reagents or conditions that produce a single-strand break (nick).
[0059] Next, the reaction mixture is contacted with a sequencing primer that can bind to the primer-binding site of the single-stranded portion of the adapter and initiate a sequencing primer extension reaction. The primer-binding site of the single-stranded portion may be in the loop portion (Figure 4). Also, the primer-binding site of the single-stranded portion may be in the region opposite the gap exposed after endonuclease cleavage (Figure 4). Primer extension proceeds to and terminates at the cleavage site (nick) of the opposite adapter of the adapter-ligated target nucleic acid.
[0060] In another embodiment, the adapter shown in FIG. 7 is used. This embodiment of the method also begins by ligating a stem-loop adapter to the ends of the double-stranded nucleic acid in the reaction mixture. In this embodiment, separate sequencing primers are not required, and the loop region of the adapter need not contain a sequencing primer binding site for the sequencing step. The adapter contains one or more deoxyuracil (dU) in the stem portion that serves as the strand cleavage site. The resulting adapter-ligated target nucleic acid is a topologically circular (closed) nucleic acid lacking free 5' and 3' ends. Any unligated target nucleic acid and unused adapter molecules can be removed from the reaction mixture, thereby concentrating the adapter-ligated target nucleic acid. For example, treatment with an exonuclease that targets free 5' or 3' ends can be used.
[0061] Next, uracil-DNA glycosylase (e.g., UNG) is used to cleave one or more uracils, exposing the apurinic / apyrimidinic site of the DNA. Further, strand breakage is performed using an AP lyase, an endonuclease (e.g., endonuclease IV or exonuclease VIII) or non-enzymatic reagents or conditions that generate a single-strand break (nick). When multiple deoxyuracils are present, gaps are generated. Next, the free 3' end of the strand in the nick (or gap) is extended in a sequencing strand extension reaction. In some embodiments, the adapter contains nucleotides that are resistant to exonuclease cleavage. These nucleotides (e.g., phosphorothioate nucleotides) are located near the cleavage site and protect the 3'-end that can be extended from exonuclease degradation. The extension proceeds to and terminates at the cleavage site (nick or gap) of the adapter on the opposite side of the adapter-ligated target nucleic acid.
Example
[0062] Example 1. Library formation and sequencing with a dU-containing stem-loop adapter. In this example, the stem-loop adapter illustrated in FIG. 3 was used. Adapters for SMRT-based sequencing systems (SEQ ID NOs: 3 and 4) and nanopore-based systems (SEQ ID NOs: 1 and 2) were created. Each adapter was created with and without deoxyuracil nucleotides in the stem region. The adapter sequence regions are shown in Table 1.
[0063]
Table 1
[0064] A fragment of the HIV genome (HIV Genewiz wild-type DNA (10 6 copies) amplicon size 1.1 kb) was used as the target nucleic acid. 2 μg of the purified PCR product was ligated to the adapter using a commercially available library preparation reagent (Kapa Biosystems, Wilmington, Mass.). After ligation, the reaction mixture was treated with exonuclease to remove any remaining linear molecules. In the case of the linear adapter (SEQ ID NO: 5), exonuclease was not used. The final library concentrations of ssDNA and dsDNA were determined using Qubit. In the next step, the adapter-ligated target nucleic acid was digested with uracil-DNA glycosylase (UDG) and DNA glycosylase-lyase, endonuclease VIII (USER™, New England Biolabs, Waltham, Mass.). In the case of adapters without uracil (SEQ ID NOs: 1, 3, and 5), USER™ was not used. After USER™ treatment, sequencing primers were added. The sequencing primer binding sites are in the loop region (FIG. 4). The resulting library was sequenced on the Pacific BioSciences RSII platform (Pacific Biosciences, Sequenced on the RSII platform (Illumina, Inc., San Diego, CA) and the nanopore platform (Roche Sequencing Solutions, Santa Clara, CA). Table 2 summarizes the quality and length of the sequencing reads on the RSII platform.
[0065]
Table 2
[0066] The quality and length of the sequencing reads on the nanopore platform are shown in FIG. 6.
[0067] Example 2 (Predictive). Library formation and sequencing with self-priming dU-containing hairpin adapters. In this example, the hairpin adapter illustrated in FIG. 7 is used. The adapter is SEQ ID NO: 6.
[0068] SEQ ID NO: 6: / 5phos / ATCTCTCTCAAATCCTCCTCCTCCGTTGGAGGAACGGAGGAGGA * G * G * AUUUGAGAGAGATT U - deoxyuracil; * - phosphorothioate nucleotide
[0069] The adapter is ligated to the target nucleic acid using a standard library preparation reagent (e.g., Kapa Biosystems). The reaction mixture containing the adapter-ligated target nucleic acid is treated with exonucleases Exo VII and Exo III to concentrate the adapter-ligated target nucleic acid. The concentrated adapter-ligated target nucleic acid is treated with UDG and endonuclease Endo IV. UDG cleaves U leaving an abasic site, and EndoIV cleaves the abasic site leaving an open 3'-OH. This step converts the hairpin adapter into a self-priming adapter. The adapter-ligated nucleic acid with the self-priming adapter is directly applicable to linear sequencing on a long-read platform, such as a nanopore platform. The sequencing polymerase binds to the 3'-OH of the self-priming adapter and extends the strand by replacing the strand forward in a linear single-pass sequencing.
[0070] Example 3 (Predictive). Library formation and sequencing using a hairpin adapter containing self-priming ribonucleotides. In this example, the hairpin adapter illustrated in FIG. 7 is used. The adapter has the sequence of SEQ ID NO: 7:
[0071] SEQ ID NO: 7: / 5phos / ATCTCTCTCTTTTCCTCCTCCTCCGTTGGAGGAACGGAGGAGGA * G * G * ArArArAGAGAGAGATT rA-ribonucleotide; * -phosphorothioate nucleotide
[0072] The adapter is ligated to the target nucleic acid using a standard library preparation reagent (e.g., Kapa Biosystems). The reaction mixture containing the adapter-ligated target nucleic acid is treated with exonucleases Exo VII and Exo III to concentrate the adapter-ligated target nucleic acid. The concentrated adapter-ligated target nucleic acid is treated with an RNaseH, such as RNaseH2, which cleaves ribonucleotides leaving the open 3'-OH. This step converts the hairpin adapter to a self-priming adapter. The adapter-ligated nucleic acid with the self-priming adapter is directly applied to linear sequencing on a long-read platform as in Example 2.
[0073] Although the invention has been described in detail with reference to specific examples, it will be apparent to those skilled in the art that various changes can be made within the scope of the invention. Accordingly, the scope of the invention is not limited by the examples described herein but is to be determined by the following claims.
Claims
1. A method for separately sequencing each strand of a target nucleic acid, comprising: a) in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly adapter-ligated target nucleic acid, wherein the adapter is a single strand forming a double-stranded stem and a single-stranded loop, the loop contains a primer binding site, the adapter further contains a chain synthesis termination site on the 3'-side of the primer binding site, the chain synthesis termination site is selected from a deoxyribonucleotide and a non-nucleotide linker, the deoxyribonucleotide is deoxyuridine, and the non-nucleotide linker is hexaethylene glycol (HEG); b) contacting the reaction mixture with an exonuclease to thereby enrich the doubly adapter-ligated target nucleic acid; c) when the chain synthesis termination site contains deoxyuridine, cleaving uracil with uracil-DNA glycosylase (UDG); d) contacting the reaction mixture with a primer capable of hybridizing to the primer binding site; and e) extending the primer to thereby separately sequence each strand of the target nucleic acid, wherein the extension terminates at the chain synthesis termination site and does not proceed to the complementary strand; The method as described above.
2. A method for creating a library for separately sequencing each strand of a target nucleic acid, comprising: a) in a reaction mixture, ligating the ends of a double-stranded target nucleic acid to an adapter to form a doubly adapter-ligated target nucleic acid, wherein the adapter is a single strand forming a double-stranded stem and a single-stranded loop, the loop contains a primer binding site, the adapter further contains a chain synthesis termination site on the 3'-side of the primer binding site, the chain synthesis termination site is selected from the group consisting of a deoxyribonucleotide and a non-nucleotide linker, the deoxyribonucleotide is deoxyuridine, and the non-nucleotide linker is hexaethylene glycol (HEG); and b) contacting the reaction mixture with an exonuclease to thereby enrich the doubly adapter-ligated target nucleic acid and form a library of doubly adapter-ligated target nucleic acids; Including, Here, when the chain synthesis termination site contains deoxyuridine, the library of double-stranded target nucleic acids is treated with uracil-DNA glycosylase (UDG) to cleave uracil at the position of deoxyuridine before contacting with a primer capable of hybridizing to the primer binding site. The method.
3. A method for determining the sequences of a library of target nucleic acids in a sample, comprising: a) forming a library of target nucleic acids by the method according to claim 2; b) when the chain synthesis termination site contains deoxyuridine, cleaving uracil with uracil-DNA glycosylase (UDG); c) contacting the library with a primer capable of hybridizing to the primer binding site; and d) extending the primer, thereby separately sequencing each strand of the target nucleic acids in the library, wherein the extension terminates at the chain synthesis termination site and does not proceed to the complementary strand, the step; The method comprising the above.
4. The method according to any one of claims 1 to 3, wherein the ligation to the adapter is by ligation.
5. The method according to any one of claims 1 to 4, wherein the adapter contains at least one barcode.
6. The method according to any one of claims 1 to 5, wherein the adapter contains exonuclease protection nucleotides.
7. The method according to any one of claims 1 to 6, wherein the adapter contains a ligand for the capture moiety. The method described therein.
8. The method according to any one of claims 1 to 7, further comprising a target enrichment step before sequencing.
9. The method according to claim 8, wherein the enrichment is by capture via a target-specific probe.
Citation Information
Patent Citations
Generation of double-stranded DNA templates for single-molecule sequencing
JP2021514646A
Oligonucleotide adaptors: compositions and methods of use
WO2012012037A1
Means and methods for amplifying nucleotide sequences
WO2017162754A1
Asymmetric templates and asymmetric method of nucleic acid sequencing
WO2018015365A1