Structure for preventing nucleic acid template from passing through nanopore during sequencing
Double hairpin adapters with self-priming capabilities address the challenge of nucleic acid translocation in nanopore sequencing by preventing strand passage and enhancing sequencing efficiency through molecular barcoding.
Patent Information
- Application Number
- JP2025146167
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-16
AI Technical Summary
Existing nucleic acid sequencing methods using nanopores face challenges in controlling or preventing the passage of nucleic acid templates through nanopores, which can complicate sequencing workflows and require additional components or steps.
The use of double hairpin adapters with self-priming capabilities, where both ends of the nucleic acid construct contain hairpin structures to prevent the displaced strand from threading through the nanopore, and include molecular barcodes for identification and sequencing.
This approach effectively controls and prevents nucleic acid translocation through nanopores, simplifying sequencing workflows by minimizing additional steps and components, while enabling efficient sequencing and identification of nucleic acid molecules.
Smart Images

Figure 2025183265000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The present invention relates to the field of nucleic acid sequencing, and more particularly to the field of forming libraries of nucleic acid targets for sequencing. [Background technology]
[0002] Background of the Invention Nucleic acid sequencing using biological and solid-state nanopores is a rapidly growing field. See Ameur, et al. (2019) Single molecule sequencing: toward clinical applications, Trends Biotech., 37:72. Some methods involve passing a nucleic acid template through a biological nanopore (U.S. Pat. No. 10,337,060) or solid-state nanopore (U.S. Pat. No. 10,288,599, U.S. Pat. App. Pub. No. 20180038001, U.S. Pat. No. 10,364,507), or a tunnel junction between two electrodes (PCT / EP2019 / 066199 and U.S. Pat. App. Pub. No. 20180217083). Other methods involve passing a detectable moiety (e.g., a label or tag) through the nanopore, rather than passing the template through the nanopore (U.S. Pat. No. 8,461,854). Methods exist to prevent or control the rate at which the template nucleic acid passes through the nanopore. For example, a "speed bump" is a complementary oligonucleotide placed outside the pore (U.S. Patent No. 10,400,278). The insertion rate of the template into the pore can be controlled by attaching a trefoil-structured tRNA to the template and interacting with a "brake" protein (U.S. Patent No. 10,131,944). Hairpins and loops can be present in the primer, or complementary probes can prevent penetration entirely in non-penetrating nanopore sequencing methods (U.S. Patent No. 9,605,309). The rate of penetration can also be controlled by using translation enzymes such as helicases (U.S. Patent Application Publication No. 20180201993). In solid-state sequencing, template nucleic acids are passed through the pore of a thin solid layer. Such penetration can be controlled using magnetic beads that "stretch" the nucleic acid strand (U.S. Patent Application Publication No. 20190317040).
[0003] There is a need for innovative and economical means to control or prevent nucleic acids from passing through biological or solid-state nanopores during sequencing, which ideally would add a minimal number of steps or components to complex sequencing workflows. Summary of the Invention
[0004] Summary of the Invention The present invention relates to the use of double hairpin adapters with self-priming capabilities. The 5'-hairpin is particularly advantageous for nanopore sequencing, preventing the displaced strand from threading through the nanopore. The present invention also includes methods for generating libraries for nucleic acid sequencing and control nucleic acid molecules for nanopore sequencing.
[0005] The target nucleic acid is ligated to a novel double hairpin adaptor (at one or both ends) such that both the 5' and 3' ends of the resulting nucleic acid construct contain hairpin structures. The 3'-hairpin has an extendable end that acts as a sequencing primer for the first strand. The second strand is displaced during primer extension, but both the first and second strands retain the hairpin and prevent it from threading through the nanopore.
[0006] In some embodiments, the present invention provides an adapter for a nucleic acid library, comprising a first strand and a second strand, wherein the first strand has a 5' portion and a 3' portion, the 5' portion forming a stem-loop structure with a loop and a stem including the 5' end of the first strand, and the 3' portion including a sequence complementary to the second strand; the second strand has a 5' portion and a 3' portion, the 3' portion forming a stem-loop structure with a loop and a stem including the 3' end of the second strand, and the 5' portion including a sequence complementary to the first strand; and the first strand and the second strand form a duplex via the 3' portion of the first strand and the 5' portion of the second strand. The 3' portion of the second strand can be extendable by a nucleic acid polymerase. The loop-forming region can be at least 4, 5, or 6 nucleotides in length and up to 20 or more nucleotides in length. The adapter can include one or more molecular barcodes selected from a sample barcode (SID) and a unique molecular identifier barcode (UID). For example, the SID can be located outside the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand, or the UID can be located within the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand. The SID and UID can comprise predefined sequences, or the UID can be a random sequence.
[0007] In some embodiments, the present invention provides a method for generating a library of nucleic acids, comprising attaching a plurality of adaptors to a plurality of double-stranded nucleic acids in a sample, each adaptor comprising: a first strand having a 5' portion and a 3' portion, the 5' portion forming a stem-loop structure with a loop and a stem including the 5' end of the first strand, and the 3' portion including a sequence complementary to the second strand; a second strand having a 5' portion and a 3' portion, the 3' portion forming a stem-loop structure with a loop and a stem including the extendible 3' end of the second strand, and the 5' portion including a sequence complementary to the first strand; and a first strand and a second strand forming a duplex via the 3' portion of the first strand and the 5' portion of the second strand. The adaptor may be attached by ligating the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand to one or both ends of the double-stranded nucleic acid. In some embodiments, prior to attachment, the plurality of nucleic acids are pretreated to form blunt ends at one or both ends of each nucleic acid. The plurality of nucleic acids may be further pretreated to add one or more non-template nucleotides to one strand at one or both ends of each nucleic acid. In some embodiments, the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand has a single-stranded overhang of one or more nucleotides. The adapter may include one or more molecular barcodes, such as a sample barcode (SID) and a unique molecular identifier barcode (UID). In some embodiments, the number of UIDs in the plurality of adapters may exceed the number of nucleic acids in the plurality of nucleic acids. In some embodiments, the number of nucleic acids in the plurality of nucleic acids exceeds the number of UIDs in the plurality of adapters.
[0008] In some embodiments, the present invention is a method of sequencing nucleic acids in a sample, the method comprising forming a library of nucleic acids as described herein and sequencing the nucleic acids by sequencing by synthesis, comprising extending the extendable 3' end of the second strand of the adapter. The sequencing by synthesis method can include nanopore-based detection.
[0009] In some embodiments, the present invention provides a control nucleic acid for use in a sequencing reaction, comprising a first strand, a second strand complementary to the first strand, and two ends, at least one of which is a 3'-overhang that forms a stem-loop structure, the 3'-end of which is extendable by a nucleic acid polymerase, and a 5'-end that forms the stem-loop structure upon displacement by the extended 3'-end. The loop-forming region of the loop formed by the 3'-end or 5'-end is at least 4, 5, or 6 nucleotides in length and at most 20 or more nucleotides in length. nucleotides in length.
[0010] In some embodiments, the invention is a method of sequencing a library of nucleic acids, the method comprising contacting the library with a control nucleic acid described herein and sequencing the library of nucleic acids by a method comprising nanopore detection. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram of a sequencing adapter with two hairpin loops. [Figure 2] FIG. 2 is a diagram of a sequencing adapter with a single hairpin loop. DETAILED DESCRIPTION OF THE INVENTION
[0012] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Sambrook et al., Molecular Cloning, A Laboratory Manual, 4 th Ed. Cold Spring Harbor Lab Press (2012).
[0013] The following definitions are provided to facilitate understanding of this disclosure.
[0014] The term "adapter" refers to a nucleotide sequence that can be added to another sequence to confer additional elements and properties to that sequence, including, but not limited to, barcodes, primer binding sites, capture moieties, labels, and secondary structures.
[0015] The term "barcode" refers to a nucleic acid sequence that can be detected and identified. Barcodes generally can be two or more nucleotides long and up to about 50 nucleotides long. Barcodes are designed to have at least a minimum number of differences from other barcodes in a population. A barcode can be unique to each molecule in a sample, unique to a sample, or shared by multiple molecules in a sample. The term "multiplex identifier," "MID," or "sample barcode" refers to a barcode that identifies a sample or the origin of a sample. Thus, all or substantially all MID-barcoded polynucleotides from a single source or sample share the same MID sequence, and all or substantially all (e.g., at least 90% or 99%) MID-barcoded polynucleotides from different sources or samples have different MID barcode sequences. Polynucleotides from different sources with different MIDs can be mixed and sequenced in parallel while maintaining the sample information encoded in the MID barcode. The term "unique molecular identifier" or "UID" refers to a barcode that identifies the polynucleotide to which it is attached. Typically, all or substantially all (eg, at least 90% or 99%) of the UID barcodes in a mixture of UID-barcoded polynucleotides are unique.
[0016] The term "DNA polymerase" refers to an enzyme that performs template-directed synthesis of polynucleotides from deoxyribonucleotides. DNA polymerases include prokaryotic Pol I, Pol II, Pol III, Pol IV, and Pol V, eukaryotic DNA polymerases, archaeal DNA polymerases, telomerases, and reverse transcriptases. The term "thermostable polymerase" refers to an enzyme that is heat-stable, thermotolerant, and retains sufficient activity to perform subsequent polynucleotide extension reactions, but Refers to an enzyme that is not irreversibly denatured (inactivated) when subjected to high temperatures for the time required to cause denaturation of double-stranded nucleic acids. In some embodiments, the following thermostable polymerases can be used: Thermococcus litoralis (Vent, GenBank: AAA72101), Pyrococcus furiosus (Pfu, GenBank: D12983, BAA02362), Pyrococcus woesii, Pyrococcus GB-D (Deep Vent, GenBank: AAA67131), Thermococcus kodakaraensis KODI (KOD, GenBank: BD175553, BAA06142; Thermococcus sp. strain KOD (Pfx, GenBank: AAE68738)), Thermococcus gorgonarius (Tgo, Pdb: 4699806), Sulfolobus solataricus (GenBank: NC002754, P26811), Aeropyrum pernix(GenBank:BAA81109), Archaeglobus fulgidus(GenBank:029753), Pyrobaculum aerophilum(GenBank:AAL63952), Pyrodictium occultum(GenBank:BAA07579, BAA07580), Thermococcus 9 degree Nm (GenBank:AAA88769, Q56366), Thermococcus fumicolans (GenBank:CAA93738, P74918), Thermococcus hydrothermalis (GenBank:CAC18555), Thermococcus sp.GE8 (GenBank:CAC12850), Thermococcus sp.JDF-3 (GenBank:AX135456;WO0132887), Thermococcus sp.TY(GenBank:CAA73475)、Pyrococcus abyssi(GenBank:P77916)、Pyrococcus glycovorans(GenBank:CAC12849)、Pyrococcus horikoshii(GenBank:NP 143776)、Pyrococcus sp.GE23(GenBank:CAA90887)、Pyrococcus sp.ST700(GenBank:CAC 12847)、Thermococcus pacificus(GenBank:AX411312.1)、Thermococcus zilligii(GenBank:DQ3366890)、Thermococcus aggregans、Thermococcus barossii、Thermococcus celer(GenBank:DD259850.1)、Thermococcus profundus(GenBank:E14137)、Thermococcus siculi(GenBank:DD259857.1) Thermococcus thioreducens, Thermococcus onnurineus NA1, Sulfolobus acidocaldarium, Sulfolobus tokodaii, Pyrobaculum calidifontis, Pyrobaculum islandicum(GenBank:AAF27815), Methanococcus jannaschii (GenBank:Q58295) , Desulfurococcus , TOK , Desulfurococcus , Pyrolobus , Pyrodictium , S taphylothermus, Vulcanisaetta, Methanococcus (GenBank:P52025) and a list of three genes of the genus B dynamite GenBank AAC62712, P956901, BAAA07579)) Manufacturer Thermus (flavus, ruber, thermophilus, lacteus, rubens, aquaticus), Bacillus stearothermophilus, Thermotoga maritima, Methanothermus fervidus, KOD, TNA1, Thermococcus. sp.9 degree of N-7, T4, T7, phi29, Pyrococcus furiosus, P. abyssi, T. gorgonarius, T. litoralis, T. zilligii, T. sp. GT, P. sp. GB-D, KOD, Pfu , T. gorgonarius, T. zilligii, T. litoralis, and Thermococcus sp. 9N-7 polymerases. In some cases, the nucleic acid (e.g., DNA or RNA) polymerase can be a modified naturally occurring A-type polymerase. Further embodiments of the invention broadly relate to methods in which the modified A-type polymerase, for example, in primer extension, end modification (e.g., terminal transferase, degradation, or polishing), or amplification reactions, can be selected from any species of the genera Meiothermus, Thermotoga, or Thermomicrobium. Another embodiment of the invention broadly relates to methods in which the polymerase, for example, in primer extension, end modification (e.g., terminal transferase, degradation, or polishing), or amplification reactions, can be isolated from any of Thermus aquaticus (Taq), Thermus thermophilus, Thermus caldophilus, or Thermus filiformis. Further embodiments of the invention broadly encompass methods in which a modified A-type polymerase may be isolated from Bacillus stearothermophilus, Sphaerobacter thermophilus, Dictoglomus thermophilum, or Escherichia coli, for example, in primer extension, end modification (e.g., terminal transferase, degradation, or polishing), or amplification reactions. In another embodiment, the invention broadly relates to methods in which a modified A-type polymerase may be a mutant Taq-E507K polymerase, for example, in primer extension, end modification (e.g., terminal transferase, degradation, or polishing), or amplification reactions. Another embodiment of the invention broadly relates to methods in which a thermostable polymerase may be used to perform amplification of a target nucleic acid.
[0017] The term "hairpin" refers to a secondary structure formed by a single strand of nucleic acid that includes at least one double-stranded region (a "stem"), where the stem-forming region is interrupted by a "loop" of single-stranded region. The size or relative sizes of the stem and loop regions are not specified.
[0018] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in single- or double-stranded form. Unless specifically limited, this term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as the explicitly indicated sequence.
[0019] The term "primer" refers to an oligonucleotide that binds to a specific region of a single-stranded template nucleic acid molecule and initiates nucleic acid synthesis via a polymerase-mediated enzymatic reaction. Typically, a primer contains fewer than about 100 nucleotides, preferably fewer than about 30 nucleotides. Target-specific primers specifically hybridize to a target polynucleotide under hybridization conditions. Such hybridization conditions may include, but are not limited to, hybridization in an isothermal amplification buffer (20 mM Tris-HCl, 10 mM (NH4)2SO4), 50 mM KCl, 2 mM MgSO4, 0.1% TWEEN® 20, pH 8.8, at 25°C) at a temperature of about 40°C to about 70°C. In addition to the target-binding region, a primer may have an additional region, typically at its 5' end. The additional region may include a universal primer binding site or a barcode.
[0020] The term "sample" refers to any biological sample containing nucleic acid molecules, typically including DNA or RNA. A sample can be tissue, cells, or extracts thereof, or can be a purified sample of nucleic acid molecules. The term "sample" refers to a sample that contains or contains a target nucleic acid. The term "sample" refers to any composition that is presumed to be a target sequence. The use of the term "sample" does not necessarily mean that the target sequence is present in the nucleic acid molecules present in the sample. A sample can be a tissue or fluid specimen isolated from an individual, such as skin, plasma, serum, cerebrospinal fluid, lymphatic fluid, synovial fluid, urine, tears, blood cells, organs, and tumors, as well as a sample of in vitro culture established from cells collected from an individual, including formalin-fixed paraffin-embedded tissue (FFPET) and nucleic acids isolated therefrom. A sample can also include acellular materials, such as cell-free blood fractions containing cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). A sample can be collected from a non-human subject or the environment.
[0021] The term "target" or "target nucleic acid" refers to a nucleic acid of interest in a sample. A sample can contain multiple targets as well as multiple copies of each target.
[0022] The term "universal primer" refers to a primer that can hybridize to a universal primer binding site, which can be a natural or artificial sequence that is typically added to a target sequence in a non-target-specific manner.
[0023] Nucleic acid sequencing using biological or solid-state nanopores is a rapidly developing field, with many technical solutions becoming available (Ameur, et al. See, e.g., (2019) Single molecule sequencing: towards clinical applications, Trends Biotech., 37:72. One common problem with nanopore sequencing is controlling the movement (threading) of nucleic acids through the pore. Some workflows involve simply controlling the speed at which the nucleic acid moves through the opening. Solutions include using helicases (US Patent Application Publication No. 20180201993) or magnetic particles (US Patent Application Publication No. 20190317040) to slow or control the movement rate. Some structures, such as "speed bump" complementary oligonucleotides (US Patent No. 10400278), can be attached to the pore. Other structures, such as tRNA structures attached to the template that interact with "brake" proteins in the pore complex (US Patent No. 10131944) or hairpins present on the sequencing primer (US Patent No. 9605309), are attached to the template molecule.
[0024] Some nanopore sequencing technologies do not require strand translocation through the nanopore. For such technologies, it is desirable to prevent translocation (threading) altogether. The present invention is suitable for both controlling and preventing translocation through the nanopore.
[0025] In addition to controlling the movement of the sequenced strand, discarding the complementary strand in the case of a double-stranded template is another issue. When a double-stranded nucleic acid template unwinds during a sequencing-by-synthesis (SBS) reaction, the non-sequenced strand must not thread through the nanopore in both threading and non-threading embodiments of nanopore sequencing.
[0026] The present invention includes methods and compositions for forming nucleic acid templates for sequencing and libraries of nucleic acid templates for sequencing that are suitable for any sequencing method but have particular advantages for nanopore sequencing.
[0027] In one embodiment, the invention is a novel adapter comprising a double-stranded portion for ligation to a nucleic acid to be sequenced, the adapter further comprising the novel feature of at least one stem-loop or hairpin structure at the opposite end of the double-stranded portion. In some embodiments, the adapter has a single stem-loop structure. In other embodiments, the adapter has two stem-loop structures. The adapter retains an extendable 3' end that can serve as a primer, e.g., a sequencing primer or an amplification primer. The adapter can include any feature useful in sequencing adapters, including a barcode, a primer binding site, or a capture moiety, for example, for the separation or purification of compatible nucleic acids. For example, the double-stranded portion of the adapter can include a unique molecular barcode (UID) that uniquely marks compatible nucleic acids, or a multiplexed sample barcode (SID) or (MID) that identically marks all nucleic acids in a sample as originating from the same source. The adapter can include a capture moiety, e.g., biotin, to capture compatible nucleic acids and separate them from non-compatible nucleic acids. A particular advantage of the adapter is the stem-loop or hairpin structure formed by at least one end of each strand in the compatible nucleic acid. The size of the loop is sufficient to inhibit or prevent the movement of the strand through the nanopore during sequencing. Depending on the nanopore used, the length of the loop-forming sequence can be 4, 5, 6, and up to 20 or more nucleotides to form a loop of sufficient size to prevent or inhibit translocation. Those skilled in the art will be able to determine the length of the loop-forming sequence because they have experimental or empirical knowledge of the sequence and secondary, tertiary, or quaternary structure of the nanopore-forming protein.
[0028] In one embodiment, the present invention is a method for preparing a library of matched nucleic acids for sequencing. A new adaptor is attached to one or both ends of each sample nucleic acid. The library can be purified or separated from unused adaptors and non-matched sample nucleic acids.
[0029] In yet another embodiment, the present invention provides a control molecule for sequencing a library of sample nucleic acids. The control molecule has ends with the novel structure described herein. In particular, the control nucleic acid molecule includes a first strand and a second strand complementary to the first strand. The control molecule further includes two ends, at least one of which includes a 3'-overhang that forms a stem-loop structure, the 3' end of which is extendable by a nucleic acid polymerase, and a recessed 5' end that also forms a stem-loop structure upon displacement by the extended 3' end. The loop size of the stem-loop structures of both strands is sufficient to inhibit or prevent the movement of the strands through the nanopore during sequencing.
[0030] The present invention involves simultaneously isolating and sequencing a target nucleic acid in a sample. In some embodiments, the sample is derived from a subject or patient. In some embodiments, the sample may include solid tissue or a fragment of a solid tumor derived from a subject or patient, for example, by biopsy. The sample may also include bodily fluids (e.g., urine, sputum, serum, plasma, or lymph, saliva, sputum, sweat, tears, cerebrospinal fluid, amniotic fluid, synovial fluid, pericardial fluid, ascites, pleural fluid, cyst fluid, bile, gastric juice, intestinal fluid, or fecal samples). The sample may include whole blood or blood fractions in which normal or tumor cells may be present. In some embodiments, the sample, particularly a liquid sample, may contain acellular material such as cell-free DNA or cell-free RNA, including cell-free tumor DNA or cell-free tumor RNA. In some embodiments, the sample is an acellular sample, for example, a cell-free blood-derived sample in which cell-free tumor DNA or cell-free tumor RNA is present. In other embodiments, the sample is a culture sample, e.g., a culture or culture supernatant containing or suspected of containing nucleic acid from cells in the culture or an infectious agent present in the culture. In some embodiments, the infectious agent is a bacterium, protozoan, virus, or mycoplasma.
[0031] A target nucleic acid is a nucleic acid of interest that may be present in a sample. Each target is characterized by its nucleic acid sequence. The present invention allows for the detection of one or more RNA or DNA targets. In some embodiments, the DNA target nucleic acid is a gene or gene fragment (including exons and introns) or an intergenic region, and the RNA target nucleic acid is a transcript or portion of a transcript to which a target-specific primer hybridizes. In some embodiments, the target nucleic acid comprises a locus of a genetic variant, e.g., a polymorphism including a single nucleotide polymorphism or single nucleotide variant (SNP or SNV), or a genetic rearrangement resulting in, for example, a gene fusion. In some embodiments, the target nucleic acid comprises a biomarker, i.e., a gene whose variants are associated with a disease or condition. For example, the target nucleic acid can be selected from a panel of disease-associated markers described in U.S. Patent Application No. 14 / 774,518, filed September 10, 2015. Such a panel is available as AVENIO ctDNA Analysis (Roche Sequencing Solutions, Pleasanton, Calif.). In other embodiments, the target nucleic acid is characteristic of a particular organism and aids in the identification of that organism or characteristics of the pathogenic organism, such as drug susceptibility or drug resistance. In yet other embodiments, the target nucleic acid is a combination of HLA or KIR sequences that define a unique characteristic of a human subject, e.g., the subject's unique HLA or KIR genotype. In yet other embodiments, the target nucleic acid is a somatic sequence, such as a rearranged immune sequence corresponding to an immunoglobulin (including IgG, IgM, and IgA immunoglobulins) or a T-cell receptor sequence (TCR). In yet another application, the target is a fetal sequence present in maternal blood, including a fetal sequence characteristic of a fetal disease or condition or a maternal condition associated with pregnancy.
[0032] In some embodiments, the target nucleic acid is RNA (including mRNA, microRNA, and viral RNA). In other embodiments, the target nucleic acid is DNA, including cellular DNA, or cell-free DNA (cfDNA), including circulating tumor DNA (ctDNA). The target nucleic acid may exist in a short or long form. Longer target nucleic acids may be fragmented. In some embodiments, the target nucleic acid is naturally fragmented, including, for example, circulating cell-free DNA (cfDNA) or chemically degraded DNA, such as that found in chemically preserved or old samples.
[0033] In some embodiments, the present invention includes a step of nucleic acid isolation. Generally, any nucleic acid extraction method that results in isolated nucleic acids, including DNA or RNA, can be used. Genomic DNA or RNA can be extracted from tissues, cells, or liquid biopsy samples (including blood or plasma samples) using solution-based or solid-phase-based nucleic acid extraction methods. Nucleic acid extraction may include detergent-based cell lysis, denaturation of nuclear proteins, and optionally removal of contaminants. Extraction of nucleic acids from preserved samples may further include a deparaffinization step. Solution-based nucleic acid extraction methods may include salting-out, organic solvent, or chaotrope methods. Solid-phase nucleic acid extraction methods include silica resin, anion exchange, or magnetic glass particles and paramagnetic beads (KAPA Pure Beads, Roche Sequencing Solutions, Pleasanton, Calif.) or AMPure beads (Beckman). Coulter, Brea, Cal.), but are not limited to these.
[0034] Typical extraction methods involve lysis of tissue material and cells present in the sample. The nucleic acids released from the lysed cells may be bound to a solid support (beads or particles) present in solution, a column, or a membrane. The nucleic acids may then undergo one or more washing steps to remove contaminants, including proteins, lipids, and their fragments, from the sample. Finally, the bound nucleic acids may be released from the solid support, column, or membrane and stored in an appropriate buffer until ready for further processing. Because both DNA and RNA must be isolated, nucleases should not be used, and care should be taken to inhibit nuclease activity during the purification process.
[0035] In some embodiments, the input DNA or RNA requires fragmentation. In such embodiments, RNA can be fragmented by a combination of heat and metal ions, such as magnesium. In some embodiments, the sample is heated to 85°-94°C for 1-6 minutes in the presence of magnesium. (See KAPA RNA HyperPrep DNA can be fragmented by physical means, such as sonication, using available equipment (Covaris, Woburn, Mass.), or by enzymatic means (KAPA Fragmentase Kit, KAPA Biosystems).
[0036] In some embodiments, the isolated nucleic acid is treated with DNA repair enzymes.In some embodiments, the DNA repair enzymes include DNA polymerases with 5'-3' polymerase activity and 3'-5' single-strand exonuclease activity, polynucleotide kinases that add 5' phosphates to dsDNA molecules, and DNA polymerases that add a single dA base to the 3' end of dsDNA molecules.End repair / A-tailing kits are available, such as Kapa Library Preparation Kits, including KAPA Hyper Prep and KAPA HyperPlus (Kapa Biosystems, Wilmington, Mass.).
[0037] In some embodiments, DNA repair enzymes target damaged bases in isolated nucleic acids. In some embodiments, the sample nucleic acid is partially damaged DNA from a preserved sample, such as a formalin-fixed paraffin-embedded (FFPET) sample. Deamination and oxidation of bases can result in incorrect base readings during the sequencing process. In some embodiments, the damaged DNA is treated with uracil N-DNA glycosylase (UNG / UDG) and / or 8-oxoguanine DNA glycosylase.
[0038] In some embodiments, the present invention includes an amplification step. The isolated nucleic acid can be amplified before further processing. This step can include linear or exponential amplification. The amplification can be isothermal or include thermocycling. In some embodiments, the amplification is exponential and includes PCR. In some embodiments, gene-specific primers are used for amplification. In other embodiments, a universal primer binding site is added to the target nucleic acid, for example, by ligating an adapter containing a universal primer binding site. All adapter-ligated nucleic acids have the same universal primer binding site and can be amplified using the same primer set. The number of amplification cycles using universal primers can be as few as 10, 20, or even about 30 or more cycles, depending on the amount of product required for subsequent steps. PCR using universal primers has reduced sequence bias, so there is no need to limit the number of amplification cycles to avoid amplification bias.
[0039] In some embodiments, the present invention utilizes an adapter nucleic acid composed of two strands having a structure described herein. In some embodiments, the adapter molecule is an artificial sequence synthesized in vitro. In other embodiments, the adapter molecule is a naturally occurring sequence synthesized in vitro. In yet other embodiments, the adapter molecule is an isolated naturally occurring molecule or an isolated non-naturally occurring molecule.
[0040] The double-stranded end of the hairpin adaptor can be ligated to a double-stranded nucleic acid molecule. The adaptor can be ligated at one or both ends of the double-stranded molecule.
[0041] The adaptor oligonucleotide has an overhang at the end that is ligated to the target nucleic acid. or blunt ends. In some embodiments, the novel adapters described herein comprise blunt ends to which blunt-end ligation of target nucleic acids can be applied. The target nucleic acid can be blunt-ended or can be made blunt by enzymatic treatment (e.g., "end repair"). In other embodiments, blunt-ended DNA is A-tailed, in which a single A nucleotide is added to the 3' end of one or both blunt ends. The adapters described herein are created so that a single T nucleotide extends from the blunt end to facilitate ligation between the nucleic acid and the adapter. Commercially available kits for performing adapter ligation include the AVENIO ctDNA Library Prep Kit or the KAPA HyperPrep and HyperPlus kits (Roche Sequencing Solutions, Pleasanton, CA). In some embodiments, the adapter-ligated DNA can be separated from excess adapters and unligated DNA.
[0042] In some embodiments, the adapter comprises additional features, such as a barcode, an amplification primer binding site, or a sequencing primer binding site.
[0043] FIG. 1 illustrates a novel adapter ligated to a nucleic acid of interest. Referring to FIG. 1, the adapter includes a first strand (top strand, "Adapter Oligo 1") and a second strand (bottom strand, "Adapter Oligo 2"). The adapter includes a double-stranded region of hybridization between the first adapter strand and the second adapter strand ("Adapter Oligo"). This region can be ligated to a target nucleic acid ("library insert"). Each strand also includes a hairpin region having a stem and a loop. Referring to the adapter illustrated in FIG. 1, the adapter is ligated to the target nucleic acid at one end and at one remaining unligated 5' end and one remaining unligated 3' end. The first (top) strand includes a non-extendable 5' end located in the double-stranded stem portion of the upper stem-loop structure. The second (bottom) strand includes a 3' end located in the double-stranded stem portion of the lower stem-loop structure. Its 3' end is extendable and can act as a primer to copy the bottom strand of the target nucleic acid without the need for a separate primer. In this embodiment, the 3' end of the adapter is a self-priming region. This self-priming region acts as a sequencing primer or an amplification primer.
[0044] In some embodiments, the adaptor-ligated nucleic acid is sequenced after adaptor ligation. In other embodiments, the adaptor-ligated target nucleic acid is amplified before sequencing. Each copy strand contains an adaptor sequence that can fold into a double hairpin structure as shown in FIG. 1. The size of the loop is sufficient to inhibit or prevent movement of either strand through the nanopore during sequencing. The length of the loop-forming region can be designed to be at least 3, 4, 5, or 6 nucleotides long and up to 20 or more nucleotides long, depending on the nanopore used for sequencing.
[0045] Figure 2 illustrates different embodiments of novel adapters ligated to a nucleic acid of interest. Referring to Figure 2, the adapter comprises a first strand (top strand, "Adapter Oligo 1") and a second strand (bottom strand, "Adapter Oligo 2"). The adapter comprises a double-stranded region of hybridization between the first adapter strand and the second adapter strand ("Adapter Oligo"). This region can be ligated to the target nucleic acid ("Library Insert"). The first (top) strand comprises a hairpin region with a stem and a loop. The second (bottom) strand comprises a capture moiety. In the embodiment shown in Figure 2, the capture moiety is a third oligonucleotide hybridized to the 3' end of the bottom adapter strand. It exists in
[0046] A capture moiety can be any moiety that can specifically interact with another capture molecule. Capture moiety-capture molecule pairs include avidin (streptavidin)-biotin, antigen-antibody, magnetic (paramagnetic) particle-magnet, or oligonucleotide-complementary oligonucleotide. The capture molecule can be bound to a solid support so that any nucleic acid presenting the capture moiety is captured on the solid support and separated from the rest of the sample or reaction mixture. In some embodiments, the capture molecule comprises a capture moiety for a secondary capture molecule. For example, the capture moiety can be an oligonucleotide complementary to a capture oligonucleotide (capture molecule). The capture oligonucleotide can be biotinylated and captured on streptavidin beads.
[0047] In some embodiments, adaptor-ligated nucleic acids are enriched by capturing the capture moiety and separating the adaptor-ligated target nucleic acids from unligated nucleic acids in the sample.
[0048] In some embodiments, the third oligonucleotide hybridized to the 3' end of the lower adaptor strand (Figure 2) serves as a sequencing primer or amplification primer. In some embodiments, the extension product of the third oligonucleotide is captured via a capture moiety. By capturing the extension product, the extension product is separated from the unligated sample nucleic acid and, optionally, from the target nucleic acid strand that does not have a capture moiety.
[0049] In some embodiments, the stem portion of the adapter includes modified nucleotides that increase the melting temperature of the capture oligonucleotide, such as 5-methylcytosine, 2,6-diaminopurine, 5-hydroxybutynyl-2'-deoxyuridine, 8-aza-7-deazaguanosine, ribonucleotides, 2'O-methylribonucleotides, or locked nucleic acids. In another aspect, the capture oligonucleotide is modified to inhibit digestion by nucleases, such as phosphorothioate nucleotides.
[0050] In some embodiments, the present invention includes an amplification step, for example, before ligating the novel adapters described herein. The primers can be target-specific. Target-specific primers contain at least a portion complementary to the target. If additional sequences, such as barcodes or second primer binding sites, are present, they are usually located in the 5' portion of the primer. The target can be a gene sequence (coding or non-coding) or a regulatory sequence present in RNA, such as an enhancer or promoter. The target can also be an intergenic sequence. In other embodiments, the primers are universal primers, capable of amplifying all nucleic acids in a sample, regardless of target sequence, for example. The universal primer anneals to a universal primer binding site added to nucleic acids in a sample by extending a primer with a universal primer binding site or by ligating an adapter (including an adapter with a novel structure described herein).
[0051] In some embodiments, the present invention utilizes barcodes. Detection of individual molecules typically requires molecular barcodes, such as those described in U.S. Patent Nos. 7,393,665, 8,168,385, 8,481,292, 8,685,678, and 8,722,368. Unique molecular barcodes are short, artificial sequences that are typically added to each molecule in a patient sample during the first steps of in vitro manipulation. The barcode marks the molecule and its progeny. Unique molecular barcodes (UIDs) have multiple uses. Barcodes can track each individual nucleic acid molecule in a sample, for example, identifying circulating tumor DNA (ctDNA) molecules in a patient's blood, to detect and monitor cancer without biopsy. This allows for the assessment of presence and quantity (Newman, A., et al., (2014) An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage, Nature Medicine doi:10.1038 / nm.3519).
[0052] When samples are mixed (multiplexed), the barcode can be a multiplexed sample ID (MID) used to identify the origin of the sample. The barcode can also serve as a unique molecular ID (UID) used to identify each original molecule and its progeny. A barcode can also be a combination of a UID and an MID. In some embodiments, a single barcode is used as both a UID and an MID. In some embodiments, each barcode contains a predefined sequence. In other embodiments, the barcode contains a random sequence. In some embodiments of the present invention, the barcodes are approximately 4-20 bases long, resulting in 96-384 different adapters (each with a different pair of identical barcodes) being added to a human genomic sample. Those skilled in the art will recognize that the number of barcodes depends on the complexity of the sample (i.e., the expected number of unique target molecules) and can generate an appropriate number of barcodes for each experiment.
[0053] Unique molecular barcodes can also be used for molecular counting and sequencing error correction. All descendants of a single target molecule are marked with the same barcode, forming a barcoded family. Sequence variations not shared by all members of the barcoded family are discarded as artifacts, not true mutations. Because the entire family represents a single molecule in the original sample (Newman, A., et al., (2016) Integrated digital error suppression for improved detection of circulating tumor DNA, Nature Biotechnology 34:547), barcodes can also be used for positional deduplication and target quantification.
[0054] In some embodiments, the number of UIDs in the plurality of adaptors can exceed the number of nucleic acids in the plurality of nucleic acids. In some embodiments, the number of nucleic acids in the plurality of nucleic acids exceeds the number of UIDs in the plurality of adaptors.
[0055] In some embodiments, the present invention is a library of target nucleic acids formed as described herein. The library comprises double-stranded nucleic acid molecules comprising nucleic acid targets present in the original sample. The nucleic acid molecules of the library further comprise novel adapters, as described herein, at one or both ends of the target nucleic acid sequences. The library nucleic acids may contain additional elements, such as barcodes and primer binding sites. In some embodiments, the additional elements are present in the adapters and are added to the library nucleic acids via adapter ligation. In other embodiments, some or all of the additional elements are present in the amplification primers and are added to the library nucleic acids by primer extension prior to adapter ligation. Amplification can be linear (comprising a single extension) or exponential, e.g., polymerase chain reaction (PCR). In some embodiments, some additional elements are added by primer extension, and the remaining additional elements are added by adapter ligation.
[0056] The utility of adapters and amplification primers for introducing additional elements into libraries of nucleic acids to be sequenced is described, for example, in U.S. Pat. Nos. 9,476,095, 9,260,753, 8,822,150, 8,563,478, 7,741,463, Nos. 8,182,989 and 8,053,192.
[0057] In some embodiments, the present invention further comprises a step of enriching for desired target nucleic acids. The desired nucleic acids can be enriched before forming a library according to the novel library formation methods described herein. Alternatively, enrichment can be performed after the library has been formed, i.e., on the molecules of the library.
[0058] In some embodiments, the method utilizes a pool of target-specific oligonucleotide probes (e.g., capture probes). Enrichment can be by subtraction, where the capture probes are complementary to undesired, abundant sequences, such as ribosomal RNA (rRNA) or abundantly expressed genes (e.g., globin). In subtraction, the undesired sequences are captured by the capture probes and removed from the target nucleic acid mixture or nucleic acid library and discarded. For example, the capture probes can contain binding moieties that can be captured on a solid support.
[0059] In other embodiments, the enrichment is capture and retention, where the capture probes are complementary to one or more target sequences, where the target sequences are captured and retained from a mixture of target nucleic acids or a library of nucleic acids by the capture probes, while the remaining solution is discarded.
[0060] In the case of enrichment, the capture probe can be free in solution or fixed on a solid support.The probe can be produced and amplified by the method described in, for example, U.S. Patent No. 9,790,543.The probe can also contain a binding moiety (e.g., biotin) and can be captured on a solid support (e.g., a support material containing avidin or streptavidin).
[0061] In some embodiments, the present invention includes an intermediate purification step. For example, any unused oligonucleotides, such as excess primers and excess adapters, are removed by a size selection method selected from gel electrophoresis, affinity chromatography, and size exclusion chromatography. In some embodiments, size selection can be performed using solid-phase reversible immobilization (SPRI) technology by Beckman Coulter (Brea, Calif.). In some embodiments, a capture moiety (FIG. 2) is used to capture and separate adapter-ligated nucleic acids from unligated nucleic acids or primer extension products from template strands.
[0062] The nucleic acid and nucleic acid library formed as described herein or its amplicon can be subjected to nucleic acid sequencing.Sequencing can be carried out by any method known in the art.Particularly advantageous is the high-throughput single-molecule sequencing method using nanopore.In some embodiments, the nucleic acid and nucleic acid library formed as described herein is sequenced by a method comprising passing through biological nanopore (US Pat. No. 10,337,060) or solid-state nanopore (US Pat. No. 10,288,599, US Patent Application Publication No. 20180038001, US Pat. No. 10,364,507).In other embodiments, sequencing comprises passing tags through nanopore (US Pat. No. 8,461,854) or any other currently existing or future DNA sequencing technology using nanopore.
[0063] In some embodiments, the sequencing step comprises analyzing the sequences. In some embodiments, the analysis comprises a step of sequence alignment. In some embodiments, alignment is used to determine a consensus sequence from multiple sequences, for example, multiple sequences with the same barcode (UID). In some embodiments, By using barcode (UID), consensus is determined from multiple sequences that all have the same barcode (UID).In another embodiment, by using barcode (UID), artifacts, that is, the variations that exist in sequences where some sequences have the same barcode (UID) but not all of them do.This artifact can be eliminated due to PCR error or sequencing error.
[0064] In some embodiments, the number of each sequence in a sample can be quantified by quantifying the relative number of sequences with each barcode (UID) in the sample. Each UID represents a single molecule in the original sample, and by counting the different UIDs associated with each sequence variant, the proportion of each sequence in the original sample can be determined. One skilled in the art can determine the number of sequence reads required to determine a consensus sequence. In some embodiments, a reasonable number is the number of reads per UID ("sequence depth") required for accurate quantification. In some embodiments, the desired depth is 5-50 reads per UID. [Example]
[0065] Example 1. Novel adapters for nucleic acid sequencing
[0066] The sequencing adapter consists of two oligonucleotides (adapter oligo 1 and adapter oligo 2) with complementary portions that allow them to form a partially double-stranded adapter. Adapter oligo 1 further contains sequenced DNA at its 5' end, which forms an intramolecular stem-loop structure under sequencing reaction conditions. The size of the loop structure can be adjusted by the length and sequence composition of the DNA to minimize the possibility of the displaced 5' end slipping through the sequencing nanopore. Adapter oligo 2 also contains a stem-loop-forming structure at its 3' end for nucleic acid polymerase binding and a free 3' end. This 3' end is extended during amplification or sequencing, and the strand ligated to adapter oligo 1 is ultimately displaced by the sequencing or amplification polymerase, while the 5' end loop restricts the displaced strand from slipping through the sequencing nanopore.
[0067] Example 2. Formation of control nucleic acids for nanopore sequencing
[0068] First, a DNA insert containing the desired endonuclease site, DNA primer annealing site, and self-complementary end sequence is cloned into a plasmid vector. The plasmid is propagated in a host bacterium; the plasmid is extracted and purified. The plasmid is then digested with an endonuclease, such as PmeI, that produces blunt-end cleavage to linearize it. Further digestion of the linearized plasmid with a nicking endonuclease, such as Nt.BbvCI, produces a short, cleaved, single-stranded fragment at the 5'-end of the linearized plasmid. The plasmid is then denatured and cooled in the presence of an oligo complementary to the cleaved fragment. The cleaved fragment hybridizes to its complementary oligo, which can be biotin-labeled to allow removal of the hybridized cleaved fragment using streptavidin bead purification. Cleavage of the 5' end of the plasmid allows the 3' end to form a secondary structure such as a hairpin (stem-loop) that prevents the single-stranded end from threading through the nanopore during DNA sequencing. The nascent 5' end of the plasmid is similarly designed so that upon displacement during polymerase extension of a primer oligo annealed to the single-stranded region at the 3' end, the 5' end also forms a secondary structure, preventing the displaced 5' end from threading through the nanopore.
Claims
1. 1. An adaptor for a nucleic acid library, comprising a first strand and a second strand, a. the first strand has a 5' portion and a 3' portion, the 5' portion forming a stem-loop structure having a loop and a stem that includes the 5' end of the first strand, and the 3' portion comprising a sequence complementary to the second strand; b. the second strand has a 5' portion and a 3' portion, the 3' portion forming a stem-loop structure having a loop and a stem that includes the 3' end of the second strand, and the 5' portion comprising a sequence complementary to the first strand; c. The first strand and the second strand form a duplex via the 3' portion of the first strand and the 5' portion of the second strand; adapter.
2. 2. The adapter of claim 1, wherein the 3' portion of the second strand is extendable by a nucleic acid polymerase.
3. 2. The adapter of claim 1, wherein one or both loop-forming regions are at least 4, 5, 6 nucleotides in length and up to 20 or more nucleotides in length.
4. 10. The adapter of claim 1, comprising one or more molecular barcodes.
5. 5. The adapter of claim 4, wherein the molecular barcode is selected from a sample barcode (SID) and a unique molecular identifier barcode (UID).
6. 5. The adapter of claim 4, wherein the SID is located outside the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand.
7. 5. The adapter of claim 4, wherein the UID is located within the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand.
8. The adapter of claim 4 , wherein the SID and the UID comprise a predefined sequence or a random sequence.
9. 1. A method for generating a library of nucleic acids, comprising attaching a plurality of adaptors to a plurality of double-stranded nucleic acids in a sample, each adaptor comprising: a. a first strand having a 5' portion and a 3' portion, the 5' portion forming a stem-loop structure having a loop and a stem that includes the 5' end of the first strand, and the 3' portion comprising a sequence complementary to a second strand; b. a second strand having a 5' portion and a 3' portion, the 3' portion forming a stem-loop structure having a loop and a stem that includes an extendible 3' end of the second strand, and the 5' portion comprising a sequence complementary to the first strand; and c. The first strand and the second strand form a duplex via the 3' portion of the first strand and the 5' portion of the second strand. A method comprising:
10. 10. The method of claim 9, wherein the adaptor is attached by ligating the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand to one or both ends of the double-stranded nucleic acid.
11. Prior to attachment, the plurality of nucleic acids may be pretreated to provide blunt ends at one or both ends of each nucleic acid. The method of claim 9 , wherein
12. 10. The method of claim 9, wherein the duplex formed by the 3' portion of the first strand and the 5' portion of the second strand has a single-stranded overhang of one or more nucleotides.
13. 10. A method of sequencing nucleic acids in a sample, comprising forming a library of nucleic acids of claim 9 and sequencing the library by sequencing by synthesis, comprising extending the extendable 3' end of the second strand of the adapter.
14. 1. A control nucleic acid for use in a sequencing reaction, comprising a first strand, a second strand complementary to the first strand, and two ends, at least one of which comprises: a. a 3'-overhang that forms a stem-loop structure, wherein the 3' end of the overhang is extendable by a nucleic acid polymerase; and b. A 5' end that forms a stem-loop structure when replaced by an extended 3' end A control nucleic acid comprising:
15. 15. A method for sequencing a library of nucleic acids, comprising contacting the library with a control nucleic acid of claim 14 and sequencing the library of nucleic acids by a method comprising detection by a nanopore.
Citation Information
Patent Citations
Direct RNA nanopore sequencing with help of a stem-loop reverse polynucleotide
US20200149101A1
Compositions and methods for selection of nucleic acids
US20200399690A1
Method for increasing throughput of single molecule sequencing by concatenating short DNA fragments
WO2018108328A1
Liquid sample workflow for nanopore sequencing
WO2020094457A1