Improvements in nucleic acid sequencing
By employing multiple hybridization and capture cycles at low DNA concentrations, the method addresses inefficiencies in high-throughput sequencing, improving nanowell occupancy and reducing waste, thus enhancing sequencing efficiency and yield.
Patent Information
- Application Number
- JP2021577991
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-09
- Filing Date
- 2021-03-09
- Publication Date
- 2026-01-08
- Estimated Expiration
- 2041-03-09
AI Technical Summary
Current high-throughput nucleic acid sequencing methods face inefficiencies with low-concentration template DNA, leading to low nanowell occupancy, increased duplication, and lower usable yields due to low template seeding concentrations, especially in high-density patterned flow cells.
A method involving multiple rounds of template hybridization and capture at low DNA concentrations, increasing the effective DNA concentration by repeating the process multiple times, thereby improving nanowell occupancy and reducing template waste.
This approach enhances the effective nucleic acid concentration, lowers library preparation requirements, and reduces template waste, while maintaining high sequencing throughput and accuracy.
Smart Images

Figure 0007795919000003 
Figure 0007795919000004 
Figure 0007795919000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to improvements in methods of high-throughput nucleic acid sequencing, and in particular to improvements in methods of preparing templates for high-throughput nucleic acid sequencing. [Background technology]
[0002] Nucleic acid sequencing methods have been known in the art for many years. Some such methods are based on the successive incorporation cycles of fluorescently labeled nucleic acid analogs. In such "sequencing by synthesis" or "cycle sequencing" methods, the identity of the added base is determined after each nucleotide addition by detecting the fluorescent label.
[0003] In particular, U.S. Patent No. 5,302,509 describes a method for sequencing a polynucleotide template that involves performing multiple extension reactions using a DNA polymerase or DNA ligase to sequentially incorporate labeled polynucleotides complementary to the template strand. In such a "sequencing by synthesis" reaction, a new paired polynucleotide strand based on the template strand is constructed in the 5' to 3' direction by sequentially incorporating individual nucleotides complementary to the template strand. The substrate nucleoside triphosphates used in the sequencing reaction are labeled at the 3' position with different 3' labels, allowing for the determination of the identity of the incorporated nucleotide as successive nucleotides are added.
[0004] In order to maximize the throughput of nucleic acid sequencing reactions, it is advantageous to be able to sequence multiple template molecules in parallel. Parallel processing of multiple templates can be achieved by using nucleic acid array technology. These arrays typically consist of a high-density matrix of polynucleotides immobilized on a solid support material.
[0005] Various methods for producing arrays of immobilized nucleic acid assays have been described in the art. Of particular interest, WO 98 / 44151 and WO 00 / 18957 both describe nucleic acid amplification methods that allow amplification products to be immobilized on a solid support to form arrays consisting of clusters or "colonies" formed from a plurality of identical immobilized polynucleotide strands and a plurality of identical immobilized complementary strands. The nucleic acid molecules present in the DNA colonies on the clustered arrays prepared according to these methods can provide templates for sequencing reactions, such as those described in WO 98 / 44152.
[0006] In current high-throughput next-generation sequencing (NGS) methods, sequencing chemistry is generally performed in a flow cell, into which nucleic acids and other reagents can be introduced. To prepare template strands, source nucleic acids are fragmented and the fragments are ligated to adapter sequences. These adapter sequences are designed to hybridize to short primer sequences immobilized on a solid support within the flow cell. The sequencing template is then amplified, resulting in clusters of identical template strands, which are then sequenced. Ideally, these clusters are of similar size and spaced apart from other clusters to achieve accurate resolution during imaging. Additional sequencing reactions are then performed in the flow cell, the results are imaged, and sequence reads are aligned to arrive at a final sequence for the template.
[0007] To maintain physical separation of different clusters, the flow cell can contain patterned nanowells formed by (for example) photolithography. Only the nanowells contain immobilized primer sequences; therefore, each cluster will form within the nanowell. The nanowell patterning determines the preferred cluster spacing and location. To maintain distinct clusters, it is desirable for a single template strand to initially hybridize within a given nanowell; therefore, template DNA is typically introduced into the flow cell in low molar amounts (as low as 10-20 pM). Additionally, amplification can be performed using exclusion amplification techniques, which allow for simultaneous seeding and amplification of template strands in the nanowells, thereby facilitating monoclonal clusters.
[0008] As NGS methods continue to improve, attempts are being made to increase the density and number of nanowells within the flow cell, enabling even greater sequencing throughput. However, the inventors have found that when using higher-density patterned flow cells, low molar concentrations of template DNA can be inefficiently used, leading to, for example, a low percentage of occupied nanowells. Low template seeding concentrations can result in increased levels of duplication and lower usable yields. Furthermore, many commonly used library preparation methods may not be amenable to modification to achieve increased template nucleic acid concentrations. To alleviate low sample concentrations, the inventors propose a modified method that utilizes the flow cell as a nucleic acid capture device to increase the effective usable concentration of low-concentration sequencing templates. Summary of the Invention
[0009] According to one aspect of the present invention, there is provided a method for preparing a template for a nucleic acid sequencing reaction, the method comprising: a) providing a solid support, the solid support comprising a plurality of nucleic acid primers immobilized thereon; b) contacting a nucleic acid library preparation with a solid support, the library preparation comprising a plurality of single-stranded fragments of a template nucleic acid to be sequenced, the template nucleic acid fragments further comprising one or more adapter nucleic acid sequences that hybridize to one or more of the nucleic acid primers, and the library preparation having a nucleic acid concentration of 400 pM or less; c) allowing the single-stranded fragment of template nucleic acid to bind to the nucleic acid primer, thereby immobilizing the single-stranded fragment on the solid support; d) repeating steps b) and c) at least three more times; Thereby, a solid support is provided having a single-stranded fragment of the template nucleic acid immobilized thereon.
[0010] According to a further aspect of the present invention there is provided a method of preparing a template for a nucleic acid sequencing reaction, the method comprising: a) providing a flow cell comprising a solid support having a plurality of nanowells formed thereon, each nanowell comprising a plurality of nucleic acid primers immobilized on the solid support, the nanowells being formed in a patterned array having a pitch of 500 nm or less; b) contacting a nucleic acid library preparation with a flow cell, the library preparation comprising a plurality of single-stranded fragments of a template nucleic acid to be sequenced, the template nucleic acid fragments further comprising one or more adapter nucleic acid sequences that hybridize to one or more of the nucleic acid primers, and the library preparation having a nucleic acid concentration of 400 pM or less; c) allowing the single-stranded fragments of the template nucleic acid to bind to the nucleic acid primer, thereby immobilizing the single-stranded fragments within the nanowell; d) repeating steps b) and c) at least three more times; Thereby, a flow cell is provided having single-stranded fragments of template nucleic acid immobilized within nanowells.
[0011] The inventors have determined that such a method addresses the drawbacks of using low-concentration nucleic acid libraries by enabling multiple "push" loading of the library. Each such push improves nanowell occupancy. Multiple rounds of template hybridization and capture at low DNA concentrations increase the effective DNA concentration, bringing it into the required range. This significantly lowers library preparation concentration requirements and reduces template waste due to system dead volume, etc.
[0012] The method can include repeating steps b) and c) at least four, five, six, seven, eight, nine, or more times. The effective nucleic acid concentration is believed to be given by the number of repeats multiplied by the original nucleic acid concentration. For example, four repeats of a 200 pM sample will result in an effective concentration of 800 pM. In a preferred embodiment of the present invention, the effective nucleic acid concentration is at least 800 pM, preferably at least 1000 pM.
[0013] The library preparation may have a nucleic acid concentration of 350 pM or less, 300 pM or less, 250 pM or less, or 200 pM or less. The library preparation may have a nucleic acid concentration of at least 50 pM, at least 100 pM, or at least 150 pM.
[0014] Repeated contacting steps can be performed with the same library preparation or with different library preparations. For example, the method can include washing unbound fragments of template nucleic acid from the flow cell and recovering and reintroducing the unbound fragments into the flow cell. However, in preferred embodiments, repeated contacting steps utilize fresh samples drawn from the same initial library preparation.
[0015] The method can include denaturing a library preparation containing double-stranded fragments of a template nucleic acid to be sequenced to obtain single-stranded fragments of the template nucleic acid to be sequenced. The denaturation step can be performed on a flow cell, for example, by introducing the double-stranded fragment library into the flow cell and subsequently denaturing it to provide a single-stranded fragment library that contacts the flow cell. In a preferred embodiment, denaturation can be performed before contacting the flow cell.
[0016] The method may further comprise amplifying the immobilized single-stranded fragment of nucleic acid, thereby generating multiple copies of the fragment. In a preferred embodiment, the amplification is carried out using exclusion amplification, e.g., as described in WO 2013 / 188582.
[0017] In a preferred embodiment of the present invention, the solid support is glass and includes a patterned nanowell array thereon. The nanowell array may be formed by photolithography. The pitch of the nanowell array (i.e., the center-to-center distance between nanowells) is preferably less than 750 nm, more preferably less than 500 nm, more preferably less than 400 nm, and most preferably 350 nm or less. The flow cell may include multiple lanes formed on the solid support, each lane including a portion of the patterned nanowell array. For example, the flow cell may be of the type manufactured by Illumina, Inc., San Diego, USA, for use in the NovaSeq6000 system. Exemplary methods for manufacturing certain types of solid supports having patterned nanowell arrays thereon, and flow cells including such solid supports, are described in EP 2961524.
[0018] The method may further comprise sequencing the immobilized single-stranded fragments of nucleic acid. DETAILED DESCRIPTION OF THE INVENTION
[0019] In its various aspects, the present invention relates generally to improvements in methods of high-throughput nucleic acid sequencing, and in particular to improvements in methods of preparing templates for high-throughput nucleic acid sequencing.
[0020] When referring to the attachment of a molecule (e.g., a nucleic acid) to a solid support, the terms "immobilized" and "attached" are used interchangeably herein, and both terms are intended to encompass direct or indirect, covalent or non-covalent attachment, unless otherwise indicated, either explicitly or by context. While covalent attachment may be preferred in certain embodiments of the present disclosure, generally, all that is required is that the molecule (e.g., a nucleic acid) remain immobilized or attached to the support under conditions under which the support is intended to be used, for example, in applications requiring nucleic acid amplification and / or sequencing. When referring to the attachment of a nucleic acid to another nucleic acid, the terms "immobilized" and "hybridized" are used herein and generally refer to hydrogen bonding between complementary nucleic acids.
[0021] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where explicit disclosure or context clearly dictates otherwise.
[0022] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquids. The substrate can be non-porous or porous. The substrate can optionally incorporate liquid (e.g., through porosity), but will typically be sufficiently rigid so that the substrate does not significantly swell when incorporating liquid and does not significantly shrink when the liquid is removed by drying. Non-porous solid supports are generally impermeable to liquids or gases. Exemplary solid support materials include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene with other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides, etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials, including silicon and modified silicon, carbon, metals, inorganic glasses, fiber optic bundles, and polymers. A particularly useful material is glass. Other suitable substrate materials can include polymeric materials, plastics, silicon, quartz (fused silica), borofloat glass, silica, silica-based materials, carbon, metals, optical fibers or fiber optic bundles, sapphire, or plastic materials such as COC and epoxy. A particular material can be selected based on properties desired for a particular use. For example, a material that is transparent to radiation of a desired wavelength is useful for analytical techniques that will utilize radiation of a desired wavelength, such as one or more of the techniques described herein. Conversely, it may be desirable to select a material that does not transmit radiation of a certain wavelength (e.g., opaque, absorbing, or reflective). This can be useful in forming masks used during the fabrication of structured substrates or for chemical reactions or analytical detection performed using structured substrates. Other properties of materials that can be utilized include inertness or reactivity to certain reagents used in downstream processes, or ease or low cost of manipulation during manufacturing processes. Further examples of materials that can be used in the structured substrates or methods of the present disclosure are described in U.S. Patent Application No. 13 / 661,524 and U.S. Patent Application Publication No. 2012 / 0316086 A1, each of which is incorporated herein by reference.
[0023] Certain embodiments of the present disclosure utilize solid supports consisting of a substrate or matrix (e.g., glass slides, polymer beads, etc.) that have been "functionalized" by the application of a layer or coating of an intermediate material containing reactive groups that allow for covalent attachment to biomolecules, such as polynucleotides. Examples of such supports include, but are not limited to, substrates such as glass. In such embodiments, biomolecules (e.g., polynucleotides) may be directly covalently attached to the intermediate material, or the intermediate material may itself be noncovalently attached to a substrate or matrix (e.g., a glass substrate). The term "covalent attachment to a solid support" should be interpreted accordingly to encompass this type of arrangement. Alternatively, substrates such as glass may be treated to allow for direct covalent attachment of biomolecules; for example, glass may be treated with hydrochloric acid, thus exposing hydroxyl groups on the glass, and phosphite-triester chemistry may be used to directly attach nucleotides to the glass via covalent bonds between the hydroxyl groups on the glass and the phosphate groups on the nucleotides.
[0024] In embodiments of the present invention, covalent attachment can be achieved through a sulfur-containing nucleophile, such as a phosphorothioate, present at the 5' end of the polynucleotide chain.
[0025] As used herein, the term "nanowell" refers to a discrete, concave feature in a solid support having a surface opening completely surrounded by an interstitial region of the surface. The well can have any of a variety of shapes at the surface opening, including, but not limited to, circular, elliptical, square, polygonal, star-shaped (with any number of vertices), and the like. The cross-section of the well taken perpendicular to the surface can be curved, square, polygonal, hyperbolic, conical, angular, and the like. In a preferred embodiment of the present invention, a nanowell array comprises a plurality of nanowells configured on a solid support, and a patterned nanowell array comprises a repeating arrangement of nanowells such that the relative arrangement of nanowells in one portion of the solid support is the same as the relative arrangement of nanowells in at least one other portion of the solid support. The pitch of a nanowell array is the center-to-center distance between two adjacent nanowells.
[0026] As will be understood by those skilled in the art, double-stranded nucleic acids are typically formed from two complementary polynucleotide strands composed of deoxyribonucleotides linked by phosphodiester bonds, but may further contain one or more ribonucleotides and / or non-nucleotide chemical moieties and / or non-naturally occurring nucleotides and / or non-naturally occurring backbone linkages. In particular, double-stranded nucleic acids may contain non-nucleotide chemical moieties, such as linkers or spacers, at the 5'-end of one or both strands. Non-limiting examples of double-stranded nucleic acids include methylated nucleotides, uracil bases, phosphorothioate groups, peptide conjugates, and the like. Such non-DNA or non-natural modifications may be included to impart certain desirable properties to the nucleic acid, such as to enable covalent attachment to a solid support or to act as a spacer to position the cleavage site at an optimal distance from the solid support. A single-stranded nucleic acid consists of one such polynucleotide strand. Even if a polynucleotide strand is only partially hybridized to a complementary strand, for example, a long polynucleotide strand hybridized to a short nucleotide primer, it may still be referred to herein as a single-stranded nucleic acid.
[0027] The template nucleic acid to be sequenced will contain a "target" region that is desired to be sequenced, either completely or partially. The nature of the target region is not limited to the present invention. It can be a previously known or unknown sequence, for example, derived from a genomic DNA fragment, cDNA, etc. The template nucleic acid molecule also contains non-target sequences, for example, at the 5' and 3' ends of one or both strands (if double-stranded), flanking the target region. If the template nucleic acid is formed by solid-phase nucleic acid amplification, these non-target sequences may be derived from the primers used in the amplification reaction. Alternatively, non-target sequences may be ligated to fragmented target sequences to incorporate them into the nucleic acid molecule.
[0028] In an embodiment of the present invention, double-stranded nucleic acid can be subjected to denaturing conditions to provide single-stranded nucleic acid. Suitable denaturing conditions will be clear to the skilled reader with reference to standard molecular biology protocols (Sambrook et al., 2001, Molecular Cloning, A Laboratory Manual, 3rd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor Laboratory Press, NY; Current Protocols, eds. Ausubel et al.).
[0029] Denaturation (and subsequent reannealing of the cleaved strand) results in the generation of a sequencing template that is partially or substantially single-stranded. The sequencing reaction can then be initiated by hybridization of a sequencing primer to the single-stranded portion of the template. In embodiments of the invention, sequencing can be carried out using a strand-displacing polymerase enzyme.
[0030] In embodiments of the present invention, the term "solid support" as used herein refers to a material to which polynucleotide molecules are attached. Suitable solid supports are commercially available and will be apparent to those skilled in the art. Supports can be made from materials such as glass, ceramic, silica, and silicon. Supports with gold surfaces can also be used. Supports typically include a flat (planar) surface, or at least a structure in which the polynucleotides under investigation are in approximately the same plane. Supports of any suitable size can be used. Alternatively, the solid support can be non-planar, for example, a microbead.
[0031] In an embodiment of the present invention, the method of the present invention can be used to prepare templates for nucleic acid sequencing. Immobilized single-stranded nucleic acids can be amplified to provide clustered arrays of nucleic acid colonies generated by solid-phase nucleic acid amplification. In this context, the term "solid-phase amplification" refers to an amplification reaction similar to standard PCR, except that forward and / or reverse amplification primers are immobilized (e.g., covalently attached) to a solid support at or near their 5' ends. Thus, the products of the PCR reaction are extended strands induced by extension of amplification primers immobilized on a solid support at or near their 5' ends. Solid-phase amplification itself can be carried out using procedures similar to those described, for example, in WO 98 / 44151 and WO 00 / 18957.
[0032] As a first step in colony generation by solid-phase amplification, a mixture of forward and reverse amplification primers can be immobilized or "grafted" onto the surface of a suitable solid support. The grafting step will generally involve covalently attaching the primers to the support at or near their 5' ends, leaving the 3' ends available for primer extension.
[0033] Amplification primers are typically oligonucleotide molecules having the following structure: Forward primer: AL-S1 Reverse primer: AL-S2
[0034] where A represents an optional moiety that allows for attachment to a solid support, L represents an optional linker moiety, and S1 and S2 are polynucleotide sequences that allow for amplification of a substrate nucleic acid molecule that includes the target region that is desired to be sequenced (fully or partially).
[0035] The mixture of primers grafted onto the solid support will generally contain substantially equal amounts of forward and reverse primers.
[0036] The A group can be any moiety (including non-nucleotide chemical modifications) that allows for attachment (preferably covalent attachment) to a solid support. In embodiments of the invention, the A group can include a sulfur-containing nucleophile, such as a phosphorothioate, present at the 5'-end of a polynucleotide chain. Alternatively, the A group can be omitted if suitable chemistries are used to directly attach either the linker or the nucleic acid to the solid support.
[0037] L represents a linker or spacer that may be included, but is not strictly necessary. The linker may be included to ensure that the cleavage site present in the immobilized polynucleotide molecule produced as a result of the amplification reaction is positioned at an optimal distance from the solid support, or the linker itself may contain the cleavage site.
[0038] The linker has the formula (CH2) n where "n" is from 1 to about 1500, for example less than about 1000, preferably less than 100, for example 2 to 50, particularly 5 to 25. However, a variety of other linkers may be used, the only restriction imposed on their structure being that the linker is stable under the conditions under which the polynucleotide is intended to be subsequently used, for example, under the conditions used in DNA amplification and sequencing.
[0039] Linkers that are not composed solely of carbon atoms may also be used, such as polyethylene glycol (PEG).
[0040] Linkers formed primarily from a chain of carbon atoms and PEG can be modified to contain chain-interrupting functional groups. Examples of such groups include ketones, esters, amines, amides, ethers, thioethers, sulfoxides, and sulfones. Separately, or in combination with the presence of such functional groups, alkenes, alkynes, aromatic or heteroaromatic moieties, or cycloaliphatic moieties (e.g., cyclohexyl) can be used. A cyclohexyl or phenyl ring can be used to link, for example, PEG or (CH2) n They can be connected to the chain through their 1 and 4 positions.
[0041] As an alternative to the above linkers, which are based on a linear chain of predominantly saturated carbon atoms, optionally interrupted by unsaturated carbon atoms or heteroatoms, other linkers based on nucleic acids or monosaccharide units (e.g., dextrose) can be envisaged. It is also within the scope of the present invention to utilize peptides as linkers.
[0042] In further embodiments, the linker may comprise one or more nucleotides. Such nucleotides may also be referred to herein as "spacer" nucleotides. Typically, 1 to 20, more preferably 1 to 15, or 1 to 10, more specifically 2, 3, 4, 5, 6, 7, 8, 9, or 10 spacer nucleotides may be included. Most preferably, the primer will comprise 10 spacer nucleotides. It is preferred to use a poly-T spacer, although other nucleotides and combinations thereof may be used. In a preferred embodiment, the primer may comprise 10T spacer nucleotides.
[0043] To allow the primer grafting reaction to proceed, a mixture of amplification primers is provided to the solid support under conditions that allow reaction between moiety A (if present) and the support, or between the nucleic acid and the support. The solid support may be suitably functionalized to allow covalent attachment via moiety A. The result of the grafting reaction is a substantially uniform distribution of primers across at least a portion of the solid support. When the solid support comprises nanowells, in preferred embodiments, the primers are restricted to the locations of the nanowells and are absent from the interstitial regions of the solid support.
[0044] The nucleic acid library preparation is typically contacted with the flow cell in free solution. The amplification reaction can then proceed substantially as described in WO 98 / 44151. Briefly, following primer binding, the solid support is contacted with the template to be amplified under conditions that allow hybridization between the template and the immobilized primer. The template is usually added to the free solution under suitable hybridization conditions, which will be apparent to the skilled reader. Typically, the hybridization conditions are, for example, 5×SSC at 40°C. Solid-phase amplification can then proceed, with the first step being a primer extension step, in which nucleotides are added to the 3′ end of the immobilized primer hybridized to the template, generating a fully extended complementary strand. This complementary strand will therefore contain a sequence at its 3′ end that can bind to a second primer molecule immobilized on the solid support. Further rounds of amplification (similar to standard PCR reactions) result in the formation of clusters or colonies of template molecules bound to the solid support. Other amplification procedures can be used and will be known to those skilled in the art, for example, amplification can be isothermal amplification using a strand-displacing polymerase or can be exclusion amplification as described in WO 2013 / 188582.
[0045] The sequences S1 and S2 in the amplification primers can be specific to a particular target nucleic acid that is desired to be amplified, but in other embodiments, the sequences S1 and S2 can be "universal" primer sequences that allow for the amplification of any target nucleic acid of known or unknown sequence that has been modified to allow amplification by a universal primer.
[0046] Suitable nucleic acids to be amplified with universal primers can be prepared by modifying a polynucleotide containing the target region to be amplified (and sequenced) by adding known adapter sequences to the 5' and 3' ends of the target polynucleotide to be amplified. The target molecule itself can be any polynucleotide molecule that is desired to be sequenced (e.g., random fragments of human genomic DNA). The adapter sequences allow for the amplification of these molecules on a solid support to form a cluster using forward and reverse primers with the general structure described above, where sequences S1 and S2 are universal primer sequences.
[0047] Adapters are typically short oligonucleotides that can be synthesized by conventional means. Adapters can be attached to the 5' and 3' ends of target nucleic acid fragments by various means (e.g., subcloning, ligation, etc.). More specifically, two different adapter sequences are attached to the amplified target nucleic acid molecule such that one adapter is attached to one end of the target nucleic acid molecule and another adapter is attached to the other end of the target nucleic acid molecule. The resulting construct containing the target nucleic acid sequence flanked by adapters may be referred to herein as a "substrate nucleic acid construct." The target polynucleotide can be advantageously size-divided prior to modification with the adapter sequences.
[0048] The adapter contains sequences that allow nucleic acid amplification using amplification primer molecules immobilized on a solid support. These sequences within the adapter may be referred to herein as "primer binding sequences." To serve as a template for nucleic acid amplification, a single strand of the template construct must contain a sequence complementary to sequence S1 in the forward amplification primer (so that the forward primer molecule can bind and prime synthesis of the complementary strand) and a sequence corresponding to sequence S2 in the reverse amplification primer molecule (so that the reverse primer molecule can bind to the complementary strand). The sequence within the adapter that allows hybridization to the primer molecules will typically be about 20-40 nucleotides in length, although the invention is not limited to sequences of this length.
[0049] The exact identity of sequences S1 and S2 in the amplification primers, and therefore the homologous sequences in the adapters, is generally not a subject of the present invention, as long as the primer molecules are able to interact with the amplification sequence to induce PCR amplification. The criteria for designing PCR primers are generally well known to those skilled in the art.
[0050] Solid-phase amplification by methods similar to either WO 98 / 44151 or WO 00 / 18957 will result in the generation of a clustered array of colonies of "bridged" amplification products. Both strands of the amplification product will be immobilized on a solid support at or near their 5' ends, and this attachment will result from the original attachment of the amplification primers. Typically, the amplification product within each colony will result from the amplification of a single template (target) molecule.
[0051] Modifications necessary to enable subsequent cleavage of the crosslinked amplification product can be advantageously included in one or both amplification primers. Such modifications can be located anywhere in the amplification primer, provided that they do not materially affect the efficiency of the amplification reaction. Thus, the cleavage-enabling modification can form part of the linker region L, or one or both of sequences S1 or S2. By way of example, amplification primers can be modified to include, inter alia, diol linkages, uracil nucleotides, ribonucleotides, methylated nucleotides, peptide linkers, PCR stoppers, or recognition sequences for restriction endonucleases. Because all nucleic acid molecules prepared by solid-phase amplification ultimately contain sequences derived from the amplification primers, any modifications in the primers will be carried over to the amplification products.
[0052] Alternative amplification methods (e.g., isothermal amplification or exclusion amplification as described in WO 2013 / 188582) can be used.
[0053] The present invention may also include a sequencing step, or aspects of the present invention may also encompass methods of sequencing a nucleic acid template generated using the methods of the present invention. Thus, the present invention provides a method of nucleic acid sequencing comprising providing a template for nucleic acid sequencing using the methods described herein, and performing a nucleic acid sequencing reaction to determine at least one region of the sequence of the template.
[0054] Sequencing can be performed using any suitable "sequencing by synthesis" technique, in which nucleotides are added sequentially to the free 3' hydroxyl group, resulting in the synthesis of a polynucleotide chain in the 5' to 3' direction. The identity of the added nucleotide is preferably determined after each addition.
[0055] The starting point for the sequencing reaction can be provided by annealing a sequencing primer to a single-stranded region of the template. Thus, the present invention encompasses a method in which a nucleic acid sequencing reaction comprises hybridizing a sequencing primer to a single-stranded fragment of a template nucleic acid immobilized in a nanowell as provided in the above-described embodiment of the present invention, sequentially incorporating one or more nucleotides into a polynucleotide strand complementary to the region of the template to be sequenced, and identifying the base present in one or more of the incorporated nucleotides, thereby determining the sequence of the region of the template.
[0056] A preferred sequencing method that can be used in the present invention relies on the use of modified nucleotides containing a 3'-blocking group that can act as a chain terminator. When a modified nucleotide is incorporated into a growing polynucleotide strand complementary to the region of the template being sequenced, the polymerase cannot add additional nucleotides because there is no free 3'-OH group available to guide further sequence extension. Once the nature of the base incorporated into the growing strand is determined, the 3'-block can be removed to allow the addition of the next successive nucleotide. By sequencing the products derived using these modified nucleotides, it is possible to deduce the DNA sequence of the DNA template. If each modified nucleotide is attached with a different label known to correspond to a specific base, to facilitate discrimination between the bases added at each incorporation step, such reactions can be performed in a single experiment. Alternatively, separate reactions can be performed, each containing a different modified nucleotide.
[0057] The modified nucleotides may carry a label to facilitate their detection. Preferably, this is a fluorescent label. Each nucleotide type may carry a different fluorescent label. However, the detectable label does not have to be a fluorescent label. Any label that allows the detection of the incorporated nucleotide may be used.
[0058] One method for detecting fluorescently labeled nucleotides involves using laser light of a wavelength specific to the labeled nucleotide, or other suitable illumination source. Fluorescence from the label on the nucleotide can be detected by a CCD camera or other suitable detection means.
[0059] The methods of the present invention are not limited to use with the sequencing methods outlined above, but can be used in conjunction with essentially any sequencing methodology that relies on the sequential incorporation of nucleotides into a polynucleotide chain. Suitable techniques include, for example, Pyrosequencing™, FISSEQ (fluorescent in situ sequencing), MPSS (massively parallel signature sequencing), and ligation-based sequencing.
[0060] The target polynucleotide to be sequenced using the method of the present invention can be any polynucleotide that is desired to be sequenced.The target polynucleotide can be of known, unknown, or partially known sequence, for example, in resequencing applications.Using the template preparation method detailed herein, it is possible to prepare a template starting from essentially any double-stranded target polynucleotide of known, unknown, or partially known sequence.By using an array, it is possible to sequence multiple targets of the same or different sequences in parallel.A particularly preferred application of this method is in sequencing fragments of genomic DNA. [Brief explanation of the drawings]
[0061] These and other aspects of the present invention will now be described with reference to the accompanying drawings. [Figure 1] 1 shows a schematic diagram of an exemplary flow cell that may be used with certain methods described herein. [Figure 2] 1 illustrates the concept of a multi-hybridization workflow. [Figure 3] Results obtained using 1 to 10 pushes (hybridization cycles) for various metrics are shown. [Figure 4] Comparing library plating efficiency using two different multihybridization workflows. [Figure 5] Results are shown for a four-push workflow using a library concentration of 200 pM.
[0062] Referring to Figure 1, a schematic diagram of an exemplary flow cell that can be used with certain methods described herein is shown. The flow cell is formed of three layers. The bottom layer 1 is formed of borosilicate glass with a depth of 1000 μm. An etched silicon channel layer (100 μm deep) is placed on top to define eight separate reaction channels. The top layer 3 (300 μm deep) contains two separate series of eight holes 4 and 4' that align with the channels in the etched silicon channel layer to provide fluid communication with the channel contents when the flow cell is assembled in use. The borosilicate glass of bottom layer 1 is etched with nanowells set in a patterned array with a 350 nm pitch and aligned with the reaction channels. The interstitial regions, i.e., between the channels, are not etched and do not contain nanowells. During use, primers for nucleic acid fragment template capture are bound to the nanowells.
[0063] The following protocol can be used to load a flow cell with a prepared library for sequencing. The initial library of double-stranded DNA containing appropriate adapters for the sequencing technology used can be prepared in any suitable manner known in the art. For example, a library preparation kit can be purchased from Illumina, Inc. (San Diego, USA) to prepare a suitable library. The examples described herein were prepared using the TruSeq Human Nano450 preparation kit.
[0064] The following workflow exemplifies a four-push multi-hybridization strategy. The total number of hybridization events can be varied up or down depending on the required use case and desired final effective concentration.
[0065] Library denaturation and dilution: 1. Dilute 20 ul of double-stranded DNA library to the appropriate concentration with HO or RSB buffer (available in the TruSeq Human Nano450 preparation kit). To utilize the multiple hybridization workflow, this working concentration is 4x lower than that required for standard single-event hybridization protocols. 2. Mix the library 1:1 with LDR denaturing reagent (100% formamide). 3. Heat to 65°C and incubate for 8 minutes to denature the double-stranded DNA template. 4. Add 160 ul of HT1 (5x SSC + 0.1% Tween 20) to dilute the denatured library to the final intended working concentration. 5. Proceed to template hybridization.
[0066] Template hybridization using the Illumina NovaSeq6000 system with a multihybridization workflow: 1. Prime / wet the cartridge lines and flow cell with BB6 buffer. 2. Heat the flow cell to 40°C. 3. Pump the denatured and diluted template from the upstream BB6 buffer into the flow cell in an initial flush aliquot large enough to completely cover the flow cell without dilution. 4. Incubate for 5 minutes. 5. After incubation, an additional flow cell volume of denatured and diluted template is drawn into the flow cell and incubated for 5 minutes. 6. Repeat step 5 two more times for a total of four hybridization events. 7. Proceed to cluster generation.
[0067] Figure 2 illustrates the concept of a multihybridization workflow. A typical single-push workflow is shown in the top line, where template hybridization is performed once, followed by cluster generation. If the template is loaded at a concentration of 800 pM, the effective concentration remains at 800 pM. The second and third lines set out alternative methods for achieving the same concentration, using two hybridization cycles at 400 pM (second line) or four hybridization cycles at 200 pM (third line). Multiple rounds of template hybridization / capture at lower DNA concentrations can be found to increase the effective DNA concentration and bring it into the required range. This, in turn, significantly lowers library prep concentration requirements, reduces template waste due to system dead volume, fluid lines, etc., and potentially allows for rerun queues for low-yield library preps.
[0068] Using the methodology described above, an initial 2 nM library was prepared and diluted to a concentration of 200 pM.
[0069] [Table 1]
[0070] Typical library yields for different applications are shown below:
[0071] [Table 2]
[0072] Figure 3 shows the results obtained using 1 to 10 pushes (hybridization cycles) for various metrics, with each workflow designed to provide an effective concentration of 800 pM. Measured metrics included cluster formation rate, occupied nanowell rate, remaining overlap rate, and usable yield. Regardless of the number of pushes, the metrics remained within a small band, and it can be seen that multiple pushes up to 10 provided metrics similar to a single push at higher concentrations. The optimum for usable yield and cluster formation was found to be within the range of 4 to 6 pushes. The flow cell used had a nanowell pitch of 350 nm.
[0073] Figure 4 compares the library seeding efficiency of two different multihybridization workflows (two-push vs. four-push). An initial, known amount of library was hybridized to the flow cell in one or more rounds, and then the hybridized template was eluted and quantified by qPCR. The initial input and uncaptured template fractions can be quantified in terms of the hybridized fraction. The graph shows that two-push and four-push provide similar numbers of ssDNA molecules per nanowell, which increases as the total DNA exposure increases up to 700 pM. The total number of pushes can be varied to obtain similar results, as long as the final DNA exposure is maintained. The two-push protocol utilizes twice the DNA concentration compared to the four-push protocol, but maintains similar final molecules per nanowell.
[0074] Figure 5 shows results from a four-push workflow using a 200 pM library concentration (800 pM effective concentration) on a 350 nm pitch flow cell using libraries prepared with the TruSeq Human Nano450 kit. It can be seen that as the total DNA concentration increases to 800 pM, the percentage of occupied nanowells increases and the percentage of duplicates decreases. It can be seen that the % pass filters remain consistent as the concentration increases to 800 pM, indicating that maintaining a lower concentration across multiple pushes results in optimal seeding.
[0075] Thus, the multiple hybridization workflow appears to be an effective protocol for efficiently loading high-density nanowell flow cells without excessively increasing the number of overlapped clusters, demonstrating that it is possible to use a sequencing flow cell as a DNA capture device to increase the effective concentration of the sample to be sequenced, especially when the sample preparation has a relatively low yield. The preferred 4x hybridization protocol reduces the library input concentration requirement by four-fold, potentially allowing for the re-running of low-yield library preparations.
Claims
1. 1. A method for preparing a template for a nucleic acid sequencing reaction, comprising: a) providing a solid support, the solid support comprising a plurality of nucleic acid primers immobilized thereon; b) contacting a nucleic acid library preparation with the solid support, the library preparation comprising a plurality of single-stranded fragments of a template nucleic acid to be sequenced, the template nucleic acid fragments further comprising one or more adapter nucleic acid sequences that hybridize to one or more of the nucleic acid primers, and the library preparation having a nucleic acid concentration of 400 pM or less; c) allowing the single-stranded fragments of the template nucleic acid to bind to the nucleic acid primer, thereby immobilizing the single-stranded fragments on the solid support; d) repeating steps b) and c) at least three more times; thereby providing a solid support having single-stranded fragments of template nucleic acid immobilized thereon; and not including a nucleic acid amplification step until steps b) and c) are repeated at least three times.
2. The method is a method for preparing a template for a nucleic acid sequencing reaction, a) providing a flow cell comprising a solid support having a plurality of nanowells formed thereon, each nanowell comprising a plurality of nucleic acid primers immobilized on the solid support, the nanowells being formed in a patterned array having a pitch of 500 nm or less; b) contacting a nucleic acid library preparation with the flow cell, the library preparation comprising a plurality of single-stranded fragments of a template nucleic acid to be sequenced, the template nucleic acid fragments further comprising one or more adapter nucleic acid sequences that hybridize to one or more of the nucleic acid primers, the library preparation having a nucleic acid concentration of 400 pM or less; c) allowing single-stranded fragments of template nucleic acid to bind to the nucleic acid primers, thereby immobilizing the single-stranded fragments within the nanowells; d) repeating steps b) and c) at least three more times; The method of claim 1 , thereby providing a flow cell having single-stranded fragments of template nucleic acid immobilized within nanowells.
3. 3. The method of claim 1 or 2, wherein the effective nucleic acid concentration, calculated as the number of cycles times the library nucleic acid concentration in step (c), is at least 800 pM.
4. The method of any one of claims 1 to 3, wherein the library preparation has a nucleic acid concentration of 250 pM or less.
5. The method of any one of claims 1 to 4, wherein the repeated contacting steps are performed with the same library preparation as used in step (b).
6. 6. The method of any one of claims 1 to 5, comprising denaturing a library preparation comprising double-stranded fragments of template nucleic acid to be sequenced to obtain the single-stranded fragments of template nucleic acid to be sequenced used in step (b).
7. 7. The method of any one of claims 1 to 6, further comprising amplifying the immobilized single-stranded fragment of nucleic acid, thereby generating multiple copies of the fragment.
8. The method of any one of claims 1 to 7, wherein the solid support is glass.
9. 10. The method of claim 1, wherein the solid support is on a flow cell, the solid support has a plurality of nanowells formed thereon, each nanowell having a plurality of nucleic acid primers immobilized on the solid support, and the nanowells are formed in a patterned array.
10. 3. The method of claim 2, wherein the pitch of the nanowell array is 350 nm or less.
11. The method of claim 1 , wherein the solid support is a microbead.
12. The method of any one of claims 1 to 11, further comprising sequencing the immobilized single-stranded fragments of nucleic acid.
Citation Information
Patent Citations
Apparatus and method for efficiently capturing nucleic acids
JP2014518639A
Multibase delivery for long reads in sequencing by synthesis protocols
US20150211062A1