Methods and Compositions for Cluster Generation by Bridge Amplification
By combining the steps of exonuclease and glycosylase into one step, using uracil DNA glycosylase and exonuclease to create a single nucleotide gap at the amplification site, solving the problems of long sequencing time and poor data quality in the NGS method, achieving a faster and more economical sequencing process.
Patent Information
- Application Number
- CN201980042543.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-05
- Filing Date
- 2019-11-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-11-29
AI Technical Summary
In the prior art, next-generation sequencing (NGS) methods require multiple separate steps in the amplification process of target nucleic acids, resulting in long sequencing time and poor signal-to-noise ratio and data quality, making it difficult to maintain efficient and high-quality data output simultaneously.
The steps of exonuclease and glycosylase are combined into one step by directly creating a single nucleotide gap at the amplification site by combining the composition with uracil DNA glycosylase and an exonuclease with 3' to 5' single-stranded DNA exonuclease activity, thereby achieving rapid amplification and sequencing of the target nucleic acid.
It shortens the sequencing time, reduces the amount of reagent used, reduces the cost, while maintaining the read quality and signal-to-noise ratio of data output, improving sequencing efficiency.
Smart Images

Figure CN112654718B_ABST
Abstract
Description
Field of the Invention
[0001] The present disclosure particularly relates to the amplification of target nucleic acids to generate amplicon clusters for sequencing, particularly in circumstances where steps involved in obtaining sequence information from a sample are reduced. Background of the Invention
[0003] Next-generation sequencing (NGS) technologies rely on the highly parallel sequencing of monoclonal populations of amplicons generated from a single target nucleic acid. NGS methods have greatly increased sequencing speed and data output, resulting in high sample throughput for current sequencing platforms. Further reducing the time required to sequence a template is highly desirable, but it is necessary to maintain a useful signal-to-noise ratio, intensity, and the percentage of clusters passing through the filter, all of which contribute to increased data output and data quality. Reducing the time required to template a sequencing can be achieved by combining separate steps; however, due to incompatibilities, separate steps often cannot be combined. For example, the activity of one enzyme can be inhibited by the product of another enzyme, so enzymes need to be used in separate steps. Summary of the Invention
[0005] Next-generation sequencing (NGS) technologies rely on the highly parallel sequencing of monoclonal populations of amplicons generated from a single target nucleic acid. Generating a monoclonal population of amplicons and performing sequencing require multiple steps, and each step increases the total time required before useful sequence data can be obtained from a sample. For example, generating a monoclonal population of amplicons requires multiple steps, including attaching and subsequently amplifying the target nucleic acid present at the amplification sites of an array. The inventors have discovered that two separate steps for generating monoclonal amplicons can be combined. In standard conventional methods, during the generation of monoclonal amplicons, the amplification sites are treated with an exonuclease and then, in a separate step, with a glycosylase that selectively creates a single nucleotide gap at a predetermined position. The steps are carried out in this order because creating a single nucleotide gap also creates a structure that inhibits exonuclease activity, namely 3'-phosphate. Unexpectedly, the inventors have found that the exonuclease and glycosylase can be combined into one step with little or no adverse effect on key metrics, read quality, dual indexing, or genome construction metrics. Since the two steps can now be carried out simultaneously, this results in faster sequencing. It also has the advantage of reducing the number and amount of reagents, thereby reducing the total cost to the consumer.
[0006] Compositions are provided herein. In one embodiment, the composition comprises uracil DNA glycosylase, an endonuclease, and an exonuclease having 3' to 5' single-stranded DNA exonuclease activity.
[0007] Methods are also provided. In one embodiment, a method is for preparing a nucleic acid for a sequencing reaction. The method includes providing an array having a plurality of amplification sites. The amplification sites include a plurality of capture nucleic acids attached to the amplification sites, wherein a first population of the plurality of capture nucleic acids includes cleavage sites. The amplification sites further include a plurality of cloned double-stranded modified target nucleic acids, wherein both strands of each double-stranded target nucleic acid are attached at their 5' ends to a capture nucleic acid, wherein one strand is attached to a capture nucleic acid that includes a cleavage site, and wherein the cleavage site is located in the double-stranded region of each double-stranded molecule. The method further includes contacting the array with a composition that includes at least one enzyme that generates an abasic site at the cleavage site and an exonuclease having 3' to 5' single-stranded DNA exonuclease activity, wherein cleavage occurs at the cleavage site, wherein the cleavage converts one strand of the double-stranded target nucleic acid into a first strand attached to the amplification site and a second strand not attached to the amplification site, and wherein the length of the single-stranded capture nucleic acid containing the free 3' end is shortened by the exonuclease.
[0008] Definitions
[0009] Unless otherwise specified, the terms used herein are to be understood as having their ordinary meaning in the relevant art. Several terms used herein and their meanings are listed below.
[0010] As used herein, the term "amplicon" when used in reference to a nucleic acid refers to the product of replication of a nucleic acid, wherein the product has a nucleotide sequence that is at least partially identical or complementary to the nucleotide sequence of the nucleic acid. An amplicon can be generated by any of a variety of amplification methods using a nucleic acid, such as a target nucleic acid or an amplicon thereof, as a template, the methods including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a specific nucleotide sequence (e.g., a polymerase extension product) or multiple copies of a nucleotide sequence (e.g., a concatameric product of RCA). The first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies created from the target nucleic acid or from the first amplicon after the first amplicon is generated. Subsequent amplicons can have a sequence that is substantially complementary or substantially identical to the target nucleic acid.
[0011] As used herein, the term "amplification site" refers to a site in or on an array at which one or more amplicons can be generated. An amplification site can be further configured to contain, hold, or attach at least one amplicon generated at the site.
[0012] As used herein, the term "array" refers to a population of sites that can be distinguished from one another based on their relative positions. Different molecules located at different sites of an array can be distinguished from one another based on the positions of the sites in the array. Individual sites of an array can include one or more molecules of a particular type. For example, a site can include a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Sites of an array can be different features located on the same substrate. Exemplary features include, but are not limited to, pores in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, or channels in a substrate. Sites of an array can be separate substrates each bearing a different molecule. Different molecules attached to separate substrates can be identified based on the substrate position on a surface associated with the substrate or based on the position of the substrate in a liquid or gel. Exemplary arrays of separate substrates located on a surface include, but are not limited to, those arrays having beads in pores.
[0013] As used herein, the term "capacity", when used in reference to a site and nucleic acid material, refers to the maximum amount of nucleic acid material, such as an amplicon derived from a target nucleic acid, that can occupy the site. For example, the term can refer to the total number of nucleic acid molecules that can occupy the site under specific conditions. Other measurements can also be used, including, for example, the total mass of the nucleic acid material or the total number of copies of a specific nucleotide sequence that can occupy the site under specific conditions. Generally, the capacity of a site for a target nucleic acid will be substantially equivalent to the capacity of a site for an amplicon of the target nucleic acid.
[0014] As used herein, the term "capture agent" refers to a material, chemical, molecule, or portion thereof that is capable of attaching to, retaining, or binding to a target molecule (such as a target nucleic acid). Exemplary capture agents include, but are not limited to, a capture nucleic acid (such as a universal capture binding sequence) that is complementary to at least a portion of a modified target nucleic acid, a member of a receptor-ligand binding pair that is capable of binding to a modified target nucleic acid (or a linking moiety attached thereto) (such as avidin, streptavidin, biotin, lectin, carbohydrate, nucleic acid binding protein, epitope, antibody, etc.), or a chemical reagent that is capable of forming a covalent bond with a modified target nucleic acid (or a linking moiety attached thereto). In one embodiment, the capture agent is a nucleic acid. The nucleic acid capture agent can also be used as an amplification primer.
[0015] When referring to nucleic acid capture agents, the terms "P5" and "P7" may be used. The terms "P5’" (P5 prime) and "P7’" (P7 prime) refer to the complements of P5 and P7, respectively. It should be understood that any suitable nucleic acid capture agent can be used in the methods presented herein, and the use of P5 and P7 is merely an exemplary embodiment. As disclosed in WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957, the use of nucleic acid capture agents such as P5 and P7 on flow cells is known in the art. Those skilled in the art will recognize that nucleic acid capture agents can also function as amplification primers. For example, any suitable nucleic acid capture agent can act as a forward amplification primer, whether immobilized or in solution, and can be used in the methods presented herein to hybridize with a sequence (e.g., a universal capture binding sequence) and amplify the sequence. Similarly, any suitable nucleic acid capture agent can act as a reverse amplification primer, whether immobilized or in solution, and can be used in the methods presented herein to hybridize with a sequence (e.g., a universal capture binding sequence) and amplify the sequence. Given the available common knowledge and the teachings of this disclosure, those skilled in the art will understand how to design and use sequences suitable for capturing and amplifying target nucleic acids as presented herein.
[0016] As used herein, the term "universal sequence" refers to a sequence region common to two or more target nucleic acids, where the molecules also have sequence regions that are different from each other. The universal sequences present in different members of a molecular collection can allow the capture of multiple different nucleic acids using a population of capture nucleic acids that are complementary to a portion of the universal sequence (e.g., a universal capture binding sequence). Non-limiting examples of universal capture binding sequences include sequences that are identical or complementary to the P5 and P7 primers. Similarly, the universal sequences present in different members of a molecular collection can allow the replication or amplification of multiple different nucleic acids using a population of universal primers that are complementary to a portion of the universal sequence, e.g., a universal primer binding site. As described herein, target nucleic acid molecules can be modified to attach universal adapters (also referred to herein as adapters) at one or both ends of, for example, different target sequences.
[0017] As used herein, the term "adapter" and its derivatives, such as universal adapters, generally refer to any linear oligonucleotide that can be ligated to a target nucleic acid. In some embodiments, the adapter is substantially non-complementary to the 3'-end or 5'-end of any target sequence present in the sample. In some embodiments, suitable adapter lengths are in the range of about 10 - 100 nucleotides, about 12 - 60 nucleotides, and about 15 - 50 nucleotides. Generally, an adapter can include any combination of nucleotides and / or nucleic acids. In some aspects, the adapter can include one or more cleavable groups at one or more positions. In another aspect, the adapter can include a sequence that is substantially identical or substantially complementary to at least a portion of a primer, such as a capture nucleic acid. In some embodiments, the adapter can include a barcode, also known as an index or tag, to assist with downstream calibration, identification, or sequencing. The terms "adapter head" and "adapter" can be used interchangeably.
[0018] As defined herein, "sample" and its derivatives are used in their broadest sense and include any sample, culture, etc. suspected of containing a target nucleic acid. In some embodiments, the sample contains DNA, RNA, PNA, LNA, chimeric or hybrid forms of nucleic acids. The sample can include any biological, clinical, surgical, agricultural, atmospheric, or aquatic sample containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, freshly frozen or formalin-fixed paraffin-embedded nucleic acid samples. Also encompassed is that the sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples from a single individual (matched), such as tumor samples and normal tissue samples, or from a single source containing two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of the nucleic acid material can include nucleic acids obtained from a neonate, such as those commonly used in neonatal screening.
[0019] As used herein, the terms "clonal population" and "monoclonal population" are used interchangeably and refer to a population of nucleic acids that is homogeneous with respect to a particular nucleotide sequence. The homologous sequences are generally at least 10 nucleotides in length, but can even be longer, including for example at least 50, at least 100, at least 250, at least 500, or at least 1000 nucleotides in length. A clonal population can be derived from a single target nucleic acid. Generally, all of the nucleic acids in a clonal population will have the same nucleotide sequence. It should be understood that a small number of mutations (e.g., due to amplification artifacts) can occur in a clonal population without loss of clonality. It should also be understood that a small number of different target nucleic acids (e.g., due to unamplified or minimally amplified target nucleic acids) can occur in a clonal population without loss of clonality.
[0020] As used herein, the term "different", when used in reference to nucleic acids, means that the nucleic acids have nucleotide sequences that are different from one another. Two or more nucleic acids can have nucleotide sequences that are different along their entire lengths. Alternatively, two or more nucleic acids can have nucleotide sequences that are different along a substantial portion of their lengths. For example, two or more nucleic acids can have portions of their target nucleotide sequences that are different from one another while also having common sequence regions that are the same as one another. As used herein, the term "different", when used in reference to amplification sites, means that the amplification sites are at distinct, separate positions on the same array.
[0021] As used herein, the term "fluidic access", when used in reference to molecules in a fluid and a site in contact with the fluid, means the ability of the molecules to move within the fluid or through the fluid to contact or enter the site. The term can also mean the ability of the molecules to separate from or leave the site to enter the solution. Fluidic access can occur when there are no obstacles that prevent the molecules from entering the site, contacting the site, separating from the site, and / or leaving the site. However, fluidic access is understood to exist as long as entry is not absolutely prevented, even if diffusion is delayed, reduced, or altered.
[0022] As used herein, the term "duplex", when used in reference to a nucleic acid molecule, means that substantially all of the nucleotides in the nucleic acid molecule are hydrogen-bonded to complementary nucleotides. A partially duplex nucleic acid can have at least 10%, at least 25%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of its nucleotides hydrogen-bonded to complementary nucleotides.
[0023] As used herein, the term "each", when used in reference to a collection of items, is intended to identify individual items in the collection but does not necessarily refer to every item in the collection unless the context clearly indicates otherwise.
[0024] As used herein, the term "excluded volume" refers to the volume of space occupied by a particular molecule that excludes other such molecules.
[0025] As used herein, the term "gap region" refers to a region in or on a substrate that is separated from other regions of the substrate or surface. For example, a gap region can separate one feature of an array from another feature of the array. Two regions that are separated from one another can be discrete and lack contact with one another. In another example, a gap region can separate a first portion of a feature from a second portion of the feature. The separation provided by the gap region can be partial or complete. A gap region will generally have a surface material that is different from the surface material of the features on the surface. For example, the amount or concentration of a capture agent that the features of an array can have can be greater than the amount or concentration present in the gap region. In some embodiments, a capture agent can be absent from the gap region.
[0026] As used herein, the term "polymerase" is intended to be consistent with its use in the art and includes, for example, enzymes that produce complementary repeats of nucleic acid molecules using a nucleic acid as a template strand. Generally, a DNA polymerase binds to the template strand and then moves down the template strand, thereby sequentially adding nucleotides to the free hydroxyl group at the 3' end of the growing strand of nucleic acid. DNA polymerases typically synthesize complementary DNA molecules from a DNA template, while RNA polymerases typically synthesize RNA molecules from a DNA template (transcription). Polymerases can use short RNA or DNA strands called primers to initiate strand growth. As described in detail herein, polymerases can be used during amplification to generate clusters of clones, polymerases can be used during sequencing reactions to determine the sequence of a nucleic acid, and different polymerases can be used in various aspects of these. Some polymerases can displace the strand upstream of the site at which they are adding bases to the strand. Such polymerases are considered strand-displacing, meaning that they have the activity of removing the complementary strand from the template strand read by the polymerase. Exemplary polymerases having strand-displacing activity include, but are not limited to, Bsu (Bacillus subtilis), Bst (Bacillus stearothermophilus) polymerase, exo-Klenow polymerase, or the large fragment of sequencing-grade T7 exo-polymerase. Some polymerases degrade the strand in front of them, effectively replacing it with the growing strand behind (5' exonuclease activity). Some polymerases have the activity of degrading the strand behind them (3' exonuclease activity). Some useful polymerases have been modified by mutation or other means to reduce or eliminate 3' and / or 5' exonuclease activity.
[0027] As used herein, the term "nucleic acid" is intended to be consistent with its use in the art and includes naturally occurring nucleic acids and functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to nucleic acids in a sequence-specific manner or are capable of serving as a template for replicating a particular nucleotide sequence. Naturally occurring nucleic acids typically have a backbone that includes phosphodiester bonds. Analog structures can have alternative backbone linkages, including any of a variety of backbone linkages known in the art. Naturally occurring nucleic acids typically have deoxyribose (e.g., found in deoxyribonucleic acid (DNA)) or ribose (e.g., found in ribonucleic acid (RNA)). Nucleic acids can contain any of a variety of analogs of these sugar moieties known in the art. Nucleic acids can include natural or unnatural bases. In this regard, naturally occurring deoxyribonucleic acid can have one or more bases selected from adenine, thymine, cytosine, or guanine, and ribonucleic acid can have one or more bases selected from uracil, adenine, cytosine, or guanine. Useful unnatural bases that can be included in nucleic acids are known in the art. When used to refer to a nucleic acid, the term "target" is intended in the context of the methods or compositions described herein as a semantic identifier of the nucleic acid and does not necessarily limit the nucleic acid structure or function beyond what is explicitly indicated otherwise. A target nucleic acid having a universal sequence at each end, e.g., a universal adapter at each end, can be referred to as a modified target nucleic acid.
[0028] As used herein, the term "transport" refers to the movement of a molecule through a fluid. The term can include passive transport, e.g., the movement of a molecule down its concentration gradient (e.g., passive diffusion). The term can also include active transport, whereby a molecule can move down or against its concentration gradient. Thus, transport can include the application of energy to move one or more molecules in a desired direction or to a desired location, e.g., an amplification site.
[0029] As used herein, the term "rate" when used to refer to transport, amplification, capture, or other chemical processes is intended to be consistent with its meaning in chemical kinetics and biochemical kinetics. The rates of two processes can be compared in terms of maximum rate (e.g., at saturation), pre-steady state rate (e.g., before equilibrium), kinetic rate constant, or other metrics known in the art. In a particular embodiment, the rate of a particular process can be determined in terms of the total time to complete the process. For example, the amplification rate can be determined in terms of the time taken to complete amplification. However, it is not necessary to determine the rate of a particular process in terms of the total time to complete the process.
[0030] The term "and / or" refers to one or all of the listed elements or to any combination of two or more of the listed elements.
[0031] The terms "preferred" and "preferably" refer to embodiments of the present invention that can provide certain benefits in certain circumstances. However, in the same or other circumstances, other embodiments may also be preferred. In addition, the recitation of one or more preferred embodiments does not mean that other embodiments are not available, and is not intended to exclude other embodiments from the scope of the present invention.
[0032] The term "comprising" and variations thereof do not have a limiting meaning when they appear in the specification and claims.
[0033] It should be understood that wherever an embodiment is described in terms of language such as "comprising", other embodiments similar in terms of being described by the terms "consisting of" and / or "consisting essentially of" are also provided.
[0034] Unless otherwise specified, "a", "an", "the", and "at least one" are used interchangeably and mean one or more than one.
[0035] A condition that is "suitable" for an event to occur, such as exonucleolytic nucleic acid digestion, or a "suitable" condition is a condition that does not prevent such an event from occurring. Thus, these conditions allow, enhance, facilitate, and / or favor the event.
[0036] As used herein, in the context of a composition, article, or nucleic acid, "providing" means preparing the composition, article, or nucleic acid, purchasing the composition, article, or nucleic acid, or otherwise obtaining the compound, composition, article, nucleic acid.
[0037] Also, in this document, the recitation of a numerical range by endpoints includes all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).
[0038] References to "one embodiment", "an embodiment", "certain embodiments", or "some embodiments", etc. throughout the specification mean that at least one embodiment of the present disclosure includes the particular feature, configuration, composition, or property described in connection with that embodiment. Thus, the appearances of these phrases throughout the specification are not necessarily referring to the same embodiment of the present disclosure. In addition, in one or more embodiments, the particular features, configurations, compositions, or properties may be combined in any suitable manner.
[0039] In the description herein, specific embodiments may be described separately for clarity. Unless otherwise explicitly stated that the features of a specific embodiment are incompatible with those of another embodiment, certain embodiments may include combinations of compatible features described herein in connection with one or more embodiments.
[0040] For any method disclosed herein that includes discrete steps, the steps may be performed in any feasible order. Also, any combination of two or more steps may be performed simultaneously, where appropriate.
[0041] The foregoing summary of the invention is not intended to describe each disclosed embodiment or every implementation of the invention. The following description more particularly exemplifies exemplary embodiments. Throughout the application, guidance is provided through lists of examples, which may be used in various combinations. In various instances, the recited lists merely serve as representative groups and should not be construed as exclusive lists. Brief Description of the Drawings
[0043] The following detailed description of exemplary embodiments of the disclosure can be best understood when read in conjunction with the following drawings.
[0044] Figure 1A-1D A schematic diagram showing an embodiment of preparing a nucleic acid for sequencing in accordance with various aspects of the disclosure presented herein.
[0045] Figure 2 Shows the effect of 3'-phosphate and DNA glycosylase on exonuclease I activity. The left figure shows the presence of DNA glycosylase and exonuclease in each lane of the flow cell. DNA glycosylase was added first, and then exonuclease. The exception is lane 4, where DNA glycosylase and exonuclease were added simultaneously. The middle figure shows the flow cell, with lanes 1-8 numbered from top to bottom, and the absence of signal in lanes 3-5. The right figure shows the fluorescence results for each lane, with essentially no fluorescence in lanes 3-5.
[0046] Figure 3 Shows the results of a sequencing run as described in Examples 1 and 3.
[0047] The schematic diagrams are not necessarily drawn to scale. The same numbers used in the drawings refer to the same components, steps, etc. However, it should be understood that the use of numbers to refer to components in a given drawing is not intended to limit the components labeled with the same numbers in another drawing. Additionally, the use of different numbers to refer to components is not intended to indicate that components with different numbers cannot be the same or similar to other numbered components. Detailed Description of the Invention
[0049] Methods and compositions related to sequencing nucleic acids are presented herein. The disclosure provides methods that include preparing a nucleic acid for a sequencing reaction, generating clonal clusters, and fabricating a nucleic acid array on a surface. In one embodiment, the method includes providing an array that includes a plurality of amplification sites. Each amplification site includes a plurality of double-stranded amplicons. For example, in Figure 1A is shown an amplification site 10 having one member of a plurality of double-stranded amplicons 11.
[0050] Multiple captured nucleic acids are attached to the surface of the amplification site. There are at least two populations of captured nucleic acids, and in some embodiments, there are three or more populations. At least one population of the captured nucleic acids includes a cleavage site. In one embodiment, the cleavage site includes uracil residues. Each double-stranded amplicon molecule is in a bridged structure, where they are attached to the captured nucleic acids at their 5' ends and not attached to the array at their 3' ends. The cleavage site is located in the double-stranded region of each double-stranded molecule. For example, as Figure 1A shown, two populations of captured nucleic acids are shown. One population is shown attached or bound to the surface of amplification site 10 at one end of each amplicon but not attached to amplicon 11. A second population of captured nucleic acids is also shown attached or bound to the surface of amplification site 10 at the other end of each amplicon but not attached to amplicon 11. Figure 1A The cleavage site (marked with an X on captured nucleic acid 13) is also shown in
[0051] The method further includes contacting the amplification site of the array with an enzyme that cleaves one DNA strand at the cleavage site and an exonuclease having 3' to 5' single-stranded DNA exonuclease activity. The exonuclease acts to digest the single-stranded captured nucleic acid containing a free 3'-OH end. For example, as Figure 1B shown, the cleavage site X in amplicon 11 is cleaved, leaving a shortened captured nucleic acid 13". The unattached captured nucleic acids 13' and 14' are no longer present at the amplification site 10.
[0052] In one embodiment, the sequence of the attached strand can be determined by using a DNA polymerase having strand displacement activity, where the 3' end of the shortened captured nucleic acid ( Figure 1B 13" in
[0053] is used as a primer to initiate DNA synthesis. In some embodiments, the enzyme that cleaves one DNA strand at the cleavage site will modify the 3' end of the shortened captured nucleic acid to terminate at 3'-phosphate. Before starting the sequencing reaction, the 3'-phosphate can be removed by a phosphatase. Figure 1B Instead of sequencing the attached strand, the method can further include subjecting the cleaved double-stranded amplicon to denaturing conditions to remove the portion of the cleaved strand that is not attached to the array ( Figure 1C 15' in
[0054] This results in immobilized single-stranded nucleic acids. For example, in Figure 1D shown, the unattached DNA strand no longer hybridizes with the attached strand 16 and has been lost.
[0055] array
[0056] The arrays of amplification sites used in the methods described herein can be present on one or more substrates. Exemplary types of substrate materials that can be used for the arrays include glass, modified glass, functionalized glass, inorganic glass, microspheres (e.g., inert and / or magnetic particles), plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, optical fibers or bundles of optical fibers, polymers, and porous (e.g., microtiter) plates. Exemplary plastics include acrylic, polystyrene, copolymers of styrene with other materials, polypropylene, polyethylene, polybutene, polyurethane, and Teflon TM . Exemplary silica-based materials include silicon and various forms of modified silicon.
[0057] In a specific embodiment, the substrate can be within or part of a container such as a well, tube, channel, dish, petri dish, bottle, etc. A particularly useful container is a flow cell, e.g., as described in U.S. Patent No. 8,241,573 or Bentley et al., Nature 456:53-59 (2008). An exemplary flow cell is a flow cell available from Illumina, Inc. (San Diego, Calif.). Another particularly useful container is a well in a porous plate or microtiter plate.
[0058] In some embodiments, the amplification sites of the array can be configured as features on a surface. These features can be present in any of a variety of desired forms. For example, the sites can be wells, pits, channels, ridges, raised areas, pins, posts, etc. In one embodiment, the amplification sites can contain beads. However, in specific embodiments, the sites need not contain beads or particles. Exemplary sites include the wells present in the substrates of commercial sequencing platforms sold by 454 Life Sciences (a subsidiary of Roche, Basel, Switzerland) or Ion Torrent (a subsidiary of Life Technologies, Carlsbad, CA, USA). Other substrates having wells include, for example, etched optical fibers and other substrates, as described in U.S. Patent No. 6,266,459; U.S. Patent No. 6,355,431; U.S. Patent No. 6,770,441; U.S. Patent No. 6,859,570; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; U.S. Patent No. 6,274,320; U.S. Patent No. 8,262,900; U.S. Patent No. 7,948,015; U.S. Patent Publication No. 2010 / 0137143; U.S. Patent No. 8,349,167, or PCT Publication No. WO 00 / 63437. In several cases, the substrates exemplified in these references illustrate the application of beads in wells. In the methods or compositions of the present disclosure, the well-containing substrates can be used with or without beads. In some embodiments, the wells of the substrate can contain a gel material (with or without beads), as described in U.S. Patent No. 9,512,422.
[0059] The amplification sites of the array can be metal features on a non-metal surface, such as glass, plastic, or other materials exemplified herein. Metal layers can be deposited on the surface using methods known in the art, such as wet plasma etching, dry plasma etching, atomic layer deposition, ion beam etching, chemical vapor deposition, vacuum sputtering, etc. Any of a variety of commercial instruments can be appropriately used, including, for example Ionfab or Optofab System (Oxford Instruments, UK). The metal layer can also be deposited by electron beam evaporation or sputtering, as described in Thornton, Ann. Rev. Mater. Sci. 7:239-60 (1977). Metal layer deposition techniques, such as those exemplified herein, can be combined with lithography techniques to create metal regions or patches on a surface. Exemplary methods for combining metal layer deposition techniques and lithography techniques are provided in U.S. Patent No. 8,778,848 and U.S. Patent No. 8,895,249.
[0060] The feature array can appear as a grid of spots or patches. The features can be positioned in a repeating pattern or in an irregular non-repeating pattern. Particularly useful patterns are hexagonal patterns, straight line patterns, grid patterns, patterns with reflection symmetry, patterns with rotational symmetry, etc. Asymmetric patterns can also be useful. The spacing between different pairs of nearest neighbor features can be the same, or there can be variation in the spacing between different pairs of nearest neighbor features. In a particular embodiment, the features of the array can each have an area greater than about 100 nm 2 、250 nm 2 、500 nm 2 、1 μm 2 、2.5 μm 2 、5 μm 2 、10 μm 2 、100 μm 2 or 500 μm 2 . Alternatively / additionally, the features of the array can each have an area less than about 1 mm 2 、500 μm 2 、100 μm 2 、25 μm 2 、10 μm 2 、5 μm 2 、1 μm 2 、500 nm 2 、or 100 nm 2 . In fact, the regions can have sizes within a range between the upper and lower limits selected from those exemplified above.
[0061] For embodiments that include a feature array on a surface, the features can be discrete, separated by interstitial regions. The size of the features and / or the spacing between regions can vary such that the array can be high density, medium density, or lower density. A high density array has features with regions having a spacing of less than about 15 μm. A medium density array has regions with a spacing of about 15 to 30 μm, while a low density array has regions with a spacing greater than 30 μm. Arrays useful in the present disclosure can have regions with a spacing less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm.
[0062] In a specific embodiment, the array can include a collection of beads or other particles. The particles can be suspended in a solution or they can be positioned on the surface of a substrate. Examples of bead arrays in solution are those commercialized by Luminex (Austin, TX, USA). Examples of arrays with beads on the surface include arrays in which the beads are located in wells, such as the BeadChip array (Illumina Inc., San Diego, CA, USA) or substrates for sequencing platforms from 454 Life Sciences (a subsidiary of Roche, Basel, Switzerland) or Ion Torrent (a subsidiary of Life Technologies, Carlsbad, CA, USA). Other arrays with beads on the surface are described in U.S. Patent No. 6,266,459; U.S. Patent No. 6,355,431; U.S. Patent No. 6,770,441; U.S. Patent No. 6,859,570; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; U.S. Patent No. 6,274,320; U.S. Patent Publication No. 2009 / 0026082A1; U.S. Patent Publication No. 2009 / 0127589A1; U.S. Patent Publication No. 2010 / 0137143A1; U.S. Patent Publication No. 2010 / 0282617A1; or PCT Publication No. WO 00 / 63437. Several of the above references describe methods for attaching a target nucleic acid to a bead prior to loading the bead into or onto an array substrate. However, it should be understood that the beads can be made to contain amplification primers and then the beads can be used to load the array, thereby forming amplification sites for the methods described herein. As previously described herein, a substrate can be used without beads. For example, amplification primers can be attached directly to a well or to a gel material in the well. Thus, the references illustrate materials, compositions, or devices that can be modified for use in the methods and compositions described herein.
[0063] The amplification sites of the array can include a variety of capture agents capable of binding to a target nucleic acid. In one embodiment, the capture agent includes a capture nucleic acid. The nucleotide sequence of the capture nucleic acid is complementary to the universal sequence of the target nucleic acid. In some embodiments, the capture nucleic acid can also serve as a primer for amplifying the target nucleic acid. In some embodiments, a population of capture nucleic acids includes a P5 primer or its complement. In some embodiments, the amplification site further includes a plurality of second capture nucleic acids, and this second capture nucleic acid can include a P7 primer or its complement. In some embodiments, the capture nucleic acid can include a cleavage site. Cleavage sites in capture nucleic acids are described in more detail herein.
[0064] In a specific embodiment, a capture agent such as a capture nucleic acid can be attached to an amplification site. For example, the capture agent can be attached to the surface of an array feature. The attachment can be through an intermediate structure such as a bead, particle, or gel. An example of attaching a capture nucleic acid to an array through a gel is described in U.S. Patent No. 8,895,249 and further exemplified by flow cells available from Illumina Inc. (San Diego, CA, USA) or described in WO 2008 / 093098. Exemplary gels that can be used in the methods and devices described herein include but are not limited to those having a colloidal structure such as agarose; a polymer network structure such as gelatin; or a cross-linked polymer structure such as polyacrylamide, SFA (e.g., see U.S. Patent Publication No. 2011 / 0059865A1) or PAZAM (e.g., see U.S. Provisional Patent Application Serial No. 61 / 753,833 and U.S. Patent No. 9,012,022). Attachment via beads can be achieved as exemplified in the descriptions and cited references previously described herein.
[0065] In some embodiments, the features on the surface of the array substrate are discontinuous and separated by gap regions of the surface. The gap regions having a substantially lower number or concentration of capture agents are advantageous compared to the features of the array. Gap regions lacking capture agents are particularly advantageous. For example, a relatively low amount or absence of capture moieties at the gap regions facilitates the localization of target nucleic acids and subsequently generated clusters to the desired features. In a specific embodiment, the features can be concave features in a surface (e.g., pores), and the features can contain a gel material. The gel-containing features can be separated from each other by gap regions on the surface where substantially no gel is present, or if present, the gel can substantially not support the localization of nucleic acids. Methods and compositions for preparing and using substrates having gel-containing features such as pores are set forth in U.S. Provisional Application No. 61 / 769,289.
[0066] target nucleic acid
[0067] The arrays used in the methods described herein include double-stranded modified target nucleic acids. The terms "target nucleic acid", "target fragment", "target nucleic acid fragment", "target molecule", and "target nucleic acid molecule" are used interchangeably and refer to a nucleic acid molecule whose nucleotide sequence is desired to be determined. The target nucleic acid can essentially be any nucleic acid of known or unknown sequence. For example, it can be a fragment of genomic DNA or cDNA. Sequencing can result in determination of the sequence of all or part of the target molecule. The target can be derived from a primary nucleic acid sample that has been randomly fragmented. In one embodiment, the target can be processed into a template suitable for amplification by placing universal amplification sequences, such as those present in universal adapters, at the ends of the respective target fragments. A target nucleic acid having universal adapters at each end can be referred to as a "modified target nucleic acid". Universal adapters are detailed herein.
[0068] The primary nucleic acid sample can be derived from the sample in the form of double-stranded DNA (dsDNA) (such as genomic DNA fragments, PCR and amplification products, etc.) or can be derived from the sample in single-stranded form (such as DNA or RNA) and converted to the dsDNA form. By way of example, mRNA molecules can be copied into double-stranded cDNA suitable for the methods described herein using standard techniques well known in the art. The exact sequence of the polynucleotide molecules from the primary nucleic acid sample is generally not important for the present disclosure and can be known or unknown.
[0069] In one embodiment, the primary polynucleotide molecule from the primary nucleic acid sample is a DNA molecule. More particularly, the primary polynucleotide molecule represents the entire genetic complement of an organism and is a genomic DNA molecule that includes both intron and exon sequences as well as non-coding regulatory sequences such as promoter and enhancer sequences. In one embodiment, a specific subset of the polynucleotide sequence or genomic DNA can be used, such as, for example, a specific chromosome. More particularly, the sequence of the primary polynucleotide molecule is unknown. Still more particularly, the primary polynucleotide molecule is a human genomic DNA molecule. The DNA target fragments can be chemically or enzymatically treated before or after any random fragmentation process and before or after the ligation of the universal adapter sequences.
[0070] Nucleic acid samples can include high molecular weight materials such as genomic DNA (gDNA). Samples can include low molecular weight materials such as nucleic acid molecules obtained from formalin-fixed paraffin-embedded or archived DNA samples. In another embodiment, the low molecular weight materials include enzymatically or mechanically fragmented DNA. Samples can contain cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinically or laboratory-obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic, or etiological sample. In some embodiments, the sample can include nucleic acid molecules obtained from animals such as human or mammalian sources. In another embodiment, the sample can include nucleic acid molecules obtained from non-mammalian sources such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecules can be an archived or depleted sample or species.
[0071] In addition, the methods and compositions disclosed herein can be used to amplify nucleic acid samples having low quality nucleic acid molecules such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, a forensic sample can include nucleic acids obtained from a forensic sample obtained from a crime scene, a missing persons DNA database, a laboratory associated with a forensic investigation, or from a law enforcement agency, one or more military branches, or any such person. The nucleic acid sample can be a purified sample or a lysate containing crude DNA, such as derived from a buccal swab, paper, fabric, or other substrate that can be impregnated with saliva, blood, or other body fluid. Thus, in some embodiments, the nucleic acid sample can include a small amount of DNA (e.g., genomic DNA) or a fragmented portion of the DNA. In some embodiments, the target sequence can be present in one or more body fluids including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence can be obtained from a victim's hair, skin, tissue sample, autopsy, or remains. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or person. In some embodiments, the target sequence can include nucleic acids obtained from non-human DNA, such as microbial, plant, or entomological DNA. In some embodiments, the target sequence or amplified target sequence is for the purpose of human identification. In some embodiments, the methods described herein can be used to characterize forensic samples. In some embodiments, the methods described herein can be used for human identification methods using one or more target-specific primers or one or more target-specific primers designed using known primer design criteria. In one embodiment, any one or more target-specific primers can be used to amplify a forensic or human identification sample containing at least one target sequence, the target-specific primers being obtained using known primer criteria.
[0072] Other non-limiting examples of biological sample sources can include whole organisms as well as samples obtained from patients. Biological samples can be obtained from any biological fluid or tissue and can be in a variety of forms, including liquid fluids and tissues, solid tissues, and preserved forms such as dried, frozen, and fixed forms. The sample can be of any biological tissue, cell, or body fluid. Such samples include, but are not limited to, sputum, blood, serum, plasma, blood cells (e.g., white blood cells), ascites, urine, saliva, tears, sputum, vaginal fluid (discharge), wash fluids obtained during medical procedures (e.g., pelvic or other lavages obtained during biopsy, endoscopy, or surgery), tissues, nipple aspirates, core or fine needle biopsy samples, cell-containing body fluids, free floating nucleic acids, peritoneal fluid, and pleural fluid, or cells from them. Biological samples can also include tissue sections, such as frozen sections or fixed sections taken for histological purposes, or microdissected cells or their extracellular portions. In some embodiments, the sample can be a blood sample, such as, for example, a whole blood sample. In another example, the sample is an untreated dried blood spot sample. In another example, the sample is a formalin-fixed paraffin-embedded sample. In yet another example, the sample is a saliva sample. In yet another example, the sample is a dried saliva spot sample.
[0073] Exemplary biological samples from which target nucleic acids can be derived include, for example, samples from eukaryotes such as mammals, e.g., rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cows, cats, dogs, primates, humans or non-human primates; plants such as Arabidopsis thaliana, maize, sorghum, oats, wheat, rice, canola or soybeans; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, bees or spiders; fish such as zebrafish; reptiles; amphibians such as frogs or Xenopus laevis; Dictyostelium discoideum; fungi such as Pneumocystis carinii, Takifugu rubripes, yeasts such as Saccharamoyces cerevisiae or Schizosaccharomyces pombe; or Plasmodium falciparum. Target nucleic acids can also be derived from prokaryotes such as bacteria, Escherichia coli, Staphylococcus or Mycoplasma pneumoniae; archaea; viruses such as human hepatitis C virus or human immunodeficiency virus; or viroids. Target nucleic acids can be derived from a homogeneous culture or population of an organism or, alternatively, from a collection of several different organisms in, for example, a community or ecosystem.
[0074] Random fragmentation refers to fragmenting polynucleotide molecules from a primary nucleic acid sample in a disordered manner by enzymatic, chemical, or mechanical means. Magnetic fragmentation methods are known in the art and use standard methods (Sambrook and Russell, Molecular Cloning, A Laboratory Manual, 3rd Edition). In one embodiment, fragmentation can be achieved using a process commonly referred to as tagmentation. Tagmentation uses a transpososome complex and combines fragmentation and ligation into one step to add universal adapters (Gunderson et al., WO 2016 / 130704). For clarity, the nucleic acid sequence of the larger fragment remains intact (i.e., the intact nucleic acid sequence is retained), and generating such smaller fragments of the larger nucleic acid segment by specific PCR amplification of the smaller fragments is not equivalent to fragmenting the larger nucleic acid segment because the larger nucleic acid sequence remains intact (i.e., is not fragmented by the PCR amplification). Additionally, random fragmentation is designed to produce fragments regardless of the sequence identity or position of the nucleotides encompassing and / or surrounding the break. More particularly, random fragmentation is by mechanical means, such as nebulization or sonication, to produce fragments ranging in length from about 50 base pairs to about 1500 base pairs, even more particularly from 50 - 700 base pairs, and even more particularly from 50 - 400 base pairs. Most particularly, the method is used to produce smaller fragments ranging in length from 50 - 150 base pairs.
[0075] Fragmenting polynucleotide molecules by mechanical means (such as nebulization, sonication, and Hydroshear) produces a heterogeneous mixture of fragments with blunt ends and 3'- and 5'-overhangs. Thus, it is desirable to use methods or kits known in the art (such as the Lucigen DNA Terminator End Repair Kit) to repair the fragment ends to produce ends that are optimal for insertion into, for example, the blunt site of a cloning vector. In one specific embodiment, the fragment ends of the nucleic acid population are blunt ends. More particularly, the fragment ends are blunt and phosphorylated. The phosphate moiety can be introduced by enzymatic treatment, such as using polynucleotide kinase.
[0076] The target nucleic acid population can have an average chain length that is desirable or suitable for a particular application of the methods or compositions described herein. For example, the average chain length can be less than about 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively / Additionally, the average chain length can be greater than about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average chain length of the target nucleic acid population can be within the range between the maximum and minimum values described herein. It should be understood that the amplicons generated (or otherwise prepared or used herein) at the amplification site can have an average chain length within the range between the upper and lower limits selected from the examples above.
[0077] In some cases, the target nucleic acid population can be generated under certain conditions or otherwise configured to have a maximum length of its members. For example, the maximum length of the members used in one or more steps of the methods described herein or present in a particular composition can be less than 100,000 nucleotides, less than 50,000 nucleotides, less than 10,000 nucleotides, less than 5,000 nucleotides, less than 1,000 nucleotides, less than 500 nucleotides, less than 100 nucleotides, or less than 50 nucleotides. Alternatively / Additionally, the target nucleic acid population can be generated under certain conditions or otherwise configured to have a minimum length of its target members. For example, the minimum length of the members used in one or more steps of the methods described herein or present in a particular composition can be greater than 10 nucleotides, greater than 50 nucleotides, greater than 100 nucleotides, greater than 100 nucleotides, greater than 500 nucleotides, greater than 1000 nucleotides, greater than 5,000 nucleotides, greater than 10,000 nucleotides, greater than 50,000 nucleotides, or greater than 100,000 nucleotides. The maximum and minimum chain lengths of the target nucleic acids in the population can be within the range between the maximum and minimum values set forth above. It should be understood that the amplicons generated (or otherwise prepared or used herein) at the amplification site can have a maximum and / or minimum chain length within the range between the upper and lower limits of the examples above.
[0078] In a specific embodiment, a target fragment sequence having a single protruding nucleotide is prepared by the activity of, for example, certain types of DNA polymerases such as Taq polymerase or Klenow exo minus polymerase, which have non-template-dependent terminal transferase activity that adds a single deoxynucleotide, such as deoxyadenosine (A), to the 3'-end of a DNA molecule (such as a PCR product). Such enzymes can be used to add a single nucleotide "A" to the blunt 3'-ends of each strand of a double-stranded target fragment. Thus, the 3'-ends of the strands can be repaired by adding "A" to each end of the double-stranded target fragment by reaction with Taq or Klenow exo minus polymerase, and the universal adaptor polynucleotide construct can be a T construct having compatible "T" overhangs present at the 3'-ends of the respective regions of the double-stranded nucleic acid of the universal adaptor. This terminal modification also prevents self-ligation of both the vector and the target, resulting in a preference for the formation of target nucleic acids having universal adaptors at each end.
[0079] In some cases, the target nucleic acid derived from such sources can be amplified prior to use in the methods or compositions herein. Any of a variety of known amplification techniques can be used, including but not limited to polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), or random prime amplification (RPA). It should be understood that amplification of the target nucleic acid is optional prior to use in the methods or compositions described herein. Thus, the target nucleic acid may not be amplified prior to use in some embodiments of the methods and compositions described herein. The target nucleic acid can optionally be derived from a synthetic library. The synthetic nucleic acid can have a natural DNA or RNA composition or can be an analogue thereof.
[0080] universal adaptor
[0081] The target nucleic acids used in the methods or compositions described herein include universal adaptors attached to each end. The target nucleic acid having universal adaptors at each end can be referred to as a "modified target nucleic acid". The methods for attaching universal adaptors to each end of the target nucleic acid used in the methods described herein are known to those skilled in the art. Ligation can be carried out by using standard library preparation techniques for ligation (U.S. Patent Publication No. 2018 / 0305753), or by tag fragmentation using a transposase complex (Gunderson et al., WO 2016 / 130704).
[0082] In one embodiment, double-stranded target nucleic acids from a sample, such as a fragmented sample, are processed by first ligating the same universal adaptor molecules ("mismatched adaptors", the general characteristics of which are defined below and further described in Gormley et al., U.S. Patent No. 7,741,463 and Bignell et al., U.S. Patent No. 8,053,192) to the 5' and 3' ends of the double-stranded target nucleic acid. In one embodiment, the universal adaptor includes universal capture binding sequences necessary to immobilize the target nucleic acid on an array for subsequent sequencing. In another embodiment, prior to immobilization and sequencing, the universal adaptors present at each end of the target nucleic acid are further modified using a PCR step. For example, an initial primer extension reaction is performed using universal primer binding sites, where extension products complementary to both strands of each individual target nucleic acid are formed and universal capture binding sequences are added. The resulting primer extension products and optionally amplified copies thereof together provide a library of modified target nucleic acids that can be immobilized and then sequenced. The term "library" refers to a collection of target nucleic acids that contain known common sequences at their 3' and 5' ends and may also be referred to as a 3' and 5' modified library.
[0083] The universal adaptors used in the methods of the present disclosure are referred to as "mismatched" adaptors because, as explained in detail herein, the adaptors contain regions of sequence mismatch, i.e., they are not formed by annealing fully complementary polynucleotide strands.
[0084] The mismatched adaptors for use herein are formed by annealing two partially complementary polynucleotide strands to provide at least one double-stranded region (also referred to as a double-stranded nucleic acid region) and at least one mismatched single-stranded region, also referred to as a region of single-stranded non-complementary nucleic acid strands, when the two strands are annealed.
[0085] The double-stranded region of the universal adaptor is a short double-stranded region, typically containing 5 or more consecutive base pairs, which is formed by annealing two partially complementary polynucleotide strands. The term refers to the nucleic acid double-stranded region where the two strands are annealed and does not imply any particular structural conformation.
[0086] It is generally advantageous for the double-stranded region to be as short as possible without loss of function. In the present context, "function" refers to the ability of the double-stranded region to form a stable duplex under the standard reaction conditions of an enzyme-catalyzed nucleic acid ligation reaction, which are well known to the skilled reader (e.g., incubation at a temperature in the range of 4°C to 25°C in a ligation buffer suitable for the enzyme), such that the two strands of the universal adaptor remain partially annealed during the ligation of the universal adaptor to the target molecule. It is not absolutely necessary for the double-stranded region to be stable under the conditions typically used in the annealing step of a primer extension or PCR reaction.
[0087] The double-stranded region of the universal adaptor is typically the same among all universal adaptors commonly used in ligation. Since the universal adaptor ligates to both ends of each target molecule, the modified target nucleic acid flanks can have complementary sequences derived from the double-stranded region of the universal adaptor. The longer the double-stranded region and the complementary sequences derived therefrom in the modified target nucleic acid construct, the greater the likelihood that the modified target nucleic acid construct will fold back on itself and base pair with itself in these regions of internal self-complementarity under the annealing conditions usable in primer extension and / or PCR. Thus, it is generally preferred that the double-stranded region be 20 or fewer, 15 or fewer, or 10 or fewer base pairs in length to reduce this effect. The stability of the double-stranded region can be increased, and thus potentially its length shortened, by including unnatural nucleotides that exhibit stronger base pairing than standard Watson-Crick base pairs.
[0088] In one embodiment, the two strands of the universal adaptor are 100% complementary in the double-stranded region. It should be understood that one or more nucleotide mismatches may be tolerated in the double-stranded region provided that the two strands can form a stable duplex under standard ligation conditions.
[0089] The universal adaptors used herein generally may include a double-stranded region that forms the "ligatable" end of the adaptor, e.g., the end that ligates to a double-stranded target nucleic acid in a ligation reaction. The ligatable end of the universal adaptor can be blunt-ended, or in other embodiments, there can be a short 5' or 3' overhang of one or more nucleotides to facilitate / promote ligation. The 5'-terminal nucleotide of the ligatable end of the universal adaptor is typically phosphorylated to effect phosphodiester ligation to the 3'-hydroxyl on the target polynucleotide.
[0090] The term "mismatch region" refers to a region of the universal adaptor, a region of single-stranded non-complementary nucleic acid strands, where the sequences of the two polynucleotide strands that form the universal adaptor exhibit a degree of non-complementarity such that under the standard annealing conditions of a primer extension or PCR reaction, the two strands cannot fully anneal to each other. Under the standard reaction conditions of an enzyme-catalyzed ligation reaction, the mismatch region can exhibit a degree of annealing provided that the two strands are converted to single-stranded form under the annealing conditions of the amplification reaction.
[0091] It should be understood that the "mismatch region" is provided by different portions of the same two polynucleotide strands that form a double-stranded region. The mismatch in the adaptor construct can take the form of one strand being longer than the other, such that there is a single-stranded region on one of the strands, or it can take the form of a sequence selected such that the two strands do not hybridize, thereby forming single-stranded regions on both strands. The mismatch can also take the form of a "bubble", where the two ends of the universal adaptor construct are able to hybridize to each other and form a duplex, while the central region cannot. Under conditions where other portions of the same two strands are annealed to form one or more double-stranded regions, the strand portions that form the mismatch region do not anneal. For the avoidance of doubt, it should be understood that in the context of the present disclosure, a single-stranded or single-base overhang at the 3' end of a polynucleotide duplex that subsequently undergoes ligation to a target sequence does not constitute a "mismatch region".
[0092] The lower limit of the length of the mismatch region is generally determined by function, such as the need to provide a sequence suitable for: i) primer binding for primer extension, PCR, and / or sequencing (e.g., binding of a primer to a universal primer binding site), or ii) binding of a universal capture binding sequence to a capture nucleic acid to immobilize the modified target nucleic acid on a surface. In theory, there is no upper limit to the length of the mismatch region, although it is generally advantageous to minimize the total length of the universal adaptor, e.g., in order to facilitate separation of unbound universal adaptor from the modified target nucleic acid construct after the ligation step. Thus, it is generally preferred that the length of the mismatch region be less than 50, or less than 40, or less than 30, or less than 25 consecutive nucleotides.
[0093] The region of the single-stranded non-complementary nucleic acid strand contains at least one universal capture binding sequence at the 3' end. The 3' end of the universal adaptor contains a universal capture binding sequence that will hybridize to a capture nucleic acid present at an array amplification site. Optionally, the 5' end of the universal adaptor contains a second universal capture binding sequence that attaches to each end of the target nucleic acid, where the second universal capture binding sequence will hybridize to a different capture nucleic acid present at the amplification site of the array.
[0094] The region of the single-stranded non-complementary nucleic acid strand generally also includes at least one universal primer binding site. A universal primer binding site is a universal sequence that can be used to amplify and / or sequence a target nucleic acid ligated to a universal linker.
[0095] Regions of single-stranded non-complementary nucleic acid strands can also include at least one index. The index can be used as a marker characteristic of the source of a particular target nucleic acid on an array (U.S. Patent No. 8,053,192). Typically, the index is a synthetic sequence of nucleotides that is part of a universal adaptor that is added to the target nucleic acid as part of a library preparation step. Thus, the index is a nucleic acid sequence attached to individual target molecules of a particular sample, the presence of which indicates or is used to identify the sample or source from which the target molecules were isolated. In one embodiment, a dual-index system can be used. In a dual-index system, the universal adaptor attached to the target nucleic acid includes two different index sequences (U.S. Patent Publication Nos. 2018 / 0305750, 2018 / 0305751, 2018 / 0305752, and 2018 / 0305753).
[0096] Preferably, the length of the index can be up to 20 nucleotides, more preferably 1-10 nucleotides, and most preferably 4-6 nucleotides in length. A 4-nucleotide index gives the possibility of multiplexing 256 samples on the same array, and a 6-base index enables processing 4096 samples on the same array.
[0097] In one embodiment, when the universal capture binding sequence is ligated to the double-stranded target fragment, it is part of the universal adaptor, and in another embodiment, after the universal adaptor is ligated to the double-stranded target fragment, a universal primer extension binding site is added to the universal adaptor. The addition can be accomplished using conventional methods, including amplification-based methods such as PCR.
[0098] The exact nucleotide sequence of the universal adaptor is generally not critical to the present disclosure and can be chosen by the user such that desired sequence elements are ultimately included in the common sequence of a plurality of different modified target nucleic acids, e.g., to provide a universal capture binding sequence and binding sites for universal amplification primers and / or sequencing primers for a particular set. Other sequence elements can be included, e.g., to provide a binding site for a sequencing primer, e.g., on a solid support, that will ultimately be used for sequencing of target nucleic acids in the library, index sequencing, or products derived from amplification of target nucleic acids in the library.
[0099] Although the exact nucleotide sequence of the universal adaptor is generally not limited to the present disclosure, the sequence of the individual strands in the mismatch region should be such that no individual strand exhibits any internal self-complementarity that could lead to self-annealing, formation of hairpin structures, etc. under standard annealing conditions. Self-annealing of the strands in the mismatch region should be avoided because it can prevent or reduce specific binding of the amplification primer to this strand.
[0100] The mismatched adaptor is preferably formed from two DNA strands, but may include a mixture of natural and unnatural nucleotides (e.g., one or more ribonucleotides) linked by a mixture of phosphodiester and non-phosphodiester backbone linkages.
[0101] ligation and amplification of universal adaptor
[0102] Ligation methods are known in the art and use standard methods. Such methods use ligases, e.g., DNA ligases, to effect or catalyze the ligation of the ends of two polynucleotide strands of a universal adaptor and a double-stranded target nucleic acid in this case, thereby forming a covalent linkage. The universal linker may contain a 5'-phosphate moiety to facilitate ligation to the 3'-OH present on the target fragment. The double-stranded target nucleic acid contains a 5'-phosphate moiety, remaining from a self-cleavage process or added using an enzymatic treatment step, and has been end-repaired and optionally extended by one or more overhanging bases to give a 3'-OH suitable for ligation. In this context, ligation refers to the covalent linkage of polynucleotide strands that were not previously covalently linked. In certain aspects of the present disclosure, such ligation occurs by forming a phosphodiester linkage between two polynucleotide strands, but other modes of covalent linkage (e.g., non-phosphodiester backbone linkages) may be used.
[0103] As discussed herein, in one embodiment, the universal adaptor for ligation is complete and includes a universal capture binding sequence and other universal sequences, such as a universal primer binding site and an index sequence. The resulting plurality of modified target nucleic acids can be used to prepare immobilized samples for sequencing.
[0104] Also as discussed herein, in one embodiment, the universal adaptor for ligation contains a universal primer binding site and an index sequence and does not contain a universal capture binding sequence. The resulting plurality of modified target nucleic acids can be further modified to include a specific sequence, such as a universal capture binding sequence. Methods for adding a specific sequence, such as a universal capture binding sequence, to a universal primer that is ligated to a double-stranded target fragment include amplification-based methods, e.g., PCR, and are known in the art and are described, for example, in: Bignell et al. (US 8,053,192) and Gunderson et al. (WO2016 / 130704).
[0105] In an embodiment where the universal linker is modified, an amplification reaction is prepared. The contents of the amplification reaction are known to those skilled in the art and include suitable substrates (such as dNTPs), enzymes (such as DNA polymerases), and buffer components required for the amplification reaction. Generally, the amplification reaction requires at least two amplification primers, commonly referred to as "forward" and "reverse" primers (primer oligonucleotides), which are capable of specifically annealing to a portion of the polynucleotide sequence to be amplified (such as the modified target nucleic acid) under the conditions encountered in the primer annealing step of each cycle of the amplification reaction. It should be understood that if a primer contains any nucleotide sequence that does not anneal to the modified target nucleic acid in the first amplification cycle, this sequence may be copied into the amplification product. For example, using a primer with a universal capture binding sequence, such as a sequence that does not anneal to the modified target nucleic acid, the universal capture binding sequence will be incorporated into the resulting amplicon.
[0106] Amplification primers are generally single-stranded polynucleotide structures. They can also contain a mixture of natural and non-natural bases and natural and non-natural backbone linkages, provided that any non-natural modifications do not preclude the function of the primer - defined as the ability to anneal to the template polynucleotide chain during the conditions of the amplification reaction and to serve as a starting point for synthesizing a new polynucleotide chain complementary to the template strand. Additionally, the primer can contain non-nucleotide chemical modifications, such as phosphorothioates to increase exonuclease resistance, again provided that the modification does not prevent primer function.
[0107] amplification to generate clusters
[0108] Methods known to those skilled in the art can be used to generate an array comprising amplification sites, each amplification site containing a clonal population (also referred to as a cluster) of double-stranded amplicons. In one embodiment, an isothermal amplification method is used and includes generating a clonal population of double-stranded amplicons from individual target nucleic acids that have seeded the site. In some embodiments, the amplification reaction is carried out until a sufficient number of amplicons are produced to fill the capacity of the corresponding amplification site. Filling the seeded site to capacity in this manner precludes subsequent target nucleic acids from landing at that site, thereby generating a clonal population of amplicons at that site. Thus, in some embodiments, it is desirable to produce amplicons at a rate that fills the amplification site capacity faster than the rate at which individual target nucleic acids are transported to individual amplification sites.
[0109] In some embodiments, the amplification methods include, but are not limited to, solid-phase amplification, polony amplification, colony amplification, emulsion PCR, bead RCA, surface RCA, or surface SDA. In some embodiments, an amplification method is used that results in the amplification of free DNA molecules in solution or that are tethered to a suitable substrate only by one end of the DNA molecule. In some embodiments, a method that relies on bridge PCR is used, in which both PCR primers are attached to a surface (see, e.g., WO 2000 / 018957, U.S. Patent No. 7,972,820; U.S. Patent No. 7,790,418, and Adessi et al., Nucleic Acids Research (2000): 28(20): E87). In some embodiments, the methods of the invention can create "polymerase colony technology" or "polony," which refers to multiplex amplification that maintains spatial clustering of identical amplicons (see Harvard Molecular Technology Group and Lipper Center for Computational Genetics website). These include, for example, in situ polony (Mitra and Church, Nucleic Acid Research 27, e34, Dec. 15, 1999), in situ rolling circle amplification (RCA) (Lizardi et al., Nature Genetics 19, 225, July 1998), bridge PCR (U.S. Patent No. 5,641,658), picoliter PCR (Leamon et al., Electrophoresis 24, 3769, November 2003), and emulsion PCR (Dressman et al., PNAS 100, 8817, July 22, 2003). In some embodiments, a method that relies on kinetic exclusion is used, in which recombinase-promoted amplification and isothermal conditions amplify a library (U.S. Patent No. 9,309,502, U.S. Patent No. 8,895,249, U.S. Patent No. 8,071,308).
[0110] In some embodiments, significant clonality can be achieved even if the amplification site is not filled to capacity before the second target nucleic acid begins amplification at that site. Under certain conditions, amplification of the first target nucleic acid can proceed to a point where a sufficient number of copies are generated to effectively outcompete or overwhelm the copy generation of the second target nucleic acid that self-transports to that site. For example, in embodiments using a bridge amplification process on circular features less than 500 nm in diameter, it has been determined that after 14 cycles of exponential amplification of the first target nucleic acid, contamination from the second target nucleic acid at the same site will produce an insufficient number of contaminating amplicons to adversely affect the synthesis sequencing analysis on an Illumina sequencing platform.
[0111] In all embodiments, the amplification sites in the array need not be completely clonal. Instead, for certain applications, individual amplification sites can be predominantly filled with amplicons from the first target nucleic acid and can also have low levels of contaminating amplicons from the second target nucleic acid. The array can have one or more amplification sites with low levels of contaminating amplicons, provided that the contamination level does not have an unacceptable effect on the subsequent use of the array. For example, when the array is to be used for a detection application, an acceptable contamination level is one that does not affect the signal-to-noise ratio or the resolution of the detection technique in an unacceptable manner. Thus, apparent clonality is generally related to the specific use or application of the array made by the methods described herein. Exemplary levels of acceptable contamination at individual amplification sites for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% contaminating amplicons. The array can include one or more amplification sites with these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites in the array can have some contaminating amplicons.
[0112] In some embodiments, the method of preparing an array useful in the methods described herein can be carried out under conditions where the target nucleic acid is transported (e.g., by diffusion) to the amplification site as amplification occurs. Thus, some amplification methods can utilize both a relatively slow transport rate relative to subsequent amplicon formation and a relatively slow production of the first amplicon. For example, the amplification reaction described herein can be carried out such that the target nucleic acid is transferred from solution to the amplification site simultaneously with (i) the production of the first amplicon, and (ii) the production of subsequent amplicons at other sites in the array. In a specific embodiment, the average rate of production of subsequent amplicons at the amplification site can exceed the average rate of transport of the target nucleic acid from solution to the amplification site. In certain cases, a sufficient number of amplicons can be produced from a single target nucleic acid at an individual amplification site to fill the capacity of the corresponding amplification site. The rate of producing amplicons to fill the capacity of the corresponding amplification site can, for example, exceed the rate of transporting the individual target nucleic acid from solution to the amplification site.
[0113] Compositions for amplifying a target nucleic acid at an amplification site, referred to herein as "amplification reagents", are generally capable of rapidly preparing copies of the target nucleic acid at the amplification site. The amplification reagents used in the methods of the present disclosure will generally include a polymerase and nucleotide triphosphates (NTPs). Any of a variety of polymerases known in the art can be used, but in some embodiments, an exonuclease-negative polymerase can be preferably used. Examples of nucleic acid polymerases suitable for the embodiments of the present invention include, but are not limited to, DNA polymerases (such as Klenow fragment, T4 DNA polymerase, Bst (Bacillus stearothermophilus) polymerase), thermostable DNA polymerases (such as Taq, Vent, Deep Vent, Pfu, Tfl, and 9°N DNA polymerase), and their genetically modified derivatives (TaqGold, VENTexo, Pfu exo). In some embodiments, the amplification reagent can also include a recombinase, accessory proteins, and single-stranded DNA binding (SSB) proteins for recombinase-promoted amplification.
[0114] For embodiments that produce DNA copies, the NTPs can be deoxyribonucleotide triphosphates (dNTPs). Generally, the four natural species dATP, dTTP, dGTP, and dCTP will be present in the DNA amplification reagent. However, analogs can be used if desired. For embodiments that produce RNA copies, the NTPs can be ribonucleotide triphosphates (rNTPs). Generally, the four natural species rATP, rUTP, rGTP, and rCTP are present in the RNA amplification reagent. However, analogs can be used if desired. The NTPs can be modified with fluorescent or radioactive groups. To increase the detectability and / or functional diversity of nucleic acids, an extremely large variety of synthetically modified nucleic acids have been developed for chemical and biological methods. These functionalized / modified molecules (such as nucleotide analogs) can be fully compatible with natural polymerases, thus maintaining the base pairing and replication properties of the natural counterparts.
[0115] Accordingly, the choice of polymerase is added to by other components of the amplification solution, and they generally correspond to compounds known in the art to be effective in supporting various polymerase activities. The concentrations of compounds such as dimethyl sulfoxide (DMSO), bovine serum albumin (BSA), polyethylene glycol (PEG), betaine, Triton X-100, denaturants (such as formamide), or MgCl2, etc. are known in the prior art to be important for optimal amplification, and thus the operator can easily adjust such concentrations for the methods of the present disclosure based on the examples given below and the knowledge generally available.
[0116] The rate of occurrence of an amplification reaction can be increased by increasing the concentration or amount of one or more active components of the amplification reaction. For example, the amount or concentration of polymerase, nucleotide triphosphates, or primers. In some cases, one or more active components of the amplification reaction whose amount or concentration is increased (or otherwise manipulated in the methods described herein) are non-nucleic acid components of the amplification reaction.
[0117] The amplification rate can also be increased in the methods described herein by modulating the temperature. For example, the amplification rate at one or more amplification sites can be increased by raising the temperature at one or more sites to the maximum temperature at which the reaction rate is reduced due to denaturation or other adverse events. The optimal or desired temperature can be determined from the known properties of the amplification components used or empirically for a given amplification reaction mixture. Such adjustments can be made based on a priori predictions or empirically of the primer melting temperature (Tm). In certain embodiments, the temperature of the amplification reaction is at least 35°C to no greater than 70°C. For example, the amplification reaction can be at least 35°C to no greater than 42°C, or at least 57°C to no greater than 63°C.
[0118] The rate of occurrence of an amplification reaction can be increased by increasing the activity of one or more amplification reagents. For example, cofactors that increase the extension rate of polymerase can be added to the reaction using that polymerase. In some embodiments, metal cofactors such as magnesium, zinc, or manganese can be added to the polymerase reaction or betaine can be added.
[0119] preparation of immobilized sample for sequencing
[0120] The result of bridge amplification is a population of clonal "bridged" amplification products at the amplification site. The two strands of the amplicon acid are immobilized at the 5' ends on the surface of the amplification site, where this attachment is derived from the original attachment of the capture nucleic acid (e.g., see Figure 1A , where the double-stranded amplicon 11 is depicted in a "bridged" orientation). The amplicons within the amplification site will be clonal and derived from the amplification of a single target nucleic acid, or have another amplicon with an acceptable level as described herein.
[0121] In addition to the free 3' ends of each strand of the bridged double-stranded amplicon, a large amount of unused capture nucleic acid remains on the surface of the amplification site after amplification to form a clonal cluster of bridged amplification products. The presence of the unused capture nucleic acid can contribute to increased noise and is therefore typically removed by contacting the array with a nuclease under conditions suitable for nuclease digestion of the unused capture nucleic acid. In one embodiment, the nuclease is an exonuclease, such as an exonuclease with 3' to 5' single-stranded DNA exonuclease activity. Examples of such exonucleases are exonuclease I. After nuclease treatment, the array is washed to remove the nuclease and the resulting nucleotides and / or nucleic acids from the amplification site.
[0122] To facilitate sequencing, one of the strands of the double-stranded bridge structure can be selectively removed from the surface to allow for efficient hybridization of the sequencing primer to the remaining immobilized strand. The selective removal of a particular strand is referred to herein as "linearization". Examples of suitable linearization methods are described herein and are described in more detail in Application No. WO 2007 / 010251 and U.S. Patent Application Publication 2012 / 0309634.
[0123] In one embodiment, linearization is achieved as follows: by cleaving one strand of the bridged double-stranded amplicon and then subjecting the resulting structure to conditions that remove the strand that is no longer attached to the surface of the amplification site. Cleavage can be accomplished by using a capture nucleic acid that contains a cleavage site. The cleavage site is typically located in a position such that most of one strand of the resulting bridge structure is free of the amplification site surface - no longer immobilized - and is prone to being lost after the removal step. For example, as Figure 1B shown, the cleavage site X in amplicon 11 is cleaved, leaving a shortened capture nucleic acid 13". One strand 16 of the bridge structure remains immobilized to the amplification site 10 at its 5' end, and the other strand 15' is no longer immobilized due to cleavage at X. In one embodiment, the 3' end of strand 16 remains annealed to the complementary bases of the shortened capture nucleic acid 13", thereby maintaining the bridge structure after linearization. The number of complementary bases between the 3' end of the strand 16 that maintains the bridge structure and the shortened capture nucleic acid 13" varies with general conditions and can be determined by one skilled in the art.
[0124] In one embodiment, the cleavage site is processed to remove nucleotides and form an abasic site. An "abasic site" is a nucleotide position in a nucleic acid from which the base moiety has been removed. Abasic sites can be chemically formed under artificial conditions or by the action of enzymes. Once formed, the abasic site can be cleaved (e.g., by using an endonuclease or other single-strand cleavage enzyme, exposure to heat or alkali treatment), thereby providing a means for site-specific cleavage of the nucleic acid.
[0125] In one embodiment, an abasic site can be created at a predetermined position on one strand of the immobilized amplicon. For example, this can be achieved by incorporating a specific nucleotide at the predetermined position.
[0126] In one embodiment, deoxyuridine (U) is incorporated into one of the capture nucleic acids attached to the surface of the amplification site. The uracil base can then be removed using uracil-DNA glycosylase (UDG), creating an abasic site on one strand. The polynucleotide strand comprising the abasic site can then be cleaved at the abasic site by treatment with an endonuclease (e.g., the DNA glycosylase-lyase endonuclease VIII), heat, or base. In a specific embodiment, the USER enzyme, commercially available from New England Biolabs (NEB#M5505S), is used to create a single nucleotide nick at the uracil base in the immobilizer. In one embodiment, the amplification site is exposed to a mixture containing a suitable glycosylase and one or more suitable endonucleases, typically at an activity ratio of at least about 2:1. Treatment with the endonuclease produces a 3'-phosphate moiety at the cleavage site, which can be removed with a suitable phosphatase such as alkaline phosphatase. For example, as Figure 1B shown, if cleavage site X is generated using the USER enzyme, the shortened capture nucleic acid 13” will terminate with a 3'-phosphate group.
[0127] In one embodiment, 8-oxoguanine is incorporated into one of the capture nucleic acids attached to the surface of the amplification site. The 8-oxoguanine base can then be removed using FPG glycosylase, creating an abasic site on one strand. In another embodiment, deoxyinosine is incorporated into one of the capture nucleic acids attached to the surface of the amplification site, and the deoxyinosine base can then be removed using the enzyme AlkA glycosylase, thereby creating an abasic site on one strand.
[0128] Advantages of this method include the option to release a free 3'-phosphate group on the cleaved strand, which can provide a starting point for sequencing a region of the complementary strand after phosphatase treatment (e.g., sequencing a region of Figure 1B strand 16). Since the cleavage reaction requires a residue that is not naturally present in DNA but otherwise does not depend on the sequence context, such as deoxyuridine, there is no possibility of glycosylase-mediated cleavage occurring elsewhere at positions in the duplex that are not desired. Another advantage obtained by cleavage of the abasic site in the duplex region of the immobilized amplicon generated by the action of UDG on uracil is that the first base incorporated in a synthesis sequencing reaction initiated at the free 3'-hydroxyl formed by cleavage will always be T. Thus, for all clone clusters at different amplification sites in an array cleaved in this manner to generate sequencing templates, the first base incorporated uniformly across the array will be T. This can provide a sequence-independent assay of the individual cluster intensities at the start of a sequencing run.
[0129] The addition of the exonuclease and linearization steps are known to those skilled in the art as necessarily separate steps. Treatment with an endonuclease generates a 3'-phosphate moiety at the cleavage site, and the presence of 3' phosphate is known to inhibit the activity of exonuclease I (Lehman and Nussbaum, 1964, J. Biol. Chem., 239:2628-2636). The inventors made the unexpected and surprising discovery that both the exonuclease and linearization steps can occur simultaneously by combining the enzymes. Simplifying these two steps into one step results in a faster sequencing run because now both steps are carried out simultaneously. In addition, combining these two steps has no adverse effect on the main metrics, read quality, dual indexing, or genome build metrics.
[0130] Abasic site generation and cleavage result in a free 5' end on the strand, which is no longer immobilized to the surface (e.g., as Figure 1B shown, one strand 16 of the bridge structure remains immobilized to the amplification site 10 at its 5' end, and the other strand 15' is no longer immobilized due to cleavage at X). This strand can be completely removed from the surface by exposing the amplification site to suitable conditions. In one embodiment, removal is by denaturation. Denaturation can be carried out thermally or isothermally, for example using chemical denaturation. The chemical denaturant can be urea, hydroxide, or formamide or other similar reagents. In another embodiment, removal can be achieved by treatment with an exonuclease having 5'-3' activity, such as lambda or T7 exonuclease. Removal of the unattached strand results in a remaining single strand, which can serve as a template for the polymerase.
[0131] Optionally, repair the 3' end of the nucleic acid at the amplification site. The exonuclease can remove some nucleotides at the 3' end of the nucleic acid after linearization. Without intending to limit, it is possible that the 3' end is slightly "breathing", resulting in a small number of nucleotides becoming single-stranded and available for exonuclease digestion. Repair can be achieved by exposing the nucleotides to a DNA polymerase, such as the DNA polymerase used for bridge amplification.
[0132] Removal of the unattached strand is optional. In one embodiment, remove the remaining 3'-phosphate group after generating the abasic site to leave a 3'-hydroxyl at the end of the cleaved capture nucleic acid ( Figure 1B , the shortened capture nucleic acid 13"). This capture nucleic acid can be used as a primer for a polymerase having strand displacement activity. Since the polymerase uses the immobilized strand ( Figure 1B , strand 16) as a template to synthesize a complementary strand, it displaces the unattached strand ( Figure 1B , strand 15').
[0133] composition
[0134] During or after the amplification clustering method described herein, different compositions can be obtained. In one embodiment, the composition comprises a glycosylase, an endonuclease, and an exonuclease. In one embodiment, the exonuclease has 3'-to-5' single-stranded DNA exonuclease activity, such as exonuclease I. In one embodiment, a class of glycosylases that can be present in the composition is uracil DNA glycosylase, and the endonuclease is DNA glycosylase-lyase endonuclease VIII. In one embodiment, a class of glycosylases that can be present in the composition is FPG glycosylase, and the endonuclease is DNA glycosylase-lyase endonuclease VIII. In one embodiment, a class of glycosylases that can be present in the composition is AlkA glycosylase, and the endonuclease is DNA glycosylase-lyase endonuclease VIII. The composition can include a double-stranded DNA substrate that includes a uracil cleavage site, an 8-oxoguanine cleavage site, or a deoxyinosine cleavage site. The composition can contain a double-stranded DNA substrate containing an abasic site. In some embodiments, the composition double-stranded DNA substrate can contain a single-stranded region, where the uracil cleavage site and the abasic site are present in the double-stranded region. Also provided is an array comprising a plurality of amplification sites, where each amplification site comprises a plurality of double-stranded DNA substrates attached to the amplification site.
[0135] for sequencing / sequencing method
[0136] Arrays of the present disclosure that have been produced by the methods described herein and that include amplified and linearized amplicons at the amplification sites can be used in any of a variety of applications. A particularly useful application is nucleic acid sequencing. An example is synthesis-based sequencing (SBS). In SBS, the extension of a nucleic acid primer along a nucleic acid template (e.g., a target nucleic acid or its amplicon) is monitored to determine the sequence of nucleotides in the template. The underlying chemical process can be a polymerization reaction (e.g., as catalyzed by a polymerase). In a particular polymerase-based SBS embodiment, fluorescently labeled nucleotides are added to the primer in a template-dependent manner (thereby extending the primer) so that the sequence of the template can be determined using detection of the order and type of nucleotides added to the primer. Multiple different templates at different sites of the arrays described herein can be subjected to SBS technology under conditions where events occurring for different templates can be distinguished due to their position in the array.
[0137] Flow cells provide a convenient format for accommodating arrays produced by the methods of the present disclosure and subjected to SBS or other detection techniques that involve repeated delivery of reagents in a cycle. For example, to initiate the first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be flowed into / onto a flow cell that houses an array of nucleic acid templates. Those sites of the array where primer extension results in incorporation of a labeled nucleotide can be detected. Optionally, the labeled nucleotides can further include reversible terminating properties that terminate further primer extension once a nucleotide has been added to the primer. For example, a nucleotide analogue having a reversible terminator moiety can be added to the primer such that subsequent extension does not occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments using reversible termination, the deblocking agent can be delivered to the flow cell (before or after detection occurs). Washing can be performed between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluid systems, and detection platforms that can be readily adapted to work with arrays produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. Patent No. 7,057,026; WO 91 / 06678; WO07 / 123,744; U.S. Patent No. 7,329,492; U.S. Patent No. 7,211,414; U.S. Patent No. 7,315,019; U.S. Patent No. 7,405,281, and U.S. Patent No. 8,343,746.
[0138] Other sequencing procedures using cyclic reactions, such as pyrosequencing, can be used. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent nucleic acid strand (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9(1996); Ronaghi, Genome Res. 11(1), 3-11(2001); Ronaghi et al., Science 281(5375), 363(1998); U.S. Patent Nos. 6,210,891; 6,258,568 and 6,274,320). In pyrosequencing, the released PPi can be detected by immediately converting it to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP produced can be detected by photons generated by luciferase. Thus, the sequencing reaction can be monitored by a luminescence detection system. The pyrosequencing procedure does not require an excitation radiation source for a fluorescence-based detection system. Useful fluid systems, detectors, and procedures for applying pyrosequencing to the arrays of the present disclosure are described, for example, in WIPO published patent application 2012 / 058096, US2005 / 0191698 A1, U.S. Patent No. 7,595,883, and U.S. Patent No. 7,244,559.
[0139] Ligation sequencing reactions are also useful, including those described, for example, in Shendure et al., Science 309:1728-1732(2005); U.S. Patent No. 5,599,675; and U.S. Patent No. 5,750,341. Some embodiments can include hybridization sequencing procedures, such as those described in Bains et al., Journal of Theoretical Biology 135(3), 303-7(1988); Drmanac et al., Nature Biotechnology 16, 54-58(1998); Fodor et al., Science 251(4995), 767-773(1995); and WO 1989 / 10977. In both ligation sequencing and hybridization sequencing procedures, repeated cycles of oligonucleotide delivery and detection are performed on the template nucleic acid (e.g., target nucleic acid or its amplicon) present at the array site. Fluid systems for the SBS method as described herein or in the references cited herein can be readily adapted to deliver reagents for ligation sequencing or hybridization sequencing procedures. Generally, the oligonucleotides are fluorescently labeled and can be detected using a fluorescence detector similar to the fluorescence detectors described for the SBS procedures herein or in the references cited herein.
[0140] Some embodiments can use methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between a polymerase carrying a fluorophore and γ-phosphate-labeled nucleotides or using zero-mode waveguides. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al., Science 299, 682-686 (2003); Lundquist et al., Opt. Lett. 33, 1026-1028 (2008); Korlach et al., Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008).
[0141] Some SBS embodiments include detecting protons released after incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and related techniques available from Ion Torrent (Guilford, Conn., a subsidiary of Life Technologies) or the sequencing methods and systems described in US2009 / 0026082A1; US2009 / 0127589 A1; US2010 / 0137143 A1; or US2010 / 0282617 A1. The methods described herein for amplifying a target nucleic acid can be readily applied to a substrate for detecting photons. More specifically, the methods described herein can be used for a clonal population of amplicons at a site of an array for detecting photons.
[0142] A useful application of the arrays of the present disclosure that have been generated by the methods described herein, for example, is gene expression analysis. Gene expression can be detected or quantified using RNA sequencing techniques such as those referred to as digital RNA sequencing. RNA sequencing techniques can be performed using sequencing methods known in the art, such as the methods described above. Gene expression can also be detected or quantified using hybridization techniques by direct hybridization to the array or using multiplex assays that detect their products on the array. For example, the arrays of the present disclosure that have been generated by the methods described herein can also be used to determine the genotype of a genomic DNA sample from one or more individuals. Exemplary methods for array-based expression and genotyping analysis that can be performed on the arrays of the present disclosure are described in U.S. Patent No. 7,582,420; 6,890,741; 6,913,884 or 6,355,431 or U.S. Patent Publication No. 2005 / 0053980A1; 2009 / 0186349A1 or US 2005 / 0181440A1.
[0143] Another useful application of the arrays generated by the methods described herein is single cell sequencing. When combined with indexing methods, single cell sequencing can be used in chromatin accessibility assays to generate profiles of active regulatory elements in thousands of single cells, and single cell whole genome libraries can be generated. Examples of single cell sequencing that can be performed on the arrays of the present disclosure are described in U.S. Published Patent Application 2018 / 0023119A1, U.S. Provisional Application Serial Nos. 62 / 673,023 and 62 / 680,259.
[0144] Advantages of the methods described herein are that they provide for the rapid and efficient creation of arrays from any of a variety of nucleic acid libraries. Accordingly, the present disclosure provides an integrated system that is capable of using one or more of the methods set forth herein to prepare an array and is also capable of detecting nucleic acids on the array using techniques known in the art (such as those exemplified herein). Thus, the integrated system of the present disclosure can include fluidic components, such as pumps, valves, reservoirs, fluid lines, etc., capable of delivering amplification reagents to an array of amplification sites. A particularly useful fluidic component is a flow cell. The flow cell can be configured and / or used in the integrated system to create the arrays of the present disclosure and to detect the arrays. Exemplary flow cells are described, for example, in US 2010 / 0111768A1 and U.S. Patent No. 8,951,781. As illustrated by the flow cell, one or more fluidic components of the integrated system can be used for both amplification methods and detection methods. By way of example of a nucleic acid sequencing embodiment, one or more fluidic components of the integrated system can be used for the amplification methods described herein and for the delivery of sequencing reagents in sequencing methods such as the sequencing methods described herein. Alternatively, the integrated system can include separate fluidic systems for performing the amplification method and for performing the detection method. Examples of integrated sequencing systems capable of creating nucleic acid arrays and determining nucleic acid sequences include, but are not limited to, the MiSeq TM , HiSeq2500 TM , NextSeq TM , MiniSeq TM , NovaSeq TM and iSeq TM sequencing platforms and the device described in U.S. Patent No. 8,951,781. Such devices can be modified according to the guidance set forth herein to prepare arrays.
[0145] Systems capable of performing the methods described herein need not be integrated with a detection device. Instead, they can also be stand-alone systems or systems integrated with other devices. Similar fluidic components to those exemplified above in the context of the integrated system can be used in such embodiments.
[0146] Whether integrated with detection capabilities or not, a system capable of performing the methods presented herein may include a system controller that is capable of executing a set of instructions to perform one or more steps of the methods, techniques, or processes described herein. For example, the instructions may direct the performance of steps to create an array under bridge amplification conditions. Optionally, the instructions may further direct the performance of steps to detect nucleic acids using the methods previously described herein. Useful system controllers may include any processor-based or microprocessor-based system, including systems using microcontrollers, reduced instruction set computers (RISC), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), logic circuits, and any other circuit or processor capable of performing the functions described herein. A set of instructions for the system controller may be in the form of a software program. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The software may take various forms, such as system software or application software. Additionally, the software may take the form of a collection of separate programs, or program modules or portions of program modules within a larger program. The software may also include modular programming in the form of object-oriented programming.
[0147] It should be understood that, for example, arrays of the present disclosure that have been produced by the methods described herein are not required to be used in detection methods. Instead, the arrays can be used to store nucleic acid libraries. Thus, the arrays can be stored in a state in which the nucleic acids are preserved. For example, the arrays can be stored in a dry state, a frozen state (e.g., in liquid nitrogen), or in a solution that protects the nucleic acids. Alternatively / additionally, the arrays can be used to replicate nucleic acid libraries. For example, the arrays can be used to create duplicate amplicons from one or more sites on the array.
[0148] Now referring to Figure 1A , a schematic diagram of an amplification site 10 with a member of a plurality of double-stranded amplicons 11 is shown. The depicted double-stranded amplicon 11 includes a first strand 15 and a second strand 16. Also shown are two populations of capture nucleic acids. A first population is shown attached to one end of the first strand 15 or bound to the surface of the amplification site 10. Also shown is a second population of capture nucleic acids attached to one end of the second strand 16 or bound to the surface of the amplification site 10. Figure 1A Also shown in [reference] are cleavage sites (marked with an X on the capture nucleic acid 13). In one embodiment, the capture nucleic acid 13 may comprise a P5 capture nucleic acid, while the other capture nucleic acid 14 may comprise a P7 capture nucleic acid.
[0149] Figure 1BShows cleavage at cleavage site X in the capture nucleic acid 13 attached to the first strand. Cleavage of the capture nucleic acid 13 results in (i) a shortened strand 15’ and (ii) a shortened capture nucleic acid 13’. The shortened strand 15’ is no longer attached to the amplification site 10. The shortened capture nucleic acid 13’ can have a 3’ phosphate terminus. In the depicted embodiment, the nucleotide present at the 3’ terminus of strand 16 remains annealed to the nucleotide present at the 3’ terminus of the shortened capture nucleic acid 13”. Due to the action of an exonuclease, such as exonuclease I, the unattached capture nucleic acids 13’ and 14’ are no longer present at the amplification site 10.
[0150] Figure 1C Shows the result of exposing Figure 1B the amplicons to denaturing conditions. The shortened strand 15’ that is not attached to the amplification site 10 is removed from strand 16, and the nucleotide at the 3’ terminus of strand 16 does not anneal to the nucleotide present at the 3’ terminus of the shortened capture nucleic acid 13”.
[0151] Figure 1D Shows the result of reannealing. The attached strand 16 reanneals to the shortened capture nucleic acid 13”.
[0152] Exemplary embodiments
[0153] Embodiment 1. A composition comprising:
[0154] Uracil DNA glycosylase,
[0155] an endonuclease, and
[0156] an exonuclease having 3’ to 5’ single-stranded DNA exonuclease activity.
[0157] Embodiment 2. The composition of Embodiment 1, wherein the exonuclease is exonuclease I.
[0158] Embodiment 3. The composition of Embodiment 1 or 2, wherein the endonuclease is DNA glycosylase-lyase endonuclease VIII.
[0159] Embodiment 4. The composition of any one of Embodiments 1-3, further comprising a double-stranded DNA substrate comprising a uracil cleavage site.
[0160] Embodiment 5. The composition of any one of Embodiments 1-4, further comprising a double-stranded DNA substrate comprising an abasic site.
[0161] Embodiment 6. The composition of any one of Embodiments 4-5, wherein the double-stranded DNA substrate comprises a single-stranded region, and wherein the uracil cleavage site and the abasic site are present in the single-stranded region.
[0162] Embodiment 7. A composition according to any one of Embodiments 4 to 6, further comprising an array comprising a plurality of amplification sites, wherein each amplification site comprises a plurality of said double-stranded DNA substrates attached to said amplification site.
[0163] Embodiment 8. A method of preparing nucleic acids for a sequencing reaction, the method comprising:
[0164] (a) providing an array comprising a plurality of amplification sites, wherein the amplification sites comprise
[0165] (i) a plurality of capture nucleic acids attached to said amplification site,
[0166] wherein a first population of said plurality of capture nucleic acids comprises a cleavage site, and
[0167] (ii) a plurality of cloned double-stranded modified target nucleic acids,
[0168] wherein the two strands of each double-stranded target nucleic acid are attached to the capture nucleic acid at their 5' ends,
[0169] wherein one strand is attached to the capture nucleic acid comprising said cleavage site, and
[0170] wherein said cleavage site is located in the double-stranded region of each double-stranded molecule;
[0171] (b) contacting the array with a composition comprising at least one enzyme that generates an abasic site at said cleavage site and an exonuclease comprising 3' to 5' single-stranded DNA exonuclease activity,
[0172] wherein cleavage occurs at said cleavage site,
[0173] wherein cleavage converts one strand of the double-stranded target nucleic acid into a first strand attached to said amplification site and a second strand not attached to said amplification site; and
[0174] wherein the length of the single-stranded capture nucleic acid comprising a free 3' end is shortened by said exonuclease.
[0175] Embodiment 9. The method of Embodiment 8, wherein said at least one enzyme that generates an abasic site at the cleavage site comprises uracil DNA glycosylase and an endonuclease.
[0176] Embodiment 10. The method of Embodiment 8 or 9, wherein said endonuclease is DNA glycosylase-lyase endonuclease VIII.
[0177] Embodiment 11. The method of any one of Embodiments 8 to 10, further comprising removing from the array the at least one enzyme that generates an abasic site at the cleavage site and the exonuclease.
[0178] Embodiment 12. The method of any one of Embodiments 8 - 11, further comprising subjecting the cleaved double-stranded target nucleic acid to conditions for removing the second strand that is not attached to the amplification site.
[0179] Embodiment 13. The method of Embodiment 12, wherein the conditions for removing the second strand comprise a denaturing agent, and wherein the denaturing agent results in immobilized single-stranded nucleic acid, the single-stranded nucleic acid comprising a target nucleic acid covalently attached to a second population of capture nucleic acids, and wherein the second population of capture nucleic acids is attached to the amplification site.
[0180] Embodiment 14. The method of any one of Embodiments 11 - 14, wherein the denaturing agent comprises formamide.
[0181] Embodiment 15. The method of any one of Embodiments 11 - 14, further comprising reannealing the immobilized single-stranded nucleic acid to a member of the first population of capture nucleic acids to produce immobilized partially single-stranded nucleic acid.
[0182] Embodiment 16. The method of any one of Embodiments 8 - 15, wherein the cleavage site is in the capture nucleic acid region of the double-stranded region of each double-stranded target nucleic acid.
[0183] Embodiment 17. The method of any one of Embodiments 8 - 16, wherein the cleavage site comprises uracil, wherein the uracil DNA glycosylase generates an abasic site, and wherein the endonuclease cleaves the abasic site.
[0184] Embodiment 18. The method of any one of Embodiments 8 - 17, wherein the exonuclease is exonuclease I.
[0185] Embodiment 19. The method of any one of Embodiments 13 or 15, further comprising hybridizing a sequencing primer to the single-stranded region of the immobilized single-stranded nucleic acid of Embodiment 13 or the immobilized partially single-stranded nucleic acid of Embodiment 15, thereby preparing single-stranded nucleic acid for a sequencing reaction.
[0186] Embodiment 20. The method of Embodiment 19, further comprising performing a sequencing reaction to determine the sequence of at least one region of the immobilized single-stranded nucleic acid or the immobilized partially single-stranded nucleic acid.
[0187] Embodiment 21. The method of Embodiment 19, wherein the sequencing reaction comprises sequencing by synthesis.
[0188] Embodiment 22. The method of any one of Embodiments 8 to 21, wherein the array is generated by amplifying a plurality of target nucleic acids using the capture nucleic acid as an amplification primer.
[0189] Embodiment 23. The method of Embodiment 22, wherein the amplification includes exclusion amplification. Examples
[0190] The present invention is illustrated by the following examples. It should be understood that the specific examples, materials, amounts, and procedures will be broadly interpreted in accordance with the scope and spirit of the present invention described herein.
[0191] Example 1
[0192] General assay methods and conditions
[0193] Unless otherwise stated, this describes the general assay conditions used in the examples described herein.
[0194] The experiment was run on a cBot (ILMN) using a v2.5 HiSeqX flow cell (ILMN). For Figure 2 , the P5 and P7 surface primers were used as substrates to test the enzyme activity, and the flow cell was used without amplifying clusters. During the experiment, various enzyme mixtures were pumped into the flow cell and incubated at 37 °C for 15 minutes. Exonuclease I and USER enzyme were both provided by New England Biolabs. After incubation, the flow cell lanes were washed with HT2 wash buffer (Illumina), and then the presence or absence of surface primers was determined by hybridization with P5' and P7' oligomers labeled with the fluorophore TET in HT1 hybridization buffer (Illumina). Fluorescent signals were detected by scanning on a Typhoon long platform imager (GE Healthcare Life Sciences).
[0195] For Figure 3 , cluster inoculation and amplification were achieved by mixing denatured DNA templates (human TruSeq Nano library) in the ExAmp mixture to a final concentration of 300 pM, and then this was pumped into the flow cell and incubated at 37 °C for 1 hour. As detailed in the figure, different lanes of the flow cell were then treated with different combinations of enzymes and treatments. "Repair" refers to typically 3 bridge amplification cycles, which are usually carried out to fill in the ends of the strands that can be chewed off during the exonuclease step. "USERExo" refers to a combined mixture of USER and ExoI. After these steps, the clusters were hybridized with sequencing primers and sequenced on the HiSeqX using standard methods and reagents (ILMN).
[0196] Example 2
[0197] The exonuclease and linearization reactions can be combined
[0198] Standard methods for generating clusters useful for genomic sequencing involve hybridizing a target nucleic acid to a cluster, amplifying the target nucleic acid, and then treating the cluster with exonuclease I to remove excess free surface primers. Then, in a separate step, the immobilized strand is linearized by creating a single nucleotide nick in a specific region of double-stranded DNA (dsDNA) to generate a template for sequencing. We tested whether the exonuclease and linearization steps could be combined. Lanes of a flow cell having amplification sites with attached surface primers were treated with exonuclease 1, an enzyme that creates a single nucleotide nick in a specific region of dsDNA (a combination of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII), or both.
[0199] Figure 2 Shown are that no treatment (lanes 1 and 8) or treatment with uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII (lane 2) does not result in loss of surface primers, while treatment with exonuclease I (lane 5) results in loss of substantially all surface primers. Treatment with uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII, followed by exonuclease I (lane 3) and treatment with a mixture of exonuclease and the combination of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII (lane 4) results in loss of essentially all surface primers. This was unexpected and surprising because treatment of dsDNA with the combination of uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII generates a 3′ phosphate, and the presence of a 3′ phosphate is known to inhibit the activity of exonuclease I (Lehman and Nussbaum, 1964, J. Biol. Chem., 239:2628-2636).
[0200] Example 3
[0201] Concurrent exonuclease treatment and linearization generates a template useful for sequencing
[0202] To determine whether concurrent use of DNA glycosylase and exonuclease has any adverse effects on sequencing metrics, a sequencing run was performed on a flow cell containing clusters generated using different combinations of DNA glycosylase and exonuclease. Quality scores for each run were determined in quadruplicate and are shown in Figure 3 As expected, a standard lane containing a sequence reaction with the P5 primer site (no exonuclease, no repair) has a Q30 value in the mid-50s (see arrow, Figure 3)。All other lanes are roughly equivalent, and these reactions carried out using DNA glycosylase and exonuclease in a single step yield optimal R2 readings with a Q30 of approximately 93%.
[0203] The complete disclosures of all patents, patent applications, and publications cited herein and electronically available materials (including, for example, nucleotide sequence submissions in, e.g., GenBank and RefSeq and amino acid sequence submissions in, e.g., SwissProt, PIR, PRF, PDB, and translations from the annotated coding regions in GenBank and RefSeq) are incorporated by reference in their entirety. Supplementary materials (such as supplementary tables, supplementary figures, supplementary materials and methods, and / or supplementary experimental data) cited in the publications are likewise incorporated by reference in their entirety. In the event of any inconsistency between the disclosure of the present application and the disclosure of any document incorporated herein by reference, the disclosure of the present application shall govern. The foregoing detailed description and examples are given for purposes of clear understanding only. No unnecessary limitations are to be understood therefrom. The present invention is not limited to the exact details shown and described, and variations that are obvious to one of ordinary skill in the art will be included within the invention as defined by the claims.
[0204] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, etc. used in the specification and claims are to be understood as being modified in all instances by the term "about". Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present invention. In any case, without limiting the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0205] While the numerical ranges and parameters setting forth the broad scope of the invention are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. However, all numerical values inherently contain a certain range necessarily resulting from the standard deviation found in their respective testing measurements.
[0206] Unless otherwise indicated, all headings are for the convenience of the reader and should not be used to limit the meaning of the text following the heading.
Claims
1. A composition comprising: uracil DNA glycosylase, DNA glycosylase-lyase endonuclease VIII, and exonuclease I; and an array comprising a plurality of amplification sites.
2. The composition of claim 1, further comprising a double-stranded DNA substrate comprising a uracil cleavage site.
3. The composition of claim 1 or claim 2, further comprising a double-stranded DNA substrate comprising an abasic site.
4. The composition of claim 3, wherein the double-stranded DNA substrate comprises a single-stranded region, and wherein the uracil cleavage site and the abasic site are present in the double-stranded region.
5. The composition of claim 2 or 4, wherein each amplification site of the array comprises a plurality of the double-stranded DNA substrates attached to the amplification site.
6. The composition of claim 3, wherein each amplification site of the array comprises a plurality of the double-stranded DNA substrates attached to the amplification site.
7. A method for preparing a nucleic acid for a sequencing reaction, the method comprising: (a) providing an array comprising a plurality of amplification sites, wherein the amplification sites comprise (i) a plurality of capture nucleic acids attached to the amplification sites, wherein a first population of the plurality of capture nucleic acids comprises a cleavage site, and (ii) a plurality of clonal double-stranded modified target nucleic acids, wherein the two strands of each double-stranded target nucleic acid are attached at their 5' ends to a capture nucleic acid, wherein one strand is attached to a capture nucleic acid comprising the cleavage site, and wherein the cleavage site is located in the double-stranded region of each double-stranded molecule; (b) contacting the array with a composition comprising at least one enzyme that generates an abasic site at the cleavage site and an exonuclease comprising 3' to 5' single-stranded DNA exonuclease activity, wherein cleavage occurs at the cleavage site, wherein cleavage converts one strand of the double-stranded target nucleic acid into a first strand attached to the amplification site and a second strand not attached to the amplification site; and wherein the length of the single-stranded capture nucleic acid comprising a free 3' end is shortened by the exonuclease; and wherein the at least one enzyme that generates an abasic site at the cleavage site comprises uracil DNA glycosylase and DNA glycosylase-lyase endonuclease VIII; and the exonuclease is exonuclease I.
8. The method of claim 7, further comprising removing the at least one enzyme that generates an abasic site at the cleavage site and the exonuclease from the array.
9. The method of claim 7 or claim 8, further comprising subjecting the cleaved double-stranded target nucleic acid to conditions for removing the second strand not attached to the amplification site.
10. The method of claim 9, wherein the conditions for removing the second strand comprise a denaturing agent, wherein the denaturing agent results in immobilized single-stranded nucleic acid comprising the target nucleic acid covalently attached to a second population of capture nucleic acids, wherein the second population of capture nucleic acids is attached to the amplification site.
11. The method of claim 10, wherein the denaturing agent comprises formamide.
12. The method of claim 10, further comprising annealing the immobilized single-stranded nucleic acid to a member of the first population of the capture nucleic acid to produce an immobilized partially single-stranded nucleic acid.
13. The method of claim 7, wherein the cleavage site is in the capture nucleic acid region of the double-stranded region of each double-stranded target nucleic acid.
14. The method of claim 7, wherein the cleavage site comprises uracil, wherein the uracil DNA glycosylase generates an abasic site, and wherein the endonuclease cleaves the abasic site.
15. The method of claim 10, further comprising hybridizing a sequencing primer to the immobilized single-stranded nucleic acid in claim 10, thereby preparing a single-stranded nucleic acid for a sequencing reaction.
16. The method of claim 12, further comprising hybridizing a sequencing primer to the single-stranded region of the immobilized partially single-stranded nucleic acid, thereby preparing a single-stranded nucleic acid for a sequencing reaction.
17. The method of claim 15 or claim 16, further comprising performing a sequencing reaction to determine the sequence of at least one region of the immobilized single-stranded nucleic acid or the immobilized partially single-stranded nucleic acid.
18. The method of claim 15 or claim 16, wherein the sequencing reaction comprises sequencing by synthesis.
19. The method of any one of claims 7, 8, and 10-16, wherein the array is produced by amplifying a plurality of target nucleic acids using the capture nucleic acid as an amplification primer.
20. The method of claim 19, wherein the amplification comprises exclusion amplification.
21. The method of claim 9, wherein the array is produced by amplifying a plurality of target nucleic acids using the capture nucleic acid as an amplification primer.
22. The method of claim 17, wherein the array is produced by amplifying a plurality of target nucleic acids using the capture nucleic acid as an amplification primer.
23. The method of claim 18, wherein the array is produced by amplifying a plurality of target nucleic acids using the capture nucleic acid as an amplification primer.
Citation Information
Patent Citations
Methods and compositions for whole genome amplification and genotyping
US20050053980A1
Nucleic acid sequencing using microsphere arrays
US20050181440A1
Nucleic acid sequencing using microsphere arrays
US20050191698A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090127589A1