Method to improve the priority of polynucleotide cluster cloning ability
Patent Information
- Application Number
- KR1020207037261
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-19
- Filing Date
- 2019-12-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2039-12-18
Smart Images

Figure 112021069960524-PCT00006_ABST
Abstract
Description
Technology Field
[0001] delete
[0002] The present disclosure relates, above all, to the generation of clusters in the sequencing of amplicons using exclusion amplification of target nucleic acids, and more specifically, to the increase in the number of monoclonal clusters. Background Technology
[0003] As next-generation sequencing (NGS) technology has improved, sequencing speed and data output have increased significantly, resulting in the high sample throughput of current sequencing platforms. About 10 years ago, the Illumina Genome Analyzer could generate up to 1 GB of sequence data per run. Today, Illumina NovaSeq TM The System Series can generate up to 2TB of data in just two days, which is a capacity increase of more than 2,000 times.
[0004] One mode of realizing this increased capacity is cluster generation. Cluster generation may involve the creation of a library in which the library members contain a universal sequence present at each end. The library is loaded into a flow cell, and the individual members of the library are captured in the lawns of surface-binding oligos complementary to the universal sequence. Subsequently, each member is amplified into distinct clone clusters through bridge amplification. Once cluster generation is complete, individual clusters may contain approximately 1,000 copies of a single member of the library, and the library is ready for sequencing.
[0005] One method of bridge amplification is exclusion amplification (ExAmp), also known as kinetic exclusion amplification. This method is a recombinase-driven amplification reaction that amplifies a library using a patterned array and isothermal conditions, generating clonal clusters in the array wells with faster amplification and less reagent usage. While the exclusion amplification method has proven highly useful for generating clonal clusters, it also produces more polyclonal wells due to conditions that result in more occupied wells.
[0006] Next-generation sequencing (NGS) technology relies on the highly parallel sequencing of monoclonal populations of amplicons generated from a single target nucleic acid. Sequencing monoclonal populations of amplicons results in a much higher signal-to-noise ratio, increased intensity, and an increased percentage of clusters passing through filters, all of which contribute to improved data output and data quality.
[0007] The exclusion amplification method enables the amplification of a single target nucleic acid per well in a patterned flow cell and the generation of a monoclonal population of amplicons in the wells. Typically, the amplification rate of the first target nucleic acid captured in the well is faster than the much slower transport and capture rate of the target nucleic acid in the well. The first target nucleic acid captured in the well can be rapidly amplified to fill the entire well, thereby preventing additional target nucleic acids from being captured in the same well. Alternatively, if a second target nucleic acid is attached to the same well after the first target nucleic acid, the rapid amplification of the first target nucleic acid often sufficiently fills the well, resulting in a signal that passes through the filter. The use of exclusion amplification can also result in super-Poisson distributions of monoclonal wells, meaning that the proportion of monoclonal wells in the array may exceed the proportion predicted by the Poisson distribution.
[0008] Increasing the super-Poisson distribution of useful clusters is highly desirable because more monoclonal wells lead to more data output, but seeding target nucleic acids into wells generally follows a spatial Poisson distribution, where the trade-off for more occupied wells is more polyclonal wells. One way to achieve a higher super-Poisson distribution is to make seeding occur rapidly and then introduce a delay between the seeded target nucleic acids. This delay, referred to as "kinetic delay" because it is believed to occur through biochemical reaction kinetics, causes one seeded target nucleic acid to start earlier than another seeded target.
[0009] Exclusion amplification functions by using a recombinase to facilitate the penetration of a primer (e.g., a primer attached to the well) into double-stranded DNA (e.g., a target nucleic acid) upon detecting a sequence match. To maximize amplification efficiency, it is standard practice for exclusion amplification to use complete identification between the penetration primer and the adapter sequence. The inventors identified a method for encoding a lag in the seeded target nucleic acid by adjusting the degree of homology between the target nucleic acid adapter and the primer attached to the well. By reducing the average homology between the penetration primer and the adapter sequence, the rate of the so-called monoclonality of the well was remarkably improved, even though the average amplification rate was reduced. Generally, as more mismatches were introduced, amplification efficiency decreased. Unexpectedly, when a mixture of adapter sequences with both high and low amplification efficiencies was used, the mixture did not exhibit the average performance of the individual components (intermediate between high and low efficiency), but it outperformed the performance of all single-type adapter sequences in both clustering and strength passing through the filter.
[0010] definition
[0011] Unless otherwise specified, the terms used herein shall be understood to have their general meanings in the relevant technology. Various terms used herein and their meanings are described below.
[0012] As used herein, the term “amplicon,” when used in reference to nucleic acids, means a replication product of a nucleic acid, wherein the product has a nucleotide sequence identical to or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon may be produced by any various amplification method using a nucleic acid, e.g., a target nucleic acid or an amplicon of a nucleic acid as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rotational amplification (RCA), ligation extension, or ligation chain reaction. An amplicon may be a nucleic acid molecule having a single replication of a specific nucleotide sequence (e.g., a polymerase extension product) or multiple replications of a nucleotide sequence (e.g., a concatameric product of an RCA). The first amplicon of the target nucleic acid is typically a complementary replication. A subsequent amplicon is a copy generated from the target nucleic acid or from the first amplicon after the generation of the first amplicon. The subsequent amplicon may be substantially complementary to the target nucleic acid or may have a sequence substantially identical to the target nucleic acid.
[0013] As used herein, the term “amplification site” refers to a site within or on an array where one or more amplicons may be generated. The amplification site may be further configured to contain, retain, or attach at least one amplicon generated at the site.
[0014] As used herein, the term “array” refers to a group of sites that can be distinguished from one another based on their relative positions. Different molecules present in different sites of the array can be distinguished from one another based on the positions of the sites within the array. Individual sites of the array may contain one or more molecules of a specific type. For example, a site may contain a single target nucleic acid molecule having a specific sequence, or a site may contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Sites of the array may be different features located on the same substrate. Exemplary features include, without limitation, wells in the substrate, beads (or other particles) in or on the substrate, protrusions from the substrate, ridges on the substrate, or channels in the substrate. Sites of the array may be distinct substrates each possessing different molecules. Different molecules attached to distinct substrates may be identified based on the position of the substrate on the surface to which the substrate is assembled, or based on the position of the substrate in a liquid or gel. An exemplary array in which separate substrates are located on the surface includes, but is not limited to, having beads within the wells.
[0015] As used herein, the term “capacity,” when used in reference to nucleic acid material, means the maximum amount of amplicon derived from the target nucleic acid that can occupy a site of the nucleic acid material, for example. For example, this term may refer to the total number of nucleic acid molecules that can occupy a site under certain conditions. For example, other measurements may also be used, including the total mass of the nucleic acid material or the total number of copies of a specific nucleotide sequence that can occupy a site under certain conditions. Typically, the capacity of a site for the target nucleic acid will be substantially the same as the capacity of a site for the amplicon of the target nucleic acid.
[0016] As used herein, the term “capture agent” refers to a substance, chemical, molecule, or moiety thereof capable of attaching, retaining, or binding to a target molecule (e.g., target nucleic acid). Exemplary capture agents include, but are not limited to, a capture nucleic acid complementary to at least a portion of the modified target nucleic acid (e.g., a universal capture binding sequence), a member of a receptor-ligand binding pair capable of binding to the modified target nucleic acid (or a linkage moiety attached thereto) (e.g., avidin, streptavidin, biotin, lectin, carbohydrate, nucleic acid binding protein, epitope, antibody, etc.), or a chemical reagent capable of forming a covalent bond with the modified target nucleic acid (or a linkage moiety attached thereto). In one embodiment, the capture agent is a nucleic acid. A nucleic acid capture agent may also be used as an amplification primer.
[0017] The terms "P5" and "P7" may be used to refer to nucleic acid capture agents. The terms "P5'" (P5 prime) and "P7'" (P7 prime) refer to the complements of P5 and P7, respectively. It will be understood that any suitable nucleic acid capture agent may be used in the method presented herein, and that the use of P5 and P7 is merely an exemplary embodiment. The use of nucleic acid capture agents such as P5 and P7 on a flow cell is known in the relevant art, as exemplified by the disclosures of International Patents WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. Those skilled in the art will recognize that nucleic acid capture agents can also function as amplification primers. For example, any suitable nucleic acid capture agent, whether immobilized or present in solution, may function as a forward amplification primer and may be useful in the methods presented herein for hybridization of a sequence (e.g., a universal capture binding sequence) and amplification of the sequence. Similarly, any suitable nucleic acid capture agent, whether immobilized or present in solution, may function as a reverse amplification primer and may be useful in the methods presented herein for hybridization of a sequence (e.g., a universal capture binding sequence) and amplification of the sequence. In light of the teachings of this disclosure and the general knowledge available, those skilled in the art will understand how to design and use sequences suitable for the capture and amplification of target nucleic acids as presented herein.
[0018] As used herein, the term “universal sequence” refers to a sequence region common to two or more target nucleic acids, wherein the molecules also have different sequence regions. A universal sequence present in different members of a molecular set may allow the capture of multiple different nucleic acids by using a group of capture nucleic acids complementary to a portion of the universal sequence, for example, a universal capture binding sequence. Non-limiting examples of universal capture binding sequences include sequences identical to or complementary to P5 and P7 primers. Other non-limiting examples of universal capture binding sequences described in detail herein include sequences having reduced identity (e.g., one or more discrepancies) or reduced complementarity with respect to P5 and P7 primers, and / or have a length shorter than that of P5 and P7 primers. Similarly, a universal sequence present in different members of a molecular set may allow the replication or amplification of multiple different nucleic acids by using a portion of the universal sequence, for example, a group of universal primers complementary to a universal primer binding site. The target nucleic acid molecule may be modified to attach a universal adapter (also referred to as an adapter herein) at one or both ends of different target sequences, for example, as described herein.
[0019] As used herein, the term "adapter" and its derivatives, e.g., universal adapter, generally refer to any linear oligonucleotide that can be ligated to a target nucleic acid. In some embodiments, the adapter is substantially non-complementary to the 3' or 5' end of any target sequence present in the sample. In some embodiments, suitable adapter length ranges are about 10 to 100 nucleotides, about 12 to 60 nucleotides, and about 15 to 50 nucleotides. Generally, the adapter may comprise any combination of nucleotides and / or nucleic acids. In some embodiments, the adapter may comprise one or more cleavable groups at one or more locations. In another embodiment, the adapter may comprise a sequence that is substantially identical to or substantially complementary to a primer, e.g., at least a portion of the capture nucleic acid. In some embodiments, the adapter may include a barcode, also referred to herein as a tag or index, to assist in downstream error testing, identification, or sequencing. The terms "adaptor" and "adapter" are used interchangeably.
[0020] As defined herein, “sample” and derivatives thereof are used in the broadest sense and include all specimens, cultures, etc. suspected of containing target nucleic acids. In some embodiments, the sample includes nucleic acids in the form of DNA, RNA, PNA, LNA, chimeric, or hybrid. The sample may include any biological, clinical, surgical, agricultural, atmospheric, or aquatic-based specimens containing one or more nucleic acids. The term also includes any isolated nucleic acid samples, such as genomic DNA, fresh frozen or formalin-fixed paraffin-embedded nucleic acid samples. Additionally, the sample may consider contaminated bacterial DNA present in a sample containing plant or animal DNA, a single individual, a collection of nucleic acid samples of genetically related members, nucleic acid samples of genetically unrelated members, nucleic acid samples of a single individual (matching), such as tumor samples and normal samples, or a single source sample containing two different forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject. In some embodiments, the source of the nucleic acid material may include nucleic acid obtained from a newborn, for example, as is typically used in newborn screening.
[0021] As used herein, the term “clonal population” refers to a group of nucleic acids that are uniform with respect to a specific nucleotide sequence. A uniform sequence is typically at least 10 nucleotides long, but can be much longer, for example, including at least about 50, at least 100, at least 250, at least 500, or at least 1000 nucleotides. A clonal population may be derived from a single target nucleic acid. Typically, all nucleic acids in a clonal population will have the same nucleotide sequence. It will be understood that a small number of mutations may occur in a clonal population (e.g., due to amplification artifacts) without deviating from clonality. It will also be understood that a small number of different target nucleic acids may occur in a clonal population (e.g., due to target nucleic acids that are amplified or not amplified to a limited extent) without deviating from clonality.
[0022] As used herein, the term “different” means that, when used in relation to nucleic acids, the nucleic acids have nucleotide sequences that are not identical to each other. Two or more nucleic acids may have different nucleotide sequences along their entire length. Alternatively, two or more nucleic acids may have different nucleotide sequences along a significant portion of their length. For example, two or more nucleic acids may have different target nucleotide sequence portions while also having the same universal sequence region.
[0023] As used herein, the term “fluidic access” refers to the ability of molecules to move in or through a fluid to come into contact with or enter a site, when used in relation to molecules in a fluid and sites in contact with the fluid. The term may also refer to the ability of molecules to separate from or exit a site and enter a solution. Fluidic access may occur when there is no barrier preventing molecules from entering a site, coming into contact with a site, separating from a site, and / or exiting a site. However, fluidic access is understood to exist even if diffusion is delayed, reduced, or altered, unless access is absolutely prevented.
[0024] As used herein, the term “double strand” means that, when used in relation to a nucleic acid molecule, substantially all nucleotides of the nucleic acid molecule are hydrogen-bonded to complementary nucleotides. A partially double-stranded nucleic acid may have at least 10%, at least 25%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of nucleotides hydrogen-bonded to complementary nucleotides.
[0025] As used herein, the term “each” is intended to identify individual items of a set when used in relation to a set of items, unless the context clearly indicates otherwise, but does not necessarily refer to all items of the set.
[0026] As used herein, the term "excluded volume" refers to the volume of space occupied by a specific molecule due to the exclusion of other molecules.
[0027] As used herein, the term "interstitial region" refers to a region on a substrate or on a surface that separates another region of the substrate or surface. For example, an interstitial region may separate one feature of an array from another feature of the array. The two separated regions may be discrete and thus lack contact with each other. In other examples, an interstitial region may separate a first portion of a feature from a second portion of a feature. The separation provided by the interstitial region may be partial or complete separation. The interstitial region typically has a surface material different from the surface material of the feature on the surface. For example, the feature of the array may have an amount or concentration of a capture agent exceeding the amount or concentration present in the interstitial region. In some embodiments, the capture agent may not be present in the interstitial region.
[0028] As used herein, the term “polymerase” is intended to be consistent with its use in the art and includes, for example, enzymes that use nucleic acid as a template strand to produce a complementary copy of a nucleic acid molecule. Typically, DNA polymerases bind to the template strand, move down the template strand, and sequentially add nucleotides to the free hydroxyl group at the 3’ end of the growing strand of nucleic acid. DNA polymerases typically synthesize complementary DNA molecules from a DNA template, and RNA polymerases typically synthesize RNA molecules from a DNA template (transcription). Polymerases may use short RNA or DNA strands called primers to initiate strand growth. Some polymerases can substitute strands upstream of the site where bases are added to the strand. Such polymerases are referred to as strand substitution, which means that the polymerase has the activity of removing the complementary strand from the template strand being read by itself. Exemplary polymerases having strand substitution activity include, without limitation, Bsu (Bacillus subtilis), Bst (Bacillus stearothermophilus) polymerase, exo-Klenow polymerase, or large fragments of sequence-grade T7 exo-polymerase. Some polymerases degrade the preceding strand and effectively substitute it with the chain growing behind it (5' exonuclease activity). Some polymerases have activity to degrade the following strand (3' exonuclease activity). Some useful polymerases are modified by mutation or other methods to reduce or eliminate 3' and / or 5' exonuclease activity.
[0029] As used herein, the term “nucleic acid” is intended to be consistent with its use in the art and includes naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs may be hybridized to nucleic acids in a sequence-specific manner or may be used as templates for the replication of specific nucleotide sequences. Naturally occurring nucleic acids generally have a backbone containing phosphodiester bonds. Analog structures may have alternative backbone linkages including any variety known in the art. Naturally occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA)) or a ribose sugar (e.g., found in ribonucleic acid (RNA)). Nucleic acids may contain any variety of analogs of these sugar moiety known in the art. Nucleic acids may contain natural or non-natural bases. In this regard, natural deoxyribonucleic acid may have one or more bases selected from adenine, thymine, cytosine, or guanine, and ribonucleic acid may have one or more bases selected from uracil, adenine, cytosine, or guanine. Useful non-natural bases that may be included in nucleic acids are known in the art. The term “target,” when used in relation to nucleic acids, is intended as a semantic identifier for the nucleic acid in the context of the method or composition described herein and does not necessarily limit the structure or function of the nucleic acid beyond what is clearly otherwise indicated. A target nucleic acid having a universal sequence at each end, for example, a universal adapter at each end, may be referred to as a modified target nucleic acid.
[0030] As used herein, the terms “recombinase loading protein” and “recombinase” are used interchangeably and are intended to be consistent with their applicable uses in the art, and include, for example, RecA protein, T4 UvsX protein, RB69 bacteriophage UvsX protein, any homologous protein or protein complex from any phyla, or functional variants thereof. The eukaryotic RecA homologue is generally referred to as Rad51, after the name of the first identified member of this group. Other non-homologous recombinases, for example, RecT or RecO, may be used instead of RecA.
[0031] As used herein, the term “single-strand binding protein,” also referred to as “SSB protein” or “SSB,” is intended to refer to any protein having a binding function to a single-strand nucleic acid, for example, to prevent premature annealing, to protect the single-strand nucleic acid from nuclease degradation, to remove secondary structures from the nucleic acid, or to promote the replication of the nucleic acid. The term is intended to include, but not be limited to, proteins formally identified as single-strand binding proteins by, for example, the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). Exemplary single-strand binding proteins include, but are not limited to, Escherichia coli SSB, T4 gp32, T7 gene 25 SSB, phage phi29 SSB, RB69 bacteriophage gp32 protein, any homologous protein or protein complex from any phylum, or functional variants thereof.
[0032] As used herein, the term “accessory protein” is intended to refer to any protein having the function of interacting with recombinase and single-strand binding proteins to assist in the nucleation of UvsX filaments on ssDNA. The terms “accessory protein,” “recombinase accessory protein,” and “recombinase helper protein” are used interchangeably. Exemplary accessory proteins include, but are not limited to, T4 UvsY, RB69 bacteriophage UvsY protein, E. coli RecO, E. coli RecR, any homologous protein or protein complex from any phylum, and functional variants thereof.
[0033] As used herein, the term “transport” refers to the movement of molecules through a fluid. The term may include passive transport, e.g., the movement of molecules along a concentration gradient (e.g., passive diffusion). The term may also include active transport, in which molecules can move along or against their own concentration gradient. Thus, transport may include applying energy to move one or more molecules in a desired direction or to a desired location, e.g., an amplification site.
[0034] As used herein, the term “rate” is intended to be consistent with its corresponding meaning in chemical kinetics and biochemical kinetics when used in relation to transport, amplification, capture, or other chemical processes. The rates for two processes may be compared in relation to the maximum rate (e.g., at saturation), the rate before steady state (e.g., before equilibrium), the kinetic rate constant, or other measurements known in the art. In specific examples, the rate of a particular process may be determined in relation to the total time for the completion of the process. For example, the amplification rate may be determined in relation to the time taken for the amplification to be completed. However, the rate of a particular process does not need to be determined in relation to the total time for the completion of the process.
[0035] The term "and / or" means one or all of the listed elements or any combination of two or more of the listed elements.
[0036] The terms “preferred” and “preferably” refer to embodiments of the present invention that may provide a specific benefit in a specific situation. However, other embodiments may also be preferred in the same or different situations. Furthermore, the mention of one or more preferred embodiments does not imply that other embodiments are not useful, nor is it intended to exclude other embodiments from the scope of the present invention.
[0037] "Complies" and variations thereof do not have a restrictive meaning where such terms exist in the description of the invention and the claims.
[0038] Where an embodiment is described in this specification with terms such as “comprising,” “comprising,” or “comprising,” it may be understood that other similar embodiments described with respect to “being made up” and / or “essentially made up” are also provided.
[0039] Unless otherwise specified, the singular expression and "at least one" are used interchangeably and mean one or more than one.
[0040] "Suitable" conditions for the occurrence of events such as the hybridization of two nucleic acid sequences are conditions that do not prevent such events from occurring. Therefore, these conditions allow, enhance, facilitate, and / or aid the event.
[0041] As used herein, with respect to a composition, article, or nucleic acid, "provision" means to manufacture the composition, article, or nucleic acid, to purchase the composition, article, or nucleic acid, or otherwise to acquire the compound, composition, article, or nucleic acid.
[0042] Additionally, in this specification, a reference to a numeric range by an endpoint includes all numbers in the sub-range within that range (e.g., 1 to 5 include 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).
[0043] Throughout this specification, references to “one embodiment,” “an embodiment,” “a predetermined embodiment,” or “some embodiment,” etc., mean that a specific feature, configuration, composition, or feature described in relation to such embodiment is included in at least one embodiment of this disclosure. Accordingly, such expressions present in various parts throughout this specification do not necessarily refer to the same embodiment of this disclosure. Furthermore, a specific feature, configuration, composition, or feature may be combined in any suitable manner in one or more embodiments.
[0044] In the case of any method disclosed herein comprising individual steps, the steps may be performed in any executable order. And, where appropriate, any combination of two or more steps may be performed simultaneously.
[0045] The above summary of the invention is not intended to describe each disclosed embodiment or all embodiments of the invention. The following description illustrates exemplary embodiments more specifically. Throughout this document, guidance is provided in several places through a list of embodiments, and these embodiments may be used in various combinations. In each case, the list mentioned functions only as a representative group and should not be interpreted as an exclusive list. Brief explanation of the drawing
[0046] The following detailed description of exemplary embodiments of the present invention is best understood when read together with the following drawings. FIGS. 1a and FIGS. 1b are schematic diagrams of an exemplary example of first and second capture sequences attached to wells of an array (Fig. 1a), and an exemplary example of a target nucleic acid having universal adapters attached to each end (Fig. 1b). FIG. 2 is a schematic diagram of an exemplary example of a hybridized single strand of a first capture nucleic acid and a target nucleic acid attached to a well of an array. FIG. 3a illustrates individual and group density pass filters of adapters. k / mm 2 is thousands per square millimeter. Lane represents the lane of the flow cell shown in Table 2 of the embodiment. FIG. 3b illustrates the ratio of final reads associated with individual mutant adapters. Figures 4a and 4b illustrate schematic diagrams of exemplary examples of strand intrusion and replication ("RPA" refers to recombinase polymerase amplification). In Figure 4a, the recombinase facilitates the intrusion of a free P7 primer into a double-strand template containing a homologous sequence (i.e., a matching P7 end). Although perfect homology is not required (indicated here by two deliberate mismatches introduced into the P7), the intrusion and amplification rates will be reduced by a reduced homologous adapter (indicated here by a small arrow to the mutant strand). In Figure 4b, recombinase-mediated intrusion from both ends occurs with an unmutated lawn primer and effectively corrects for mutations in the daughter strand, converting it back into a perfect adapter. However, since the homology between the original strand and the lawn strand is reduced, the time delay until the first replication occurs is proportional to the number and degree of mutations. Figures 5a and 5b illustrate the effect of short and mutant adapter libraries on amplification rates. In Figure 5a, a successful replication converts each template into a perfect template. However, the time constant for this conversion depends on the degree of non-homology to be overcome (larger proportions are indicated by thick arrows). In Figure 5b, slow amplification rates in short and mutant adapter libraries are indicated by a right shift of the real-time amplification curve. FIG. 6 illustrates competition between different templates for clone dominance in individual pads. Seeded templates are indicated with amplification bias (i.e., motion delay), where 1 = fastest and 6 = slowest. The same molar ratio of the templates is not required or desirable. A larger number of faster templates is desirable. However, the slowest template (6) can also fill the pad with a single clone cluster if there is no competition in the pad. Schematic drawings are not necessarily to scale. Similar numbers used in drawings represent similar components, steps, etc. However, you will understand that using numbers to refer to components in a given drawing is not intended to restrict components in other drawings labeled with the same number. Furthermore, using different numbers to refer to components is not intended to indicate that the component with that different number cannot be the same as or similar to the component with yet another number. Specific details for implementing the invention
[0047] The present invention provides a composition and a method related to increasing the production of monoclonal clusters that can be used for sequencing analysis.
[0048] The present disclosure provides a method for amplifying nucleic acids and a method for determining nucleic acid sequences. In one embodiment, the method comprises the step of providing an amplification reagent comprising (i) an array of amplification sites and (ii) a solution having a plurality of different target nucleic acids. The amplification sites comprise at least two groups of capture nucleic acids. One group, the first group, comprises a first capture sequence, and the second group comprises a second capture sequence. The different target nucleic acids comprise a first universal capture binding sequence at the 3' end. In one embodiment, the target nucleic acid is double-stranded. The first universal capture binding sequence has less affinity for the first capture sequence than the first universal capture binding sequence which has 100% complementarity with the first capture sequence. For example, as illustrated in FIG. 1a, the nucleic acid (100) of the first capture nucleic acid group comprises a first capture sequence (110), wherein the nucleic acid (100) is attached to the surface of an amplification site (120). FIG. 1b illustrates a double-stranded target nucleic acid (130) comprising a universal adapter (140) at each end and a first universal capture binding sequence (150) at the 3' end of each universal adapter (140).
[0049] Optionally, different target nucleic acids also include a second universal capture binding sequence at the 5' end. The complement of the second universal capture binding sequence has less affinity for the second capture sequence than the second universal capture binding sequence having a complement that has 100% complementarity to the second capture sequence. For example, as shown in FIG. 1a, the nucleic acid (160) of the second capture nucleic acid group includes a first capture sequence (170), wherein the nucleic acid (160) is attached to the surface of the amplification site (120). FIG. 1b shows a double-stranded target nucleic acid (130) including a universal adapter (140) at each end and a second universal capture binding sequence (180) at the 5' end of each universal adapter (140).
[0050] The method further comprises the step of reacting an amplification reagent to generate a plurality of amplification sites, each having a clone population of amplicons from individual target nucleic acids, from a solution. The reaction comprises transporting different target nucleic acids to the amplification sites and amplifying the target nucleic acids at the amplification sites. For example, as illustrated in FIG. 2, the nucleic acid (200) of the first capture nucleic acid population comprises a first capture sequence (210), wherein the nucleic acid (200) is attached to the surface of the amplification site (220). One strand of the target nucleic acid (230) comprising a first universal capture binding sequence (250) at the 3' end of the single strand is hybridized to the first capture sequence (210) of the nucleic acid (200). The first universal capture binding sequence (250) includes an 'X' indicating the presence of a mismatch between the first universal capture binding sequence (250) and the first capture sequence (210). This allows for cluster amplification to be performed, for example, through bridge amplification, thereby generating clusters.
[0051] delete
[0052] Array
[0053] The array of amplification sites used in the method described herein may exist as one or more substrates. Exemplary types of substrate materials that may be used for the array include glass, modified glass, functionalized glass, inorganic glass, microspheres (e.g., inert and / or magnetic particles), plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, optical fibers or optical fiber bundles, polymers, and multiwell (e.g., microtitration) plates. Exemplary plastics include acrylic, polystyrene, copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, and Teflon TM Exemplary silica-based materials include various forms of silicon and modified silicon.
[0054] In a specific embodiment, the substrate may be in or part of a vessel, such as a well, tube, channel, cuvette, petri plate, bottle, etc. Particularly useful vessels are, for example, flow cells as described in U.S. Patent No. 8,241,573 or in the literature (Bentley et al, Nature 456:53-59 (2008)). An exemplary flow cell is commercially available from Illumina, Inc. (San Diego, California). Other particularly useful vessels are wells in a multi-well plate or microtitration plate.
[0055] In some embodiments, regions of the array may be configured as features on the surface. The features may exist in any various desired formats. For example, the regions may be wells, pits, channels, ridges, raised regions, pegs, posts, etc. As previously mentioned, the regions may contain beads. However, in certain embodiments, the regions do not need to contain beads or particles. Exemplary regions include wells present in substrates used in commercial sequencing platforms sold by 454 LifeSciences (a subsidiary of Roche (Basel, Switzerland)) or Ion Torrent (a subsidiary of Life Technologies (Carlsbad, California)). Other substrates having wells are, for example, U.S. Patent No. 6,266,459; U.S. Patent No. 6,355,431; U.S. Patent No. 6,770,441; U.S. Patent No. 6,859,570; U.S. Patent No. 6,210,891; It includes etched optical fibers and other substrates as described in U.S. Patent No. 6,258,568; U.S. Patent No. 6,274,320; U.S. Patent No. 8,262,900; U.S. Patent No. 7,948,015; U.S. Patent Publication No. 2010 / 0137143; U.S. Patent No. 8,349,167, or PCT Publication WO No. 00 / 63437. In some examples, the substrate is exemplified in these references for applications using beads in wells. The well-containing substrate may be used with or without beads in the method or composition of the present disclosure. In some embodiments, the wells of the substrate may comprise a gel material (with or without beads) as described in U.S. Patent No. 9,512,422.
[0056] The portion of the array may be a metallic feature on a non-metallic surface, e.g., glass, plastic, or other material exemplified above. The metal layer may be deposited on the surface using methods known in the art, such as wet plasma etching, dry plasma etching, atomic layer deposition, ion beam etching, chemical vapor deposition, vacuum sputtering, etc. Any various commercial instruments, such as the FlexAL® OpAL® Ionfab 300plus® or Optofab®3000 system (Oxford Instruments, UK), may be appropriately used. The metal layer may also be deposited by e-beam evaporation or sputtering as described in the literature (Thornton, Ann Rev Mater Sci 7:239-60 (1977)). A metal layer deposition technique, such as that exemplified above, may be combined with photolithography techniques to create a metal region or patch on the surface. Exemplary methods for combining metal layer deposition technology and photolithography technology are provided in U.S. Patent No. 8,778,848 and U.S. Patent No. 8,895,249.
[0057] An array of features may appear as a grid of patches or spots. The features may be positioned in a repeating pattern or in an irregular non-repeating pattern. Particularly useful patterns include hexagonal patterns, straight line patterns, grid patterns, patterns with reflection symmetry, and patterns with rotational symmetry. Asymmetric patterns may also be useful. The pitch may be the same between different pairs of nearest neighbor features, or the pitch may be variable between different pairs of nearest neighbor features. In a specific embodiment, each feature of the array is approximately 100 nm 2 , 250 nm 2 , 500 nm 2 , 1 μm 2 , 2.5 μm 2 , 5 μm 2 , 10 μm 2 , 100 μm2 , or 500 μm 2 It may have an area exceeding. Alternatively, or additionally, each feature of the array is approximately 1 mm 2 , 500 μm 2 , 100 μm 2 , 25 μm 2 , 10 μm 2 , 5 μm 2 , 1 μm 2 , 500 nm 2 , or 100 nm 2 It may have an area less than that. In practice, the area may have a size within a range between an upper limit and a lower limit selected from the examples above.
[0058] For an embodiment comprising an array of features on a surface, the features may be individual, separated by intervening regions. The size of the features and / or the spacing between regions may be variable, so that the array may be high-density, medium-density, or low-density. A high-density array is characterized by having regions separated by less than about 15 μm. A medium-density array has regions separated by about 15 to about 30 μm, while a low-density array has regions separated by more than 30 μm. An array useful in the present disclosure may have regions separated by 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or less than about 0.5 μm.
[0059] In a specific embodiment, the array may comprise a set of beads or other particles. The particles may be suspended in a solution or located on the surface of a substrate. An example of a bead array in a solution is commercialized by Luminex (Austin, Texas). An example of an array having beads located on a surface includes a substrate used in a sequencing platform from 454 LifeSciences (a subsidiary of Roche (Basel, Switzerland)) or Ion Torrent (a subsidiary of Life Technologies (Carlsbad, California)), where the beads are located in wells. Other arrays having beads located on a surface are U.S. Patent No. 6,266,459; U.S. Patent No. 6,355,431; U.S. Patent No. 6,770,441; U.S. Patent No. 6,859,570; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; U.S. Patent No. 6,274,320; U.S. Patent Publication No. 2009 / 0026082 A1; U.S. Patent Publication No. 2009 / 0127589 A1; U.S. Patent Publication No. 2010 / 0137143 A1; U.S. Patent Publication No. 2010 / 0282617 A1 or PCT Publication WO No. 00 / 63437. Some of the aforementioned references describe a method for attaching a target nucleic acid to a bead on or before loading beads on an array substrate. However, it will be understood that the bead may be manufactured to include an amplification primer, and the array may then be loaded using the bead, thereby forming an amplification site for use in the method described herein. As previously stated herein, the substrate may be used without a bead. For example, the amplification primer can be attached directly to the well or to the gel material within the well.Accordingly, the references illustrate materials, compositions, or devices that can be modified for use in the methods and compositions described herein.
[0060] The amplification sites of the array may include a plurality of capture agents capable of binding to the target nucleic acid. In one embodiment, the capture agent comprises a capture nucleic acid. Under typical conditions used to prepare the array for sequencing, the nucleotide sequence of the capture nucleic acid is complementary to the sequence of one or more capture nucleic acids. In contrast, the nucleotide sequence of the capture nucleic acid of the present disclosure is not completely complementary to the sequence of one or more target nucleic acids. The nucleotide sequence of the capture nucleic acid useful for the method presented in the present disclosure is described in detail herein. In some embodiments, the capture nucleic acid may also function as a primer for target nucleic acid amplification (whether or not it contains a universal sequence). In some embodiments, one group of capture nucleic acids comprises a P5 primer or its complement, and a second group of capture nucleic acids comprises a P7 primer or its complement.
[0061] In certain embodiments, a capture agent, e.g., a capture nucleic acid, may be attached to the amplification site. For example, the capture agent may be attached to the surface of a feature portion of the array. Attachment may be via an intermediate structure, e.g., a bead, a particle, or a gel. Attachment of the capture nucleic acid to the array via a gel is described in U.S. Patent No. 8,895,249 and is also exemplified by a flow cell commercially available from Illumina Inc. (San Diego, California) or described in WO No. 2008 / 093098. Exemplary gels that may be used in the methods and apparatus described herein include a colloidal structure, e.g., agarose; a polymer mesh structure, e.g., gelatin; or having a cross-linked polymer structure, e.g., polyacrylamide, SFA (e.g., see U.S. Patent Application Publication No. 2011 / 0059865 A1), or PAZAM (e.g., see U.S. Provisional Patent Application Series No. 61 / 753,833 and U.S. Patent No. 9,012,022), but is not limited thereto. Attachment through beads may be achieved as exemplified in the description and cited literature provided herein.
[0062] In some embodiments, features on the surface of an array substrate are separated by intervening regions on the surface and are not adjacent. Intervening regions having a significantly smaller amount or concentration of capture agent compared to the features of the array are advantageous. Intervening regions without a capture agent are particularly advantageous. For example, a relatively small amount or absence of capture moiety in the intervening regions favors localization of the target nucleic acid and the subsequently generated cluster to the desired features. In certain embodiments, the features may be concave features (e.g., wells) on the surface, and the features may contain a gel material. Gel-containing features may be separated from one another by intervening regions on the surface where the gel is substantially absent or, if present, where the gel cannot substantially support the localization of the nucleic acid. Methods and compositions for manufacturing and using a substrate having gel-containing features, e.g., wells, are described in U.S. Patent No. 9,512,422.
[0063] Target nucleic acid
[0064] The solution of the amplification reagent used in the method described herein comprises a target nucleic acid. The terms “target nucleic acid,” “target fragment,” “target nucleic acid fragment,” “target molecule,” and “target nucleic acid molecule” are used interchangeably to refer to a nucleic acid molecule that is desirable to sequence, for example, on an array. The target nucleic acid may be any nucleic acid of a known or unknown sequence. This may be, for example, a fragment of genomic DNA or cDNA. Sequencing may determine the sequence of the target molecule in whole or in part. The target may be derived from a randomly fragmented primary nucleic acid sample. In one embodiment, the target may be treated as a template suitable for amplification by placing a universal amplification sequence, for example, a sequence present in a universal adapter, at the end of each target fragment.
[0065] The primary nucleic acid sample may be derived from double-stranded DNA (dsDNA) forms from the sample (e.g., genomic DNA fragments, PCR and amplification products, etc.), or may be derived from the sample as single-stranded DNA or RNA and converted into dsDNA forms. For example, mRNA molecules may be replicated into double-stranded cDNA suitable for use in the methods described herein using standard techniques well known in the art. The exact sequence of polynucleotide molecules from the primary nucleic acid sample is generally not important in this disclosure and may or may not be known.
[0066] In one embodiment, the primary polynucleotide molecule from the primary nucleic acid sample is a DNA molecule. More specifically, the primary polynucleotide molecule is a genomic DNA molecule representing the entire genetic complement of an organism and containing non-coding regulatory sequences, such as promoter and enhancer sequences, as well as intron and exon sequences. In one embodiment, a specific subset of polynucleotide sequences or genomic DNA, e.g., a specific chromosome, may be used. More specifically, the sequence of the primary polynucleotide molecule is unknown. More specifically, the primary polynucleotide molecule is a human genomic DNA molecule. The DNA target fragment may be chemically or enzymatically treated before or after any random fragmentation process, and before or after the ligation of the universal adapter sequence.
[0067] The nucleic acid sample may include high molecular weight materials such as genomic DNA (gDNA). The sample may include low molecular weight materials such as nucleic acid molecules obtained from FFPE or stored DNA samples. In another embodiment, the low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample may include cell-free circulating DNA. In some embodiments, the sample may include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser-captured micro-resection, surgical resection, and other clinical or laboratory-acquired samples. In some embodiments, the sample may be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, the sample may include nucleic acid molecules obtained from animals such as human or mammalian sources. In another embodiment, the sample may include nucleic acid molecules obtained from non-mammalian sources such as plants, bacteria, viruses, or fungi. In some embodiments, the source of nucleic acid molecules may be stored or extinct samples or species.
[0068] Additionally, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory involved in a forensic investigation, or forensic samples obtained by a law enforcement agency, one or more military agencies, or any of such personnel. The nucleic acid sample may be unpurified DNA comprising a lysate derived from a buccal swab, paper, fabric, or other substrate that may be impregnated with a purified sample, for example, saliva, blood, or other body fluids. As such, in some embodiments, the nucleic acid sample may contain a small amount or fragmented portion of DNA, such as genomic DNA. In some embodiments, the target sequence may be present in one or more body fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from hair, skin, tissue samples, autopsy, or the remains of a victim. In some embodiments, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence may include nucleic acids obtained from non-human DNA, such as microorganism, plant, or entomological DNA. In some embodiments, the target sequence or amplified target sequence is intended for the purpose of human identification. In some embodiments, the present disclosure generally relates to a method for identifying characteristics of a forensic sample. In some embodiments, the present disclosure generally relates to a method for human identification using one or more target-specific primers disclosed herein, or one or more target-specific primers designed using the primer design criteria outlined herein.In one embodiment, a forensic or human identification sample containing at least one target sequence may be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0069] Non-limiting additional examples of sources for biological samples may include samples obtained from whole organisms and patients. Biological samples may be obtained from any biological fluid or tissue and may be in various forms, including liquid fluids and tissues, solid tissues, and preserved forms such as dried, frozen, and fixed forms. Samples may be biological tissues, cells, or fluids. Such samples include, but are not limited to, sputum, blood, serum, plasma, blood cells (e.g., leukocytes), ascites fluid, urine, saliva, tears, sputum, vaginal fluid (discharge), lavages obtained during medical procedures (e.g., pelvic or other lavages obtained during biopsy, endoscopy, or surgery), tissues, papillary aspirates, core or fine-needle biopsy samples, cell-containing body fluids, free-floating nucleic acids, peritoneal fluid, pleural fluid, or cells from these. Biological samples may also include sections of tissue, such as frozen or fixed sections taken for histological purposes, or microdissected cells or extracellular parts thereof. In some embodiments, the sample may be a blood sample, such as, for example, a whole blood sample. In another example, the sample is an untreated dried blood spot (DBS) sample. In another example, the sample is a formalin-fixed paraffin-embedded (FFPE) sample. In another example, the sample is a saliva sample. In another example, the sample is a dried saliva spot (DSS) sample.
[0070] Exemplary biological samples from which the target nucleic acid may be derived include eukaryotes, e.g., mammals, e.g., rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cattle, cats, dogs, primates, human or non-human primates; plants, e.g., Arabidopsis thaliana, corn, sorghum, oats, wheat, rice, canola, or soybeans; algae, e.g., Chlamydomonas reinhardtii; nematodes, e.g., Caenorhabditis elegans; insects, e.g., Drosophila melanogaster, mosquitoes, fruit flies, bees, or spiders; fish, e.g., zebrafish; reptiles; amphibians, e.g., frogs or Xenopus laevis; It includes dictyostelium discoideum; fungi, e.g., pneumocystis carinii, Takifugu rubripes, yeast, Saccharomyces cerevisiae, or Schizosaccharomyces pombe; or those from plasmodiwn falciparum. The target nucleic acid may also be derived from prokaryotes, e.g., bacteria, Escherichia coli, Staphylococcus, or Mycoplasma pneurnomae; archaea; viruses, e.g., hepatitis C virus or human immunodeficiency virus; or viroids. The target nucleic acid may be derived from a homogeneous culture or population of the aforementioned organisms, or alternatively, from a set of several different organisms in a community or ecosystem, for example.
[0071] Random fragmentation refers to the fragmentation of polynucleotide molecules from primary nucleic acid samples in an unordered manner by enzymatic, chemical, or mechanical means. Such fragmentation methods are known in the art and use standard methods (Sambrook and Russell, Molecular Cloning, A Laboratory Manual, third edition). In one embodiment, fragmentation may be achieved using a process also referred to as tagmentation. Tagmentation involves the addition of a universal adapter using a transposome complex combined with single-step fragmentation and ligation (Gunderson et al., WO 2016 / 130704). For clarity, generating small fragments of a large nucleic acid through specific PCR amplification of these small fragments is not equivalent to fragmenting a large nucleic acid, because the nucleic acid sequence of the large fragment is preserved intact (i.e., not fragmented by PCR amplification). Furthermore, random fragmentation is designed to generate fragments regardless of the sequence identity or position of nucleotides containing or surrounding the rupture. More specifically, random fragmentation is carried out by mechanical means, such as atomization or ultrasonic decomposition, to produce fragments of about 50 base pairs to about 1500 base pairs in length, more specifically 50 to 700 base pairs in length, and even more specifically 50 to 400 base pairs in length. Most specifically, the above method is used to produce small fragments of 50 to 150 base pairs in length.
[0072] Fragmentation of polynucleotide molecules by mechanical means (e.g., atomization, sonication, and / or hydroshear) results in fragments having a heterogeneous mixture of blunt and 3'- and 5'-protruding ends. Therefore, it is desirable to repair the fragment ends using methods or kits known in the art (e.g., Lucigen DNA Terminator End Repair Kit) to produce ends that are optimal for insertion into the blunt site of a cloning vector, for example. In a specific embodiment, the fragment ends of the nucleic acid group are blunt ends. More specifically, the fragment ends are blunt ends and are phosphorylated. A phosphate moiety can be introduced through enzymatic treatment, for example, using a polynucleotide kinase.
[0073] A group of target nucleic acids or amplicons thereof may have an average strand length that is necessary or appropriate for a specific application of the method or composition described herein. For example, the average strand length may be about 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or less than 50 nucleotides. Alternatively or additionally, the average strand length may be about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or more than 100,000 nucleotides. The average strand length for a group of target nucleic acids or their amplicons may be in the range between the maximum and minimum values described above. It will be understood that the amplicons generated at the amplification site (or otherwise manufactured or used herein) may have an average strand length in the range between the upper and lower limits selected from those exemplified above.
[0074] In some cases, a group of target nucleic acids may be produced under conditions having a maximum length for a member of the target nucleic acid, or otherwise configured to have such a maximum length. For example, the maximum length for a member used in one or more steps of the method described herein or present in a particular composition may be 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or less than 50 nucleotides. Alternatively or additionally, a group of target nucleic acids or amplicons thereof may be produced under conditions having a minimum length for the corresponding member, or otherwise configured to have such a minimum length. For example, the minimum length for a member used in one or more steps of the method described herein or present in a particular composition may exceed 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The maximum and minimum strand lengths for the target nucleic acid in the group may be within the range between the maximum and minimum values described above. It will be understood that the amplicon generated at the amplification site (or otherwise manufactured or used herein) may have a maximum and / or minimum strand length within the range between the upper and lower limits exemplified above.
[0075] In a specific embodiment, the target nucleic acid is sized for a region of the amplification site to facilitate exclusion amplification, for example. For example, the region for each region of the array may be larger than the diameter of the excluded volume of the target nucleic acid to achieve exclusion amplification. For example, in an embodiment using an array of features on a surface, the region for each feature may be larger than the diameter of the excluded volume of the target nucleic acid transported to the amplification site. The excluded volume and its diameter for the target nucleic acid may be determined, for example, from the length of the target nucleic acid. Methods for determining the excluded volume and the diameter of the excluded volume of the nucleic acid are, for example, U.S. Patent No. 7,785,790; Rybenkov et al, Proc Natl Acad Sci USA 90: 5307-5311 (1993); Zimmerman et al, J Mol Biol 222:599-620 (1991); Or it is described in Sobel et al, Biopolymers 31:1559-1564 (1991).
[0076] In a specific embodiment, the target fragment sequence is prepared using a single overhang nucleotide by the activity of a specific type of DNA polymerase, e.g., Taq polymerase or Klenow exo minus polymerase (which has non-template-dependent end-transferase activity that adds a single deoxynucleotide, e.g., deoxyadenosine (A), to the 3' end of a DNA molecule, e.g., a PCR product). Using these enzymes, a single nucleotide 'A' can be added to the blunt-terminal 3' end of each strand of the double-stranded target fragment. Thus, 'A' can be added to the 3' end of each strand of the double-stranded target fragment by a reaction using Taq or Klenow exo minus polymerase, whereas the universal adapter polynucleotide construct may be a T-structure having a compatible 'T' overhang present on the 3' end of each region of the double-stranded nucleic acid of the universal adapter. These end deformations also prevent self-ligation of both the vector and the target, so that there is a bias toward the formation of combined ligated adapter-target-adapter molecules.
[0077] In some cases, target nucleic acids derived from such sources may be amplified before being used in the methods or compositions described herein. Any various known amplification techniques, including but not limited to polymerase chain reaction (PCR), rotational ring amplification (RCA), multiple displacement amplification (MDA), or random prime amplification (RPA), may be used. It will be understood that amplification of the target nucleic acid before use in the methods or compositions described herein is optional. As such, the target nucleic acid will not be amplified before being used in some embodiments of the methods and compositions described herein. The target nucleic acid may be optionally derived from a synthetic library. The synthetic nucleic acid may have a natural DNA or RNA composition, or may be an analogue thereof.
[0078] Universal adapter
[0079] The target nucleic acid used in the method or composition described herein comprises a universal adapter attached to each end. A method of attaching a universal adapter to each end of the target nucleic acid used in the method described herein is known to those skilled in the art. Attachment may be achieved through a standard library preparation technique using ligation (Chesney et al. US Pat. Pub. No. 2018 / 0305753 A1) or through tagging using a transposase complex (Gunderson et al., WO 2016 / 130704).
[0080] In one embodiment, a double-stranded target nucleic acid from a sample, e.g., a fragmented sample, is first processed by ligating the same universal adapter molecule ('mismatched adapter'; its general characteristics are defined below and further described in Gormley et al., US 7,741,463, and Bignell et al., US 8,053,192) to the 5' and 3' ends of the double-stranded target nucleic acid (which may have known, partially known, or unknown sequences). In one embodiment, the universal adapter includes a universal capture binding sequence necessary to immobilize the target nucleic acid on an array for subsequent sequencing. In another embodiment, a PCR step is used to further modify the universal adapter present at each end of the target nucleic acid prior to immobilization and sequencing. For example, an initial primer extension reaction is performed using a universal primer binding site that adds a universal capture binding sequence and forms an extension product complementary to both strands of each individual target nucleic acid. The generated primer extension product and optionally amplified replicas thereof collectively provide a modified target nucleic acid library that can be fixed and then sequenced. The term library refers to a set of target nucleic acids containing known common sequences at the 3' and 5' ends, and may also be referred to as a 3' and 5' modified library. The 3' end and optionally the 5' end of the universal adapter attached to the target nucleic acid may comprise a homogeneous or heterogeneous population of the universal capture binding sequences described herein.
[0081] The general adapter used in the method of the present disclosure is called a ‘mismatched’ adapter because, as described in detail herein, the adapter includes a sequence mismatch region, that is, it is not formed by the annealing of a completely complementary polynucleotide strand.
[0082] A mismatch adapter for use herein is formed by the annealing of two partially complementary polynucleotide strands such that when the two partially complementary polynucleotide strands are annealed, at least one double-stranded region, also referred to as a double-stranded nucleic acid region, and at least one mismatched single-stranded region, also referred to as a single-stranded non-complementary nucleic acid strand region, are provided.
[0083] The 'double-stranded region' of a universal adapter is a short double-stranded region containing typically five or more consecutive base pairs, formed by the annealing of two partially complementary polynucleotide strands. This term refers to a double-stranded region of nucleic acid in which two strands have been annealed and does not imply any specific structural conformation. As used herein, the term "double strand," when used in relation to a nucleic acid molecule, means that substantially all nucleotides of the nucleic acid molecule are hydrogen-bonded to complementary nucleotides. A partially double-stranded nucleic acid may have at least 10%, 25%, 50%, 60%, 70%, 80%, 90%, or 95% of nucleotides hydrogen-bonded to complementary nucleotides.
[0084] Generally, it is advantageous for the double-stranded region to be as short as possible without loss of function. In this context, 'function' refers to the ability of the double-stranded region to form a stable duplex under standard reaction conditions for enzyme-catalyzed nucleic acid ligation reactions, which is well known to those skilled in the art (e.g., incubation in a ligation buffer suitable for the enzyme at a temperature ranging from 4°C to 25°C), whereby the two strands forming the universal adapter remain in a partially annealed state during the ligation of the universal adapter to the target molecule. The double-stranded region does not need to be stable under conditions typically used in primer extension or the annealing step of a PCR reaction.
[0085] The double-stranded region of a universal adapter is typically identical across all universal adapters used for ligation. Since the universal adapter is ligated to both ends of each target molecule, the complementary sequence derived from the universal adapter's double-stranded region is located laterally within the modified target nucleic acid. The longer the double-stranded region, and consequently the longer the complementary sequence derived therefrom within the modified target nucleic acid construct, the greater the likelihood that the modified target nucleic acid construct will fold within this internal self-complementary region and form a base pair with itself under the annealing conditions used for primer extension and / or PCR. Therefore, to reduce this effect, it is generally desirable for the double-stranded region to have a length of 20 or fewer, 15 or fewer, or 10 or fewer base pairs. The stability of the double-stranded region can be increased by including non-natural nucleotides that exhibit stronger base pairing than standard Watson-Crick base pairing, which may consequently potentially reduce its length.
[0086] In one embodiment, the two strands of the universal adapter are 100% complementary in the double-strand region. It will be understood that if the two strands can form a stable duplex under standard ligation conditions, one or more nucleotide mismatches within the double-strand region may be allowed.
[0087] A universal adapter for use herein generally comprises a 'ligable' end of the adapter, for example, a double-stranded region that forms a end connected to a double-stranded target nucleic acid in a ligation reaction. The ligable end of the universal adapter may be blunt, or in other embodiments, a short 5' or 3' overhang of one or more nucleotides may be present to facilitate / promote ligation. The 5' terminal nucleotide at the ligable end of the universal adapter is typically phosphorylated to enable phosphodiester bonding to the 3' hydroxyl group on the target polynucleotide.
[0088] The term 'mismatched region' refers to a region of a universal adapter that is a single-stranded non-complementary nucleic acid strand, wherein the sequences of the two polynucleotide strands forming the universal adapter exhibit a degree of non-complementarity such that the two strands cannot be fully annealed to each other under standard annealing conditions for primer extension or PCR reactions. The mismatched region(s) may exhibit some degree of annealing under standard reaction conditions for enzyme-catalyzed ligation reactions, provided that the two strands revert to a single-stranded form under annealing conditions during amplification reactions.
[0089] It should be understood that a ‘mismatched region’ is provided by different portions of the same two polynucleotide strands forming a double-stranded region(s). The mismatch of the adapter structure may take the form where one strand is longer than the other, where a single-stranded region exists on one of the strands, or where a sequence selected so that the two strands do not hybridize and thus form a single-stranded region on both strands exists. The mismatch may also take the form of a ‘bubble,’ where the two ends of the universal adapter structure(s) may hybridize with each other to form a dimer, but the central region does not. The portion of the strand(s) forming the mismatched region is not annealed under conditions where different portions of the same two strands are annealed to form one or more double-stranded regions. To avoid doubt, it should be understood that a single strand or single base overhang at the 3' end of the polynucleotide dimer subsequently ligated to a target sequence does not constitute a ‘mismatched region’ in the context of this disclosure.
[0090] The lower limit for the length of the mismatch region is typically determined by the need to provide a suitable sequence for function, for example, i) primer extension, binding of a primer for PCR and / or sequencing (e.g., binding a primer to a universal primer binding site), or ii) binding of a universal capture binding sequence to a capture sequence to immobilize a modified target nucleic acid on the surface. Theoretically, there is no upper limit for the length of the mismatch region, but exceptionally, it is generally advantageous to minimize the total length of the universal adapter to facilitate, for example, separating the unbound universal adapter from the modified target nucleic acid construct following a ligation step. Therefore, it is generally desirable for the mismatch region to be a continuous nucleotide length of less than 50, less than 40, less than 30, or less than 25.
[0091] A region of a single-stranded non-complementary nucleic acid strand comprises at least one universal capture binding sequence at the 3' end (see FIG. lb, universal capture binding sequence 150). The 3' end of the universal adapter comprises a first universal capture binding sequence that hybridizes to a first capture sequence present in the capture nucleic acid. For example, as illustrated in FIG. 2, the nucleic acid (200) of the first capture nucleic acid group comprises the first capture sequence (210). One strand of the modified target nucleic acid (230) comprising the first universal capture binding sequence (250) at the 3' end of the single strand is illustrated as being hybridized to the first capture sequence (210). It is the interaction between the first universal capture binding sequence (250) and the first capture sequence (210) that is modified to reduce affinity and encode a delay in movement to the target nucleic acid seeded in the well. Standard ExAmp methods use a capture sequence and a universal capture binding sequence that are fully complementary over the entire length of the capture sequence. The ExAmp methods described herein use a universal capture binding sequence that includes one or more mismatches, has a reduced length, or a combination thereof. The result of the mismatch(s) and / or reduced length is a reduced affinity between the two sequences compared to the affinity between the two fully complementary full-length sequences. Reduced affinity causes a reduction in amplification efficiency, where the resulting amplification efficiency is generally a function of the number of differences between the universal capture binding sequence and the capture sequence.
[0092] Optionally, the 5' end of the universal adapter includes a second universal capture binding sequence attached to each end of the target nucleic acid, wherein the second universal capture binding sequence hybridizes to the second capture sequence present in the capture nucleic acid. For example, as illustrated in FIG. 1b, there is a universal capture binding sequence (180). Thus, unless otherwise noted, the following discussion on how universal capture binding sequences are adjusted to reduce affinity applies to both the 3' and 5' universal capture binding sequences.
[0093] The 3' end of the capture sequence functions as a starting point for DNA synthesis by DNA polymerase in the method described herein. Those skilled in the art will recognize that the nucleotide at the 3' end of the capture sequence and the corresponding nucleotide in the universal capture binding sequence must be complementary to preserve the ability of the DNA polymerase to initiate DNA synthesis.
[0094] The universal capture binding sequence may include one or more nucleotides that are non-complementary to the capture sequence. In one embodiment, the universal capture binding sequence may include one to five mismatch nucleotides (also referred to as non-complementary nucleotides) relative to the capture sequence used in the amplification reaction described herein, for example, at least one, at least two, at least three, at least four, or five mismatch nucleotides. The mismatch nucleotides may be wobble mismatches or true mismatches.
[0095] A wobble mismatch refers to a location in a group of universal capture binding sequences where all four nucleotides appear. For example, if N is a wobble nucleotide of ACTNGC, the group of universal capture binding sequences includes ACTTGC, ACTAGC, ACTCGC, and ACTGGC, and 25% of the universal capture binding sequences in the group are complementary to the corresponding nucleotides of the capture sequences. In one embodiment, the universal capture binding sequence may include one to five wobble nucleotides, for example, at least one, at least two, at least three, at least four, or five wobble nucleotides, relative to the capture sequence used in the amplification reaction described herein. In one embodiment, the wobble nucleotides may be located anywhere throughout the universal capture binding sequence.
[0096] A true mismatch refers to a position where only three of the four nucleotides appear at a specific location in a group of universal capture binding sequences. For example, if G is the location of the true true mismatch nucleotide in ACTTGC, the group of universal capture binding sequences includes ACTTCC, ACTTTC, and ACTTAC, and none of the universal capture binding sequences in the group will be non-complementary to the corresponding nucleotide of the capture sequence, in this case, C. In one embodiment, the universal capture binding sequence may include one to five mismatch nucleotides, e.g., at least one, at least two, at least three, at least four, or five wobble nucleotides, relative to the capture sequence used in the amplification reaction described herein. In one embodiment, the wobble nucleotides may be located anywhere throughout the universal capture binding sequence.
[0097] Those skilled in the art will recognize that the use of wobble mismatch or true mismatch provides greater control over altering the affinity of universal capture binding sequences. By using a universal capture binding sequence having only a single wobble nucleotide, it is expected that 25% of universal capture binding sequences that are complementary at that position will have greater affinity and higher amplification efficiency than the remaining 75%. By using a universal capture binding sequence having only one true mismatch nucleotide, all universal capture binding sequences will be non-complementary at that position, resulting in reduced affinity and reduced amplification efficiency.
[0098] In another embodiment, the universal capture binding sequence has a shortened length resulting in an affinity smaller than the affinity between the full-length universal capture binding sequence and the capture sequence. The capture sequence useful for the standard amplification method described herein may be longer or shorter as needed, but typically has a length of about 20 to about 30 nucleotides. The universal capture binding sequence useful for the method described herein may have a length that is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides shorter than the capture sequence used in the amplification reaction described herein. In one embodiment, the length of the universal capture binding sequence is reduced by removing nucleotides from the 3' end of the first universal capture binding sequence and / or the 5' end of the second universal capture binding sequence.
[0099] The amplification reaction described herein may use a heterogeneous population of universal capture binding sequences (e.g., a plurality of different target nucleic acids may comprise a heterogeneous population of universal capture binding sequences present at the 3' end and optionally at the 5' end). In one embodiment, the heterogeneous population comprises individual universal capture binding sequences having mismatched nucleotides. In one embodiment, the universal capture binding sequence has one, two, three, four, or five mismatched nucleotides. The mismatched nucleotides may be wobble mismatches, true mismatches, or a combination thereof.
[0100] In one embodiment, the heterogeneous group comprises individual universal capture binding sequences having shortened lengths. In one embodiment, the heterogeneous group comprises individual universal capture binding sequences having lengths that are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotide lengths shorter than the capture sequence used in the amplification reaction described herein.
[0101] In one embodiment, the heterogeneous group comprises individual universal capture binding sequences having a combination of one or more mismatched nucleotides and shortened lengths. The number of missing nucleotides and the number of mismatched nucleotides in the universal capture binding sequences may exist in any combination, for example, the number of mismatched nucleotides and the number of missing nucleotides are independent.
[0102] The heterogeneous population may also include individual target nucleic acids having universal capture binding sequences that are 100% complementary to the capture sequence at the 3' end and optionally at the 5' end. The molar ratios of different universal capture binding sequences in the heterogeneous population may be the same or may change. In embodiments where the molar ratios are not the same, a high molar ratio of universal capture binding sequences having high amplification efficiency is preferred. Accordingly, in such embodiments where the heterogeneous population includes universal capture binding sequences that are 100% complementary to the capture sequence, the universal capture binding sequence having 100% complementarity may be present in a larger proportion than any other member of the heterogeneous population.
[0103] The region of the single-stranded non-complementary nucleic acid strand typically also includes at least one universal primer binding site. The universal primer binding site is a universal sequence that can be used for the amplification and / or sequencing of the target nucleic acid ligated to the universal adapter.
[0104] A region of a single-stranded non-complementary nucleic acid strand may also include at least one index. The index can be used as a characteristic marker of a source of a specific target nucleic acid on the array. Generally, the index is a synthetic sequence of nucleotides that is part of a universal adapter added to the target nucleic acid as part of a library preparation step. Accordingly, the index is a nucleic acid sequence attached to each target molecule of a specific sample, and its presence is used to indicate or identify the sample or source from which the target molecule was isolated.
[0105] Preferably, the index may be up to 20 nucleotide lengths, more preferably 1 to 10 nucleotides, and most preferably 4 to 8 nucleotides. For example, a 4-nucleotide index is 256 (4 on the same array 4 While it provides the possibility to multiplex ) samples, the 6-base index allows 4,096(4) samples on the same array. 6 Enables processing ) samples.
[0106] In one embodiment, a universal capture binding sequence is part of the universal adapter when ligated to a double-stranded target fragment, and in another embodiment, a universal primer extension binding site is added to the universal adapter after the universal adapter is ligated to a double-stranded target fragment. The addition can be achieved using routine methods including PCR-based methods.
[0107] The exact nucleotide sequence of the universal adapter is generally not important in the present invention, and, for example, to provide a universal capture binding sequence and binding site for a specific set of universal amplification primers and / or sequencing primers, the user may select desired sequence elements so that they are ultimately included in the common sequence of a plurality of different modified target nucleic acids. For example, additional sequence elements may be included to provide a binding site for a sequencing primer to be ultimately used for sequencing a target nucleic acid in a library, or, for example, a product derived from the amplification of a target nucleic acid in a library on a solid support.
[0108] The exact nucleotide sequence of the universal adapter is generally not limited to that disclosed herein, but the sequence of the individual strands of the mismatched region must be such that the individual strands do not exhibit any internal self-complementarity that could lead to self-annealing, the formation of a hairpin structure, etc. under standard annealing conditions. Self-annealing of the strands in the mismatched region must be avoided, as it can prevent or reduce the specific binding of the amplification primer to these strands.
[0109] The mismatch adapter may comprise a mixture of natural and non-natural nucleotides (e.g., one or more ribonucleotides) connected by a mixture of phosphodiester and non-phosphodiester backbone linkages, preferably formed from two strands of DNA.
[0110] Ligation and amplification
[0111] Ligation methods are known in the art and standard methods are used. These methods use a ligase enzyme, such as DNA ligase, to influence or catalyze the terminal ligation of the two polynucleotide strands of the universal adapter and the double-stranded target nucleic acid in this case so that a covalent bond is formed. The universal adapter may include a 5'-phosphate moiety to facilitate ligation to the 3'-OH present on the target fragment. The double-stranded target nucleic acid includes a 5'-phosphate moiety that is retained from a shearing process or added using an enzymatic treatment step, is terminally repaired, and is selectively extended by an overhanging base or bases to provide a 3'-OH suitable for ligation. In this context, ligation refers to the covalent linkage of previously uncovalently linked polynucleotide strands. In certain embodiments of the present disclosure, such linkage occurs by the formation of a phosphodiester linkage between the two polynucleotide strands, but other means of covalent linkage (e.g., non-phosphodiester backbone linkage) may be used.
[0112] As described herein, in one embodiment, the universal adapter used for ligation is complete and includes a universal capture binding sequence and other universal sequences, e.g., a universal primer binding site and an index sequence. The plurality of target nucleic acids produced can be used to prepare immobilized samples for sequencing analysis.
[0113] Additionally, as discussed herein, in one embodiment, the universal adapter used for ligation comprises a universal primer binding site and an index sequence, and does not comprise a universal capture binding sequence. A plurality of modified target nucleic acids resulting from this may be further modified to comprise a specific sequence, such as a universal capture binding sequence. A method for adding a specific sequence, such as a universal capture binding sequence, to a universal primer ligated to a double-stranded target fragment comprises PCR-based methods and is known in the art, such as those described, for example, in Bignell et al. (US 8,053,192) and Gunderson et al. (WO2016 / 130704).
[0114] In an embodiment in which the universal adapter is modified, an amplification reaction is prepared. The contents of the amplification reaction are known to those skilled in the art and include a suitable substrate (e.g., dNTP), an enzyme (e.g., DNA polymerase), and a buffer component required for the amplification reaction. Generally, the amplification reaction requires at least two amplification primers, often referred to as 'forward' and 'reverse' primers (primer oligonucleotides), which can be specifically annealed to a portion of the polynucleotide sequence to be amplified, e.g., the target nucleic acid, under the conditions faced during the primer annealing step of each cycle of the amplification reaction. It will be recognized that if the primers contain any nucleotide sequence that is not annealed to the modified target nucleic acid in the first amplification cycle, this sequence may be replicated into the amplification product. For example, if a primer having a universal capture binding sequence, i.e., a sequence that is not annealed to the modified target nucleic acid, is used, the universal capture binding sequence will be incorporated into the amplicon produced.
[0115] Amplification primers are generally single-stranded polynucleotide structures. They may also include a mixture of natural and non-natural bases and natural and non-natural backbone linkages, provided that any non-natural modification defined as the ability to anneale the template polynucleotide strand during the amplification reaction conditions and function as a starting point for the synthesis of a new polynucleotide strand complementary to the template strand does not exclude function as a primer. Primers may also further include non-nucleotide chemical modifications, for example, phosphorothioates to increase exonuclease resistance, provided that the modification does not interfere with primer function.
[0116] Preparation of immobilized samples for sequence analysis
[0117] The method of the present disclosure may include the step of reacting an amplification reagent (an array of amplification sites and a plurality of different modified target nucleic acids) to generate a plurality of amplification sites, each comprising a clonal population of amplicons from individual target nucleic acids seeded at the sites. In a standard reaction, exclusion amplification occurs due to a relatively slow rate of target nucleic acid seeding (e.g., a relatively slow diffusion or transport) versus a relatively fast rate at which amplification occurs to fill the sites with a replication of the nucleic acid seed. In the method described herein, exclusion amplification may occur due to a kinetic delay in the formation of a first replication of the target nucleic acid seeded at the sites versus a relatively fast rate at which subsequent replications are made to fill the sites. For example, individual sites may be seeded with several different target nucleic acids each having different universal capture binding sequences (e.g., a plurality of different modified target nucleic acids include heterogeneous populations of universal capture binding sequences). However, the formation of a first replication for any given target nucleic acid is expected to depend on the amplification efficiency of its universal capture binding sequence, where the average rate of the first replication is relatively slow compared to the rate at which subsequent replications are generated. In this case, although individual sites may be seeded with multiple different target nucleic acids, only one begins amplification first, and exclusion amplification typically ensures that only that target nucleic acid can fill the amplification site. More specifically, once the first target nucleic acid begins amplification, the site is rapidly filled with its replication, preventing replication of the second target nucleic acid from being produced at that site.
[0118] In some embodiments, apparent cloning performance can be achieved even if the amplification site is not filled to capacity before the second target nucleic acid begins amplification at the site. Under some conditions, the amplification of the first target nucleic acid may proceed to a point where a sufficient number of copies are produced to effectively outpace or overwhelm the production of copies from the second target nucleic acid transported to the site. For example, in an embodiment using a bridge amplification process for a circular feature with a diameter of less than 500 nm, it was determined that after 14 cycles of exponential amplification for the first target nucleic acid, contamination from the second target nucleic acid at the same site produced an insufficient number of contaminating amplicons that adversely affected the synthetic sequencing analysis on the Illumina sequencing platform.
[0119] The amplification sites of the array do not need to be completely clonal in all embodiments. Rather, in some applications, individual amplification sites may be predominantly filled with amplicons from a first target nucleic acid and may also have low levels of contaminating amplicons from a second target nucleic acid. The array may have one or more amplification sites having low levels of contaminating amplicons, provided that the level of contamination does not have an unacceptable effect on the subsequent use of the array. For example, when the array is used in a detection application, an acceptable level of contamination is one that does not affect the signal-to-noise or resolution of the detection technology in an unacceptable manner. Accordingly, apparent clonality will generally be relevant to the specific use or application of the array produced by the method presented herein. Exemplary levels of contamination acceptable in individual amplification sites for a specific application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% of contaminating amplicons. The array may include one or more amplification sites having such exemplary levels of contaminated amplicons. For example, some contaminated amplicons may be present in up to 5%, 10%, 25%, 50%, 75%, or 100% of the amplification sites of the array.
[0120] Although the use of differentially active primers to induce the formation of a first amplicon and a subsequent amplicon at different rates has been illustrated above for an embodiment in which the target nucleic acid is present at the amplification site prior to amplification, the method may also be performed under conditions in which the target nucleic acid is transported to the amplification site (e.g., via diffusion) as amplification occurs. Thus, exclusion amplification may utilize both a relatively slow transport rate and the generation of a first amplicon that is relatively slow compared to the formation of a subsequent amplicon. Accordingly, the amplification reaction presented herein may be performed such that the target nucleic acid is transported from solution to the amplification site simultaneously with (i) the generation of a first amplicon and (ii) the generation of a subsequent amplicon at another site of the array. In certain embodiments, the average rate at which a subsequent amplicon is generated at the amplification site may exceed the average rate at which the target nucleic acid is transported from solution to the amplification site. In some cases, a sufficient number of amplicons may be generated from a single target nucleic acid at individual amplification sites to fill the capacity of each amplification site. The rate at which amplicons are generated to fill the capacity of each amplification site may exceed, for example, the rate at which individual target nucleic acids are transported from solution to the amplification site.
[0121] The amplification reagent used in the method presented herein is preferably capable of rapidly inducing replication of the target nucleic acid at the amplification site. Typically, the amplification reagent used in the method of the present disclosure comprises a polymerase and a nucleotide triphosphate (NTP). Any various polymerase known in the art may be used, but in some embodiments, it may be preferable to use an exonuclease-negative polymerase. The NTP may be a deoxyribonucleotide triphosphate (dNTP) in embodiments where DNA replication is produced. Typically, four natural species, dATP, dTTP, dGTP, and dCTP, are present in the DNA amplification reagent, but analogs may be used if necessary. The NTP may be a ribonucleotide triphosphate (rNTP) in embodiments where RNA replication is produced. Typically, four natural species, rATP, rUTP, rGTP, and rCTP, are present in the RNA amplification reagent, but analogs may be used if necessary.
[0122] Amplification reagents may contain additional components that facilitate amplicon formation and, in some cases, increase the rate of amplicon formation. An example is a recombinase loading protein. The recombinase can facilitate amplicon formation by allowing repeated intrusion / extension. More specifically, the recombinase can facilitate the intrusion of the target nucleic acid by the polymerase and the extension of the primer by the polymerase by using the target nucleic acid as a template for amplicon formation. This process can be repeated as a chain reaction in which the amplicon generated from the intrusion / extension of each round functions as a template in subsequent rounds. Since denaturation cycles (e.g., through heating or chemical denaturation) are not required, the process can occur faster than standard PCR. As such, recombinase-assisted amplification can be performed isothermally. Generally, it is desirable to include ATP or other nucleotides (or, in some cases, their non-hydrolyzable analogs) in the recombinase-assisted amplification reagent to facilitate amplification. Mixtures of recombinase, single-stranded binding (SSB) proteins, and auxiliary proteins are particularly useful. Exemplary formulations for recombinase-driven amplification include those marketed as the TwistAmp kit by TwistDx (Cambridge, UK). Useful components and reaction conditions for recombinase-driven amplification reagents are described in U.S. Patent No. 5,223,414 and U.S. Patent No. 7,399,590.
[0123] Another example of a component that can be included in amplification reagents to promote amplicon formation and, in some cases, increase the rate of amplicon formation is helicase. Helicase can facilitate amplicon formation by allowing a chain reaction of amplicon formation. Since denaturation cycles (e.g., through heating or chemical denaturation) are not required, the process can occur faster than standard PCR. As such, helicase-assisted amplification can be performed isothermally. Mixtures of helicase and single-stranded binding (SSB) proteins are particularly useful because the SSB can further facilitate amplification. Exemplary formulations for helicase-assisted amplification include those marketed as the IsoAmp kit by Biohelix (Beverly, Massachusetts). Additionally, examples of useful formulations containing helicase proteins are described in U.S. Patent No. 7,399,590 and U.S. Patent No. 7,829,284.
[0124] Another example of a component that can be included in amplification reagents to facilitate amplicon formation and, in some cases, increase the rate of amplicon formation is origin-binding protein.
[0125] Molecular crowding reagents present in solution can be used to aid exclusion amplification. Examples of useful molecular crowding reagents include, but are not limited to, polyethylene glycol (PEG), Ficoll®, dextran, or polyvinyl alcohol. Exemplary molecular crowding reagents and formulations are described in U.S. Patent No. 7,399,590.
[0126] The rate at which an amplification reaction occurs can be increased by increasing the concentration or amount of one or more active components of the amplification reaction. For example, the amplification rate can be increased by increasing the amount or concentration of a polymerase, nucleotide triphosphate, primer, recombinase, helicase, or SSB. In some cases, one or more active components of the amplification reaction whose amount or concentration has been increased (or otherwise manipulated in the methods described herein) are non-nucleic components of the amplification reaction.
[0127] The amplification rate can also be increased in the method presented herein by adjusting the temperature. For example, the amplification rate at one or more amplification sites can be increased by increasing the temperature of the site(s) up to the maximum temperature at which the reaction rate decreases due to denaturation or other adverse effects. The optimal or desired temperature can be determined from the known characteristics of the amplification component in use or empirically for a given amplification reaction mixture. Such adjustment is based on the primer melting temperature (T m It can be based on prior predictions of ) or empirically.
[0128] The rate at which the amplification reaction occurs can be increased by increasing the activity of one or more amplification reagents. For example, a cofactor that increases the elongation rate of the polymerase can be added to the reaction in which the polymerase is being used. In some embodiments, a metal cofactor such as magnesium, zinc, or manganese can be added to the polymerase reaction, or betaine can be added.
[0129] In some embodiments of the method presented herein, it is preferable to use a double-stranded target nucleic acid population. It has been observed that the formation of amplicons in an array of sites under exclusion amplification conditions is efficient for double-stranded target nucleic acids. For example, multiple amplification sites having a cloned population of amplicons can be generated more efficiently from double-stranded target nucleic acids (compared to single-stranded target nucleic acids at the same concentration) in the presence of a recombinase and a single-stranded binding protein. Nevertheless, it will be understood that single-stranded target nucleic acids may be used in some embodiments of the method presented herein.
[0130] The method described herein may use any of the various amplification techniques. Exemplary techniques that may be used include, but are not limited to, polymerase chain reaction (PCR), rotational amplification (RCA), multiple displacement amplification (MDA), or random prime amplification (RPA). In some embodiments, amplification may be performed in solution, for example, where the amplification site may contain an amplicon in a volume having a desired capacity. Preferably, the amplification technique used under exclusion amplification conditions in the method of the present disclosure is performed on a solid. For example, one or more primers used for amplification may be attached to a solid at the amplification site. In PCR embodiments, one or both of the primers used for amplification may be attached to a solid. Because the double-stranded amplicon forms a bridge-like structure between two surface-attached primers on the sides of the replicated template sequence, a format using two types of surface-attached primers is often called bridge amplification. Exemplary reagents and conditions that may be used for bridge amplification are described, for example, in U.S. Patent No. 5,641,658, U.S. Patent Publication No. 2002 / 0055100, U.S. Patent No. 7,115,400, U.S. Patent Publication No. 2004 / 0096853, U.S. Patent Publication No. 2004 / 0002090, U.S. Patent Publication No. 2007 / 0128624, and U.S. Patent Publication No. 2008 / 0009420. Solid-phase PCR amplification may also be performed with a second primer in solution and one of the amplification primers attached to a solid support. An exemplary format using a combination of a surface-attached primer and a soluble primer is, for example, Dressman et al., Proc. Natl. Acad. Sci. It is an emulsion PCR as described in USA 100:8817-8822 (2003), WO 05 / 010145 or U.S. Patent Publication No. 2005 / 0130173 or No. 2005 / 0064460.It will be understood that emulsion PCR is illustrative of the format, and that the use of an emulsion for the method presented herein is optional and, in practice, an emulsion is not used in many embodiments. The described PCR technique may be modified for non-cyclic amplification (e.g., isothermal amplification) by using components exemplified elsewhere in this invention to accelerate or increase the amplification rate. Accordingly, the described PCR technique may be used under exclusion amplification conditions.
[0131] RCA techniques may be modified for use in the methods of the present disclosure. Exemplary components that may be used in the RCA reaction and the principle by which RCA generates an amplicon are described, for example, in Lizardi et al., Nat. Genet. 19:225-232 (1998), and U.S. Patent Publication 2007 / 0099208 A1. Primers used in the RCA may be in solution or attached to the surface of a solid support at the amplification site. The RCA techniques exemplified in the above references may be modified according to the teachings of this specification to increase the amplification rate, for example, to suit a specific application. Thus, RCA techniques may be used under exclusion amplification conditions.
[0132] MDA technology may be modified for use in the method of the present disclosure. Some basic principles and useful conditions for MDA are described, for example, in Dean et al., Proc Natl. Acad. Sci. USA 99:5261-66 (2002); Lage et al., Genome Research 13:294-307 (2003); Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995; Walker et al., Nucl. Acids Res. 20:1691-96 (1992); U.S. Patent No. 5,455,166; U.S. Patent No. 5,130,238; and U.S. Patent No. 6,214,587. Primers used in MDA may be in solution or attached to the surface of a solid support at the amplification site. The MDA techniques exemplified in the above references may be modified, for example, in accordance with the teachings of this specification to increase the amplification speed to suit a specific application. Accordingly, the MDA techniques may be used under exclusion amplification conditions.
[0133] In a specific embodiment, an array may be formed under exclusion amplification conditions using a combination of the described amplification techniques. For example, RCA and MDA may be used in a combination where RCA is used to generate a chain of amplicons in solution (e.g., using a solution-phase primer). Subsequently, the amplicons may be used as a template for the MDA using a primer attached to the surface of a solid support at the amplification site. In this example, the amplicons generated after the RCA and MDA combination step are attached to the surface of the amplification site.
[0134] As illustrated in the various embodiments above, the method of the present disclosure does not require the use of cyclic amplification techniques. For example, amplification of the target nucleic acid may be performed at an amplification site without denaturation cycles. An exemplary denaturation cycle includes the introduction of a chemical denaturant into the amplification reaction and / or an increase in the temperature of the amplification reaction. Accordingly, amplification of the target nucleic acid does not require the step of replacing the amplification solution with a chemical reagent that denatures the target nucleic acid and the amplicon. Similarly, amplification of the target nucleic acid does not require heating the solution to a temperature that denatures the target nucleic acid and the amplicon. Accordingly, amplification of the target nucleic acid at the amplification site may be performed isothermally during the duration of the method presented herein. In practice, the amplification method presented herein may occur without one or more cyclic operations performed for some amplification techniques under standard conditions. Additionally, in some standard solid-phase amplification techniques, washing is performed after the target nucleic acid is loaded onto a substrate and before amplification begins. However, in embodiments of the present method, it is not necessary to perform a washing step between the transport of the target nucleic acid to the reaction site and the amplification of the target nucleic acid at the amplification site. Instead, to provide exclusion amplification, transport and amplification are made to occur simultaneously (e.g., through diffusion).
[0135] In some embodiments, it may be desirable to repeat the amplification cycle occurring under exclusion amplification conditions. Thus, while replication of the target nucleic acid may be produced at individual amplification sites without cycle manipulation, the array of amplification sites may be processed periodically to increase the number of sites containing amplicons after each cycle. In certain embodiments, the amplification conditions may be modified per cycle. For example, one or more of the conditions described above may be adjusted between cycles to change the transport rate or the amplification rate. Thus, the transport rate may be increased per cycle, the transport rate may be decreased per cycle, the amplification rate may be increased per cycle, or the amplification rate may be decreased per cycle.
[0136] composition
[0137] Different compositions may be produced during or after the amplification clustering method described herein. In one embodiment, the composition comprises an array of amplification sites. Each site comprises a first capture nucleic acid and a second capture nucleic acid, each comprising a first capture sequence and a second capture sequence, respectively, wherein the first and second capture nucleic acids are bound to the surface of the site. Different sites of the array comprise target nucleic acids hybridized to the first capture sequence of the first capture nucleic acid. Each of the target nucleic acids at different sites comprises a universal capture binding sequence hybridized to the capture sequence at the 3' end. There exists a universal capture binding sequence having less affinity for the capture sequence than a universal capture binding sequence having 100% complementarity with the first capture sequence. In one embodiment, different universal capture binding sequences exist at each site, for example, a first heterogeneous group of universal capture binding sequences exists. The first heterogeneous group may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 different universal capture combination sequences.
[0138] In one embodiment, the first universal capture binding sequence comprises one to five nucleotides that are non-complementary to the first capture sequence. The composition may comprise some target nucleic acids having a universal capture binding sequence that is 100% complementary to the first capture sequence. In one embodiment, members of the first heterogeneous group having 100% complementary to the first capture sequence exist in greater numbers than other members of the first heterogeneous group. The first heterogeneous group may also comprise individual first universal capture binding sequences having a length less than the length of the first capture sequence, such as a length that is one to 12 nucleotides shorter than the length of the first capture sequence. Individual members of the first heterogeneous group may have a length less than the length of the first capture sequence and may comprise one to five nucleotides that are non-complementary to the sequence of the first capture sequence or may comprise 100% complementary to the sequence of the first capture sequence.
[0139] The 5' end may optionally include a second universal capture binding sequence having a complement having a lower affinity for the second capture sequence than a second universal capture binding sequence having a complement having 100% complementation for the second capture sequence. In one embodiment, the complement of the second universal capture binding sequence comprises one to five nucleotides that are non-complementary to the second capture sequence. The composition may include some target nucleic acid having the second universal capture binding sequence together with a complement having 100% complementation for the second capture sequence. In one embodiment, different second universal capture binding sequences are present at each site, for example, a second heterogeneous group of second universal capture binding sequences is present. The second heterogeneous group may include at least two, at least three, at least four, at least five, at least six, at least seven, or at least eight different second universal capture binding sequences. In one embodiment, members of the second heterogeneous group having a complement having 100% complementarity with the second capture sequence exist in greater numbers than other members of the second heterogeneous group. The second heterogeneous group of universal capture binding sequences at the 5' end may also include individual second universal capture binding sequences having a length less than the length of the second capture sequence, such as a length one to 12 nucleotides shorter than the length of the second capture sequence. Individual members of the second heterogeneous group may have a length less than the length of the second capture sequence and may include one to five nucleotides that are non-complementary to the sequence of the second capture sequence, or may include a complement having 100% complementarity with the sequence of the second capture sequence.
[0140] Other compositions that may be produced comprise a solution containing different double-stranded target nucleic acids from a single sample or source, e.g., a library, wherein each target nucleic acid comprises a universal adapter attached to each end. The universal adapter comprises a universal capture binding sequence, and the universal capture binding sequence is a heterogeneous population. The heterogeneous population may comprise at least two, at least three, at least four, at least five, at least six, at least seven, or at least eight different universal capture binding sequences.
[0141] Use in sequencing / sequencing methods
[0142] For example, an array of the present disclosure comprising a target nucleic acid generated by the method presented herein and amplified at an amplification site can be used in any and all various applications. A particularly useful application is nucleic acid sequencing. An example is sequencing-by-synthesis (SBS). In SBS, the nucleotide sequence of a template is determined by monitoring the extension of a nucleic acid primer along a nucleic acid template (e.g., a target nucleic acid or its amplicon). The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme). In a specific polymerase-based SBS embodiment, a fluorescently labeled nucleotide is added to the primer in a template-dependent manner (thus extending the primer) so that the sequence of the template can be determined by detecting the sequence and type of nucleotides added to the primer. Multiple different templates at different sites of the array described herein may be subjected to SBS technology under conditions where events occurring for different templates can be distinguished by their location in the array.
[0143] A flow cell provides a convenient format for receiving an array that receives SBS or other detection technology, which is generated by the method of the present disclosure and involves the repeated delivery of reagents periodically. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc., may flow into / through a flow cell receiving an array of nucleic acid templates. Sites of such an array where labeled nucleotides are incorporated due to primer extension may be detected. Optionally, the nucleotides may further possess a reversible termination property that terminates further primer extension once the nucleotides are added to the primer. For example, a nucleotide analog having a reversible termination moiety may be added to the primer, such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, in an embodiment using reversible termination, a deblocking reagent may be delivered to the flow cell (before or after detection occurs). Washing may be performed between various delivery steps. Subsequently, the cycle can be repeated n times to extend the primer by n nucleotides, thereby enabling the detection of a sequence of length n. Exemplary SBS procedures, fluid systems, and detection platforms that can be easily configured for use with arrays generated by the method of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. Patent No. 7,057,026; WO 91 / 06678; WO 07 / 123,744; U.S. Patent No. 7,329,492; U.S. Patent No. 7,211,414; U.S. Patent No. 7,315,019; U.S. Patent No. 7,405,281, and U.S. Patent No. 8,343,746.
[0144] Other sequencing procedures using circular reactions, such as pyresequencing, can be used. Pyresequencing detects the release of inorganic pyrophosphate (PPi) as specific nucleotides are incorporated into the initial nucleic acid strand (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568, and U.S. Patent No. 6,274,320). In pyresequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulferylase, and the level of the generated ATP can be detected via luciferase-generated photons. Therefore, the sequencing reaction can be monitored through a luminescence detection system. In a fluorescence-based detection system The excitation radiation source used is not required for the pyrosequencing procedure. Useful fluid systems, detectors, and procedures that may be used to apply pyrosequencing to the arrays of the present disclosure are described, for example, in WIPO Patent Application No. 2012 / 058096, U.S. Patent Publication No. 2005 / 0191698 A1, U.S. Patent No. 7,595,883, and U.S. Patent No. 7,244,559.
[0145] Sequencing reactions by ligation are also useful, including those described in, for example, Shendure et al., Science 309:1728-1732 (2005); U.S. Patent No. 5,599,675; and U.S. Patent No. 5,750,341. Some embodiments may include sequencing procedures by hybridization as described in, for example, Bains et al., Journal of Theoretical Biology 135(3), 303-7 (1988); Drmanac et al., Nature Biotechnology 16, 54-58 (1998); Fodor et al., Science 251(4995), 767-773 (1995); and WO 1989 / 10977. In both sequencing by ligation and sequencing by hybridization procedures, template nucleic acids (e.g., target nucleic acids or their amplicons) present in a region of the array undergo repeated cycles of oligonucleotide delivery and detection. A fluid system for the SBS method described herein or in the references cited herein can be easily configured to deliver reagents for sequencing by ligation or sequencing by hybridization procedures. Typically, oligonucleotides are fluorescently labeled and can be detected using a fluorescent detector similar to those described herein or in the references cited herein in relation to the SBS procedure.
[0146] Some embodiments may use methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation may be detected via fluorescence resonance energy transfer (FRET) interactions between the fluorophore-containing polymerase and γ-phosphate-labeled nucleotides or by using zero-mode waveguides (ZMWs). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008).
[0147] Some SBS embodiments involve the detection of protons emitted when nucleotides are incorporated into the extension product. For example, sequencing based on the detection of emitted protons may use electrodetectors and related technologies commercially available from Ion Torrent (a subsidiary of Guilford, Conn., Life Technologies), or sequencing methods and systems described in US 2009 / 0026082 A1; US 2009 / 0127589 A1; US 2010 / 0137143 A1; or US 2010 / 0282617 A1. The method presented herein for amplifying target nucleic acids using exclusion amplification can be easily applied to the substrate used to detect protons. More specifically, using the method presented herein, a clonal population of amplicons can be generated at a site of the array used to detect protons.
[0148] For example, a useful application for the array of the present disclosure generated by the method presented herein is gene expression analysis. Gene expression can be detected or quantified using RNA sequencing techniques, such as digital RNA sequencing. RNA sequencing techniques can be performed using sequencing methods known in the art as described above. Gene expression can also be detected or quantified using hybridization techniques performed by direct hybridization on the array, or using multiplex analysis in which products are detected on the array. For example, the array of the present disclosure generated by the method presented herein can also be used to determine the genotype of genomic DNA samples from one or more individuals. Exemplary methods for array-based expression and genotyping analysis that can be performed on the array of the present disclosure include U.S. Patent No. 7,582,420; No. 6,890,741; It is described in US Patent Publication No. 6,913,884 or No. 6,355,431 or US Patent Publication 2005 / 0053980 A1; 2009 / 0186349 A1 or US 2005 / 0181440 A1.
[0149] Another useful application for arrays generated by the method presented herein is single-cell sequencing. In combination with indexing methods, single-cell sequencing can be used for chromatin accessibility analysis to generate profiles of active regulatory elements from thousands of single cells and to generate single-cell whole-genome libraries. Examples of single-cell sequencing that can be performed on arrays of the present disclosure are described in U.S. Patent Publication 2018 / 0023119 A1, and U.S. Provisional Patent Applications 62 / 673,023 and 62 / 680,259.
[0150] The advantage of the method presented herein is that it provides for the rapid and efficient generation of an array from any of various nucleic acid libraries. Accordingly, the present disclosure provides an integrated system capable of generating an array using one or more of the methods presented herein and further detecting nucleic acids on the array using techniques known in the art, such as those exemplified above. Accordingly, the integrated system of the present disclosure may include fluid components capable of delivering amplification reagents to the array of amplification sites, such as pumps, valves, reservoirs, and fluid lines. A particularly useful fluid component is a flow cell. A flow cell may be configured and / or used in the integrated system to generate the array of the present disclosure and to detect the array. Exemplary flow cells are described, for example, in U.S. Patent Publication 2010 / 0111768 A1 and U.S. Patent No. 8,951,781. As exemplified with respect to the flow cell, one or more of the fluid components of the integrated system may be used in the amplification method and the detection method. As an example of an embodiment of nucleic acid sequencing, one or more of the fluid components of the integrated system may be used for the delivery of sequencing reagents in the amplification method presented herein and in the sequencing method as exemplified above. Alternatively, the integrated system may include separate fluid systems to perform the amplification method and the detection method. An example of an integrated sequencing system capable of generating a nucleic acid array and also determining a nucleic acid sequence is MiSeq TM , HiSeq TM , NextSeq TM , MiniSeq TM , NovaSeq TM , and iSeq TMPlatform (Illumina, Inc., San Diego, California), and devices described in U.S. No. 8,951,781, but not limited thereto. Such devices may be modified to form an array using exclusion amplification in accordance with the instructions described herein.
[0151] A system capable of performing the method described herein does not need to be integrated with a detection device. Rather, a standalone system or a system integrated with other devices is also possible. In the context of an integrated system, fluid components similar to those exemplified above may be used in such embodiments.
[0152] A system capable of performing the method described herein may include a system controller capable of executing a series of instructions to perform one or more steps of the method, technique, or process described herein, regardless of whether it is integrated with a detection function. For example, instructions may instruct the performance of a step of generating an array under exclusion amplification conditions. Optionally, instructions may further instruct the performance of a nucleic acid detection step using the method described above herein. A useful system controller may include any processor-based or microprocessor-based system, including a system using a microcontroller, RISC, ASIC, FPGA, logic circuit, and any other circuit or processor capable of performing the functions described herein. A series of instructions to the system controller may be in the form of a software program. As used herein, the terms “software” and “firmware” are interchangeable and include any computer program stored in memory to be executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. Software may take various forms, such as system software or application software. Additionally, software may take the form of a collection of separate programs, a program module within a larger program, or a part of a program module. Software may also include modular programming in the form of object-oriented programming.
[0153] Several applications for the arrays of the present disclosure have been illustrated above in the context of ensemble detection, where multiple amplicons present in each amplification site are detected together. In alternative embodiments, a single nucleic acid, whether a target nucleic acid or its amplicon, may be detected at each amplification site. For example, the amplification site may be configured to contain a single nucleic acid molecule having the target nucleotide sequence to be detected and a plurality of filler nucleic acids. In this example, the filler nucleic acids function to charge the capacity of the amplification site and are not necessarily intended to be detected. The single molecule to be detected may be detected by a method capable of distinguishing the single molecule from the background of the filler nucleic acids. For example, any of the various single molecule detection techniques may be used, including variations of the ensemble detection techniques presented above to detect sites at increased gain or to use more sensitive labels. Other examples of single molecule detection methods that may be used are U.S. Patent Publication 2011 / 0312529 A1; It is described in U.S. Patent No. 9,279,154 and U.S. Patent Publication 2013 / 0085073 A1.
[0154] Arrays useful for detecting single-molecule nucleic acids may be generated using one or more of the methods presented herein with modifications as follows. A plurality of different target nucleic acids may be configured to include both a target nucleotide sequence to be detected and one or more filler nucleotide sequences to be amplified to produce a filler amplicon. A plurality of different target nucleic acids may be included in an amplification reagent such as that presented elsewhere in this invention and may react with an array of amplification sites under exclusion amplification conditions such that the filler nucleotide sequence(s) fill the amplification sites. An exemplary configuration that may be used to amplify the filler sequence while prohibiting the amplification of the target sequence comprises, for example, a single target molecule having a first region having a filler sequence adjacent to a binding site for an amplification primer present in the amplification site and a second region having a target sequence outside the adjacent region. In other configurations, the target nucleic acid may comprise distinct molecules or strands carrying the target sequence and the filler sequence(s), respectively. Separate molecules or strands can be attached to particles or formed as arms of nucleic acid dendrimers or other branched structures.
[0155] In a specific embodiment, an array having amplification sites containing a filler sequence and a target sequence, respectively, can be detected using primer extension analysis or synthetic sequencing techniques. In this case, specific extensions in the target nucleotide sequence can be achieved in contrast to a large amount of filler sequences by using appropriately placed primer binding sites. For example, the binding site for the sequencing primer may be placed upstream of the target sequence and may not be present in any filler sequence. Alternatively or additionally, the target sequence may include one or more non-natural nucleotide analogs that cannot hydrogen bond to standard nucleotides. The non-natural nucleotide(s) may be located downstream of the primer binding site (e.g., the target sequence or the region interposed between the target sequence and the primer binding site) and thus will prevent extension or synthetic sequencing until an appropriate nucleotide partner (i.e., a partner capable of hydrogen bonding to the non-natural analog(s) in the target sequence) is added. The nucleotide analogs isocytosine (isoC) and isoguanine (isoG) are particularly useful because they specifically pair with each other but do not pair with other standard nucleotides used in most extension and synthetic sequencing techniques. An additional advantage of using isoC and / or isoG on or upstream of the target sequence is to prevent unwanted amplification of the target sequence during the amplification step by omitting each partner from the nucleotide mixture used for amplification.
[0156] For example, it will be understood that the array of the present disclosure generated by the method presented herein does not need to be used in a detection method. Rather, the array can be used to store a nucleic acid library. Accordingly, the array can be stored with the nucleic acid preserved inside. For example, the array can be stored in a dried state, a frozen state (e.g., in liquid nitrogen), or in a solution that protects the nucleic acid. Alternatively or additionally, the array can be used to replicate the nucleic acid library. For example, the array can be used to generate a replicated amplicon from one or more sites on the array.
[0157] Various embodiments of the present disclosure are illustrated herein in relation to transporting a target nucleic acid to an amplification site of an array and creating a replica of the target nucleic acid captured at the amplification site. Similar methods may be used for non-nucleic acid target molecules. Thus, the methods presented herein may be used with other target molecules instead of the illustrative target nucleic acid. For example, the methods of the present disclosure may be performed to transport individual target molecules from a population of different target molecules. Each target molecule may be transported to an individual site of the array to initiate a reaction at the capture site (in some cases, it may be captured at that individual site). The reaction at each site may, for example, produce a replica of the captured molecule, or the reaction may change the site to separate or isolate the captured molecule. In both cases, the end result may be sites of the array that are pure in relation to the type of target molecule present from the population containing different types of target molecules.
[0158] In a specific embodiment using target molecules rather than nucleic acids, a library of different target molecules can be generated using a method utilizing exclusion amplification. For example, a target molecule array can be created under conditions where sites of the array are randomly seeded with target molecules from solution and copies of the target molecules are generated to fill each seeded site to capacity. According to the exclusion amplification method of the present disclosure, the seeding and replication processes may proceed simultaneously under conditions where the rate at which replication is generated exceeds the seeding rate. Thus, the relatively fast rate at which replication is generated at a site seeded by a first target molecule effectively excludes a second target molecule from seeding the site. In some cases, seeding of the target molecule initiates a reaction that fills the site to capacity by a process other than replication of the target molecule. For example, when a target molecule is captured at one site, a chain reaction is initiated, eventually preventing the site from capturing a second target molecule. The chain reaction may occur under exclusion amplification conditions at a rate exceeding the rate at which the target molecule is captured.
[0159] As exemplified with respect to the target nucleic acid, exclusion amplification, when applied to other target molecules, may utilize a relatively slow rate to initiate a repetitive reaction (e.g., a chain reaction) at a site in the array versus a relatively fast rate to continue the repetitive reaction once initiated. In the example of the previous paragraph, exclusion amplification occurs due to a relatively slow rate of target molecule seeding (e.g., relatively slow diffusion) versus a relatively fast rate at which a reaction occurs to charge the site with a replica of the target molecule seed. In another exemplary embodiment, exclusion amplification may occur due to a delay in the formation of the first replica of the target molecule seeding the site (e.g., delayed or slow activation) versus a relatively fast rate at which subsequent replicas are generated to charge the site. In this example, individual sites may have been seeded with multiple different target molecules. However, the formation of the first replica for any given target molecule may be randomly activated, such that the average rate of the first replica formation is relatively slow compared to the rate at which subsequent replicas are generated. In this case, individual sites may have been seeded with multiple different target molecules, but exclusion amplification will ensure that only one of those target molecules is replicated.
[0160] Accordingly, the present disclosure provides a method for generating an array of molecules, the method comprising: (a) providing a reagent comprising (i) an array of sites and (ii) a solution having a plurality of different target molecules; and (b) reacting the reagent to generate a plurality of sites each having a single target molecule from a plurality of sites, or generating a plurality of sites each having a pure population of copies from individual target molecules from a solution, wherein the number of target molecules in the solution exceeds the number of sites of the array, the different target molecules have fluid access to the plurality of sites, and each site has a capacity for multiple target molecules within the plurality of different target molecules, and the reaction step simultaneously comprises (i) transporting different molecules to the sites at an average transport rate and (ii) initiating a reaction to fill the sites to capacity at an average reaction rate, wherein the average reaction rate exceeds the average transport rate. In some embodiments, step (b) may instead be performed by reacting a reagent to produce a plurality of sites each having a single target molecule from a plurality of sites or by producing a plurality of sites each having a pure population of copies from individual target molecules from a solution, wherein the reaction step comprises (i) initiating a repetitive reaction (e.g., a chain reaction) to form a product from the target molecule at each site and (ii) continuing the reaction at each site to form a subsequent product, and the average rate at which the reaction occurs at the site exceeds the average rate at which the reaction begins at the site.
[0161] In the above-described non-nucleic acid embodiment, the target molecule may be an initiator of a repetitive reaction occurring at each site of the array. For example, the repetitive reaction may form a polymer that prevents other target molecules from occupying the site. Alternatively, the repetitive reaction may form one or more polymers that constitute molecular replicas of the target molecule transported to the site.
[0162] Exemplary implementation example
[0163] Implementation Example 1.
[0164] (a) A step of providing an amplification reagent,
[0165] The amplification reagent comprises (i) an array of amplification sites and (ii) a solution containing multiple different modified double-stranded target nucleic acids, and
[0166] The amplification site comprises two groups of capture nucleic acids each containing a capture sequence, the first group containing a first capture sequence, and the second group containing a second capture sequence.
[0167] The different modified target nucleic acid comprises, at the 3' end, a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence, and
[0168] (b) reacting an amplification reagent to generate a plurality of amplification sites, each containing a clonal population of amplicons from individual target nucleic acids from solution.
[0169] A method for amplifying nucleic acids, including
[0170] Implementation Example 2.
[0171] (a) A step of providing an amplification reagent,
[0172] The above amplification reagent comprises (i) an array of amplification sites and (ii) a solution comprising a plurality of different modified target nucleic acids, and
[0173] The amplification site comprises two groups of capture nucleic acids each containing a capture sequence, the first group containing a first capture sequence, and the second group containing a second capture sequence.
[0174] The different modified target nucleic acid comprises, at the 3' end, a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence, and
[0175] (b) a step of reacting an amplification reagent to generate a plurality of amplification sites, each comprising a clone population of an amplicon from an individual target nucleic acid from a solution,
[0176] The above reaction is,
[0177] (i) generating a first amplicon from individual target nucleic acids transported to each of the amplification sites, and
[0178] (ii) generating a subsequent amplicon from the individual target nucleic acid transported to each of the amplification sites or from the first amplicon
[0179] A step including, wherein the average speed at which the subsequent amplicon is generated in the amplification section is slower than the average speed at which the first amplicon is generated in the amplification section.
[0180] A method for amplifying nucleic acids, including
[0181] Embodiment 3. A method for determining a nucleic acid sequence comprising the step of performing a sequencing procedure to detect an apparent clonal population of amplicons at each of a plurality of amplicon sites on an array,
[0182] The above array is,
[0183] (a) A step of providing an amplification reagent,
[0184] The amplification reagent comprises a solution containing (i) a plurality of amplification sites and (ii) a plurality of different modified target nucleic acids, and
[0185] The amplification site comprises two groups of capture nucleic acids each containing a capture sequence, the first group containing a first capture sequence, and the second group containing a second capture sequence.
[0186] The different modified target nucleic acid comprises, at the 3' end, a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence, and
[0187] (b) Step of reacting the amplification reagent
[0188] A method that is created by a process including:
[0189] Embodiment 4. In any one of Embodiments 1 to 3, the number of different modified target nucleic acids in the solution exceeds the number of amplification sites of the array, and
[0190] Different modified target nucleic acids have fluid access to multiple amplification sites, and
[0191] A method in which each of the amplification sites contains a capacity for several nucleic acids among a plurality of different nucleic acids.
[0192] Example 5. In any one of Examples 1 to 4, the reaction is,
[0193] (i) a step of transporting different modified target nucleic acids to an amplification site at an average transport rate, and
[0194] (ii) a step of amplifying the target nucleic acid in the amplification site at an average amplification rate
[0195] A method comprising simultaneously, wherein the average amplification speed is slower than the average transport speed.
[0196] Embodiment 6. In any one of Embodiments 1 to 5, the plurality of different modified target nucleic acids in the solution are,
[0197] (i) transporting different modified target nucleic acids from solution to an amplification site, and
[0198] (ii) amplifying the target nucleic acid in the amplification site at an amplification rate to generate an array of amplicon sites each containing an apparent clonal population of the amplicon.
[0199] A method at a concentration that simultaneously generates
[0200] Embodiment 7. A method in any one of Embodiments 1 to 6, wherein the first universal capture binding sequence has less than 100% complementarity with the first capture sequence.
[0201] Embodiment 8. A method in any one of Embodiments 1 to 7, wherein the first universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the first capture sequence.
[0202] Embodiment 9. In any one of Embodiments 1 to 7, the different modified target nucleic acids comprise a heterogeneous population of the first universal capture binding sequence, and
[0203] A method in which a heterogeneous population comprises (i) one, two, or three nucleotides that are non-complementary to the first capture sequence, or (ii) an individual first universal capture binding sequence that is 100% complementary to the first capture sequence.
[0204] Embodiment 10. A method in any one of Embodiments 1 to 9, wherein the members of a heterogeneous group having 100% complementarity with the first capture sequence exist in a greater number than the remaining members of the heterogeneous group.
[0205] Embodiment 11. A method in any one of Embodiments 1 to 10, wherein the first general-purpose capture binding sequence has a length less than the length of the first capture sequence.
[0206] Embodiment 12. A method in any one of Embodiments 1 to 11, wherein the first universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the first capture sequence.
[0207] Embodiment 13. In any one of Embodiments 1 to 12, the different modified target nucleic acids comprise a heterogeneous population of the first universal capture binding sequence, and
[0208] A method in which the heterogeneous population comprises individual first universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the first capture sequence.
[0209] Embodiment 14. A method in any one of Embodiments 1 to 13, wherein the heterogeneous group further comprises (iii) an individual first universal capture combination sequence having a length less than the length of the first capture sequence.
[0210] Embodiment 15. A method in any one of Embodiments 1 to 14, wherein the individual first universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the first capture sequence.
[0211] Embodiment 16. A method in any one of Embodiments 1 to 15, wherein an individual member of a heterogeneous group having a length less than the length of the first capture sequence comprises one, two, or three nucleotides that are non-complementary to the sequence of the first capture sequence, or has 100% complementarity with the sequence of the first capture sequence.
[0212] Embodiment 17. A method in any one of Embodiments 1 to 16, wherein the different modified target nucleic acid comprises, at the 5' end, a second universal capture binding sequence having a complement having an affinity for the second capture sequence lower than that of the second universal capture binding sequence having a complement having 100% complementary to the second capture sequence.
[0213] Embodiment 18. A method in any one of Embodiments 1 to 17, wherein the complement of the second universal capture binding sequence has less than 100% complementarity with the second capture sequence.
[0214] Embodiment 19. A method in any one of Embodiments 1 to 18, wherein the complement of the second universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the second capture sequence.
[0215] Embodiment 20. In any one of Embodiments 1 to 19, the different modified target nucleic acid comprises a heterogeneous population of the second universal capture binding sequence, and
[0216] A method wherein the heterogeneous population comprises (i) one, two, or three nucleotides that are non-complementary to the second capture sequence, or (ii) an individual second universal capture binding sequence comprising a complement that is 100% complementary to the second capture sequence.
[0217] Embodiment 21. A method in any one of Embodiments 1 to 20, wherein the members of a heterogeneous group comprising a complement having 100% complementarity with the second capture sequence exist in a greater number than the remaining members of the heterogeneous group.
[0218] Embodiment 22. A method in any one of Embodiments 1 to 21, wherein the second universal capture binding sequence has a length less than the length of the second capture sequence.
[0219] Embodiment 23. A method in any one of Embodiments 1 to 22, wherein the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence.
[0220] Embodiment 24. In any one of Embodiments 1 to 23, the different modified target nucleic acid comprises a heterogeneous population of the second universal capture binding sequence, and
[0221] A method in which the heterogeneous population comprises individual second universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the second capture sequence.
[0222] Embodiment 25. A method in any one of Embodiments 1 to 24, wherein the heterogeneous group further comprises (iii) an individual second universal capture combination sequence having a length less than the length of the second capture sequence.
[0223] Embodiment 26. A method in any one of Embodiments 1 to 25, wherein the individual second universal capture binding sequence has a length of 1 to 12 nucleotides less than the length of the second capture sequence.
[0224] Embodiment 27. A method in any one of Embodiments 1 to 26, wherein an individual member of a heterogeneous group having a length less than the length of the second capture sequence comprises one, two, or three nucleotides that are non-complementary to the sequence of the second capture sequence, or comprises a complement having 100% complementarity with the sequence of the second capture sequence.
[0225] Embodiment 28. A method in any one of Embodiments 1 to 27, wherein the target nucleic acid is DNA.
[0226] Embodiment 29. A method in any one of Embodiments 1 to 28, wherein the array of amplification sites comprises an array of feature portions on a surface.
[0227] Embodiment 30. A method in any one of Embodiments 1 to 29, wherein the area for each feature portion is larger than the diameter of the excluded volume of the target nucleic acid transported to the amplification site.
[0228] Embodiment 31. A method in any one of Embodiments 1 to 30, wherein the feature portion is discontinuous and separated by an interstitial region of a surface free of a capture agent.
[0229] Embodiment 32. A method in any one of Embodiments 1 to 31, wherein each feature part comprises a bead, a well, a channel, a ridge, a protrusion, or a combination thereof.
[0230] Embodiment 33. A method in which, in Embodiment 1 or Embodiment 2, the array of amplification sites comprises beads in the solution or beads on the surface.
[0231] Embodiment 34. A method in any one of Embodiments 1 to 3, wherein the amplification of the target nucleic acid transported to the amplification site occurs isothermally.
[0232] Embodiment 35. A method in any one of Embodiments 1 to 34, wherein the amplification of different modified target nucleic acids transported to the amplification site does not include a denaturation cycle.
[0233] Embodiment 36. A method in any one of Embodiments 1 to 35, wherein a plurality of amplification sites comprising a clone population of an amplicon exceeds 40% of the amplification sites where different modified target nucleic acids had fluid access during step (b).
[0234] Embodiment 37. A method in any one of Embodiments 1 to 36, wherein a sufficient number of amplifiers are generated from each individual target nucleic acid at each individual amplification site to fill the capacity of each amplification site during step (b).
[0235] Embodiment 38. A method in any one of Embodiments 1 to 37, wherein the rate at which an amplifier is generated to fill the capacity of each amplification site is less than the rate at which individual target nucleic acids are transported to individual amplification sites.
[0236] Embodiment 39. A method in any one of Embodiments 1 to 38, wherein the transport comprises passive diffusion.
[0237] Embodiment 40. A method in any one of Embodiments 1 to 39, wherein the amplification reagent further comprises a polymerase and a recombinant enzyme.
[0238] delete
[0239] delete
[0240] delete
[0241] delete
[0242] delete
[0243] delete
[0244] delete
[0245] delete
[0246] delete
[0247] delete
[0248] delete
[0249] delete
[0250] delete
[0251] delete
[0252] delete
[0253] Embodiment 49. A composition comprising an array of amplification sites and at least one target nucleic acid bound to the amplification sites, wherein
[0254] The amplification site comprises two groups of capture nucleic acids, each group comprising a capture sequence, the first group comprising a first capture sequence, and the second group comprising a second capture sequence, and
[0255] The target nucleic acid includes a first universal capture binding sequence at the 3' end having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence, and
[0256] A composition in which the target nucleic acid universal capture binding sequence hybridizes to the first capture sequence.
[0257] Embodiment 50. A composition according to Embodiment 49, wherein the first universal capture binding sequence has less than 100% complementarity with the first capture sequence.
[0258] Embodiment 51. A composition according to Embodiment 49 or Embodiment 50, wherein the first universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the first capture sequence.
[0259] Embodiment 52. A composition in any one of Embodiments 49 to 51, wherein at least 30% of the amplification sites of the array are occupied by at least one target nucleic acid.
[0260] Embodiment 53. In any one of Embodiments 49 to 52, the first universal capture binding sequence comprises a heterogeneous group, and
[0261] The heterogeneous population comprises (i) one, two, or three nucleotides that are non-complementary to the first capture sequence, or (ii) an individual first universal capture binding sequence that is 100% complementary to the first capture sequence, and
[0262] A composition in which members of a heterogeneous group are combined at different amplification sites.
[0263] Embodiment 54. A composition in any one of Embodiments 49 to 53, wherein the members of a heterogeneous group having 100% complementarity with the first capture sequence exist in a greater number than the remaining members of the heterogeneous group.
[0264] Embodiment 55. A composition in any one of Embodiments 49 to 54, wherein the second universal capture binding sequence has a length less than the length of the second capture sequence.
[0265] Embodiment 56. A composition in any one of Embodiments 49 to 55, wherein the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence.
[0266] Embodiment 57. In any one of Embodiments 49 to 56, the composition comprises a plurality of different target nucleic acids, and
[0267] Different target nucleic acids include heterogeneous populations of the first universal capture binding sequence, and
[0268] A composition comprising a heterogeneous population including an individual first universal capture binding sequence having 1 to 12 fewer nucleotides than the length of the first capture sequence.
[0269] Embodiment 58. A composition in any one of Embodiments 49 to 57, wherein the heterogeneous group further comprises (iii) an individual first universal capture binding sequence having a length less than the length of the first capture sequence.
[0270] Embodiment 59. A composition in any one of Embodiments 49 to 58, wherein the individual first universal capture binding sequence having a reduced length has a length that is 1 to 12 nucleotides shorter than the length of the first capture sequence.
[0271] Embodiment 60. A composition in any one of Embodiments 49 to 59, wherein individual members of a heterogeneous group having a reduced length have one, two, or three nucleotides that are non-complementary to the sequence of the first capture sequence, or have 100% complementarity with the sequence of the first capture sequence.
[0272] Embodiment 61. In any one of Embodiments 49 to 60, the composition comprises a plurality of different target nucleic acids, and
[0273] A composition comprising a different target nucleic acid having a second universal capture binding sequence at the 5' end having a complement having an affinity for the second capture sequence lower than that of a second universal capture binding sequence having a complement having 100% complementation with the second capture sequence.
[0274] Embodiment 62. A composition in any one of Embodiments 49 to 61, wherein the complement of the second universal capture binding sequence has less than 100% complementarity with the second capture sequence.
[0275] Embodiment 63. A composition in any one of Embodiments 49 to 62, wherein the complement of the second universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the second capture sequence.
[0276] Embodiment 64. In any one of Embodiments 49 to 63, the composition comprises a plurality of different target nucleic acids, and
[0277] Different target nucleic acids include heterogeneous populations of second universal capture binding sequences, and
[0278] A composition comprising a heterogeneous population having (i) one, two, or three nucleotides non-complementary to the second capture sequence, or (ii) an individual second universal capture binding sequence comprising a complement having 100% complementarity with the second capture sequence.
[0279] Embodiment 65. A composition in any one of Embodiments 49 to 64, wherein the members of a heterogeneous group comprising a complement having 100% complementarity with the second capture sequence exist in a greater number than the remaining members of the heterogeneous group.
[0280] Embodiment 66. A composition in any one of Embodiments 49 to 65, wherein the target nucleic acid comprises a second universal capture binding sequence having a length less than the length of the second capture sequence at the 5' end.
[0281] Embodiment 67. A composition in any one of Embodiments 49 to 66, wherein the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence.
[0282] Embodiment 68. In any one of Embodiments 49 to 67, the composition comprises a plurality of different target nucleic acids, and
[0283] Different target nucleic acids include heterogeneous populations of second universal capture binding sequences, and
[0284] A composition comprising a heterogeneous population including individual second universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the second capture sequence.
[0285] Embodiment 69. A composition in any one of Embodiments 49 to 68, wherein the heterogeneous group further comprises (iii) an individual second universal capture binding sequence having a length less than the length of the second capture sequence.
[0286] Embodiment 70. A composition in any one of Embodiments 49 to 69, wherein the individual second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence.
[0287] Embodiment 71. A composition in any one of Embodiments 49 to 70, wherein an individual member of a heterogeneous group having a length less than the length of the second capture sequence comprises one, two, or three nucleotides that are non-complementary to the sequence of the second capture sequence, or comprises a complement having 100% complementarity with the sequence of the second capture sequence.
[0288] Embodiment 72. A composition in any one of Embodiments 49 to 71, wherein the target nucleic acid is DNA.
[0289] Examples
[0290] The present invention is illustrated by the following examples. It should be understood that specific examples, materials, quantities, and procedures should be interpreted broadly in accordance with the scope and spirit of the invention as described herein.
[0291] Example 1
[0292] General analysis methods and conditions
[0293] Unless otherwise noted, this describes the general analytical conditions used in the embodiments described herein.
[0294] Standard Nexter to introduce the universal part of the adapter through the tagging of human gDNATM Starting with library preparation, a nucleic acid library was generated. Subsequently, this universal tagging was divided into individual reactions, with one reaction for each of the different adapter pairs (PCR1-PCR21). In each of these reactions, the modified adapter was introduced via 12 cycles of PCR by replacing the standard P5 / P7 adapter with the modified adapter using Nextera™ XT library preparation reagents (Illumina, Inc., San Diego, California). The modified adapter was designed using modifications from the standard P5 / P7 sequence (as outlined below) and synthesized by Integrated DNA Technologies (IDT Inc., Skokie, Illinois).
[0295] Example 2
[0296] Evaluation of adapter mutants for exercise delay
[0297] Using various modified adapters, libraries different from the standard Nextera™ library (Table 1) were generated. These adapters were slightly shorter than the standard (-4 bp or -9 bp) or had one, two, or three mismatches ('wobble' bases: 1W, 2W, or 3W, respectively) introduced along the length of the P5 / P7 region. The range of mutations was from a complete P5 & P7 sequence (PCR 1) to three mismatches at both ends (PCR 21).
[0298] Next, the concentration of each library was quantified, and the libraries were normalized to the same concentration. Then, libraries with different modified adapters were amplified separately on a qPCR instrument (BioRad CFX384 Real-Time System) using the KAPA Library Quantification Kit for Illumina Platform (Kapa Biosystems) with custom primers to simulate flow cell conditions and generate the results in Table 1. Efficiency was calculated from the number of cycles required to reach Ct (critical cycle), as standardly defined for qPCR.
[0299] FCP5 FCP7 Efficiency Adapter number Adapter identity qPCR 3 1 PCR 1 - Nex Control 1.5079 2 PCR 2 αβ-4 1.4875 3 PCR 3 αβ-9 1.5229 4 PCR 4 - Nex P5 + 1W 1.4305 5 PCR 5 - Nex P5 + 2W 1.4160 6 PCR 6 - Nex P5 + 3W 1.3716 7 PCR 7 αβ-4 + 1W 1.3577 8 PCR 8 αβ-4 + 2W 1.3673 9 PCR 9 αβ-4 + 3W 1.3273 10 PCR 10 αβ-9 + 1W 1.4236 11 PCR 11 αβ-9 + 2W 1.4179 12 PCR 12 αβ-9 + 3W 1.3691 13 PCR 13 - 1MM + 1MM 1.3692 14 PCR 14 - 1MM + 2W 1.3789 15 PCR 15 - 1MM + 3W 1.3488 16 PCR 16 - 2W + 1MM 1.3504 17 PCR 17 - 2W + 2W 1.2971 18 PCR 18 - 2W + 3W 1.2980 19 PCR 19 - 3W + 1MM 1.2861 20 PCR 20 - 3W + 2W 1.2498 21 PCR 21 - 3W + 3W 1.2261
[0300] αβ-4 refers to 4 base pairs removed from the adapter, αβ-9 refers to 9 base pairs removed from the adapter, 1W, 2W, and 3W refer to 1, 2, or 3 wobble mismatches, respectively, existing only along the length of the region binding to the capture nucleic acid present on the surface of the array, 1MM refers to a true mismatch, and FCP5 and FCP7 refer to full-length P5 and P7 primers, respectively.
[0301] Example 3
[0302] Evaluation of adapter mutants for exercise delay using sequencing analysis
[0303] Different mutant adapters were selected and run individually or in groups on the HiSeq™ flowcell (Table 2). All sequencing was performed using the Illumina HiSeq standard reagent kit. TM It was performed on device X.
[0304] Flowcell lane Adapter used in the lane Lane 1 PCR 1 Lane 2 PCR 2 Lane 3 PCR 18 Lane 4 PCR 6 Lane 5 PCR 10 Lane 6 PCR 1-2, 4-6, 10 Lane 7 PCR 1-2, 7-8, 10, 11 Lane 8 PCR 2, 6, 8, 12
[0305] As shown in the results in 3a, lanes 1, 2, and 4 were reactions using adapters with low mismatch and high efficiency, and as expected, the percentage and strength of clusters passing through the filter increased. Lanes 3 and 5 were reactions using adapters with high mismatch and low efficiency, and as expected, the percentage and strength of clusters passing through the filter were low. Lanes 6, 7, and 8 were a mixture of high-efficiency and low-efficiency adapters. Counterintuitively, although the mixture did not achieve the average performance of the individual components (e.g., the middle ground between high and low efficiency), it outperformed all single-type libraries in both the number of clusters passing through the filter and the strength. Thus, the surprising result is that reducing average homology improved the speed of the so-called monoclonal performance of the nanowell, even though the average amplification rate decreased. A novel point is that some degree of variability was introduced into the adapter sequences, resulting in varying efficiencies among the template populations. In this way, when multiple molds are seeded onto a pad, the mold that generally has an advantage over the others is present to clearly dominate the pad. Additionally, the reduced homology is corrected in the sub-replications so that a delay is introduced only in the first replication without affecting the efficiency of subsequent amplification.
[0306] Different mutant adapters were selected and run either individually or in groups on a HiSeq™ flow cell (Table 3). As such, all sequencing was performed on an Illumina HiSeq™ instrument using standard reagent kits.
[0307] Flowcell lane Adapter used in the lane Lane 1 PCR 1 Lane 2 PCR 3 Lane 3 PCR 10 Lane 4 PCR 16 Lane 5 PCR 1, 3, 6, 8, 9, 10, 16-18 Lane 6 PCR 21 Lane 7 PCR 1, 3, 6, 8, 9, 10, 16-18 Lane 8 PCR 1, 3, 10
[0308] When groups of mutant adapters were run in a mixture (e.g., lane 7), they were combined at the same concentration. In conventional sequencing runs, i.e., when the proposed method is not used, using libraries of mixed adapters at the same concentration results in the same ratio of reads in the flow cell. As illustrated in Fig. 3b, the mixture of different mutant adapters resulted in a redefinition of the final read count, which was proportional to the efficiency rather than the seed concentration, thereby demonstrating the efficacy of the proposed method; that is, adapters with lower affinity had longer movement delays, and consequently, a lower ratio of final reads.
[0309] The foregoing detailed description and examples are provided for clarity only. Unnecessary limitations should not be understood from this. The present disclosure is not limited to the exact details shown and described, and variations obvious to those skilled in the art will be included in the disclosure as defined by the claims.
[0310] Unless otherwise specified, all numbers expressing quantities such as ingredients, molecular weights, etc., used in the specification and claims shall be understood in all cases as modified by the term “about.” Accordingly, unless otherwise indicated, numerical parameters described in the specification and claims are approximations that may vary depending on the desired characteristics to be obtained by the present disclosure. At least, without attempting to limit the principle of equivalence for the claims, each numerical parameter shall be interpreted by taking into account at least the reported significant digits and by applying general rounding techniques.
[0311] Although the numerical ranges and parameters describing the broad scope of this disclosure are approximations, the numerical values described in specific embodiments are reported as accurately as possible. However, all numerical values inherently include the range that inevitably arises from the standard deviation found in each test measurement.
[0312] All headings are for the convenience of the reader and must not be used to restrict the meaning of the text following them unless otherwise specified.
Claims
Claim 1 (a) a step of providing an amplification reagent, wherein the amplification reagent comprises a solution comprising (i) an array of amplification sites and (ii) a plurality of different modified double-stranded target nucleic acids, wherein the amplification sites comprise two groups of capture nucleic acids each comprising a capture sequence, wherein the first group comprises a first capture sequence and the second group comprises a second capture sequence, and the different modified target nucleic acids comprise a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence at the 3' end; and (b) a step of reacting the amplification reagent to generate a plurality of amplification sites each comprising a clone group of amplicons from individual target nucleic acids from the solution. Claim 2 A method according to claim 1, wherein the reaction comprises (i) generating a first amplifier from individual target nucleic acids transported to each of the amplification sites, and (ii) generating a subsequent amplifier from individual target nucleic acids transported to each of the amplification sites or from the first amplifier, wherein the average rate at which the subsequent amplifier is generated at the amplification site is slower than the average rate at which the first amplifier is generated at the amplification site. Claim 3 A method for determining a nucleic acid sequence, comprising the step of performing a sequencing procedure to detect a clone population of an amplicon at each of a plurality of amplicon sites on an array, wherein the array is produced by a process comprising: (a) a step of providing an amplification reagent, wherein the amplification reagent comprises a solution comprising (i) a plurality of amplification sites and (ii) a plurality of different modified target nucleic acids, wherein the amplification sites each comprise two populations of capture nucleic acids comprising a capture sequence, wherein the first population comprises a first capture sequence and the second population comprises a second capture sequence, and the different modified target nucleic acids each comprise a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence, and (b) a step of reacting the amplification reagent. Claim 4 A method according to any one of claims 1 to 3, wherein the number of different modified target nucleic acids in the solution exceeds the number of amplification sites of the array, the different modified target nucleic acids have fluid access to a plurality of amplification sites, and each of the amplification sites contains a capacity for several nucleic acids among a plurality of different nucleic acids. Claim 5 A method according to claim 1, wherein the reaction simultaneously comprises (i) a step of transporting a different modified target nucleic acid to an amplification site at an average transport rate, and (ii) a step of amplifying a target nucleic acid at an average amplification rate, wherein the average amplification rate is slower than the average transport rate. Claim 6 A method according to claim 3, wherein a plurality of different modified target nucleic acids in a solution are at a concentration that simultaneously causes (i) transporting different modified target nucleic acids from the solution to an amplification site, and (ii) amplifying the target nucleic acids in the amplification site at an amplification rate to produce an array of amplicon sites each comprising a clonal population of amplicons. Claim 7 A method according to any one of claims 1 to 3, wherein the first universal capture binding sequence has less than 100% complementarity with the first capture sequence. Claim 8 A method according to claim 7, wherein a) the first universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the first capture sequence; or b) different modified target nucleic acids comprise heterogeneous populations of the first universal capture binding sequence, wherein the heterogeneous populations have (i) one, two, or three nucleotides that are non-complementary to the first capture sequence, or (ii) individual first universal capture binding sequences having 100% complementarity with the first capture sequence. Claim 9 A method according to claim 8, wherein the number of members of a heterogeneous group having 100% complementarity with the first capture sequence is greater than the number of members of the heterogeneous group. Claim 10 A method according to claim 8, wherein the heterogeneous group further comprises (iii) individual first universal capture combination sequences having a length less than the length of the first capture sequence. Claim 11 A method according to any one of claims 1 to 3, wherein the first universal capture combination sequence has a length less than the length of the first capture sequence. Claim 12 A method according to claim 11, wherein the first universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the first capture sequence. Claim 13 A method according to claim 11, wherein different modified target nucleic acids comprise heterogeneous populations of a first universal capture binding sequence, and the heterogeneous population comprises individual first universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the first capture sequence. Claim 14 A method according to claim 10, wherein individual members of a heterogeneous group having a length less than the length of the first capture sequence include one, two, or three nucleotides that are non-complementary to the sequence of the first capture sequence, or have 100% complementarity with the sequence of the first capture sequence. Claim 15 A method according to any one of claims 1 to 3, wherein the different modified target nucleic acid comprises at the 5' end a second universal capture binding sequence having a complement having an affinity for the second capture sequence lower than that of the second universal capture binding sequence having a complement having 100% complementary to the second capture sequence. Claim 16 A method according to claim 15, wherein the complement of the second universal capture binding sequence has less than 100% complementarity with the second capture sequence. Claim 17 A method according to claim 16, wherein a) the complement of the second universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the second capture sequence; or b) different modified target nucleic acids comprise heterogeneous populations of the second universal capture binding sequence, wherein the heterogeneous population comprises individual second universal capture binding sequences that (i) have one, two, or three nucleotides that are non-complementary to the second capture sequence, or (ii) have a complement that is 100% complementary to the second capture sequence. Claim 18 A method according to claim 17, wherein members of a heterogeneous group including a complement having 100% complementarity with the second capture sequence exist in greater numbers than the remaining members of the heterogeneous group. Claim 19 A method according to claim 17, wherein the heterogeneous group further comprises (iii) individual second universal capture combination sequences having a length less than the length of the second capture sequence. Claim 20 A method according to claim 15, wherein the second universal capture combination sequence has a length less than the length of the second capture sequence. Claim 21 A method according to claim 20, wherein the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence. Claim 22 A method according to claim 20, wherein different modified target nucleic acids comprise heterogeneous populations of a second universal capture binding sequence, and the heterogeneous population comprises individual second universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the second capture sequence. Claim 23 A method according to any one of claims 1 to 3, wherein the array of amplification sites comprises an array of feature portions on the surface. Claim 24 A method according to paragraph 23, wherein the area for each feature part is larger than the diameter of the excluded volume of the target nucleic acid transported to the amplification site. Claim 25 A method according to claim 24, wherein the feature portion is discontinuous and separated by an interstitial region of a surface free of a capture agent. Claim 26 A composition comprising an array of amplification sites and at least one target nucleic acid bound to the amplification sites, wherein the amplification sites comprise two groups of capture nucleic acids, each group comprising a capture sequence, the first group comprising a first capture sequence, and the second group comprising a second capture sequence, wherein the target nucleic acid comprises a first universal capture binding sequence having an affinity for the first capture sequence lower than that of a first universal capture binding sequence having 100% complementarity with the first capture sequence at the 3' end, and wherein the target nucleic acid universal capture binding sequence is hybridized to the first capture sequence. Claim 27 A composition according to claim 26, wherein a) the first universal capture binding sequence has less than 100% complementarity with the first capture sequence, b) the first universal capture binding sequence contains one, two, or three nucleotides that are non-complementary to the first capture sequence, or c) at least 30% of the amplification sites of the array are occupied by at least one target nucleic acid. Claim 28 A composition according to claim 27, wherein the first universal capture binding sequence comprises a heterogeneous population, and the heterogeneous population has (i) one, two, or three nucleotides that are non-complementary to the first capture sequence, or (ii) individual first universal capture binding sequences that are 100% complementary to the first capture sequence, and members of the heterogeneous population are bound to different amplification sites. Claim 29 A composition according to claim 28, wherein members of a heterogeneous group having 100% complementarity with the first capture sequence exist in greater numbers than the remaining members of the heterogeneous group. Claim 30 A composition according to claim 28, wherein the heterogeneous group further comprises (iii) individual first universal capture combination sequences having a length less than the length of the first capture sequence. Claim 31 A composition according to claim 30, wherein a) an individual first universal capture binding sequence having a reduced length has a length of 1 to 12 nucleotides less than the length of the first capture sequence, or b) an individual member of a heterogeneous group having a reduced length contains 1, 2, or 3 nucleotides that are non-complementary to the sequence of the first capture sequence, or has 100% complementarity with the sequence of the first capture sequence. Claim 32 A composition according to claim 26, wherein the composition comprises a plurality of different target nucleic acids, and the different target nucleic acids comprise a second universal capture binding sequence having a complement having an affinity for the second capture sequence lower than that of the second universal capture binding sequence having a complement having 100% complementary to the second capture sequence at the 5' end. Claim 33 A composition according to claim 32, wherein the complement of the second universal capture binding sequence has less than 100% complementarity with the second capture sequence. Claim 34 A composition according to claim 33, wherein a) the complement of the second universal capture binding sequence comprises one, two, or three nucleotides that are non-complementary to the second capture sequence, or b) the composition comprises a plurality of different target nucleic acids, wherein the different target nucleic acids comprise a heterogeneous population of the second universal capture binding sequence, and the heterogeneous population comprises (i) one, two, or three nucleotides that are non-complementary to the second capture sequence, or (ii) individual second universal capture binding sequences comprising a complement that is 100% complementary to the second capture sequence, or c) the number of members of the heterogeneous population comprising a complement that is 100% complementary to the second capture sequence is greater than the number of members of the heterogeneous population. Claim 35 A composition according to claim 26, wherein the target nucleic acid comprises a second universal capture binding sequence having a length less than the length of the second capture sequence at the 5' end. Claim 36 A composition according to claim 35, wherein a) the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence, or b) the composition comprises a plurality of different target nucleic acids, wherein the different target nucleic acids comprise heterogeneous groups of the second universal capture binding sequence, and the heterogeneous groups comprise individual second universal capture binding sequences that are 1 to 12 nucleotides shorter than the length of the second capture sequence. Claim 37 A composition according to claim 34, wherein the heterogeneous group further comprises (iii) individual second universal capture combination sequences having a length less than the length of the second capture sequence. Claim 38 A composition according to claim 37, wherein the individual second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence. Claim 39 A composition according to claim 38, wherein individual members of a heterogeneous group having a length less than the length of the second capture sequence comprise one, two, or three nucleotides that are non-complementary to the sequence of the second capture sequence, or comprise a complement having 100% complementarity with the sequence of the second capture sequence. Claim 40 A composition according to claim 32, wherein the second universal capture binding sequence has a length less than the length of the second capture sequence. Claim 41 A composition according to claim 40, wherein the second universal capture binding sequence has a length that is 1 to 12 nucleotides shorter than the length of the second capture sequence. Claim 42 A composition according to claim 41, wherein the composition comprises a plurality of different target nucleic acids, the different target nucleic acids comprise a heterogeneous group of a first universal capture binding sequence, and the heterogeneous group comprises individual first universal capture binding sequences having 1 to 12 fewer nucleotides than the length of the first capture sequence. Claim 43 delete Claim 44 delete Claim 45 delete Claim 46 delete Claim 47 delete Claim 48 delete Claim 49 delete Claim 50 delete Claim 51 delete Claim 52 delete Claim 53 delete Claim 54 delete Claim 55 delete Claim 56 delete Claim 57 delete Claim 58 delete Claim 59 delete Claim 60 delete Claim 61 delete Claim 62 delete Claim 63 delete Claim 64 delete Claim 65 delete Claim 66 delete Claim 67 delete Claim 68 delete Claim 69 delete Claim 70 delete Claim 71 delete Claim 72 delete
Citation Information
Patent Citations
Nucleic acid detection methods using universal priming
EP1990428A1
Kinetic exclusion amplification of nucleic acid libraries
KR1020140140083A
Kinetic exclusion amplification of nucleic acid libraries
KR1020170083643A
Methods and compositions for nucleic acid library normalization
KR1020180043357A
Amplicon preparation and sequencing on solid supports
US20150197798A1