Methods of preparation of a next generation sequencing (NGS) library and NGS target enrichment
Truncated surface primer binding regions in NGS library preparation methods enable precise hybridization and amplification, addressing inefficiencies in target enrichment and enhancing sequencing outcomes.
Patent Information
- Application Number
- PCT/US2025/044160
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Existing next-generation sequencing (NGS) library preparation methods face inefficiencies in target enrichment due to indiscriminate capture by universal capture primers, leading to suboptimal hybridization and amplification processes.
The use of truncated surface primer binding regions in adapters that hybridize or anneal to surface primers at a lower temperature than traditional methods, allowing precise hybridization and annealing to target capture nucleic acids on a solid support, followed by temperature-controlled primer extension.
Enhances the specificity and efficiency of NGS library preparation by ensuring targeted hybridization and amplification, improving the quality of sequencing results.
Smart Images

Figure US2025044160_05032026_PF_FP_ABST
Abstract
Description
PATENTMETHODS OF PREPARATION OF A NEXT GENERATION SEQUENCING (NGS) LIBRARY AND NGS TARGET ENRICHMENTRELATED APPLICATION
[0001] This application claims priority from U.S. Provisional ApplicationNo. 63 / 688,594, filed August 29, 2024, the subject matter of which is incorporated herein by reference in its entirety.SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on August 29, 2025 is named BIO-033536WO ORD.st.26 and is 18,695 bytes in size.BACKGROUND
[0003] Next generation sequencing (NGS) has enabled whole genome sequencing and whole genome analysis. Next generation sequencing methods typically rely on the universal amplification of genomic fragments that are first equipped with universal amplification regions and then captured indiscriminately by universal capture primers on a solid surface. The universal capture primers mediate both polynucleotide capture and bridge amplification, a key element in next generation sequencing methods
[0004] In the process of library preparation for next generation sequencing (NGS), long genomic DNA is first fragmented to manageable sizes with a fragmentation enzyme or by sonication. Then a solution of DNA polymerase and dNTPs is used to repair the ends of the DNA fragments with an A sticky end. The DNA fragments are ligated with adapters to both ends. The adapters are composed of surface primer binding regions, indexes, and sequencing primers for the templates and indexes. After the library preparation, the library is denatured, loaded on a flow cell, bound to a surface modified with primers, amplified to form clusters of copies of the same DNA strands, and sequenced.SUMMARY
[0005] This disclosure describes methods of preparing a next generation sequencing (NGS) library that can be used with target enrichment methods, and, particularly, on-chip in- situ NGS target enrichment methods. During library preparation, such as for NGS, targetnucleic acid molecules, which are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support of a flow cell at a first annealing temperature (Tai), are ligated with adapters that include truncated surface primer binding regions instead full-length surface primer binding regions, which have nucleic acid sequences substantially complementary to the surface primers. The truncated surface primer binding regions are configured to hybridize or anneal respectively to first surface primers or second surface primers on the surface of the solid support at a second annealing temperature (Ta2) below or lower than the first annealing temperature (Tai) and not at the first annealing temperature (Tai). Library nucleic acid templates formed using the truncated surface primer binding regions do not hybridize with surface primers tethered to the solid support at a higher temperatures, such as Tai, yet can be primed by surface primers if the temperature is lowered to, for example Ta2.
[0006] In some embodiments, a next generation sequencing (NGS) library can include a plurality of single stranded nucleic acid templates that include target nucleic acid sequences ligated at 5’ and 3’ ends with adapters that include truncated surface primer binding regions. The single stranded nucleic acid templates are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support at a first annealing temperature (Tai). The truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal, respectively, to first surface primers or second surface primers on the surface of the solid support at a second annealing temperature (Ta2) lower than or below the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
[0007] In some embodiments, the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers. For example, the truncated surface primer binding regions can have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the nucleic acid sequences of the first surface primers or the second surface primers.
[0008] In some embodiments, the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first annealing temperature (Tai).
[0009] In some embodiments, the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support have a first melting temperature (Tml). The truncated surface primer binding regions when hybridized or annealed to the first surface primers or second surface primers have a second melting temperature (Tm2). The second melting temperature (Tm2) is lower than the first melting temperature (Tml ).
[0010] In some embodiments, the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0011] In some embodiments, the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is lower than Tml.
[0012] In some embodiments, the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0013] In some embodiments, a nucleic acid sequence complementary to full length nucleic acid sequence of the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third melting temperature (Tm3). The third melting temperature (Tm3) is higher than the second melting temperature (Tm2).
[0014] In some embodiments, ligated adapters further include sequencing primer binding domains and indexes.
[0015] In some embodiments, the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
[0016] Other embodiments described herein relate to a method of forming a NGS library. The method includes providing a plurality of fragmented subsets of double stranded nucleic acids. 3’ ends of the fragmented subsets of double stranded nucleic acids are end- repaired and A-tailed. Adapters are then ligated to the 5’ and 3’ ends of the double stranded nucleic acids. The ligated double stranded nucleic acids are denatured to form the library of single strand nucleic acid templates.
[0017] In some embodiments, the ligated adapters include truncated surface primer binding regions. The single stranded nucleic acid templates are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support at a first annealing temperature (Tai). The truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal respectively to first surface primers and second surface primers on the surface of the solid support at a second annealing temperature (Ta2) lower than or below the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
[0018] In some embodiments, the plurality of fragmented subsets of double stranded nucleic acids include fragmented subsets of double stranded DNA, preferably double stranded genomic DNA.
[0019] In some embodiments, the fragmented subsets of double stranded nucleic acids are from about 100 to about 1000 nucleic acids in length.
[0020] In some embodiments, the double stranded nucleic acids are end-repaired with polymerase and nucleotide triphosphates (NTPs) to form an end with one A at both 3’ ends.
[0021] In some embodiments, the ligated adapters further include sequencing primer binding domains and indexes.
[0022] In some embodiments, the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
[0023] In some embodiments, the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers. For example, the truncated surface primer binding regions can have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the nucleic acid sequence of the first surface primers or the second surface primers.
[0024] In some embodiments, the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first annealing temperature (Tai).
[0025] In some embodiments, the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support have a first melting temperature (Tml). The truncated surface primerbinding regions when hybridized or annealed to the first surface primers or second surface primers have a second melting temperature (Tm2). the second melting temperature (Tm2) is lower than the first melting temperature (Tml).
[0026] In some embodiments, the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tm l).
[0027] In some embodiments, the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is lower than Tml.
[0028] In some embodiments, the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0029] In some embodiments, a nucleic acid sequence complementary to full length nucleic acid sequence of the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers, has a third melting temperature (Tm3). The third melting temperature (Tm3) is higher than the second melting temperature (Tm2).
[0030] Still other embodiments described herein relate to a method of in-situ next generation sequencing (NGS) target enrichment. The method includes providing a solid support that includes a plurality of first surface capture primers and second surface capture primers tethered to a surface of the solid support. A plurality of nucleic acid target probes are provided that each include a linker and a target capture nucleic acid that hybridizes to and captures a target nucleic acid sequence of a library of single stranded nucleic acid templates at a first annealing temperature (Tai). A plurality of single stranded nucleic acid templates is provided that include target nucleic acid sequences ligated at 5’ and 3’ ends with adapters that include truncated surface primer binding regions. The single stranded nucleic acid templates are configured to hybridize or anneal to the target capture nucleic acids of the nucleic acid target probes on the surface of the solid support at the first annealing temperature (Tai). The truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal respectively to the first surface primers or the second surface primers onthe surface of the solid support at a second annealing temperature (Ta2) lower than the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
[0031] In some embodiments, the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers. For example, the truncated surface primer binding regions have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the nucleic acid sequences of the first surface primers or the second surface primers.
[0032] In some embodiments, the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first annealing temperature (Tai).
[0033] In some embodiments, the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support have a first melting temperature (Tml). The truncated surface primer binding regions when hybridized or annealed to the first surface primers or second surface primers have a second melting temperature (Tm2). The second melting temperature (Tm2) is lower than the first melting temperature (Tml).
[0034] In some embodiments, the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0035] In some embodiments, the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is lower than Tml.
[0036] In some embodiments, the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0037] In some embodiments, a nucleic acid sequence complementary to full length nucleic acid sequence of the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third meltingtemperature (Tm3). The third melting temperature (Tm3) is higher than the second melting temperature (Tm2).
[0038] In some embodiments, the method further includes tethering or linking the plurality of nucleic acid target probes to the surface of solid support. The library is loaded onto the solid support at or below the first melting temperature (Tml) and above the second melting temperature (Tm2) such that target nucleic acid sequences of the single stranded nucleic acid templates hybridize and are captured by the probes and the truncated surface binding regions do not bind to the first surface capture primers and second surface capture primers. The temperature on the solid support is lowered to or below the second melting temperature (Tm2) such that one of truncated surface primer binding regions of the captured single stranded nucleic acid templates hybridizes to the first surface capture primer or second surface capture primer. The surface capture primers with captured single stranded nucleic acid templates are extended to form a plurality of complementary templates tethered to the surface and hybridized to the plurality of single stranded capture nucleic acid templates.
[0039] In some embodiments, captured single stranded nucleic acid templates are removed from the plurality of complementary templates tethered to the surface by, for example, denaturing.
[0040] In some embodiments, the denatured captured single stranded nucleic acid are removed from the solid support by washing.
[0041] In some embodiments, the method further includes amplifying the plurality of complementary templates tethered to provide amplicons or clusters of the complementary templates on the solid support.
[0042] In some embodiments, the method further includes sequencing the clusters of the complementary templates.
[0043] In some embodiments, single stranded nucleic acid templates not captured by the nucleic acid target probes after loading on the solid support are removed from the solid support by, for example, washing the solid support.
[0044] In some embodiments, 3’ ends of the nucleic acid target probes are blocked to prevent extension during extending of the surface capture primers.
[0045] In some embodiments, the nucleic acid target probes include a nucleic acid linker with a 3’ end ligated to a 5’ end of the target capture nucleic acid.
[0046] In some embodiments, the plurality of tethered complementary templates are amplified by bridge amplification.
[0047] In some embodiments, the library of nucleic acid strand templates is provided by providing a plurality of fragmented subsets of double stranded nucleic acids. 3’ ends of the fragmented subsets of double stranded nucleic acids are end-repaired and A-tailed. Adapters are ligated to the 5’ and 3’ ends of the double stranded nucleic acids. The adapters include the truncated surface primer binding regions.
[0048] In some embodiments, the method further includes denaturing the ligated double stranded nucleic acids to form the library of single strand nucleic acid templates and loading the single strand nucleic acid templates onto the surface of the solid support of a flow cell.
[0049] In some embodiments, the plurality of fragmented subsets of double stranded nucleic acids include fragmented subsets of double stranded DNA, preferably double stranded genomic DNA.
[0050] In some embodiments, the fragmented subsets of double stranded nucleic acids are from about 100 to about 1000 nucleic acids in length.
[0051] In some embodiments, the double stranded nucleic acids are end-repaired with polymerase and nucleotide triphosphates (NTPs) to form an end with one A at both 3’ ends.
[0052] In some embodiments, the ligated adapters further include sequencing primer binding domains and indexes.
[0053] In some embodiments, the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
[0054] In some embodiments, an end of each of the nucleic acid target probes is tethered to the surface of the solid support prior to loading the library of single stranded nucleic acid templates onto the solid support.
[0055] In some embodiments, an end of each of the nucleic acid target probes includes a surface primer binding region that hybridizes to one of the surface capture primers.
[0056] In some embodiments, a 3’ end of each of the surface primer binding regions of the nucleic acid target probes is connected to the linker via blocking region that prevents extension of the nucleic acid target probe during extension of the surface capture primers.
[0057] In some embodiments, the plurality of nucleic acid targeting probes are loaded on the solid support along with the library of single stranded nucleic acids.
[0058] In some embodiments, the nucleic target probes hybridized to surface capture primers are removed from the solid support along with the captured single stranded nucleic acid templates by, for example, denaturing the nucleic target probes hybridized to surface capture primers and washing the denatured nucleic acid target probes from the solid support.BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Fig. 1 illustrates a schematic of steps in the preparation of the NGS library to be used with a target capturing method described herein.
[0060] Fig. 2 illustrates a schematic of a process of on-chip in-situ NGS target enrichment.
[0061] Fig. 3 illustrates a schematic of an alternate method of loading target probes on a flow cell.DETAILED DESCRIPTION
[0062] Terms used herein will be understood to take on their ordinary meaning in the relevant art unless specified otherwise. Several terms used herein and their meanings are set forth below.
[0063] As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a biomarker” includes a mixture of two or more biomarkers, and the like.
[0064] The term “about,” particularly in reference to a given quantity, is meant to encompass deviations of plus or minus five percent.
[0065] The terms “includes,” “including,” “includes,” “including,” “contains,” “containing,” and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, product-by-process, or composition of matter that includes, includes, or contains an element or list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, product- by-process, or composition of matter.
[0066] The term “adapter” refers to linear oligonucleotide sequence that can be fused to a nucleic acid molecule, for example, by ligation. Suitable adapter lengths may range from about 10 nucleotides to about 100 nucleotides, or from about 12 nucleotides to about 60nucleotides, or from about 15 nucleotides to about 50 nucleotides. The adapter may include any combination of nucleotides and / or nucleic acids. In some examples, the adapter can include a sequence that is complementary to at least a portion of a primer, for example, a primer including a universal nucleotide sequence (such as a P5 or P7 sequence). As an example, the adapter at one end of a fragment includes a sequence that is complementary to at least a portion of a first flow cell primer, and the adapter at the other end of the fragment includes a sequence that is identical to at least a portion of a second flow cell primer. The complementary adapter can hybridize to the first flow cell primer, and the identical adapter is a template for its complementary copy, which can hybridize to the second flow cell primer during clustering. In some examples, the adapter can include a sequencing primer sequence or a sequencing binding site. Combinations of different adapters may be incorporated into a nucleic acid molecule, such as a DNA fragment.
[0067] The term “amplicon,” when used in reference to a nucleic acid, means the product of copying the nucleic acid, wherein the product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be produced by any of a variety of amplification methods that use the nucleic acid, or an amplicon thereof, as a template including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of the nucleotide sequence. A first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies that are created, after generation of the first amplicon, from the target nucleic acid or from the first amplicon. A subsequent amplicon can have a sequence that is substantially complementary to the target nucleic acid or substantially identical to the target nucleic acid.
[0068] The term “array” refers to a population of features or sites that can be differentiated from each other according to relative location. Different molecules that are at different sites of an array can be differentiated from each other according to the locations of the sites in the array. An individual site of an array can include one or more molecules of a particular type. For example, a site can include a single nucleic acid molecule having a particular sequence or a site can include several nucleic acid molecules having the same sequence (and / or complementary sequence, thereof). The sites of an array can be differentfeatures located on the same substrate. Exemplary features include without limitation, wells in a substrate, beads (or other particles) in or on a substrate, projections from a substrate, ridges on a substrate or channels in a substrate. The sites of an array can be separate substrates each bearing a different molecule. Different molecules attached to separate substrates can be identified according to the locations of the substrates on a surface to which the substrates are associated or according to the locations of the substrates in a liquid or gel.
[0069] The term “attached” refers to the state of two things being joined, fastened, adhered, connected or bound to each other. For example, an analyte, such as a nucleic acid, can be attached to a material, such as a solid support, by a covalent or non-covalent bond. A covalent bond is characterized by the sharing of pairs of electrons between atoms. A non- covalent bond is a chemical bond that does not involve the sharing of pairs of electrons and can include, for example, hydrogen bonds, ionic bonds, van der Waals forces, hydrophilic interactions and hydrophobic interactions.
[0070] The term "capture primers" is intended to mean an oligonucleotide having a nucleotide sequence that is capable of specifically annealing to a single stranded polynucleotide sequence to be analyzed or subjected to a nucleic acid interrogation under conditions encountered in a primer annealing step of, for example, an amplification or sequencing reaction.
[0071] The term “index”, “index sequence”, or “barcode sequence” are used interchangeably and are intended to refer to a series of nucleotides in a nucleic acid that can be used to identify the nucleic acid, a characteristic of the nucleic acid, or a manipulation that has been carried out on the nucleic acid. The index or barcode sequence can be a naturally occurring sequence or a sequence that does not occur naturally in the organism from which the barcoded nucleic acid was obtained. An index or barcode sequence can be unique to a single nucleic acid species in a population or a barcode sequence can be shared by several different nucleic acid species in a population. For example, each nucleic acid probe in a population can include different index or barcode sequences from all other nucleic acid probes in the population. Alternatively, each nucleic acid probe in a population can include different index or barcode sequences from some or most other nucleic acid probes in a population. For example, each probe in a population can have an index or barcode that is present for several different probes in the population even though the probes with the common barcode differ from each other at other sequence regions along their length. Inparticular embodiments, one or more barcode sequences that are used with a biological sample are not present in the genome, transcriptome or other nucleic acids of the biological sample. For example, index or barcode sequences can have less than 80%, 70%, 60%, 50% or 40% sequence identity to the nucleic acid sequences in a particular biological sample.
[0072] The term “biological sample” is intended to mean one or more cells, tissues, organisms, or portions thereof. A biological sample can be obtained from any of a variety of organisms. Exemplary organisms include, but are not limited to, a mammal, such as a rodent, mouse, rat, rabbit, guinea pig, ungulate, horse, sheep, pig, goat, cow, cat, dog, primate (i.e., human or non-human primate); a plant; an algae; a nematode; an insect, such as mosquito, fruit fly, honey bee or spider; a fish such as zebrafish; a reptile; an amphibian such as a frog; a fungi, or yeast. Target nucleic acids can also be derived from a prokaryote such as a bacterium; an archaea; a virus, such as Hepatitis C virus or human immunodeficiency virus; or a viroid. Samples can be derived from a homogeneous culture or population of the above organisms or alternatively from a collection of several different organisms, for example, in a community or ecosystem.
[0073] The term “cluster,” when used in reference to nucleic acids, refers to a population of the nucleic acids that is attached to a solid support to form a feature or site. The nucleic acids are generally members of a single species, thereby forming a monoclonal cluster. A “monoclonal population” of nucleic acids is a population that is homogeneous with respect to a particular nucleotide sequence. Clusters need not be monoclonal. Rather, for some applications, a cluster can be predominantly populated with amplicons from a first nucleic acid and can also have a low level of contaminating amplicons from a second nucleic acid. For example, when an array of clusters is to be used in a detection application, an acceptable level of contamination would be a level that does not impact signal to noise or resolution of the detection technique in an unacceptable way. Accordingly, apparent clonality will generally be relevant to a particular use or application of an array made by the methods set forth herein. Exemplary levels of contamination that can be acceptable at an individual cluster include, but are not limited to, at most 0.1%, 0.5%, 1 %, 5%, 10%, 5 25%, or 35% contaminating amplicons. The nucleic acids in a cluster are generally covalently attached to a solid support, for example, via their 5' ends, but in some cases other attachment means are possible. The nucleic acids in a cluster can be single stranded or double stranded. In some but not all embodiments, clusters are made by a solid-phase amplification methodknown as bridge amplification. Exemplary configurations for clusters and methods for their production are set forth, for example, in U.S. Pat. No. 5,641,658; U.S. Patent Publ. No. 2002 / 0055100; U.S. Pat. No. 7,115,400; U.S. Patent Publ. No. 2004 / 0096853; U.S. Patent Publ. No. 2004 / 0002090; U.S. Patent Publ. No. 2007 / 0128624; and U.S. Patent Publ. No. 2008 / 0009420, each of which is incorporated herein by reference.
[0074] The term “different”, when used in reference to nucleic acids, means that the nucleic acids have nucleotide sequences that are not the same as each other. Two or more nucleic acids can have nucleotide sequences that are different along their entire length. Alternatively, two or more nucleic acids can have nucleotide sequences that are different along a substantial portion of their length. For example, two or more nucleic acids can have target nucleotide sequence portions that are different for the two or more molecules while also having a universal sequence portion that is the same on the two or more molecules.
[0075] The term “each,” when used in reference to a collection of items, is intended to identify an individual item in the collection but does not necessarily refer to every item in the collection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise.
[0076] The term “extend,” when used in reference to a nucleic acid, is intended to mean addition of at least one nucleotide or oligonucleotide to the nucleic acid. In particular embodiments one or more nucleotides can be added to the 3' end of a nucleic acid, for example, via polymerase catalysis (e.g., DNA polymerase, RNA polymerase or reverse transcriptase). Chemical or enzymatic methods can be used to add one or more nucleotide to the 3' or 5' end of a nucleic acid. One or more oligonucleotides can be added to the 3' or 5’ end of a nucleic acid, for example, via chemical or enzymatic (e.g., ligase catalysis) methods. A nucleic acid can be extended in a template directed manner, whereby the product of extension is complementary to a template nucleic acid that is hybridized to the nucleic acid that is extended.
[0077] The term “flow cell” is intended to mean a vessel having a chamber where a reaction can be carried out, an inlet for delivering reagents to the chamber and an outlet for removing reagents from the chamber. In some embodiments the chamber is configured for detection of the reaction that occurs in the chamber. For example, the chamber can include one or more transparent surfaces allowing optical detection of biological samples, optically labeled molecules, or the like in the chamber. Exemplary flow cells include, but are not limited to those used in a nucleic acid sequencing apparatus such as flow cells for theGenome Analyzer, MiSeq, NextSeq or HiSeq platforms commercialized by Illumina, Inc. (San Diego, Calif.); or for the SOLiD or Ion Torrent sequencing platform commercialized by Life Technologies (Carlsbad, Calif.). Exemplary flow cells and methods for their manufacture and use are also described, for example, in WO 2014 / 142841 Al; U.S. Pat. App. Pub. No. 2010 / 0111768 Al and U.S. Pat. No. 8,951,781, each of which is incorporated herein by reference.
[0078] The terms “nucleic acid” and “nucleotide” are intended to be consistent with their use in the art and to include naturally occurring species or functional analogs thereof. Particularly useful functional analogs of nucleic acids are capable of hybridizing to a nucleic acid in a sequence specific fashion or capable of being used as a template for replication of a particular nucleotide sequence. Naturally occurring nucleic acids generally have a backbone containing phosphodiester bonds. An analog structure can have an alternate backbone linkage including any of a variety of those known in the art. Naturally occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA)) or a ribose sugar (e.g., found in ribonucleic acid (RNA)). A nucleic acid can contain nucleotides having any of a variety of analogs of these sugar moieties that are known in the art. A nucleic acid can include native or non-native nucleotides. In this regard, a native deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine, thymine, cytosine or guanine and a ribonucleic acid can have one or more bases selected from the group consisting of uracil, adenine, cytosine or guanine. Useful non-native bases that can be included in a nucleic acid or nucleotide are known in the art.
[0079] The term "plurality" refers to a population of two or more members, such as polynucleotide members or other referenced molecules. In some embodiments, the two or more members of a plurality of members are the same members. For example, a plurality of polynucleotides can include two or more polynucleotide members having the same nucleic acid sequence. In some embodiments, the two or more members of a plurality of members are different members. For example, a plurality of polynucleotides can include two or more polynucleotide members having different nucleic acid sequences. A plurality includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or a 100 or more different members. A plurality can also include 200, 300, 400, 500, 1000, 5000, 10000, 50000, 1x10s, 2xl05, 3x10s, 4x10s, 5x10s, 6x10s, 7x10s, 8 xlO5, 9x10s, IxlO6, 2xl06, 3xl06, 4xl06, 5xl06, 6xl06, 7xl06, 8xl06, 9xl06orIxlO7or more different members. A plurality includes all integer numbers in between the above exemplary plurality numbers.
[0080] The terms “probe” or “target,” when used in reference to a nucleic acid or sequence of a nucleic acid, are intended as semantic identifiers for the nucleic acid or sequence in the context of a method or composition set forth herein and does not necessarily limit the structure or function of the nucleic acid or sequence beyond what is otherwise explicitly indicated.
[0081] The term “poly T or poly A,” when used in reference to a nucleic acid sequence, is intended to mean a series of two or more thiamine (T) or adenine (A) bases, respectively. A poly T or poly A can include at least about 2, 5, 8, 10, 12, 15, 18, 20 or more of the T or A bases, respectively. Alternatively, or additionally, a poly T or poly A can include at most about, 30, 20, 18, 15, 12, 10, 8, 5 or 2 of the T or A bases, respectively.
[0082] The term “random” can be used to refer to the spatial arrangement or composition of locations on a surface. For example, there are at least two types of order for an array described herein, the first relating to the spacing and relative location of features (also called “sites”) and the second relating to identity or predetermined knowledge of the particular species of molecule that is present at a particular feature. Accordingly, features of an array can be randomly spaced such that nearest neighbor features have variable spacing between each other. Alternatively, the spacing between features can be ordered, for example, forming a regular pattern such as a rectilinear grid or hexagonal grid. In another respect, features of an array can be random with respect to the identity or predetermined knowledge of the species of analyte (e.g., nucleic acid of a particular sequence) that occupies each feature independent of whether spacing produces a random pattern or ordered pattern. An array set forth herein can be ordered in one respect and random in another. For example, in some embodiments set forth herein a surface is contacted with a population of nucleic acids under conditions where the nucleic acids attach at sites that are ordered with respect to their relative locations but ‘randomly located’ with respect to knowledge of the sequence for the nucleic acid species present at any particular site. Reference to “randomly distributing” nucleic acids at locations on a surface is intended to refer to the absence of knowledge or absence of predetermination regarding which nucleic acid will be captured at which location (regardless of whether the locations are arranged in an ordered pattern or not).
[0083] The term “solid support” refers to a rigid substrate that is insoluble in aqueous liquid. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g., due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon?, cyclic olefins, polyimides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solid supports for some embodiments are located within a flow cell apparatus.
[0084] The term "target polynucleotide" or “target nucleic acid” is intended to mean a polynucleotide that is the object of an analysis or action. The analysis or action includes subjecting the polynucleotide to copying, amplification, sequencing and / or other procedure for nucleic acid interrogation. A target polynucleotide can include nucleotide sequences additional to the target sequence to be analyzed. For example, a target polynucleotide can include one or more adapters, including an adapter that functions as a primer binding site, that flank(s) a target polynucleotide sequence that is to be analyzed. A target polynucleotide hybridized to a capture oligonucleotide or capture primer can contain nucleotides that extend beyond the 5' or 3' end of the capture oligonucleotide in such a way that not all of the target polynucleotide is amenable to extension. In particular embodiments, as set forth in further detail below, a plurality of target polynucleotides includes different species that differ in their target polynucleotide sequences but have adapters that are the same for two or more of the different species. The two adapters that can flank a particular target polynucleotide sequence can have the same sequence or the two adapters can have different sequences. Accordingly, a plurality of different target polynucleotides can have the same adapter sequence or two different adapter sequences at each end of the target polynucleotide sequence. Thus, species in a plurality of target polynucleotides can include regions of known sequence that flank regions of unknown sequence that are to be evaluated by, for example, sequencing. In cases where the target polynucleotides carry an adapter at a single end, the adapter can be located at either the 3' end or the 5' end the target polynucleotide. Target polynucleotides can be usedwithout any adapter, in which case a primer binding sequence can come directly from a sequence found in the target polynucleotide.
[0085] The term "target specific" when used in reference to a capture primer or other oligonucleotide is intended to mean a capture primer or other oligonucleotide that includes a nucleotide sequence specific to a target polynucleotide sequence, namely a sequence of nucleotides capable of selectively annealing to an identifying region of a target polynucleotide. Target specific capture primers can have a single species of oligonucleotide, or it can include two or more species with different sequences. Thus, the target specific capture primers can be two or more sequences, including 3, 4, 5, 6, 7, 8, 9 or 10 or more different sequences. The target specific capture oligonucleotides can include a target specific capture primer sequence and universal capture primer sequence. Other sequences such as sequencing primer sequences and the like also can be included in a target specific capture primer.
[0086] In comparison, the term "universal" when used in reference to a capture primer or other oligonucleotide sequence is intended to mean a capture primer or other oligonucleotide having a common nucleotide sequence among a plurality of capture primers. A common sequence can be, for example, a sequence complementary to the same adapter sequence. Universal capture primers are applicable for interrogating a plurality of different polynucleotides without necessarily distinguishing the different species whereas target specific capture primers are applicable for distinguishing the different species.
[0087] The term “universal sequence” refers to a series of nucleotides that is common to two or more nucleic acid molecules even if the molecules also have regions of sequence that differ from each other. A universal sequence that is present in different members of a collection of molecules can allow capture of multiple different nucleic acids using a population of universal capture nucleic acids that are complementary to the universal sequence. Similarly, a universal sequence present in different members of a collection of molecules can allow the replication or amplification of multiple different nucleic acids using a population of universal primers that are complementary to the universal sequence. Thus, a universal capture nucleic acid or a universal primer includes a sequence that can hybridize specifically to a universal sequence. Target nucleic acid molecules may be modified to attach universal adapters, for example, at one or both ends of the different target sequences.
[0088] This disclosure describes methods of preparing a next generation sequencing (NGS) library that can be used with target enrichment methods, and, particularly, on-chip in- situ NGS target enrichment methods. During library preparation, such as for NGS, target nucleic acid molecules, which are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support of a flow cell at a first annealing temperature (Tai), are ligated with adapters that include truncated surface primer binding regions instead full-length surface primer binding regions, which have nucleic acid sequences substantially complementary to the surface primers. The truncated surface primer binding regions are configured to hybridize or anneal respectively to first surface primers or second surface primers on the surface of the solid support at a second annealing temperature (Ta2) below or lower than the first annealing temperature (Tai) and not at the first annealing temperature (Tai). Library nucleic acid templates formed using the truncated surface primer binding regions do not hybridize with surface primers tethered to the solid support at a higher temperatures, such as Tai, yet can be primed by surface primers if the temperature is lowered to, for example Ta2.
[0089] Referring to Fig. 1 , a library of single stranded nucleic acid templates for next generation sequencing (NGS) can be generated by providing a plurality of fragmented subsets of double stranded nucleic acids. The plurality of fragmented subsets of double stranded nucleic acids used to form the library of nucleic acid templates can be obtained from essentially any nucleic acid sample of known or unknown sequence. It may be, for example, a fragment of genomic DNA or cDNA. The nucleic acids can be derived from a primary nucleic acid sample that has been randomly fragmented.
[0090] The primary nucleic acid sample may originate in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, PCR and amplification products and the like) from a sample or may originate in single-stranded form from a sample, as DNA or RNA, and been converted to dsDNA form. By way of example, mRNA molecules may be copied into double-stranded cDNAs suitable for use in a method described herein using standard techniques well known in the art. The precise sequence of the nucleotide molecules from a primary nucleic acid sample is generally not material to the disclosure, and may be known or unknown.
[0091] In one embodiment, the polynucleotide molecules from a nucleic acid sample are DNA molecules. More particularly, the polynucleotide molecules represent the entiregenetic complement of an organism, and are genomic DNA molecules, which include both intron and exon sequences, as well as non-coding regulatory sequences such as promoter and enhancer sequences. In one embodiment, particular subsets of polynucleotide sequences or genomic DNA can be used, such as, for example, particular chromosomes. Yet more particularly, the sequence of the polynucleotide molecules is not known. Still yet more particularly, the polynucleotide molecules are human genomic DNA molecules. The DNA fragments may be treated chemically or enzymatically either prior or subsequent to any random fragmentation processes, and prior or subsequent to the ligation of the adapter sequences.
[0092] In some embodiments, the nucleic acid sample can include a high molecular weight material, such as genomic DNA (gDNA). The sample can also include a low molecular weight material, such as nucleic acid molecules obtained from formalin-fixed paraffin-embedded or archived DNA samples. In another embodiment, the low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, the sample can include nucleic acid molecules obtained from an animal, such as a human or mammalian source. In another embodiment, the sample can include nucleic acid molecules obtained from a nonmammalian source, such as a plant, a bacterium, a virus, or a fungus. In some embodiments, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0093] Further, the nucleic acid sample can have low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, from a missing person's DNA database, from a laboratory associated with a forensic investigation, or from forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example, derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may include low amounts of, or fragmented portionsof DNA, such as genomic DNA. In some embodiments, nucleic acid sequences can be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, nucleic acid sequences can be obtained from hair, skin, tissue samples, autopsy, or remains of a victim. In some embodiments, nucleic acids can be obtained from a deceased animal or human. In some embodiments, nucleic acid sequences can include nucleic acids obtained from non-human DNA, such as a microbial, plant or entomological DNA. In some embodiments, nucleic acid sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, a method described herein can be used for identifying characteristics of a forensic sample.
[0094] Examples of biological samples from which nucleic acids can be derived include, for example, those from a eukaryote, for instance a mammal, such as a rodent, mouse, rat, rabbit, guinea pig, ungulate, horse, sheep, pig, goat, cow, cat, dog, primate, human or non-human primate; a plant, such as Arabidopsis thaliana, corn, sorghum, oat, wheat, rice, canola, or soybean; an algae, such as Chlamydomonas reinhardtii', a nematode such as Caenorhabditis elegans; an insect, such as Drosophila melanogaster, mosquito, fruit fly, honey bee or spider; a fish, such as zebrafish; a reptile; an amphibian, such as a frog or Xenopus laevis', a Dictyostelium discoideunr, a fungi, such as Pneumocystis carinii, Takifugu rubripes, yeast, such as Saccharamoyces cerevisiae or Schizosaccharomyces pombe', or Plasmodium falciparum. Target nucleic acids can also be derived from a prokaryote such as a bacterium, Escherichia coli, staphylococci or Mycoplasma pneumoniae', an archaeon; a virus such as Hepatitis C virus or human immunodeficiency virus; or a viroid. Target nucleic acids can be derived from a homogeneous culture or population of organisms or alternatively from a collection of several different organisms, for example, in a community or ecosystem.
[0095] In some embodiments, the plurality of fragmented subsets of double stranded nucleic acids can be randomly fragmented. Random fragmentation refers to the fragmentation of a polynucleotide molecule from a primary nucleic acid sample in a non-ordered fashion by enzymatic, chemical, or mechanical methods. Such fragmentation methods are known in the art and use standard methods (Sambrook and Russell, Molecular Cloning, A Laboratory Manual, third edition). For the sake of clarity, generating smaller fragments of a larger piece of nucleic acid via specific PCR amplification of such smaller fragments is not equivalent tofragmenting the larger piece of nucleic acid because the larger piece of nucleic acid sequence remains in intact (i.e., is not fragmented by the PCR amplification). Moreover, random fragmentation is designed to produce fragments irrespective of the sequence identity or position of nucleotides comprising and / or surrounding the break. More particularly, the random fragmentation is by mechanical means such as nebulization or sonication to produce fragments of about 50 base pairs in length to about 1500 base pairs in length, still more particularly 50-700 base pairs in length, yet more particularly 50-400 base pairs in length. Most particularly, the method is used to generate smaller fragments of from 50-150 base pairs in length
[0096] In some embodiments, a population of fragmented or unfragmented nucleic acids can have an average strand length that is desired or appropriate for a particular application of the methods or compositions set forth herein. For example, the average strand length can be less than about 100,000 nucleotides, less than about 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively or additionally, the average strand length can be greater than about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average strand length for a population of nucleic acids can be in a range between a maximum and minimum value set forth herein.
[0097] In some embodiments, a population of nucleic acids can be produced under conditions or otherwise configured to have a maximum length for its members. For example, the maximum length for the members that are used in one or more steps of a method set forth herein or that are present in a particular composition can be less than 100,000 nucleotides, less than 50,000 nucleotides, less than 10,000 nucleotides, less than 5,000 nucleotides, less than 1,000 nucleotides, less than 500 nucleotides, less than 100 nucleotides, or less than 50 nucleotides. Alternatively or additionally, a population of nucleic acids can be produced under conditions or otherwise configured to have a minimum length for its members. For example, the minimum length for the members that are used in one or more steps of a method set forth herein or that are present in a particular composition can be more than 10 nucleotides, more than 50 nucleotides, more than 100 nucleotides, more than 500 nucleotides, more than 1,000 nucleotides, more than 5,000 nucleotides, more than 10,000 nucleotides, more than 50,000 nucleotides, or more than 100,000 nucleotides. The maximum andminimum strand length for nucleic acids in a population can be in a range between a maximum and minimum value set forth above.
[0098] In some embodiments, the plurality of fragmented subsets of double stranded nucleic acids include fragmented subsets of double stranded DNA. The fragmented subsets of double stranded nucleic acids can be from about 50 to about 10,000 nucleic acids in length, for example, about 50 to about 5,000 nucleic acids in length, about 50 to about 4,000 nucleic acids in length, about 50 to about 3,000 nucleic acids in length, about 50 to about 2,000 nucleic acids in length, about 50 to about 1 ,000 nucleic acids in length, about 50 to about 900 nucleic acids in length, about 50 to about 800 nucleic acids in length, about 50 to about700 nucleic acids in length, about 50 to about 600 nucleic acids in length, about 50 to about500 nucleic acids in length, about 50 to about 400 nucleic acids in length, about 50 to about300 nucleic acids in length, about 50 to about 200 nucleic acids in length, about 100 to about10,000 nucleic acids in length, about 200 to about 10,000 nucleic acids in length, about 300 to about 10,000 nucleic acids in length, about 400 to about 10,000 nucleic acids in length, about 500 to about 10,000 nucleic acids in length, about 600 to about 10,000 nucleic acids in length, about 700 to about 10,000 nucleic acids in length, about 800 to about 10,000 nucleic acids in length, about 900 to about 10,000 nucleic acids in length, about 1000 to about 10,000 nucleic acids in length, about 2000 to about 10,000 nucleic acids in length, about 3000 to about 10,000 nucleic acids in length, about 4000 to about 10,000 nucleic acids in length, or about 5000 to about 10,000 nucleic acids in length.
[0099] In some embodiments, the nucleic acids that are derived from such sources can be amplified prior to use in a method or composition herein. Any of a variety of known amplification techniques can be used including, but not limited to, polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), or random prime amplification (RPA). It will be understood that amplification of nucleic acids prior to use in a method or composition set forth herein is optional. As such, nucleic acids will not be amplified prior to use in some embodiments of the methods and compositions set forth herein. Nucleic acids can optionally be derived from synthetic libraries. Synthetic nucleic acids can have native DNA or RNA compositions or can be analogs thereof.
[0100] In some embodiments, fragmentation of polynucleotide molecules by mechanical means (e.g., nebulization, sonication, and Hydroshear) can result in fragments with a heterogeneous mix of blunt and 3'- and 5 '-overhanging ends.
[0101] The fragmented double stranded nucleic acids, which are provided, are end- repaired with polymerase and nucleotide triphosphates (NTPs) to form an end with one A at both 3’ ends. In a particular embodiment, the fragmented double stranded nucleic acids are prepared with single overhanging nucleotides by, for example, activity of certain types of DNA polymerase, such as Taq polymerase or Klenow exo minus polymerase, which has a non-template-dependent terminal transferase activity that adds a single deoxynucleotide, for example, deoxy adenosine (A), to the 3' ends of a DNA molecule. Such enzymes can be used to add a single nucleotide ‘A’ to the blunt ended 3' terminus of each strand of the doublestranded target fragments. Thus, an ‘A’ could be added to the 3' terminus of each end repaired strand of the double-stranded target fragments by reaction with Taq or Klenow exo minus polymerase
[0102] The end-repaired and A-tailed subsets of double stranded nucleic acids are ligated with adapters at the 5’ and 3’ ends of the double stranded nucleic acid. The adapters can include sets or pairs of truncated surface primer binding regions on opposite flanking 5’ and 3’ ends of the adapter ligated strands of nucleic acids that are complementary to and / or can hybridize or anneal to oligonucleotide surface primer pairs provided on a solid support surface of a microarray of a microfluidic device, such as a flow cell. The truncated surface primer binding regions at the 5 ’ and 3 ’ ends are configured to hybridize or anneal respectively to first surface primers or second surface primers on the surface of the solid support at a second annealing temperature (Ta2). The second annealing temperature is below or lower than the first annealing temperature (Tai), which is the temperature at which the single stranded nucleic acid templates are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support.
[0103] In some embodiments, the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first annealing temperature (Tai).
[0104] In some embodiments, the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support have a first melting temperature (Tml). The truncated surface primer binding regions when hybridized or annealed to the first surface primers or second surfaceprimers have a second melting temperature (Tm2). The second melting temperature (Tm2) is lower than the first melting temperature (Tml).
[0105] In some embodiments, the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0106] In some embodiments, the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is lower than Tml.
[0107] In some embodiments, the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C lower than the first melting temperature (Tml).
[0108] In some embodiments, a nucleic acid sequence complementary to full length nucleic acid sequence of the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third melting temperature (Tm3). The third melting temperature (Tm3) is lower than the second melting temperature (Tm2).
[0109] In some embodiments, the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers. For example, the truncated surface primer binding regions can have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the nucleic acid sequences of the first surface primers or the second surface primers.
[0110] In some embodiments, the truncated surface primer binding regions can have a nucleic acid sequence that can hybridize with an Illumina® capture primer P5 (5'- AATGATACGGCGACCACCGA-3') (SEQ ID NO: 1) or P7 (5'- CAAGCAGAAGACGGCATACGA-3') (SEQ ID NO: 2). In certain embodiments, the truncated surface primer binding regions are complementary to a portion of the Illumina® capture primer P5 or P7. For example, the truncated surface primer binding regions have a complementary nucleic acid sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than nucleic acid sequences of P5 or P7. In one embodiment, a truncated anti-P5 surface primer binding region, i.e., “P5-trunc”, can have a nucleic acid sequence of5'- CGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 3), 5'- GGTGGTCGCCGTATCATT-3' (SEQ ID NO: 4), 5'- GTGGTCGCCGTATCATT-3' (SEQ ID NO: 5), or 5'- TGGTCGCCGTATCATT-3' (SEQ ID NO: 6). In another embodiment, a truncated anti-P7 surface primer binding region, i.e., “P7-trunc”, can have a nucleic acid sequence of 5'- CGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 7), 5'-GTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 8), 5'- TATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 9), or 5'- ATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 10).
[0111] In other embodiments, the truncated surface primer binding regions can have a nucleic acid sequence that can hybridize with Illumina® capture primers P5 (paired end) (5'- AATGATACGGCGACCACCGAGAUCTACAC-3') (SEQ ID NO: 11) or P7(paired end) (5'-CAAGCAGAAGACGGCATACGA(8-oxo-G)AT-3') (SEQ ID NO: 12). For example, the truncated surface primer binding regions have a complementary nucleic acid sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than nucleic acid sequences of P5 (paired end) or P7 (paired end). For example, the truncated surface primer binding regions have a complementary nucleic acid sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than nucleic acid sequences of P5 (paired end) or P7 (paired end). In one embodiment, a truncated anti-P5 surface primer binding region, i.e., “P5-trunc”, can have a nucleic acid sequence of 5'- TGTAGATCTCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 13), 5'- GTAGATCTCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 14), 5'- TAGATCTCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 15), or 5'- AGATCTCGGTGGTCGCCGTATCATT-3' (SEQ ID NO: 16). In another embodiment, a truncated anti-P7 surface primer binding region, i.e., “P7-trunc”, can have a nucleic acid sequence of 5'-TCTCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 17), 5'- CTCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 18), 5'- TCGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 19), or 5'- CGTATGCCGTCTTCTGCTTG-3' (SEQ ID NO: 20).
[0112] In some embodiments, the adapters can further include at least one sequencing primer binding domain or universal sequencing primer binding site. A universal sequencing primer binding domain is a universal sequence that can be used for amplification and / or sequencing of a target or template nucleic acid ligated to the adapter. In some embodiments, the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
[0113] The adapters can further include index or barcode sequences. An index sequence can be used as a marker characteristic of the source of particular nucleic acids on an array. Generally, the index is a synthetic sequence of nucleotides that is part of the adapter, which is attached to each nucleic acid strand of the double stranded nucleic acid, and the presence of which is indicative of, or is used to identify the nucleic acid strand or spatial location of the nucleic strand on a flow cell surface.
[0114] In some embodiments, an index may be up to 20 nucleotides in length, more preferably 1-10 nucleotides, and most preferably 4-6 nucleotides in length. A four nucleotide index gives a possibility of multiplexing 256 samples on the same array, a six base index enables 4096 samples to be processed on the same array.
[0115] Methods for ligating an universal adapter to each end of a target or template nucleic acid used in a method described herein are known to the person skilled in the art. Such methods use ligase enzymes such as DNA ligase to effect or catalyze joining of the ends of the two polynucleotide strands of, in this case, the adapters and the double-stranded target nucleic acids, such that covalent linkages are formed. The adapters may contain a 5'- phosphate moiety to facilitate ligation to the 3'-OH present on the target fragment. The double-stranded target nucleic acid contains a 5 '-phosphate moiety, either residual from the shearing process, or added using an enzymatic treatment step, and has been end repaired, and optionally extended by an overhanging base or bases, to give a 3'-OH suitable for ligation. In this context, joining means covalent linkage of polynucleotide strands, which were not previously covalently linked. In a particular aspect of the disclosure, such joining takes place by formation of a phosphodiester linkage between the two polynucleotide strands, but other means of covalent linkage (e.g. non-phosphodiester backbone linkages) may be used.
[0116] After ligation of the adapters to the 5’ and 3’ ends of the double stranded nucleic acids, the adapter ligated double stranded nucleic acids are denatured to form the library of single stranded nucleic acid templates. Each of the single stranded nucleic acid templates include a target nucleic acid from the fragmented nucleic acids, such as fragmented genomic DNA, and adapters with truncated surface primer binding regions, ligated to opposite ends of the target nucleic acid.
[0117] The generated library of nucleic acid templates can be used in a method of in- situ next generation sequencing (NGS) target enrichment. Referring to Fig. 2, the method can include providing a plurality of first surface capture primers and second surface captureprimers that are immobilized or tethered to a surface of a solid support of a flow cell. The term “immobilized” as used herein is intended to mean direct or indirect attachment to a solid support via covalent or non-covalent bond(s). In some embodiments, covalent attachment can be used, but all that is required is that the oligonucleotide primer pairs remain stationary or attached to a support under conditions in which it is intended to use the support, for example, in applications requiring nucleic acid amplification and / or sequencing. Oligonucleotides to be used as capture and / or amplification primer pairs can be immobilized such that a 3 '-end is available for enzymatic extension and at least a portion of the sequence is capable of hybridizing to a complementary sequence.
[0118] Any of a variety of solid supports can be used in the method described herein. Particularly useful solid supports are those used for nucleic acid arrays. Examples include glass, modified glass, functionalized glass, inorganic glasses, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, an optical fiber or optical fiber bundles, polymers and multiwell (e.g., microtiter) plates. Examples of plastics include acrylics, polystyrene, copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes and Teflon. Examples of silica-based materials include silicon and various forms of modified silicon.
[0119] In particular embodiments, a solid support can be within or part of a vessel such as a well, tube, channel, cuvette, Petri plate, bottle or the like. In some embodiments, the solid support can be provided in a flow cell. A flow cell can provide a convenient apparatus for use in a method set forth herein. For example, a flow cell is a convenient apparatus for housing a solid support that will be treated with multiple fluidic reagents such as the repeated fluidic deliveries used for some nucleic acid sequencing protocols or some nucleic acid hybridization protocols.
[0120] Examples of a flow-cell are described in WO 2014 / 142841 Al; U.S. Pat. App. Pub. No. 2010 / 0111768 Al and U.S. Pat. No. 8,951,781 or Bentley et al., Nature 456:53-59 (2008), each of which is incorporated herein by reference. Other examples of flow-cells are those that are commercially available from Illumina, Inc. (San Diego, Calif.) for use with a sequencing platform such as a Genome Analyzer®, MiSeq®, NextSeq® or HiSeq® platform. Another particularly useful vessel is a well in a multiwell plate or microtiter plate.
[0121] Optionally, a solid support can include a gel coating. Attachment of nucleic acids to a solid support via a gel is exemplified by flow cells available commercially fromIllumina Inc. (San Diego, Calif.) or described in US Pat. App. Pub. Nos. 2011 / 0059865 Al, 2014 / 0079923 Al, or 2015 / 0005447 Al ; or PCT Publ. No. WO 2008 / 093098, each of which is incorporated herein by reference. Exemplary gels that can be used in the methods and apparatus set forth herein include, but are not limited to, those having a colloidal structure, such as agarose; polymer mesh structure, such as gelatin; or cross-linked polymer structure, such as polyacrylamide, SFA (see, for example, US Pat. App. Pub. No. 2011 / 0059865 A l , which is incorporated herein by reference) or PAZAM (see, for example, US Pat. App. Publ. Nos. 2014 / 0079923 Al, or 2015 / 0005447 Al, each of which is incorporated herein by reference).
[0122] In some embodiments, a solid support can be configured as an array of features to which nucleic acids can be attached. The features can be present in any of a variety of desired formats. For example, the features can be wells, pits, channels, ridges, raised regions, pegs, posts or the like. In some embodiments, the features can contain beads. However, in particular embodiments the features need not contain a bead or particle. Exemplary features include wells that are present in substrates used for commercial sequencing platforms sold by 454 LifeSciences (a subsidiary of Roche, Basel Switzerland) or Ion Torrent (a subsidiary of Life Technologies, Carlsbad Calif.). Other substrates having wells include, for example, etched fiber optics and other substrates described in U.S. Pat Nos. 6,266,459; 6,355,431; 6,770,441 ; 6,859,570; 6,210,891 ; 6,258,568; 6,274,320; US Pat app. Publ. Nos.2009 / 0026082 Al ; 2009 / 0127589 Al; 2010 / 0137143 Al ; 2010 / 0282617 Al or PCT Publication No. WO 00 / 63437, each of which is incorporated herein by reference. In some embodiments, wells of a substrate can include gel material (with or without beads) as set forth in US Pat. App. Publ. No. 2014 / 0243224 Al, which is incorporated herein by reference.
[0123] Features can appear on a solid support as a grid of spots or patches. The features can be located in a repeating pattern or in an irregular, non-repeating pattern. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. Asymmetric patterns can also be useful. The pitch can be the same between different pairs of nearest neighbor features or the pitch can vary between different pairs of nearest neighbor features.
[0124] In particular embodiments, features on a solid support can each have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or500 pnr. Alternatively or additionally, features can each have an area that is smaller than about 1 mm2, 500 (im2, 100 gm2, 25 m2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2.
[0125] In some embodiments, the oligonucleotide surface primer pairs immobilized on the solid support surface may include nucleic acid sequences, which are specific for the truncated surface primer binding regions on the 5’ end or 3’ end of the single stranded nucleic acid templates of the library. The sequences of the surface primers will generally be of known sequence and will therefore include a portion complementary to the truncated surface primer binding regions of the single stranded nucleic acid templates.
[0126] The length of the surface primer oligonucleotides can be, for example, 16-50 nucleotides, more particularly 16-40 nucleotides, and yet more particularly 20-30 nucleotides in length. The desired length of the primer oligonucleotides will depend upon a number of factors. However, the primers are typically long (complex) enough so that the likelihood of annealing to sequences other than a truncated surface primer binding region of the single stranded nucleic acid templates is very low.
[0127] The primer pairs with portions complementary to the truncated surface primer binding regions are capable of annealing or hybridizing specifically to the truncated surface primer binding regions under conditions encountered in a primer annealing step at a second annealing temperature (Ta2) described herein. In some embodiments, a first primer of a primer pair can include a first universal capture region with a portion complementary to a first truncated surface primer binding region of the single stranded nucleic acid template, and a second capture primer of the primer pair can include a second universal capture region with a portion complementary to a second truncated surface primer binding region of the single stranded nucleic acid template. As discussed previously, the first primer of a primer pair can include an Illumina® P5 primer nucleotide sequence, and the second primer of the primer pair can include an Illumina® P7 primer nucleotide sequence. It will be appreciated that other known universal primer pairs can be used.
[0128] Surface primer pairs may additionally comprise non-nucleotide chemical modifications, for example, to facilitate covalent attachment of the primers to a solid support. Certain chemical modifications may themselves improve the function of the molecule as a primer or may provide some other useful functionality, such as providing a cleavage site that enables the primer (or an extended polynucleotide strand derived therefrom) to be cleaved from a solid support. Useful chemical modifications can also provide reversible modificationsthat prevent hybridization or extension of the primer until the modification is removed or reversed. Similarly, other molecules attached to a surface can include cleavable linker moieties and or reversible modifications that alter a particular chemical activity of function of the molecule.
[0129] In certain embodiments, primer pairs are immobilized by covalent attachment to the solid support at or near the 5' end of the primer, such that a portion of the primer is free to anneal to its single stranded nucleic acid template and the 3' hydroxyl group is free to function in primer extension. In certain embodiments, a subset of modified primers can be provided that are prevented from hybridization and / or extension until the modification is removed, reversed or altered. In particular embodiments, the primer oligonucleotides will be incapable of hybridization to the initial single stranded nucleic acid probe templates.
[0130] The chosen attachment chemistry for attaching the primer pairs to solid support surface can depend on the nature of the solid support and any functionalization or derivatization applied to it. The primer itself may include a moiety which may be a nonnucleotide chemical modification to facilitate attachment. For example, the primer may include a sulfur containing nucleophile, such as a phosphorothioate or thiophosphate at the 5' end. In the case of solid supported polyacrylamide hydrogels, this nucleophile may bind to a bromoacetamide group present in the hydrogel. In one embodiment, the means of attaching primers to the solid support is via 5' phosphorothioate attachment to a hydrogel comprised of polymerized acrylamide and N-(5-bromoacetamidylpentyl) acrylamide (BRAPA).
[0131] A uniform, homogeneously distributed ‘lawn’ of immobilized surface primer pairs may be formed by coupling (grafting) a solution of oligonucleotide primer pairs onto the solid support. The solution can contain a homogeneous population of primer pairs. Each surface that is exposed to the solution therefore reacts with the solution to create a uniform density of immobilized sequences over the whole of the exposed solid support. A suitable density of oligonucleotide primer pairs is at least about 1 fmol / mm2(6xlO10per cm2), or more optimally, at least about 10 fmol / mm2(6xlOnper cm2). The density of the oligonucleotide primer pairs can be controlled to give an optimum cluster density.
[0132] The method can further include providing a plurality of nucleic acid target probes that each include a linker and a target capture nucleic acid that hybridizes and captures a target nucleic acid sequence of the library of single stranded nucleic acids at or below a first annealing temperature (Tai). The nucleic acid linker can include an end immobilized ortethered to the surface of the solid support and an end ligated to an end of the target capture nucleic acid such that the plurality of nucleic acid target probes can be tethered or linked to the surface of solid support. For example, 5’ ends of the nucleic acid target probes can be tethered to the surface of the solid support prior to loading the library of single stranded nucleic acids onto the solid support.
[0133] The nucleic acid target probes can include different target capture sequences that hybridize to and capture different target nucleic acid sequences of the nucleic acid templates of the library. Different target capture sequences can be used to selectively bind to one or more desired target nucleic acids from the library. In some cases, the different nucleic acid target probes can include a target capture sequence that is common to all or a subset of the probes. For example, the nucleic acid probes on a solid support can have a poly A or poly T sequence. Such probes can hybridize to mRNA molecules, cDNA molecules or amplicons thereof that have poly A or poly T tails. Although the mRNA or cDNA species will have different target sequences, capture will be mediated by the common poly A or poly T sequence regions.
[0134] Any of a variety of target nucleic acids can be captured and analyzed in a method set forth herein, including but not limited to, messenger RNA (mRNA), copy DNA (cDNA), genomic DNA (gDNA), ribosomal RNA (rRNA) or transfer RNA (tRNA). Particular target sequences can be selected from databases, and appropriate capture sequences designed using techniques and databases known in the art.
[0135] In some embodiments, where the 5’ ends of the nucleic acid target probes are linked to the surface of the solid support, the 3’ ends of the nucleic acid target probes can be blocked to prevent extension with DNA polymerase and dNTPs during extension of the surface capture primers. Alternatively, 3’ ends of the nucleic acid target probes can be immobilized or tethered to the surface of the solid support. This prevents or blocks extension of the target nucleic acid probes with DNA polymerase and dNTPs during extension of the surface capture primers without the use of a 3 ’ end block.
[0136] The library of single stranded nucleic acids ligated at the 5’ and 3’ ends with adapters that include truncated surface primer binding regions can then be loaded on the solid support of the flow cell and incubated at or below the first melting temperature (Tml) of the target capture nucleic acid and above the second melting temperature (Tm2) so that target nucleic acid sequences of the single stranded nucleic acid templates hybridize and arecaptured by target capture nucleic acids of the probes and the truncated surface binding regions do not bind to the first surface capture primers and second surface capture primers.
[0137] The single stranded nucleic acid templates not targeted by the target capture nucleic acids of the nucleic acid target probes do not bind or hybridize to the surface primers through primer hybridization and are free-floating in solution. Free-floating single stranded nucleic acid templates not captured by the nucleic acid target probes after loading on the solid support are then removed from the solid support by, for example, washing the solid support.
[0138] The temperature on the solid support is subsequently lowered to at or below the second melting temperature (Tm2) such that one of the truncated surface primer binding regions of the captured single stranded nucleic acid templates hybridizes to the first surface capture primers or second surface capture primers.
[0139] The surface capture primers hybridized or annealed to the truncated surface primer binding regions of captured single stranded nucleic acid templates are extended to form a plurality of complementary templates tethered to the surface and hybridized to the plurality of single stranded capture nucleic acid templates.
[0140] It will be appreciated that the extension of the surface primers can be carried out using methods known in the art for amplification of nucleic acids or sequencing of nucleic acids. Primers hybridized with a single stranded nucleic acid template can be extended to form a plurality of immobilized complementary single stranded nucleic acid templates hybridized to the plurality of single stranded nucleic acid templates. An extension reaction may be carried out wherein the primer oligonucleotide is extended by sequential addition of nucleotides to generate a complementary copy of the single stranded nucleic acid template attached to the solid support. In some embodiments, the nucleic acid sequences are extended by adding nucleotide triphosphates (NTPs) and polymerase to the surface of the microarray.
[0141] In particular embodiments one or more nucleotides can be added to the 3' end of a nucleic acid, for example, via polymerase catalysis (e.g., DNA polymerase, RNA polymerase or reverse transcriptase). Chemical or enzymatic methods can be used to add one or more nucleotide to the 3' or 5' end of a nucleic acid. One or more oligonucleotides can be added to the 3' or 5' end of a nucleic acid, for example, via chemical or enzymatic (e.g., ligase catalysis) methods. A nucleic acid can be extended in a template directed manner, whereby the product of extension is complementary to a target nucleic acid that is hybridized to the nucleic acid that is extended. In some embodiments, a DNA primer is extended by areverse transcriptase using an RNA template, thereby producing a cDNA. Thus, an extended surface primer made in a method set forth herein can be a reverse transcribed DNA molecule. Exemplary methods for extending nucleic acids are set forth in US Pat. App. Publ. No. US 2005 / 0037393 Al or U.S. Pat. No. 8,288,103 or 8,486,625, each of which is incorporated herein by reference.
[0142] All or part of a target nucleic acid that is hybridized to a nucleic acid probe can be copied by extension. For example, an extended surface primer can include at least 1, 2, 5,10, 25, 50, 100, 200, 500, 1000 or more nucleotides that are copied from a target nucleic acid. The length of the extension product can be controlled, for example, using reversibly terminated nucleotides in the extension reaction and running a limited number of extension cycles. The cycles can be run as exemplified for sequencing by synthesis techniques, and the use of labeled nucleotides is not necessary. Accordingly, an extended surface primer produced in a method set forth herein can include no more than 1000, 500, 200, 100, 50, 25, 10, 5, 2 or 1 nucleotides that are copied from a target nucleic acid. Of course, extended surface primers can be any length within or outside of the ranges set forth above.
[0143] The resulting extended nucleic sequence of the surface capture primers include the target nucleic acid sequences and adapter sequences from the captured single strand nucleic acid templates (albeit in complementary form). It will be understood that other sequence elements that are present in the captured single strand nucleic acid templates can also be included in the extended surface primers. Such elements include, for example, primer binding sites, cleavage sites, other tag sequences (e.g., sample identification tags), capture sequences, recognition sites for nucleic acid binding proteins or nucleic acid enzymes, or the like.
[0144] If the target probes used for target capturing are not blocked at the 3’ ends, they would also be extended at this step, forming strands composed partially with the original target probes and partially with copied library. But this strand would only have one of the two surface primer binding regions, and thus cannot be exponentially amplified before sequencing. So, they might have some negative effects on the signal to noise ratio but are not expected to totally ruin the sequencing process.
[0145] Extension of the surface primers also causes the single stranded nucleic acid templates to separate from the target capture nucleic acids of the target probes, so that the single stranded nucleic acids are hybridized only to their respective extended surface primeror the complementary nucleic acid template is tethered to the surface of the solid support. The terms ‘separate’ and ‘separating’ when used in reference to strands of a nucleic acid, refer to the physical dissociation of the DNA bases that interact within, for example, a Watson-Crick DNA-duplex of the single stranded nucleic acid templates and its complement.
[0146] After the extension reaction, captured single stranded nucleic acid templates can be removed or separated from the extended nucleic sequence of the surface capture primers, and hence, strand separation can result in loss of one of the strands from the surface. In some embodiments, the single stranded nucleic acid templates are removed by denaturing the templates from the immobilized complementary single stranded nucleic acids and washing the single stranded nucleic acid templates from the solid support, leaving a copy of each targeted library molecule tethered to the surface of the solid support.
[0147] Fig. 3 illustrates another strategy or method of loading nucleic acid target probes on the surface of the solid support of a flow cell. The method shown in Fig. 2 has a disadvantage that once the surface of the solid support of the flow cell is modified with specific nucleic acid target probes, the flow cell cannot be used for other panels of nucleic acid target probes. Using the method of Fig. 2, every specific panel of nucleic acid target probes must have a specific type of flow cell manufactured.
[0148] In the method of Fig. 3, the panel of nucleic acid target probes are single stranded nucleic acids, such as ssDNA, that each include a target capture nucleic acid at a first end (e.g., 5’ end) and a surface primer binding region at an opposite second end (e.g., 3’ end). The target capture nucleic acid can hybridize and capture a target nucleic acid sequence of the library of single stranded nucleic acid templates at or below a first annealing temperature (Tai ). The surface primer binding region can hybridize to a surface primer on the surface of the solid support of the flow cell. The surface primer binding region can be linked to the target capture nucleic acid with a nucleic acid linker. Optionally, at a connection of the surface primer binding region to the linker or rest of the strand, the nucleic acid target probe can include a blocker or modification that can stop or prevent DNA polymerase from copying the target probe or extending the surface primer to which the target probe is hybridized during an extension reaction. If the probes do not have the optional modification, surface primers binding to the target probes can also be extended during an extension reaction. The resulting extended strands would not have a surface primer bindingregion so could not be amplified in a later clustering step, and would not significantly affect sequencing steps.
[0149] The surface primer binding region can have a nucleic acid sequence complementary to a full-length nucleic acid sequence of the first surface primers or second surface primers such that the surface primer binding region of the target probe when hybridized or annealed to the first surface primers or second surface primers, has a third melting temperature (Tm3) and third annealing temperature (Ta3) that is the higher than the second melting temperature (Tm2) and second annealing temperature (Ta2) of the truncated surface binding region of the single stranded nucleic templates. At higher temperatures above the second melting temperature (Tm2) of the truncated surface primer binding region of the single stranded nucleic acid templates, the target probes can anneal to the surface primers without annealing of the single stranded nucleic acid templates to the surface primers.
[0150] The nucleic acid target probes, similar to the target probes of the method of Fig. 2, can include different target capture nucleic acid sequences that hybridize to and capture different target nucleic acid sequences from the library. Different target capture sequences can be used to selectively bind to one or more desired target nucleic acids from the library. In some cases, the different nucleic acid probes can include a target capture sequence that is common to all or a subset of the probes.
[0151] Similar to the method of Fig. 2, the method of Fig. 3 includes providing a plurality of first surface capture primers and second surface capture primers that are immobilized or tethered to a surface of a solid support of a flow cell.
[0152] The nucleic acid target probes and the library of single stranded nucleic acids ligated at the 5 ’ and 3 ’ ends with adapters that include truncated surface primer binding regions can then be loaded on the solid support of the flow cell by providing a fluid that contains a mixture of nucleic acid target probes and the library of single stranded nucleic acid templates and contacting this fluidic mixture with the primer pairs immobilized on the solid support.
[0153] The library and nucleic acid target probes loaded onto the solid support are then incubated at or below the first melting temperature (Tml) of the target capture nucleic acid and third melting temperature (Tm3) of the surface primer binding region of the target probe and above the second melting temperature (Tm2) so that target nucleic acid sequences of thesingle stranded nucleic acid templates hybridize and are captured by target capture nucleic acids of the probes, the surface primer binding regions of the target probe hybridize to the surface primers, and the truncated surface binding regions do not bind to the first surface capture primers and second surface capture primers. The nucleic acid target probes can have random access to primer pairs on the surface. Accordingly, the target probes and captured single stranded nucleic acid templates can be randomly located on the solid support surface.
[0154] When hybridizing target probes to immobilized primers on a patterned flow cell, hybridization conditions can be adjusted such that only a target probe hybridizes with an immobilized surface primer. Methods of hybridization for formation of stable duplexes between complementary sequences by way of Watson-Crick base-pairing are known in the art.
[0155] The single stranded nucleic acid templates not targeted by the target capture nucleic acids of the nucleic acid target probes do not bind or hybridize to the surface primers through primer hybridization and are free-floating solution. Free-floating single stranded nucleic acid templates not captured by the nucleic acid target probes after loading on the solid support are then removed from the solid support by, for example, washing the solid support.
[0156] The temperature on the solid support is subsequently lowered to at or below the second melting temperature (Tm2) such that one of truncated surface primer binding regions of the captured single stranded nucleic acid templates hybridizes to the first surface capture primer or second surface capture primer.
[0157] The surface capture primers hybridized or annealed to the truncated surface primer binding regions of captured single stranded nucleic acid templates are extended to form a plurality of complementary templates tethered to the surface and hybridized to the plurality of single stranded capture nucleic acid templates. An extension reaction may be carried out wherein surface primer oligonucleotides are extended by sequential addition of nucleotides to generate a complementary copy of the single stranded nucleic acid template attached to the solid support. In some embodiments, the nucleic acid sequences are extended by adding nucleotide triphosphates (NTPs) and polymerase to the surface of the microarray.
[0158] The resulting extended nucleic sequence of the surface capture primers includes the target nucleic acid sequences and adapter sequences from the captured single strand nucleic acid templates (albeit in complementary form). It will be understood that other sequence elements that are present in the captured single strand nucleic acid templates canalso be included in the extended surface primers. Such elements include, for example, primer binding sites, cleavage sites, other tag sequences (e.g., sample identification tags), capture sequences, recognition sites for nucleic acid binding proteins or nucleic acid enzymes, or the like.
[0159] Extension of the surface primers also causes the single stranded nucleic acid templates to separate from the target capture nucleic acids of the target probes so that the single stranded nucleic acids are hybridized only to their respective extended surface primer or the complementary nucleic acid template is tethered to the surface of the solid support.
[0160] After the extension reaction, captured single stranded nucleic acid templates and the targeting probes can be removed or separated, respectively, from the extended nucleic sequence of the surface capture primers tethered to the surface and the surface primers so that only the plurality of complementary templates are tethered to the surface. In some embodiments, the single stranded nucleic acid templates and target probes are removed by denaturing the templates from the immobilized complementary single stranded nucleic acids and the target probes from the surface primers. The denatured single stranded nucleic acid templates and target probes can be washed from the solid support, leaving a copy of each targeted library molecule tethered to the surface of the solid support.
[0161] Following removal of the denatured single stranded nucleic acid templates in the method of Fig. 2 or the removal of the denatured single stranded nucleic acid templates and target probes in the method of Fig. 3, the immobilized complementary single stranded nucleic acids, i.e., the extended nucleic sequence of the surface capture primers on the surface of the solid support, can be amplified to provide clusters of nucleic acid strands. For example, the immobilized complementary single stranded nucleic acids can be amplified by, for example, bridge amplification or PCR, to form amplicons or clusters of amplified single stranded nucleic acids attached to the surface of the substrate. Any of a variety of amplification techniques can be used to form the amplicons or clusters of the amplified single stranded nucleic acids. Examples of amplification techniques include polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), or random prime amplification (RPA). In some embodiments, the amplification can be carried out on solid phase. For example, the 3 '-end of the immobilized complementary single stranded nucleic acid can hybridize with a non-extended immobilized capture primer via its complementary 3 '-terminal surface primer binding domain, thereby forming a bridgestructure and amplified by bridge PCR or bridge amplification. Exemplary reagents and conditions that can be used for bridge amplification are described, for example, in U.S. Pat. Nos. 5,641,658, 7,115,400, or 8,895,249; or U.S. Pat. Publ. Nos. 2002 / 0055100 Al, 2004 / 0096853 Al, 2004 / 0002090 Al, 2007 / 0128624 Al or 2008 / 0009420 Al, each of which is incorporated herein by reference. In some embodiments, one or more rounds of bridge amplification are conducted to form a monoclonal clusters of different amplified single stranded nucleic acids.
[0162] Solid-phase PCR amplification can also be carried out with one of the primers attached to the solid support and the second primer in solution. An exemplary format that uses a combination of a surface attached primer and soluble primer is the format used in emulsion PCR as described, for example, in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, or U.S. Pat. App. Publ. Nos. 2005 / 0130173 Al or 2005 / 0064460 Al, each of which is incorporated herein by reference.
[0163] Amplification sites or clusters in an array need not be entirely clonal in all embodiments. Rather, for some applications, an individual amplification site or cluster can be predominantly populated with amplicons from a first amplified single stranded nucleic acid and can also have a low level of contaminating amplicons from a second amplified single stranded nucleic acid. An array can have one or more amplification sites that have a low level of contaminating amplicons so long as the level of contamination does not have an unacceptable impact on a subsequent use of the array. For example, when the array is to be used in a detection application, an acceptable level of contamination would be a level that does not impact signal to noise or resolution of the detection technique in an unacceptable way. Accordingly, apparent clonality will generally be relevant to a particular use or application of an array made by the methods set forth herein. Exemplary levels of contamination that can be acceptable at an individual amplification site for particular applications include, but are not limited to, at most 0.1%, 0.5%, 1%, 5%, 10% or 25% contaminating amplicons. An array can include one or more amplification sites having these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites in an array can have some contaminating amplicons.
[0164] A composition for amplifying single stranded nucleic acids at amplification sites, referred to herein as an “amplification reagent,” is typically capable of rapidly making copies of the single stranded nucleic acids at amplification sites. An amplification reagentused in a method described herein will generally include a polymerase and nucleotide triphosphates (NTPs). Any of a variety of polymerases known in the art can be used, but in some embodiments it may be preferable to use a polymerase that is exonuclease negative. Examples of nucleic acid polymerases that can be used include, but are not limited to, DNA polymerase (such as Klenow fragment, T4 DNA polymerase, Bst (Bacillus stearothermophilus) polymerase), thermostable DNA polymerases (such as Taq, Vent, Deep Vent, Pfu, TH, and 9° N DNA polymerases) as well as their genetically modified derivatives (TaqGold, VENTexo, Pfu exo). In some embodiments, an amplification reagent can also include recombinase, accessory protein, and single-stranded DNA binding (SSB) protein for recombinase-facilitated amplification.
[0165] The NTPs can be deoxyribonucleotide triphosphates (dNTPs) for embodiments where DNA copies are made. Typically the four native species, dATP, dTTP, dGTP and dCTP, will be present in a DNA amplification reagent; however, analogs can be used if desired. The NTPs can be ribonucleotide triphosphates (rNTPs) for embodiments where RNA copies are made. Typically the four native species, rATP, rUTP, rGTP and rCTP, will be present in a RNA amplification reagent; however, analogs can be used if desired. NTPs can be modified with a fluorescent or radioactive group. A large variety of synthetically modified nucleic acids have been developed for chemical and biological methods in order to increase the detectability and / or the functional diversity of nucleic acids. These functionalized / modified molecules (e.g., nucleotide analogs) can be fully compatible with natural polymerizing enzymes, maintaining the base pairing and replication properties of the natural counterparts.
[0166] The rate at which an amplification reaction occurs can be increased by increasing the concentration or amount of one or more of the active components of an amplification reaction. For example, the amount or concentration of polymerase, nucleotide triphosphates, or primers. In some cases, the one or more active components of an amplification reaction that are increased in amount or concentration (or otherwise manipulated in a method set forth herein) are non-nucleic acid components of the amplification reaction.
[0167] Amplification rate can also be increased in a method set forth herein by adjusting the temperature. For example, the rate of amplification at one or more amplification sites can be increased by increasing the temperature at the site(s) up to a maximumtemperature where reaction rate declines due to denaturation or other adverse events. Optimal or desired temperatures can be determined from known properties of the amplification components in use or empirically for a given amplification reaction mixture. Such adjustments can be made based on a priori predictions of primer melting temperature (Tm) or empirically. In certain embodiments the temperature of an amplification reaction are at least 35°C to no greater than 70°C. For instance, an amplification reaction can be at least 35°C to no greater than 42°C, or at least 57°C to no greater than 63°C.
[0168] The result of bridging amplification is a population of clonal “bridged” amplification products or amplified single stranded nucleic acids clusters at the amplification sites. Both strands of the amplicon are immobilized on the surface of an amplification site at the 5' ends, where this attachment is derived from the original attachment of the oligonucleotide primer pairs. The amplicons within amplification sites will be clonal and derived from amplification of a single complementary template, or with acceptable levels of another amplicon as described herein.
[0169] To facilitate identification of nucleic acid strands of the clusters, one of the strands of the double stranded bridged structure can be selectively removed from the surface to allow efficient hybridization of a complementary identification probe or nucleic acid for sequencing. The selective removal of a specific strand is referred to herein as “linearization.” Examples of suitable methods for linearization are described herein and are described in more detail in application number WO 2007 / 010251 and U.S. Pat. Application Pub. 2012 / 0309634.
[0170] In one embodiment, linearization is achieved by cleaving one strand of the bridged double stranded amplicons and then subjecting the resulting structure to conditions that remove the strand that is no longer attached to the amplification site surface. Cleavage can be accomplished through the use of a primer that includes a cleavage site. The cleavage site is typically in a location that results in a substantial portion of one strand of the bridged structure to be free of the surface of the amplification site — no longer immobilized — and susceptible to loss after the removal step.
[0171] In one embodiment, a cleavage site is treated to remove a nucleotide and make an abasic site. An “abasic site” is a nucleotide position in a nucleic acid from which the base component has been removed. Abasic sites can be formed chemically under artificial conditions or by the action of enzymes. Once formed, abasic sites may be cleaved (e.g. bytreatment with an endonuclease or other single- stranded cleaving enzyme, exposure to heat or alkali), providing a means for site-specific cleavage of a nucleic acid.
[0172] In one embodiment, an abasic site may be created at a pre-determined position on one strand of an immobilized amplicon. This can be achieved, for example, by incorporating a specific nucleotide at the pre-determined position.
[0173] In one embodiment, a deoxyuridine (U) is incorporated in one of the primers attached to the surface of an amplification site. The enzyme uracil DNA glycosylase (UDG) can then be used to remove the uracil base, generating an abasic site on one strand. The polynucleotide strand including the abasic site can then be cleaved at the abasic site by treatment with endonuclease (e.g., DNA glycosylase-lyase Endonuclease VIII), heat or alkali. In a particular embodiment, the USER reagent available from New England Biolabs (NEB # M5505S) is used for the creation of a single nucleotide gap at a uracil base in an immobilized capture primer. In one embodiment, the amplification sites are exposed to a mixture containing the appropriate glycosylase and one or more suitable endonucleases. Treatment with endonuclease enzymes gives rise to a 3'-phosphate moiety at the cleavage site, which can be removed with a suitable phosphatase such as alkaline phosphatase.
[0174] Treatment with endonuclease enzymes gives rise to a 3 '-phosphate moiety at the cleavage site, and the presence of 3' phosphate is known to inhibit the activity of exonuclease I (Lehman and Nussbaum, 1964, J. Biol. Chem., 239: 2628-2636). Both exonuclease and linearization steps can occur at the same time by combining the enzymes. The reduction of these two steps into one results in faster sequencing runs as two steps are now preformed simultaneously. Moreover, combining both steps does not have a detrimental effect on primary metrics, read quality, dual indexing, or genome build metrics.
[0175] Abasic site generation and cleavage results in an end of amplified single stranded nucleic acid that is no longer immobilized to the surface. This strand can be completely removed from the surface by exposing the amplification site to suitable conditions. In one embodiment, removal is by denaturation. The denaturation can be performed thermally or isothermally, for example, using chemical denaturation. The chemical denaturant may be urea, hydroxide, or formamide or other similar reagents. In another embodiment, removal can be achieved by treatment with an exonuclease with 5 '-3' activity, such as lambda or T7 exonuclease. Removal of the unattached strand results in a remaining single strand that can act as a capture probe for a target nucleic acid.
[0176] Optionally, the 3' ends of the nucleic acid strands at the amplification sites are repaired. The exonuclease can remove some of the nucleotides at the 3' ends of the nucleic acids after the linearization. Without intending to be limiting, it is possible the 3' ends are “breathing” slightly resulting in a small number of nucleotides becoming single stranded and available to the exonuclease for digestion. Repair can be achieved by exposing the nucleotides to a DNA polymerase, such as the DNA polymerase used for the bridging amplification.
[0177] The clusters of the amplified single stranded nucleic acids can be sequenced by hybridization of sequencing primers to the sequencing primer binding domain of the amplified single stranded nucleic acids, sequentially incorporating one or more nucleotides into a polynucleotide strand complementary to the region of the amplified single stranded nucleic acids to be sequenced, identifying the base present in one or more of the incorporated nucleotide (s) and thereby determining the sequence of a region of the amplified single stranded nucleic acids.
[0178] One sequencing method that can be used is sequencing-by-synthesis (SBS). In SBS, extension of a nucleic acid primer along a nucleic acid template (e.g., amplified single stranded nucleic acids or amplicon thereof) is monitored to determine the sequence of nucleotides in the template. The underlying chemical process can be polymerization e.g., as catalyzed by a polymerase enzyme). In a particular polymerase-based SBS embodiment, fluorescently labeled nucleotides are added to a primer (thereby extending the primer) in a template dependent fashion such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. A plurality of different amplified single stranded nucleic acids at different sites of an array set forth herein can be subjected to an SBS technique under conditions where events occurring for different amplified single stranded nucleic acids can be distinguished due to their location in the array.
[0179] Each nucleotide type may thus carry a different fluorescent label, for example, as described in US Provisional Application No. 60 / 801,270 (Novel dyes and the use of their labelled conjugates), published as WO07135368, the contents of which are incorporated herein by reference in their entirety. The detectable label need not, however, be a fluorescent label. Any label can be used which allows the detection of an incorporated nucleotide.
[0180] One method for detecting fluorescently labeled nucleotides comprises using laser light of a wavelength specific for the labelled nucleotides, or the use of other suitablesources of illumination. The fluorescence from the label on the nucleotide may be detected by a CCD camera or other suitable detection means. Suitable instrumentation for recording images of clustered arrays is described in U.S. Provisional Application No. 60 / 788,248 (Systems and devices for sequence by synthesis analysis), published as WO07123744, the contents of which are incorporated herein by reference in their entirety. The method need not be limited to use of the sequencing method outlined above, as essentially any sequencing methodology which relies on successive incorporation of nucleotides into a polynucleotide chain can be used. Suitable alternative techniques include, for example, Pyrosequencing™, FISSEQ (fluorescent in situ sequencing), MPSS and sequencing by ligation-based methods, for example as described in U.S. Pat. No. 6,306,597 which is incorporated herein by reference.
[0181] The amplified single stranded nucleic acids may be further analyzed to obtain a second read from the opposite end of the fragment. Methodology for sequencing both ends of a cluster are described in W007010252, PCTGB2007 / 003798 and US 20090088327, the contents of which are incorporated by reference herein in their entirety.
[0182] Other sequencing procedures that use cyclic reactions can be used, such as pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into a nascent nucleic acid strand (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 1 1(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Pat. Nos. 6,210,891 ; 6,258,568 and 6,274,320). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated can be detected via luciferase-produced photons. Thus, the sequencing reaction can be monitored via a luminescence detection system. Excitation radiation sources used for fluorescence-based detection systems are not necessary for pyrosequencing procedures. Useful fluidic systems, detectors and procedures that can be used for application of pyrosequencing to arrays of the present disclosure are described, for example, in WIPO Published Pat. App. 2012 / 058096, US 2005 / 0191698 Al, U.S. Pat. Nos. 7,595,883, and 7,244,559.
[0183] Sequencing-by-ligation reactions are also useful including, for example, those described in Shendure et al. Science 309:1728-1732 (2005); U.S. Pat. Nos. 5,599,675; and 5,750,341. Some embodiments can include sequencing-by-hybridization procedures as described, for example, in Bains et al., Journal of Theoretical Biology 135(3), 303-7 (1988);Drmanac et al., Nature Biotechnology 16, 54-58 (1998); Fodor et al., Science 251(4995), 767-773 (1995); and WO 1989 / 10977. In both sequencing-by-ligation and sequencing-by- hybridization procedures, template nucleic acids (e.g., a target nucleic acid or amplicons thereof) that are present at sites of an array are subjected to repeated cycles of oligonucleotide delivery and detection. Fluidic systems for SBS methods as set forth herein or in references cited herein can be readily adapted for delivery of reagents for sequencing-by-ligation or sequencing-by-hybridization procedures. Typically, the oligonucleotides are fluorescently labeled and can be detected using fluorescence detectors similar to those described with regard to SBS procedures herein or in references cited herein.
[0184] Some embodiments can use methods involving the real-time monitoring of DNA polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate-labeled nucleotides, or with zeromode waveguides (ZMWs). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008).
[0185] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, Conn., a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al ; US 2010 / 0137143 Al; or US 2010 / 0282617 Al. Methods set forth herein for amplifying target nucleic acids can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons at the sites of the arrays that are used to detect protons.
[0186] An array of the sequenced clusters, having been produced by a method described herein, can be used for gene expression analysis. Gene expression can be detected or quantified using RNA sequencing techniques, such as those referred to as digital RNA sequencing. RNA sequencing techniques can be carried out using sequencing methodologies known in the art such as those set forth above. Gene expression can also be detected or quantified using hybridization techniques carried out by direct hybridization to an array or using a multiplex assay, the products of which are detected on an array. An array of thepresent disclosure, for example, having been produced by a method set forth herein, can also be used to determine genotypes for a genomic DNA sample from one or more individual. Exemplary methods for array-based expression and genotyping analysis that can be carried out on an array of the present disclosure are described in U.S. Pat. Nos. 7,582,420; 6,890,741 ; 6,913,884 or 6,355,431 or US Pat. Pub. Nos. 2005 / 0053980 Al; 2009 / 0186349 Al or US 2005 / 0181440 A l .
[0187] Another useful application for an array having been produced by a method set forth herein is single-cell sequencing. When combined with indexing methods single cell sequencing can be used in chromatin accessibility assays to produce profiles of active regulatory elements in thousands of single cells, and single cell whole genome libraries can be produced. Examples for single-cell sequencing that can be carried out on an array of the present disclosure are described in U.S. Published Patent Application 2018 / 0023119 Al, U.S. Provisional Application Ser. No. 62 / 673,023 and Ser. No. 62 / 680,259.
[0188] An advantage of the methods set forth herein is that they provide for rapid and efficient creation of arrays from any of a variety of nucleic acid libraries. Accordingly the present disclosure provides integrated systems capable of making an array using one or more of the methods set forth herein and further capable of detecting nucleic acids on the arrays using techniques known in the art such as those exemplified herein. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents to an array of amplification sites such as pumps, valves, reservoirs, fluidic lines and the like. A particularly useful fluidic component is a flow cell. A flow cell can be configured and / or used in an integrated system to create an array of the present disclosure and to detect the array. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and U.S. Pat. No. 8,951,781. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those described herein. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating arrays of nucleic acids and also determining the sequence of the nucleic acids include, without limitation, theMiSeg™, HiSeg2500™, NextSeg™, MiniSeg™, NovaSeg™ and iSeg™ sequencing platforms from Illumina, Inc. (San Diego, Calif.) and devices described in U.S. Pat. No. 8,951,781. Such devices can be modified to make arrays in accordance with the guidance set forth herein.
[0189] A system capable of carrying out a method set forth herein need not be integrated with a detection device. Rather, a stand-alone system or a system integrated with other devices is also possible. Fluidic components similar to those exemplified above in the context of an integrated system can be used in such embodiments.
[0190] A system capable of carrying out a method set forth herein, whether integrated with detection capabilities or not, can include a system controller that is capable of executing a set of instructions to perform one or more steps of a method, technique or process set forth herein. For example, the instructions can direct the performance of steps for creating an array under bridge amplification conditions. Optionally, the instructions can further direct the performance of steps for detecting nucleic acids using methods set forth previously herein. A useful system controller may include any processor-based or microprocessor-based system, including systems using microcontrollers, reduced instruction set computers (RISC), application specific integrated circuits (ASICs), field programmable gate array (FPGAs), logic circuits, and any other circuit or processor capable of executing functions described herein. A set of instructions for a system controller may be in the form of a software program. As used herein, the terms “software” and “firmware” are interchangeable, and include any computer program stored in memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The software may be in various forms such as system software or application software. Further, the software may be in the form of a collection of separate programs, or a program module within a larger program or a portion of a program module. The software also may include modular programming in the form of object-oriented programming.
[0191] It will be understood that an array of the present disclosure, for example, having been produced by a method set forth herein, need not be used for a detection method. Rather, the array can be used to store a nucleic acid library. Accordingly, the array can be stored in a state that preserves the nucleic acids therein. For example, an array can be stored in a desiccated state, frozen state (e.g. in liquid nitrogen), or in a solution that is protective of nucleic acids. Alternatively or additionally, the array can be used to replicate a nucleic acidlibrary. For example, an array can be used to create replicate amplicons from one or more of the sites on the array.
[0192] From the above description of the invention, those skilled in the art will perceive improvements, changes and modifications. Such improvements, changes and modifications within the skill of the art are intended to be covered by the appended claims. All references, publications, and patents cited in the present application are herein incorporated by reference in their entirety.
Claims
Having described the invention, we claim:
1. A next generation sequencing (NGS) library comprising: a plurality of single stranded nucleic acid templates that include target nucleic acid sequences ligated at 5’ and 3’ ends with adapters that include truncated surface primer binding regions, wherein the single stranded nucleic acid templates are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support at a first annealing temperature (Tai), and the truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal respectively to first surface primers or second surface primers on the surface of the solid support at a second annealing temperature (Ta2) below the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
2. The NGS library of claim 1, wherein the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers.
3. The NGS library of claim 2, wherein the truncated surface primer binding regions have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the nucleic acid sequence of first surface primers or the second surface primers.
4. The NGS library of any of claims 1 to 3, wherein the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first annealing temperature (Tai).
5. The NGS library of any of claims 1 to 4, wherein the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support have a first melting temperature (Tml), and the truncated surface primer binding regions when hybridized or annealed to the first surfaceprimers or second surface primers have a second melting temperature (Tm2), and wherein the second melting temperature (Tm2) is below the first melting temperature (Tml).
6. The NGS library of claim 5, wherein the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
7. The NGS library of claim 5 or claim 6, wherein the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is below than Tml.
8. The NGS library of claim 5 or claim 6, wherein the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
9. The NGS library of any claims 5 to 8, a nucleic acid sequence complementary to full length nucleic sequence of the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third melting temperature (Tm3), and wherein the second melting temperature (Tm2) is below the third melting temperature (Tm3).
10. The NGS library of any of claims 1 to 9, wherein ligated adapters further include sequencing primer binding domains and indexes.
11. The NGS of claim 10, wherein the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
12. A method of forming a NGS library, the method comprising: providing a plurality of fragmented subsets of double stranded nucleic acids; end-repairing and A-tailing 3’ ends of the fragmented subsets of double stranded nucleic acids; ligating adapters to the 5’ and 3’ ends of the double stranded nucleic acids, wherein the adapters include the truncated surface primer binding regions; and denaturing the ligated double stranded nucleic acids to form the library of single strand nucleic acid templates; wherein the ligated adapters include truncated surface primer binding regions, and wherein the single stranded nucleic acid templates are configured to hybridize or anneal to target capture nucleic acids of nucleic acid target probes on a surface of a solid support at a first annealing temperature (Tai), and the truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal respectively to first surface primers and second surface primers on the surface of the solid support at a second annealing temperature (Ta2) below the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
13. The method of claim 12, wherein the plurality of fragmented subsets of double stranded nucleic acids include fragmented subsets of double stranded DNA, preferably double stranded genomic DNA.
14. The method of claim 12 or 13, wherein the fragmented subsets of double stranded nucleic acids are from about 100 to about 1000 nucleic acids in length.
15. The method of any of claims 12 to 13, wherein the double stranded nucleic acids are end-repaired with polymerase and nucleotide triphosphates (NTPs) to form an end with one A at both 3’ ends.
16. The method of any of claims 12 to 15, wherein ligated adapters further include sequencing primer binding domains and indexes.
17. The method of claim 16, wherein the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
18. The method of any of claims 12 to 17, wherein the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers.
19. The method of any of claims 12 to 18, wherein the truncated surface primer binding regions have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the first surface primers or the second surface primers.
20. The method of any of claims 12 to 19, wherein the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first annealing temperature (Tai).
21. The method of any of claims 12 to 20, wherein the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support has a first melting temperature (Tml), and the truncated surface primer binding regions when hybridized or annealed to the first surface primers or second surface primers has a second melting temperature (Tm2), and wherein the second melting temperature (Tm2) is below the first melting temperature (Tml).
22. The method of claim 21, wherein the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
23. The method of claim 21 or claim 22, wherein the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is below than Tml.
24. The method of any of claims 21 to 23, wherein the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
25. The method of any claims 21 to 24, a nucleic acid sequence complementary to full length the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third melting temperature (Tm3), and wherein the second melting temperature (Tm2) is below the third melting temperature (Tm3).
26. A method of in-situ next generation sequencing (NGS) target enrichment, the method comprising: providing a solid support that includes a plurality of first surface capture primers and second surface capture primers tethered to a surface of the solid support; providing a plurality of nucleic acid target probes that each include a linker and a target capture nucleic acid that hybridizes to and captures a target nucleic acid sequence of a library of single stranded nucleic acid templates at a first annealing temperature (Tai); and providing a plurality of single stranded nucleic acid templates that include target nucleic acid sequences ligated at 5’ and 3’ ends with adapters that include truncated surface primer binding regions, wherein the single stranded nucleic acid templates are configured to hybridize or anneal to the target capture nucleic acids of the nucleic target probes on the surface of the solid support at the first annealing temperature (Tai), and the truncated surface primer binding regions at the 5’ and 3’ ends are configured to hybridize or anneal respectively to the first surface primers or the second surface primers on the surface of the solid support at a second annealing temperature (Ta2) below the first annealing temperature (Tai) and not at the first annealing temperature (Tai).
27. The method of claim 26, wherein the truncated surface primer binding regions have a nucleic sequence that is complementary to a portion and not a full length of a nucleic acid sequence of the first surface primers or the second surface primers.
28. The method of claim 27 or claim 28, wherein the truncated surface primer binding regions have a complementary nucleic sequence at least 2, at least 3, at least 4, or at least 5 bases shorter in length than the first surface primers or the second surface primers.
29. The method of any of claims 27 to 28, wherein the second annealing temperature (Ta2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first annealing temperature (Tai).
30. The method of any of claims 27 to 29, wherein the single stranded nucleic acid templates when hybridized or annealed to the target capture nucleic acids of nucleic target probes on a surface of a solid support has a first melting temperature (Tml), and the truncated surface primer binding regions when hybridized or annealed to the first surface primers or second surface primers has a second melting temperature (Tm2), and wherein the second melting temperature (Tm2) is below the first melting temperature (Tml).
31. The method of claim 30, wherein the second melting temperature (Tm2) is at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
32. The method of claim 30 or claim 31 , wherein the truncated surface primer binding regions do not hybridize or anneal to the first surface primers or second surface primers at the first melting temperature (Tml), preferably Ta2 is below than Tml.
33. The method of any of claims 30 to 32, wherein the second annealing temperature (Ta2) is at least about 2°C, at least about 3°C, at least about 4°C, at least about 5°C, at least about 6°C, at least about 7°C, at least about 8°C, at least about 9°C, or at least about 10°C below the first melting temperature (Tml).
34. The method of any claim 30 to 33, a nucleic acid sequence complementary to full length the first surface primers or second surface primers when hybridized or annealed to the first surface primers or second surface primers has a third melting temperature (Tm3), and wherein the second melting temperature (Tm2) is below the third melting temperature (Tm3).
35. The method of any of claims 30 to 34, further comprising: tethering or linking the plurality of nucleic acid target probes to the surface of solid support; loading the library onto the solid support at or below the first melting temperature (Tml) and above the second melting temperature (Tm2) such that target nucleic acid sequences of the single stranded nucleic acid templates hybridize and are captured by the probes and the truncated surface binding regions do not bind to the first surface capture primers and second surface capture primers; lowering the temperature on the solid support to at or below the second melting temperature (Tm2) such that one of truncated surface primer binding regions of the captured single stranded nucleic acid templates hybridizes to the first surface capture primer or second surface capture primer; and extending the surface capture primers with captured single stranded nucleic acid templates to form a plurality of complementary templates tethered to the surface and hybridized to the plurality of single stranded capture nucleic acid templates.
36. The method of claim 35, further comprising removing captured single stranded nucleic acid templates from the plurality of complementary templates tethered to the surface by, for example, denaturing.
37. The method of claim 36, further comprising washing the denatured captured single stranded nucleic acid from the solid support.
38. The method of claim 36 or claim 37, further comprising amplifying the plurality of complementary templates tethered to provide amplicons or clusters of the complementary templates on the solid support.
39. The method of claim 38, further comprising sequencing the clusters of the complementary templates.
40. The method of any of claims 35 to 39, wherein single stranded nucleic acid templates not captured by the nucleic acid target probes after loading on the solid support are removed from the solid support by, for example, washing the solid support.
41. The method of any of claims 35 to 40, wherein 3’ ends of the nucleic acid target probes are blocked to prevent extension during extension of the surface capture primers.
42. The method of any of claims 35 to 40, wherein the nucleic acid target probes include a nucleic acid linker with a 3’ end ligated to a 5’ end of the target capture nucleic acid.
43. The method of any of claims 38 to 42, wherein the plurality of complementary templates tethered are amplified by bridge amplification.
44. The method of any of claims 26 to 43, wherein the library of nucleic acid strand templates is provided by: providing a plurality of fragmented subsets of double stranded nucleic acids; end-repairing and A-tailing 3’ ends of the fragmented subsets of double stranded nucleic acids; and ligating adapters to the 5’ and 3’ ends of the double stranded nucleic acids, wherein the adapters include the truncated surface primer binding regions.
45. The method of claim 44, further comprising denaturing the ligated double stranded nucleic acids to form the library of single strand nucleic acid templates and loading the single strand nucleic acid templates onto the surface of the solid support of a flow cell.
46. The method of claims 44 or 45, wherein the plurality of fragmented subsets of double stranded nucleic acids include fragmented subsets of double stranded DNA, preferably double stranded genomic DNA.
47. The method of any of claims 44 to 46, wherein the fragmented subsets of double stranded nucleic acids are from about 100 to about 1000 nucleic acids in length.
48. The method of any of claims 44 to 47, wherein the double stranded nucleic acids are end-repaired with polymerase and nucleotide triphosphates (NTPs) to form an end with one A at both 3’ ends.
49. The method of any of claims 26 to 48, wherein ligated adapters further include sequencing primer binding domains and indexes.
50. The method of claim 49, wherein the sequencing primer binding domains include those for the indexes and reading primers from both ends for paired-end sequencing.
51. The method of any of claims 26 to 50, wherein a 5’ end of each of the nucleic acid target probes is tethered to the surface of the solid support prior to loading the library of single stranded nucleic acid templates onto the solid support.
52. The method of any of claims 26 to 50, wherein a 5’ end of each of the nucleic acid target probes includes a surface primer binding region that hybridizes to one of the surface capture primers.
53. The method of claim 52, wherein a 3’ end of each of the surface primer binding regions of the nucleic acid target probes is connected to the linker via blocking region that prevents extension of the nucleic acid target probe during extension of the surface capture primers.
54. The method of claims 52 or 53, wherein the plurality of nucleic acid targeting probes are loaded on the solid support along with the library of single stranded nucleic acids.
55. The method of any of claims 52 to 54, wherein the nucleic target probes hybridized to surface capture primers are removed from the solid support along with the captured single stranded nucleic acid templates by, for example, denaturing the nucleic target probes hybridized to surface capture primers and washing the denatured nucleic acid target probes from the solid support.
Citation Information
Patent Citations
Methods and systems for processing polynucleotides
US20220349003A1
Methods and devices of generating clusters of amplicons
US20240167088A1