Method for obtaining correctly assembled nucleic acids
Through split-merging enzymatic DNA synthesis and high-throughput sequencing technology, the addition of indexed sequences on the beads is solved, and efficient and rapid nucleic acid sequence identification and amplification are achieved.
Patent Information
- Application Number
- CN202480006788.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-06
- Filing Date
- 2024-01-05
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is inefficient in obtaining a perfect copy of long-sequence nucleic acids, and traditional methods are labor-intensive and time-consuming, making it difficult to efficiently screen and amplify the correct nucleic acid sequences.
Using split-merged enzymatic DNA synthesis (EDS) technology, bead-specific index sequences are added to beads, combined with high-throughput sequencing, and the correct nucleic acid sequence is selectively amplified by adaptor ligation and clonal amplification.
It realizes efficient and rapid identification and amplification of the correct nucleic acid sequence, reduces manual operation, and improves the efficiency of obtaining perfect copying of nucleic acid sequences.
Smart Images

Figure CN120418451A_ABST
Abstract
Description
Technical Field
[0001] In some aspects, the present invention relates to assembling perfect sequence nucleic acids from nucleic acid segments. Background of the Invention
[0003] Obtaining perfect sequence copies of nucleic acids is a key problem in synthetic biology. Generally, the process of synthesizing a nucleic acid sequence of interest begins with preparing synthetic oligonucleotides, which are designed to assemble into the desired sequence of interest. Then, a series of laboratory operations are performed to combine the oligonucleotides into a full-length nucleic acid. Errors occur during this process, such that the likelihood of obtaining a perfect copy of the nucleic acid of interest is inversely proportional to the length of the desired nucleic acid sequence. Thus, perfect copies of long sequences are typically produced as a small fraction of the total population of assembled nucleic acids.
[0004] Current state-of-the-art techniques use picking clones from a large population of clones and sequencing individual clones. For example, the assembled nucleic acid is inserted into a plasmid that is expressed in a host (such as Escherichia coli ( E.coli ))). The host cells are plated on a suitable growth medium to produce individual clones. Then, the clone-specific plasmids are sequenced separately to find the plasmids containing perfect sequence inserts of the desired full length. Longer assembled sequences require more clone screening than shorter assembled sequences. The method of picking clones, purifying plasmids, and sequencing each clone individually is labor-intensive and time-consuming. Thus, there is a need for new methods for obtaining perfect copies of nucleic acid sequences of interest. Summary of the Invention
[0005] Cell-free methods, systems, and devices are provided for identifying and amplifying one or more correctly assembled nucleic acids containing a sequence of interest from a nucleic acid mixture. In particular, the methods of the present invention use split-pool enzymatic DNA synthesis (EDS) to add bead-specific index sequences to bead-bound, clonally amplified assembled nucleic acids. High-throughput sequencing is used to identify one or more correctly assembled nucleic acids containing the sequence of interest, wherein each bead-specific index sequence associated with a correctly assembled nucleic acid allows selective amplification of the correctly assembled nucleic acid using a primer having a sequence complementary to the bead-specific index sequence.
[0006] In some embodiments, a cell-free method for identifying and amplifying correctly assembled nucleic acids containing a sequence of interest is provided. The method includes: ligating adapters to the 5' and 3' ends of a plurality of assembled nucleic acids, wherein each adapter contains an adapter primer binding site and an adapter capture sequence; binding the plurality of assembled nucleic acids to a plurality of beads, wherein each bead contains a capture oligonucleotide that specifically hybridizes to the adapter capture sequence, preferably, wherein no more than one assembled nucleic acid binds to each bead; performing clonal amplification of the plurality of assembled nucleic acids bound to the beads using a first primer that hybridizes to the adapter capture sequence and a second primer that hybridizes to the adapter primer binding site, wherein amplicons generated by clonal amplification of the assembled nucleic acids bind to the beads, wherein the beads are separated during clonal amplification such that the amplicons of each bead are separated from all the amplicons of other beads, such that amplicons cannot migrate from one bead to another; adding a bead-specific index sequence to the free 3' end of each clonally amplified assembled nucleic acid and its amplicon using split-and-pool enzymatic DNA synthesis to generate a plurality of indexed assembled nucleic acids bound to the beads, wherein the bead-specific index sequence of each bead contains a unique identifier sequence that is different from the unique identifier sequences of other beads; adding a reverse primer binding site to the free 3' end of each indexed assembled nucleic acid bound to the beads using enzymatic DNA synthesis; amplifying the plurality of indexed assembled nucleic acids bound to the beads using a primer that hybridizes to the adapter primer binding site and a primer that hybridizes to the reverse primer binding site added by enzymatic DNA synthesis; sequencing at least a portion of the amplified indexed assembled nucleic acids to identify correctly assembled nucleic acids containing the sequence of interest; and selectively amplifying at least one correctly assembled nucleic acid containing the sequence of interest using a primer that hybridizes to the bead-specific index sequence of at least one correctly assembled nucleic acid containing the sequence of interest.
[0007] In some preferred embodiments, binding the plurality of assembled nucleic acids to the beads is performed under the condition that no more than one assembled nucleic acid binds to each bead. In some embodiments, the binding of the plurality of assembled nucleic acids to the beads is performed at a nucleic acid / bead ratio ranging from 0.08 to 0.30.
[0008] In some embodiments, the method further includes performing deep sequencing of the amplified indexed assembled nucleic acids on one or more beads.
[0009] In some embodiments, the method further includes adding unique molecular identifiers to assembled nucleic acids having different sequences using EDS.
[0010] In some embodiments, the beads are separated during clonal amplification by separating each bead in a separate emulsion droplet. In some embodiments, the beads are separated during clonal amplification by separating each bead in a separate compartment of a multi-compartment container.
[0011] In some embodiments, the method further includes storing the amplified indexed assembled nucleic acids for a period of time before performing nucleic acid amplification on the correctly assembled nucleic acids containing the target sequence.
[0012] In some embodiments, the method further includes analyzing the sequences of the amplified indexed assembled nucleic acids to identify one or more additional correctly assembled nucleic acids containing the target sequence; and identifying the bead-specific index sequences of the one or more additional correctly assembled nucleic acids containing the target sequence; and performing nucleic acid amplification using primers having sequences that hybridize to the bead-specific index sequences of the one or more additional correctly assembled nucleic acids containing the target sequence, wherein the one or more additional correctly assembled nucleic acids containing the target sequence are selectively amplified.
[0013] In some embodiments, the adaptor is a linear adaptor, a Y-shaped adaptor, a stubby adaptor, or a hairpin adaptor. In some embodiments, a Y-shaped adaptor is used, wherein the double-stranded stem of the Y-shaped adaptor is ligated to the 5' and 3' ends of multiple assembled nucleic acids.
[0014] In some embodiments, the adaptor further includes a barcode.
[0015] In some embodiments, the method further includes removing the adaptor from the correctly assembled nucleic acids containing the target sequence. In some embodiments, the assembled nucleic acids contain restriction enzyme cleavage sites, and the ligated adaptor can be removed by cleavage with an endonuclease at the restriction enzyme cleavage site.
[0016] In some embodiments, the method further includes removing the unique identifier sequence from the correctly assembled nucleic acids containing the target sequence.
[0017] In some embodiments, multiple assembled nucleic acids are assembled using an isothermal assembly method. In some embodiments, the isothermal method uses an exonuclease, a DNA polymerase, or a ligase or any combination thereof. Exemplary assembly methods include, but are not limited to, Gibson assembly, polymerase chain assembly (PCA), Goldengate assembly, BioBrick assembly, or oligonucleotide ligation.
[0018] In some embodiments, performing clonal amplification includes performing emulsion polymerase chain reaction.
[0019] In some embodiments, the beads are magnetic beads.
[0020] In some embodiments, the method further includes preparing multiple assembled nucleic acids for ligation of the adaptor by blunt-ending the 3' and 5' ends of the assembled nucleic acids, phosphorylating the 5' end of the assembled nucleic acids, or adding a poly A tail to the 3' end of the assembled nucleic acids or a combination thereof.
[0021] In some embodiments, prior to the ligation linker, the method further comprises: assembling a plurality of oligonucleotides into a plurality of polynucleotides to produce a plurality of assembled nucleic acids, wherein each oligonucleotide comprises a portion of a target sequence. In some embodiments, the oligonucleotides range in length from 20 base pairs to 400 base pairs, including any length within that range, such as lengths of 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 base pairs. In some embodiments, the oligonucleotides are less than 300 base pairs in length, less than 200 base pairs in length, less than 150 base pairs in length, less than 100 base pairs in length, or less than 70 base pairs in length. In some embodiments, the oligonucleotides are more than 40 base pairs in length, more than 50 base pairs in length, more than 75 base pairs in length, more than 100 base pairs in length, or more than 200 base pairs in length. In some embodiments, the oligonucleotides have a length ranging from about 60 base pairs to about 80 base pairs, including any length within that range, such as 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 base pairs. In some embodiments, the oligonucleotides have a length ranging from 20 base pairs to 50 base pairs, including any length within that range, such as 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 base pairs.
[0022] In some embodiments, the nucleic acid comprising the target sequence includes an endogenous gene sequence encoding a protein or RNA, an exon, an intron, a wild-type sequence, a major allele, a minor allele, an allele associated with a disease, a nucleic acid comprising a mutation (e.g., a deletion, an insertion, a substitution, a single nucleotide polymorphism, a frameshift mutation, a missense mutation, or a nonsense mutation), or an engineered nucleic acid.
[0023] In some embodiments, sequencing is performed using next-generation sequencing (NGS) methods. Exemplary NGS methods include, but are not limited to, second-generation NGS technologies such as those sold by Illumina, or Ion Torrent Technology from ThermoFisher, systems from Gynapsys, Omniome, Element Biosciences, SingularGenomics, third-generation NGS technologies (such as single-molecule real-time (SMRT) sequencing from Pacific Biosystems), and fourth-generation single-molecule NGS technologies such as nanopore sequencing from Oxford Nanopore.
[0024] In some embodiments, sequencing includes sequencing adapter sequences, bead-specific index sequences, and assembled nucleic acid sequences.
[0025] In another aspect, provided is a system that includes: a device for performing split-and-pool enzymatic DNA synthesis (EDS), a processor, and a sequencer, wherein the processor is programmed to instruct the device to perform split-and-pool EDS to add (i) a bead-specific index sequence and (ii) a reverse priming site to the free 3' end of bead-bound clonally amplified assembled nucleic acids, and wherein the sequencer is connected to the processor such that the processor receives sequencing data from the sequencer.
[0026] In some embodiments, the processor is provided by a computer or a handheld device (e.g., a cell phone or a tablet).
[0027] In some embodiments, the sequencer is inside a device for performing split-and-pool EDS.
[0028] In some embodiments, the processor is further programmed to receive the sequence of a nucleic acid containing a target sequence; design the sequences of a plurality of polynucleotides that can be assembled into a nucleic acid containing the target sequence, wherein each polynucleotide contains a portion of the target sequence; and instruct a synthesizer to synthesize the plurality of polynucleotides.
[0029] In some embodiments, the system further includes a chamber for assembling the plurality of polynucleotides. In some embodiments, the chamber further contains an exonuclease, a DNA polymerase, and a ligase for performing the assembly of the plurality of polynucleotides.
[0030] In some embodiments, the processor is further programmed to instruct the device to perform the assembly of the plurality of polynucleotides.
[0031] In some embodiments, the system further includes a memory storage component operably connected to the processor, wherein the memory storage component is configured to store sequencing data.
[0032] In some embodiments, the system further comprises a plurality of Y-junction adapters, a set of polymerase chain reaction (PCR) primers having sequences complementary to the sequences in the 5' and 3' arms of the Y-junction adapter, a PCR primer having a sequence complementary to the reverse primer binding site, reagents for performing emulsion PCR, beads, or ligases, or any combination thereof.
[0033] In some embodiments, the system further comprises magnetic beads.
[0034] In some embodiments, the system further comprises reagents for preparing a plurality of assembled nucleic acids for ligation of adapters, wherein the reagents comprise reagents for blunt-ending the 3' and 5' ends of the assembled nucleic acids, phosphorylating the 5' end of the assembled nucleic acids, or adding a poly A tail to the 3' end of the assembled nucleic acids, or a combination thereof.
[0035] In some embodiments, the system further comprises reagents for performing gene assembly. In some embodiments, the system comprises reagents for performing Gibson assembly, polymerase chain assembly, Goldengate assembly, BioBrick assembly, or oligonucleotide ligation.
[0036] In another aspect, there is provided a kit comprising the system described herein, a packaging for the system, and instructions for using the system to cell-free assemble a nucleic acid comprising a target sequence and identify the correctly assembled nucleic acid comprising the target sequence.
[0037] In another aspect, there is provided a computer-implemented method for assembling a nucleic acid comprising a target sequence, the computer performing steps comprising: receiving the sequences of nucleic acids to be assembled from a plurality of polynucleotides; designing the sequences of the plurality of polynucleotides, wherein each polynucleotide comprises a portion of the sequence of the nucleic acid comprising the target sequence, wherein the plurality of polynucleotides can be assembled into the full-length sequence of the nucleic acid comprising the target sequence; instructing a synthesizer to synthesize the plurality of polynucleotides; instructing a device to add a bead-specific index sequence and a reverse primer site to the free 3' ends of the plurality of bead-bound assembled nucleic acids, the plurality of bead-bound assembled nucleic acids being generated by clonal amplification from the plurality of polynucleotides using split-and-pool enzymatic DNA synthesis (EDS), wherein the bead-specific index sequence of each bead comprises a different unique identifier sequence; receiving the indexed assembled nucleic acid sequences; analyzing the sequences of the indexed assembled nucleic acids to identify the bead-specific index sequences of the correctly assembled nucleic acids comprising the target sequence; and displaying the bead-specific index sequences of the correctly assembled nucleic acids comprising the target sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figures 1 to 5 : Schematic overview of the cell-free identification protocol for correctly assembled nucleic acids.Figure 1 Shows the assembly steps, including Assembly 1 (oligonucleotides assembled into small fragments or exons) and Assembly 2 (small fragments assembled into genes). Figure 2 Shows the adaptor ligation and cloning amplification steps. Figure 3 Shows the addition of bead-specific barcodes to the assembled nucleic acids using split and merge EDS. Figure 4 Shows the addition of PCR primer sites to the indexed assembled nucleic acids using EDS. Figure 5 Shows NGS sequencing to identify the bead-specific barcodes of the assembled nucleic acids with the correct sequence, and selective amplification with primers complementary to the bead-specific barcodes. Figure 6 Schematically depicts the split and merge method, which is illustrated by adding three different alternatives after the split step and before the merge step (each step adds a light grey circle, a dark grey circle, or a black circle to the growing strand constructed from the starting point (i.e., the white circle)). Detailed Description
[0039] Methods, systems, and devices are provided for cell-free identification and amplification of correctly assembled nucleic acids that contain a target sequence from a nucleic acid mixture. The subject methods use split-merge enzymatic DNA synthesis (EDS) to add bead-specific index sequences to bead-bound, clonally amplified assembled nucleic acids. The assembled nucleic acids are sequenced to identify one or more correctly assembled nucleic acids that contain the target sequence, wherein the bead-specific index sequence of each correctly assembled nucleic acid allows for selective amplification of the correctly assembled nucleic acid from the pooled sequencing library by using a primer having a sequence complementary to the bead-specific index sequence.
[0040] The invention is not strictly limited to the specific methods, devices, or compositions described herein, as some variations can also be implemented. Additionally, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, as the scope of the invention will be defined only by the authorized claims.
[0041] Where a numerical range is provided, it is understood that each intermediate value, to the tenth of the unit of the lower limit, between the upper and lower limits of the range is also specifically disclosed. Each smaller range between any of the stated numerical values or intermediate values within the stated range is included within the invention. The upper and lower limits of these smaller ranges can be independently included in or excluded from the range, and any range that includes either, neither, or both of the limits of the smaller ranges is also included within the invention, subject to any specific exclusions stated within the range. Where the range includes one or both of the limits, ranges excluding one or both of the included limits are also included in the invention.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferred methods and materials are described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials related to the cited publications. It should be understood that, in case of conflict, the present disclosure supersedes any disclosure of the incorporated publications.
[0043] As will be apparent to those skilled in the art upon reading the present disclosure, each of the various embodiments described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any one of several other embodiments without departing from the scope or spirit of the present invention. Any of the methods described may be performed in the order of the recited events or in any other order that is logically possible.
[0044] The publications discussed herein are provided solely for the purpose of illustrating their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such disclosure by virtue of a prior invention. Further, the provided publication dates may be different from the actual publication dates that may need to be independently confirmed.
[0045] Definitions
[0046] The term "about", particularly when referring to a quantity, means including a deviation of ±5%.
[0047] As used herein, the terms "polynucleotide", "oligonucleotide", "nucleic acid", and "nucleic acid molecule" include polymeric forms of nucleotides (such as ribonucleotides and / or deoxyribonucleotides) of any length. Thus, these terms can be used to refer to triple-stranded, double-stranded, and single-stranded forms of DNA, as well as triple-stranded, double-stranded, and single-stranded forms of RNA. These terms also include modified and unmodified forms (such as by methylation and / or by capping). More specifically, the terms "polynucleotide", "oligonucleotide", "nucleic acid", and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), any other type of polynucleotide (which is an N- or C-glycoside of a purine or pyrimidine base), and other polymers containing non-nucleotide backbones, such as polyamides (e.g., peptide nucleic acids (PNAs)) and polymorpholino oligonucleotides (commercially available from AntiVirals, Inc., Corvallis, Oregon, Neugene), and other synthetic sequence-specific nucleic acid polymers, provided that the polymers contain configurations of nucleobases that permit base pairing and base stacking as found in DNA and RNA. The terms "polynucleotide", "oligonucleotide", "nucleic acid", and "nucleic acid molecule" are not deliberately distinguished by length, and these terms will be used interchangeably unless they have a specific meaning based on their context of use. Thus, these terms include, for example, 3'-deoxy-2',5'-DNA, oligodeoxyribonucleotide N3' P5' phosphoramidates, 2'-O-alkyl-substituted RNAs, double-stranded DNA and single-stranded DNA, as well as double-stranded RNA and single-stranded RNA, DNA:RNA hybrids, and hybrids between PNA and DNA or RNA, and also include known types of modifications, such as, for example, labels known in the art, methylation, "capping", replacement of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as those having uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), those having negatively charged linkages (e.g., phosphorothioates, dithiophosphates, etc.), and those having positively charged linkages (e.g., aminoalkylphosphoramidates, aminoalkylphosphotriesters), those containing pendant moiety modifications, such as, for example, proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those containing intercalator modifications (e.g., acridine, psoralen, etc.), those containing chelator modifications (e.g., metals, radioactive metals, boron, oxidized metals, etc.), those containing alkylator modifications, those having modified linkages (e.g., α-anomeric nucleic acids, etc.), and unmodified forms of polynucleotides or oligonucleotides. In particular, DNA is deoxyribonucleic acid.
[0048] As used herein, the term "target nucleic acid" refers to a nucleic acid molecule containing a "target sequence" to be amplified. The target nucleic acid can be single-stranded or double-stranded and can include additional nucleotides on one or both sides of the target sequence. The term "target sequence" refers to a specific nucleotide sequence of the target nucleic acid to be amplified. In some embodiments, the target sequence includes a probe hybridization region to which a probe forms a stable hybridization under suitable hybridization conditions. The "target sequence" can also include a sequence complementary to one or more primers, which can be extended by a polymerase or ligase using the target sequence as a template. If the target nucleic acid is initially single-stranded (e.g., the "sense strand"), the term "target sequence" also refers to the sequence complementary to the "target sequence" present in the target nucleic acid. If the "target nucleic acid" is initially double-stranded, the term "target sequence" refers to the positive (+) strand and the negative (-) strand (or the sense strand and the antisense strand).
[0049] As used herein, the term "primer" or "oligonucleotide primer" refers to an oligonucleotide that hybridizes to a nucleic acid template strand and initiates the synthesis of a nucleic acid strand complementary to the template strand in the presence of nucleotides and a polymerization inducer such as DNA or RNA polymerase, and under suitable temperature, pH, metal concentration, and salt conditions. Primers are typically single-stranded, but can also be double-stranded or contain double-stranded portions. If double-stranded, the primer can be treated to separate the strands before being used to form a primer extension product. This denaturation step is typically carried out by heating, but can also be done using a base followed by pH neutralization. Thus, a "primer" is complementary to a template sequence, and the hybridization complex forms a primer / template complex by hydrogen bonding or hybridization to the template, and primer extension is initiated by a polymerase to add nucleotides to the 3' end of the primer complementary to the corresponding template sequence. Typically, nucleic acids are amplified using at least one forward primer and at least one reverse primer that are capable of hybridizing to nucleic acid regions flanking a portion of the nucleic acid sequence to be amplified.
[0050] The term "amplicon" refers to an amplified nucleic acid product (e.g., rolling circle amplification and isothermal amplification methods such as recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), and nicking enzyme amplification reaction (NEAR), etc.) produced by a PCR reaction or other nucleic acid amplification process that generates copies of a single-stranded or double-stranded nucleic acid sequence. The amplicon can include RNA or DNA, depending on the technique used for amplification.
[0051] The term "hybridize / hybridization" refers to the formation of a complex between nucleotide sequences that are sufficiently complementary to form a complex through Watson-Crick base pairing. When a primer "hybridizes" to a target (template), such a complex (hybrid) is stable enough to perform the priming function required to initiate DNA synthesis (e.g., by DNA polymerase). It should be understood that the sequences do not need to have perfect complementarity to provide a stable hybrid. In many cases, stable hybrids can be formed where there are fewer than about 10% mismatched bases, ignoring loops of four or more nucleotides. Thus, as used herein, the term "complementary" refers to an oligonucleotide that forms a stable duplex with its "complement" under appropriate conditions, typically having about 90% or greater homology.
[0052] The "melting temperature" or "Tm" of double-stranded DNA is defined as the temperature at which half of the helical structure of the DNA is lost due to heating or other dissociating treatments of the hydrogen bonds between base pairs (e.g., by alkali treatment, etc.). The Tm of a DNA molecule depends on its length and base composition. A DNA molecule with a high GC base pair content has a higher Tm than a DNA molecule with a lower GC base pair content. When the temperature is lowered below the Tm, the separated DNA complementary strands spontaneously re-associate or anneal to form double-stranded DNA. The Tm can be estimated using the following relationship: Tm = 69.3 + 0.41 (GC)% (Marmur et al. (1962) J. Mol. Biol. 5 :109-118).
[0053] The term "Y-shaped adaptor" refers to an adaptor having a Y-shaped structure that comprises two partially complementary strands that hybridize to each other such that the adaptor has a stem region and two single-stranded arms. Each of the two strands in the Y-shaped adaptor has a 5'-end region and a 3'-end region such that the 5'-end region of the first strand is complementary to and hybridizes with the 3'-end region of the other (second) strand, forming a double-stranded stem region ("stem"). The 3'-end region of the first strand and the 5'-end region of the second strand provide single-stranded 3'- and 5'-arms as they are not complementary to each other and thus they do not substantially hybridize with each other. By ligating the free 3'- and 5'-ends of the complementary strand regions of the double-stranded stem of the Y-shaped adaptor to the 5'- and 3'-ends of each nucleic acid, the Y-shaped adaptor can be ligated to a double-stranded nucleic acid containing a target sequence of interest. The use of Y-shaped adaptors allows different adaptor sequences to be added simultaneously to the 5'- and 3'-ends of double-stranded nucleic acids in a DNA library.
[0054] The terms "hairpin adaptor", "circular adaptor", and "hairpin-loop adaptor" refer to adaptors that include a hairpin-loop structure. After the hairpin adaptor is ligated to one or more ends of a double-stranded nucleic acid, the hairpin loop can be cleaved to generate strands having non-complementary sequences at the ends, thereby generating a Y-shaped adaptor having a single-stranded 5' arm and a single-stranded 3' arm. In some cases, the loop of the hairpin adaptor can contain uracil residues, where the loop can be cleaved using uracil DNA glycosylase and endonuclease VIII.
[0055] The terms "connected" or "coupled" are used in an operative sense and need not be limited to direct connection or coupling. For example, two devices or components can be directly connected, or connected via one or more intermediate media or devices. As another example, devices can be connected in such a way that information or data can be transferred between them while not sharing any physical connection with each other. In some cases, two devices or components can be connected to each other by wired or wireless means.
[0056] Cell-free methods for identifying and amplifying correctly assembled nucleic acids
[0057] The methods, systems, and devices described herein allow for cell-free identification and amplification of correctly assembled nucleic acids, or one or more correctly assembled nucleic acids, that contain a target sequence from a nucleic acid mixture. The methods generally include: ligating adapters to the 5' and 3' ends of a plurality of assembled nucleic acids, wherein the adapters contain adapter primer binding sites and adapter capture sequences; binding the plurality of assembled nucleic acids to beads, wherein the beads contain capture oligonucleotides attached to the bead surface, the capture oligonucleotides being complementary to and specifically hybridizing with the adapter capture sequences; performing clonal amplification of the plurality of assembled nucleic acids bound to the beads using (i) the capture oligonucleotides attached to the beads as a first primer and (ii) a second primer that hybridizes to the adapter primer binding site, wherein the amplicons generated by the clonal amplification of the assembled nucleic acids bind to the beads, and wherein the beads are partitioned during clonal amplification such that the amplicons on each bead are separated from the other amplicons on all other beads, such that amplicons cannot migrate from one bead to another; adding bead-specific index sequences to the free 3' ends of each clonally amplified assembled nucleic acid and its amplicons using split-and-pool enzymatic DNA synthesis (EDS) to produce a plurality of indexed assembled nucleic acids bound to the beads, wherein the bead-specific index sequence for each bead contains a different unique identifier sequence; adding a reverse priming site to the free 3' ends of each indexed assembled nucleic acid bound to the beads using EDS; amplifying the plurality of indexed assembled nucleic acids bound to the beads using a primer that hybridizes to the adapter primer binding site and a primer that hybridizes to the reverse priming site added by EDS; sequencing at least a portion of the amplified indexed assembled nucleic acids to identify the correctly assembled nucleic acids that contain the target sequence; identifying the bead-specific index sequences of the correctly assembled nucleic acids that contain the target sequence; and amplifying the nucleic acids using a primer that hybridizes to the bead-specific index sequence of the correctly assembled nucleic acids that contain the target sequence, wherein the correctly assembled nucleic acids that contain the target sequence are selectively amplified.
[0058] In some embodiments, the beads are partitioned during clonal amplification by separating each bead in a separate emulsion droplet or in a separate compartment of a multi-compartment container.
[0059] In some embodiments, the method further includes adding unique molecular identifiers to the assembled nucleic acids having different sequences using enzymatic DNA synthesis.
[0060] In some embodiments, the method further includes storing the amplified indexed assembled nucleic acids for a period of time before performing nucleic acid amplification of the correctly assembled nucleic acids that contain the target sequence.
[0061] In some embodiments, the method further comprises analyzing the sequences of the amplified indexed assembled nucleic acids to identify one or more additional correctly assembled nucleic acids that contain the sequence of interest; and identifying the bead-specific index sequences of the one or more additional correctly assembled nucleic acids that contain the sequence of interest; and performing nucleic acid amplification using primers that have sequences that are complementary to and hybridize with the bead-specific index sequences of the one or more additional correctly assembled nucleic acids that contain the sequence of interest, wherein the one or more additional correctly assembled nucleic acids that contain the sequence of interest are selectively amplified.
[0062] In some embodiments, prior to ligating the adaptor, the method further comprises generating a plurality of assembled nucleic acids. Nucleic acids can be assembled from multiple nucleic acid fragments comprising a portion of the target sequence using different methods, as further described below. In some cases, the nucleic acid comprising the target sequence is assembled from a plurality of oligonucleotides each comprising a different portion of the target sequence. In some embodiments, the oligonucleotides used for assembling the nucleic acids range in length from 20 base pairs to 400 base pairs, including any length within that range, such as 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 base pairs in length. In some embodiments, the oligonucleotides are less than 300 base pairs in length, less than 200 base pairs in length, less than 150 base pairs in length, less than 100 base pairs in length, or less than 70 base pairs in length. In some embodiments, the oligonucleotides are more than 40 base pairs in length, more than 50 base pairs in length, more than 75 base pairs in length, more than 100 base pairs in length, or more than 200 base pairs in length. In some embodiments, the oligonucleotides have a length ranging from about 60 base pairs to about 80 base pairs, including any length within that range, such as 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 base pairs. In some embodiments, the oligonucleotides have a length ranging from 20 base pairs to 50 base pairs, including any length within that range, such as 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 base pairs.
[0063] Beads for binding and assembling nucleic acids can be magnetic beads (e.g., superparamagnetic beads) or non-magnetic beads. Magnetic beads have the advantage of allowing the separation of the bound nucleic acids by magnetic separation techniques. In some embodiments, the beads further comprise a coating. For example, the surface of the beads can be coated or functionalized with silica to facilitate the binding of the assembled nucleic acids. In some embodiments, a capture oligonucleotide is attached to the surface of the bead, wherein the capture oligonucleotide has a sequence complementary to a universal adapter sequence or a sequence present in all or part of the assembled nucleic acid. Amplifying the assembled nucleic acid using the capture oligonucleotide as a primer allows the newly synthesized strand to be covalently linked to the bead, thereby immortalizing the strand and generating a resource that can support the amplification of perfect copies at any time.
[0064] The binding of the assembled nucleic acid to the beads is preferably carried out under suitable conditions to allow no more than one assembled nucleic acid to bind to each bead. The solution containing the assembled nucleic acid can be diluted to provide a nucleic acid / bead ratio such that no more than one assembled nucleic acid binds to each bead. If the nucleic acid / bead ratio is too high, some beads will have multiple different assembled nucleic acids bound to them. If the nucleic acid / bead ratio is too low, many beads will not have any assembled nucleic acid bound to them. The nucleic acid / bead ratio can be optimized such that most beads contain a single bound assembled nucleic acid. In some embodiments, the nucleic acid / bead ratio ranges from 0.08 to 0.30, including any ratio within this range, such as 0.08, 0.081, 0.082, 0.083, 0.084, 0.085, 0.09, 0.095, 0.1, 0.2, or 0.3. In some embodiments, the nucleic acid / bead ratio ranges from 0.08 to 0.09.
[0065] The adapter and index sequences can be removed from the final correctly assembled nucleic acid product by different methods. For example, PCR can be carried out with primers designed to amplify only the target sequence and not the other added sequences. Alternatively, restriction enzyme cleavage sites can be added to allow the cleavage of the adapter from the target sequence by a restriction endonuclease. In the case of using Gibson assembly, the adapter can be eliminated by using the adapter as a target to be removed by recombination with a plasmid.
[0066] The nucleic acids assembled by the methods described herein can contain any type of target sequence. In some embodiments, the assembled nucleic acid contains an endogenous gene, exon, intron, wild-type sequence, major allele, minor allele, disease-related allele, nucleic acid containing a mutation (e.g., deletion, insertion, substitution, single nucleotide polymorphism, frameshift mutation, missense mutation, or nonsense mutation), or engineered nucleic acid that encodes a protein or RNA.
[0067] Assembly
[0068] The assembly of nucleic acids typically proceeds in at least two steps: First, short DNA fragments (e.g., oligonucleotides) containing portions of the desired nucleic acid sequence are assembled into longer DNA fragments (e.g., partially assembled nucleic acid intermediates, which may also be referred to as "polynucleotides" in this example); Next, the longer DNA fragments are assembled to produce the complete nucleic acid. In some cases, the assembly of nucleic acids includes multiple steps, with each assembly step producing multiple partially assembled nucleic acid intermediates of gradually increasing length. DNA fragments with overlapping sequences (e.g., oligonucleotides or polynucleotides) can be joined by annealing and ligation and / or polymerase-mediated reactions. For example, DNA fragments can be ligated together without any template, as is typically required by DNA polymerase. Endonucleases that cleave nucleic acids into fragments can be used to direct DNA assembly. Any suitable method for assembling nucleic acids from multiple DNA fragments can be used to produce the complete nucleic acid containing the sequence of interest. Exemplary methods for nucleic acid assembly from DNA fragments include, but are not limited to, Gibson assembly, polymerase chain assembly (PCA), Golden Gate assembly, BioBrick assembly, and oligonucleotide ligation.
[0069] Gibson assembly can be used to combine DNA fragments under isothermal conditions using exonuclease, DNA polymerase, and ligase. Adjacent DNA fragments must contain overlapping regions (usually about 20 to 40 base pairs). The overlapping regions can be added to adjacent DNA fragments, for example, using PCR or EDS. First, DNA is cleaved from the 5' end of the DNA fragment using 5' exonuclease, generating single-stranded overhangs. Next, the single-stranded regions on adjacent DNA fragments are annealed, and nucleotides are added using DNA polymerase to fill any gaps in the annealed single-stranded regions. Then, DNA ligase is used to covalently join the adjacent DNA fragments and seal any nicks in the annealed DNA. Compared to conventional cloning methods, Gibson assembly has the advantage that there is no need for restriction enzyme digestion of DNA fragments after PCR. For a description of Gibson assembly, see, for example, Gibson et al. (2009) Nat.Methods 6:343-345, Gibson et al. (2010) Science 329, 52-56; the above is incorporated herein by reference.
[0070] Polymerase chain assembly (also known as polymerase cycling assembly or PCA) is similar to PCR but uses multiple single-stranded overlapping oligonucleotides that anneal to each other. Each oligonucleotide is designed to be part of the sense or antisense strand of the double-stranded target DNA to be assembled. The single-stranded oligonucleotides are designed to cover most of the sequence of both strands of the full-length nucleic acid sequence to be assembled. The overlapping end regions of the oligonucleotides must have sufficient complementarity to assemble into the final complete sequence. Preferably, when annealed to the complementary paired strand, the complementary overlapping end regions have similar melting temperatures, are hairpin-free, and are not GC-rich. Forward and reverse primers are used to initiate, allowing DNA polymerase to fill in the entire template sequence. Primers used for PCA are longer than those used for PCR, typically having a length of 40 to 50 nucleotides to ensure proper hybridization. During the PCA cycle, the oligonucleotides anneal to complementary oligonucleotides and are then extended by polymerase in a template-dependent manner until the 5' end of the template strand is reached. Each cycle increases the length of the different oligonucleotides that anneal to each other. Forward and reverse PCR primers are used to selectively amplify the complete target sequence from the product mixture, which also contains shorter incomplete fragments. The complete target sequence can then be separated using gel electrophoresis or chromatography. A typical reaction consists of oligonucleotides with lengths of about 40 to 50 base pairs, where each oligonucleotide has an overlapping region of about 20 base pairs. The PCA reaction is carried out with all oligonucleotides, typically for about 30 cycles, followed by about 23 additional cycles with PCR primers that selectively amplify the desired assembled complete DNA product. For a description of PCA, see, e.g., Stemmer; et al. (1995) Gene 164(1):49-53, Smith et al. (2003) Proc. Natl. Acad. Sci. USA 100(26):15440-15445, Marchand et al. (2012) Methods Mol. Biol. 8:3-10; the foregoing is incorporated herein by reference.
[0071] Golden Gate assembly uses type IIS restriction endonucleases and DNA ligases (such as T4 DNA ligase) to directionally assemble multiple DNA fragments into a complete DNA construct. For Golden Gate assembly, type IIS restriction endonucleases that generate 4-base overhangs are typically used. Exemplary restriction endonucleases that can be used for Golden Gate assembly include, but are not limited to, BsaI, BsmBI, BsmBI-v2, PaqCI, BsaI-HF, BbsI, BbsI-HF, and Esp3I. The type IIS restriction endonuclease selected should cut DNA outside its recognition site to generate non-palindromic overhangs. Using such type IIS restriction endonucleases allows Golden Gate assembly to be scarless because the final ligation product does not retain the type IIS restriction endonuclease recognition site and cannot be cut again by the restriction endonuclease for assembly. The type IIS restriction endonuclease is selected to generate a combination of overhang sequences that allow the directional assembly of multiple DNA fragments, which can then be ligated together to produce the final assembled DNA product. In some cases, multiple DNA fragments are generated with the same type IIS restriction endonuclease. For example, BsaI can generate 256 different four-base pair overhangs depending on the sequence of the DNA. The overhangs of the DNA fragments should be designed so that the DNA fragments can be ligated together without a scar sequence between them. The vectors used in Golden Gate assembly contain the DNA fragments for assembly, flanked by type IIS restriction endonuclease recognition sites. After digestion with the type IIS restriction endonuclease, each DNA fragment for assembly has a unique overhang that anneals and ligates to the next DNA fragment, generating the complete assembled DNA. The sequence of the DNA fragment overhangs allows multiple fragments to be assembled simultaneously in the correct order to produce the final assembled DNA product. Because no restriction sites are present in the final ligation product, digestion and ligation can be performed simultaneously. Variants of Golden Gate assembly include MoClo and Golden Braid assembly, which require DNA multi-layer assembly from components constructed in each round of assembly. These methods can be used to assemble multi-gene or multi-part constructs with multiple transcriptional units.For a description of Golden Gate assembly and its variants, see, for example, Engler, et al. (2014) "Golden Gate cloning" Methods in Molecular Biology 1116:119-131, Engler et al. (2008) PLoS ONE 3:e3647, Engler, C., et al. (2009) PLoS ONE 4:e5553, Weber et al. (2011) PLOS ONE 6(2): e16765, Lee et al, (1996) Genet. Anal. 13(6):139-145, Padgett et al. (1996) Gene 168(1):31-35, Werner, S., et al. (2012) Bioeng Bugs 3(1):38-43, Andreou et al. (2018) PLoS ONE 13(1):e0189892, Pryor et al. (2020) PLOS ONE.15 (9): e0238592, Wu et al. (2018) Mol. Plant Pathol. 19(6):1511-1522, Moore et al. (2016) ACS Synthetic Biology. 5 (10):1059-1069, Crozet et al. (2018) ACS Synthetic Biology. 7 (9): 2074-2086; the foregoing is incorporated herein by reference.
[0072] BioBrick assembly uses DNA sequences that conform to the Restriction Enzyme Assembly Standard. Each BioBrick part has a DNA sequence contained in a circular plasmid. Flanking the BioBrick parts are universal upstream and downstream sequences that contain restriction sites for specific restriction enzymes, with at least two of the restriction enzymes being isocaudomers. By using restriction enzymes that cut at the recognition sites to remove the restriction sites between the parts, the BioBrick parts can be ligated together in any desired order. The BioBrick parts do not contain any restriction sites in the universal upstream and downstream sequences. BioBrick assembly can be used for the assembly of any DNA molecule, but is particularly suitable for applications in synthetic biology to combine DNA fragments with different functions. For a description of the BioBrick standard and assembly methods, see, e.g., Knight (2003). "Idempotent Vector Design for Standard Assembly of Biobricks" (hdl:1721.1 / 21168), Knight et al. (2008) J. Biol. Eng. 2:5, Shetty, et al. (2008). J BiolEng. 2008 2:5, Shetty et al. (2011) Methods Enzymol. 498:311-326, Zucca et al.(2013) J. Biol. Eng. 7(1):12, Yamazaki et al. (2017) Synth Biol (Oxf) 2(1):ysx003, Smolke et al. (2009) Nature Biotechnology. 27 (12):1099-1102, Sleightet al. (2010) Nucleic Acids Research. 38 (8): 2624-2636, and Rokke et al.(2014) “BioBrick Assembly Standards and Techniques and Associated SoftwareTools” Methods in Molecular Biology Vol. 1116, Humana Press pp. 1-24; the foregoing is incorporated herein by reference.
[0073] Adapter
[0074] Adapter oligonucleotides containing known sequences are added to the 5' and 3' ends of nucleic acids to facilitate amplification and / or sequencing. Adapters can be designed with primer binding sites having sequences suitable for hybridizing with primers for primer-dependent amplification and / or sequencing. The adapter can also include sites that allow the nucleic acid to attach to a solid support. For facilitating multiplex detection, the adapters can be barcoded.
[0075] Ligases can be used to ligate adapter oligonucleotides to the ends of nucleic acids. Any suitable ligase can be used, including but not limited to phage ligases (such as T4 or T7 DNA ligase), archaeal ligases, or bacterial ligases. The adapters ligated to either end of the nucleic acid can be the same or different. The adapter can be attached to the end of DNA by blunt-end ligation or sticky-end ligation. Ligation of the ends of two DNA molecules involves the formation of a phosphodiester bond between the 3'-hydroxyl group at the 3' end of one DNA molecule and the 5'-phosphoryl group at the 5' end of the other DNA molecule. The ends of the DNA molecule can be prepared for ligation by blunt-ending the DNA ends and phosphorylating the 5' end. Blunt-ending involves removing single-stranded overhangs (e.g., which can be generated by restriction endonucleases) by adding nucleotides to the complementary strand using the overhang as a polymerization template, or removing the overhang using an exonuclease. The DNA ends can be "blunt-ended" to allow the joining of incompatible ends by ligation. DNA polymerases, such as the Klenow fragment of DNA polymerase I and T4 DNA polymerase, can be used to fill in nucleotides or digest 3' overhangs. Nucleases (such as mung bean nuclease) can be used to remove 5' overhangs. In some cases, the adapter is designed to have a poly T overhang, which allows the adapter to be ligated to the 3' end of DNA having a poly A overhang. For example, by treating the DNA with T4 polynucleotide kinase, T4 DNA polymerase, and the Klenow large fragment, the ends of the nucleic acid can be blunt-ended and phosphorylated at the 5' end. A poly A tail can be added to the 3' end of the nucleic acid using Taq polymerase or the Klenow large fragment.
[0076] Solid-phase amplification of polynucleotides is typically carried out by first ligating known adapter sequences to each end of the target polynucleotide. The double-stranded polynucleotide is then denatured to form single-stranded template molecules immobilized on a solid support (e.g., the flow cell surface of the Illumina platform or the beads of the Ion Torrent platform). The adapter sequence on the 3' end of the template hybridizes to an extension primer and is amplified by the extension primer. In some aspects, the sequencing platform adapter construct includes one or more nucleic acid domains selected from: domains that specifically bind surface-attached sequencing platform oligonucleotides (e.g., P5 or P7 oligonucleotides attached to the flow cell surface of an Illumina® sequencing system) (e.g., a "capture site" or "capture sequence"); a sequencing primer binding domain (e.g., a domain to which the Read 1 or Read 2 primer of the Illumina® platform can bind); a barcode domain (e.g., a domain that can uniquely identify the source of the nucleic acid sample being sequenced to enable sample multiplexing by labeling each molecule from a given sample with a specific barcode or "tag"); a barcode sequencing primer binding domain (a domain that binds a primer for sequencing the barcode); or any combination of these domains.
[0077] The adapters selected for preparing the sequencing library should be compatible with the sequencing system to be used. The polynucleotide is incorporated into the sequencing library by ligation to a sequencing adapter that contains specific sequences designed to work with the sequencing platform (i.e., the sequencing platform adapter domain). When present in the adapter, the sequencing platform adapter domain can include one or more nucleic acid domains of any length and with sequences suitable for the intended sequencing platform. In some embodiments, the length of the nucleic acid domain is from 4 to 200 nucleotides. For example, the length of the nucleic acid domain can be from 4 to 100 nucleotides, such as 6 to 75 nucleotides, 8 to 50 nucleotides, or 10 to 40 nucleotides. According to some embodiments, the sequencing platform adapter construct includes a nucleic acid domain having a length of 2 to 8 nucleotides, such as a length of 9 to 15 nucleotides, 16 to 22 nucleotides, 23 to 29 nucleotides, or 30 to 36 nucleotides.
[0078] The nucleic acid domain can have a length and sequence that enable a polynucleotide (e.g., an oligonucleotide) employed by a target sequencing platform to specifically bind to the nucleic acid domain, e.g., for solid-phase amplification and / or sequencing. Exemplary nucleic acid domains include P5 (5'-AATGATACGGCGACCACCGA-3') (SEQ ID NO:1), P7 (5'-CAAGCAGAAGACGGCATACGAGAT-3') (SEQ ID NO:2), Read 1 primer (5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3') (SEQ ID NO:3), and Read 2 primer (5'-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT-3') (SEQ ID NO:4) domains used on Illumina®-based sequencing platforms. Other exemplary nucleic acid domains include the A adapter (5'-CCATCTCATCCCTGCGTGTCTCCGACTCAG-3') (SEQ ID NO:5) and P1 adapter (5'-CCTCTCTATGGGCAGTCGGTGAT-3') (SEQ ID NO:6) domains used on Ion Torrent™-based sequencing platforms. The nucleotide sequence of the nucleic acid domain for sequencing on a target sequencing platform can vary and / or change over time. Adapter sequences are typically provided by the manufacturer of the sequencing platform (e.g., in technical documentation provided with the sequencing system and / or available on the manufacturer's website). Based on such information, any sequencing platform adapter domain, amplification primer, etc., similar sequences can be designed to include all or part of one or more nucleic acid domains, the configuration of which can sequence nucleic acid inserts (corresponding to assembled nucleic acid products) on the target platform.
[0079] Any suitable type of adapter can be ligated to the assembled nucleic acid, including but not limited to linear adapters, Y-shaped adapters, blunt-ended adapters, and hairpin adapters. In some embodiments, two different linear adapters are ligated to the 5' and 3' ends of each member of the library. The disadvantage of using linear double-stranded adapters is that if two different adapters (referred to as A and B) are ligated to double-stranded DNA in the library, a mixture of ligation products is produced, which includes approximately 50% of mis-ligated products with the same adapter at both ends (A-insert-A, B-insert-B), and only 50% have correctly ligated adapters (e.g., A-insert-B or B-insert-A). DNA inserts with A-A and B-B adapters cannot be amplified by PCR.
[0080] The Y - shaped adaptor has a Y - shaped structure, which has a single - stranded 5' arm, a single - stranded 3' arm, and a double - stranded stem. The arms of the Y have non - complementary sequences, and the stem is a double - stranded with DNA complementary strands hybridizing to each other. The Y - shaped adaptor is ligated to the assembled nucleic acid by ligating the double - stranded stem of the Y - shaped adaptor to the 5' end and 3' end of the assembled nucleic acid. The use of the Y - shaped adaptor allows different adaptor sequences to be added simultaneously to the 5' end and 3' end of a DNA library. The advantage of using a Y - shaped adaptor instead of a linear adaptor is that the Y - shaped adaptor provides an efficient way to ligate two different adaptor sequences to the 5' end and 3' end of each assembled nucleic acid in a library with a single adaptor. After ligating the Y - shaped adaptor to the two DNA ends, primers with sequences complementary to the sequences of the 5' arm and 3' arm of the Y - shaped adaptor can be used to amplify the library. The use of the Y - shaped adaptor also allows for the sequencing of both DNA strands.
[0081] Exemplary Y - shaped adaptors compatible with the Ion Torrent sequencing platform include: a 5' arm containing the nucleotide sequence CCATCTCATCCCTGCGTGTCTCCGACTCAG (SEQ ID NO:7), which includes: a primer - binding site containing the nucleotide sequence CCATCTCATCCCTGCGTGTCTCCGAC (SEQ ID NO:8), a sequence key (TCAG) for identifying library molecules and signal normalization, and a barcode sequence with a barcode termination signal (GAT); a 3' arm containing the nucleotide sequence GGAGAGATACCCGTCAGCCACTATCTGA (SEQ ID NO:9), which includes: a primer - binding site containing the nucleotide sequence GGAGAGATACCCGTCAGCCACTA (SEQ ID NO:10); and a stem, which includes: a first DNA strand containing the nucleotide sequence AGCACGAATCGAT (SEQ ID NO:11) and a second DNA strand containing the complementary nucleotide sequence TCGTGCTTAGCTSEQ ID (NO:12), as described in Forth et al. (BioTechniques (2019) 67:229 - 237); which is incorporated herein by reference. Exemplary Y - shaped adaptors compatible with the Illumina sequencing platform include: a first strand containing the nucleotide sequence AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO:13) and a second strand containing the nucleotide sequence GATCGGAAGAGCACACGTCTGAACTCCAGTCAC NNNNNNThe second strand of ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO:14), said second strand including a 6 nucleotide index sequence (NNNNNN), as underlined, which varies in the sequence. The xGen Stubby adaptor from Integrated DNA Technologies, Inc. (Coralville, Iowa) is a short Y-shaped adaptor that can ligate to fragments with A overhangs and is used with a proprietary unique dual-index (UDI) primer pair.
[0082] In some embodiments, hairpin adaptors are used. Hairpin adaptors generally include a double-stranded "stem" region and a single-stranded "loop" region. In some embodiments, a hairpin adaptor includes a single strand (i.e., a continuous strand) capable of adopting a hairpin structure, wherein the hairpin adaptor includes a self-complementary palindromic region forming the stem and a non-complementary region forming the loop of the hairpin adaptor. Additionally, hairpin adaptors can include various components of the adaptors described herein, including but not limited to amplification primer sites, barcode sequences, and sequences specifying the adaptor domain structure of a sequencing platform (e.g., P5 and P7 or A and P1 adaptor sequences). A hairpin adaptor can further contain one or more cleavage sites that are capable of being cleaved under cleavage conditions. In some embodiments, the cleavage site is located in the loop region. Cleavage at the cleavage site generates two separate strands from the hairpin adaptor. In some embodiments, cleavage at the cleavage site in the loop region generates a partially double-stranded adaptor having a double-stranded stem forming a "Y" structure and two unpaired strands. Cleavage sites can include, for example, uracil and / or deoxyuridine bases, which can be cleaved, for example, using DNA glycosylases, endonucleases, ribonucleases, etc. and combinations thereof. In some embodiments, the loop of the hairpin adaptor contains uracil residues and can be cleaved using uracil DNA glycosylase and endonuclease VIII. The advantage of using hairpin adaptors is that such adaptors allow for the continuous sequencing of both strands of a double-stranded DNA molecule by covalently linking one strand to the other. Hairpin adaptors also help minimize the formation of adaptor dimers during adaptor ligation.
[0083] Hairpin adapters can be purchased commercially from New England Biolabs (Ipswich, MA), such as NEBNext® adapters, which have sequences compatible with the Illumina sequencing platform. The Oxford Nanopore MinION sequencing platform uses Y-shaped adapters and hairpin adapters. The Y-shaped adapter is ligated to one end of the double-stranded DNA molecule, which provides the connection of the DNA molecule to the sequencing nanopore. Sequentially ligating a hairpin-like adapter to the other end of the double-stranded DNA molecule allows for the sequencing of both strands of the DNA molecule. Exemplary Oxford Nanopore Y-shaped adapters comprise a first strand containing the nucleotide sequence GGTTGTTTCTGTTGGTGCTGATATTGCGGCGTCTGCTTGGGTGTTTAACCT (SEQ ID NO:15) and a second strand containing the nucleotide sequence GGTTAAACACCCAAGCAGACGCCGCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA (SEQ ID NO:16). Exemplary Oxford Nanopore hairpin adapters comprise the nucleotide sequence CGTTCTGTTTATGTTTCTTGGACACTGATTGACACGGTTTAGTAGAAC (SEQ ID NO:17)-4(C3 spacer)-28(T)-CAAGAAACATAAACAGAACGT (SEQ ID NO:18), as described by Karamitros et al. (2015) Nucleic Acids Res. 43(22): e152; which is incorporated herein by reference.
[0084] As described above, bead-specific index sequences and reverse primer sites are added to the 3’ end of the assembled nucleic acids bound to each bead using split-ligation enzymatic EDS, where the bead-specific index sequence for each bead comprises a different unique identifier sequence. The indexed assembled nucleic acids bound to the beads can be amplified using a forward primer complementary to the adapter sequence and a reverse primer complementary to the reverse primer site. For example, if a Y-shaped adapter is used, the indexed assembled nucleic acids can be amplified using a forward primer complementary to one of the Y-shaped adapter arm sequences and a reverse primer complementary to the reverse primer site added by EDS.
[0085] Nucleic Acid Amplification
[0086] Any primer-dependent amplification method can be used for nucleic acid amplification, including but not limited to polymerase chain reaction (PCR), rolling circle amplification, and isothermal amplification methods such as recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), and nicking enzyme amplification reaction (NEAR), etc.
[0087] In some cases, amplification can be carried out by polymerase chain reaction (PCR). PCR primers should be of sufficient length to hybridize with the complementary template DNA under annealing conditions. The primer length is usually at least 6 bp, including but not limited to, for example, at least 10 bp, at least 15 bp, at least 16 bp, at least 17 bp, at least 18 bp, at least 19 bp, at least 20 bp, at least 21 bp, at least 22 bp, at least 23 bp, at least 24 bp, at least 25 bp, at least 26 bp, at least 27 bp, at least 28 bp, at least 29 bp, at least 30 bp, and the length can be as long as 60 bp or longer, where the primer length range is usually 18 bp to 50 bp, including but not limited to, for example, about 20 bp to 35 bp. In some cases, the template DNA can be contacted with a single primer or a set of two primers (forward and reverse primers), depending on whether primer extension, linear or exponential amplification of the template DNA is required. PCR methods that can be used in the methods of the present invention include but are not limited to U.S. Patent Nos. 4,683,202; 4,683,195; 4,800,159; 4,965,188, and 5,512,462, the disclosures of which are incorporated herein by reference.
[0088] In addition to the above components, the PCR reaction mixture can include a polymerase and deoxynucleoside triphosphates (dNTPs). The required polymerase activity can be provided by one or more different polymerases. In many embodiments, the reaction mixture includes at least a family A polymerase, and representative family A polymerases for the purpose include but are not limited to: Thermus aquaticus ( Thermus aquaticus ) polymerase, including the naturally occurring polymerase (Taq) and its derivatives and homologs, such as Klentaq (as described in Proc. Natl. Acad. Sci USA (1994) 91:2216-2220); Thermus thermophilus ( Thermus thermophilus ) polymerase, including the naturally occurring polymerase (Tth) and its derivatives and homologs Pol Θ (polymerase ), such as those described in, for example, WO 2019 / 030149 and WO 2022 / 086866. In some embodiments where the amplification reaction is a high-fidelity reaction, the reaction mixture can contain a polymerase with 3'-to-5' exonuclease activity, for example, which can be provided by a family B polymerase, where the family B polymerases of interest include but are not limited to: Thermococcus litoralis Pyrococcus woesei DNA polymerase (Vent) (e.g., as described in Perler et al., Proc. Natl. Acad. Sci. USA (1992) 89:5577); Pyrococcus Pyrococcus sp. strain GB-D (Deep Vent); Pyrococcus furiosus Pyrococcus furiosus DNA polymerase (Pfu) (e.g., as described in Lundberg et al., Gene (1991) 108: 1-6); Pyrococcus woesei Pyrococcus woesei, Pwo), etc. Typically, the reaction mixture includes four different types of dNTPs corresponding to the four naturally occurring bases, namely dATP, dTTP, dCTP, and dGTP, and in some cases, can include one or more modified nucleotide dNTPs.
[0089] Moreover, polymerases that preferentially use dUTP instead of dTTP can be used for PCR. Such polymerases include archaeal family B DNA polymerases, such as Nanoarchaeum equitans Nanoarchaeum equitans type B DNA polymerase, which can use deaminated bases (such as uracil and hypoxanthine) and perform PCR with higher fidelity than Thermus aquaticus (Taq) DNA polymerase (e.g., as described in Choi et al. (2008) Appl. Environ. Microbiol. 74(21): 6563-6569, which is incorporated herein by reference). In addition, engineered polymerases, such as Q5U Hot Start High-Fidelity DNA polymerase from New England Biolabs (Ipswich, MA) and Phusion U DNA polymerase from Thermo Fisher Scientific (Waltham, MA), contain mutations in the nucleotide-binding pocket such that these polymerases can amplify templates containing uracil and inosine bases and can be used for PCR using dUTP. Polymerases that use UTP can be used to prevent residual contamination in different PCR runs. The uracil-containing amplicon products of such polymerases can be digested by uracil-DNA glycosylase to remove residual products from previous PCR amplifications and inhibit template contamination between runs. Such methods are described, for example, in PCT Publication No. WO 92 / 01814.
[0090] In addition, one or more PCR additives or enhancers may be included to increase the yield of the amplification reaction, for example, by reducing secondary structure or mispriming events in the nucleic acid. Such additives or enhancers include, but are not limited to, dimethyl sulfoxide (DMSO), N,N,N-trimethylglycine (betaine), formamide, glycerol, nonionic detergents (such as Triton X-100, Tween 20, and Nonidet P-40 (NP-40)), 7-deaza-2'-deoxyguanosine, bovine serum albumin, T4 gene 32 protein, polyethylene glycol, 1,2-propanediol, and tetramethylammonium chloride.
[0091] The PCR reaction is typically carried out by cycling the reaction mixture between appropriate temperatures for annealing, extension / elongation, and denaturation for a specific time. Such temperatures and times will vary depending on the specific components of the reaction (including, for example, the polymerase and primers) and the expected length of the resulting PCR product. In some cases, for example, in the case of nested PCR or two-step PCR, the cycling reaction can be carried out in stages, for example, by cycling according to a first stage with a specific cycling program or using a specific temperature, and subsequently cycling according to a second stage with a specific cycling program or using a specific temperature.
[0092] The multi-step PCR method may or may not include adding one or more reagents after the start of the amplification. For example, in some cases, the amplification can be initiated by using the extension of a polymerase, and after the initial stage of the reaction, additional reagents (such as one or more additional primers, additional enzymes, etc.) can be added to the reaction to facilitate the second stage of the reaction. In some cases, the amplification can be initiated with a first primer or a first set of primers, and after the initial stage of the reaction, additional reagents (such as one or more additional primers, additional enzymes, etc.) can be added to the reaction to facilitate the second stage of the reaction. In some embodiments, the initial stage of the amplification may be referred to as "pre-amplification".
[0093] In particular, the method of the present invention can be applied to digital PCR technology. For digital PCR, a sample containing nucleic acid is separated into a large number of partitions before performing PCR. Partitioning can be achieved in a variety of ways, such as by using a microplate, capillary, emulsion, microchamber array, or nucleic acid binding surface. The separation of the sample can involve dispensing any suitable portion, including up to the entire sample, into the partitions. Each partition includes a fluid volume that is isolated from the fluid volumes of other partitions. The partitions can be isolated from each other by a fluid phase (such as the continuous phase of an emulsion), by a solid phase (such as at least one wall of a container), or a combination thereof. In some embodiments, the partitions can include droplets arranged in a continuous phase such that the droplets and the continuous phase together form an emulsion.
[0094] Compartments can be formed by any suitable procedure, in any suitable manner, and with any suitable characteristics. For example, compartments can be formed by a fluid dispenser (such as a pipette), a droplet generator, by agitating the sample (such as shaking, stirring, sonication, etc.), and so on. Thus, compartments can be formed continuously, in parallel, or in batches. Compartments can have any suitable one or more volumes. Compartments can have a substantially uniform volume or can have different volumes. Exemplary compartments having substantially the same volume are monodisperse droplets. Exemplary volumes of compartments include an average volume of less than about 100 μL, 10 μL, or 1 μL, less than about 100 nL, 10 nL, or 1 nL, or less than about 100 pL, 10 pL, or 1 pL, etc.
[0095] After separating the sample, PCR is performed in the compartments. When forming the compartments, one or more reactions can be carried out in the compartments. Alternatively, one or more reagents can be added after the compartments are formed to enable the reactions to occur. The reagents can be added by any suitable means, such as by a fluid dispenser, droplet fusion, etc.
[0096] In some embodiments, nucleic acids are amplified by emulsion PCR to compartmentalize the amplification reaction of individual DNA molecules. An aqueous PCR mixture having forward and reverse primers is mixed with oil to produce an emulsion. Preferably, each droplet in the water-in-oil emulsion contains a bead and a template DNA molecule (e.g., a single assembled nucleic acid of a sequencing library), such that individual molecules are amplified in separate emulsion droplets. After amplification, the emulsion is disrupted, for example, by vortexing with isopropanol and a detergent. In some embodiments, the gene fragment library and the sequencing library are bound to magnetic beads or superparamagnetic beads before amplification, and magnetic separation of the beads is performed after the amplification and disruption of the emulsion. For a description of emulsion PCR, see, for example, Kanagal-Shamanna et al. (2016) Methods Mol Biol. 1392:33-42, Zhu et al. (2012) Anal Bioanal Chem. 403(8):2127-43, Zhang et al. (2020) Lab Chip 20(13):2328-2333, Siu et al. (2021) Talanta 221:121593, Zheng et al. (2011) Nat. Protoc. 6(9):1367-1376, and Kojima et al. (2015) Methods Mol. Biol. 2015;1347:87-100; the above are incorporated herein by reference.
[0097] After PCR amplification, nucleic acids can be quantified by counting partitions containing PCR amplicons. Partitioning of the sample allows for quantification of the number of different molecules by assuming that the molecular population follows a Poisson distribution. For a description of digital PCR methods, see, e.g., Hindson et al. (2011) Anal. Chem. 83(22):8604-8610; Pohl and Shih (2004) Expert Rev. Mol. Diagn. 4(1):41-47; Pekin et al. (2011) Lab Chip 11 (13): 2156-2166; Pinheiro et al. (2012) Anal. Chem. 84 (2): 1003-1011; Day et al. (2013) Methods 59(1):101-107; the foregoing is incorporated herein by reference.
[0098] In some cases, amplification can be carried out under isothermal conditions, e.g., by means of isothermal amplification. Isothermal amplification methods generally utilize enzymatic means for separating DNA strands to facilitate amplification at a constant temperature, such as strand-displacing polymerases or helicases, thereby obviating the need for thermal cycling to denature the DNA. Any convenient and suitable isothermal amplification method can be employed in the methods, including but not limited to: recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), and nicking enzyme amplification reaction (NEAR), etc.
[0099] LAMP generally uses multiple primers, e.g., 4 to 6 primers, which can recognize multiple different regions of the target DNA, e.g., 6 to 8 different regions. Synthesis is typically initiated by a strand-displacing DNA polymerase, where two primers form a loop structure to facilitate subsequent amplification cycles. LAMP is rapid and sensitive. In addition, magnesium pyrophosphate generated during the LAMP amplification reaction can in some cases be visualized without the use of specialized equipment, e.g., by eye.
[0100] RPA combines isothermal recombinase-mediated primer targeting with strand-displacement DNA synthesis (Piepenburg et al. (2006) PLOS Biology. 4 (7): e204; incorporated herein by reference). This technique uses two primers along with recombinase, single-stranded DNA-binding protein, and strand-displacement polymerase for amplification. Unlike PCR, denaturation of DNA strands does not require heating. Instead, the recombinase-primer complex is used for local strand exchange to place the oligonucleotide primer on the homologous sequence of the DNA template. The single-stranded DNA-binding protein binds to the displaced template strand to prevent the primer from being displaced by branch migration. Dissociation of the recombinase makes the 3’ end of the primer accessible for binding by the strand-displacement DNA polymerase (e.g., the large fragment of Bacillus subtilis Pol I), which catalyzes primer extension. Repetition of the process cycles results in exponential amplification.
[0101] SDA generally involves initiation at a nick generated by a strand-restricted restriction endonuclease or nicking enzyme at a primer-containing site using a strand-displacement DNA polymerase (e.g., Bst DNA polymerase, large (Klenow) fragment polymerase, Klenow fragment (3’-5’ exonuclease-deficient), etc.). In SDA, the nick site is typically regenerated with each polymerase displacement step, resulting in exponential amplification.
[0102] HDA generally employs: a helicase that unwinds double-stranded DNA to separate it into strands; primers, such as two primers, that can anneal to the unwound DNA; and a strand-displacement DNA polymerase for extension.
[0103] NEAR generally includes a strand-displacement DNA polymerase that initiates extension at a nick (generated by a nicking enzyme). NEAR is rapid and sensitive, rapidly generating many short nucleic acids from the target sequence.
[0104] In some cases, entire amplification methods can be combined, or aspects of various amplification methods can be recombined to produce a combined amplification method. For example, in some cases, aspects of PCR can be used, such as to generate an initial template or amplicon or the first round or multiple rounds of amplification, and subsequently an isothermal amplification method can be employed for further amplification. In some cases, an isothermal amplification method or aspects of an isothermal amplification method can be employed, followed by PCR to further amplify the product of the isothermal amplification reaction. In some cases, a first amplification method can be used to pre-amplify a sample, and a second amplification method can be used to further process the sample, including, for example, further amplification or analysis. As a non-limiting example, a sample can be pre-amplified by PCR and further analyzed by qPCR.
[0105] In some cases, the method further includes monitoring the amplification of the target DNA molecule, such as in real-time PCR (also referred to herein as quantitative PCR (qPCR)).
[0106] DNA synthesis
[0107] Polynucleotides, adapters, and primers can be synthesized by any suitable technique, such as by solid-phase synthesis using phosphoramidite chemistry, as disclosed in U.S. Patent Nos. 4,458,066 and 4,415,732, which are incorporated herein by reference; and Adams et al., J. Amer. Chem. Soc. (1983) 105 :661-663, Froehler et al., Tetrahedron Lett. (1983) 24 :3171-3174, Beaucage et al., Tetrahedron (1992) 48 :2223-2311; and as disclosed in Applied Biosystems User Bulletin No. 13 (1 April 1987). Other chemical synthesis methods include, for example, the phosphotriester method described by Narang et al., Meth. Enzymol. (1979) 68 :90 and the phosphodiester method described by Brown et al., Meth. Enzymol. (1979) 68 :109. Poly(A) or poly(C), or other non-complementary nucleotide extensions can be incorporated into the polynucleotide using these same methods. Hexaethylene glycol ether extensions can be ligated to the polynucleotide by methods known in the art. Cload et al., J. Am. Chem. Soc. (1991) 113 :6324-6326; U.S. Patent No. 4,914,210 to Levenson et al.; Durand etal., Nucleic Acids Res. (1990) 18 :6353-6359; and Horn et al., Tet. Lett. (1986) 27 :4705-4708. Polynucleotides, adapters, and primers can be synthesized in the laboratory, for example, using an automated synthesizer or by enzymatic DNA synthesis (EDS), as further described below.
[0108] In some embodiments, polynucleotides are synthesized that have only nucleotides that occur naturally in DNA, such as adenine, thymine, guanine, and cytosine. In other embodiments, unnatural bases, nucleotide analogs, or unnatural base pairs (UBPs) are incorporated into the polynucleotide during synthesis. Unnatural nucleobases can have base pairing and base stacking properties different from those of natural nucleobases. Artificial nucleic acids include peptide nucleic acid (PNA), morpholino nucleic acid, locked nucleic acid (LNA), glycol nucleic acid (GNA), threose nucleic acid (TNA), and hexitol nucleic acid (HNA). Modifications can include, but are not limited to, N3-methylation modification of cytosine, O6-methylation modification of guanine, and N-acetylation modification of guanine. Exemplary nucleotide analogs include, but are not limited to, dideoxynucleotides, deazanucleotides, aminoallyl nucleotides, thiol-containing nucleotides, biotin-containing nucleotides, furan-modified bases, and fluorescent base analogs (e.g., 2-aminopurine, 1,3-diaza-2-oxophenoxazine, 3-MI, 6-MI, 6-MAP, and pyrrolo-dC). Thionophosphorothioate oligonucleotides (OPS) can also be synthesized that have an oxygen atom replaced by sulfur in the phosphate moiety. Examples of UBPs include d5SICS and dNaM, which have hydrophobic nucleobases and two fused aromatic rings, forming the (d5SICS–dNaM) complex or base pair in DNA. Various modified nucleotides that can be used for the enzymatic synthesis of polynucleotides in the absence of a template by terminal deoxynucleotidyl transferase or its variant enzymes are described in U.S. Patent Nos. 11,059,849 and 5,763,594; which are incorporated herein by reference in their entirety. In some cases, modification of the nucleotide can block polymerization of the nucleotide and / or allow the nucleotide to interact with another molecule, such as a protein. In some cases, the polynucleotide can be chemically modified after DNA synthesis.
[0109] Enzymatic DNA synthesis
[0110] Enzymatic DNA synthesis (EDS) uses an enzyme catalyst for the polymerization of nucleotides. For example, for the purpose of adding an index sequence, barcode, primer sequence, restriction site, or any other desired sequence to a polynucleotide or assembling nucleic acids, EDS can be used to generate primers and other polynucleotides, as well as to add nucleotides to the 3’ end of a polynucleotide. In some embodiments, EDS is performed with an enzyme that catalyzes the addition of a nucleotide-5’-triphosphate to the 3’ end of a DNA molecule. More specifically, the method uses an enzyme that generates a phosphodiester bond between (i) a free 3’-OH group of a nucleic acid, such as a terminal 3’-OH group, and (ii) the 5’-phosphate group of the nucleotide to be added by the enzyme.
[0111] The polynucleotide to which nucleotides are added by EDS can be single-stranded or double-stranded, or can include single-stranded and double-stranded portions. When the polynucleotide chain contains a double-stranded portion, the polynucleotide chain can contain a 3'-overhang, a recessed 3'-end (such that the complementary strand has a 5'-overhang), or a blunt-ended 3'-end (i.e., no 3' or 5' overhang), and nucleotides are added to the polynucleotide by EDS.
[0112] In some embodiments, EDS is performed with an enzyme capable of catalyzing nucleotide polymerization without using a template strand. In some embodiments, the enzyme for enzymatic DNA synthesis is a member of the polX family polymerase, such as DNA polymerase β (Pol β), λ (Pol λ), μ (Pol μ), yeast IV (Pol IV), and terminal deoxynucleotidyl transferase (TdT). In some embodiments, the enzyme is an η or ζ type translesion DNA polymerase, polynucleotide phosphorylase (PNPase), template-independent RNA polymerase, ligase, template-independent DNA polymerase, reverse transcriptase, 9°N DNA polymerase, or terminal deoxynucleotidyl transferase (TdT). Many enzymes available for EDS are expressed by cells of living organisms and can be extracted from these cells or purified from recombinant cultures. These enzymes can have a naturally occurring amino acid sequence (natural for their source organism) or can contain one or more modifications relative to their native sequence, such as amino acid insertions, deletions, truncations, substitutions, and / or other modifications.
[0113] In some embodiments, engineered terminal deoxynucleotidyl transferase is used to perform EDS. For this purpose, various variants of terminal deoxynucleotidyl transferase have been developed. See, e.g., PCT Publication No. 2022 / 0002687; incorporated herein by reference in its entirety. In some embodiments, engineered reverse transcriptase is used to perform EDS. For example, human immunodeficiency virus type 1 and Moloney murine leukemia virus reverse transcriptases can be used. Engineered Moloney murine leukemia virus reverse transcriptase variants are commercially available, such as SuperScript IV reverse transcriptase from Thermo Fisher (Waltham, MA) and SMARTScribe reverse transcriptase from Clonetech (Mountain View, Calif.).
[0114] In some embodiments, engineered 9°N DNA polymerase is used to perform EDS. Engineered 9°N DNA polymerase variants are commercially available, including those from Centrillion Technology Holdings Corporation (Grand Cayman, KY) and TherminatorThermococcus sp duplases. DNA polymerases from New England Biolabs (Ipswich, MA) (Ipswich, MA). See also, e.g., Hoff et al. (2020) ACSSynth Biol 9(2):283-293; Gardner et al. (2019) Front. Mol. Biosci. 6:28; the foregoing are incorporated herein by reference in their entirety.
[0115] The polymerization reaction can be carried out with natural nucleotides, modified nucleotides, or a combination thereof. The nucleotides employed generally consist of a cyclic sugar moiety (which contains at least one chemical group at the 5'-end and one chemical group at the 3'-end) and a natural or modified nitrogenous base. The conditions of the reaction medium are selected, particularly temperature, pressure, pH, optional buffer medium, and other reagents, so that the polymerase functions optimally possible while ensuring the integrity of the molecular structures of the different reagents present.
[0116] In some cases, the free ends of one or more nucleotides added to a nucleic acid advantageously contain protective chemical groups to prevent multiple additions of nucleotides to the same nucleic acid. Nucleotides protected by the addition of a protecting group at their 3'-end prevent subsequent polymerization and thus limit the risk of uncontrolled polymerization. In addition to preventing the formation of phosphodiester bonds, the protecting group can have other functions, such as allowing interaction with a solid support (e.g., a bead) or with other reagents in the reaction medium. For example, the solid support can be covered with molecules (such as proteins) such that non-covalent bonds between these molecules and the protective chemical group are possible to allow immobilization on the solid support. Only nucleotides with a protecting group can interact with the solid support. Thus, nucleic acids without the addition of a protecting group can be removed based on the lack of interaction between their 3'-end and the solid support. If a protecting group is used, deprotection is necessary for satisfactory progress of the synthesis. Deprotection can be effected by chemical reactions, electromagnetic interactions, enzymatic reactions, and / or chemical or protein interactions.
[0117] EDS can be used to add nucleotides to multiple nucleic acids simultaneously. The EDS reaction can be carried out in parallel for simultaneous synthesis of the desired sequences on a variety of different nucleic acids. Split-and-pool EDS can be used to add index sequences to assembled nucleic acids while immobilizing the assembled nucleic acids on beads to provide a "bead index" for all the molecular clones amplified on the beads.
[0118] In some embodiments, the nucleic acid is immobilized on a solid support, which can be washed by an automated liquid handling device and exposed to enzymes and buffers. For example, the solid support can be a microtiter plate, and reagents are dispensed into and removed from the microtiter plate by a liquid handling robot. Alternatively, the solid support can be magnetic beads, which can be magnetically separated from a suspension and then resuspended in a new reagent solution in a multi-well plate. Another possible solid support can be the inner surface of a microfluidic device, which dispenses reagents to a location in an automated manner. Immobilizing the nucleic acid on a solid support prevents loss of the nucleic acid during washing.
[0119] The SYNTAX EDS system is commercially available from DNAScript (South San Francisco, CA). Currently available systems use engineered terminal deoxynucleotidyl transferase to provide automated parallel EDS in 96-well plates. For further description of EDS, see, for example, U.S. Patent Nos. 10,913,964; 11,268,091; 11,059,849 and U.S. Patent Application Publication No. 2021 / 0254114; the foregoing are incorporated herein by reference in their entirety.
[0120] By split-and-pool synthesis indexing
[0121] In some embodiments of the present invention, bead-specific index sequences (also referred to as "barcode" sequences) are added to the bead-bound, clonally amplified assembled nucleic acids using a "split and pool" method. The bead-bound assembled nucleic acids are divided into multiple compartments to physically separate individual beads carrying the clonally amplified assembled nucleic acids such that each bead is separate from every other bead. When the bead-bound assembled nucleic acids are partitioned (where each compartment receives a different nucleotide to add to the 3' end of the nucleic acid strand on the bead), nucleotides are added to the bead-bound assembled nucleic acids for indexing, the bead-bound assembled nucleic acids are combined, and then the process is repeated as many times as needed. If the partitioning is random, the assembled nucleic acids attached to different beads should receive different sets of nucleotides and thus different bead-specific index sequences at the end of the process. All copies of the assembled nucleic acids clonally amplified on the same bead should acquire the same bead-specific index.
[0122] Because the total number of index combinations grows exponentially with the number of split and merge cycles, split and merge methods can be used to index assembled nucleic acids bound to beads in almost unlimited numbers. The longer the gene to be synthesized, the more relevant this becomes. For example, finding a perfect copy of a 100 bp gene requires a bar code (i.e., shorter in length) of lower complexity than finding a 10,000 bp gene. One of ordinary skill in the art can calculate the probability of recovering a perfect copy based on the DNA synthesis and gene synthesis methods employed, and then adjust the number of split and merge cycles in a manner that ensures the use of a bar code of sufficient length to ensure a successful experimental outcome. In all cases, there is no risk in using a bar code longer than needed. However, as the length of the synthesized bar code decreases, the risk of independently and randomly synthesizing the same bar code on two different beads increases (the random probability of beads with the same bar code is higher for short lengths and drops to zero as the length increases).
[0123] Multi-compartment plates can be used to split samples containing multiple bead-bound assembled nucleic acids. In some embodiments, a device is used to automate the split and merge process, where the device is capable of toggling back and forth between the following states: (i) a merge state, where the split samples between multiple compartments are merged; and (ii) a split state, where the samples merged in (i) are split into multiple compartments. The device can be used to repeatedly split and merge samples containing multiple bead-bound assembled nucleic acids. The device is capable of easily toggling back and forth between the split state, where the samples merged in the split state are split into multiple compartments, and the merge state, where the samples split between multiple compartments are merged. In the split state, the previously merged samples are split into at least 2, at least 4, at least 8, at least 16, at least 48, or at least 96 compartments. In the merge state, the previously separated samples are merged into a smaller number of compartments. For example, in the merge state, samples that have been divided into n compartments (where n = at least 2, at least 4, at least 8, at least 16, at least 48, or at least 96, etc.) can be merged into a smaller number of compartments (e.g., n / 2, n / 4, or n / 8 compartments). In some cases, in the split state, the previously merged samples can be split into at least 8 compartments, and in the merge state, the samples (which were previously split into at least 8 compartments) can be merged into a single compartment. The device can easily toggle between the two states, allowing the sample to be split and merged many times as needed to provide a unique identifier sequence for each indexed assembled nucleic acid. In some embodiments, the compartments can have a volume in the range of 5 μl to 1 ml, such as 10 μl to 500 μl, or 20 μl to 200 μl.
[0124] Nucleic acids can be added to the assembled nucleic acids bound to beads by enzymatic catalysis (e.g., by EDS). In certain embodiments, an EDS device can be used to index samples using a split-and-pool method, the overall goal being to add bead-specific indices containing unique identifier sequences to each assembled nucleic acid in the sample.
[0125] In some embodiments, split-pool EDS is used to add nucleotides one at a time to the 3' end of partitioned bead-bound assembled nucleic acids. The bead-bound assembled nucleic acids are partitioned into multiple compartments, different nucleotides for bead indexing are added to each partition, and then the bead-bound assembled nucleic acids are combined. Subsequently, the partitioning, nucleotide addition, and combining steps are repeated as needed until a sufficient number of nucleotides have been added to provide unique bead indices to the clonally amplified assembled nucleic acids. Typically, no more than 12 to 25 cycles of split-pool EDS are required to provide unique bead indices to all beads bound to clonally amplified assembled nucleic acids generated from the assembly of a gene fragment library.
[0126] Bead index sequences can be added to the assembled nucleic acids as needed, thereby allowing identification of the sequencing data for each assembled nucleic acid and selective amplification of the indexed assembled nucleic acids with primers having sequences complementary to the index sequences. For example, primers complementary to the index sequences that uniquely identify the beads to which the assembled nucleic acids are bound can be used to selectively amplify the indexed assembled nucleic acids containing perfect copies of the sequences of interest from the combined mixture.
[0127] The sample can be input into the split-pool device when the sample is in a split state, a combined state, or in another state (e.g., a "loaded" state). Next, the method includes switching the device to the split state. In this step, the sample is partitioned into multiple compartments. After splitting, different nucleotides are added to the bead-bound assembled nucleic acids while splitting the sample (usually one nucleotide per compartment, where different compartments receive different nucleotides). Next, the method includes switching the device to the combined state, thereby combining the separated samples. Then the partitioning, addition, and combining steps can be repeated as needed until the bead-bound assembled nucleic acids are indexed. In these embodiments, the method can include repeating the partitioning and addition steps multiple times, with a combining step after each repetition, except that the combining step is optional in the last step. After repeating these steps, the indexed sample can be collected. Obviously, appropriate mixing and / or washing steps can be performed between any steps of the method or during any step of the method.
[0128] Sequencing
[0129] Any high-throughput technique for sequencing can be used to implement the present invention. For example, DNA sequencing techniques include dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and slab or capillary gel electrophoresis, sequencing by synthesis using reversible terminator-labeled nucleotides, pyrosequencing, 454 sequencing, sequencing by synthesis (e.g., using a clonal library labeled with allele-specific hybridization, followed by ligation and real-time monitoring of the incorporation of labeled nucleotides during the polymerization step), polony sequencing, SOLID sequencing, etc.
[0130] Some high-throughput sequencing methods include the step of spatially isolating individual molecules on a solid surface, on which they are sequenced in parallel. Such solid surfaces can include, as in Solexa sequencing (e.g., as described in Bentley et al, Nature, 456: 53-59 (2008)), or Complete Genomics sequencing, e.g., (as described in Drmanac et al, Science, 327: 78-81 (2010)), a poreless surface, can include a bead-bound or particle-bound template pore array (such as 454 (e.g., as described in Margulies et al, Nature, 437: 376-380 (2005)) or Ion Torrent sequencing (e.g., as described in U.S. Patent Publication No. 2010 / 0137143 or 2010 / 0304982), a microfabricated thin film (such as SMRT sequencing by Pacific Biosciences (e.g., as described in Eid et al, Science, 323: 133-138 (2009))), or a bead array (such as SOLiD sequencing or Polony sequencing, e.g., as described in Kim et al, Science, 316: 1481-1414 (2007)). Such methods can include amplifying the nucleic acid before or after spatially isolating it on the solid surface. Amplification can include emulsion-based amplification, such as emulsion PCR, or rolling circle amplification.
[0131] In some embodiments, sequencing can be performed using the Illumina MiSeq, NextSeq, or HiSeq platforms, which use techniques for sequencing by reversible terminator synthesis (see, e.g., Shen et al. (2012) BMC Bioinformatics 13:160; Junemann et al. (2013) Nat. Biotechnol. 31(4):294-296; Glenn (2011) Mol. Ecol. Resour. 11(5):759-769; Thudi et al. (2012) Brief Funct. Genomics 11(1):3-11; the foregoing is incorporated herein by reference); the MinION, GridION, and PromethION nanopore sequencing platforms of Oxford Nanopore Technologies Inc, which can be used to determine the sequence of DNA or RNA by monitoring changes in current as nucleic acids pass through a protein nanopore (see, e.g., Lu et al. (2016) Genomics Proteomics Bioinformatics 14(5):265-279, Petersen et al. (2019) J. Clin. Microbiol. 58(1):e01315-19, Kono et al. (2019) Dev Growth Differ. 61(5):316-326, Deamer et al. (2016) Nat. Biotechnol. 34(5):518-24, Madoui et al. (2015) BMC Genomics 16:327, Szalay et al. (2015) Nat. Biotechnol 33, 1087-1091; the foregoing is incorporated herein by reference); the PacBIO single molecule, real-time (SMRT) sequencing platforms, including the Sqequel, HiFi, and RS II sequencing platforms (see, e.g., Ardui et al. (2018) Nucleic Acids Res. 46(5):2159-2168, An et al. (2018) Genes (Basel) 9(1):43), Nakano et al. (2017) Hum Cell. 30(3):149-161; the foregoing is incorporated herein by reference), the Omniome combined sequencing short read sequencing (SBB®) platform using a high-fidelity plasmonic nanopore array (see, e.g., Cetin et al.(2018) ACS Sens. 3(3):561-568; which is incorporated herein by reference), the Gynapsys compact DNA sequencer (which uses a complementary metal-oxide semiconductor (CMOS) sequencing chip for electronic data detection and sequencing by synthesis (SBS) chemistry), the Singular Genomics G4 benchtop sequencing platform (which uses SBS chemistry), and the Element Biosciences AVITI™ benchtop sequencer (which uses a modified form of SBS chemistry that reduces reagent consumption).
[0132] Thus, these sequencing methods can be used to sequence a sequencing library to identify the unique bead index sequences of the assembled nucleic acids that are clonally amplified to contain perfect copies of the sequence of interest. The short index sequences can also be used for multiplexed sequencing of pooled library samples. Thus, primers complementary to the unique bead index sequences of the assembled nucleic acids that are clonally amplified to contain perfect copies of the sequence of interest can subsequently be used to amplify the correctly assembled nucleic acids containing perfect copies of the sequence of interest from the pooled library samples.
[0133] In some embodiments, the methods of the invention further comprise performing deep sequencing of the amplified indexed assembled nucleic acids on one or more beads. Sequence errors may be introduced at various stages. For example, errors may occur during the clonal nucleic acid amplification of different amplification cycles. Subsequent errors affect the assembled nucleic acids present in fewer numbers on the beads, while errors that occur earlier during amplification affect the assembled nucleic acids present in greater numbers on the beads. Performing deep sequencing on the assembled nucleic acid products from the beads can resolve these situations and identify the perfect copies of the sequence of interest to be used as the desired template, as well as identify the amplified nucleic acids that contain an unacceptable proportion of non-perfect copies.
[0134] System
[0135] The present disclosure also provides a system that can be used to implement the method of the present invention. The system is used for cell-free identification of correctly assembled nucleic acids containing a target sequence. In some embodiments, the system may include: a device and a processor for performing split-and-pool EDS, wherein the processor is programmed to instruct the device to perform split-and-pool EDS to add bead-specific index sequences and add reverse primer sites to the free 3' ends of bead-bound clonal amplified assembled nucleic acids. In some embodiments, the system further includes a sequencer, wherein the sequencer is connected to the processor such that the processor receives sequencing data from the sequencer. In some embodiments, the sequencer is inside the device for performing split-and-pool EDS. In other embodiments, the sequencer is a separate device. In some embodiments, the system further includes a memory storage component operably connected to the processor, wherein the memory storage component is configured to store sequencing data. In some embodiments, the processor is further programmed to receive the sequence of a nucleic acid containing a target sequence; design the sequences of a plurality of polynucleotides, wherein each polynucleotide contains a portion of the target sequence; and instruct a synthesizer to synthesize the plurality of polynucleotides.
[0136] In some embodiments, the system further includes a chamber for assembling a plurality of polynucleotides. In some embodiments, the chamber further includes an exonuclease, a DNA polymerase, and a ligase for performing the assembly of the plurality of polynucleotides. In some embodiments, the chamber includes reagents for performing Gibson assembly, polymerase chain assembly (PCA), Golden Gate assembly, BioBrick assembly, or oligonucleotide ligation. In some embodiments, the processor is further programmed to instruct the device to perform the assembly of the plurality of polynucleotides.
[0137] In some embodiments, the system further includes a plurality of Y-shaped adapters, a set of polymerase chain reaction (PCR) primers having sequences complementary to the 5' arm and 3' arm of the Y-shaped adapter, a PCR primer having a sequence complementary to the reverse primer site; reagents, beads, or ligases for performing emulsion PCR, or any combination thereof. In some embodiments, the beads are non-magnetic beads or magnetic beads (e.g., superparamagnetic beads). In some embodiments, the system further includes reagents for performing Gibson assembly, polymerase chain assembly, Golden Gate assembly, BioBrick assembly, or oligonucleotide ligation.
[0138] In some embodiments, a computer-implemented method is used to assemble a nucleic acid comprising a target sequence. A processor can be programmed to perform the steps of the computer-implemented method, the method comprising: receiving a nucleic acid sequence to be assembled from a plurality of polynucleotides; designing the sequences of the plurality of polynucleotides, wherein each polynucleotide comprises a portion of the nucleic acid sequence containing the target sequence, and the plurality of polynucleotides can be assembled into the full-length sequence of the nucleic acid containing the target sequence; instructing a synthesizer to synthesize the plurality of polynucleotides; instructing a device to add a bead-specific index sequence and a reverse primer site to the free 3' end of the plurality of bead-bound clonal amplified assembled nucleic acids generated from the plurality of polynucleotides using split-and-pool EDS, wherein the bead-specific index sequence of each bead comprises a different unique identifier sequence; receiving the sequence of the indexed assembled nucleic acid; analyzing the sequence of the indexed assembled nucleic acid to identify the bead-specific index sequence of the correctly assembled nucleic acid containing the target sequence; and displaying the bead-specific index sequence of the correctly assembled nucleic acid containing the target sequence. In some embodiments, the computer-implemented method further comprises displaying the sequences of the plurality of polynucleotides. In some embodiments, the computer-implemented method further comprises displaying the sequence of the indexed assembled nucleic acid. In some embodiments, the computer-implemented method further comprises instructing a device to generate a sequencing library comprising a plurality of assembled nucleic acids generated by assembling the plurality of polynucleotides.
[0139] The methods described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware. The methods herein can be implemented as one or more computer program products (i.e., one or more modules of computer program instructions encoded on a computer-readable medium) to perform or control the operation of a data processing apparatus by a data processing device. The computer-readable medium can be a machine-readable memory storage device, a machine-readable memory storage substrate, a memory device, a substance composition implementing a machine-readable propagated signal, or any combination thereof.
[0140] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including a compiled or interpreted language) and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program need not correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communication network.
[0141] In a further aspect, as described, a system for performing a computer-implemented method can include a computer that includes a processor, a memory storage component, a display component, and other components typically present in a general-purpose computer. The memory storage component stores information accessible by the processor, including instructions executable by the processor and data retrievable, manipulable, or storable by the processor.
[0142] The memory storage component can include instructions. For example, the memory storage component includes instructions for analyzing sequencing data of index-assembled nucleic acids to identify correctly assembled nucleic acids that contain perfect copies of a target sequence; and identifying bead-specific index sequences of correctly assembled nucleic acids that contain the target sequence. The computer processor is connected to the memory storage component and is configured to execute the instructions stored in the memory storage component to receive a target sequence, assemble nucleic acids that contain the target sequence from multiple polynucleotides, and identify correctly assembled nucleic acids that contain the target sequence, as described herein. The display component can display information about the target sequence, polynucleotide fragment sequences that contain a portion of the target sequence, sequenced assembled nucleic acid sequences, and / or bead-specific index sequences of correctly assembled nucleic acids that contain the target sequence.
[0143] The memory storage component can be any type capable of storing information accessible by the processor, such as a hard disk drive, memory card, ROM, RAM, DVD, CD-ROM, USB flash drive, writable and read-only memory. The processor can be any well-known processor, such as a processor from Intel Corporation. Alternatively, the processor can be a dedicated controller, such as an ASIC.
[0144] The instructions can be any set of instructions directly executable (such as machine code) or indirectly executable (such as a script) by the processor. In this regard, the terms "instructions", "steps", and "program" can be used interchangeably herein. The instructions can be stored in object code form for direct processing by the processor or in any other computer language, including scripts or collections of independent source code modules that are interpreted or pre-compiled as needed.
[0145] Data can be retrieved, stored, or modified by the processor according to the instructions. For example, although the system is not limited to any particular data structure, the data can be stored in a computer register of a relational database as a table with multiple different fields and records, an XML document, or a flat file. The data can also be formatted in any computer-readable format, such as but not limited to binary values, ASCII, or Unicode. In addition, the data can include any information sufficient to identify relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories (including other network locations), or information for calculating relevant data by a function.
[0146] In some embodiments, the processor and memory storage components can include multiple processor and memory storage components, which may or may not be stored within the same physical housing. For example, some instructions and data can be stored in a removable CD-ROM and other CD-ROMs within a read-only computer chip. Some or all of the instructions and data can be stored at a location physically remote from the processor but still accessible by the processor. Similarly, the processor can include a collection of processors that may or may not operate in parallel.
[0147] Components of a system for performing the methods of the present disclosure are further described in the embodiments below.
[0148] Kit
[0149] Also provided are kits for practicing the methods described herein. The kits can include adapters, reagents for amplifying the assembled nucleic acids (including primers having sequences complementary to the primer binding sites of the adapters), reagents for amplifying the sequencing library of the assembled nucleic acids (including primers complementary to the primer binding sites of the adapters (e.g., one of the arm sequences of a Y-shaped adapter) and primers complementary to the index sequences added to the assembled nucleic acids by EDS), nucleotides, and polymerases (e.g., Taq DNA polymerase). The kits can further include reagents for performing emulsion PCR (emPCR), such as oil for generating the emulsion, and reagents for disrupting the emulsion (e.g., isopropanol, detergents). In some embodiments, the kits include a plurality of double-stranded oligonucleotides / polynucleotides, each oligonucleotide / polynucleotide containing a portion of the target sequence. Different types of kits can be provided according to the needs of the user. In particular, different oligonucleotides / polynucleotides can be provided in different kits according to the desired sequences to be synthesized. The kits can also include reagents for synthesizing fragment libraries of oligonucleotides / polynucleotides by EDS. The kits can further include reagents for adding bead index sequences containing unique identifier sequences and reverse primer sites to the assembled nucleic acids using EDS. The EDS reagents can include a reaction medium that includes an enzyme (e.g., terminal deoxynucleotidyl transferase) for adding nucleotides to the 3' end of the nucleic acids described herein and natural and / or unnatural nucleotides that can be used as substrates by the enzyme. In some embodiments, the kit further includes a device for performing split-and-pool EDS. The kits can further include reagents for nucleic acid assembly, such as ligases, restriction endonucleases, polymerases, vectors, and other reagents suitable for the selected assembly method, such as Gibson assembly, polymerase chain assembly, Goldengate assembly, BioBrick assembly, oligonucleotide ligation, or other assembly methods. The kits can further include buffers, wash solutions, solid supports (e.g., magnetic beads / superparamagnetic beads), etc. Different types of kits can be provided according to the functions of automatic or non-automatic use.
[0150] In addition to the above components, the subject kit can further include (in some embodiments) instructions for practicing the subject methods using the kit. In some embodiments, instructions for cell-free assembly of a nucleic acid comprising a target sequence and for identifying correctly assembled nucleic acids comprising the target sequence are provided in the kit. These instructions can be present in the subject kit in a variety of forms, and one or more of these instructions can be present in the kit. One form in which these instructions can be present is as information printed on a suitable medium or substrate, such as printed on paper, in the packaging of the kit, in a package insert, etc. Another form of these instructions is a computer-readable medium, e.g., a disk, a compact disc (CD), a DVD, a flash drive, an SD drive, etc., on which information has been recorded. Yet another form in which these instructions can be present is a website address through which information on a deleted site can be accessed via the Internet.
[0151] In some embodiments, the kit can include a non-transitory computer-readable medium having program instructions that, when executed by a processor in a computer, cause the processor to perform the aforementioned method for assembling a nucleic acid comprising a target sequence. The kit can further include instructions for assembling a nucleic acid comprising a target sequence from multiple polynucleotides, and / or a device for performing split-merge enzymatic DNA synthesis (EDS), and / or multiple Y-shaped adapters, a set of polymerase chain reaction (PCR) primers having sequences complementary to the 5' and 3' arms of the Y-shaped adapter, reagents for performing emulsion PCR, beads, ligases, exonucleases, DNA polymerases, or any combination thereof. In some embodiments, the beads are magnetic beads, preferably superparamagnetic beads. In further embodiments, the kit can further include reagents for performing gene assembly.
[0152] Utility
[0153] The methods, devices, and systems of the present disclosure can be used to assemble one or more nucleic acids comprising a target sequence and / or to identify correctly assembled nucleic acids. The methods are cell-free and allow for the selective amplification and isolation of assembled nucleic acids comprising perfect copies of the target sequence from a pooled mixture after oligonucleotide assembly. Perfect genes, even in very low copy numbers, can be readily identified and recovered by the methods disclosed herein. The methods are particularly suitable for isolating perfect copies of long target genes that are typically produced at low frequencies. Sequentially verified and indexed nucleic acids can be stored in a pooled mixture, which allows for easy selection and amplification of any individual nucleic acid or subgroup of nucleic acids as needed.
[0154] The disclosed method can be used for in vitro gene synthesis to generate nucleic acids with any desired sequence. For example, nucleic acids containing endogenous genes, exons, introns, wild-type sequences, major alleles, minor alleles, alleles related to diseases that encode proteins or RNAs, nucleic acids having mutations (such as deletions, insertions, substitutions, single nucleotide polymorphisms, frameshift mutations, missense mutations or nonsense mutations), and engineered nucleic acids having desired functions can be easily assembled and sequence-verified, and nucleic acids with perfect copies of the desired sequence can be easily amplified and isolated from the pooled mixture using the disclosed method. The disclosed method can be used in conjunction with commonly used gene assembly methods, including but not limited to Gibson assembly, polymerase chain assembly (PCA), Goldengate assembly, BioBrick assembly or other oligonucleotide ligation or polymerase-based methods.
[0155] Examples
[0156] The following examples provide further guidance on how to prepare and use the disclosed subject matter and are not intended to limit the scope of the invention. One or more aspects of these examples can be modified by those skilled in the art without departing from the spirit of the invention.
[0157] Example 1
[0158] The present disclosure can be used to discover perfectly assembled nucleic acid sequences and preferentially amplify them to obtain pure or purified nucleic acid sequences of interest. A series of steps can be carried out by an automated system designed to deliver sequence-perfect nucleic acids without using colony picking or in vivo cloning propagation. The methods herein do not require DNA to proliferate within a host. The methods herein are cell-free.
[0159] Exemplary protocol for assembling nucleic acids containing gene sequences
[0160] (a) First, construct a gene fragment library for assembling the target gene sequence. Any method can be used to generate the gene fragment library, such as but not limited to Gibson assembly ® , polymerase chain assembly (PCA), Goldengate assembly, BioBrick ® assembly or oligonucleotide ligation.
[0161] (b) Next, a Y-shaped adaptor is ligated to the fragments of the library, generating a library of fragments ligated with adaptors (also referred to as "adapted fragment library"). Preferably, to improve the ligation efficiency, the fragments of the fragment library are "improved" by end repair and / or addition of T and / or A tails prior to ligating the adaptor to the fragments. The Y-shaped adaptor can be designed to be suitable for emulsion polymerase chain reaction (emPCR), use of unique identifiers (UIDs) and / or Twinstrand Biosciences (Seattle, WA) duplex detection technology. One or more restriction sites can be included in the fragments ligated with adaptors, so that the adaptors can be excised from the correctly assembled nucleic acids using restriction endonucleases. Sample tracking barcodes on the adaptors are not required.
[0162] (c) Dilute the adapted fragment library and bind the adaptor-ligated fragments to beads. Preferably, each bead binds to only one adaptor-ligated fragment / bead.
[0163] (d) Perform single molecule emPCR on the adaptor-ligated fragments using primer pairs having sequences complementary to the A and B sequences of the Y-shaped adaptor. emPCR generates approximately 10 5 copies of each fragment on each individual bead. The emPCR conditions can be optimized for the length of the fragment library. For example, to amplify 1.5 kb fragments, longer extension cycles are used than for amplifying 3 kb fragments. This emPCR step generates an amplified fragment bead library, where each bead is ligated to a set of cloned copies of the fragment from the original fragment library.
[0164] (e) Disrupt the emulsion and purify the beads.
[0165] (f) Add unique index sequences to the fragments on the beads using split-and-pool enzymatic DNA synthesis. The number of cycles of split-and-pool synthesis depends on the number of beads in the amplified fragment bead library. The number of beads in this indexing step depends on the expected frequency of perfect cloned copies of the target sequence to be assembled. The expected frequency is inversely proportional to the length of the sequence. Longer assembled sequences require more beads and more cycles of split-and-pool synthesis (since the frequency of perfect sequences is lower). For example, adding index sequences to the fragments requires approximately 12 to 15 cycles and no more than 25 cycles (4 25 index sequences), which is substantially greater than the number of beads used (e.g., 1 × 10 16 beads). An excess of index sequences relative to the number of beads helps ensure that each bead will contain a unique index sequence.
[0166] (g) Adding a reverse primer primer extension sequence (primer binding site) to the 3' end of each indexed assembled nucleic acid on the beads using EDS synthesis. Although longer or shorter sequence lengths can be used, a fixed sequence length is typically 20 to 25 nucleotides.
[0167] (h) Amplifying the plurality of indexed assembled nucleic acids using a primer that hybridizes to the Y-adapter primer binding site (forward (F) primer extension sequence) and a primer that hybridizes to the reverse primer binding site. For example, the Y-adapter sequence can be amplified with a Pacific Biosciences sequencing primer (Menlo Park, CA) or the Y-adapter sequence can be amplified with an Oxford Nanopore Technologies adapter primer (Oxford, UK).
[0168] (i) Sequencing the amplified library using any suitable next-generation sequencing (NGS) technology. For example, the nanopore sequencing platform of Oxford Nanopore Technologies (Oxford, UK) or the single molecule real-time (SMRT) sequencing platform of Pacific Biosciences (Menlo Park, CA) can be used for sequencing. Sequencing of a small portion of the library is typically sufficient to identify at least one perfect copy of the nucleic acid sequence of interest. Preferably, the sequencing reads span the full length of the construct of the double-stranded amplified library, including (i) the forward and reverse primer sequences at the 5' end, and (ii) the forward and reverse primer extension sequences at the 3' end. The remainder of the amplified library can be stored intact on the indexed beads for subsequent recovery of the perfect sequence nucleic acids.
[0169] (j) Analyzing the sequencing results to find one or more perfect copies of the assembled nucleic acid sequence of interest. The index sequences identify which beads contain the perfect sequence copies of the desired nucleic acid sequence.
[0170] (k) Amplify the perfect clones using primers that hybridize to the bead-index sequences such that only the desired index sequences are selectively amplified. More specifically, provide first and second PCR primers to hybridize to and prime amplification of (i) one or more bead-specific index sequences associated with the perfect sequence nucleic acids, and (ii) the immobilized ends of the Y-shaped adaptors that are bound to each bead. The primers can be synthesized by any method, such as using EDS or phosphoramidite chemistry. In some embodiments, the perfect clones bound to beads associated with the perfect clones are amplified. In some embodiments, the perfect clones not bound to beads are amplified. For example, perfect clones that have been released from their previously associated beads can be amplified. In some embodiments, singleplex PCR amplification can be performed using primers that selectively hybridize to only one index sequence associated with one clone to selectively amplify a single perfect clone. In some embodiments, multiple singleplex PCR amplifications can be performed in separate reaction mixtures to amplify multiple selected clones separately, where different primers are used to hybridize to different index sequences for each different clone to be amplified. Alternatively, using multiple different primers, multiple PCR amplifications (multiplex PCR) can be performed simultaneously in a single reaction mixture, where each different primer is used to hybridize to a different index sequence associated with each clone to be amplified. One benefit of performing singleplex amplification is that primer competition can be reduced relative to multiplex PCR because some index sequences perform better than others, which can result in biases in the final amplification products. Also, if one or more of the singleplex reactions contain sequence errors introduced during PCR amplification, performing multiple singleplex PCR reactions helps increase the likelihood of obtaining perfect sequence amplification products.
[0171] (1) The perfect clones can be amplified without amplifying the sequences that include the Y-shaped adaptor sequences or the index sequences. For example, this can be accomplished using PCR with forward and reverse primers that amplify the desired sequence without amplifying the flanking adaptors or index sequences. Alternatively, a restriction site can be included in the Y-shaped adaptor such that the adaptor can be excised from the ligated construct. Another option is to use Gibson assembly into a plasmid such that the Y-shaped adaptor sequence is used as a target for recombination of the clone into the corresponding recombination site in the plasmid vector. The advantage of this option is that Escherichia coli or another host organism can be used to prepare large amounts of the perfect clones.
[0172] Additional advantages
[0173] The method of the present invention is cell-free - the workflow does not require in-cell proliferation. The entire workflow can be automated. The present invention simplifies and enables the identification, selective amplification, and isolation of ultra-rare assembled nucleic acids (1 in 1 billion, or 1 in 1 trillion, etc.). The present invention does not have the drawbacks of the "split-and-pool" strategy, which requires consecutive oligonucleotide ligation events and complex bridging templates (i.e., splint oligonucleotides as in the method described in O'Huallachain et al. (2020) Commun. Biol. 3(1):213).
[0174] The above examples are provided to illustrate the present invention, but do not limit the scope of the present invention. Other variations of the present invention will be apparent to those of ordinary skill in the art. All publications, materials, references, databases, and patents cited herein are hereby incorporated by reference for all purposes.
Claims
1. A cell-free method for identifying and amplifying a correctly assembled nucleic acid comprising a target sequence, the method comprising: ligating adapters to the 5' and 3' ends of a plurality of assembled nucleic acids, wherein the adapters comprise adapter primer binding sites and adapter capture sequences; binding the plurality of assembled nucleic acids to beads, wherein the beads comprise capture oligonucleotides attached to the bead surface, and the capture oligonucleotides specifically hybridize with the adapter capture sequences; performing clonal amplification of the plurality of assembled nucleic acids bound to the beads using the capture oligonucleotides as a first primer and a second primer that hybridizes with the adapter primer binding site, wherein amplicons generated by clonal amplification of the assembled nucleic acids bind to the beads, and wherein the beads are partitioned during clonal amplification such that amplicons on each bead are separated from other amplicons on other beads, so that amplicons cannot migrate from one bead to another; adding bead-specific index sequences to the free 3' ends of each clonally amplified assembled nucleic acid and its amplicons using split-and-pool enzymatic DNA synthesis (EDS) to generate a plurality of indexed assembled nucleic acids bound to the beads, wherein the bead-specific index sequence of each bead comprises a different unique identifier sequence; adding a reverse priming site to the free 3' end of each indexed assembled nucleic acid bound to the beads using EDS; amplifying the plurality of indexed assembled nucleic acids bound to the beads using a primer that hybridizes with the adapter primer binding site and a primer that hybridizes with the reverse priming site added by EDS; sequencing at least a portion of the amplified indexed assembled nucleic acids to identify correctly assembled nucleic acids comprising the target sequence; identifying the bead-specific index sequences of the correctly assembled nucleic acids comprising the target sequence; and performing nucleic acid amplification using a primer that hybridizes with the bead-specific index sequence of the correctly assembled nucleic acid comprising the target sequence, wherein the correctly assembled nucleic acid comprising the target sequence is selectively amplified.
2. The method according to claim 1, wherein binding the plurality of assembled nucleic acids to the beads is performed under suitable conditions, wherein no more than one assembled nucleic acid binds to each bead, and preferably, wherein binding the plurality of assembled nucleic acids to the beads is performed in a nucleic acid / bead ratio range of 0.08 to 0.
30.
3. The method according to any one of claims 1 to 2, further comprising performing deep sequencing of the amplified indexed assembled nucleic acids on one or more beads.
4. The method according to any one of claims 1 to 3, further comprising adding unique molecular identifiers to assembled nucleic acids having different sequences using EDS.
5. The method according to any one of claims 1 to 4, wherein the beads are partitioned during clonal amplification by separating each bead in a separate emulsion droplet or in a separate compartment of a multi-compartment container.
6. The method according to any one of claims 1 to 5, wherein the adapter is a linear adapter, a Y-shaped adapter, a stubby adapter, or a hairpin adapter, and preferably, wherein the double-stranded stem of the Y-shaped adapter is linked to the 5' and 3' ends of the plurality of assembled nucleic acids.
7. The method according to any one of claims 1 to 6, wherein before ligating the adaptor, the method further comprises: assembling a plurality of oligonucleotides into a plurality of polynucleotides, wherein each oligonucleotide comprises a portion of the target sequence; and assembling the plurality of polynucleotides to produce a plurality of assembled nucleic acids.
8. The method according to any one of claims 1 to 7, wherein a plurality of assembled nucleic acids are produced using an isothermal assembly method, preferably, wherein the isothermal method uses an exonuclease, a DNA polymerase, and a ligase.
9. A system, comprising: a device for performing split-ligation enzymatic DNA synthesis (EDS); a processor, wherein the processor is programmed to instruct the device to perform split-ligation EDS to add a bead-specific index sequence and a reverse primer site to the free 3' end of the cloned and amplified assembled nucleic acid bound to the bead; and a sequencer, wherein the sequencer is connected to the processor such that the processor receives sequencing data from the sequencer, preferably, wherein the sequencer is inside the device for performing split-ligation EDS.
10. The system according to claim 9, wherein the processor is further programmed to receive a nucleic acid sequence comprising a target sequence; design sequences of a plurality of polynucleotides that can be assembled into a nucleic acid comprising the target sequence, wherein each polynucleotide comprises a portion of the target sequence; and instruct a synthesizer to synthesize the plurality of polynucleotides.
11. The system according to any one of claims 9 to 10, further comprising a chamber for assembling a plurality of polynucleotides, preferably, wherein the chamber further comprises an exonuclease, a DNA polymerase, and a ligase for performing the assembly of the plurality of polynucleotides.
12. The system according to any one of claims 9 to 11, wherein the processor is further programmed to instruct the device to perform the assembly of the plurality of polynucleotides.
13. The system according to any one of claims 9 to 12, further comprising a memory storage component operably connected to the processor, wherein the memory storage component is configured to store sequencing data.
14. The system according to any one of claims 9 to 13, further comprising a plurality of Y-shaped adaptors, a set of polymerase chain reaction (PCR) primers having sequences complementary to the 5' arm and 3' arm of the Y-shaped adaptor, a PCR primer having a sequence complementary to the reverse primer site, reagents for performing emulsion PCR, beads, or a ligase or any combination thereof, preferably, wherein the beads are magnetic beads.
15. The system according to any one of claims 9 to 14, further comprising reagents for preparing a plurality of assembled nucleic acids ligated to the adaptor, wherein the reagents include reagents for blunt-ending the 3' end and 5' end of the assembled nucleic acid, reagents for phosphorylating the 5' end of the assembled nucleic acid, or reagents for adding a poly A tail to the 3' end of the assembled nucleic acid or a combination thereof.
16. The system according to any one of claims 9 to 15, further comprising reagents for performing gene assembly.
Citation Information
Patent Citations
Modified nucleotides for synthesis of nucleic acids, a kit containing such nucleotides and their use for the production of synthetic nucleic acid sequences or genes
US11059849B2
Methods and apparatus for measuring analytes
US20100137143A1
Scaffolded nucleic acid polymer particles and methods of making and using
US20100304982A1
Efficient product cleavage in template-free enzymatic synthesis of polynucleotides
US20210254114A1
Phosphoramidite compounds and processes
US4415732A