Multiplex unbiased nucleic acid amplification method
Patent Information
- Application Number
- JP2023570209
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-13
- Filing Date
- 2022-05-12
- Publication Date
- 2025-05-19
AI Technical Summary
Traditional multiplex PCR requires unique primers for each target, leading to limitations such as limited flexibility due to PCR thermodynamics, formation of primer-primer dimers, bias due to primer-induced variability, complexity of primer design, and high costs associated with synthesizing large numbers of custom oligos.
A method using a fusion protein composed of a dead CRISPR-associated protein (dCAS) linked to Tn5 transposase, which binds to target DNA, cuts it, and attaches universal primers, allowing for simultaneous amplification of multiple targets with a single primer pair.
This approach eliminates the limitations of traditional multiplex PCR by ensuring uniform amplification, reducing primer-induced bias, and simplifying primer design, while maintaining high sensitivity and specificity, enabling efficient detection of multiple targets with a single primer pair.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 189,021, filed May 14, 2021, and U.S. Provisional Patent Application No. 63 / 243,449, filed September 13, 2021. The contents of these related applications are incorporated herein by reference in their entireties. Reference to sequence listing This application is submitted with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 68EB_317327_WO_Sequence_Listing, created on May 5, 2022, and is 56.0 kilobytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.
[0002] The present disclosure relates generally to the field of molecular biology, including methods, compositions, kits and systems for multiplex, unbiased nucleic acid amplification. [Background technology]
[0003] Polymerase chain reaction (PCR) is a molecular biology technique that uses target-specific primers to exponentially amplify small, specific segments of a DNA amplicon. Multiplex PCR is the exponential amplification of more than one DNA target simultaneously, and traditional multiplex PCR requires a unique primer pair for each target, typically maximizing at 5-10 targets, e.g., 10 target DNA strands, 20 primers. The use of multiple primers creates many problems, such as limited flexibility of the target region due to PCR thermodynamics, primer-primer dimer formation, bias due to primer-induced variability, complexity of primer design, the cost and long lead time associated with the synthesis of many custom oligonucleotide primers, and the complexity of manual labor. There is a need for methods, compositions, kits, and systems for multiplex unbiased nucleic acid amplification. There is a need for methods, compositions, kits, and systems that allow the use of a single primer or a single primer pair to amplify multiple nucleic acid targets. Summary of the Invention
[0004] Disclosed herein includes compositions. In some embodiments, the compositions include a first protein complex and a second protein complex. In some embodiments, the first protein complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to a first binding site on a target double-stranded DNA (dsDNA). In some embodiments, the second protein complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to a second binding site on a target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. Disclosed herein includes compositions. In some embodiments, the compositions include a plurality of protein complex pairs, each of the plurality of protein complex pairs includes a first protein complex and a second protein complex. In some embodiments, the first protein complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to a first binding site on a target dsDNA. In some embodiments, the second protein complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to a second binding site on a target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the plurality of protein complex pairs are different from each other, and / or the second binding sites for each of the plurality of protein complex pairs are different from each other. In some embodiments, all of the plurality of protein complex pairs have the same transposome.
[0005] In some embodiments, the target dsDNA for two or more of the plurality of protein complex pairs is different. In some embodiments, the plurality of protein complex pairs includes at least five protein complex pairs. In some embodiments, the plurality of protein complex pairs includes about 5 to about 3000 protein complex pairs. In some embodiments, the adapter is a dsDNA or a DNA / RNA duplex. In some embodiments, the adapter is about 5 to about 200 base pairs in length. In some embodiments, the transposase is a Tn5 transposase, a Tn7 transposase, a mariner Tc1-like transposase, a Himar1C9 transposase, or a Sleeping Beauty transposase. In some embodiments, the transposase is a hyperactive transposase.
[0006] In some embodiments, the first programmable DNA binding unit comprises a nuclease-deficient CRISPR-associated protein (dCAS protein) and a first guide RNA (gRNA) capable of specifically binding to a first binding site of the target dsDNA, and the second programmable DNA binding unit comprises a dCAS protein and a second gRNA capable of specifically binding to a second binding site of the target dsDNA. In some embodiments, the transposome is linked to the first programmable DNA binding unit, the second programmable DNA binding unit, or both, via a linker that connects the transposase and the dCAS protein. In some embodiments, the linker comprises a peptide linker, a chemical linker, or both. In some embodiments, the transposase is present in a fusion protein comprising the dCAS protein of the first programmable DNA binding unit, the dCAS protein of the second programmable DNA binding unit, or both. The dCAS protein can be dCAS9, dCAS12, dCAS13, or dCAS14. In some embodiments, the dCAS13 protein is dCAS13a, dCAS13b, dCAS13c, or dCAS13d.
[0007] In some embodiments, the first programmable DNA binding unit comprises a first endonuclease-deficient zinc finger nuclease (ZFN) or a first endonuclease-deficient transcription activator-like effector nuclease (TALEN) capable of specifically binding to a first binding site of the target dsDNA, and the second programmable DNA binding unit comprises a second endonuclease-deficient ZFN or a second endonuclease-deficient TALEN capable of specifically binding to a second binding site of the target dsDNA. In some embodiments, the transposome is linked to the first programmable DNA binding unit, the second programmable DNA binding unit, or both, via a linker connecting the transposase and the ZFN or TALEN. In some embodiments, the linker comprises a peptide linker, a chemical linker, or both. In some embodiments, the transposase is present in a fusion protein that includes a first programmable DNA-binding unit, a ZFN or TALEN, a second programmable DNA-binding unit, a ZFN or TALEN, or both.
[0008] In some embodiments, the first programmable DNA binding unit comprises a first endonuclease-deficient meganuclease capable of specifically binding to a first binding site of the target dsDNA, and the second programmable DNA binding unit comprises a second endonuclease-deficient meganuclease capable of specifically binding to a second binding site of the target dsDNA. In some embodiments, the transposome is linked to the first programmable DNA binding unit, the second programmable DNA binding unit, or both, via a linker connecting the transposase and the endonuclease-deficient meganuclease. In some embodiments, the linker comprises a peptide linker, a chemical linker, or both. In some embodiments, the transposase is present in a fusion protein comprising the endonuclease-deficient meganuclease of the first programmable DNA binding unit, the endonuclease-deficient meganuclease of the second programmable DNA binding unit, or both.
[0009] In some embodiments, the second binding site is 1 to about 50,000 nucleotides upstream or downstream of the first binding site on the target dsDNA. In some embodiments, the second binding site is 100 to 500 nucleotides upstream or downstream of the first binding site on the target dsDNA. In some embodiments, the distance between the first binding site and the second binding site on each target dsDNA is substantially the same. In some embodiments, the distance between the first binding site and the second binding site on at least two target dsDNAs is different. In some embodiments, the composition comprises a third protein complex, the third protein complex comprising a transposome and a third programmable DNA binding unit capable of specifically binding to a third binding site on the target dsDNA. In some embodiments, the third binding site is (i) 1-50000 nucleotides upstream or downstream of the first binding site on the target dsDNA, (ii) 1-50000 nucleotides upstream or downstream of the second binding site on the target dsDNA, and / or (iii) located between the first binding site on the target dsDNA and the second binding site on the target dsDNA.
[0010] Disclosed herein includes a reaction mixture. In some embodiments, the reaction mixture comprises the composition disclosed herein; and a sample nucleic acid suspected to contain target dsDNA. In some embodiments, the reaction mixture comprises a DNA polymerase; and a plurality of dNTPs.
[0011] In some embodiments, the reaction mixture comprises one or more of the plurality of oligonucleotide probes, a buffer, and MgCl 2In some embodiments, the adaptor is covalently attached to the target dsDNA or a fragment thereof. In some embodiments, the reaction mixture comprises a plurality of dsDNA fragments comprising adaptors at both ends. In some embodiments, the sample nucleic acid comprises bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or a combination thereof. In some embodiments, the target dsDNA is genomic DNA, mitochondrial DNA, plasmid DNA, or a combination thereof. In some embodiments, the sample nucleic acid is from a biological sample, optionally the biological sample comprises feces, sputum, peripheral blood, plasma, serum, lymph nodes, respiratory tissue, exudate, or a combination thereof.
[0012] Disclosed herein includes a method for simultaneously detecting multiple target nucleic acids. In some embodiments, the method includes contacting a sample nucleic acid suspected of containing multiple target dsDNAs with multiple protein complex pairs to form a reaction mixture, each of the multiple target dsDNAs includes a target sequence adjacent to a first binding site on the target dsDNA and a second binding site on the target dsDNA, and each of the protein complex pairs includes a first protein complex and a second protein complex. In some embodiments, the first complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to the first binding site on the target dsDNA. In some embodiments, the second complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to the second binding site on the target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the multiple protein complex pairs are different from each other, the second binding sites for each of the multiple protein complex pairs are different from each other, or both. In some embodiments, all of the plurality of protein complex pairs comprise the same transposome. In some embodiments, the method includes incubating a reaction mixture to generate a plurality of dsDNA fragments, each of which comprises an adaptor and a target sequence on both ends. In some embodiments, the method includes amplifying the plurality of dsDNA fragments with a primer capable of binding to one strand of the adaptor to generate an amplified product. In some embodiments, the method includes detecting the presence of the target sequence in the amplified product as an indication of the presence of the plurality of target dsDNA. In some embodiments, detecting the presence of the target sequence in the amplified product includes contacting the amplified product with an oligonucleotide probe, each of which can specifically bind to a target sequence.
[0013] In some embodiments, the second binding site is about 1-50,000 base pairs upstream or downstream of the first binding site. In some embodiments, the adapter is a dsDNA or a DNA / RNA duplex. In some embodiments, the adapter is about 5-200 base pairs in length. In some embodiments, the primer is about 5-80 nucleotides in length. In some embodiments, the multiple target dsDNAs include genomic DNA, mitochondrial DNA, plasmid DNA, or combinations thereof. In some embodiments, the multiple target dsDNAs are from one or more organisms, one or more genes, or combinations thereof. The multiple target dsDNAs can include bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or combinations thereof. In some embodiments, the multiple target dsDNAs include genomic DNA from at least two different organisms. In some embodiments, the multiple target dsDNAs include DNA from at least five different genes.
[0014] The method can include generating a plurality of target dsDNAs from a plurality of target RNAs using a reverse transcriptase. In some embodiments, contacting the plurality of target dsDNAs with a plurality of protein complex pairs is performed at about 25°C to about 80°C. In some embodiments, incubating the reaction mixture includes incubating the reaction mixture at about 37°C to about 55°C. In some embodiments, the plurality of protein complex pairs and the plurality of target dsDNAs are present in the reaction mixture at a molecular ratio of about 2:1 to about 2,000:1. In some embodiments, the plurality of protein complex pairs and the plurality of target dsDNAs are present in the reaction mixture at a molecular ratio of about 2:1 to about 200:1.
[0015] In some embodiments, amplifying a plurality of dsDNA fragments using primers is performed using polymerase chain reaction (PCR). In some embodiments, PCR is loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinant polymerase amplification (RPA), strand displacement amplification (SDA), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple displacement amplification (MDA), branching (RAM), circular helicase-dependent amplification (cHDA), single primer isothermal amplification (SPIA), signal-mediated amplification of RNA technology (SMART), self-sustained sequence replication (3SR), genomic exponential amplification reaction (GEAR), or isothermal multiple displacement amplification (IMDA). In some embodiments, PCR is real-time PCR or quantitative real-time PCR (QRT-PCR).
[0016] In some embodiments, the method comprises labeling one or both ends of one or more of the plurality of dsDNA fragments. In some embodiments, the method comprises differentially labeling two ends of one or more of the plurality of dsDNA fragments. In some embodiments, the label comprises an anionic label, a cationic label, a neutral label, an electrochemical label, a protein label, a fluorescent label, a magnetic label, or a combination thereof. In some embodiments, the sample nucleic acid is from a biological sample. In some embodiments, the biological sample includes feces, sputum, peripheral blood, plasma, serum, lymph nodes, respiratory tissue, exudate, or a combination thereof. In some embodiments, the transposase is Tn5 transposase, Tn7 transposase, mariner Tc1-like transposase, Himar1C9 transposase, or Sleeping Beauty transposase. In some embodiments, the first programmable DNA binding unit includes a nuclease-deficient CRISPR-associated protein (dCAS protein) and a first guide RNA (gRNA) that can specifically bind to a first binding site of the target dsDNA, and the second programmable DNA binding unit includes a dCAS protein and a second gRNA that can specifically bind to a second binding site of the target dsDNA. In some embodiments, the transposome is linked to the first programmable DNA binding unit, the second programmable DNA binding unit, or both, via a linker that connects the transposase and the dCAS protein. In some embodiments, the linker comprises a peptide linker, a chemical linker, or both. In some embodiments, the transposase is present in a fusion protein comprising a dCAS protein of a first programmable DNA binding unit, a dCAS protein of a second programmable DNA binding unit, or both. In some embodiments, the dCAS protein is dCAS9, dCAS12, dCAS13, or dCAS14. In some embodiments, amplifying the multiple dsDNA fragments does not use any primer other than the primer that can bind to one strand of the adapter. [Brief description of the drawings]
[0017] [Fig. 1A-1G]Figure 1A shows a non-limiting and exemplary embodiment of a highly multiplexed non-biased DNA amplification method using universal primers. Figure 1A shows a non-limiting and exemplary embodiment of a dCAS9 protein linked to Tn5 transposase. Fusion protein: dCAS9 linked to Tn5 transposase (dCAS9-Tn5). Figure 1B shows a non-limiting and exemplary embodiment of a guide RNA complementary to a specific DNA target is attached to dCAS9. The dCAS9 portion of dCAS9-Tn5 binds to a customized gRNA specific to that DNA target. Figure 1C shows a non-limiting and exemplary embodiment of a DNA adapter is attached to the linked Tn5 transposase. Figure 1D shows a non-limiting and exemplary embodiment of a dCAS9 binds to a complementary sequence in the genomic DNA of the target specified by the guide RNA. The Tn5 transposase then cleaves the DNA and covalently attaches the adapter to the cleavage site. The ligated dCAS9 protein binds to the target, and Tn5 transposase makes a site-specific double-stranded cut on the targeted DNA and attaches a bound adapter to the cut site. FIG. 1E shows a non-limiting and exemplary embodiment showing that the resulting modified DNA region has a covalently attached adapter at each cut site. FIG. 1F shows a non-limiting and exemplary embodiment showing that a second dCAS9-Tn5 molecule with a guide RNA targeting a region downstream of the first cut site binds to the targeted DNA. Tn5 transposase again cleaves the DNA and covalently attaches an adapter to the cut site. The second dCAS9-Tn5 unit targets a region of the desired number of base pairs upstream or downstream of the first cut site, it binds to the targeted region, and Tn5 makes its double-stranded cut in the targeted DNA and attaches a bound adapter to this cut site. FIG. 1G shows a non-limiting and exemplary embodiment showing that the result is a piece of DNA with the same primer sequence at each end of the molecule. The isolated targeted DNA segment is truncated to a specific length in a programmed manner and identical primers are attached to both ends.This same process occurs simultaneously for each unique target, resulting in multiple unique target DNA segments, each bound to a universal primer sequence. [Diagram 2] FIG. 1 shows a non-limiting, exemplary schematic diagram of customized locus-specific library preparation (CLLP). [Diagram 3] FIG. 1 shows a non-limiting exemplary embodiment showing targeted sequencing using a genome editing tool (Cas9). [Figure 4A-4B] 4A-4C are non-limiting and exemplary embodiments showing a highly multiplexed, unbiased DNA amplification method. FIG. 4A shows a non-limiting and exemplary embodiment showing a single-tube reaction that can have several Cas9-Tn5 molecules, each targeting a unique region in one or more genomes. FIG. 4B shows a non-limiting and exemplary embodiment showing that the result of this reaction is several DNA molecules from all targeted regions with identical primer sequences at both ends of the molecule. Simultaneous DNA PCR amplification of targets as shown herein can be easily performed using a single primer pair. [Diagram 5] FIG. 1 shows a non-limiting, exemplary schematic diagram of a plasmid construct (3XFlag-Cas9-Fl26-Tn5; SEQ ID NO:1) for use in generating the protein complexes provided herein. [Figure 6] FIG. 1 shows a non-limiting, exemplary schematic diagram of a plasmid construct (3XFlag-Cas9-xTen-Tn5; SEQ ID NO:2) for use in generating the protein complexes provided herein. [Figure 7] FIG. 1 shows a non-limiting, exemplary schematic diagram of a plasmid construct (pET-Tn5-xTen-dCas9; SEQ ID NO:3) for use in generating the protein complexes provided herein. [Figure 8] FIG. 1 shows the relative binding sites of exemplary sgRNAs for the S. enterica InvA gene. [Figure 9] FIG. 1 shows the relative binding sites of exemplary sgRNAs for the S. enterica FliC gene. [Figure 10] FIG. 1 shows a graph of exemplary bioanalyzer data demonstrating that cleavage of genomic DNA is specific to the expected size and indicates that the guide RNA for Salmonella enterica is functional. See also Table 3. [Figure 11] 1 shows a graph of a tape station analysis showing amplification of a Tn5 generated fragment using adapter A as a primer, indicating that adapters were added to the 5' and 3' ends of the cleaved molecule. [Figure 12] 1 shows a graph of a tape station analysis showing amplification of a Tn5 generated fragment using adapter B as a primer, indicating that adapters were added to the 5' and 3' ends of the cleaved molecule. [Figure 13] FIG. 1 shows an exemplary SDS-PAGE gel analysis of recombinantly expressed and purified dCas9-Fl26-Tn5 fusion protein. The arrow points to the fusion protein band. [Figure 14] FIG. 1 shows a bioanalyzer analysis of an exemplary electrophoretic gel of recombinantly expressed and purified dCas9-Fl26-Tn5 fusion protein. [Figure 15] FIG. 1 shows an exemplary SDS-PAGE gel analysis of recombinantly expressed and purified dCas9-xTen-Tn5 fusion protein. The arrow points to the fusion protein band. [Figure 16] FIG. 1 shows bioanalyzer data from an exemplary electrophoretic analysis of recombinantly expressed and purified dCas9-xTen-Tn5 fusion proteins. [Figure 17] FIG. 1 shows an exemplary SDS-PAGE gel analysis of recombinantly expressed and purified Tn5-Fl26-dCas9 fusion protein. The arrow points to the fusion protein band. [Figure 18] FIG. 1 shows an exemplary SDS-PAGE gel analysis of recombinantly expressed and purified Tn5-xTen-dCas9 fusion proteins. The arrows point to the fusion protein bands. [Figure 19] Figure 1 shows a tape station analysis of an amplification reaction using catalytically active Cas9 alone (no fusion protein). No amplification was observed, suggesting that Cas9 itself is unable to add adapters to the 5' and 3' ends of the digested fragments. Visible signal is from samples incubated with Cas9 but not subjected to PCR. The lower peak is the 100 bp size marker and the upper peak is genomic DNA. [Figure 20] Figure 14. Tape station analysis of amplification reactions following tagmentation reactions with dCas9-Fl26-Tn5. Arrows indicate signal from reactions subjected to PCR conditions after incubation with Cas9-Tn5 fusion protein. This reaction did not contain gRNA, resulting in a broad peak indicative of random tagmentation. The lower peak is a 100 bp size marker and the upper peak is genomic DNA. [Figure 21] Figure 1 shows an exemplary tape station analysis of an amplification reaction following a tagmentation reaction with dCas9-xTen-Tn5. The arrows indicate the signal from a reaction that was subjected to PCR conditions after incubation with Cas9-Tn5 fusion protein. This reaction did not contain gRNA, resulting in a broad peak indicative of random tagmentation. The lower peak is a 100 bp size marker and the upper peak is genomic DNA. [Figure 22] Figure 1 shows a tape station analysis of amplification reactions after tagmentation with 100 nM dCas9-Fl26-Tn5 fusion protein. The arrows indicate the signal from the reaction subjected to PCR conditions after incubation with Cas9-Tn5 fusion protein. The lower peak is the 100 bp size marker and the upper peak is genomic DNA. [Figure 23]Figure 1 shows a tape station analysis of an amplification reaction following a tagmentation reaction with 1 nM dCas9-Fl26-Tn5 fusion protein. The arrows indicate the signal from the reaction subjected to PCR conditions after incubation with Cas9-Tn5 fusion protein. The lower peak is a 100 bp size marker and the upper peak is genomic DNA. [Figure 24] Figure 1 shows a tape station analysis of the amplification reaction after tagmentation reaction with 100 pM dCas9-Fl26-Tn5 fusion protein. The arrows indicate the signal from the reaction subjected to PCR conditions after incubation with Cas9-Tn5 fusion protein. The lower peak is the 100 bp size marker and the upper peak is genomic DNA. [Diagram 25] FIG. 25 shows a zoomed-in tape station analysis of an amplification reaction following a tagmentation reaction with 100 pM dCas9-Fl26-Tn5 fusion protein from FIG. 24. [Figure 26] FIG. 13 shows a tape station analysis of a tagmentation reaction with 100 pM dCas9-xTen-Tn5 fusion protein followed by an amplification reaction. Lower, 100 bp lower marker. [Figure 27] FIG. 13 shows a tape station analysis of an amplification reaction following a tagmentation reaction with 10 pM dCas9-xTen-Tn5 fusion protein. Lower, 100 bp lower marker. [Figure 28] FIG. 13 shows tape station analysis of an amplification reaction following a tagmentation reaction with 1 pM dCas9-xTen-Tn5 fusion protein. [Figure 29] FIG. 13 shows Bioanalyzer analysis of amplification from a library prepared by Tn5-only tagmentation, loaded with only one adaptor (adapter B). [Diagram 30]Bioanalyzer analysis of amplification from a library prepared by dCas9-Fl26-Tn5-induced tagmentation loaded with only one adapter (adapter B). In this experiment, a shorter incubation protocol was used. [Diagram 31] Bioanalyzer analysis of amplification from a library prepared by dCas9-Fl26-Tn5-induced tagmentation loaded with only one adapter (adapter B). In this experiment, a longer incubation protocol was used. [Diagram 32] FIG. 1 shows an exemplary bioanalyzer analysis of amplification from a library prepared by dCas9-Fl26-Tn5-guided tagmentation loaded with both adapters A and B. In this experiment, a longer incubation protocol was used. [Diagram 33] FIG. 1 shows an exemplary bioanalyzer analysis of amplification from a library prepared by dCas9-Fl26-Tn5-guided tagmentation loaded with both adapters A and B. In this experiment, a shorter incubation protocol was used. [Diagram 34] FIG. 1 shows an exemplary embodiment of DNA fragments labeled with NGS sequence adapters using the CasTn-NEBNext ligation-based library method disclosed herein. [Diagram 35] FIG. 13 shows an exemplary tape station analysis of PCR amplification from S. enterica genomic DNA samples loaded with S. enterica sgRNA and incubated with dCas9-xTen-Tn5 lower, 100 bp marker. [Diagram 36] FIG. 13 shows an exemplary tape station analysis of PCR amplification from S. enterica samples incubated with dCas9-xTen-Tn5 without sgRNA. Lower, lower 100 bp marker. [Figure 37]FIG. 1 shows an example of a dCas9-Tn5 generated fragment using a single adaptor (e.g., adaptor B). [Figure 38] FIG. 1 shows dCas9-Tn5 fragments generated from reactions in which Tn5 was loaded with two different adaptors (e.g., adaptor A and adaptor B). [Figure 39A-39B] FIG. 33 shows an example of NEBNext ligation-based library preparation for next generation sequencing. Symbols shown in key labeling portions of adapter and primer sequences. Fragments generated by dCas9-Tn5 tagmentation with NEBNext library preparation are shown in FIG. 34. [Diagram 40] FIG. 1 shows an example of tagmentation-based Nextera library preparation for next generation sequencing. [Diagram 41] FIG. 1 shows tagmentation-based library preparation using dCas9-Tn5-guided tagmentation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It is readily understood that aspects of the present disclosure, as generally described herein and illustrated in the drawings, may be arranged, substituted, combined, separated, and designed in a wide variety of different forms, all of which are expressly contemplated herein and may form part of the disclosure herein. All patents, published patent applications, other publications, and sequences from GenBank and other databases mentioned herein are hereby incorporated by reference in their entirety with respect to the relevant art. Disclosed herein includes compositions. In some embodiments, the compositions include a first protein complex and a second protein complex. In some embodiments, the first protein complex includes a transposome and a first programmable DNA binding unit that can specifically bind to a first binding site on a target double-stranded DNA (dsDNA). In some embodiments, the second protein complex includes a transposome and a second programmable DNA binding unit that can specifically bind to a second binding site on a target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor.
[0019] Disclosed herein includes compositions. In some embodiments, the compositions include a plurality of protein complex pairs, each of the plurality of protein complex pairs includes a first protein complex and a second protein complex. In some embodiments, the first protein complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to a first binding site on a target double-stranded DNA (dsDNA). In some embodiments, the second protein complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to a second binding site on a target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the plurality of protein complex pairs are different from each other, and / or the second binding sites for each of the plurality of protein complex pairs are different from each other. In some embodiments, all of the plurality of protein complex pairs have the same transposome.
[0020] Disclosed herein includes a reaction mixture. In some embodiments, the reaction mixture comprises the composition disclosed herein and a sample nucleic acid suspected to contain target dsDNA. In some embodiments, the reaction mixture comprises a DNA polymerase; and a plurality of dNTPs.
[0021] Disclosed herein includes a method for simultaneous detection of multiple target nucleic acids. In some embodiments, the method includes contacting a sample nucleic acid suspected of containing multiple target dsDNAs with multiple protein complex pairs to form a reaction mixture, each of the multiple target dsDNAs includes a target sequence adjacent to a first binding site on the target dsDNA and a second binding site on the target dsDNA, and each of the protein complex pairs includes a first protein complex and a second protein complex. In some embodiments, the first complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to the first binding site on the target dsDNA. In some embodiments, the second complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to the second binding site on the target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the multiple protein complex pairs are different from each other, the second binding sites for each of the multiple protein complex pairs are different from each other, or both. In some embodiments, all of the plurality of protein complex pairs comprise the same transposome. In some embodiments, the method includes incubating a reaction mixture to generate a plurality of dsDNA fragments, each of which comprises an adaptor and a target sequence on both ends. In some embodiments, the method includes amplifying the plurality of dsDNA fragments with a primer capable of binding to one strand of the adaptor to generate an amplified product. In some embodiments, the method includes detecting the presence of the target sequence in the amplified product as an indication of the presence of the plurality of target dsDNA. In some embodiments, detecting the presence of the target sequence in the amplified product includes contacting the amplified product with an oligonucleotide probe, each of which can specifically bind to a target sequence.
[0022] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs.See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989).For the purposes of this disclosure, the following terms are defined below.
[0023] As used herein, the term "adapter" can refer to a sequence that can facilitate amplification or sequencing of the nucleic acid to which it is attached. The attached nucleic acid can include a target nucleic acid. The attached nucleic acid can include one or more of a spatial label, a target label, a sample label, an indexing label, or a barcode sequence (e.g., a molecular label). The adapter can be linear. The adapter can be a pre-adenylated adapter. The adapter can be double-stranded or single-stranded. One or more adapters can be located at the 5' or 3' end of the nucleic acid. When the adapter includes a known sequence at the 5' and 3' ends, the known sequence can be the same or different sequences. The adapter located at the 5' and / or 3' end of the polynucleotide can hybridize to one or more oligonucleotides immobilized on a surface. The adapter can include a universal sequence in some embodiments. The universal sequence can be a region of nucleotide sequence common to two or more nucleic acid molecules. The two or more nucleic acid molecules can also have regions of different sequences. Thus, for example, the 5' adaptors may comprise the same and / or universal nucleic acid sequences, and the 3' adaptors may comprise the same and / or universal sequences. A universal sequence that may be present in different members of a plurality of nucleic acid molecules may allow for the duplication or amplification of a plurality of different sequences using a single universal primer that is complementary to the universal sequence. Similarly, at least one, two (e.g., a pair) or multiple universal sequences that may be present in different members of a population of nucleic acid molecules may allow for the duplication or amplification of a plurality of different sequences using at least one, two (e.g., a pair) or multiple single universal primers that are complementary to the universal sequence. Thus, a universal primer comprises a sequence that can hybridize to such a universal sequence. A molecule having a target nucleic acid sequence may be modified to attach a universal adaptor (e.g., a non-target nucleic acid sequence) to one or both ends of the different target nucleic acid sequences. One or more universal primers attached to the target nucleic acid may provide a site for hybridization of the universal primer.The one or more universal primers attached to the target nucleic acid can be identical to one another or different.
[0024] As used herein, the term "associated" or "associated with" can mean that two or more species are identifiable as coexisting at a time. Association can mean that two or more species are or were in the same container. Association can be an informatic association. For example, digital information about two or more species can be stored and used to determine that one or more species coexisted at a time. Association can also be a physical association. In some embodiments, two or more associated species are "tethered," "attached," or "anchored" to each other or to a common solid or semi-solid surface. Association can refer to a covalent or non-covalent means for attaching a label to a solid or semi-solid support such as a bead. Association can be a covalent bond between a target and a label. Association can include hybridization between two molecules (such as a target molecule and a label).
[0025] As used herein, the term "complementary" can refer to the ability for precise pairing between two nucleotides. For example, if a nucleotide at a given position in a nucleic acid can hydrogen bond with a nucleotide in another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial," where only a portion of the nucleotides bind, or can be complete, where there is total complementarity between the single-stranded molecules. A first nucleotide sequence can be said to be the "complement" of a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be the "reverse complement" of a second sequence if the first nucleotide sequence is complementary to the reverse (i.e., the order of the nucleotides is reversed) sequence of the second sequence. As used herein, a "complementary" sequence can refer to the "complement" or "reverse complement" of a sequence. It is understood from the present disclosure that when a molecule is capable of hybridizing to another molecule, it may be complementary or partially complementary to the hybridizing molecule.
[0026] As used herein, the term "digital counting" can refer to a method for estimating the number of target molecules in a sample. Digital counting can include determining the number of unique labels associated with targets in a sample. This methodology can be probabilistic in nature, and transforms the problem of counting molecules from a problem of finding and identifying identical molecules to a series of yes / no digital questions regarding the detection of a set of predefined labels.
[0027] As used herein, the term "nucleic acid" refers to a polynucleotide sequence, or a fragment thereof. A nucleic acid may comprise nucleotides. A nucleic acid may be exogenous or endogenous to a cell. A nucleic acid may be present in a cell-free environment. A nucleic acid may be a gene or a fragment thereof. A nucleic acid may be DNA. A nucleic acid may be RNA. A nucleic acid may comprise one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholinose, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and uyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" can be used interchangeably.
[0028] Nucleic acids can include one or more modifications (e.g., base modifications, backbone modifications) to provide the nucleic acid with new or enhanced characteristics (e.g., improved stability). Nucleic acids can include a nucleic acid affinity tag. A nucleoside can be a combination of a base and a sugar. The base portion of a nucleoside can be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleoside that further includes a phosphate group covalently linked to the sugar portion of the nucleoside. In those nucleosides that include a pentofuranosyl sugar, the phosphate group can be linked to the 2', 3', or 5' hydroxyl portion of the sugar. In forming a nucleic acid, the phosphate group can covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound can then be further connected to form a circular compound, although linear compounds are generally suitable. Additionally, linear compounds can have internal nucleotide base complementarity and thus can fold to generate a fully or partially double-stranded compound. Within nucleic acids, the phosphate groups may be commonly referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3' to 5' phosphodiester linkage.
[0029] The nucleic acids can contain modified backbones and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, such as 3'-amino phosphoramidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates, and boranophosphates having normal 3'-5' linkages, 2'-5' linked analogs, and those having reverse polarity where one or more internucleotide linkages are 3' to 3', 5' to 5', or 2' to 2' linkages.
[0030] Nucleic acids can contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; riboacetyl backbones; alkylene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and mixed N, O, S and CH 2 It may include those having other having component parts.
[0031] Nucleic acids can include nucleic acid mimetics. The term "mimetics" can be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups, and replacement of only the furanose ring can also be referred to as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety can be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of the polynucleotide can be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides can be retained and are directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone. The backbone of a PNA compound can include two or more linked aminoethylglycine units that give the PNA an amide-containing backbone. The heterocyclic base moiety can be directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone. The nucleic acid can include a morpholino backbone structure. For example, the nucleic acid can include a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidates or other non-phosphodiester internucleoside linkages can replace the phosphodiester linkages.
[0032] Nucleic acids can include linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to morpholino rings. Linking groups can link morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds can have less undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino class can be connected using different linking groups. A further class of polynucleotide mimics can be referred to as cyclohexenyl nucleic acids (CeNAs). The furanose rings normally present in nucleic acid molecules can be replaced with cyclohexenyl rings. CeNA DMT-protected phosphoramidite monomers can be prepared and used in oligomeric compound synthesis using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid strands can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to the native complex. Further modifications can include Locked Nucleic Acids (LNAs) in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage is a methylene (-CH 2 ) group, where n is 1 or 2. LNA and LNA analogs can exhibit very high duplex thermal stability with complementary nucleic acids (Tm=+3 to +10° C.), stability to 3′-exonuclease degradation, and good solubility.
[0033] Nucleic acids may also include nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C) and uracil (U)). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and other alkynyl derivatives of cytosine and pyrimidine bases, 6-azouracil, 6-amino ... , cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines, such as phenoxazine cytidines (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-ones), phenothiazine cytidines (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-ones), G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-ones), phenothiazine cytidines (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-ones), and G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-ones), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-ones), pyridoindole cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-ones).
[0034] As used herein, the term "target" can refer to a nucleic acid of interest (e.g., a target dsDNA). In some embodiments, a target can be associated with an adapter and / or a barcode. Exemplary targets suitable for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. A target can be single-stranded or double-stranded. In some embodiments, a target can be a protein, peptide, or polypeptide. In some embodiments, a target is a lipid. As used herein, "target" can be used interchangeably with "species."
[0035] As used herein, the term "reverse transcriptase" can refer to a group of enzymes that have reverse transcriptase activity (i.e., catalyze the synthesis of DNA from an RNA template). In general, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and diversity-generating retroelements, among others).
[0036] As used herein, the term "isolating nucleic acid" can refer to the purification of nucleic acid from one or more cellular components. Those skilled in the art will understand that a sample processed to "isolate nucleic acid" therefrom may contain components and impurities other than nucleic acid. A sample containing isolated nucleic acid can be prepared from a specimen using any acceptable method known in the art. For example, cells can be lysed using known lysis agents, and nucleic acid can be purified or partially purified from other cellular components. Suitable reagents and protocols for DNA and RNA extraction can be found, for example, in US Patent Application Publication No. 20100009351 and US Patent Application Publication No. 20090131650, respectively (each of which is incorporated herein by reference in its entirety). As used herein, a "template" can refer to all or a portion of a polynucleotide that contains at least one target nucleotide sequence.
[0037] As used herein, a "primer" can refer to a polynucleotide that can serve to initiate a nucleic acid chain extension reaction. The length of a primer can vary, for example, from about 5 to about 100 nucleotides, from about 10 to about 50 nucleotides, from about 15 to about 40 nucleotides, or from about 20 to about 30 nucleotides. The length of a primer can be about 10 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, or a range between any two of these values. In some embodiments, the primers have a length of 10 to about 50 nucleotides, i.e., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In some embodiments, the primers have a length of 18 to 32 nucleotides.
[0038] As used herein, a "probe" can refer to a polynucleotide that can hybridize (e.g., specifically) to a target sequence in a nucleic acid under conditions that allow hybridization, thereby allowing detection of the target sequence or amplified nucleic acid. The "target" of a probe generally refers to a sequence or a subset thereof within an amplified nucleic acid sequence that specifically hybridizes to at least a portion of a probe oligomer by standard hydrogen bonding (i.e., base pairing). A probe can include a target-specific sequence and other sequences that contribute to the three-dimensional conformation of the probe. A sequence is "sufficiently complementary" if, under appropriate hybridization conditions of the probe oligomer, it allows stable hybridization to a target sequence that is not completely complementary to the target-specific sequence of the probe. The length of the probe can vary, for example, from about 5 to about 100 nucleotides, from about 10 to about 50 nucleotides, from about 15 to about 40 nucleotides, or from about 20 to about 30 nucleotides. The length of the probe can be about 10 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 50 nucleotides, about 100 nucleotides, or a range between any two of these values. In some embodiments, the probe has a length of 10 to about 50 nucleotides. For example, the primers and probes can be at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In some embodiments, the probe can be non-sequence specific.
[0039] Preferably, the primers and / or probes can be 8-45 nucleotides in length. For example, the primers and probes can be at least 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or more nucleotides in length. Primers and probes can be modified to contain additional nucleotides at the 5' or 3' end, or both. Those skilled in the art will appreciate that the additional bases to the 3' end of the amplification primer (but not necessarily the probe) are generally complementary to the template sequence. Primer and probe sequences can also be modified to remove nucleotides at the 5' or 3' end. One of skill in the art will appreciate that to function for amplification, a primer or probe has a minimum length and annealing temperature as disclosed herein.
[0040] The primers and probes have a melting temperature (T m As used herein, "T" refers to a temperature that is less than 100° C. m " and "melting temperature" are interchangeable terms that refer to the temperature at which 50% of a population of double-stranded polynucleotide molecules becomes dissociated into single strands. m The formula for calculating T is well known in the art. m is expressed by the following formula: m The T of the hybrid polynucleotide can be calculated by: T = 69.3 + 0.41 × (G + C) % - 6 - 50 / L, where L is the length of the probe in nucleotides. m can also be estimated using the formula adopted from the hybridization assay in 1 M salt, T for PCR primers: mThe formula commonly used to calculate T is: [(number of A+T) x 2°C + (number of G+C) x 4°C]. See, e.g., CR Newton et al. PCR, 2nd ed., Springer-Verlag (New York: 1997), p. 24, which is incorporated herein by reference in its entirety. Other, more sophisticated computations are m There are techniques in the art that take into account structural and sequence characteristics for the calculation of the melting temperature of an oligonucleotide. The melting temperature of an oligonucleotide may depend on the complementarity between the oligonucleotide primer or probe and the binding sequence, and on the salt conditions. In some embodiments, the oligonucleotide primers or probes provided herein have a T of less than about 90° C. in 50 mM KCl, 10 mM Tris-HCl buffer. m For example, about 89° C., 88, 87, 86, 85, 84, 83, 82, 81, 80 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 53, 52, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39° C. or less, including ranges between any two of the recited values.
[0041] In some embodiments, the primers disclosed herein, e.g., amplification primers, can be provided as an amplification primer pair, e.g., comprising a forward primer and a reverse primer (a first amplification primer and a second amplification primer). Preferably, the forward primer and the reverse primer have a T that differs by no more than 10° C., e.g., by less than 10° C., less than 9° C., less than 8° C., less than 7° C., less than 6° C., less than 5° C., less than 4° C., less than 3° C., less than 2° C., or less than 1° C. m has.
[0042] Primer and probe sequences can be modified by having nucleotide substitutions in the oligonucleotide sequence (relative to the target sequence), provided that the oligonucleotide contains sufficient complementarity to specifically hybridize to the target nucleic acid sequence. In this manner, at least 1, 2, 3, 4, or up to about 5 nucleotides can be substituted. As used herein, the term "complementary" can refer to sequence complementarity between regions of two polynucleotide strands, or between two regions of the same polynucleotide strand. A first region of a polynucleotide is complementary to a second region of the same or different polynucleotide if at least one nucleotide of the first region can base pair with a base of the second region when the two regions are arranged antiparallel. Thus, it is not necessary for two complementary polynucleotides to base pair at every nucleotide position. "Fully complementary" can refer to a first polynucleotide that is 100% or "fully" complementary to a second polynucleotide, and thus base pairs at every nucleotide position. "Partially complementary" can also refer to a first polynucleotide that is not 100% complementary (e.g., 90%, or 80%, or 70% complementary) and contains mismatched nucleotides at one or more nucleotide positions. In some embodiments, the oligonucleotide comprises a universal base.
[0043] As used herein, the term "sufficiently complementary" can refer to a contiguous nucleobase sequence that can hybridize to another base sequence by hydrogen bonding between a series of complementary bases. Complementary base sequences can be complementary at every position of the oligomer sequence by using standard base pairing (e.g., G:C, A:T, or A:U), or can be non-complementary (including non-basic positions), but contain one or more residues such that the entire complementary base sequence can specifically hybridize to another base sequence under suitable hybridization conditions. Contiguous bases can be at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% complementary to the sequence to which the oligomer is intended to hybridize. A substantially complementary sequence can refer to a sequence with a percentage identity of 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 75, 70 or less, or any number of ranges therebetween, relative to a reference sequence. Those skilled in the art can easily select appropriate hybridization conditions, which can be predicted based on the base sequence composition, or can be determined using routine testing (see, for example, Green and Sambrook, Molecular Cloning, A Laboratory Manual, 4th ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2012)). As used herein, the term "multiplex PCR" refers to a type of PCR in which more than one primer set is included in the reaction, allowing one single target, or two or more different targets, to be amplified in a single reaction vessel (e.g., tube). Multiplex PCR can be, for example, real-time PCR.
[0044] Provided herein include nucleic acid amplification methods that allow for highly multiplexed single primer PCR reactions that are non-biased, sensitive and specific. In some embodiments, the methods, compositions, kits and systems disclosed herein can allow for single tube reactions with unlimited DNA targets. In some embodiments, a fusion protein composed of dead CRISPR associated protein (dCAS) linked with transposase (Tn5) can be used to generate many unique, custom, easily PCR amplified DNA targets with a single universal primer, thus eliminating many limitations of traditional multiplex PCR. PCR is a molecular biology technique that exponentially amplifies small specific segments of a DNA amplicon using target specific primers. Multiplex PCR is the exponential amplification of multiple DNA targets simultaneously, traditional multiplex PCR requires a unique primer pair for each target, typically 5-10 targets, e.g., 10 target DNA strands, maxing out at 20 primers. The use of multiple primers creates many problems, as explained below.
[0045] The present disclosure describes methods, compositions, kits and systems for highly multiplexed PCR reactions that occur in a manner that completely eliminates most of the limitations of conventional multiplexed PCR. The methods allow for optimized highly multiplexed PCR reactions by utilizing a single primer pair for unlimited targets, for example, by using a fusion protein composed of dCAS protein linked to a Tn5 transposase preloaded with primers. In some embodiments, the methods, compositions, kits and systems can be used to amplify 2, 5, 10, 15, 20, 30, 35, or any number or range between any two of these numbers of DNA targets using only one primer pair.
[0046] In some embodiments, the dCAS9 protein binds to the target DNA, and its attached Tn5 transposase uses preloaded primers to make a DNA double-strand break and bind to both cut ends. Instead of floating paired primers until they encounter a match to bind, the method disclosed herein takes advantage of the high specificity and sensitivity of the dCAS9 protein programmed with guide RNA to rapidly search for their specific DNA target. For example, if dCAS9 identifies its target, it binds to its specific site. When the transposase is activated, it efficiently cuts the targeted site of the DNA strand and attaches primers to both cut ends. All targets are bound by the same primer, and the dCAS9 pair can be programmed to identify and bind to a desired number of unique targets. Because the primers bind to the Tn5 transposase, the limitations due to dimers (primer-primer binding) are eliminated. Because all targets are bound to a single primer, amplification can be truly optimized and uniform.
[0047] For amplification of any nucleic acid (e.g., DNA) target, two dCAS9-Tn5 binding proteins can be configured to identify, cleave, and apply primers to that target. The first binding protein unit is programmed with a specific gRNA to identify and bind the desired DNA target. The second dCAS9-Tn5 binding protein unit is programmed with a complementary gRNA that targets the same region as the first gRNA but is programmed to bind to that target at a space of a few base pairs (e.g., <300 bp) upstream or downstream of the first dCAS9-Tn5 binding unit. Regardless of how the dCAS9 protein pairs are programmed to their unique targets, all Tn5 transposase molecules can be loaded with two short DNA primers in the same way. The end result of the action of the binding protein units is a selected DNA segment with a primer attached to each end. (Figure 1A-Figure 1G).
[0048] [Table 1]
[0049] Challenges and limitations to conventional multiplex PCR are well known in the art and are described, for example, at https: / / www.lcsciences.com / discovery / overcome-common-challenges-to-multiplex-pcr-with-innovative-relay-pcr-and-omega-primer-technologies / . Non-limiting examples of advantages of the multiplex PCR methods disclosed herein, including those based on the dCAStellaTn5 (constellation) method, are described below.
[0050] Limited flexibility of the target region due to PCR thermodynamics This is a problem related to the various optimal functions of each specific primer when multiple primers are used simultaneously.Under current practice, this challenge requires maintaining similar melting temperatures across all primers, avoiding complementary or similar DNA target sequences, and minimizing cross-hybridization in target selection due to primer non-specificity.Since the method disclosed herein allows the use of a single primer, in some embodiments, due to the high sensitivity and specificity of dCAS9, users can select for any number of targets and for any arbitrary target.
[0051] Primer-primer dimer formation This is a challenging limitation of PCR because dimers clog the gears of PCR in many ways. Primer-primer dimers are when primers bind to each other instead of to the DNA target. When primers bind to each other, they are unavailable for target amplification. Amplification amplifies the dimers (primer-primer strands) but not the uncomplemented target, contributing to uneven amplification. Furthermore, the presence of these dimers in the amplified sample creates a kind of "noise" like static that obscures the resulting image. The primers used in the methods disclosed herein are not free floating (e.g., floating alone in the reaction solution) but are bound to a protein (e.g., transposase-like Tn5), so dimers do not form. In some embodiments, the dCAS9 protein only binds to its nucleic acid (e.g., DNA) target; once bound, the complex (e.g., dCAS9-Tn5) units remain bound, so no additional "noise" is introduced.
[0052] Bias due to primer-induced variability The methods disclosed herein, in some embodiments, require only one primer or one primer pair, so that the thermodynamics of the reaction can be optimized for a single primer (or primer pair). This is also important because traditional multiplexing requires a large number of unique and delicate primers to be exposed to a single temperature, resulting in uneven functionality, contributing to uneven amplification, or even the primers failing to target specific regions that require conditions outside of a preset range. With dCAStellaTn5, any region can be targeted, and reaction conditions can be set to optimize the function of that one single primer. PCR is believed to produce quantifiable results, but inherent, primer-induced variability limits this ability. With universal primers, the consistent reliability of results enhances the ability to quantitate. The complexity of primer design The methods disclosed herein allow the user to select the primers of their choice and also allow for the selection of any and as many targets as desired without the compromise of allowing multiple primers to work under the same conditions. Costs and long lead times associated with synthesizing large numbers of custom oligos Again, with the compositions disclosed herein (e.g., dCAStellaTn5), users can choose the simplest, cheapest, and most preferred primers. Although the guide RNA that programs the dCAS9 protein must be synthesized, customizable guide RNAs are readily and widely available commercially.
[0053] Complexity of on-site procedures In the absence of primers in solution, the clean-up step that is often part of current multiplex PCR is eliminated.
[0054] The highly multiplexed PCR method disclosed herein, without many of the limitations of conventional multiplexed PCR currently available, opens the door to a myriad of embodiments and is broadly transformative across many industries. The method is generally applicable in many fields, e.g., medical diagnostics, in view of its broad applicability, high sensitivity and specificity, and ability to detect a large number of targets. The method easily evaluates a broad menu of specimen types for ID for any pathogen that has DNA (bacteria, fungi, protozoa, and DNA viruses) and fits the specimen types used for major infectious disease (ID) syndromes. The method can be used with DNA extraction, amplification, and / or detection steps, as well as with similar platforms. Currently, the entire process from sample to amplification / detection result can take approximately 90 minutes. The method disclosed herein can be utilized with any PCR amplification and detection platform, the advantage of which is that it is easily accessible to acute care customers, but with comparable quality and reliability in addition to any range of infinite DNA targets. Existing platforms are compatible, so instrumentation is not a barrier in terms of space and cost. The method can be integrated into platforms to optimize the end-to-end user experience. Its broad accessibility to existing and potentially customized platforms, combined with inherent improvements in specificity, sensitivity, and elimination of the limitations of conventional multiplex PCR, provides diagnostic power and confidence.
[0055] Nucleic acid amplification techniques that the methods disclosed herein can use to amplify multiple DNA targets can vary, including, but not limited to, isothermal DNA amplification techniques, such as LAMP, RPA, SDA, and HDA. A variety of transposases (including hyperactive transposases) can be used in the methods disclosed herein, including, for example, Tn5 transposase, mariner Tc1-like transposon, Himar1C9 transposase, Sleeping Beauty transposase, Tn7 transposon, and combinations thereof. The inserted primer region can be labeled, for example, with one or more anionic, cationic, neutral, fluorescent, optical, or magnetic particles. The labeled molecules can have, for example, two different tags per molecule generated (e.g., a magnetic tag on one end of the DNA molecule and a fluorescent tag on the other end for single molecule separation and visualization). As another example, an avidin tag can be present on one end of the DNA molecule and a fluorescent tag can be present on the other end for single molecule capture and imaging. Or similarly, for capture and chemical imaging, an avidin tag is present at one end of the DNA molecule and a ferrocene molecule at the other end. Adding different tags to each side of the molecule can lead to many different variations and options. The molecules can be separated by gel or capillary electrophoresis, and the color is detected as in qPCR. Alternatives to dCas9 proteins can be used for programmable DNA binding activity, including zinc fingers that are not bound to FOK1 nuclease. In some embodiments, TALEN molecules that do not contain FOK1 nuclease can be used. Further non-limiting examples of CAS proteins that can be used in the methods, compositions, kits and systems disclosed herein include, but are not limited to, CAS12, CAS13 and CAS14. Recombinases in combination with sequence-specific primers can be used as programmable DNA binding molecules in some embodiments.The method can include the use of genome editing tools as programmable tools to target specific regions of the genome and the use of transposases to cut and paste the adapters required to create a sequencing library (see FIG. 2 for an example of an exemplary embodiment using the Oxford nanopore system). In some embodiments, targeted sequencing can be used without the assistance of transposases by using genome editing tools (e.g., Cas proteins, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), Argonaute proteins). This can further result in a programmable fragmentation method of nucleic acids (FIG. 3) that can be used to create locus-specific sequencing libraries.
[0056] Provided herein include a DNA amplification method for unbiased, highly multiplexed, single primer reactions made possible by using a fusion protein composed of dead CRISPR-associated (dCAS) protein linked to Tn5 transposase to generate custom, easily PCR amplified DNA targets with a single universal primer. Each of the following patent application publications and references is incorporated by reference in its entirety: U.S. Patent Application Publication No. 200377881, U.S. Patent Application Publication No. 20200190487, U.S. Patent Application Publication No. 20190093090, and U.S. Patent Application Publication No. 20180305683; Sway P. Chen and Harris H. Wang, An Engineered Cas-Transposon System for Programmable and Site-Directed DNA Transpositions, The CRISPR Journal Volume 2, Number 6, 2019; Hatice S. Kaya-Okur Et Al, CUT&Tag for efficient epigenomic profiling of small samples and single cells, Nature Communications, (2019) 10:1930; Simone Picelli Et Al, Tn5 transposase and tagmentation procedures for massively scaled sequencing projects, Genome Research, 24:2033-2040.
[0057] In some embodiments, a composition is provided. In some embodiments, the composition comprises a first protein complex and a second protein complex. In some embodiments, the first protein complex comprises a transposome and a first programmable DNA binding unit capable of specifically binding to a first binding site on a target double-stranded DNA (dsDNA). In some embodiments, the second protein complex comprises a transposome and a second programmable DNA binding unit capable of specifically binding to a second binding site on a target dsDNA. In some embodiments, the transposome comprises a transposase and two copies of an adaptor.
[0058] In some embodiments, a composition is provided. In some embodiments, the composition comprises a plurality of protein complex pairs, each of the plurality of protein complex pairs comprising a first protein complex and a second protein complex. In some embodiments, the first protein complex comprises a transposome and a first programmable DNA binding unit capable of specifically binding to a first binding site on a target dsDNA. In some embodiments, the second protein complex comprises a transposome and a second programmable DNA binding unit capable of specifically binding to a second binding site on a target dsDNA. In some embodiments, the transposome comprises a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the plurality of protein complex pairs are different from each other, and / or the second binding sites for each of the plurality of protein complex pairs are different from each other. In some embodiments, all of the plurality of protein complex pairs have the same transposome. Some embodiments provide a reaction mixture. In some embodiments, the reaction mixture comprises a composition disclosed herein; and a sample nucleic acid suspected of containing a target dsDNA. In some embodiments, the reaction mixture comprises a DNA polymerase; and a plurality of dNTPs. In some embodiments, the reaction mixture comprises one or more of a plurality of oligonucleotide probes, a buffer, and MgCl 2 In some embodiments, the reaction mixture comprises a plurality of dsDNA fragments comprising adaptors at both ends.
[0059] In some embodiments, a method for simultaneous detection of multiple target nucleic acids is provided. In some embodiments, the method includes contacting a sample nucleic acid suspected of containing multiple target dsDNAs with multiple protein complex pairs to form a reaction mixture, each of the multiple target dsDNAs includes a target sequence adjacent to a first binding site on the target dsDNA and a second binding site on the target dsDNA, and each of the protein complex pairs includes a first protein complex and a second protein complex. In some embodiments, the first complex includes a transposome and a first programmable DNA binding unit capable of specifically binding to the first binding site on the target dsDNA. In some embodiments, the second complex includes a transposome and a second programmable DNA binding unit capable of specifically binding to the second binding site on the target dsDNA. In some embodiments, the transposome includes a transposase and two copies of an adaptor. In some embodiments, the first binding sites for each of the multiple protein complex pairs are different from each other, the second binding sites for each of the multiple protein complex pairs are different from each other, or both. In some embodiments, all of the plurality of protein complex pairs comprise the same transposome. In some embodiments, the method includes incubating a reaction mixture to generate a plurality of dsDNA fragments, each of which comprises an adaptor and a target sequence on both ends. In some embodiments, the method includes amplifying the plurality of dsDNA fragments with a primer capable of binding to one strand of the adaptor to generate an amplification product. In some embodiments, the method includes detecting the presence of the target sequence in the amplified product as an indication of the presence of the plurality of target dsDNA.
[0060] Contacting the multiple target dsDNAs with the multiple protein complex pairs can be performed at about 25 to about 85°C (e.g., 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, or a number or range between any two of these values). Incubating the reaction mixture can include incubating the reaction mixture at about 37 to about 55°C (e.g., 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, or a number or range between any two of these values).
[0061] The multiple protein complex pairs and the multiple target dsDNAs may be from about 2:1 to about 2,000:1 (e.g., 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58 :1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1 , 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, or a number or range between any two of these values).In some embodiments, the plurality of protein complex pairs and the plurality of target dsDNAs are from about 2:1 to about 200:1 (e.g., 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21: 1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51 :1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, 200:1, or a number or range between any two of these values).
[0062] At least two binding sites of the plurality of protein complexes can be on the same target dsDNA. At least two binding sites of the plurality of protein complexes can be separated by about 1 to about 50,000 nucleotides on the same target dsDNA. In some embodiments, at least two binding sites of the plurality of protein complexes are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21 The sequence can be separated by 6, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 70000, 80000, 9000, 100000 nucleotides or by about these values, or by a value or range between any two of these values.In some embodiments, at least two binding sites of the plurality of protein complexes are present on the same target dsDNA at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 102, 104, 105, 106, 107, 108, 10 The sequences can be separated by 8, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 nucleotides. The distance between the binding sites of a pair of the multiple protein complexes can be substantially the same as the distance between the binding sites of another pair of the multiple protein complexes. The distance between the binding sites of a pair of the multiple protein complexes can be different from the distance between the binding sites of another pair of the multiple protein complexes. The distance between the first binding site and the second binding site on each target dsDNA can be substantially the same. The distance between the first binding site and the second binding site on at least two target dsDNAs can be different.
[0063] The first binding site and the second binding site can be on the same strand of the target dsDNA. The first binding site and the second binding site can be on different strands of the target dsDNA. In some embodiments, the second binding site is at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 , 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 upstream or downstream.
[0064] In some embodiments, the composition comprises a third protein complex. The third protein complex can comprise a transposome and a third programmable DNA binding unit capable of specifically binding to a third binding site on a target dsDNA. In some embodiments, the third binding site is at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800 , 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 nucleotides upstream or downstream. In some embodiments, the third binding site is located between the first binding site on the target dsDNA and the second binding site on the target dsDNA.
[0065] The number of protein complex pairs can be different in different embodiments. In some embodiments, the number of protein complex pairs is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 101, 102, 103, 104, 105, 106, 107, 108, 9, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values.
[0066] At least two binding sites of the plurality of protein complexes can be on different strands of the target dsDNA. At least two of the plurality of protein complexes can specifically bind to different target dsDNA. The plurality of protein complexes can specifically bind to about 2 to 5000 target dsDNA. In some embodiments, the plurality of protein complexes can specifically bind to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 71 70, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710 , 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3250, 3500, 3750, 4000, 4250, 4500, 4750, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or a number or range between any two of these values.
[0067] Transposome In some embodiments, the transposome comprises a transposase and two copies of an adaptor. At least two of the plurality of protein complexes can comprise the same transposome. All of the plurality of protein complexes can comprise the same transposome. All of the plurality of protein complexes can comprise the same transposase. The transposase can be a Tn5 transposase, a Tn7 transposase, a mariner Tc1-like transposase, a Himar1C9 transposase, or a Sleeping Beauty transposase. The transposase can be a hyperactive transposase.
[0068] In some embodiments, the transposase is Tn5, Tn7, MuA, or Vibrio harveyi transposase, or an active mutant thereof. In other embodiments, the transposase is Tn5 transposase or a mutant thereof. In some embodiments, the Tn5 transposase is a hyperactive Tn5 transposase, or an active mutant thereof. In some embodiments, the Tn5 transposase is a Tn5 transposase as described in WO 2015 / 160895, which is incorporated herein by reference. In some embodiments, the Tn5 transposase is a hyperactive Tn5 having mutations at positions 54, 56, 372, 212, 214, 251, and 338 relative to wild-type Tn5 transposase. In some embodiments, the Tn5 transposase is a hyperactive Tn5 with the following mutations relative to wild-type Tn5 transposase: E54K, M56A, L372P, K212R, P214R, G251R, and A338V. In some embodiments, the Tn5 transposase is a fusion protein. In some embodiments, the Tn5 transposase fusion protein comprises a fusion elongation factor Ts (Tsf) tag. In some embodiments, the Tn5 transposase is a hyperactive Tn5 transposase that comprises mutations at amino acids 54, 56, and 372 relative to the wild-type sequence. In some embodiments, the hyperactive Tn5 transposase is a fusion protein. In some embodiments, the recognition site is a Tn5-type transposase recognition site (Goryshin and Reznikoff, J. Biol. Chem., 273:7367, 1998).
[0069] The transposase may comprise a single protein or may comprise multiple protein subunits. The transposase may be an enzyme capable of forming a functional complex with a transposon end or a transposon end sequence. In some embodiments, the transposase complex comprises a transposase (e.g., Tn5 transposase) dimer comprising a first monomer and a second monomer. In some embodiments, the transposome complex comprises a dimer of two molecules of transposase.
[0070] The transposase and / or transposome may vary depending on the embodiment. The transposase may include Tn5 transposase. Transposases include Tn transposases (e.g., Tn3, Tn5, Tn7, Tn10, Tn552, Tn903), MuA transposase, Vibhar transposase (e.g., from Vibrio herberii), Ac-Ds, Ascot-1, Bs1, Cin4, Copia, En / Spm, F element, hobo, Hsmar1, Hsmar2, IN(HIV), IS1, IS2, IS3, IS4, IS5, IS6, IS10, IS21, IS30, IS50, IS51, It may be IS150, IS256, IS407, IS427, IS630, IS903, IS911, IS982, IS1031, ISL2, L1, mariner, P element, Tam3, Tc1, Tc3, Tel, THE-1, Tn / O, TnA, Tn3, Tn5, Tn7, Tn10, Tn552, Tn903, Tol1, Tol2, Tn10, Ty1, any prokaryotic transposase, or any transposase related to and / or derived from those listed above. In some embodiments, a transposase related to and / or derived from a parent transposase can comprise a peptide fragment having at least about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% amino acid sequence identity to a corresponding peptide fragment of the parent transposase. The peptide fragment can be at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, or 500 amino acids in length. For example, a Tn5-derived transposase can contain a peptide fragment that is 50 amino acids in length and is about 80% homologous to the corresponding fragment in the parent Tn5 transposase. In some cases, insertion can be promoted and / or induced by the addition of one or more cations. Cations can be, for example, Ca 2+ , Mg 2+ and Mn2+ The cation may be a divalent cation such as
[0071] adapter The transposome can include a transposase and two copies of an adapter. The adapter can be a dsDNA or a DNA / RNA duplex. The adapter can be about 3-200 base pairs in length (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200 nucleotides in length, or a number or range between any two of these values). In some embodiments, the adapter can be 3-500 base pairs in length (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 300, 500 nucleotides in length, or a number or range between any two of these values). In some embodiments, the adapter comprises a barcode (e.g., a stochastic barcode). In some embodiments, the adapter comprises a universal sequence. In some embodiments, the adapter has a single-stranded portion and / or a double-stranded portion. In some embodiments, the adapter comprises a transposon end sequence that binds to a transposase. The transposon end sequence can be double-stranded. In some embodiments, the transposon end sequence is a mosaic end (ME) sequence. In certain embodiments, the transposon end is a mosaic end, or a hyperactive form of the transposon end. The adapter sequence can be attached to one of the two transposon end sequences. Thus, in some embodiments, the adapter can include an ME sequence or an ME' sequence. In some embodiments, the adapter can have a structure that can enhance suppression (repressive structure of the adapter) or inhibit suppression (permissive structure of the adapter). For example, the adapter can be tuned, for example, by changing its sequence to affect the level of suppression. The level of suppression can be related to the amount of amplification of the artifact in the sample. The adapter can include a palindromic sequence.
[0072] The methods provided herein can generate multiple dsDNA fragments, each of which includes an adapter at both ends of a target sequence. The adapter can be covalently attached to the target dsDNA or a fragment thereof in the methods provided herein. The multiple dsDNA fragments can include an adapter at both ends. The adapter at each end can be the same or different (e.g., different sequences, different linker functionalities, different detectable moieties, etc.). In some embodiments of the compositions and methods provided herein, the protein complex pair each includes the same adapter, while in other embodiments, the protein complex pair includes different adapters. The transposome of the first protein complex and the transposome of the second protein complex of the protein complex pair can include the same adapter or different adapters. The transposome can include the same adapter or multiple adapters that differ with respect to at least one characteristic (e.g., different sequences, different linker functionalities, different detectable moieties, etc.).
[0073] The adaptor can include a detectable moiety (e.g., a detectable label) as provided herein. The adaptor can be labeled, for example, with one or more anionic, cationic, neutral, fluorescent, optical, or magnetic particles. The adaptor can include one or more nucleotides (or analogs thereof) that are modified or otherwise non-naturally occurring. For example, the adaptor can include one or more nucleotide analogs (e.g., LNA, FANA, 2'-O-Me RNA, 2'-fluoro RNA, etc.), linkage modifications (e.g., phosphorothioate, 3'-3' and 5'-5' back linkages), 5' and / or 3' end modifications (e.g., 5' and / or 3' amino, biotin, DIG, phosphate, thiol, dye, quencher, etc.), one or more fluorescently labeled nucleotides, or any other feature that provides a desired functionality. The method can include labeling one or both ends of one or more of the plurality of dsDNA fragments (e.g., with a detectable label). The method can include differentially labeling two ends of one or more of the plurality of dsDNA fragments. The labeling can include labeling with a detectable label (e.g., an anionic label, a cationic label, a neutral label, an electrochemical label, a protein label, a fluorescent label, a magnetic label, or a combination thereof). The method can include concentrating the labeled dsDNA fragments, capturing the labeled dsDNA fragments, isolating the labeled dsDNA fragments, and / or visualizing the labeled dsDNA fragments. The method can include monitoring the detectable label (e.g., chemical monitoring). The adaptor can include a linker functional group individually selected from the group consisting of biotin, streptavidin, a primary amine, an aldehyde, a ketone, and any combination thereof. In some embodiments, a solid support (e.g., a synthetic particle and / or a planar surface) is provided. The solid support can have magnetic properties. The solid support can include a support functional group individually selected from the group consisting of C6, biotin, streptavidin, a primary amine, an aldehyde, a ketone, and any combination thereof. In some embodiments, the adaptor and the solid support are attached to each other.In some embodiments, the supporting functional group and the linker functional group are attached to one another.
[0074] Some embodiments of the methods provided herein include generating a plurality of dsDNA fragments that include a distinct adapter at each end (e.g., opposite ends). The plurality of dsDNA fragments can include a plurality of labeled dsDNA fragments, where the labeled dsDNA fragments include a label on the adapter at one or both ends. The labeled dsDNA fragments that include different target sequences can be distinguished based on size (first level of multiplexing) and label (second level of multiplexing) (e.g., size / label profile). In some such embodiments, labeled dsDNA fragments of a particular size (first level of multiplexing) and a particular tag (e.g., chemical and fluorescent, second level of multiplexing) are generated for a particular target sequence. For example, if fluorescent labeling is used as an example of the second level of multiplexing, the method can include generating a first labeled dsDNA fragment of 100 base pairs labeled with a blue fluorophore, a second labeled dsDNA fragment of 200 base pairs labeled with a yellow fluorophore, a third labeled dsDNA fragment of 300 base pairs with a red fluorophore, a fourth labeled dsDNA fragment of 400 base pairs with a green fluorophore, a fifth labeled dsDNA fragment of 500 base pairs with a blue fluorophore, etc. Thus, the first labeled dsDNA fragment, the second labeled dsDNA fragment, the third labeled dsDNA fragment, the fourth labeled dsDNA fragment, and the fifth labeled dsDNA fragment can be distinguished from each other based on their size / labeling profiles. For this embodiment, the first level of multiplexing includes increasing the size of 100 bp and alternating four fluorophores (second level of multiplexing), while both levels of multiplexing can be adjusted according to the needs of the user. The number of labeled dsDNA fragments that contain different target sequences that can be distinguished from one another can vary in different embodiments.In some embodiments, the number of labeled dsDNA fragments having a characteristic size / labeling profile is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, The number of dsDNA fragments may be 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or about these values, or a number or range between any two of these values. The labeled dsDNA fragments can be detected using the methods and compositions provided herein and known to those skilled in the art. In some embodiments, detecting the labeled dsDNA fragments includes the use of electrophoresis (e.g., gel electrophoresis, capillary electrophoresis). Gel electrophoresis involves the separation of nucleic acids through a matrix, typically a cross-linked polymer, using electromotive forces that pull molecules through the matrix. The molecules migrate through the matrix at different rates, causing a separation between products that can be visualized and interpreted through a number of methods, including but not limited to autoradiography, phosphorimaging, and staining with nucleic acid chelating dyes. Capillary gel electrophoresis (CGE) is a combination of traditional gel electrophoresis and liquid chromatography, using a medium such as polyacrylamide in a narrow-bore capillary to produce rapid, highly efficient separation of nucleic acid molecules down to single-base resolution. CGE can be combined with laser-induced fluorescence (LIF) detection, which can detect as few as six molecules of stained DNA. CGE / LIF detection generally involves the use of fluorescent DNA intercalating dyes, including ethidium bromide, YOYO, and SYBR® Green 1, and can include the use of fluorescent DNA derivatives in which the fluorescent dye is covalently attached to DNA.Using this method, simultaneous identification of several different target sequences (eg, products from a multiplex reaction) can be achieved.
[0075] The adaptor provided herein can include a barcode, e.g., a stochastic barcode, and can include one or more labels. For example, Fu et al., Proc Natl Acad Sci USA, 2011 May 31, 108(22):9026-31; US Patent Application Publication No. 2011 / 0160078; Fan et al., Science, 2015 February 6, 347(6222):1258367; US Patent Application Publication No. 2015 / 0299784; and International Publication No. 2015 / 031691 describe barcoding, such as stochastic barcoding, the contents of each of which, including any supporting or supplementary information or materials, are incorporated herein by reference in their entirety. In some embodiments, the barcode disclosed herein can be a stochastic barcode, which can be a polynucleotide sequence that can be used to stochastically label (e.g., barcode, tag) a target. A barcode may be referred to as a probabilistic barcode if the ratio of the number of distinct barcode sequences of the probabilistic barcode to the number of occurrences of any of the targets to be labeled may be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or about these values, or a value or range between any two of these values. The targets may be mRNA species that include mRNA molecules with identical or nearly identical sequences. A barcode can be referred to as a probabilistic barcode if the ratio of the number of different barcode sequences of the probabilistic barcode to the number of occurrences of any of the targets to be labeled is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1. The barcode sequences of the probabilistic barcode can be referred to as molecular labels.
[0076] The adapter and / or barcode may include one or more universal labels. In some embodiments, the one or more universal labels may be the same for all barcodes and / or adapters. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a sequencing primer. The sequencing primer may be used to sequence the barcode that includes the universal label. The sequencing primer (e.g., a universal sequencing primer) may include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a PCR primer. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a sequencing primer and a PCR primer. The nucleic acid sequence of the universal label that can hybridize to a sequencing or PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the barcode. The universal label may include a sequence that can be used to extend the barcode or a region within the barcode. A universal label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or about these values, or a number or range between any two of these values, nucleotides in length. For example, a universal label can include at least about 10 nucleotides. A universal label can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.
[0077] A barcode, e.g., a probabilistic barcode, can include one or more labels. Exemplary labels can include a universal label, a cell label, a barcode sequence (e.g., a molecular label), a sample label, a plate label, a spatial label, and / or a pre-spatial label. A barcode can include a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label. The order of the different labels (including but not limited to the universal label, the dimensional label, the spatial label, the cell label, and the molecular label) in a barcode can vary. For example, the universal label can be the 5'-most label and the molecular label can be the 3'-most label. The spatial label, the dimensional label, and the cell label can be in any order. In some embodiments, the universal label, the spatial label, the dimensional label, the cell label, and the molecular label are in any order. In some embodiments, the labels of a barcode (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence) may be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0078] Labels, e.g., cellular labels, may comprise a set of unique nucleic acid subsequences of defined length, e.g., seven nucleotides each (equal to the number of bits used in some Hamming error-correcting codes), and may be designed to provide error-correcting capabilities. A set of error-correcting subsequences comprising seven nucleotide sequences may be designed such that any pairwise combination of sequences in the set exhibits a defined "genetic distance" (or number of mismatched bases), e.g., a set of error-correcting subsequences may be designed to exhibit a genetic distance of three nucleotides. In this case, the recitation of the error-correcting sequences in a set of sequence data for a labeled target nucleic acid molecule (described more fully below) may allow amplification or sequencing errors to be detected or corrected. In some embodiments, the length of the nucleic acid subsequences used to create the error-correcting code may vary, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50, or about these numbers, or a number or range between any two of these values, of nucleotides in length. In some embodiments, nucleic acid subsequences of other lengths may be used to create error-correcting codes.
[0079] CRISPR-associated proteins The programmable DNA binding unit can include a nuclease-deficient CRISPR-associated protein (dCAS protein) and a guide RNA (gRNA) that can specifically bind to the binding site of the target dsDNA.The dCAS protein can be dCAS9, dCAS12, dCAS13, dCAS14, or SpRY dCAS.The dCAS13 protein can be dCAS13a, dCAS13b, dCAS13c, or dCAS13d.
[0080] In some embodiments, the Cas9 protein has an inactive (e.g., inactivated) DNA cleavage domain. Nuclease-inactivated Cas9 proteins can be interchangeably referred to as "dCas9" proteins (for nuclease-dead Cas9). Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science.337:816-821 (2012); Qi et al., Cell.28; 152(5): 1173-83 (2013), the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand that is complementary to the gRNA, while the RuvCl subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., and Qi et al.).
[0081] The programmable DNA binding unit can include a suitable nuclease-deficient Cas protein that can still bind to the guide RNA. The programmable DNA binding unit can include a class 2 type II Cas protein. The class 2 type II Cas protein can be a mutated Cas protein compared to its wild-type counterpart. The mutated Cas protein can be nuclease-deficient. The mutated Cas protein can be a mutated Cas9. The mutated Cas9 can be Cas9D10A. Other examples of mutations in Cas9 include H820A, D839A, H840A, N863A, or any combination thereof, such as D10A / H820A, D10A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A. The mutations described herein refer to SpCas9 and also include similar mutations in CRISPR proteins other than SpCas9. The programmable DNA binding units include Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb 1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, C2c1, C2c3, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, derivatives thereof, or any combination thereof. Cas9 molecules of various species can be used in the methods and compositions described herein. S. pyogenes and S. aureus (S.aureus Cas9 molecules are the subject of most of the disclosure herein, however, Cas9 molecules of, derived from, or based on the Cas9 proteins of other species listed herein can be used as well. These include, for example, Acidovorax avenae, Actinobacillus pleuropneumoniae, Actinobacillus succinogenes, Actinobacillus suis, Actinomyces sp., cycliphilus denitrificans, Aminomonas paucivorans, Bacillus cereus, Bacillus smithii, Bacillus thuringiensis, Bacteroides sp. sp., Blastopirellula marina, Bradyrhizobium sp., Brevibacillus laterosporus, Campylobacter coli, Campylobacter jejuni, Campylobacter lari, Candidatus Puniceispirillum, Clostridium cellulolyticum, Clostridium perfringens, Corynebacterium accolens, Corynebacterium diphtheriae diphtheria, Corynebacterium maturkotimatruchotii, Dinoroseobacter shibae, Eubacterium dolichum, gamma proteobacterium, Gluconacetobacter diazotrophicus, Haemophilus parainfluenzae, Haemophilus sputorum, Helicobacter canadensis, Helicobacter cinaedi, Helicobacter mustelae, Ilyobacter polytropus, Kingella kingae, Lactobacillus crispatus crispatus, Listeria ivanovii, Listeria monocytogenes, Listeria bacterium, Methylocystis sp., Methylosinus trichosporium, Mobiluncus mulieris, Neisseria bacilliformis, Neisseria cinerea, Neisseria flavescens, Neisseria lactamica, Neisseria meningitidis, Neisseria sp., Neisseria waswolchii wadsworthii, Nitrosomonas sp., Parvibaculum lavamentivoranslavamentivorans, Pasteurella multocida, Phascolarctobacterium succinatutens, Ralstonia syzygii, Rhodopseudomonas palustris, Rhodovulum sp., Simonsiella muelleri, Sphingomonas sp., Sporolactobacillus vineae, Staphylococcus lugdunensis, Streptococcus sp., Subdoligranulum sp. Examples of suitable Cas9 molecules include those derived from T. sp., Tistrella mobilis, Treponema sp., or Verminephrobacter eiseniae. Catalytically inactivating mutations and means for assessing the nuclease activity of such mutants are known to those of skill in the art.
[0082] Cas proteins may require recognition of short sequence motifs adjacent to the target site, such as protospacer adjacent motifs (PAMs). This requirement may adversely limit target site recognition to a subset of sequences. In some embodiments provided herein, the Cas protein is engineered to reduce or eliminate the PAM recognition requirement. In some embodiments of the compositions and methods disclosed herein, the programmable DNA binding unit comprises a quasi-PAMless SpCas9 variant named SpRY, or a variant or derivative thereof. Engineering of quasi-PAMless Cas9 variants is described in Walton et al. ("Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants." Science 368.6488 (2020): 290-296), the contents of which are incorporated herein by reference in their entirety.
[0083] The programmable DNA binding unit can include a guide molecule. A guide RNA molecule (sgRNA or gRNA) can be composed of two separate molecules: a target-specific crRNA and a tracrRNA that binds to a Cas molecule. In some embodiments, the crRNA and tracrRNA are provided as separate molecules, one that must anneal to become a functional sgRNA. As used herein, the terms "guide sequence" and "guide molecule" in the context of the CRISPR-Cas system include any polynucleotide sequence that has sufficient complementarity with a selected binding site to hybridize with the selected binding site, and the direct sequence-specific binding of the programmable DNA binding unit to the selected binding site. A gRNA molecule can refer to a nucleic acid that facilitates specific targeting or homing of the gRNA molecule / Cas9 molecule complex to a target binding site. A gRNA molecule can be unimolecular (having a single RNA molecule) (e.g., chimeric) or modular (comprising more than one, typically two separate RNA molecules). The guide sequences generated using the methods disclosed herein can be full length guide sequences, truncated guide sequences, full length sgRNA sequences, truncated sgRNA sequences, or E+F sgRNA sequences. In some embodiments, the degree of complementarity of a guide sequence to a given binding site, when optimally aligned using a suitable alignment algorithm, is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more. In certain exemplary embodiments, the guide molecule comprises a guide sequence that can be designed to have at least one mismatch with the binding site such that an RNA duplex forms between the guide sequence and the binding site. Thus, the degree of complementarity is preferably less than 99%. For example, if the guide sequence consists of 24 nucleotides, the degree of complementarity is more specifically about 96% or less. In certain embodiments, the guide sequence is designed to have a stretch of two or more adjacent mismatched nucleotides such that the degree of complementarity over the entire guide sequence is further reduced.For example, if the guide sequence consists of 24 nucleotides, the degree of complementarity is more particularly about 96% or less, more particularly about 92% or less, more particularly about 88% or less, more particularly about 84% or less, more particularly about 80% or less, more particularly about 76% or less, more particularly about 72% or less, depending on whether the stretch of two or more mismatched nucleotides comprises 2, 3, 4, 5, 6 or 7 nucleotides, etc. In some embodiments, the degree of complementarity when optimally aligned using a suitable alignment algorithm, excluding the stretch of one or more mismatched nucleotides, is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), Clustal W, Clustal X, Clustal Omega, BLAT, Novoalign (Novocraft Technologies; available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of the guide sequence (within the nucleic acid-targeting guide RNA) to direct sequence-specific binding of the programmable DNA binding unit to a selected binding site can be assessed by any suitable assay. In some embodiments, the guide sequence is an RNA sequence of 10-50 nt in length, but more particularly about 20-30 nt, advantageously about 20 nt, 23-25 nt or 24 nt in length. The guide sequence can be selected to ensure that it will hybridize to a selected binding site.
[0084] Killing guide sequence The programmable DNA binding unit can include a CRISPR-associated protein (CAS protein) and a guide RNA (gRNA or sgRNA) that can specifically bind to a binding site of a target dsDNA. In some embodiments, the guide sequence is modified in a manner that allows the formation of a CRISPR Cas complex and successful binding to the binding site while at the same time not allowing successful nuclease activity. Such modified guide sequences are referred to as "dead guides" or "dead guide sequences". These dead guides or dead guide sequences can be considered catalytically inactive or conformationally inactive with respect to nuclease activity. The programmable DNA binding unit can include a functional Cas protein and a guide RNA (gRNA) or crRNA, where the gRNA or crRNA includes a dead guide sequence, such that the gRNA can hybridize to a selected binding site such that the Cas protein is directed to the selected binding site without detectable cleavage activity of the non-mutant Cas protein. The ability of the dead guide sequence to direct sequence-specific binding to the binding site of the CRISPR complex can be evaluated by any suitable assay.The dead guide sequence can typically be shorter than the respective guide sequence that causes active cleavage.In certain embodiments, the dead guide is 5%, 10%, 20%, 30%, 40% or 50% shorter than the respective guide that is directed to the same binding site.
[0085] Protein Components The programmable DNA binding unit can include a protein component that can specifically bind to a binding site on a target dsDNA. The protein component can include an endonuclease-deficient zinc finger nuclease (ZFN), an endonuclease-deficient transcription activator-like effector nuclease (TALEN), an Argonaute protein, an endonuclease-deficient meganuclease, a recombinase, or a combination thereof. In some embodiments, the programmable DNA binding unit does not have a nuclease domain. In some embodiments, the programmable DNA binding unit has a nuclease domain that is made catalytically inactive through one or more mutations. Catalytically inactivating mutations and means for evaluating the nuclease activity of the mutants are known to those skilled in the art.
[0086] Transcription activator-like effector (TALE) The programmable DNA binding unit can comprise an endonuclease-deficient transcription activator-like effector nuclease (TALEN), its functional fragment, or its variant. Transcription activator-like effector (TALE) can be engineered to bind to virtually any desired DNA sequence. Exemplary methods of targeting using the TALEN system can be found, for example, in Cermak T. Doyle EL. Christian et al. Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic Acids Res. 2011;39:e82; Zhang et al. Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nat Biotechnol. 2011;29:149-153, U.S. Patent No. 8,450,471, U.S. Patent No. 8,440,431, and U.S. Patent No. 8,440,432, all of which are specifically incorporated by reference.
[0087] The programmable DNA binding unit can comprise a TALE polypeptide. TALEs are transcription factors from the plant pathogen Xanthomonas and can be easily engineered to bind new DNA targets. In some embodiments provided herein, the TALE is not linked to a catalytic domain of an endonuclease (e.g., Fokl). In some embodiments provided herein, the programmable DNA binding unit can comprise a TALEN whose endonuclease domain is catalytically inactive. TALE polypeptides comprise a nucleic acid binding domain composed of tandem repeats of highly conserved monomeric polypeptides, predominantly 33, 34 or 35 amino acids in length, differing from each other predominantly at amino acid positions 12 and 13. As used herein, the term "polypeptide monomer" or "TALE monomer" is used to refer to the highly conserved repeated polypeptide sequence within the TALE nucleic acid binding domain, and the term "repeated variable diresidue" or "RVD" is used to refer to the highly variable amino acids at positions 12 and 13 of the polypeptide monomer. A TALE monomer has a nucleotide binding affinity determined by the identity of the amino acids in its RVD. For example, a polypeptide monomer with an RVD of NI preferentially binds adenine (A), a polypeptide monomer with an RVD of NG preferentially binds thymine (T), a polypeptide monomer with an RVD of HD preferentially binds cytosine (C), and a polypeptide monomer with an RVD of NN preferentially binds both adenine (A) and guanine (G). In yet another embodiment provided herein, a polypeptide monomer with an RVD of IG preferentially binds T. Thus, the number and order of polypeptide monomer repeats in the nucleic acid binding domain of a TALE determines its nucleic acid target specificity. In yet a further embodiment provided herein, a polypeptide monomer with an RVD of NS can recognize all four base pairs and bind A, T, G, or C.The structure and function of TALEs are further described, for example, in Moscou et al., Science 326:1501 (2009); Boch et al., Science 326:1509-1512 (2009); Zhang et al., Nature Biotechnology 29: 149-153 (2011), each of which is incorporated by reference in its entirety. Programmable DNA-binding units can include polypeptide monomer repeats designed to target specific nucleic acid sequences.
[0088] As described in Zhang et al., Nature Biotechnology 29:149-153 (2011), TALE polypeptide binding efficiency can be increased by including an amino acid sequence from a "capping region" that is directly N- or C-terminal to the DNA-binding region of a naturally occurring TALE into the engineered TALE at a position N- or C-terminal to the engineered TALE DNA-binding region. Thus, in some embodiments, the TALE polypeptides described herein further comprise an N-terminal capping region and / or a C-terminal capping region. As used herein, the DNA binding domain comprising a repeating TALE monomer and a C-terminal capping region, in a given "N-terminus" to "C-terminus" orientation of the N-terminal capping region, provides the structural basis for the organization of different domains in the TALEs or polypeptides provided herein. The entire N- and / or C-terminal capping region is not required to enhance the binding activity of the DNA-binding region. Thus, in some embodiments, fragments of the N- and / or C-terminal capping region are included in the TALE polypeptides described herein.
[0089] In some embodiments, the TALE polypeptide comprises an N-terminal capping region fragment comprising at least 10, 20, 30, 40, 50, 54, 60, 70, 80, 87, 90, 94, 100, 102, 110, 117, 120, 130, 140, 147, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, or 270 amino acids of the N-terminal capping region. In some embodiments, the N-terminal capping region fragment amino acids are at the C-terminus (DNA-binding region proximal end) of the N-terminal capping region. As described in Zhang et al., Nature Biotechnology 29:149-153 (2011), an N-terminal capping region fragment containing the C-terminal 240 amino acids has enhanced binding activity equivalent to the full-length capping region, while a fragment containing the C-terminal 147 amino acids retains greater than 80% of the effectiveness of the full-length capping region, and a fragment containing the C-terminal 117 amino acids retains greater than 50% of the activity of the full-length capping region.
[0090] In some embodiments, the TALE polypeptide comprises a C-terminal capping region fragment comprising at least 6, 10, 20, 30, 37, 40, 50, 60, 68, 70, 80, 90, 100, 110, 120, 127, 130, 140, 150, 155, 160, 170, 180 amino acids of the C-terminal capping region. In some embodiments, the C-terminal capping region fragment amino acids are at the N-terminus (DNA-binding region proximal end) of the C-terminal capping region. As described in Zhang et al., Nature Biotechnology 29: 149-153 (2011), a C-terminal capping region fragment comprising the C-terminal 68 amino acids enhances binding activity equivalent to the full-length capping region, while a fragment comprising the C-terminal 20 amino acids retains more than 50% of the effectiveness of the full-length capping region.
[0091] Zinc finger (ZF) proteins The programmable DNA binding unit can include a zinc finger (ZF) nuclease, a functional fragment thereof, or a variant thereof. The programmable DNA binding unit can include an endonuclease-deficient ZF nuclease, a functional fragment thereof, or a variant thereof, in which the endonuclease domain (e.g., Fokl) is catalytically inactive or absent. The programmable DNA binding unit can include a ZF protein (ZFP). The ZFP can be engineered to bind to a selected target site. For example, Beerli et al. (2002) Nature Biotechnol. 20: 135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nature Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct Biol. 10:411-416; US Patent Nos. 6,453,242; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,030,215; 6,794,136; See US Patent Application Publication No. 2005 / 0064474; US Patent Application Publication No. 2007 / 0218528; US Patent Application Publication No. 2005 / 0267061. ZFPs can include an array of ZF modules that target desired DNA binding sites. Each finger module of the ZF array can target three DNA bases. Customized arrays of individual zinc finger domains can be assembled into ZFPs.
[0092] Meganuclease The programmable DNA binding unit can be an endonuclease-deficient meganuclease, a functional fragment thereof, or a variant thereof. The DNA binding domain of the meganuclease can have a double-stranded DNA target sequence of 12-45 bp. In some embodiments, the meganuclease is either a dimeric enzyme, with each meganuclease domain on a monomer, or a monomeric enzyme that contains the two domains on a single polypeptide. Not only wild-type meganucleases, but also various meganuclease variants have been generated by protein engineering to cover a myriad of unique sequence combinations. In some embodiments, chimeric meganucleases can be used, with a recognition site composed of a half-site of meganuclease A and a half-site of protein B. Specific examples of such chimeric meganucleases include the protein domains of I-DmoI and I-CreI. Examples of meganucleases include homing endonucleases from the LAGLIDADG family. "LAGLIDADG meganuclease" refers to a homing endonuclease from the LAGLIDADG family or an engineered variant comprising a polypeptide sharing at least 80%, 85%, 90%, 95%, 97.5%, 99% or more identity or similarity with said naturally occurring homing endonuclease. Such engineered LAGLIDADG meganucleases can be derived from monomeric or dimeric meganucleases. When derived from dimeric meganucleases, such engineered LAGLIDADG meganucleases can be single-stranded or dimeric endonucleases. Meganucleases can be targeted to specific sequences by modifying their recognition sequences using techniques well known to those skilled in the art. See, e.g., Epinat et al., 2003, Nuc. Acid Res., 31(ll):2952-62 and Stoddard, 2005, Quarterly Review of Biophysics, pp. 1-47.
[0093] LAGLIDADG meganucleases include I-SceI, I-ChuI, I-CreI, I-CsmI, PI-SceI, PI-TliI, PI-MtuI, I-CeuI, I-SceII, I-SceIII, HO, PI-CivI, PI- CtrI, PI-AaeI, PI-BsuI, PI-DhaI, PI-DraI, PI-MavI, PI-MchI, PI-MfuI, PI-MflI, PI-MgaI, PI-MgoI, PI-MinI, PI-MkaI, PI-MleI, The LAGLIDADG meganuclease may be PI-MmaI, PI-MshI, PI-MsmI, PI-MthI, PI-MtuI, PI-MxeI, PI-NpuI, PI-PfuI, PI-RmaI, PI-SpbI, PI-SspI, PI-FacI, PI-MjaI, PI-PhoI, PI-TagI, PI-Thyl, PI-TkoI, PI-TspI, or I-MsoI; or a functional mutant or variant thereof, whether homodimeric, heterodimeric, or monomeric. In some embodiments, the LAGLIDADG meganuclease is an I-CreI derivative. In some embodiments, the LAGLIDADG meganuclease shares at least 80% similarity with the native I-CreI LAGLIDADG meganuclease. In some embodiments, the LAGLIDADG meganuclease shares at least 80% similarity with residues 1-152 of the native I-CreI LAGLIDADG meganuclease. In some embodiments, the LAGLIDADG meganuclease may consist of two monomers that share at least 80% similarity with residues 1-152 of the native I-CreI LAGLIDADG meganuclease linked together, with or without a linker peptide.
[0094] Argonaute Proteins In some embodiments, the programmable DNA binding unit comprises a nuclease-inactive Argonaute. In some embodiments, the programmable DNA binding unit comprises an Argonaute protein from Natronobacterium gregoryi (NgAgo), a functional fragment thereof, or a variant thereof. NgAgo is a ssDNA-guided endonuclease. NgAgo binds to approximately 24 nucleotide 5' phosphorylated ssDNA (gDNA) to reach the target site and creates a DNA double-strand break at the gDNA site. In some embodiments, the programmable DNA binding unit comprises a nuclease-inactive NgAgo (dNgAgo). The characterization and use of NgAgo is described in Gao et al, Nat Biotechnol. Epub 2016 May 2. PubMed PMID: 27136078; Swarts et al, Nature. 507(7491) (2014):258-61; Swarts et al, Nucleic Acids Res. 43(10) (2015):5120-9, each of which is incorporated herein by reference. The NgAgo-based programmable DNA binding unit can include at least one guide DNA element, or a nucleic acid comprising a nucleic acid sequence encoding a guide DNA element, and can achieve specific targeting or recognition of the binding site via direct base pairing with the DNA of the binding site. Prokaryotic homologs of Argonaute proteins are known and are described, for example, in Makarova K., et al., "Prokaryotic homologs of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile genetic elements", Biol. Direct. 2009 Aug. 25; 4:29. doi: 10.1186 / 1745-6150-4-29, incorporated herein by reference.In some embodiments, the programmable DNA-binding unit is a Marinitog a piezophila Argonaute (MpAgo) protein, a functional fragment thereof, or a variant thereof.
[0095] Recombinase In some embodiments, the programmable DNA binding unit comprises a recombinase that is configured to bind to a binding site on target dsDNA. Site-specific recombinases are well known in the art and can be generally referred to as invertases, resolvases, or integrases. Non-limiting examples of site-specific recombinases include, but are not limited to, lambda integrase, Cre, Int, IHF, Xis, Flp, Fis, Hin, Gin, phiC31, Cin, Tn3 resolvases, TndX, XerC, XerD, TnpX, Hjc, Gin, SpCCEl, and ParA.
[0096] Linker The transposome may be linked to the programmable DNA binding unit through a linker connecting the transposase and the dCAS protein. The linker may include a peptide linker, a chemical linker, or both. The transposase may be present in a fusion protein that includes the dCAS protein. The transposome may be linked to the programmable DNA binding unit through a linker connecting the transposase and the protein component. The peptide linker may include multiple glycines, serine, threonine, alanine, lysine, glutamine, or combinations thereof. The peptide linker may include a GS linker. The peptide linker may be an XTEN linker. The protein component may be present in a fusion protein that includes the transposase. The term "linker" as used herein refers to a molecule that facilitates interactions between molecules or portions of molecules. In one embodiment, the linker is a polypeptide linker. In another embodiment, the linker is a chemical linker. The term "peptide linker" or "polypeptide linker" as used herein refers to a peptide or polypeptide that includes two or more amino acid residues connected by peptide bonds. Such peptide or polypeptide linkers are well known in the art. The linker can comprise naturally occurring and / or non-naturally occurring peptides or polypeptides. The linker can be attached to the C-terminus and / or N-terminus of the transposase and / or the programmable DNA binding unit. The linker can be a chemical linker or a peptide linker. Thus, embodiments relate to polypeptides conjugated to other molecules via peptide bonds, and to polypeptides conjugated to other molecules via chemical conjugation.
[0097] A peptide linker with some degree of flexibility can be used. The peptide linker can have virtually any amino acid sequence, keeping in mind that suitable peptide linkers generally have sequences that result in flexible peptides. The use of small amino acids such as glycine and alanine are used to create flexible peptides. The creation of such sequences is routine for those skilled in the art.
[0098] Suitable linkers can be readily selected and can be of any suitable length, for example, from 1 amino acid (e.g., Gly) to 50 amino acids, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 amino acids, or a number or range between any two of these values (or any derivable range therein).
[0099] Preferred peptide linker sequences adopt flexible extended conformations and do not tend to develop ordered secondary structures. In some embodiments, the linker can be a chemical moiety that can be monomeric, dimeric, multimeric, or polymeric. Preferably, the linker comprises amino acids. Exemplary amino acids in flexible linkers include Gly, Asn, and Ser. Thus, in certain embodiments, the linker comprises one or more combinations of Gly, Asn, and Ser amino acids. Other near-neutral amino acids, such as Thr and Ala, can also be used in the linker sequence. Examples of flexible linkers include glycine polymers (G)n (SEQ ID NO:32), glycine-serine polymers (e.g., (GS)n (SEQ ID NO:33), (GSGGS)n (SEQ ID NO:34), (GS)n (SEQ ID NO:35), and (GGGS)n (SEQ ID NO:36), where n is an integer of at least 1. In some embodiments, n is at least, at most, or exactly 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 (or any derivable range therein). Glycine-alanine polymers, alanine-serine polymers, and other flexible linkers known in the art. Glycine and glycine-serine polymers can be used, where both Gly and Ser are relatively unstructured and therefore neutral between the components. It can act as a tether. Glycine polymers can be used, with glycine accessing significantly more phi-psi space than alanine and being much less restricted than residues with longer side chains. Exemplary spacers can include amino acid sequences including, but not limited to, GGSG (SEQ ID NO: 37), GGSGG (SEQ ID NO: 38), GSGSG (SEQ ID NO: 39), GSGGG (SEQ ID NO: 40), GGGSG (SEQ ID NO: 41), GSSSG (SEQ ID NO: 42), and the like. Other near-neutral amino acids such as Thr and Ala can also be used in the linker sequence. The length of the linker sequence can be varied without significantly affecting the function or activity of the fusion protein (see, e.g., U.S. Patent No. 6,087,329).In some embodiments, the linker can be at least, at most, or exactly 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acid residues (or any range derivable therein).
[0100] In some embodiments, the polypeptide linker is an XTEN linker. In some embodiments, the linker is an XTEN linker or a variation of an XTEN linker, such as SGSETPGTSESA (SEQ ID NO: 43), SGSETPGTSESATPES (SEQ ID NO: 44) or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 45). XTEN linkers are described, for example, in Schellenberger et al. (2009), Nature Biotechnology 27: 1186-1190, the entire contents of which are incorporated herein by reference.
[0101] Suitable linkers for use in the methods provided herein are well known to those skilled in the art and include, but are not limited to, straight or branched carbon linkers, heterocyclic carbon linkers, or peptide linkers. However, as used herein, linkers can also be covalent bonds (carbon-carbon or carbon-heteroatom bonds). In certain embodiments, linkers are used to separate the transposome and the programmable DNA binding unit by a distance sufficient to ensure that each protein retains its required functional properties.
[0102] A linker can be used to fuse two protein partners to form a fusion protein. A "linker" can be a chemical group or molecule that links two molecules or moieties, such as two domains of a fusion protein. Typically, a linker is placed between (adjacent to) two groups, molecules, domains, or other moieties and is connected to each via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer (e.g., a non-natural polymer, a non-peptide polymer), or chemical moiety. In another embodiment, the linker is a direct bond or an atom, such as oxygen (O) or sulfur (S), a unit, such as -NR- (where R is hydrogen or alkyl, -C(O)-, -C(O)O-, -C(O)NH-, SO, SO 2 , -SO 2 NH-), or a chain of atoms, such as substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, arylalkyl, heteroarylalkyl. In some embodiments, one or more methylenes in the chain of atoms can be O, S, S(O), SO 2 , -SO 2 NH-, -NR-, -NR 2, -C(O)-, -C(O)O-, -C(O)NH-, a cleavable linking group, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, and substituted or unsubstituted heterocycle. Exemplary linkers may also include chemical moieties and conjugating agents, such as sulfo-succinimidyl derivatives (sulfo-SMCC, sulfo-SMPB), disuccinimidyl suberate (DSS), disuccinimidyl glutarate (DSG), and disuccinimidyl tartarate (DST). Exemplary linkers further include linear carbon chains, such as CN (where N=1-100 carbon atoms). In some embodiments, the linker may be a dipeptide linker, such as a valine-citrulline (val-cit), a phenylalanine-lysine (phe-lys) linker, or a maleimidocaproic-valine-citrulline-p-aminobenzyloxycarbonyl (vc) linker. In some embodiments, the linker is sulfosuccinimidyl-4-[N-maleimidomethyl]cyclohexane-1-carboxylate (smcc). Sulfo-smcc conjugation occurs via a maleimide group that reacts with sulfhydryls (thiols, -SH), while its sulfo-NHS ester is reactive towards primary amines (found in lysines and protein or peptide N-termini). Additionally, the linker can be maleimidocaproyl (me). In some embodiments, covalent linkage can be achieved through the use of Traut's reagent.
[0103] amplification The methods provided herein can include amplifying a plurality of dsDNA fragments with a primer capable of binding to one strand of an adaptor to generate an amplification product. The amplification can generate an amplification product. The primer can bind to all or a portion of the adaptor strand. The primer can include a 5' overhang (e.g., a sequence that does not hybridize to the adaptor and / or the dsDNA fragment). In some embodiments, amplifying a plurality of dsDNA fragments does not use any primer other than the primer capable of binding to one strand of the adaptor. In some embodiments, the amplifying step includes the use of a single primer. In some embodiments, the amplifying step includes the use of a single primer pair. The primers provided herein can be about 5-80 nucleotides in length (e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80 nucleotides in length, or a number or range between any two of these values). Amplifying multiple dsDNA fragments with primers can be performed using PCR. PCR can be loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinase polymerase amplification (RPA), strand displacement amplification (SDA), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple displacement amplification (MDA), branching (RAM), circular helicase-dependent amplification (cHDA), single primer isothermal amplification (SPIA), signal-mediated amplification of RNA technology (SMART), self-sustained sequence replication (3SR), genomic exponential amplification reaction (GEAR), ligase chain reaction (LCR), self-sustained sequence replication (3SR), rolling circle amplification, transcription-mediated amplification (TMA), or isothermal multiple displacement amplification (IMDA). PCR can be real-time PCR or quantitative real-time PCR (QRT-PCR).
[0104] For example, LCR amplification uses at least four separate oligonucleotides to amplify a target and its complementary strand by using multiple cycles of hybridization, ligation, and denaturation. SDA amplifies by using a primer that contains a recognition site for a restriction endonuclease that nicks one strand of a hemi-modified DNA duplex containing the target sequence, followed by amplification in a series of primer extension and strand displacement steps.
[0105] PCR is a method well known in the art for the amplification of nucleic acids. PCR involves the amplification of a target sequence using two or more extendible sequence-specific oligonucleotide primers that flank the target sequence. A nucleic acid containing the target sequence of interest is subjected to a program of multiple rounds of thermal cycling (denaturation, annealing and extension) in the presence of primers, a thermostable DNA polymerase (e.g., Taq polymerase) and various dNTPs, resulting in the amplification of the target sequence. PCR uses multiple rounds of primer extension reactions in which complementary strands of defined regions of a DNA molecule are simultaneously synthesized by a thermostable DNA polymerase. At the end of each cycle, each newly synthesized DNA molecule serves as a template for the next cycle. During these repeated rounds of reactions, the number of newly synthesized DNA strands increases exponentially, so that after 20-30 reaction cycles, the initial template DNA is replicated thousands or millions of times. PCR can generate double-stranded amplification products suitable for post-amplification processing. If desired, the amplification products can be detected by visualization via agarose gel electrophoresis, enzyme immunoassay formats using probe-based colorimetric detection, fluorescence techniques, or other detection means known in the art.
[0106] Examples of PCR methods include, but are not limited to, real-time PCR, end-point PCR, amplified fragment length polymorphism PCR (AFLP-PCR), Alu-PCR, asymmetric PCR, colony PCR, DD-PCR, degenerate PCR, hot-start PCR, in situ PCR, inverse PCR, long-PCR, multiplex PCR, nested PCR, PCR-ELISA, PCR-RFLP, PCR-single-strand conformation polymorphism (PCR-SSCP), quantitative competitive PCR (QC-PCR), rapid amplification of cDNA ends-PCR (RACE-PCR), random amplification of polymorphic DNA-PCR (RAPD-PCR), real-time PCR, repetitive extragenic palindrome PCR (Rep-PCR), reverse transcriptase PCR (RT-PCR), TAIL-PCR, touchdown PCR, and Vectorette PCR.
[0107] Real-time PCR, also called quantitative real-time polymerase chain reaction (QRT-PCR), can be used to simultaneously quantify and amplify specific portions of a given nucleic acid molecule. It can be used to determine whether a particular sequence is present in a sample, and if so, the number of copies of the sequence present. The term "real-time" can refer to periodic monitoring during PCR. Certain systems, such as the ABI 7700 and 7900HT Sequence Detection Systems (Applied Biosystems, Foster City, Calif.), perform monitoring during each thermal cycle at pre-determined or user-defined points. Real-time analysis of PCR using fluorescence resonance energy transfer (FRET) probes measures the change in fluorescent dye signal from cycle to cycle, preferably minus any internal control signal. Real-time techniques follow the general pattern of PCR, but the nucleic acid is quantified after each round of amplification. Two examples of methods of quantification are the use of fluorescent dyes (e.g., SYBRGreen) that intercalate into double-stranded DNA, and modified DNA oligonucleotide probes that fluoresce when hybridized with complementary DNA. The intercalating agent has a relatively low fluorescence when unbound and a relatively high fluorescence when bound to double-stranded nucleic acid. Therefore, the intercalating agent can be used to monitor the accumulation of double-stranded nucleic acid during a nucleic acid amplification reaction. Examples of such non-specific dyes useful in the embodiments disclosed herein include intercalating agents such as SYBR Green I (Molecular Probes), propidium iodide, and ethidium bromide.
[0108] 5-7 show non-limiting and exemplary schematic diagrams of plasmid constructs 3XFlag-Cas9-Fl26-Tn5 (SEQ ID NO: 1), 3XFlag-Cas9-xTen-Tn5 (SEQ ID NO: 2), and pET-Tn5-xTen-dCas9 (SEQ ID NO: 3), respectively, for use in generating protein complexes provided herein. The protein complexes, linkers, programmable DNA binding units, and / or transposases disclosed herein are at least about 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 102%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, In some embodiments, the nucleic acid sequence may be encoded by a nucleotide sequence that is 6%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% identical, or a number or range between any two of these values.
[0109] Oligonucleotide Probes Detecting the presence of a target sequence in the amplified products may include contacting the amplified products with oligonucleotide probes, each capable of specifically binding to a target sequence. The oligonucleotide probes may, in some embodiments, comprise a detectable moiety. For example, the oligonucleotide probes disclosed herein may comprise a radioactive label. Non-limiting examples of radioactive labels include: 3 H, 14 C. 32 P, and 35In some embodiments, the oligonucleotide probe can include one or more non-radioactive detectable markers or moieties, including but not limited to ligands, fluorophores, chemiluminescent agents, enzymes, and antibodies. Other detectable markers for use with the probe that can increase the sensitivity of the method of the present invention include biotin and radioactive nucleotides. It will be clear to one of skill in the art that the choice of a particular label will determine the manner in which it is attached to the probe. For example, the oligonucleotide probe is labeled with one or more dyes such that a detectable change in fluorescence is produced upon hybridization to the template nucleic acid. While non-specific dyes may be desirable for some applications, sequence-specific probes can provide a more accurate measurement of amplification. One configuration of a sequence-specific probe can include one end of the probe tethered to a fluorophore and the other end of the probe tethered to a quencher. When the probe is not hybridized, the probe can maintain a stem-loop structure, and the fluorophore is quenched by the quencher, thus preventing the fluorophore from fluorescing. When the probe is hybridized to the template nucleic acid sequence, the probe is linearized, moving the fluorophore away from the quencher, and thus causing the fluorophore to fluoresce. Another configuration of sequence-specific probes can include a first probe tethered to a first fluorophore of a FRET pair and a second probe tethered to a second fluorophore of a FRET pair. The first and second probes can be configured to hybridize to sequences of the amplicon that are in sufficient proximity to allow energy transfer by FRET when the first and second probes hybridize to the same amplicon.
[0110] In some embodiments, the probe is a TaqMan probe. The TaqMan probe can include a fluorophore and a quencher. The quencher molecule can quench the fluorescence emitted by the fluorophore when excited via Förster resonance energy transfer (FRET) by the light source of the cycler. As long as the fluorophore and the quencher are in close proximity, the quenching can inhibit any detectable (e.g., fluorescent) signal. The TaqMan probes provided herein can be designed such that they anneal within the DNA region amplified by the primers provided herein. Without being bound to a particular theory, in some embodiments, as a PCR polymerase (e.g., Taq) extends the primer and synthesizes a nascent strand on the single-stranded template, the 5' to 3' exonuclease activity of the PCR polymerase degrades the probe annealed to the template. Degradation of the probe can then release the fluorophore and cleave it in close proximity to the quencher, thereby relieving the quenching effect and allowing the fluorophore to fluoresce. Thus, the fluorescence detected in a quantitative PCR thermal cycler can, in some embodiments, be directly proportional to the amount of released fluorophore and DNA template present in the PCR.
[0111] In some embodiments, the sequence-specific probe comprises an oligonucleotide disclosed herein conjugated to a fluorophore. In some embodiments, the probe is conjugated to two or more fluorophores. Examples of fluorophores include xanthene dyes, such as fluorescein and rhodamine dyes, such as fluorescein isothiocyanate (FITC), 2-[ethylamino]-3-(ethylimino)-2-7-dimethyl-3H-xanthen-9-yl]benzoic acid ethyl ester monohydrochloride (R6G) (emitting response radiation at a wavelength in the range of about 500-560 nm), 1,1,3,3,3',3'-hexafluorobenzoate (EtOH ... Methylindodicarbocyanine iodide (HIDC) (emitting responsive radiation at wavelengths in the range of about 600-660 nm), 6-carboxyfluorescein (commonly known by the abbreviations FAM and F), 6-carboxy-2',4',7',4,7-hexachlorofluorescein (HEX), 6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein (JOE or J), N,N,N',N '-Tetramethyl-6-carboxyrhodamine (TAMRA or T), 6-carboxy-X-rhodamine (ROX or R), 5-carboxyrhodamine-6G (R6G5 or G5), 6-carboxyrhodamine-6G (R6G6 or G6), and rhodamine 110; cyanine dyes, such as Cy3, Cy5 and Cy7 dyes; coumarins, such as umbelliferone; benzimide dyes, such as Hoechst 33258; phenanthridine dyes, such as Texas Red; ethidium dyes; acridine dyes; carbazole dyes; phenoxazine dyes; porphyrin dyes; polymethine dyes, such as cyanine dyes, such as Cy3 (emitting responsive radiation at a wavelength that is in the range of about 540-580 nm); Cy5 (emitting responsive radiation at a wavelength that is in the range of about 640-680 nm), etc; BODIPY dyes and quinoline dyes.Specific fluorophores of interest include pyrene, coumarin, diethylaminocoumarin, FAM, fluorescein chlorotriazinyl, fluorescein, R110, eosin, JOE, R6G, HIDC, tetramethylrhodamine, TAMRA, Lissamine, ROX, naphthofluorescein, Texas Red, naphthofluorescein, Cy3, and Cy5, CAL Fluor Orange, and the like. Other examples of fluorescein dyes include 6-carboxyfluorescein (6-FAM), 2',4',1,4-tetrachlorofluorescein (TET), 2',4',5',7',1,4-hexachlorofluorescein (HEX), 2',7'-dimethoxy-4',5'-dichloro-6-carboxyrhodamine (JOE), 2'-chloro-5'-fluoro-7',8'-fused phenyl-1,4-dichloro-6-carboxyfluorescein (NED), and 2'-chloro-7'-phenyl-1,4-dichloro-6-carboxyfluorescein (VIC). The probe can include SpC6, or functional equivalents and derivatives thereof. The probe can include a spacer moiety. The spacer moiety can include an alkyl group of at least 2 carbons to about 12 carbons. The probe can include a spacer that includes a non-basic unit. The probe may comprise a spacer selected from the group including idSp, iSp9, iS18, iSpC3, iSpC6, iSpC12, or any combination thereof.
[0112] In some embodiments, the probe is conjugated to a quencher. The quencher can absorb electromagnetic radiation and dissipate it as heat, thus retaining a dark color. Exemplary quenchers include Dabcyl, NFQ, such as BHQ-1 or BHQ-2 (Biosearch), IOWA BLACK FQ (IDT), and IOWA BLACK RQ (IDT). In some embodiments, the quencher is selected to pair with the fluorophore so as to absorb the electromagnetic radiation emitted by the fluorophore. Fluorophore / quencher pairs useful in the compositions and methods disclosed herein are well known in the art and can be found, for example, in Marras, "Selection of Fluorophore and Quencher Pairs for Fluorescent Nucleic Acid Hybridization Probes," available at www.molecular-beacons.org / download / marras,mmb06%28335%293.pdf. Examples of quencher moieties include, but are not limited to, dark quenchers, Black Hole Quenchers® (BHQ®) (e.g., BHQ-0, BHQ-1, BHQ-2, BHQ-3), Qxl quenchers, ATTO quenchers (e.g., ATTO 540Q, ATTO 580Q, and ATTO 612Q), dimethylaminoazobenzenesulfonic acid (Dabsyl), IOWA Black RQ, IOWA Black FQ, IRDye QC-1, QSY dyes (e.g., QSY 7, QSY 9, QSY 21), AbsoluteQuencher, Eclipse, and metal clusters such as gold nanoparticles. Examples of ATTO quenchers include, but are not limited to, ATTO 540Q, ATTO 580Q, and ATTO 612Q. Examples of Black Hole Quenchers® (BHQ®) include, but are not limited to, BHQ-0 (493 nm), BHQ-1 (534 nm), BHQ-2 (579 nm) and BHQ-3 (672 nm).
[0113] In some embodiments, the detectable label is an Alexa Fluor® dye (e.g., Alexa Fluor® 350, Alexa Fluor® 405, Alexa Fluor® 430, Alexa Fluor® 488, Alexa Fluor® 500, Alexa Fluor® 514, Alexa Fluor® 532, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 610, Alexa Fluor® 633, Alexa Fluor® 635, Alexa Fluor® 647, Alexa Fluor® 660, Alexa Fluor® 680, Alexa Fluor® 700, Alexa Fluor® 750, Alexa Fluor® 790), an ATTO dye (e.g., ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 514, ATTO 520, ATTO 532, ATTO Rho6G, ATTO 542, ATTO 550, ATTO 565, ATTO Rho3B, ATTO Rhol l, ATTO Rhol2, ATTO Thiol 2, ATTO RholOl, ATTO 590, ATTO 594, ATTO Rhol3, ATTO 610, ATTO 620, ATTO Rhol4, ATTO 633, ATTO 647, ATTO 647N, ATTO 655, ATTO Oxal2, ATTO 665, ATTO 680, ATTO 700, ATTO 725, ATTO 740), DyFight dyes, cyanine dyes (e.g., Cy2, Cy3, Cy3.5, Cy3b, Cy5, Cy5.5, Cy7, Cy7.5) is a fluorescent label selected from FluoProbes dyes, Sulfo Cy dyes, Seta dyes, IRIS dyes, SeTau dyes, SRfluor dyes, Square dyes, fluorescein (FITC), tetramethylrhodamine (TRITC), Texas Red, Oregon Green, Pacific Blue, Pacific Green, Pacific Orange, quantum dots, and tethered fluorescent proteins.
[0114] In some embodiments, the fluorophore is attached to the first end of the probe and the quencher is attached to the second end of the probe. In some embodiments, the probe can include two or more fluorophores. In some embodiments, the probe can include two or more quencher moieties. In some embodiments, the probe can include one or more quencher moieties and / or one or more fluorophores. The quencher moiety or fluorophore can be attached to any part of the probe (e.g., the 5' end, the 3' end, the middle of the probe). Any probe nucleotide can include a fluorophore or quencher moiety, such as, for example, BHQ1dT. The attachment can include a covalent bond and optionally include at least one linker molecule disposed between the probe and the fluorophore or quencher. In some embodiments, the fluorophore is attached to the 5' end of the probe and the quencher is attached to the 3' end of the probe. In some embodiments, the fluorophore is attached to the 3' end of the probe and the quencher is attached to the 5' end of the probe. Examples of probes that can be used for quantitative nucleic acid amplification include molecular beacons, SCORPION™ probes (Sigma), TAQMAN™ probes (Life Technologies), etc. Other nucleic acid detection technologies useful in embodiments disclosed herein include, but are not limited to, nanoparticle probe technology (see Elghanian, et al. (1997) Science 277:1078-1081) and Amplifluor probe technology (see U.S. Patent Nos. 5,866,366; 6,090,592; 6,117,635; and 6,117,986).
[0115] sign The method can include labeling one or both ends of one or more of the plurality of dsDNA fragments (e.g., with a detectable label). The method can include differentially labeling two ends of one or more of the plurality of dsDNA fragments. The labeling can include labeling with a detectable label (e.g., an anionic label, a cationic label, a neutral label, an electrochemical label, a protein label, a fluorescent label, a magnetic label, or a combination thereof). The method can include concentrating the labeled dsDNA fragments, capturing the labeled dsDNA fragments, isolating the labeled dsDNA fragments, and / or visualizing the labeled dsDNA fragments. The method can include monitoring the detectable label (e.g., chemical monitoring).
[0116] In some embodiments, the detectable moiety (e.g., detectable label) comprises an optical moiety, a luminescent moiety, an electrochemically active moiety, a nanoparticle, or a combination thereof. In some embodiments, the luminescent moiety comprises a chemiluminescent moiety, an electroluminescent moiety, a photoluminescent moiety, or a combination thereof. In some embodiments, the photoluminescent moiety comprises a fluorescent moiety, a phosphorescent moiety, or a combination thereof. In some embodiments, the fluorescent moiety comprises a fluorescent dye. In some embodiments, the nanoparticle comprises a quantum dot. In some embodiments, the method comprises performing a reaction to convert a detectable moiety precursor to a detectable moiety. In some embodiments, performing a reaction to convert a detectable moiety precursor to a detectable moiety comprises contacting the detectable moiety precursor with a substrate. In some such embodiments, contacting the detectable moiety precursor with the substrate produces a detectable by-product of the reaction between the two molecules.
[0117] Detection and quantification of target sequences in amplification products Some of the methods provided herein include amplifying a plurality of dsDNA fragments to generate a nucleic acid amplification product. The methods described herein may further include detecting and / or quantifying the nucleic acid amplification product, or a product thereof. Detecting the presence of a target sequence in the amplified product may include contacting the amplified product with an oligonucleotide probe, each capable of specifically binding to the target sequence. The amplification product, or a product thereof, may be detected and / or quantified by any suitable detection and / or quantification method, including, for example, any detection or quantification method described herein.Non-limiting examples of detection and / or quantification methods include molecular beacons (e.g., real-time, end-point), lateral flow, fluorescence resonance energy transfer (FRET), fluorescence polarization (FP), surface capture, 5' to 3' exonuclease hydrolysis probes (e.g., TAQMAN), intercalating dyes / binding dyes, absorbance methods (e.g., colorimetric, turbidity), electrophoresis (e.g., gel electrophoresis, capillary electrophoresis), mass spectrometry, nucleic acid sequencing, digital amplification, primer extension methods (e.g., iPLEX™), Affymetrix or similar. Molecular Inverse Probe (MIP) technology, Restriction Fragment Length Polymorphism (RFLP analysis), Allele-specific Oligonucleotide (ASO) analysis, Methylation-specific PCR (MSPCR), Pyrosequencing analysis, Acycloprime analysis, Reverse Dot Blot, GeneChip microarray, Dynamic Allele-specific Hybridization (DASH), Peptide Nucleic Acid (PNA) and Locked Nucleic Acid (LNA) probes, AlphaScreen, SNP Stream, Gene Bit Analysis (GBA), Multiplex Minisequencing, SNaPshot These include GOOD assay, microarray miniseq, arrayed primer extension (APEX), microarray primer extension, Tag array, coded microsphere, template-directed integration (TDI), colorimetric oligonucleotide ligation assay (OLA), sequence-coded OLA, microarray ligation, ligase chain reaction, padlock probe, invader assay, hybridization with at least one probe, hybridization with at least one fluorescently labeled probe, cloning and sequencing, use of hybridization probes and quantitative real-time polymerase chain reaction (QRT-PCR), nanopore sequencing, chips, and combinations thereof. In some embodiments, detecting the nucleic acid amplification product includes the use of real-time detection methods (i.e., the product is detected and / or continuously monitored during the amplification process). In some embodiments, detecting the nucleic acid amplification product includes the use of end-point detection methods (i.e., the product is detected after completing or stopping the amplification process).Nucleic acid detection methods can also employ the use of labeled nucleotides, either directly incorporated into the target sequence or incorporated into a probe containing a sequence complementary to the target. Such labels can be radioactive and / or fluorescent in nature and can be resolved by any method discussed herein. In some embodiments, quantification of nucleic acid amplification products can be achieved using one or more detection methods described below. In some embodiments, detection methods can be used in conjunction with measuring signal intensity and / or generating (or referencing) standard curves and / or look-up tables for quantification of nucleic acid amplification products.
[0118] The detection of nucleic acid amplification products can include the use of molecular beacon technology. The term molecular beacon generally refers to a detectable molecule, where the detectable property of the molecule is detectable under certain conditions, thereby allowing the molecule to function as a specific and informative signal. Non-limiting examples of detectable properties include optical properties (e.g., fluorescence), electrical properties, magnetic properties, chemical properties, and time or speed through an aperture of known size. Molecular beacons for detecting nucleic acid molecules can be, for example, hairpin-shaped oligonucleotides that contain a fluorophore at one end and a quenching dye at the opposite end. The loop of the hairpin can contain a probe sequence that is complementary to the target sequence, and the stem is formed by annealing of complementary arm sequences located on either side of the probe sequence. The fluorophore and quenching molecules can be covalently linked at both ends of each arm. Under conditions that prevent the oligonucleotide from hybridizing to its complementary target, or when the molecular beacon is free in solution, the fluorescent and quenching molecules are in close proximity to each other, preventing FRET. When a molecular beacon encounters a target molecule (e.g., a nucleic acid amplification product), hybridization can occur, converting the loop structure into a stable, more rigid conformation, causing separation of the fluorophore and quencher molecules, leading to fluorescence. Due to the specificity of the probe, the generation of fluorescence is generally exclusively due to the synthesis of the intended amplified product. In some embodiments, the molecular beacon probe sequence hybridizes to a sequence in the amplification product that is identical or complementary to a sequence in the target nucleic acid. In some embodiments, the molecular beacon probe sequence hybridizes to a sequence in the amplification product that is not identical or complementary to a sequence in the target nucleic acid (e.g., hybridizes to a tail amplification primer or a sequence added to the amplification product by ligation). Molecular beacons can also be synthesized with different colored fluorophores and different target sequences, allowing for the simultaneous detection of several products in the same reaction (e.g., a multiplex reaction).In a quantitative amplification process, molecular beacons can specifically bind to the amplified target after each cycle of amplification, and since unhybridized molecular beacons are dark, it is not necessary to isolate the probe-target hybrid to quantitatively determine the amount of amplified product. The signal obtained is proportional to the amount of amplified product. Detection using molecular beacons can be performed in real time or as an end-point detection method.
[0119] The detection of nucleic acid amplification products can include the use of lateral flow. Lateral flow devices can generally include a solid-phase fluid-permeable channel through which fluid flows by capillary forces. Exemplary devices include, but are not limited to, dipstick assays and thin-layer chromatography plates with various suitable coatings. Immobilized on the channel are various binding reagents for the sample, binding partners or conjugates including binding partners for the sample and signal generation systems. Detection can be accomplished in several ways, including, for example, enzyme detection, nanoparticle detection, colorimetric detection, and fluorescent detection.
[0120] In some embodiments, detecting nucleic acid amplification products involves the use of FRET, an energy transfer mechanism between two chromophores: donor and acceptor molecules. Briefly, a donor fluorophore molecule is excited with a specific excitation wavelength. The excitation energy released from the donor molecule upon returning to the ground state can then be transferred to the acceptor molecule via long-range dipole-dipole interactions. The emission intensity of the acceptor molecule can be monitored and is a function of the distance between the donor and acceptor, the overlap of the donor emission spectrum with the acceptor absorption spectrum, and the orientation of the donor emission dipole moment and the acceptor absorption dipole moment. FRET can be useful to quantify molecular dynamics, for example, in DNA-DNA interactions as described for molecular beacons. To monitor the generation of a specific product, a probe can be labeled with a donor molecule at one end and an acceptor molecule at the other. Probe-target hybridization results in a change in the distance or orientation of the donor and acceptor, and a FRET change is observed.
[0121] In some embodiments, detecting nucleic acid amplification products generally involves the use of FP technology, which is based on the principle that when excited by linearly polarized light, a fluorescently labeled compound emits fluorescence with a degree of polarization that is inversely proportional to its rotation speed. Thus, when a molecule such as a fluorescently labeled tracer-nucleic acid conjugate is excited by linearly polarized light, the emitted light remains highly polarized because the fluorophore is restricted from rotating during the time that light is absorbed and emitted. When a free tracer compound (i.e., not bound to a nucleic acid) is excited by linearly polarized light, its rotation is much faster than that of the corresponding tracer-nucleic acid conjugate, the molecule is more randomly oriented, and the emitted light is therefore depolarized. Thus, fluorescence polarization provides a quantitative means for measuring the amount of tracer-nucleic acid conjugate produced in an amplification reaction. In some embodiments, detection of nucleic acid amplification products involves the use of surface capture, which can be achieved by immobilizing specific oligonucleotides on a surface to generate highly sensitive and selective biosensors. In some embodiments, the detection of nucleic acid amplification products involves the use of 5' to 3' exonuclease hydrolysis probes (e.g., TAQMAN). For example, TAQMAN probes are hydrolysis probes that can increase the specificity of quantitative amplification methods (e.g., quantitative PCR). The TAQMAN probe principle relies on 1) the 5' to 3' exonuclease activity of Taq polymerase to cleave a dual-labeled probe during hybridization to a complementary target sequence, and 2) fluorophore-based detection. The resulting fluorescent signal allows for quantitative measurement of the accumulation of amplification products during the exponential phase of amplification.
[0122] In some embodiments, detecting the nucleic acid amplification product includes the use of an intercalating dye and / or a binding dye. In some embodiments, detecting the nucleic acid amplification product includes the use of a dye that specifically stains nucleic acid. For example, an intercalating dye exhibits enhanced fluorescence when bound to DNA or RNA. The dye may include DNA or RNA intercalating fluorophores, such as SYTO® 82, acridine orange, ethidium bromide, Hoechst dyes, PicoGreen®, propidium iodide, SYBR® I (asymmetric cyanine dye), SYBR® II, TOTO (thiaxol orange dimer) and YOYO (oxazole yellow dimer).
[0123] In some embodiments, detection of nucleic acid amplification products includes the use of absorbance methods (e.g., colorimetry, turbidity). In some embodiments, detection and / or quantification of nucleic acids can be achieved, for example, by directly converting absorbance (e.g., UV absorbance measurements at 260 nm) to concentration. Direct measurements of nucleic acids can be converted to concentration using the Beer Lambert law, which relates absorbance to concentration using the path length and extinction coefficient of the measurement. In some embodiments, detection of the nucleic acid amplification products comprises the use of electrophoresis (e.g., gel electrophoresis, capillary electrophoresis), mass spectrometry, nucleic acid sequencing, digital amplification (e.g., digital PCR), or any combination thereof.
[0124] target nucleic acid The target nucleic acid detected using the compositions and methods provided herein can include a genetic signature of interest, such as a mutation of interest (e.g., a biomarker). The multiple target dsDNAs can include a genetic signature of interest (e.g., a biomarker signature). The genetic signature of interest can include one or more mutations of interest (e.g., a biomarker). The one or more mutations of interest can include a point mutation, an inversion, a deletion, an insertion, a translocation, a duplication, a copy number variation, or a combination thereof. The one or more mutations of interest can include a nucleotide substitution, a deletion, an insertion, or a combination thereof. The genetic signature of interest can be indicative of antibiotic resistance or antibiotic sensitivity of the organism from which the target dsDNA is derived. The genetic signature of interest can be indicative of a cancer state of the organism from which the target dsDNA is derived. The genetic signature of interest can be indicative of a genetic disease state of the organism from which the target dsDNA is derived. The genetic disease can be a single gene disorder. The genetic disease may be cystic fibrosis, Huntington's disease, sickle cell anemia, hemophilia, Duchenne muscular dystrophy, thalassemia, fragile X syndrome, familial hypercholesterolemia, polycystic kidney disease, neurofibromatosis type I, hereditary spherocytosis, Marfan syndrome, Tay-Sachs disease, phenylketonuria, mucopolysaccharidoses, lysosomal acid lipase deficiency, glycogen storage disease, galactosemia, or hemochromatosis. Genetic signatures of interest (e.g., biomarker signatures) can be detected using the methods and compositions provided herein. Diagnostic assessments can be performed using the methods and compositions provided herein.
[0125] The diagnostic assessment is based on a biomarker signature (e.g., a gene signature of interest), alone or in combination with other assessments or factors, as described herein. Provided herein are compositions and methods for assessing the risk of developing a disease or condition, predicting the disease, diagnosing the disease or condition, monitoring the progression or regression of the disease or condition, or assessing the effectiveness of a treatment, or identifying compounds that can ameliorate or treat the disease or condition, based on a biomarker signature (e.g., a gene signature of interest).
[0126] Diseases and Conditions The methods provided herein can be applied to a variety of diseases or conditions based on biomarker signatures (e.g., gene signatures of interest) associated with the various diseases or conditions. Exemplary diseases or conditions having gene signatures of interest of interest of interest of the disclosed compositions and methods include a cardiovascular disease or condition, a kidney-related disease or condition, a prenatal or pregnancy-related disease or condition, a neurological or neuropsychiatric disease or condition, an autoimmune or immune-related disease or condition, a cancer, an infectious disease or condition, a pediatric disease, disorder or condition, a mitochondrial disorder, a respiratory-digestive tract disease or condition, a reproductive system disease or condition, an ophthalmic disease or condition, a musculoskeletal disease or condition, or a skin disease or condition.
[0127] sample The sample can include eukaryotic DNA, bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or a combination thereof. The multiple target dsDNA can include genomic DNA, mitochondrial DNA, plasmid DNA, or a combination thereof. The multiple target dsDNA can be from one or more organisms, one or more genes, or a combination thereof. The sample can be or be from a biological sample, a clinical sample, an environmental sample, or a combination thereof. The multiple target dsDNA can include DNA from at least two (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 species, or any number or range between these values) different organisms. The multiple target dsDNAs can include DNA from at least two (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range between these values) different genes. The method can include generating multiple target dsDNAs from multiple target RNAs using reverse transcriptase. The multiple target dsDNAs can include target dsDNAs generated from target RNAs using reverse transcriptase. The sample nucleic acid can include eukaryotic DNA, bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or combinations thereof. The target dsDNA can be genomic DNA, mitochondrial DNA, plasmid DNA, or combinations thereof. The sample nucleic acid can be from a biological sample, a clinical sample, an environmental sample, or combinations thereof. The biological sample can include stool, sputum, peripheral blood, plasma, serum, lymph nodes, respiratory tissue, exudates, bodily fluids, or combinations thereof.
[0128] The nucleic acids utilized in the methods described herein can be obtained from any suitable biological sample, and are often isolated from a sample obtained from a subject, which can be any living or non-living organism, including, but not limited to, a human, a non-human animal, a plant, a bacterium, a fungus, a virus, or a protist. Any human or non-human animal can be selected, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), swine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry, dog, cat, mouse, rat, fish, dolphin, whale, and shark. The subject can be male or female, and the subject can be of any age (e.g., embryo, fetus, infant, child, adult).
[0129] A sample or test sample can be any specimen isolated or obtained from a subject or a part thereof. Non-limiting examples of specimens include bodily fluids or tissues from a subject, including but not limited to blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, bone marrow, chorion, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, stomach, peritoneal, duct, ear, arthroscope), biopsy samples, serocentesis samples, cells (e.g., blood cells) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), female genital tract washings, urine, stool, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, breast fluid, hard tissues (e.g., liver, spleen, kidney, lung, or ovary), etc., or combinations thereof. The term blood, as conventionally defined, includes whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc. Blood plasma refers to the fraction of whole blood resulting from centrifugation of blood that has been treated with an anticoagulant. Blood serum refers to the aqueous portion of the fluid remaining after a blood sample has clotted. Samples of body fluids or tissues are often collected according to standard protocols typically followed by hospitals or clinics. For blood, an appropriate amount of peripheral blood (e.g., 3-40 milliliters) is often collected and may be stored according to standard procedures before or after preparation.
[0130] The sample or test sample may include a sample containing nucleic acid from spores, viruses, cells, prokaryotes or eukaryotes, or any free nucleic acid. For example, the methods described herein may be used to detect nucleic acid outside of spores (e.g., without the need for lysis). The sample may be isolated from any material suspected of containing the target sequence, such as from a subject as described above. In some embodiments, the target sequence is present in air, plants, soil, or other material suspected of containing biological organisms.
[0131] Nucleic acids can be derived (e.g., isolated, extracted, purified) from one or more sources by methods known in the art. Any suitable method can be used to isolate, extract and / or purify nucleic acids from biological samples, non-limiting examples of which include methods of DNA preparation in the art and various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis), GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ), and the like, or combinations thereof.
[0132] In some embodiments, a cell lysis technique is performed. Cell lysis can be performed before the initiation of the reactions provided herein. Cell lysis techniques and reagents are known in the art and can generally be performed by chemical lysis (e.g., detergents, hypotonic solutions, enzymatic techniques, etc., or combinations thereof), physical lysis (e.g., pressurized cell disruption, sonication, etc.), or electrolytic lysis. Any suitable lysis technique can be utilized. For example, chemical methods generally employ lysis agents to disrupt cells and extract nucleic acids from the cells, followed by treatment with chaotropic salts. In some embodiments, cell lysis includes the use of detergents (e.g., ionic, nonionic, anionic, zwitterionic). In some embodiments, cell lysis includes the use of ionic detergents (e.g., sodium dodecyl sulfate (SDS), sodium lauryl sulfate (SLS), deoxycholate, cholate, sarkosyl). Physical methods such as the use of freeze / thaw followed by crushing, cell squeezing, etc. can also be useful. High salt lysis techniques can also be used. For example, alkaline lysis techniques can be utilized. The latter approach traditionally incorporates the use of phenol-chloroform solutions, and an alternative phenol-chloroform-free approach involving three solutions can be utilized. In the latter approach, one solution can contain 15 mM Tris, pH 8.0; 10 mM EDTA and 100 μg / ml RNase A; a second solution can contain 0.2 N NaOH and 1% SDS; and a third solution can contain, for example, 3 M KOAc, pH 5.5. In some embodiments, a cell lysis buffer is used with the methods and components described herein.
[0133] The nucleic acid may be provided to perform the methods described herein without processing a sample containing the nucleic acid. For example, in some embodiments, the nucleic acid is provided to perform the amplification methods described herein without prior nucleic acid purification. In some embodiments, the target sequence is amplified directly from the sample (e.g., without any nucleic acid extraction, isolation, purification and / or partial purification steps). In some embodiments, the nucleic acid is provided to perform the methods described herein after processing a sample containing the nucleic acid. For example, the nucleic acid may be extracted, isolated, purified, or partially purified from the sample. The term "isolated" generally refers to a nucleic acid that has been removed from its original environment (e.g., the natural environment if it is naturally occurring, or the host cell if it is exogenously expressed) and thus is altered by human intervention (e.g., "by the hand of man") from its original environment. The term "isolated nucleic acid" may refer to a nucleic acid that has been removed from a subject (e.g., a human subject). An isolated nucleic acid may provide less non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount of those components present in the source sample. A composition containing an isolated nucleic acid may be about 50% to more than 99% free of non-nucleic acid components. A composition containing an isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of non-nucleic acid components. The term "purified" generally refers to a nucleic acid that contains less non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount of non-nucleic acid components present before the nucleic acid is subjected to a purification procedure. A composition containing a purified nucleic acid may be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of other non-nucleic acid components.
[0134] Nucleic acids can be provided for carrying out the methods described herein without modifying the nucleic acid, including, for example, denaturation, digestion, nicking, unwinding, incorporation and / or ligation of heterogeneous sequences, addition of epigenetic modifications, labeling (e.g., 32 P, 33 P,125 I or 35 These may include the addition of a radioactive label such as S; an enzyme label such as alkaline phosphatase; a fluorescent label such as fluorescein isothiocyanate (FITC); or other labels such as biotin, avidin, digoxigenin, antigens, haptens, fluorescent dyes, etc. Thus, in some embodiments, the unmodified nucleic acid is amplified.
[0135] The disclosed methods for detecting a target nucleic acid sequence (single- or double-stranded DNA and / or RNA) in a sample can detect the target nucleic acid sequence (e.g., DNA or RNA) with a high degree of sensitivity. In some embodiments, the disclosed methods can be used to detect a target RNA / DNA present in a sample containing multiple RNAs / DNAs (including a target RNA / DNA and multiple non-target RNAs / DNAs), where the target RNA / DNA is present in an amount of 1 or more copies / 10 7 non-target RNA / DNA (e.g., 1 copy or more / 10 6 Non-targeted RNA / DNA, ≥1 copy / 10 5 Non-targeted RNA / DNA, ≥1 copy / 10 4 Non-targeted RNA / DNA, ≥1 copy / 10 3 Non-targeted RNA / DNA, ≥1 copy / 10 2 non-target RNA / DNA, 1 or more copies per 50 non-target RNA / DNA, 1 or more copies per 20 non-target RNA / DNA, 1 or more copies per 10 non-target RNA / DNA, or 1 or more copies per 5 non-target RNA / DNA. In some embodiments, the methods of the present disclosure can be used to detect a target RNA / DNA present in a sample that contains multiple RNAs / DNAs (including a target RNA / DNA and multiple non-target RNA / DNAs), where the target RNA / DNA is present at 1 or more copies per 10 18 non-target RNA / DNA (e.g., 1 copy or more / 10 15 Non-targeted RNA / DNA, ≥1 copy / 10 12 Non-targeted RNA / DNA, ≥1 copy / 10 9 Non-targeted RNA / DNA, ≥1 copy / 10 6 Non-targeted RNA / DNA, ≥1 copy / 105 Non-targeted RNA / DNA, ≥1 copy / 10 4 Non-targeted RNA / DNA, ≥1 copy / 10 3 Non-targeted RNA / DNA, ≥1 copy / 10 2 non-targeted RNA / DNA, 1 or more copies / 50 non-targeted RNA / DNA, 1 or more copies / 20 non-targeted RNA / DNA, 1 or more copies / 10 non-targeted RNA / DNA, or 1 or more copies / 5 non-targeted RNA / DNA). As used herein, the terms "RNA / DNA" and "RNAs / DNAs" shall be given their ordinary meanings and shall refer to DNA, or RNA, or a combination of DNA and RNA.
[0136] In some embodiments, the disclosed methods can detect target RNA / DNA present in a sample, and the target RNA / DNA is at least 1 copy / 10 7 1 copy / 10 non-target RNA / DNA to 1 copy / 10 non-target RNA / DNA (e.g., 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 3 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 4 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 6 Non-targeted RNA / DNA, 1 copy / 10 6 non-targeted RNA / DNA ~ 1 copy / 10 non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 3 Non-targeted RNA / DNA, 1 copy / 106 Non-targeted RNA / DNA ~1 copy / 10 4 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 5 non-targeted RNA / DNA ~ 1 copy / 10 non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 3 non-targeted RNA / DNA or 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 4 non-target RNA / DNA).
[0137] In some embodiments, the disclosed methods can detect target RNA / DNA present in a sample, and the target RNA / DNA is at least 1 copy / 10 18 1 copy / 10 non-target RNA / DNA to 1 copy / 10 non-target RNA / DNA (e.g., 1 copy / 10 18 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 15 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 12 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 9 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 3 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 4Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 6 Non-targeted RNA / DNA, 1 copy / 10 6 non-targeted RNA / DNA ~ 1 copy / 10 non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 3 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 4 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 5 non-targeted RNA / DNA ~ 1 copy / 10 non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 3 non-targeted RNA / DNA or 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 4 non-target RNA / DNA).
[0138] In some embodiments, the methods of the present disclosure can detect a target RNA / DNA (e.g., a target nucleic acid sequence) present in a sample, and the target RNA / DNA is present at a concentration of 1 copy / 10 7 1 copy / 10 non-target RNA / DNA to 1 copy / 100 non-target RNA / DNA (e.g., 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 103 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 4 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 7 Non-targeted RNA / DNA ~1 copy / 10 6 Non-targeted RNA / DNA, 1 copy / 10 6 ~1 copy / 100 non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 3 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 4 Non-targeted RNA / DNA, 1 copy / 10 6 Non-targeted RNA / DNA ~1 copy / 10 5 Non-targeted RNA / DNA, 1 copy / 10 5 ~1 copy / 100 non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 2 Non-targeted RNA / DNA, 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 3 non-targeted RNA / DNA or 1 copy / 10 5 Non-targeted RNA / DNA ~1 copy / 10 4 non-target RNA / DNA).
[0139] In some embodiments, the detection threshold for the subject methods for detecting a target RNA / DNA (e.g., a target nucleic acid sequence) in a sample is 10 nM or less. The term "detection threshold" is used herein to describe the minimum amount of target RNA / DNA that must be present in a sample for detection to occur. Thus, as an illustrative example, if the detection threshold is 10 nM, a signal can be detected when the target RNA / DNA is present in the sample at a concentration of 10 nM or more. In some embodiments, the disclosed methods have a detection threshold of 5 nM or less. In some embodiments, the disclosed methods have a detection threshold of 1 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.5 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.1 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.05 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.01 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.005 nM or less. In some embodiments, the disclosed methods have a detection threshold of 0.001 nM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 0.0005 nM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 0.0001 nM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 0.00005 nM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 0.00001 nM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 10 pM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 1 pM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 500 fM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 250 fM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 100 fM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 50 fM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 500 aM (attomolar) or less. In some embodiments, the methods of the present disclosure have a detection threshold of 250 aM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 100 aM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 50 aM or less.In some embodiments, the methods of the present disclosure have a detection threshold of 10 aM or less. In some embodiments, the methods of the present disclosure have a detection threshold of 1 aM or less.
[0140] In some embodiments, the detection threshold (for detecting a target RNA / DNA in a subject method) is in the range of 500 fM to 1 nM (e.g., 500 fM to 500 pM, 500 fM to 200 pM, 500 fM to 100 pM, 500 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, or 1 pM to 10 pM) (concentration refers to the threshold concentration of the target RNA / DNA at which the target RNA / DNA can be detected). In some embodiments, the disclosed methods have a detection threshold in the range of 800 fM to 100 pM. In some embodiments, the disclosed methods have a detection threshold in the range of 1 pM to 10 pM. In some embodiments, the disclosed methods have a detection threshold in the range of 10 fM to 500 fM, e.g., 10 fM to 50 fM, 50 fM to 100 fM, 100 fM to 250 fM, or 250 fM to 500 fM.
[0141] In some embodiments, the minimum concentration at which a target RNA / DNA (e.g., a target nucleic acid sequence) can be detected in a sample is in the range of 500 fM to 1 nM (e.g., 500 fM to 500 pM, 500 fM to 200 pM, 500 fM to 100 pM, 500 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, or 1 pM to 10 pM). In some embodiments, the minimum concentration at which a target RNA / DNA can be detected in a sample is in the range of 800 fM to 100 pM. In some embodiments, the minimum concentration at which the target RNA / DNA can be detected in a sample ranges from 1 pM to 10 pM.
[0142] In some embodiments, the detection threshold (for detecting target RNA / DNA in the subject methods) is between 1 aM and 1 nM (e.g., between 1 aM and 500 pM, between 1 aM and 200 pM, between 1 aM and 100 pM, between 1 aM and 10 pM, between 1 aM and 1 pM, between 100 aM and 1 nM, between 100 aM and 500 pM, between 100 aM and 200 pM, between 100 aM and 100 pM, between 100 aM and 10 pM, 00aM~1pM, 250aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10pM, 250aM~1pM, 500a M~1nM, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1pM, 750aM~1nM, 750aM~50 0pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1nM, 1fM~500pM, 1fM~200pM, 1fM ~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500fM~100pM, 500fM~10pM, 500fM~1pM, 800fM to 1nM, 800fM to 500pM, 800fM to 200pM, 800fM to 100pM, 800fM to 10pM, 800fM to 1pM, 1pM to 1nM, 1pM to 500pM, 1pM to 200pM, 1pM to 100pM, or 1pM to 10pM) (wherein the concentration refers to the threshold concentration of the target RNA / DNA at which the target RNA / DNA can be detected). In some embodiments, the method of the present disclosure has a detection threshold in the range of 1aM to 800aM. In some embodiments, the method of the present disclosure has a detection threshold in the range of 50aM to 1pM. In some embodiments, the method of the present disclosure has a detection threshold in the range of 50aM to 500fM.
[0143] In some embodiments, the minimum concentration at which a target RNA / DNA (e.g., a target nucleic acid sequence) can be detected in a sample is between 1 aM and 1 nM (e.g., between 1 aM and 500 pM, between 1 aM and 200 pM, between 1 aM and 100 pM, between 1 aM and 1 pM, between 100 aM and 1 nM, between 100 aM and 500 pM, between 100 aM and 200 pM, between 100 aM and 100 pM). , 100aM~10pM, 100aM~1pM, 250aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10 pM, 250aM~1pM, 500aM~1nM, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1 pM, 750aM~1nM, 750aM~500pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1n M, 1fM~500pM, 1fM~200pM, 1fM~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500 In some embodiments, the minimum concentration at which the target RNA / DNA can be detected in the sample is in the range of 1 aM to 500 pM. In some embodiments, the minimum concentration at which the target RNA / DNA can be detected in the sample is in the range of 100 aM to 500 pM. In some embodiments, the disclosed compositions or methods exhibit attomolar (aM) detection sensitivity. In some embodiments, the disclosed compositions or methods exhibit femtomolar (fM) detection sensitivity. In some embodiments, the disclosed compositions or methods exhibit picomolar (pM) detection sensitivity. In some embodiments, the disclosed compositions or methods exhibit nanomolar (nM) detection sensitivity.
[0144] The disclosed samples include sample nucleic acids (e.g., multiple sample nucleic acids). The term "multiple" is used herein to mean two or more. Thus, in some embodiments, a sample includes two or more (e.g., three or more, five or more, ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more) sample nucleic acids (e.g., RNA). The disclosed methods can be used as highly sensitive methods for detecting target nucleic acids present in a sample (e.g., in a complex mixture of nucleic acids such as RNA). In some embodiments, a sample includes five or more DNAs (e.g., ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more) that differ from each other in sequence. In some embodiments, a sample includes ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, ten or more, 3 More than seeds, 5×10 3 More than 10 species 4 More than seeds, 5×10 4 More than 10 species 5 More than seeds, 5×10 5 More than 10 species 6 More than seeds, 5×10 6 More than 10 species 7 In some embodiments, the sample contains 10-20, 20-50, 50-100, 100-500, 500-10 3 seeds, 10 3 ~5×10 3 Seeds, 5x10 3 ~10 4 seeds, 10 4 ~5×10 4 Seeds, 5x10 4 ~10 5 seeds, 10 5 ~5×10 5 Seeds, 5x10 5 ~10 6 seeds, 10 6 ~5×10 6 Seeds or 5x10 6 ~10 7 Seeds or 10 7 In some embodiments, the sample contains 5 to 10 7RNA of different species (e.g., different sequences from each other) (e.g., 5-10 6 Seeds, 5-10 5 species, 5-50,000 species, 5-30,000 species, 10-10 6 Seeds, 10~10 5 species, 10-50,000 species, 10-30,000 species, 20-10 6 Seeds, 20~10 5 In some embodiments, the sample comprises 20 or more species of RNA that differ from one another in sequence. In some embodiments, the sample comprises RNA from a cell lysate (e.g., a eukaryotic cell lysate, a mammalian cell lysate, a human cell lysate, a prokaryotic cell lysate, a plant cell lysate, etc.). For example, in some embodiments, the sample comprises DNA from a cell, such as a eukaryotic cell, e.g., a mammalian cell, such as a human cell.
[0145] The term "sample", as used herein, shall be given its ordinary meaning and shall include any sample containing RNA and / or DNA (e.g., to determine whether target DNA and / or target RNA is present in a population of RNA and / or DNA). A sample may be from any source, e.g., a sample may be a synthetic combination of purified DNA and / or RNA; a sample may be a cell lysate, a DNA / RNA enriched cell lysate, or DNA / RNA isolated and / or purified from a cell lysate. A sample may be from a patient (e.g., for diagnostic purposes). A sample may be from permeabilized cells. A sample may be from crosslinked cells. A sample may be in a tissue section. A sample may be from tissue prepared by crosslinking followed by delipidation and adjustment to create a uniform refractive index.
[0146] Suitable samples include, but are not limited to, saliva, blood, serum, plasma, urine, aspirates, and biopsy samples. Thus, the term "sample" with respect to a patient includes blood and other liquid samples of biological origin, solid tissue samples, such as biopsy specimens or tissue cultures or cells derived therefrom, and their progeny. This definition also includes samples that have been manipulated in any way after their procurement, such as by treatment with reagents, washing, or enrichment for certain cell populations, such as cancer cells. This definition also includes samples enriched for specific types of molecules, such as RNA. The term "sample" includes biological samples, such as clinical samples, such as blood, plasma, serum, aspirates, cerebrospinal fluid (CSF), and also includes tissue obtained by surgical resection, tissue obtained by biopsy, cultured cells, cell supernatants, cell lysates, tissue samples, organs, bone marrow, etc. "Biological sample" includes biological fluids derived therefrom (e.g., cancerous cells, infected cells, etc.), such as samples containing RNA obtained from such cells (e.g., cell lysates or other cell extracts containing RNA).
[0147] In some embodiments, the source of the sample is (or is suspected of being) a diseased cell, body fluid, tissue, or organ. In some embodiments, the source of the sample is a normal (non-disease) cell, body fluid, tissue, or organ. In some embodiments, the source of the sample is (or is suspected of being) a cell, tissue, or organ infected with a pathogen. For example, the source of the sample can be an individual who may or may not be infected, and the sample can be any biological sample taken from the individual (e.g., blood, saliva, biopsy, plasma, serum, bronchoalveolar lavage, sputum, fecal sample, cerebrospinal fluid, fine needle aspirate, swab sample (e.g., oral swab, cervical swab, nasal swab), interstitial fluid, synovial fluid, nasal secretion, tears, buffy coat, mucosal sample, epithelial cell sample (e.g., epithelial cell peeling), etc.). In some embodiments, the sample is a cell-free liquid sample. In some embodiments, the sample is a liquid sample that can contain cells. Pathogens include viruses, fungi, helminths, protozoa, malarial parasites, Plasmodium parasites, Toxoplasma parasites, Schistosoma parasites, and the like. "Helminths" include roundworms, heart worms, and plant nematodes, trematodes, Acanthocephala, and cestoda. Protozoal infections include infections from Giardia spp., Trichomonas spp., African trypanosomiasis, amebic dysentery, babesiosis, balantidial dysentery, Chagas' disease, coccidiosis, malaria, and toxoplasmosis. Examples of pathogens, such as parasitic / protozoal pathogens, include, but are not limited to, Plasmodium falciparum, Plasmodium vivax, Trypanosoma cruzi, and Toxoplasma gondii. Fungal pathogens include, but are not limited to, Cryptococcus neoformans, Histoplasma capsulatum, Coccidioides immitis, and the like.immitis, Blastomyces dermatitidis, Chlamydia trachomatis, and Candida albicans. Pathogenic viruses include, for example, immunodeficiency viruses (e.g., HIV), influenza viruses, dengue, West Nile virus, herpes viruses, yellow fever viruses, hepatitis C virus, hepatitis A virus, hepatitis B virus, papilloma viruses, etc. Pathogenic viruses include DNA viruses, such as papovaviruses (e.g., human papillomavirus (HPV) and polyomavirus); hepadnaviruses (e.g., hepatitis B virus (HBV)); herpes viruses (e.g., herpes simplex virus (HSV)), varicella zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpes lymphotropic virus, and pityriasis rosea. Rosea), Kaposi's sarcoma-associated herpesvirus); adenoviruses (e.g., atadenovirus, aviadenovirus, ictadenovirus, mastadenovirus, siaadenovirus); poxviruses (e.g., smallpox virus, vaccinia virus, cowpox virus, monkeypox virus, orf virus, pseudopox virus, bovine papular stomatitis virus; tanapox virus, yabasa tumor virus; molluscum contagiosum virus (MCV)); parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, human parv4G1; Geminiviridae; Nanoviridae; Phycodnaviridae, etc. Pathogens may include, for example, DNA viruses [e.g., papovaviruses (e.g., human papillomavirus (HPV), polyomaviruses); hepadnaviruses (e.g., hepatitis B virus (HBV)); herpes viruses (e.g., herpes simplex virus (HSV), varicella zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpes lymphotropic virus, pityriasis rosea, Kaposi's sarcoma-associated herpes virus); adenoviruses (e.g., Attaviridae, adenoviruses, aviadenoviruses, ictadenoviruses, mastadenoviruses, siaadenoviruses; poxviruses (e.g., smallpox virus, vaccinia virus, cowpox virus, monkeypox virus, orf virus, pseudopox virus, bovine papular stomatitis virus; tanapox virus, yabasa tumor virus; molluscum contagiosum virus (MCV)); parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, human parv4) G1; Geminiviridae; Nanoviridae; Phycodnaviridae, etc.], Mycobacterium tuberculosis, Streptococcus agalactiae, Methicillin-resistant Staphylococcus aureus, Legionella pneumophila, Streptococcus pyogenes, Escherichia coli, Neisseria gonorrhoeae, Neisseria meningitidis, Streptococcus pneumophila, Cryptococcus neoformans, Histoplasma capsulatum, Haemophilus influenzae B, Treponema pallidum, Lyme disease spirochete, Pseudomonas aeruginosa, Mycobacterium leprae leprae, Brucella abortusabortus), rabies virus, influenza virus, cytomegalovirus, herpes simplex virus I, herpes simplex virus II, human serum parvo-like virus, respiratory syncytial virus, varicella-zoster virus, hepatitis B virus, hepatitis C virus, measles virus, adenovirus, human T-cell leukemia virus, Epstein-Barr virus, murine leukemia virus, mumps virus, vesicular stomatitis virus, Sindbis virus, lymphocytic choriomeningitis virus, wart virus, blue tongue virus, Sendai virus, feline leukemia virus, reovirus, poliovirus, simian virus 40, mouse mammary tumor virus, dengue virus, rubella virus, West Nile virus, Plasmodium falciparum, Plasmodium vivax, Toxoplasma gondii, Trypanosoma rangeli, Trypanosoma cruzi, Trypanosoma rhodesiens rhodesiense, Trypanosoma brucei, Schistosoma mansoni, Schistosoma japonicum, Babesia bovis, Eimeria tenella, Onchocerca volvulus, Leishmania tropica, Mycobacterium tuberculosis, Trichinella spiralis, Theileria parva, Taenia hydatigena, Taenia ovis, Taenia saginata, Echinococcus granulosus, Mesocestoides corti, Mycoplasma arthritidis, M. hyorhinis, M. orale, M. arginini, Acholeplasma laidlawii, M. salivariumThe pathogenic virus may include one or more of SARS-CoV-2, influenza A, influenza B, and / or influenza C.
[0148] The sample may be a biological sample, e.g., a clinical sample. In some embodiments, the sample is taken from a biological source, e.g., vagina, urethra, penis, anus, throat, cervix, fermentation broth, cell culture, etc. The sample may include fluids and cells from, e.g., fecal samples. Biological samples can be used (i) as obtained directly from a subject or source, or (ii) after pretreatment to modify the characteristics of the sample. Thus, a test sample can be pretreated before use, e.g., by disrupting cells or viral particles, preparing liquids from solid materials, diluting viscous fluids, filtering liquids, concentrating liquids, inactivating interfering components, adding reagents, purifying nucleic acids, etc. Thus, as used herein, a "biological sample" includes nucleic acids (DNA, RNA, or total nucleic acids) extracted from clinical or biological specimens. Sample preparation can also include using solutions containing buffers, salts, detergents, etc., used to prepare the sample for analysis. In some embodiments, the sample is processed prior to molecular testing. In some embodiments, the sample is analyzed directly and is not pretreated prior to testing. The sample can be, for example, a fecal sample. In some embodiments, the sample is a fecal sample from a patient with clinical symptoms of acute gastroenteritis.
[0149] In some embodiments, the sample to be tested is processed before performing the methods disclosed herein. For example, in some embodiments, the sample can be isolated, concentrated, or subjected to various other processing steps before performing the methods disclosed herein. For example, in some embodiments, the sample can be processed to isolate nucleic acid from the sample before contacting the sample with an oligonucleotide as disclosed herein. In some embodiments, the methods disclosed herein are performed on the sample without culturing the sample in vitro. In some embodiments, the methods disclosed herein are performed on the sample without isolating nucleic acid from the sample before contacting the sample with an oligonucleotide as disclosed herein.
[0150] A sample can contain one or more nucleic acids (e.g., multiple nucleic acids). As used herein, the term "multiple" can refer to two or more. Thus, in some embodiments, a sample contains two or more (e.g., three or more, five or more, ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more) nucleic acids (e.g., gDNA, mRNA). The disclosed methods can be used as highly sensitive methods for detecting target nucleic acids present in a sample (e.g., in a complex mixture of nucleic acids such as gDNA). In some embodiments, a sample contains five or more nucleic acids (e.g., ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more RNA) that differ from each other in sequence. In some embodiments, a sample contains ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, ten or more, 3 More than seeds, 5×10 3 More than 10 species 4 More than seeds, 5×10 4 More than 10 species 5 More than seeds, 5×10 5 More than 10 species 6 More than seeds, 5×10 6 More than 10 species 7 It contains one or more species of nucleic acid.
[0151] In some embodiments, the samples include 10-20 species, 20-50 species, 50-100 species, 100-500 species, 500-10 3 seeds, 10 3 ~5×10 3 Seeds, 5x10 3 ~10 4 seeds, 10 4 ~5×10 4 Seeds, 5x10 4 ~10 5 seeds, 10 5 ~5×10 5 Seeds, 5x10 5 ~10 6 seeds, 10 6 ~5×10 6 Seeds or 5x10 6 ~10 7 Seeds or 10 7 In some embodiments, the sample contains 5 to 10 7 Nucleic acids of different species (e.g., differing in sequence from each other) (e.g., 5-10 6 Seeds, 5-10 5 species, 5-50,000 species, 5-30,000 species, 10-10 6 Seeds, 10~10 5 species, 10-50,000 species, 10-30,000 species, 20-10 6 Seeds, 20~10 5 In some embodiments, the sample comprises 20 or more nucleic acids that differ from one another in sequence.
[0152] The sample can be any sample that contains nucleic acid (e.g., to determine if a target nucleic acid is present in a population of nucleic acids). The sample can be derived from any source, e.g., the sample can be a synthetic combination of purified nucleic acid; the sample can be a cell lysate, a DNA-enriched cell lysate, or nucleic acid isolated and / or purified from a cell lysate. The sample can be from a patient (e.g., for diagnostic purposes). The sample can be from permeabilized cells. The sample can be from crosslinked cells. The sample can be in a tissue section. The sample can be from tissue that has been prepared by crosslinking followed by delipidation and adjustment to create a uniform refractive index.
[0153] The sample can include a target nucleic acid and a plurality of non-target nucleic acids. In some embodiments, the target nucleic acid is 1 copy / 10 non-target nucleic acids, 1 copy / 20 non-target nucleic acids, 1 copy / 25 non-target nucleic acids, 1 copy / 50 non-target nucleic acids, 1 copy / 100 non-target nucleic acids, 1 copy / 500 non-target nucleic acids, 1 copy / 10 3 Non-target nucleic acid, 1 copy / 5×10 3 Non-target nucleic acid, 1 copy / 10 4 Non-target nucleic acid, 1 copy / 5×10 4 Non-target nucleic acid, 1 copy / 10 5 Non-target nucleic acid, 1 copy / 5×10 5 Non-target nucleic acid, 1 copy / 10 6 Non-target nucleic acid, <1 copy / 10 6 non-target nucleic acids, or a value or range between any two of these values. In some embodiments, the target nucleic acids are present in the sample at 1 copy / 10 non-target nucleic acids to 1 copy / 20 non-target nucleic acids, 1 copy / 20 non-target nucleic acids to 1 copy / 50 non-target nucleic acids, 1 copy / 50 non-target nucleic acids to 1 copy / 100 non-target nucleic acids, 1 copy / 100 non-target nucleic acids to 1 copy / 500 non-target nucleic acids, 1 copy / 500 non-target nucleic acids to 1 copy / 10 3 Non-target nucleic acid, 1 copy / 10 3 Non-target nucleic acid ~1 copy / 5×10 3 Non-target nucleic acid, 1 copy / 5×10 3Non-targeted nucleic acid ~1 copy / 10 4 Non-target nucleic acid, 1 copy / 10 4 Non-targeted nucleic acid ~1 copy / 10 5 Non-target nucleic acid, 1 copy / 10 5 Non-targeted nucleic acid ~1 copy / 10 6 non-target nucleic acid, or 1 copy / 10 6 Non-targeted nucleic acid ~1 copy / 10 7 non-target nucleic acids are present in the sample, or a value or range between any two of these values.
[0154] Suitable samples include, but are not limited to, saliva, blood, serum, plasma, urine, aspirates, and biopsy samples. Thus, the term "sample" with respect to a patient encompasses blood and other liquid samples of biological origin, solid tissue samples such as biopsy specimens or tissue cultures or cells derived therefrom, and their progeny. This definition also includes samples that have been manipulated in any way after their procurement, such as by treatment with reagents, washing, or enrichment for certain cell populations, such as cancer cells. This definition also includes samples enriched for specific types of molecules, such as nucleic acids. The term "sample" encompasses biological samples, such as clinical samples, such as blood, plasma, serum, aspirates, cerebrospinal fluid (CSF), and also includes tissues obtained by surgical resection, tissues obtained by biopsy, cultured cells, cell supernatants, cell lysates, tissue samples, organs, bone marrow, and the like. A "biological sample" includes biological fluids derived therefrom (e.g., cancerous cells, infected cells, etc.), such as samples containing nucleic acids obtained from such cells (e.g., cell lysates or other cell extracts containing nucleic acids).
[0155] Suitable samples for use in the methods disclosed herein include any conventional biological sample obtained from an organism or part thereof, such as a plant, animal, bacteria, etc. In certain embodiments, the biological sample is obtained from an animal subject, such as a human subject. A biological sample is any solid or fluid sample obtained, excreted, or secreted from any organism, including, but not limited to, unicellular organisms, multicellular organisms, such as bacteria, yeast, protozoa, and amoebas, among others, including samples from healthy or apparently healthy human subjects, such as plants or animals, or human patients suffering from a condition or disease of interest to be diagnosed or investigated, such as infection with a pathogenic microorganism, such as a pathogenic bacterium or virus. For example, the biological sample can be, for example, blood, plasma, serum, urine, stool, sputum, mucus, lymphatic fluid, synovial fluid, bile, ascites, pleural fluid, seroma, saliva, cerebrospinal fluid, aqueous humor, or vitreous fluid, or any secretion, exudate, transudate (e.g., fluid obtained from an abscess, or any other site of infection or inflammation), or fluid obtained from a joint (e.g., rheumatoid arthritis, osteoarthritis, gout, or septic arthritis), or a biological fluid obtained from a swab of the skin or mucosal surface.
[0156] The sample may be a sample obtained from any organ or tissue (including a biopsy or autopsy specimen, e.g., a tumor biopsy), or may include cells (either primary or cultured cells) or media conditioned by any cell, tissue, or organ. Exemplary samples include, but are not limited to, cells, cell lysates, blood smears, cytocentrifugation preparations, cytology smears, bodily fluids (e.g., blood, plasma, serum, saliva, sputum, urine, bronchoalveolar lavage, semen, etc.), tissue biopsies (e.g., tumor biopsies), fine needle aspirates, and / or tissue sections (e.g., cryostat tissue sections and / or paraffin-embedded tissue sections). In other examples, the sample includes circulating tumor cells (which can be identified by cell surface markers). In particular examples, the sample may be used directly (e.g., fresh or frozen) or may be manipulated prior to use, for example, by fixing (e.g., using formalin) and / or embedding in wax (e.g., formalin-fixed paraffin-embedded (FFPE) tissue samples, etc.). Any method of obtaining tissue from a subject can be utilized, and it will be understood that the choice of method used will depend on a variety of factors, such as the type of tissue, the age of the subject, or the procedures available to the practitioner. Standard techniques for obtaining such samples are available in the art. The sample may be an environmental sample, such as water, soil, or a surface, such as an industrial or medical surface. Due to the increased sensitivity of the embodiments disclosed herein, in certain exemplary embodiments, the assays and methods can be performed on crude samples, or samples in which the target molecule to be detected has not been further fractionated or purified from the sample.
[0157] Cells can be lysed to release target molecules (e.g., target dsDNA). Cell lysis can be achieved by a variety of means, for example, chemical or biochemical means, osmotic shock, or thermal, mechanical, or optical lysis. Cells can be lysed by the addition of a cell lysis buffer containing a detergent (e.g., SDS, Li-dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association of the target with the barcode, the diffusion rate of the target molecule can be altered, for example, by lowering the temperature and / or increasing the viscosity of the lysate. In some embodiments, the sample may be lysed using filter paper, which may have a lysis buffer soaked on top of it, which may be applied to the sample with pressure that may facilitate lysis of the sample and hybridization of the sample to the target substrate.
[0158] In some embodiments, lysis can be performed by mechanical lysis, heat lysis, optical lysis, and / or chemical lysis. Chemical lysis can include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer can include Tris HCl. The lysis buffer can include at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris HCl. The lysis buffer can include at most about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris HCL. The lysis buffer can include about 0.1 M Tris HCl. The pH of the lysis buffer can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The pH of the lysis buffer can be at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer can include a salt (e.g., LiCl). The concentration of the salt in the lysis buffer can be at least about 0.1, 0.5, or 1 M or more. The concentration of the salt in the lysis buffer can be at most about 0.1, 0.5, or 1 M or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer can include a detergent (e.g., SDS, Li-dodecyl sulfate, triton X, tween, NP-40). The concentration of the detergent in the lysis buffer can be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. The concentration of detergent in the lysis buffer can be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. In some embodiments, the concentration of detergent in the lysis buffer is about 1% Li-dodecyl sulfate. The time used in the lysis method can vary depending on the amount of detergent used. In some embodiments, the more detergent used, the less time is required for lysis.The lysis buffer can include a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer can be at least about 1, 5, 10, 15, 20, 25, or 30 mM or more. The concentration of the chelating agent in the lysis buffer can be at least about 1, 5, 10, 15, 20, 25, or 30 mM or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer can include a reducing reagent (e.g., beta-mercaptoethanol, DTT). The concentration of the reducing reagent in the lysis buffer can be at least about 1, 5, 10, 15, or 20 mM or more. The concentration of the reducing reagent in the lysis buffer can be at most about 1, 5, 10, 15, or 20 mM or more. In some embodiments, the concentration of the reducing reagent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer can include about 0.1 M Tris HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.
[0159] Lysing can be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing can be performed for about 1, 5, 10, 15, or 20 minutes or more. Lysed cells can contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. Lysed cells can contain at most about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.
[0160] kit The kits described herein can include a plurality of protein complexes. In some embodiments, each of the plurality of protein complexes includes a transposome and a programmable DNA binding unit capable of specifically binding to a binding site on a target double-stranded DNA (dsDNA). In some embodiments, the transposome includes a transposase and two copies of an adapter. In some embodiments, the binding sites for each of the plurality of protein complexes are different from each other. In some embodiments, the kit includes at least one component that provides real-time detection activity for the nucleic acid amplification product. The real-time detection activity can be provided by a molecular beacon. The kit can include a reverse transcriptase and / or a reverse transcription primer. The kit can include one or more primers that can bind to one strand of the adapter. The kit may include one or more polymerases and one or more primers, and optionally one or more reverse transcriptases and / or reverse transcription primers, for example as described herein. If one target is amplified, one pair of primers (forward and reverse) may be included in the kit. If multiple target sequences are amplified, multiple primer pairs may be included in the kit. The kit may include a control polynucleotide, and if multiple target sequences are amplified, multiple control polynucleotides may be included in the kit.
[0161] The kits can also include one or more components in any number of separate vessels, chambers, containers, packets, tubes, vials, microtiter plates, etc., or the components can be combined in various combinations in such containers. The components of the kit can be, for example, in one or more containers. In some embodiments, all components are provided in one container. In some embodiments, the enzymes (e.g., polymerase and / or reverse transcriptase) can be provided in a separate container from the primers. The components can be, for example, lyophilized, heat dried, freeze dried, or in a stable buffer. In some embodiments, the polymerase and / or reverse transcriptase are in lyophilized or heat dried form in a single container, and the primers are lyophilized, heat dried, lyophilized, or in a buffer in a different container. In some embodiments, the polymerase and / or reverse transcriptase and the primers are in a single container in lyophilized or heat dried form. The kit may further include, for example, dNTPs used in the reaction, or modified nucleotides used in the reaction, containers, cuvettes or other vessels, or vials of water or buffer for rehydrating lyophilized or heat-dried components. The buffer used may, for example, be suitable for both the polymerase and primer annealing activities.
[0162] The kits can also include instructions for performing one or more of the methods described herein and / or a description of one or more of the components described herein. The instructions and / or instructions can be in printed form and can be included in a kit insert. The kits can also include a written description of an internet location that provides such instructions or instructions. The kit may further comprise reagents used in detection methods, such as reagents used in FRET, lateral flow devices, dipsticks, fluorescent dyes, colloidal gold particles, latex particles, molecular beacons, or polystyrene beads.
[0163] FIGS. 1A-1G, 2, 3, and 4A-4B of the present disclosure were generated using BioRender.com. EXAMPLES
[0164] Certain aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not intended to limit the scope of the disclosure. Example 1
[0165] Detection of bacterial vaginosis in vaginal swab specimens The examples demonstrate the detection of bacterial vaginosis (BV) in vaginal swab samples using the non-limiting exemplary detection methods described herein. Vaginal swab samples are collected from women with clinical symptoms of vaginosis. The samples are lysed and contacted with five protein complex pairs, each specific for one of the five BV-associated pathogens: (1) Atopobium vaginae, (2) BVAB-2, (3) Megasphaera type 1, (4) Gardnerella vaginalis, and (5) Lactobacillus species (Lactobacillus crispatus and Lactobacillus jensenii), to form a reaction mixture.
[0166] (1) The A. vaginae-specific protein complex pair includes: (a) a first Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to an sgRNA such that dCas9 can bind to a first binding site on the 16S rRNA gene of A. vaginae; and (b) a second Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to another sgRNA such that dCas9 can bind to a second binding site about 100-500 nucleotides downstream of the first binding site on the A. vaginae genome. The region between the first binding site and the second binding site is the target sequence of A. vaginae.
[0167] (2) The BVAB-2-specific protein complex pair includes: (a) a first Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to an sgRNA such that dCas9 can bind to a first binding site on the 16S rRNA gene of BVAB-2; and (b) a second Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to another sgRNA such that dCas9 can bind to a second binding site about 100 to 500 nucleotides downstream of the first binding site on the BVAB-2 genome. The region between the first and second binding sites is the target sequence of BVAB-2.
[0168] (3) The pair of megasphaera type 1-specific protein complexes includes: (a) a first Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to an sgRNA such that dCas9 can bind to a first binding site on the 16S rRNA gene of megasphaera type 1; and (b) a second Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to another sgRNA such that dCas9 can bind to a second binding site about 100 to 500 nucleotides downstream of the first binding site on the megasphaera type 1 genome. The region between the first and second binding sites is the target sequence of megasphaera type 1.
[0169] (4) The G. vaginalis-specific protein complex pair includes: (a) a first Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to an sgRNA such that dCas9 can bind to a first binding site on the vly gene of G. vaginalis; and (b) a second Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to another sgRNA such that dCas9 can bind to a second binding site about 100 to 500 nucleotides downstream of the first binding site on the G. vaginalis genome. The region between the first binding site and the second binding site is a target sequence of G. vaginalis. (5) The Lactobacillus species-specific protein complex pair includes (a) a first Tn5-dCAS9 fusion protein complex, in which Tn5 is linked to two copies of adapter A, and dCas9 is linked to an sgRNA such that dCas9 can bind to a first binding site on the 16S rRNA gene of the Lactobacillus species, and (b) a second Tn5-dCAS9 fusion protein complex, in which T5 is linked to two copies of adapter A, and dCas9 is linked to another sgRNA such that dCas9 can bind to a second binding site on the Lactobacillus species genome that is about 100-500 nucleotides downstream of the first binding site. The region between the first and second binding sites is the target sequence of the Lactobacillus species.
[0170] The reaction mixture is incubated to generate dsDNA fragments each containing adapter A at both ends of a target sequence from (1) Atopobium vaginae, (2) BVAB-2, (3) Megasphaera type 1, (4) Gardnerella vaginalis, or (5) Lactobacillus species (Lactobacillus crispatus and Lactobacillus jensenii), and the dsDNA fragments are amplified with a primer capable of binding to one strand of adapter A to generate an amplification product. The target sequence of each of these BV-related species in the amplified product is detected with a probe capable of specifically binding to the target sequence from (1) Atopobium vaginae, (2) BVAB-2, (3) Megasphaera type 1, (4) Gardnerella vaginalis, or (5) Lactobacillus species (Lactobacillus crispatus and Lactobacillus jensenii). The presence of a BV-related species target sequence in the amplified product is used to diagnose BV.
[0171] Example 2 Fusion protein and guide RNA (sgRNA) design and validation Four constructs were designed to generate fusion proteins: dCAS9-Fl26-Tn5, dCAS9-xTen-Tn5, Tn5-Fl26-dCas9, and Tn5-xTen-dCas9 (see, e.g., Figures 5-7). These constructs have either dCas9 or Tn5 sequences at the N-terminus of the fusion protein separated by a Fl26 linker or xTen linker. Plasmid design, in some embodiments, is based on "Chen, SP & Wang, HH (2019). An Engineered Cas-Transposon System for Programmable and Site-Directed DNA Transposition. The CRISPR Journal. Vol 2, Number 6. DOI: 10.1089 / crispr.2019.0030 and Picelli S., Bjorklund, AK, Reinius, B., Sgasser, S., Wingerb, G., & Sandbert, R. (2014)"; and "Tn5 transposase and tagmentation procedures for massively scaled sequencing projects. Genome Research. 24:2033-2040. ISSN 1088-9051 / 14".
[0172] sgRNA design sgRNAs were designed to target the InvA and FliC genes of Salmonella enterica. The sequence of Salmonella enterica strain ATCC 13311 was used. sgRNAs were designed using tools from Integrated DNA Technologies (IDT) (Table 2). The relative positions of sgRNAs in the InvA and FliC genes are shown in Figure 8 and Figure 9, respectively.
[0173] [Table 2] For InvA, 264 bp, 8 bp, 148 bp, 292 bp, 458 bp, and 195 bp fragments were predicted, whereas for FliC, approximately 130 bp, 82 bp, and 232 bp fragments were predicted.
[0174] Validation of Salmonella enterica sgRNA To verify the specificity of the sgRNA, genomic samples were cleaved by Cas9. Adapters were ligated to the Cas9-cleaved DNA, and PCR-amplified fragments were visualized by bioanalyzer.
[0175] [Table 3] Figure 10 and Table 3 show that cleavage of gDNA was specific to the expected size (compare the "Bioanalyzer predicted size [bp] column" with the "Actual [bp]" column in Table 3), thus indicating that the guide RNA for Salmonella enterica is functional. Next, sgRNAs targeting human genes EXT1, BCL9, HOXA13, HOXD11, and OLIG2 were designed for a total of 10 sgRNAs (Tables 4A to 4C). The sgRNAs were designed using the GenScript tool.
[0176] [Table 4]
[0177] [Table 5]
[0178] [Table 6] gRNAs were also designed to target the Chlamydia trachomatis gene polymorphic membrane protein A (pmp A) (Table 5). A total of five sgRNAs were designed using tools from IDT.
[0179] [Table 7]
[0180] Validation of transposase Tn5 Figures 11-12 show that Tn5 can ligate the designed adapter A (5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3', SEQ ID NO: 27) and adapter B (5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3', SEQ ID NO: 28) to DNA fragments for PCR amplification, demonstrating functionality. First, gDNA from S. enterica was cleaved with the custom adapters using Tn5 and pasted. Then, the labeled fragments were amplified by PCR. The data in Figures 11-12 show that Tn5 transposase was loaded with the custom adapters.
[0181] Fusion protein validation dCAS9-Fl26-Tn5, dCAS9-xTen-Tn5, Tn5-Fl26-dCas9, Tn5-xTen-dCas9 were recombinantly expressed and then purified. In some embodiments, recombinant proteins were isolated using a self-cleaving moiety (intein) on a chitin column. The purified fusion proteins were analyzed for expected size and purity on SDS-PAGE gels (Figures 13 to 18). SDS-PAGE analysis of dCAS9-Fl26-Tn5 is shown in Figure 13. The sample was observed to be >80% pure. In some embodiments, the fusion protein may also contain an intein domain. Bioanalyzer analysis in Figure 14 shows that a portion of the generated protein (peak at 44.91) is the correct size (without the intein). SDS-PAGE analysis of dCAS9-xTen-Tn5 is shown in Figure 15. The sample was observed to be >70% pure. In some embodiments, the fusion protein may also contain an intein domain, resulting in a larger than expected size. Bioanalyzer analysis in Figure 16 shows that a portion of the generated protein (peak at 44.62) is the correct size (without the intein). Figure 17 shows SDS-PAGE analysis of recombinantly expressed and purified Tn5-Fl26-dCas9. Figure 18 shows SDS-PAGE analysis of recombinantly expressed and purified Tn5-xTen-dCas9. Samples were observed to be >65% pure.
[0182] Testing the fusion protein for functionality Cas9-Fl26-Tn5 and dCas9-xTen-Tn5 were tested for functionality, and the protocol was as follows: (1) loading of sgRNA and adapter into the fusion protein (human sgRNA was used unless otherwise noted), (2) guided tagmentation, (3) cleanup, (4) PCR amplification, (5) quality control (QC), and (6) result analysis. Loading sgRNA and adapter into fusion protein The fusion proteins were loaded at a ratio of 1:1:2 (1 dCas9-Tn5 to 1 sgRNA to 2 adapters). The mixture was incubated at 24° C. for 30 min.
[0183] Guided tagmentation 100 mM dCas9-Tn5 (6.02 e10 molecules) and 500 ng human gDNA (1.52 e5 molecules) were combined at a ratio of 1 to 3.95 e5 of gDNA to dCas9-Tn5. The mixture was incubated at 37 °C for 60 minutes and at 55 °C for 60 minutes to generate tagged fragments. Several incubation methods were tried, and in some embodiments, dCas9 can function in the range of 25 °C to 42 °C. Tn5 can function in the range of 37 °C to 60 °C. The PCR amplification program is shown in Table 6.
[0184] [Table 8] Figure 19 shows data for the Cas9 only control reaction. The visible line shows the tape station analysis of DNA after Cas9 digestion. Sample analysis after PCR amplification reaction shows no signal. This data showed that Cas9 itself is not able to add adapters to the 5' or 3' ends of DNA fragments.
[0185] Figures 20-21 show the results of PCR amplification after digestion and ligation of adapters with dCas9-Fl26-Tn5 or dCas9-xTen-Tn5, respectively. The arrows in the graphs indicate the signal from the samples after PCR. PCR amplification was detected, but this is only possible if both fusion proteins (dCas9-Fl26-Tn5 and dCas9-xTen-Tn5) are able to perform transposition (e.g., adding adapters (adapter B) to the 5' and 3' ends of the DNA molecule). result The results show that Tn5 can add custom adapters to human gDNA. Cas9-only controls showed that this process requires Tn5 for amplification. These results demonstrated the functionality of Tn5 fused to dCas9.
[0186] Fusion protein to DNA ratio test Next, the effect of decreasing the gDNA to Cas9-Tn5 ratio was tested. The DNA concentration was kept constant while decreasing the Cas-Tn fusion protein concentration: 100 nM (194,071 molecules of dCas9-Tn5 to 1 genome copy of DNA), 1 nM (1,940:1), 100 pM (194:1), 10 pM (19.4:1), 1 pM (1.94:1). The results are shown in Figures 22-28. Figure 22 shows the results of PCR amplification after an induced tagmentation reaction using a 194,071:1 ratio of dCas9-Tn5, showing a broad peak after PCR, indicating non-specific tagmentation. The decrease in the amount of dCas9-Tn5 (Figures 23-28) resulted in the production of a detectable peak from the PCR reaction, indicating that decreasing the ratio of fusion protein to DNA adds specificity to the tagmentation. result The results show that Tn5 was able to add the custom adapter to human gDNA. A Cas9-only control showed that this process required Tn5 to amplify the DNA. Tn5 was shown to be functional and there was evidence of induced transposition. Thus, there is evidence of a fusion protein containing both dCas9 and Tn5 activity. Fusion proteins and sgRNAs on S. enterica Figures 35-36 show guided tagmentation using S. enterica sgRNA on dCas9-xTen-Tn5. The data shows that the addition of sgRNA adds specificity. Figure 36 shows that guided tagmentation without sgRNA is random. Figure 35 shows that the addition of sgRNA confers specificity.
[0187] Example 3 Sample library preparation Guided Tagmentation Libraries Described herein are methods and compositions for generating libraries for sequencing on an Illumina NextSeq. Three libraries were generated using a ligation-based method (Figure 34, Figure 37, Figure 39A-B), where the NEBNext sequencing adapter was added after the tagmentation step with a single adapter (e.g., adapter B with either Tn5 alone or dCas9-Tn5 fusion), and two libraries were generated using an induced tagmentation-based method (Figure 38, Figure 40-1), where the sequences required for NGS were included in the induced tagmentation step on adapters A and B. All libraries were prepared with human sgRNA. For induced tagmentation, dCas9-Fl26-Tn5 fusion protein was used. In these experiments, DNA was incubated with dCas-Tn5 under long or short incubation protocols. In the short protocol, the reaction was incubated at 30°C for 30 minutes, then at 37°C for 30 minutes. In the long protocol, reactions were incubated at 30°C for 30 min, followed by incubation at 38°C for 60 min, then at 55°C for 60 min. Figure 29 shows highly multiplexed single primer DNA amplification using only Tn5. Bioanalyzer analysis shows non-specific DNA amplification by PCR, indicating that DNA can be amplified using only one primer (adapter B).
[0188] Evidence supporting highly multiplexed single primer DNA amplification using dCas9-Tn fusion proteins is shown in Figure 30 (short incubation protocol) and Figure 31 (longer incubation protocol). Bioanalyzer analysis of the PCR amplifications showed simultaneous specific amplification of several DNA fragments using only one primer (adapter B). Evidence supporting customized locus-specific sequence library preparation is shown in Figure 32 (longer incubation protocol) and Figure 33 (shorter incubation protocol). Bioanalyzer analysis shows that a sequencing library is generated. Addition of adapters A and B required for sequencing on the Illumina platform shows that a sequencing library can be generated using guided tagmentation.
[0189] In at least some of the above-described embodiments, one or more elements used in an embodiment may be used interchangeably in another embodiment, unless such substitution is technically not feasible. Those skilled in the art will appreciate that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and variations are intended to fall within the scope of the subject matter, as defined by the appended claims.
[0190] With respect to the use of virtually any plural and / or singular term herein, one of ordinary skill in the art may translate from plural to singular and / or from singular to plural as appropriate to the context and / or application. Various singular / plural permutations may be expressly set forth herein for clarity. As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Any reference to "or" herein is intended to encompass "and / or" unless specifically stated otherwise.
[0191] In general, it will be understood by those skilled in the art that the terms used herein, and particularly the terms used in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," and the term "including" should be interpreted as "including, but not limited to"). Furthermore, it will be understood by those skilled in the art that, if a specific number of introduced claim recitations is intended, such intent is expressly set forth in the claim, and, absent such recitation, no such intent exists. For example, to aid in understanding, the following appended claims may contain the use of the introductory terms "at least one" and "one or more" to introduce the claim recitations. However, the use of such phrases should not be construed as meaning that the introduction of a claim recitation with the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to embodiments containing only one of such recitations, even if the same claim includes the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" or "an" should be construed to mean "at least one" or "one or more"); the same applies to the use of definite articles used to introduce claim recitations. Moreover, even if a specific number of recitations of an introduced claim are explicitly recited, one of ordinary skill in the art will recognize that such recitation should be construed to mean at least the recited number (e.g., the literal recitation of "two recitations" without other modifiers means at least two recitations, or more than two recitations).Furthermore, where a convention similar to "at least one of A, B, and C, etc." is used, such a structure is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, a combination of A and B, a combination of A and C, a combination of B and C, and / or a combination of A, B, and C, etc.). Where a convention similar to "at least one of A, B, or C, etc." is used, such a structure is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having only A, only B, only C, a combination of A and B, a combination of A and C, a combination of B and C, and / or a combination of A, B, and C, etc.). It will be further understood by those skilled in the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms or both terms. Furthermore, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also described in terms of any individual members or subgroups of members of the Markush group.
[0192] As will be understood by one of ordinary skill in the art, for any and all purposes, e.g., in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges thereof, as well as combinations of subranges thereof. Any recited range can be readily recognized as fully descriptive and enabling that same range to be divided into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily divided into a lower third, a middle third, and an upper third, etc. Also, as will be understood by one of ordinary skill in the art, all language such as "up to," "at least," "greater than," "less than," etc., includes the recited numbers and refers to a range that can be substantially divided into the subranges discussed above. Finally, as will be understood by one of ordinary skill in the art, a range includes each individual member. Thus, for example, a group having 1-3 items refers to a group having 1, 2, or 3 items. Similarly, a group having 1-5 items refers to a group having 1, 2, 3, 4, or 5 items, etc.
[0193] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those of ordinary skill in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting. The true scope and spirit are indicated by the following claims.
Claims
1. A composition comprising a first protein complex and a second protein complex, The first protein complex comprises a transposome and a first programmable DNA-binding unit capable of specifically binding to a first binding site on a target double-stranded DNA (dsDNA); the second protein complex comprises a transposome and a second programmable DNA-binding unit capable of specifically binding to a second binding site on the target dsDNA; A composition comprising a transposase and two copies of an adaptor, the transposome comprising:
2. A composition comprising a plurality of protein complex pairs, each of the plurality of protein complex pairs comprising a first protein complex and a second protein complex; a first protein complex comprising a transposome and a first programmable DNA-binding unit capable of specifically binding to a first binding site on a target double-stranded DNA (dsDNA); a second protein complex comprising the transposome and a second programmable DNA-binding unit capable of specifically binding to a second binding site on the target dsDNA; The transposome contains a transposase and two copies of an adapter, the first binding sites for each of the plurality of protein complex pairs are distinct from one another, and / or the second binding sites for each of the plurality of protein complex pairs are distinct from one another; A composition, wherein multiple protein complex pairs all have the same transposome. (i) the target dsDNA for two or more of the multiple protein complex pairs is different; and / or (ii) the plurality of protein complex pairs comprises at least 5 protein complex pairs, and optionally, may comprise about 5-100 protein complex pairs; The composition of claim 2.
4. (i) The adapter comprises: (a) is a dsDNA or a DNA / RNA duplex; and / or (b) is about 5 to 80 base pairs in length; and / or (ii) the transposase is (a) a Tn5 transposase, a Tn7 transposase, a mariner Tc1-like transposase, a Himar1C9 transposase, or a Sleeping Beauty transposase; and / or (b) is a hyperactive transposase; The composition according to any one of claims 1 to 3.
5. a first programmable DNA binding unit comprising: (1) a nuclease-deficient CRISPR associated protein (dCAS protein) and a first guide RNA (gRNA) capable of specifically binding to a first binding site of a target dsDNA, and a second programmable DNA binding unit comprising the dCAS protein and a second gRNA capable of specifically binding to a second binding site of the target dsDNA; Optional (i) the transposome may be linked to the first programmable DNA-binding unit, the second programmable DNA-binding unit, or both, via a linker connecting the transposase and the dCAS protein, and optionally, the linker may comprise a peptide linker, a chemical linker, or both; (ii) the transposase may be present in a fusion protein comprising the dCAS protein of the first programmable DNA-binding unit, the dCAS protein of the second programmable DNA-binding unit, or both; and / or (iii) the dCAS protein may be dCAS9, dCAS12, dCAS13, or dCAS14, and optionally the dCAS13 protein may be dCAS13a, dCAS13b, dCAS13c, or dCAS13d; or (2) a first endonuclease-deficient zinc finger nuclease (ZFN) or a first endonuclease-deficient transcription activator-like effector nuclease (TALEN) capable of specifically binding to a first binding site of a target dsDNA, and a second programmable DNA-binding unit comprises a second endonuclease-deficient ZFN or a second endonuclease-deficient TALEN capable of specifically binding to a second binding site of the target dsDNA; Optional (i) the transposome may be linked to the first programmable DNA-binding unit, the second programmable DNA-binding unit, or both, via a linker connecting the transposase and the ZFN or TALEN, and optionally, the linker may comprise a peptide linker, a chemical linker, or both; and / or (ii) the transposase may be present in a fusion protein comprising the first programmable DNA-binding unit ZFN or TALEN, the second programmable DNA-binding unit ZFN or TALEN, or both; and / or (3) a first endonuclease-deficient meganuclease capable of specifically binding to a first binding site of a target dsDNA, and a second programmable DNA-binding unit comprising a second endonuclease-deficient meganuclease capable of specifically binding to a second binding site of the target dsDNA; Optionally, (i) the transposome may be linked to the first programmable DNA-binding unit, the second programmable DNA-binding unit, or both, via a linker connecting the transposase and the endonuclease-deficient meganuclease, and further optionally, the linker may comprise a peptide linker, a chemical linker, or both; and / or (ii) the transposase may be present in a fusion protein comprising the endonuclease-deficient meganuclease of the first programmable DNA-binding unit, the endonuclease-deficient meganuclease of the second programmable DNA-binding unit, or both; The composition according to any one of claims 1 to 3. (i) the second binding site is 1 to 1000 nucleotides upstream or downstream of the first binding site on the target dsDNA, and optionally the second binding site may be 100 to 500 nucleotides upstream or downstream of the first binding site on the target dsDNA; (ii) the distance between the first binding site and the second binding site on each target dsDNA is substantially the same; and / or (iii) the distance between the first binding site and the second binding site on at least two of the target dsDNAs is different; The composition according to any one of claims 1 to 3.
7. a third protein complex, the third protein complex comprising a transposome and a third programmable DNA-binding unit capable of specifically binding to a third binding site on a target dsDNA; 4. The composition of any one of claims 1 to 3, wherein optionally the third binding site is (i) 1 to 1000 nucleotides upstream or downstream of the first binding site on the target dsDNA, (ii) 1 to 1000 nucleotides upstream or downstream of the second binding site on the target dsDNA, and / or (iii) located between the first binding site on the target dsDNA and the second binding site on the target dsDNA.
8. A composition according to any one of claims 1 to 3, a sample nucleic acid suspected of containing a target dsDNA; A DNA polymerase; Multiple dNTPs and A reaction mixture comprising: (i) a reaction mixture comprising one or more oligonucleotide probes, a buffer, and MgCl 2 Further comprising: (ii) the adaptor is covalently attached to the target dsDNA or a fragment thereof; (iii) the reaction mixture comprises a plurality of dsDNA fragments containing adapters at both ends; (iv) the sample nucleic acid comprises bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or a combination thereof; (v) the target dsDNA is genomic DNA, mitochondrial DNA, plasmid DNA, or a combination thereof; and / or (vi) the sample nucleic acid is derived from a biological sample, optionally the biological sample may comprise stool, sputum, peripheral blood, plasma, serum, lymph node, respiratory tissue, exudate, or a combination thereof; The reaction mixture of claim 8.
10. 1. A method for simultaneously detecting multiple target nucleic acids, comprising: contacting a sample nucleic acid suspected of containing a plurality of target double-stranded DNAs (dsDNAs) with a plurality of pairs of protein complexes to form a reaction mixture; Each of the plurality of target dsDNAs comprises a target sequence adjacent to a first binding site on the target dsDNA and a second binding site on the target dsDNA; Each of the protein complex pairs comprises a first protein complex and a second protein complex, The first complex comprises a transposome and a first programmable DNA-binding unit capable of specifically binding to a first binding site on a target dsDNA; the second complex comprises a transposome and a second programmable DNA-binding unit capable of specifically binding to a second binding site on the target dsDNA; The transposome contains a transposase and two copies of an adapter. the first binding sites for each of the plurality of pairs of protein complexes are distinct from one another, the second binding sites for each of the plurality of pairs of protein complexes are distinct from one another, or both; The plurality of pairs of protein complexes all comprise the same transposome; incubating the reaction mixture to generate a plurality of dsDNA fragments each containing an adaptor and a target sequence on both ends; amplifying a plurality of dsDNA fragments using a primer capable of binding to one strand of the adapter to generate an amplification product; detecting the presence of the target sequence in the amplified product as an indication of the presence of the plurality of target dsDNAs; The method includes: (i) the second binding site is about 1 to 1000 base pairs downstream of the first binding site; (ii) the adapter is (a) a dsDNA or a DNA / RNA duplex, and / or (b) is about 5-80 base pairs in length; (iii) the primer is about 5 to 80 nucleotides in length; (iv) a plurality of target dsDNAs, (a) contains genomic DNA, mitochondrial DNA, plasmid DNA, or a combination thereof; (b) derived from one or more organisms, one or more genes, or a combination thereof; (c) contains bacterial DNA, viral DNA, fungal DNA, protozoan DNA, or a combination thereof; and / or (d) containing genomic DNA from at least two different organisms and / or from at least five different genes; (v) the sample nucleic acid is derived from a biological sample, optionally the biological sample may comprise stool, sputum, peripheral blood, plasma, serum, lymph node, respiratory tissue, exudate, or a combination thereof; and / or (vi) the transposase is a Tn5 transposase, a Tn7 transposase, a mariner Tc1-like transposase, a Himar1C9 transposase, or a Sleeping Beauty transposase; The method of claim 10.
12. The method of claim 1, further comprising generating a plurality of target dsDNAs from a plurality of target RNAs using a reverse transcriptase. (ii) detecting the presence of a target sequence in the amplified products comprises contacting each amplified product with an oligonucleotide probe capable of specifically binding to a target sequence; (iii) contacting the plurality of target dsDNAs with the plurality of protein complex pairs is performed at about 25° C. to about 80° C.; and / or (iv) incubating the reaction mixture comprises incubating the reaction mixture at about 37° C. to about 55° C.; The method according to any one of claims 10 to 11. (a) amplifying the plurality of dsDNA fragments using primers is carried out using polymerase chain reaction (PCR); Optionally, (i) PCR may be loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), recombinant polymerase amplification (RPA), strand displacement amplification (SDA), nucleic acid sequence-based amplification (NASBA), transcription-mediated amplification (TMA), nicking enzyme amplification reaction (NEAR), rolling circle amplification (RCA), multiple displacement amplification (MDA), branching amplification (RAM), circular helicase-dependent amplification (cHDA), single primer isothermal amplification (SPIA), signal-mediated amplification of RNA technology (SMART), self-sustained sequence replication (3SR), genomic exponential amplification reaction (GEAR), or isothermal multiple displacement amplification (IMDA); and / or (ii) PCR may be real-time PCR or quantitative real-time PCR (QRT-PCR); and / or (b) the step of amplifying the plurality of dsDNA fragments does not use any primer other than a primer capable of binding to one strand of the adapter; The method according to any one of claims 10 to 11. (i) labeling one or both ends of one or more of the plurality of dsDNA fragments; and / or (ii) differentially labeling two ends of one or more of the plurality of dsDNA fragments; The method of any one of claims 10 to 11, wherein the label comprises an anionic label, a cationic label, a neutral label, an electrochemical label, a protein label, a fluorescent label, a magnetic label, or a label with a combination thereof.
15. The first programmable DNA-binding unit comprises a nuclease-deficient CRISPR-associated protein (dCAS protein) and a first guide RNA (gRNA) capable of specifically binding to a first binding site of a target dsDNA, and the second programmable DNA-binding unit comprises a dCAS protein and a second gRNA capable of specifically binding to a second binding site of a target dsDNA; Optionally, (i) the transposome may be linked to the first programmable DNA-binding unit, the second programmable DNA-binding unit, or both, via a linker connecting the transposase and the dCAS protein, and optionally, the linker may comprise a peptide linker, a chemical linker, or both; (ii) the transposase may be present in a fusion protein comprising the dCAS protein of the first programmable DNA-binding unit, the dCAS protein of the second programmable DNA-binding unit, or both; and / or (iii) the dCAS protein may be dCAS9, dCAS12, dCAS13, or dCAS14; The method according to any one of claims 10 to 11.