Methods and compositions for sequencing library normalization
By contacting the polynucleotide sequencing library with ribonucleoprotein and using its binding to the adaptor sequence to extract the polynucleotide, the problem of inaccurate and long time-consuming normalization of nucleic acid concentrations between polynucleotide sequencing libraries in the prior art is solved, and efficient and accurate normalization of nucleic acid concentrations is achieved.
Patent Information
- Application Number
- CN202380074242.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-27
- Filing Date
- 2023-10-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems of inaccuracy and long-term use when normalizing the nucleic acid concentrations between polynucleotide sequencing libraries.
By contacting the polynucleotide sequencing library with a predetermined amount of ribonucleoprotein, the ribonucleoprotein is used to bind to the adaptor sequence in the library and the polynucleotide containing the adaptor sequence is extracted to achieve normalization of the nucleic acid concentration.
This method can effectively normalize the polynucleotide concentration between polynucleotide sequencing libraries, improving the accuracy and efficiency of sequencing results.
Smart Images

Figure CN120077129A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 380,488, filed on October 21, 2022, titled "Methods for Sequencing Library Normalization" and U.S. Provisional Application No. 63 / 516,033, filed on July 27, 2023, titled "Methods and Compositions for Sequencing Library Normalization" under 35 U.S.C. § 119(e), the entire contents of each of which are incorporated herein by reference.
[0003] Reference to Electronic Sequence Listing
[0004] The content of the electronic sequence listing (W109470000WO00-SEQ-ARM.xml; size: 53,547 bytes; and creation date: October 20, 2023) is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION
[0005] Massively parallel deep sequencing, also known as next-generation sequencing (NGS), enables large-scale DNA and RNA sequencing. These methods involve parallel multiplex analysis of large-scale nucleic acid sequences, allowing millions to billions of sequences from individual strands to be analyzed separately but simultaneously.
[0006] Multiplex analysis can be challenging because the amount of nucleic acid present in different samples to be analyzed is usually different, which can result in inaccuracies where low-concentration samples are under-sequenced and high-concentration samples are over-sequenced. Therefore, various methods have been adopted to normalize the nucleic acid concentration between samples. Spectrophotometry, electrophoresis, fluorometry, and quantitative PCR (qPCR) have been used to detect the nucleic acid concentration in samples in order to normalize the concentration between samples. Each of these methods has drawbacks, such as inaccuracies introduced by manual adjustments, lack of sensitivity, and multiple steps involving a large amount of time and / or cost. SUMMARY OF THE INVENTION
[0007] In some aspects, the present disclosure describes methods and compositions for normalizing the concentration of polynucleotides between two or more polynucleotide sequencing libraries. Polynucleotide preparation for sequencing (e.g., next-generation sequencing (NGS)) generally involves adding adapter polynucleotide sequences to the polynucleotides. Certain techniques (such as ILLUMINA sequencing) use adapter sequences to sequence polynucleotides. It is recognized herein that these adapter sequences can be utilized to normalize the concentration of polynucleotides between different polynucleotide sequencing libraries. Specifically, it is recognized that polynucleotide sequencing library normalization can be achieved by contacting a polynucleotide sequencing library with a predetermined amount of ribonucleoprotein (e.g., a catalytically inactive CRISPR-associated protein (dCAS)-guide RNA (gRNA) complex), which binds to the adapter sequence of the polynucleotide library, and extracting the polynucleotides that (1) contain the adapter and (2) are bound by the ribonucleoprotein complex. It has been demonstrated that using this method for multiple different polynucleotide sequencing libraries normalizes the concentration of polynucleotides between the polynucleotide sequencing libraries.
[0008] Accordingly, in some aspects, the present disclosure provides a method for normalizing the concentration of target polynucleotides between at least two samples, each of the at least two samples comprising a target polynucleotide, the method comprising, for each of the at least two samples:
[0009] (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence;
[0010] (ii) generating a solution, the generating comprising combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adapter sequence of the target polynucleotide of the sample; and (c) a dCas or dArgonaute at a predetermined concentration, the dCas or the dArgonaute comprising an affinity tag, wherein the predetermined concentration of the dCas or the dArgonaute is homologous to the guide polynucleotide;
[0011] (iii) contacting the solution with a solid phase comprising an affinity tag-binding molecule capable of binding to the affinity tag;
[0012] (iv) separating the solution from the solid phase; and
[0013] (v) Extract the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples. In some embodiments, the targeting region is complementary to the static region of the adaptor sequence. In some embodiments, the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide. In some embodiments, the static region is the static region of a next-generation sequencing adaptor sequence. In some embodiments, the targeting region is complementary to the adaptor sequence shown in any one of SEQ ID NOs: 27-36. In some embodiments, the targeting region is complementary to the adaptor sequence shown in any one of SEQ ID NOs: 27-28. In some embodiments, the targeting region is complementary to the adaptor sequence shown in any one of SEQ ID NOs: 29-30. In some embodiments, the targeting region is complementary to the adaptor sequence shown in any one of SEQ ID NOs: 31-32. In some embodiments, the targeting region is complementary to the adaptor sequence shown in any one of SEQ ID NOs: 33-34. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 35. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 36.
[0014] In some embodiments, the Cas gRNA polynucleotide targeting region is a homologous region. In some embodiments, the Cas gRNA polynucleotide comprises the homologous region shown in any one of SEQ ID NOs: 1-26. In some embodiments, the Cas gRNA polynucleotide comprises the homologous region shown in SEQ ID NO: 1.
[0015] In some embodiments, the Argonaute guide polynucleotide is siRNA, miRNA, piRNA, shRNA, or siDNA.
[0016] In some embodiments, the guide polynucleotide is a DNA polynucleotide. In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide comprises a modified nucleic acid. In some embodiments, the modified nucleic acid comprises 2'F RNA, 2'OMe RNA, and / or phosphorothioate bond (PS).
[0017] In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not contain a modified nucleic acid at the position where it interacts with the Cas protein, and the Cas protein and the gRNA are homologous. In some embodiments, the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX or CasY protein. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO:37.
[0018] In some embodiments, the dArgonaute is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo.
[0019] In some embodiments, the affinity tag-binding molecule comprises Ni2+ and the affinity tag comprises a His tag.
[0020] In some embodiments, the affinity tag-binding molecule comprises biotin and the affinity tag comprises avidin. In some embodiments, the affinity tag-binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag. In some embodiments, the affinity tag and the corresponding affinity tag-binding molecule are selected from Table 1.
[0021] In some embodiments, the solid phase comprises magnetic beads. In some embodiments, separating the solution from the solid phase comprises immobilizing the solid phase and washing the solid phase. In some embodiments, extracting the target polynucleotide from the solid phase comprises combining the solid phase and a protease. In some embodiments, extracting the target polynucleotide from the solid phase comprises combining the solid phase and proteinase K in a solution sufficient to produce proteinase K activity. In some embodiments, proteinase K digests the dCas or dArgonaute bound to the solid phase, thereby extracting the target polynucleotide.
[0022] In some embodiments, steps (i)-(v) are performed in sequence.
[0023] In some embodiments, normalization comprises bringing the concentration of the target polynucleotide between the at least two samples within 15% of each other after normalization.
[0024] In some aspects, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples, each of the at least two samples comprising a target polynucleotide, and for each of the at least two samples, the method comprises:
[0025] (i) Obtain the sample, wherein the target polynucleotide of the sample comprises an adaptor sequence;
[0026] (ii) Generate a solution, the generating including combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adaptor sequence of the target polynucleotide of the sample, and (c) dCas or dArgonaute at a predetermined concentration, the dCas or the dArgonaute comprising an affinity tag; and
[0027] (iii) Contact the solution with a predetermined amount of catalytically active Cas or Argonaute protein to normalize the concentration of the target polynucleotide between the two or more samples.
[0028] In some embodiments, the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY protein. In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO:37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO:37.
[0029] In some embodiments, the dArgonaute is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo.
[0030] In some embodiments, step (ii) includes the binding of the dCas or dArgonaute to the adaptor sequence of at least some of the target polynucleotides in the target polynucleotide. In some embodiments, the catalytically active Cas9 or Argonaute digests the target polynucleotides not bound by the dCas9 or the dArgonaute. In some embodiments, normalization includes bringing the concentration of the target polynucleotide between the at least two samples within 15% of each other after normalization. In some embodiments, normalization includes bringing the concentration of the target polynucleotide between the at least two samples within 10% of each other after normalization. In some embodiments, the target polynucleotide comprises a first adaptor sequence and a second adaptor sequence, and wherein the guide polynucleotide targeting region is complementary to the first adaptor sequence.
[0031] In some embodiments, the method further comprises, after step (i) and before step (ii):
[0032] (a) contacting the sample with a primer encoding a nucleic acid sequence complementary to a portion of the second adaptor sequence, wherein the portion of the adaptor sequence is located in the proximal portion of the adaptor sequence; and
[0033] (b) contacting the sample with a DNA polymerase under conditions sufficient to promote primer extension.
[0034] In some embodiments, the portion of the adaptor sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34.
[0035] In some embodiments, the second adaptor sequence is a 3' adaptor sequence. In some embodiments, the second adaptor sequence is a 5' adaptor sequence.
[0036] In some aspects, the present disclosure provides a guide polynucleotide comprising a targeting region that is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 28 and 30, 32, 34 - 36. In some embodiments, the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
[0037] In some embodiments, the static region is the static region of a next-generation sequencing adaptor sequence. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 28. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 30. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 32. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 34. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 35. In some embodiments, the targeting region is complementary to the adaptor sequence shown in SEQ ID NO: 36.
[0038] In some embodiments, the Cas gRNA polynucleotide targeting region is a homologous region. In some embodiments, the Cas gRNA polynucleotide comprises a homologous region shown in any one of SEQ ID NOs: 1 - 26. In some embodiments, the Cas gRNA polynucleotide comprises the homologous region shown in SEQ ID NO: 1.
[0039] In some embodiments, the Argonaute guide polynucleotide is siRNA, miRNA, piRNA, shRNA, or siDNA.
[0040] In some embodiments, the guide polynucleotide is a DNA polynucleotide. In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide comprises a modified nucleic acid. In some embodiments, the modified nucleic acid comprises 2′F RNA, 2′Ome RNA, and / or phosphorothioate bond (PS). In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not comprise a modified nucleic acid at a position that interacts with a Cas protein, and the Cas protein and the gRNA are homologous.
[0041] In some embodiments, the present disclosure provides herein a kit comprising (a) a guide polynucleotide as described herein, and (b) a Cas protein or an Argonaute protein, or a polynucleotide sequence encoding a Cas protein or an Argonaute protein. In some embodiments, the Cas protein is a catalytically inactive Cas protein (dCas). In some embodiments, the Cas protein is a catalytically inactive Cas9 protein (dCas9). In some embodiments, the Cas9 protein comprises D10A and H840A mutations. In some embodiments, the Cas protein is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, or CasY protein. In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO: 37.
[0042] In some embodiments, the Argonaute protein is catalytically inactive (dArgonaute). In some embodiments, the Argonaute protein is a catalytically inactive CbAgo protein. In some embodiments, the Argonaute protein is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo protein. In some embodiments, the dArgonaute protein comprises an amino acid sequence having at least 95% identity to any one of SEQ ID NOs: 39-45. In some embodiments, the dArgonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 39-45.
[0043] In some embodiments, the kit further comprises a catalytically active Cas protein or a catalytically active Argonaute protein. In some embodiments, the catalytically active Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY. In some embodiments, the catalytically active Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo. In some embodiments, the guide polynucleotide is capable of binding to the Cas protein or the Argonaute protein to form a ribonucleoprotein complex, and the ribonucleoprotein complex is capable of binding to the adaptor sequence.
[0044] In some embodiments, the kit further comprises a primer complementary to a portion of the adaptor sequence. In some embodiments, the portion of the adaptor sequence is located in the proximal portion of the adaptor sequence. In some embodiments, the portion of the adaptor sequence is located in the proximal portion of the adaptor sequence, and the adaptor sequence comprises the nucleic acid sequence of any one of SEQ ID NO: 28, 30, 32 and 34. In some embodiments, the adaptor sequence is a 3' adaptor sequence. In some embodiments, the adaptor sequence is a 5' adaptor sequence. In some embodiments, the primer is not complementary to the adaptor sequence, and the adaptor sequence is complementary to the guide polynucleotide targeting region.
[0045] In some aspects, the present disclosure provides a reaction mixture, the reaction mixture comprising: (i) a plurality of target polynucleotides, wherein the target polynucleotide comprises an adaptor sequence shown in any one of SEQ ID NO: 28, 30, 32 and 34-36; (ii) a predetermined concentration of a guide polynucleotide, the guide polynucleotide comprising a targeting region complementary to the adaptor sequence; and (iii) a predetermined concentration of a dCas protein or a dArgonaute protein.
[0046] In some embodiments, the predetermined concentration of the guide polynucleotide or the dCas protein or the dArgonaute protein is lower than the concentration of the target polynucleotide in the reaction mixture.
[0047] In some embodiments, the dArgonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo or a catalytically inactivated variant thereof, and the guide polynucleotide is homologous to the dArgonaute protein. In some embodiments, the dCas protein comprises catalytically inactivated Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY or a variant thereof, and the guide polynucleotide is homologous to the dCas protein. In some embodiments, the linker sequence comprises the polynucleotide sequence of any one of SEQ ID NO: 28, 30, 32, and 34-36. In some embodiments, the reaction mixture further comprises a catalytically active Cas protein or a catalytically active Argonaute protein.
[0048] In some embodiments, the present disclosure provides a ribonucleoprotein (RNP) complex, the RNP complex comprising: (i) a guide polynucleotide as described herein; (ii) a Cas protein or an Argonaute protein, the Cas protein or the Argonaute protein being homologous to the guide polynucleotide; and (iii) a linker sequence, wherein the targeting region of the guide polynucleotide is complementary to the linker sequence, and wherein the guide polynucleotide is homologous to the Cas protein or the Argonaute protein. In some embodiments, the Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo or a catalytically inactivated variant thereof. In some embodiments, the Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY or a catalytically inactivated variant thereof. In some embodiments, the linker sequence comprises the polynucleotide sequence of any one of SEQ ID NO: 28, 30, 32, and 34-36. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Shows a representative schematic diagram of the polynucleotide capture process by binding of CAS-gRNA RNP to metal beads.
[0050] Figure 2Normalization of two different samples each containing a polynucleotide with an i5 - adapter is shown. Two different samples of i5 - containing DNA fragments were normalized with the same mass (125 femtomoles, fmol) of dCas9. After elution following protease digestion, both samples contained an equal number of molecules, demonstrating the concept of normalization. As expected, the molecular ratio between the samples changed from 2:1 to 1:1. If these samples were pooled using equal volumes and sequences by NGS, it was expected that both samples would generate an equal number of clusters.
[0051] Figure 3 The ability to regulate the amount of polynucleotide extracted from a sample by titrating the amount of ribonucleoprotein that binds to the adapter of the polynucleotide (e.g., RNP) added to the sample is shown. Samples containing a constant amount of DNA (160 fmol) were titrated with RNP (dCas9 - gRNA complex) followed by bead binding and elution. Increasing the amount of RNP caused a linear increase in the amount of DNA bound (and ultimately recovered). At the 160 fmol DNA input level, there was a linear relationship for RNP in the range of 125 femtomoles to 1000 fmol. These data demonstrate that the amount of the extracted DNA library can be regulated by adjusting the amount of RNP added to the sample.
[0052] Figure 4 Specific and stoichiometric targeting and retention of target DNA molecules driven by interaction with dCas9 RNP and subsequent binding of 6His - tagged dCas9 to HisPur beads is shown. Known amounts of DNA fragments containing the i5 - sequence were bound to HisPur beads after contact with RNP (RNP+) or without RNP (RNP -). Most of the library remained bead - bound only when exposed to RNP and eluted after proteinase K digestion (comparing RNP+ flow - through with RNP+ eluate). In the absence of RNP, all DNA was found in the flow - through and none was bead - bound (comparing RNP - flow - through with RNP - eluate).
[0053] Figure 5 Normalization of four different samples of a PCR - amplified ILLUMINA library using the dCas9 - gRNA complex is shown. Four different samples of a PCR - amplified ILLUMINA library (input) were normalized with the same mass of dCas9. After elution following protease digestion (elution), all samples contained a significantly more uniform number of molecules. The highest - concentration sample was 25% more concentrated than the lowest - concentration sample after normalization (4.68 nM eluate vs. 5.89 nM eluate) compared to 150% before normalization (4 nM input vs. 10 nM input). Additionally, most of the unbound library remained in the flow - through of all libraries.
[0054] Figure 6 A schematic diagram showing an exemplary embodiment of the disclosed normalization method is presented. In this exemplary embodiment, two samples are provided, which include end-labeled target nucleic acids at different concentrations (library A at a concentration of 9 fmol and library B at a concentration of 18 fmol). In this exemplary embodiment, a reaction mixture is produced by adding a predetermined concentration of 3 fmol of guide polynucleotide and catalytically inactivated Cas9 protein (which may be pre-assembled) to library A and library B and allowing the protein to bind to the adapter sequence at one end of the end-labeled target nucleic acid in the sample, thereby protecting the ends of the target nucleic acid from nuclease digestion. An exonuclease is added to the reaction mixture and digests the unprotected nucleic acids. The resulting library A and library B each contain 3 fmol of target nucleic acid. In other embodiments, the exonuclease in this schematic diagram can be replaced with a nucleic acid-guided nuclease (e.g., Cas9 nuclease), a nucleic acid-guided nickase (e.g., Cas9 nickase), a restriction enzyme nuclease or nickase, or a transcription activator-like effector nuclease (TALEN) or another similar nuclease or nickase.
[0055] Figures 7A - 7D A schematic diagram showing the denaturation of library molecules and partial sequence extension to generate double-stranded PAM sites is presented. In Figure 7A this, the library molecules are ligated to two Y-shaped adapters at the 3'-end and 5'-end. In Figure 7B this, the library molecules are denatured and annealed to a partial primer at the 3'-end. In Figure 7C this, the annealed primer is extended with a polymerase to generate a complementary strand (dashed line), which does not span the entire 3'-adapter but does contain a double-stranded, full-length 5'-adapter, and the complementary strand contains a PAM site for dCas9 targeting. Figure 7D The binding of the dCas9 molecule to the double-stranded PAM site is shown.
[0056] Figure 8 is a bar graph showing the femtomoles (fmol output) of DNA captured after normalization and amplification for different amounts of starting DNA (fmol input) and two concentrations of ribonucleoprotein (RNP fmol).
[0057] Figure 9 A graph showing the relationship between the fmol (fmol output) of captured DNA and different concentrations of ribonucleoprotein (RNP fmol) is presented. The concentration of the input DNA is shown as 3,400 fmol.
[0058] Figure 10Shows the relationship between the fmol of captured DNA (fmol output) and different concentrations of ribonucleoprotein (RNP fmol) and single-guide (bottom line) or duplex-guide (top line). Detailed Description
[0059] Guide polynucleotide
[0060] In some aspects, the present disclosure describes a guide polynucleotide comprising a targeting region complementary to an adaptor sequence.
[0061] "Polynucleotide" refers to a polymer of nucleotides (e.g., deoxyribonucleotides, ribonucleotides, and modified nucleotides). The polymer can be in single-stranded or double-stranded form. Unless otherwise specified, the term encompasses polynucleotides containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have binding properties similar to reference nucleic acids, and which are metabolized in a manner similar to naturally occurring nucleotides. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidates, methylphosphonates, chiral methylphosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs).
[0062] "Guide polynucleotide" refers to a polynucleotide that (1) includes a targeting region complementary to a target (e.g., an adaptor sequence) and (2) is capable of facilitating binding of a ribonucleoprotein complex comprising the guide polynucleotide to the target. In some embodiments, the guide polynucleotide is a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide. In some embodiments, the Cas gRNA polynucleotide refers to a two polynucleotide system comprising a tracrRNA and a crRNA, e.g., as described in Karvelis T et al., RNA Biol. May 2013;10(5):841-51. PMID:23535272. The tracrRNA includes a sequence encoding a stem-loop structure that associates with a Cas protein (e.g., dCas9). The crRNA includes a homologous region complementary to the target (e.g., an adaptor sequence) and a region complementary to the tracrRNA. The crRNA and tracrRNA can form a complex with a Cas protein, which in turn binds to a target DNA or RNA polynucleotide (depending on the Cas type). In some embodiments, the Cas gRNA polynucleotide refers to a single guide RNA (sgRNA) polynucleotide, e.g., as described in Jinek M et al., Science. Aug 17, 2012;337(6096):816-21. doi:10.1126 / science.1225829. The sgRNA includes a sequence encoding a stem-loop structure that associates with a Cas protein and a homologous region. Many Cas proteins bind to a polynucleotide (e.g., double-stranded DNA) at a position adjacent to a protospacer adjacent motif (PAM) site. In some embodiments, the homologous region is complementary to a portion of the adaptor sequence adjacent to the PAM site (e.g., sufficiently adjacent such that the Cas can bind to the adaptor sequence). Methods for preparing and using gRNAs are known in the art, e.g., as described in Mohr SE et al., FEBS J. Sep 2016;283(17):3232-8. PMCID:PMC5014588.
[0063] An Argonaute-guide polynucleotide refers to a polynucleotide encoding small interfering RNA (siRNA), microRNA (miRNA), P-element induced wimpy testis (PIWI)-interacting RNA (piRNA), and small interfering DNA (siDNA), for example, as described in Wu J et al., Journal of Advanced Research (J Adv Res.) April 29, 2020; 24:317-324, PMID: 32455006. In some embodiments, the Argonaute-guide polynucleotide comprises a targeting region complementary to an adaptor sequence as described herein. In some embodiments, the guide polynucleotide is single-stranded (e.g., an RNA Cas gRNA polynucleotide). In some embodiments, the guide polynucleotide is double-stranded (e.g., a double-stranded DNA encoding a Cas gRNA polynucleotide, or a dsRNA using an Argonaute-guide polynucleotide).
[0064] In some embodiments, the guide polynucleotide is an RNA molecule. In some embodiments, the guide polynucleotide is a DNA molecule. In some embodiments, the guide polynucleotide comprises modified nucleic acids (e.g., 2′F RNA, 2′OMeRNA, and / or phosphorothioate bonds (PS)). The nucleic acids in the guide polynucleotide (e.g., an RNA guide polynucleotide such as a Cas guide RNA) can be modified to increase the stability of the guide polynucleotide (e.g., reduce nuclease digestion). A Cas gRNA (e.g., a Cas9 gRNA) can comprise modified nucleic acids (e.g., 2′OMe) in the gRNA region that does not interact with the Cas protein (e.g., Cas9) and still maintain the ability to direct the Cas protein to the target, as described in Yin H et al., Nature Biotechnology 2017 Dec; 35(12):1179-1187. In some embodiments, the Cas guide RNA polynucleotide comprises a modification of e-sgRNA, as described in Yin et al. 2017. In some embodiments, the Cas9 sgRNA does not comprise modified nucleic acids (e.g., 2′OH modification) at one or more positions counted 22-27, 43-45, 47, 49, 51, 58, 59, 62, 63-65, 68-69, or 82 from the 5′ end of the sgRNA (i.e., the sgRNA referring to (GGGCGAGGAGCUGUUCACCGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAG GCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU, SEQ ID NO:46)).
[0065] In some embodiments, the sgRNA comprises 2'-O-methyl RNA bases. In 2'-O-methyl RNA bases, the 2'-hydroxyl (-OH) of the ribose in the RNA molecule is replaced by a methyl group (-CH3). This modification affects the ribose, and the nitrogenous bases (adenine, guanine, cytosine, uracil) themselves are generally not modified in this context. The 2'-O-methyl RNA bases are 2'O-methyladenosine, 2'O-methylguanosine, 2'O-methylcytidine, and 2'O-methyluridine. In some embodiments, the sgRNA comprises the following sequence:
[0066] 5'-mA*mG*mA*rUrCrG rGrArA rGrArG rCrGrU rCrGrU rGrUrG rUrUrU rUrArG rArArA rUrArG rCrArA rGrUrU rArArA rArUrA rArGrG rCrUrA rGrUrCrCrGrU rUrArU rCrArA rCrUrU rGrArA rArArA rGrUrG rGrCrA rCrCrG rArGrU rCrGrGrUrGrC mU*mU*mU*rU-3'(SEQ ID NO:48)
[0067] where mA and mG are 2'-O-methyladenosine and 2'O-methylguanosine RNA bases.
[0068] The "targeting region" of the guide polynucleotide refers to the region of the guide polynucleotide that is complementary to the target polynucleotide sequence (e.g., the adapter sequence). In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide comprising a homologous region. The homologous region of the CasgRNA polynucleotide can comprise a series of consecutive amino acids that are complementary to the target polynucleotide sequence (e.g., the adapter sequence). In some embodiments, the length of the homologous region is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides. In some embodiments, the length of the homologous region is 18-26 nucleotides. In some embodiments, the length of the homologous region is 17-24 nucleotides. In some embodiments, the length of the homologous region is 18-22 nucleotides. In some embodiments, the length of the homologous region is 20 nucleotides.
[0069] In some embodiments, the guide polynucleotide is an Argonaute guide polynucleotide. The Argonaute guide polynucleotide can be an RNA interference (RNAi) polynucleotide. In some embodiments, the Argonaute guide polynucleotide comprises a targeting region of small interfering RNA (siRNA), microRNA (miRNA), P-element induced wimpy testis (PIWI)-interacting RNA (piRNA), and small interfering DNA (siDNA), e.g., as described in Wu J et al., Journal of Frontiers in Research, Apr 29, 2020; 24:317-324, PMID: 32455006.
[0070] In some embodiments, the targeting region is complementary to an adaptor sequence as described herein. In some embodiments, the targeting region is complementary to a static region of the adaptor sequence (e.g., generally does not change between adaptors). In some embodiments, the targeting region is complementary to a sequence comprising a region added to the sequence for the purpose of being complementary to the guide polynucleotide targeting region. For example, for the purpose of being complementary to the guide polynucleotide targeting region (e.g., for library normalization), additional sequences can be added to the adaptor sequence. Such additional sequences can be sequences that are specifically bound by the Cas protein-gRNA complex or the Argonaute guide polynucleotide complex as compared to other sequences in the polynucleotide sequencing library.
[0071] In some embodiments, the targeting region is complementary to a next-generation sequencing adapter sequence. In some embodiments, the static region is the static region of a next-generation sequencing adapter sequence. In some embodiments, the targeting region is complementary to a p5 adapter or a p7 adapter. In some embodiments, the p5 adapter and the p7 adapter are ILLUMINA sequencing adapters. In some embodiments, the targeting region is complementary to the i5 polynucleotide sequence of SEQ ID NO:27 or SEQ ID NO:28. In some embodiments, the targeting region is complementary to the i7 adapter sequence shown in SEQ ID NO:29 or SEQ ID NO:30. In some embodiments, the targeting region is complementary to a NEXTERA read 1 adapter (e.g., SEQ ID NO:31 or SEQ ID NO:32) or a NEXTERA read 2 adapter (e.g., SEQ ID NO:33 or SEQ ID NO:34). In some embodiments, the targeting region is complementary to an ION TORRENT A adapter or an ION TORRENT P1 adapter. In some embodiments, the targeting region is complementary to the ION TORRENT A adapter shown in SEQ ID NO:35. In some embodiments, the targeting region is complementary to the ION TORRENT P1 adapter shown in SEQ ID NO:36. In some embodiments, the gRNA polynucleotide comprises a homology region shown in any of SEQ ID NOs: 1-26. In some embodiments, the gRNA polynucleotide comprises the homology region shown in SEQ ID NO:1.
[0072] As used herein, the term "complementary" refers to the degree of expected Watson-Crick base pairing between a first polynucleotide (e.g., a guide polynucleotide targeting region) and a second polynucleotide (e.g., an adaptor sequence). Complementary nucleotides are generally A and T (or A and U) and G and C. In some embodiments, complementary refers to 100% complementarity between two polynucleotides (e.g., a guide polynucleotide targeting region and an adaptor sequence). In some embodiments, complementarity refers to 70%, 80%, 90%, 95% or 99% complementarity between two polynucleotides. For example, a targeting region may be complementary to an adaptor sequence when one, two or three nucleic acids in the targeting region are not complementary to the adaptor sequence. In some embodiments, complementary refers to a sufficient degree of Watson-Crick base pairing between an RNP (e.g., dCas9 bound to a guide RNA) guide polynucleotide targeting region and an adaptor sequence that is for binding of the RNP to the adaptor sequence. In some embodiments, complementary refers to a sufficient degree of Watson-Crick base pairing between an Argonaute guide polynucleotide (e.g., miRNA, siRNA, pwRNA, shRNA or siDNA) and a target sequence (e.g., an adaptor sequence) to direct binding of an Argonaute protein comprising the Argonaute guide polynucleotide to the target.
[0073] "Adaptor sequence" refers to a polynucleotide that is added (e.g., ligated) to the ends (e.g., 3' and / or 5') of a target polynucleotide for polynucleotide sequencing. In some embodiments, an adaptor sequence refers to a nucleic acid sequence encoding an adaptor. In some embodiments, an adaptor sequence can be used for sequencing using a specific sequencing platform. For example, ILLUMINA i5 / p5 and i7 / p7 adaptor sequences can be used for sequencing using the ILLUMINA sequencing platform. In some embodiments, an adaptor sequence includes a static region (e.g., generally does not change between adaptors) and a dynamic region (e.g., can change between adaptors). In some embodiments, an adaptor sequence includes a distal region (generally static), an index region (generally dynamic) and a proximal region (generally static). For example, an ILLUMINA adaptor sequence can include a static p5 region, a dynamic index region and a static i5 region.
[0074] In some embodiments, the guide polynucleotide targeting region is complementary to the static region of the adaptor sequence. In some embodiments, the guide polynucleotide targeting region is complementary to the dynamic region of the adaptor sequence. In some embodiments, the guide polynucleotide comprises a targeting region that is complementary to the static region (e.g., the distal or proximal region of the adaptor sequence) of the adaptor sequence. In some embodiments, the guide polynucleotide comprises a targeting region that is complementary to the dynamic region (e.g., an index) of the adaptor sequence. In some embodiments, at least some of the target polynucleotides in the target polynucleotides comprise the same static adaptor sequence (e.g., the same distal region and / or proximal region). In some embodiments, most of the target polynucleotides comprise the same static adaptor sequence. In some embodiments, all of the target polynucleotides comprise the same static adaptor sequence.
[0075] In some embodiments, the adaptor sequence is modified to comprise a Cas protein protospacer adjacent motif (PAM) or the reverse complement of a PAM. A "protospacer adjacent motif (PAM)" is a nucleotide motif (usually 3 contiguous nucleotides in a polynucleotide) required for Cas protein binding to a target polynucleotide. The canonical Cas9 PAM sequence is 5'-NGG, where N is any nucleotide A, G, C, or T. In some embodiments, the adaptor sequence is modified to comprise a PAM or the reverse complement of a PAM to facilitate binding of the Cas-gRNA complex to the target polynucleotide. For example, a PAM sequence can be added to the 3' end or the 5' end of the adaptor sequence. In some embodiments, the distal sequence can be modified to comprise a protospacer adjacent motif (PAM) or the reverse complement thereof.
[0076] In some embodiments, the adapter sequence is an adapter sequence for sequencing by ILLUMINA, ION TORRENT, PACBIO, ELEMENT, ULTIMA, OMNIOME, SINGULAR, or MGI. In some embodiments, the adapter sequence is an adapter sequence from any of the following kits: WATCHMAKER DNA Library Preparation Kit with Fragmentation or WATCHMAKER RNA Library Preparation Kit with Polaris Depletion, ILLUMINA TruSeq PCR-Free Library Preparation Kit, TruSeq Nano DNA Library Preparation Kit, NEXTERA DNA Library Preparation Kit, NEXTERA DNA XT Library Preparation Kit, NEXTERA Rapid Capture Exome Kit, NEXTERA Rapid Capture Expanded Exome Kit, AmpliSeq for ILLUMINA Library Preparation Kit, ILLUMINA RNA Preparation and Enrichment Kit, ILLUMINA Stranded mRNA Preparation Kit, TruSeq RNA Library Preparation Kit, TruSeq Stranded Total RNA Kit, TruSeq Stranded mRNA Kit, TruSeq Small RNA Kit. In some embodiments, the adapter sequence is an adapter sequence for any of the following kits: NEBNEXT Rapid DNA for ION TORRENT, NEBNEXT Rapid DNA Fragmentation and Library Preparation Set for ION TORRENT, THERMOFISHER Precision ID Library Kit, THERMOFISHER Ion Xpress Plus Fragment Library Kit, THERMOFISHER Ion Xpress Barcode Adapters 1-96 Kit, and / or THERMOFISHER Ion AmpliSeq Transcriptome Human Gene Expression Kit. In some embodiments, the adapter sequence is an adapter sequence for any of the following kits: PACBIO SMRTbell Template Preparation Kit 1.0, SMRTbell Barcoded Adapter Complete Preparation Kit - 96, or Barcoded Adapter Kit. In some embodiments, the target region can be complementary to the adapter sequence for the following kits: QIAGEN QIAseq Stranded RNA Library Kit or QIAseq UPX 3' Transcriptome Kit, PerkinElmer NEXTFLEX Rapid Directional RNA-Seq Kit or NEXTFLEX Small RNA-Seq Kit, or Takara Bio SMART-Seq mRNA Kit or SMART-Seq mRNA LP Kit.
[0077] In some embodiments, the adaptor sequence comprises a shared Y-adaptor sequence. In certain embodiments, the shared Y-adaptor sequence is a 13 bp sequence. In other embodiments, the shared Y-adaptor sequence is a 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, 12 bp, 14 bp, 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, 20 bp, 21 bp, 22 bp, 23 bp, 24 bp, 25 bp, 26 bp, 27 bp, 28 bp, 29 bp or 30 bp sequence.
[0078] In some embodiments, the adaptor sequence can comprise the sequence of any one of SEQ ID NOs: 27-36.
[0079] Kit
[0080] In some aspects, the present disclosure describes a kit comprising: (a) a guide polynucleotide as described herein; and (b) a Cas protein or an Argonaute protein, or a polynucleotide sequence encoding a Cas protein or an Argonaute protein (e.g., as described herein).
[0081] In some embodiments, the kit comprises a Cas protein. "Cas protein" refers to a clustered regularly interspaced short palindromic repeats (CRISPR)-associated protein. In some embodiments, the Cas protein has nuclease activity. In some embodiments, the Cas protein does not have nuclease activity (dCas). In some embodiments, the Cas protein has nickase activity (nickase).
[0082] In some embodiments, the Cas protein is a Cas9 protein. "Cas9" refers to a Cas9 protein or fragment thereof present in any bacterial species encoding a type II CRISPR / Cas9 system. See, e.g., Makarova et al., Nature Reviews, Microbiology, 9:467-477 (2011), including supplementary information. Cas9 homologs have been found in a variety of eubacteria, including but not limited to bacteria having the following taxonomic groups: Actinobacteria, Aquificae, Bacteroidetes-Chlorobi, Chlamydiae-Verrucomicrobia, Chloroflexi, Cyanobacteria, Firmicutes, Proteobacteria, Spirochaete, and Thermotogae. An exemplary Cas9 protein is Streptococcus pyogenes Cas9 protein. Additional Cas9 proteins and their homologs are described in, e.g., Chylinski et al., 2013, RNA Biology 10(5):726–37; Hou et al., 2013, Proc. Natl. Acad. Sci. USA 110(39):15644-49; Sampson et al., 2013, 497(7448):254-57; and Jinek et al., 2012, Science 337(6096):816-21. Full-length Cas9 is an RNA-guided endonuclease that includes a recognition domain and two nuclease domains (HNH and RuvC, respectively). In the amino acid sequence, HNH is linearly continuous, while RuvC is split into three regions, one to the left of the recognition domain and the other two to the right of the recognition domain flanking the HNH domain. Cas9 from Streptococcus pyogenes is targeted to a genomic locus in a cell by interaction with a guide RNA that hybridizes to a DNA sequence of 20 nucleotides immediately preceding the NGG motif recognized by Cas9.
[0083] In some embodiments, Cas9 is catalytically inactivated or nuclease-dead Cas9 (dCas9) that has been modified to inactivate Cas9 nuclease activity. The modifications include, but are not limited to, altering one or more amino acids to inactivate nuclease activity or the nuclease domain. For example, and not by way of limitation, D10A, H840A, and / or R1335K mutations can be made in Cas9 from Streptococcus pyogenes to inactivate Cas9 nuclease activity. In some embodiments, the Cas9 protein comprises the D10A mutation and the H840A mutation. Other modifications include removing all or a portion of the nuclease domain of Cas9 such that no sequence exhibiting nuclease activity is present in Cas9. Thus, catalytically inactivated Cas9 can include polypeptide sequences that have been modified to inactivate nuclease activity or that have one or more polypeptide sequences removed to inactivate nuclease activity. Catalytically inactivated Cas9 retains the ability to bind to DNA even though nuclease activity has been inactivated. Thus, dCas9 includes one or more polypeptide sequences required for DNA binding but includes a modified nuclease sequence or lacks a nuclease sequence responsible for nuclease activity.
[0084] In some embodiments, the catalytically inactivated Cas9 protein is a full-length Cas9 sequence from Streptococcus pyogenes, the Cas9 protein lacking the polypeptide sequence of the RuvC nuclease domain and / or the HNH nuclease domain and retaining DNA binding function. In other embodiments, the catalytically inactivated Cas9 protein sequence has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to a Cas9 polypeptide sequence lacking the RuvC nuclease domain and / or the HNH nuclease domain and retaining DNA binding function. In some embodiments, the Cas protein is Alt-R TM S.p. dCas9 protein V3 (IDT).
[0085] In some embodiments, the Cas protein is a Cpf1 (Cas12a), C2c1, C2c3, C2c2, CasX or CasY protein. In some embodiments, the Cas protein has been modified to inactivate Cas9 nuclease activity. The modification includes, but is not limited to, altering one or more amino acids to inactivate the nuclease activity or nuclease domain of the Cas protein. For example, D908, D832, E993, R1226 and / or D1235 mutations can be made in Cpf1 from Acidaminacoccus sp. BV3L6, Lachnospiraceae ND2006 or Francisella tularensis subsp. novicida U112 to inactivate Cpf1 nuclease activity. Other modifications include removing all or a portion of the nuclease domain such that the sequence that exhibits nuclease activity is absent. Thus, the catalytically inactivated Cas protein can include a polypeptide sequence that has been modified to inactivate nuclease activity or has had one or more polypeptide sequences removed to inactivate nuclease activity. The catalytically inactivated Cas protein retains the ability to bind DNA even though the nuclease activity has been inactivated. Thus, the catalytically inactivated Cas protein includes one or more polypeptide sequences required for DNA binding, but includes a modified nuclease sequence or lacks a nuclease sequence responsible for nuclease activity. In some embodiments, the catalytically inactivated Cas protein is a full-length Cas sequence that lacks a nuclease domain polypeptide sequence and retains DNA binding function. In other embodiments, the catalytically inactivated Cas protein sequence has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to a Cas polypeptide sequence that lacks a nuclease domain and retains DNA binding function.
[0086] In some embodiments, the Cas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO:37 or SEQ ID NO:38. In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO:37 or SEQ ID NO:38.
[0087] In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO:37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO:37.
[0088] In some embodiments, the Cas protein comprises a nuclease localization sequence.
[0089] In some embodiments, the kit comprises an Argonaute protein. An "Argonaute protein" refers to a protein that binds to small non-coding nucleic acids (e.g., an Argonaute guide polynucleotide) and uses them to effect guide cleavage of a complementary nucleic acid target or indirect gene silencing by recruiting additional factors. A catalytically active Argonaute protein is capable of nucleic acid guide binding to a complementary nucleic acid target (e.g., DNA) and cleaving the nucleic acid target. See, e.g., Kaya et al., 2016, Proceedings of the National Academy of Sciences of the United States of America (PNAS) 113(5):4057-62. The prokaryotic Argonaute (AGO) gene family encodes several domains: an N-terminal (N), PAZ, MID, and C-terminal PIWI domain. The MID and PAZ domains are responsible for binding the 5'-end and 3'-end of the guide polynucleotide, respectively. In contrast to eukaryotic Ago, which uses only small RNA guides, most characterized prokaryotic Ago binds to single-stranded DNA (ssDNA) guides. The Ago PIWI domain contains nuclease activity. In some embodiments, the Argonaute protein is catalytically inactive. In some embodiments, the Argonaute protein is a catalytically inactive CbAgo protein. In some embodiments, the Argonaute protein is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo protein. In certain embodiments, the Argonaute protein has been modified to inactivate nuclease activity or inherently has reduced nuclease activity (see, e.g., Kaya et al., 2016, Proceedings of the National Academy of Sciences 113(5):4057-62).
[0090] Modifications can include, but are not limited to, altering one or more amino acids to inactivate the nuclease activity or nuclease domain of the Argonaute protein. For example, in some embodiments, the catalytically inactivated Argonaute protein is the CbAgo protein comprising the D541A mutation and the D611A mutation. See, e.g., Hegge et al., 2019, Nuc. Acids Res. 47(11):5809-21. Alternatively, all or a portion of the nuclease domain can be removed such that the sequence exhibiting nuclease activity is absent. Thus, the catalytically inactivated Argonaute protein can include a polypeptide sequence that is modified to inactivate nuclease activity or a polypeptide sequence that has one or more polypeptide sequences removed to inactivate nuclease activity. The catalytically inactivated Argonaute protein retains the ability to bind DNA even though the nuclease activity has been inactivated. Thus, the catalytically inactivated Argonaute protein includes one or more polypeptide sequences required for DNA binding, but includes a modified nuclease sequence or lacks the nuclease sequence responsible for nuclease activity. In some embodiments, the catalytically inactivated Argonaute protein is a full-length Argonaute sequence that lacks the polypeptide sequence of the nuclease domain and retains DNA binding function. In other embodiments, the catalytically inactivated Argonaute protein sequence has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to an Argonaute polypeptide sequence that lacks the nuclease domain and retains DNA binding function.
[0091] In some embodiments, the dArgonaute protein comprises an amino acid sequence having at least 95% identity to any one of SEQ ID NOs: 40-45. In some embodiments, the dArgonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 40-45.
[0092] In some embodiments, the kit comprises a catalytically inactivated Cas protein (dCas protein) or a catalytically inactivated Argonaute protein (e.g., dArgonaute) as described herein. In some embodiments, the kit comprises a catalytically inactivated Cas protein (dCas protein) and a catalytically active Cas protein as described herein. In some embodiments, the kit comprises a catalytically inactivated Argonaute protein (e.g., dArgonaute) and a catalytically active Argonaute protein. In some embodiments, the kit comprises dCas and a homologous dCas guide RNA polynucleotide. In some embodiments, the kit comprises dCas9 and a homologous dCas9 guide RNA. In some embodiments, the kit comprises dArgonaute and a homologous dArgonaute guide polynucleotide.
[0093] In some embodiments, the kit comprises a Cas protein comprising an affinity tag or an Argonaute protein comprising an affinity tag. In some embodiments, the affinity tag is albumin binding protein (ABP), alkaline phosphatase (AP), AU1 epitope, AU5 epitope, bacteriophage T7 epitope (T7 tag), bacteriophage V5 epitope (V5 tag), biotin-carboxyl carrier protein (BCCP), bluetongue virus tag (B tag), calmodulin binding peptide (CBP), chloramphenicol acetyltransferase (CAT), cellulose binding domain (CBP), chitin binding domain (CBD), choline binding domain (CBD), dihydrofolate reductase (DHFR), E2 epitope, FLAG epitope, galactose binding protein (GBP), green fluorescent protein (GFP), Glu-Glu (EE-tag), Glu-Glu (EE-tag), human influenza hemagglutinin (HA), Histidine affinity tag (HAT), horseradish peroxidase (HRP), HSV epitope, ketosteroid isomerase (KSI), KT3 epitope, LacZ, luciferase, maltose binding protein (MBP), Myc epitope, Nus, PDZ domain, PDZ ligand, polyarginine (Arg-tag), polyaspartate ester (Asp-tag), polyhistidine (His-tag), polyphenylalanine (Phe-tag), Profinity eXact, protein C, S1-tag, S-tag, streptavidin-binding peptide (SBP), staphylococcal protein A (protein A), staphylococcal protein G (protein G), Strep-tag, streptavidin, small ubiquitin-like modifier (SUMO), tandem affinity purification (TAP), T7 epitope, thioredoxin (Trx), TrpE, ubiquitin, Universal and VSV-G, as described in Kimple ME et al., Current Protocols in Protein Science, September 24, 2013; 73:9.9.1-9.9.23. doi:10.1002 / 0471140864.ps0909s73.
[0094] In some embodiments, the affinity tag and corresponding affinity tag-binding molecule in the kit or for use in the method are selected from Table 1.
[0095] Table 1: Affinity tags and affinity tag-binding molecules.
[0096] Affinity Tag Fusion MW Affinity Tag - Binding Molecule GST 27 kDa Glutathione Beads Protein A 49 kDa IgG - Conjugated Beads Protein G 65 kDa IgG - Conjugated Beads Streptavidin 16.5 kDa Biotin Beads Biotin Post - Translation Streptavidin Beads Cell - Surface Vimentin (CSV) 57 kDa CSV Monoclonal Ab Beads Human PSMA 82.5 kDa PMSA Monoclonal Ab Beads IgG Fc Heavy Chain Approximately 25 - 50 kDa Protein A / G Beads Chitin - Binding Domain 27 kDa Chitin Beads Maltose - Binding Protein (MBP) 42.5 kDa Amylose or Anti - MBP Beads Protein L ‘B’ Repeat 36 kDa IgG - Conjugated Beads
[0097] In some embodiments, the kit comprises a guide polynucleotide that is capable of binding to the Cas protein of the kit (e.g., the guide polynucleotide is homologous to the Cas protein) or an Argonaute protein to form a ribonucleoprotein complex, and the ribonucleoprotein complex is capable of binding to an adapter sequence.
[0098] The term "homologous" as used in the context of a guide polynucleotide and a Cas protein or an Argonaute protein refers to a guide polynucleotide that is compatible with the Cas protein or an Argonaute protein (e.g., the guide polynucleotide is capable of guiding the Cas protein or an Argonaute protein to a target polynucleotide). For example, a Cas9 sgRNA is homologous to Cas9 and dCas9.
[0099] The term "bind" or "specifically bind" and similar terms refer to a molecule (e.g., a Cas-gRNA complex or an Argonaute-guide polynucleotide complex) that binds to a target nucleic acid with an affinity at least 2-fold greater than that for a non-target nucleic acid when measured under the same binding affinity assay conditions. For example, the affinity for the target nucleic acid is at least any one of 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 25-fold, 50-fold, 100-fold, 1,000-fold, 10,000-fold or greater than that for an unrelated nucleic acid. In some embodiments, as used herein, the term "bind" or "specifically bind" can be demonstrated, for example, by a molecule (e.g., a Cas-gRNA complex or an Argonaute-guide polynucleotide complex) having an equilibrium dissociation constant KD of, for example, 10 -2 M or less, such as 10 -3 M, 10 -4 M, 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M, 10 - 9 M, 10 -10 M, 10 -11 M or 10 -12 M for the target nucleic acid. In some embodiments, the KD of the antibody is less than 10 nM or less than 100 nM.
[0100] In some embodiments, the kit further comprises a primer (e.g., as described herein) that is complementary to a portion of the adapter sequence. In some embodiments, the primer is complementary to the proximal portion of the adapter sequence. The primer can be used to generate a double-stranded target nucleotide that contains a double-stranded protospacer adjacent motif (PAM) site to which the dCas protein (e.g., as described herein) can bind. In some embodiments, the primer binds to the proximal portion of the adapter sequence such that the new polynucleotide (which is complementary to the target polynucleotide) does not contain the distal portion of the adapter sequence and thus cannot be sequenced using next-generation sequencing (e.g., ILLUMINA sequencing). Without being bound by theory, this method may be advantageous compared to PCR-based methods because errors (e.g., mutations) introduced during primer extension are not sequenced, but the double-stranded target polynucleotide can still be generated and bound by the dCas protein.
[0101] In some embodiments, the primer is a DNA primer. In some embodiments, the DNA primer comprises a nucleic acid modification (e.g., as described herein). In some embodiments, the portion of the adapter sequence is located in the proximal portion of the adapter sequence. In some embodiments, the proximal portion of the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NO: 28, 30, 32, and 34. In some embodiments, the primer is complementary to the 3' adapter sequence. In some embodiments, the primer is complementary to the 5' adapter sequence. In some embodiments, the primer is not complementary to the adapter sequence that is complementary to the guide polynucleotide targeting region. In some embodiments, the primer does not extend the entire adapter sequence when binding to and extending the adapter sequence.
[0102] Reaction mixture
[0103] In some aspects, the present disclosure describes a reaction mixture comprising: (i) a plurality of target polynucleotides, wherein the target polynucleotide comprises an adapter sequence; (ii) a predetermined concentration of a guide polynucleotide, the guide polynucleotide comprising a targeting region complementary to the adapter sequence; and (iii) a predetermined concentration of a dCas protein or a dArgonaute protein. In some embodiments, the reaction mixture comprises an aqueous solution (e.g., a buffer suitable for performing the methods described herein). In some embodiments, the reaction mixture comprises a predetermined concentration of a dCas9 protein. In some embodiments, the reaction mixture comprises a predetermined concentration of a dArgonaute protein. In some embodiments, the reaction mixture comprises a target polynucleotide comprising the adapter sequence shown in any one of SEQ ID NO: 27-36. In some embodiments, the reaction mixture comprises a predetermined concentration of a guide polynucleotide of any one of SEQ ID NO: 27-36.
[0104] In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 50 femtomoles (fmol) to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 200 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 300 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 500 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 750 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 2,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 50 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides, the plurality of target polynucleotides comprising an adaptor sequence at a concentration of 500 fmol to 3,400 fmol.
[0105] In some embodiments, the reaction mixture comprises a dCas9 of SEQ ID NO:38 at a predetermined concentration, a target polynucleotide comprising an adaptor sequence shown in SEQ ID NO:28, and a guide polynucleotide (e.g., an RNA guide polynucleotide) at a predetermined concentration, the guide polynucleotide comprising a homologous region complementary to the adaptor sequence (e.g., the homologous region comprises SEQ ID NO:37).
[0106] In some embodiments, the reaction mixture comprises a dArgonaute of any one of SEQ ID NOs: 40-45 at a predetermined concentration (e.g., a catalytically inactivated variant of any one of SEQ ID NOs: 40-45), a target polynucleotide comprising the adaptor sequence shown in SEQ ID NO: 28, and a dArgonaute guide polynucleotide at a predetermined concentration, wherein the dArgonaute guide polynucleotide comprises a targeting region complementary to the adaptor sequence.
[0107] In some embodiments, the reaction mixture further comprises a catalytically active nuclease as described herein.
[0108] Ribonucleoprotein (RNP) complex
[0109] In some aspects, the present disclosure provides a ribonucleoprotein (RNP) complex, the RNP complex comprising: (i) a guide polynucleotide; (ii) a Cas protein (e.g., dCas) or an Argonaute protein (e.g., dArgonaute); and (iii) an adaptor sequence; wherein the targeting region of the guide polynucleotide is complementary to the adaptor sequence.
[0110] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide; (ii) a Cas protein as described herein; and (iii) an adaptor sequence; wherein the homology region of the RNA Cas gRNA polynucleotide is complementary to the adaptor sequence.
[0111] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide comprising a homology region shown in any one of SEQ ID NOs: 1-26; (ii) a dCas protein of SEQ ID NO: 37; and (iii) an adaptor sequence shown in any one of SEQ ID NOs: 27-36; wherein the homology region of the RNA Cas gRNA polynucleotide is complementary to the adaptor sequence.
[0112] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide comprising the homology region shown in SEQ ID NO: 1, (ii) a dCas protein of SEQ ID NO: 37, and (iii) an adaptor sequence shown in SEQ ID NO: 28; wherein the homology region of the RNA Cas gRNA polynucleotide is complementary to the adaptor sequence.
[0113] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide, (ii) an Argonaute protein, and (iii) an adaptor sequence; wherein the targeting region of the Argonaute guide polynucleotide is complementary to the adaptor sequence.
[0114] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide (e.g., siDNA), (ii) an Argonaute protein of any one of SEQ ID NOs: 39-45, and (iii) the adaptor sequence shown in SEQ ID NO: 28; wherein the targeting region of the Argonaute guide polynucleotide is complementary to the adaptor sequence.
[0115] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide (e.g., siDNA), (ii) an Argonaute protein of any one of SEQ ID NOs: 40-45, and (iii) the adaptor sequence shown in SEQ ID NO: 28; wherein the targeting region of the Argonaute guide polynucleotide is complementary to the adaptor sequence.
[0116] In some embodiments of the RNP complexes provided herein, the adaptor sequence is unmodified. In some embodiments, the adaptor sequence is unmodified at the 5' end.
[0117] In some embodiments, the RNP complex is present at a concentration of from 250 femtomoles (fmol) to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of from 300 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of from 2,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of from 6,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of from 10,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 250 fmol. In some embodiments, the RNP complex is present at a concentration of at least 300 fmol. In some embodiments, the RNP complex is present at a concentration of at least 2,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 6,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 10,000 fmol. In some embodiments, the RNP complex is present at a concentration of 250 fmol. In some embodiments, the RNP complex is present at a concentration of 300 fmol. In some embodiments, the RNP complex is present at a concentration of 2,000 fmol. In some embodiments, the RNP complex is present at a concentration of 6,000 fmol. In some embodiments, the RNP complex is present at a concentration of 10,000 fmol. In some embodiments, the RNP complex is present at a concentration of 14,000 fmol.
[0118] Sequence
[0119] Table 2: Sequences
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129] Methods for normalizing polynucleotide libraries
[0130] In some aspects, the present disclosure provides methods for normalizing the concentration of target polynucleotides between two or more samples (e.g., normalizing between two or more polynucleotide libraries).
[0131] As used herein, "normalizing," "normalization," and like terms refer to the process of generating a subsequent sample having a desired concentration of a target polynucleotide from an initial sample that has an initial concentration of the target polynucleotide different from that of the subsequent sample. In some embodiments, "normalizing" two or more samples results in the generation of two or more subsequent samples, each of which has a more similar ratio of the concentration of the target polynucleotide to each other after the normalization method than before the normalization method. For example, a first sample (e.g., a first polynucleotide library) may contain 5 times more target polynucleotide than a second sample (e.g., a second polynucleotide library). In this example, after normalization of the first and second samples, the concentration difference between the first sample and the second sample will be less than 5-fold (e.g., less than 4-fold, less than 3-fold, or less than 2-fold). In some embodiments, normalizing two or more samples results in the generation of two or more subsequent samples having equimolar (i.e., 1:1) concentrations. In some embodiments, the normalization method results in a difference in the concentration of the target polynucleotide between two or more samples of less than 50% (e.g., less than 40%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, or less than 2.5%). In some embodiments, the normalization method results in a difference in the concentration of the target polynucleotide between two or more samples that is 5%-40%, 5%-30%, 5%-20%, 5%-10%, 10%-40%, 10%-30%, or 10%-20% in concentration. In some embodiments, the difference in the starting concentration of the target polynucleotide between two or more samples may be relatively small (e.g., less than 5-10%) prior to performing the normalization methods described herein. In such embodiments, although the normalization method may be performed, the two or more samples may not have more similar detectable concentrations after normalization than before normalization.
[0132] "Target polynucleotide" or "target polynucleotides" refers to a polynucleotide that comprises a nucleic acid encoding an adaptor sequence or a static region of an adaptor sequence.
[0133] In some embodiments, the target polynucleotide comprises a polynucleotide that includes (1) a nucleic acid that encodes an adaptor sequence or a sequence that has been added to the polynucleotide for normalization purposes using the methods described herein (e.g., a zinc finger binding domain or a Talen binding domain), and (2) a nucleic acid that encodes a sequence of interest. In some embodiments, the target polynucleotide comprises a nucleic acid that encodes an adaptor sequence described herein and a nucleic acid that encodes a sequence of interest. The sequence of interest can be any sequence for which normalization is desired. For example, the sequence of interest can include, but is not limited to, DNA, genomic DNA, circulating tumor DNA, RNA, rRNA, mRNA, miRNA, or cDNA. The adaptor sequence can be linked to the 5' end and / or the 3' end of the sequence of interest. In some embodiments, the target polynucleotide comprises a first adaptor sequence (e.g., a p5 adaptor) located at the 5' end of the sequence of interest and a second adaptor sequence (e.g., a p7 adaptor) located at the 3' end of the target sequence.
[0134] "Sample" refers to a composition or solution that contains one or more target polynucleotides. In some embodiments, the sample contains multiple target polynucleotides. "Multiple" means two or more. For example, multiple target polynucleotides include two or more target polynucleotides. In some embodiments, the multiple target polynucleotides include multiple copies of a single target polynucleotide sequence (e.g., at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, or at least 1,000,000, or at least 10,000,000, at least 100,000,000 target polynucleotides). In some embodiments, the multiple target polynucleotides include different target polynucleotide sequences (e.g., at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, or at least 1,000,000, or at least 10,000,000, at least 100,000,000 different target polynucleotides), but at least some of the target polynucleotides (e.g., all target polynucleotides) contain the same adapter sequence. In some embodiments, the sample contains target polynucleotides and other molecules (e.g., polynucleotides that are not target polynucleotides and do not contain an adapter sequence or contain an adapter sequence different from the adapter sequence complementary to the guide polynucleotide). In some embodiments, the sample contains target polynucleotides that contain nucleic acids encoding genomic DNA. In some embodiments, the sample contains target polynucleotides that contain nucleic acids encoding RNA. In some embodiments, the sample contains target polynucleotides that contain nucleic acids encoding cDNA (e.g., of mRNA or lncRNA). In some embodiments, the sample contains target polynucleotides that contain nucleic acids encoding exons. In some embodiments, the sample contains target polynucleotides that contain nucleic acids encoding introns.
[0135] In some embodiments, the plurality of target polynucleotides in the sample are present at a concentration of 50 femtomoles (fmol) to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 200 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 300 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 500 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 750 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 2,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 50 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 500 fmol to 3,400 fmol.
[0136] In some embodiments, the normalization method is performed on at least 2 samples (e.g., at least 3 samples, at least 4 samples, at least 5 samples, at least 6 samples, at least 7 samples, at least 8 samples, at least 9 samples, at least 10 samples, at least 25 samples, at least 50 samples, at least 75 samples, at least 100 samples, at least 150 samples, or at least 200 samples, at least 300 samples, at least 384 samples, at least 400 samples, at least 500 samples, at least 600 samples, at least 700 samples, at least 800 samples, at least 900 samples, at least 1,000 samples, at least 1,500 samples, at least 2,000 samples, at least 3,000 samples, or at least 5,000 samples or more). In some embodiments, the normalization method is performed on 2 - 384 samples. In some embodiments, the normalization method is performed on 24 - 384 samples. In some embodiments, the normalization method is performed on 24 - 500 samples. In some embodiments, the normalization method is performed on 24 - 750 samples. In some embodiments, the normalization method is performed on 24 - 1,000 samples. In some embodiments, the normalization method is performed on 24 - 2,000 samples. In some embodiments, the normalization method is performed on 24 - 3,000 samples.
[0137] In some embodiments, the sample comprises target polynucleotides having concentrations within 70-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 60-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 50-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 40-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 30-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 20-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 10-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 5-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 2-fold of each other prior to normalization. In some embodiments, the sample comprises target polynucleotides having concentrations within 1.5-fold of each other prior to normalization.
[0138] In some embodiments, the present disclosure provides a method for normalizing the concentration of target polynucleotides between at least two samples (e.g., two polynucleotide libraries), each of the at least two samples comprising a target polynucleotide, the method comprising, for each of the at least two samples: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adaptor sequence; (ii) generating a solution, the generating comprising combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to a static region of the adaptor sequence of the target polynucleotide of the sample, and (c) dCas or dArgonaute at a predetermined concentration; (iii) contacting the solution with a solid phase comprising a dCas or dArgonaute binding molecule; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0139] In some embodiments, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples (e.g., two polynucleotide libraries), each of the at least two samples containing the target polynucleotide. For each of the at least two samples, the method includes: (i) obtaining the sample, wherein the target polynucleotide of the sample contains an adaptor sequence; (ii) generating a solution, the generating including combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide containing a targeting region complementary to the adaptor sequence of the target polynucleotide of the sample, and (c) dCas or dArgonaute at a predetermined concentration; (iii) contacting the solution with a solid phase containing a dCas or dArgonaute binding molecule; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0140] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes the step of obtaining a sample containing the target polynucleotide. The sample can be obtained from any suitable source or generated in vitro, the suitable sources including but not limited to cells (e.g., eukaryotic cells or prokaryotic cells), tissues (e.g., brain, heart, lung, liver, kidney, fat, skin, gallbladder or mammary gland), diseased tissues (e.g., tumors), body fluids (e.g., saliva, sweat, blood, urine, mucus, or cerebrospinal fluid). In some embodiments, obtaining the sample includes obtaining the polynucleotide of interest (e.g., extracting the polynucleotide of interest from a biological sample) and modifying the polynucleotide of interest to contain an adaptor sequence (e.g., modifying the polynucleotide of interest to become a target polynucleotide). Modifying the polynucleotide of interest to contain an adaptor sequence can be performed using any suitable method, including but not limited to using a kit for ligating an adaptor sequence to the polynucleotides described herein. In some embodiments, the sample contains 1-100 nM target polynucleotide. In some embodiments, the sample contains 1-1000 nM target polynucleotide.
[0141] In some embodiments, in the context of the normalization method, generating a solution refers to combining a sample with a guide polynucleotide at a predetermined concentration and a dCas or dArgonaute at a predetermined concentration in an aqueous solution (e.g., a buffer suitable for the method). In some embodiments, generating a solution includes combining a predetermined amount of dCas protein or dArgonaute protein and a predetermined amount of guide polynucleotide prior to combining with the sample. In some embodiments, generating a solution includes combining a predetermined amount of guide RNA polynucleotide, which is RNA, with a predetermined amount of dCas9 protein or dArgonaute protein. In some embodiments, generating a solution includes combining a predetermined amount of guide polynucleotide, which is DNA, with a predetermined amount of dArgonaute protein. In some embodiments, generating a solution includes combining a predetermined amount of guide RNA comprising a homology region shown in any of SEQ ID NO: 1-26 with a predetermined amount of dCas9 of SEQ ID NO: 37.
[0142] "Predetermined concentration" refers to the concentration (e.g., nanomolar) or amount (e.g., nanomolar) of a reagent (e.g., a guide polynucleotide, dCas protein, or dArgonaute protein) that has been selected to extract a specific or consistent amount of a target polynucleotide from a sample. In some embodiments, the same or a similar amount of a dCas protein or dArgonaute protein at a predetermined concentration is added to each sample to be normalized. In some embodiments, the same or a similar amount of a guide RNA polynucleotide at a predetermined concentration is added to each sample to be normalized. In some embodiments, a dCas protein or dArgonaute protein at a predetermined concentration is added to a sample, and a predetermined amount of guide RNA polynucleotide added to the sample exceeds the predetermined amount of dCas protein or dArgonaute protein added to the sample. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 1 - 2000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 60 - 2000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 60 - 1000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 125 - 1000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is at least 60 femtomoles (e.g., at least 60 femtomoles, at least 125 femtomoles, at least 250 femtomoles, at least 500 femtomoles, at least 750 femtomoles, or at least 1000 femtomoles). In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 50 femtomoles, 60 femtomoles, 125 femtomoles, 250 femtomoles, 750 femtomoles, 1000 femtomoles, 1500 femtomoles, or 2000 femtomoles. In some embodiments, the predetermined concentration of the guide RNA polynucleotide exceeds the predetermined concentration of the dCas9 protein or dArgonaute protein.
[0143] In some embodiments, the ribonucleoprotein (RNP) complex (e.g., a guide RNA, a dCas protein or a dArgonaute protein, and a linker sequence) is present in the sample at a concentration of from 250 femtomoles (fmol) to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of from 300 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of from 2,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of from 6,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of from 10,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 250 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 300 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 2,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 6,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 10,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 250 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 300 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 2,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 6,000. In some embodiments, the RNP complex is present in the sample at a concentration of 10,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 14,000 fmol.
[0144] In some embodiments, the method includes generating a solution comprising a dCas protein at a predetermined concentration, the dCas protein comprising an affinity tag, or a dArgonaute protein comprising an affinity tag as described herein.
[0145] In some embodiments, the affinity tag-binding molecule comprises a metal ion (e.g., Ni2+) and the affinity tag comprises a His tag. In some embodiments, the affinity tag-binding molecule comprises biotin and the affinity tag comprises avidin. In some embodiments, the affinity tag-binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag. In some embodiments, the affinity tag-binding molecule and the corresponding affinity tag are selected from Table 1.
[0146] In some embodiments, the dArgonaute or dCas of the method does not contain an affinity tag. In some embodiments, the method includes contacting a solution with a solid phase that includes a molecule that binds to dCas or dArgonaute (e.g., an antibody that binds to dCas or dArgonaute).
[0147] In some aspects, the method includes contacting a solution with a solid phase. In some embodiments, the solid phase includes a molecule that binds to dArgonaute or dCas (e.g., an antibody or an affinity tag-binding molecule). In some embodiments, the solid phase includes beads (e.g., microparticles or nanoparticles). In some embodiments, the beads are metal beads, polymer beads, protein beads, or lipid beads. In some embodiments, the beads (e.g., metal beads) include metal ions (e.g., Ni2+, Co2+, Cu2+, and / or Zn2+ ions) on the surface of the beads. In some embodiments, the metal beads are magnetic (e.g., paramagnetic, diamagnetic, or ferromagnetic).
[0148] In some embodiments, the method includes incubating the solid phase and the solution. In some embodiments, the incubation is carried out at room temperature. In some embodiments, the incubation is carried out at 30 - 40 °C. In some embodiments, the incubation is carried out at 35 - 38 °C. In some embodiments, the incubation is carried out at 37 °C or about 37 °C. In some embodiments, the incubation is at least 15 minutes (e.g., at least 30 minutes, at least 45 minutes, at least 60 minutes, or at least 2 hours). In some embodiments, the incubation is at least 60 minutes. In some embodiments, the incubation is 15 minutes to 2 hours.
[0149] In some embodiments, the method further includes separating the solution from the solid phase. "Separation" as used in the context of this method refers to removing the solution from the solid phase or removing the solid phase from the solution. In some embodiments, the separation is not a perfect separation, e.g., there may be a small amount of solution (e.g., less than about 1 microliter of solution) and the solid phase still in contact with each other after separation. In some embodiments, the purpose of the separation step is to separate unbound target polynucleotides in the solution from target polynucleotides bound to the solid phase, which is part of normalizing the concentration. This can be accomplished by: (1) capturing magnetic beads bound to the target polynucleotide (e.g., using a magnet); (2) removing the solution from the solid phase (e.g., using a pipette); and (3) repeatedly rinsing the solid phase with a buffer (e.g., a buffer that does not denature the Cas protein or the Argonaute protein). In some embodiments, the method includes washing the solid phase at least 1 time (e.g., at least 2 times, at least 3 times, at least 4 times, or at least 5 times or more).
[0150] In some embodiments, the method further includes extracting the target polynucleotide from the solid phase. In some embodiments, the extraction includes contacting the solid phase with a protease that proteolytically cleaves the Cas protein or Argonaute protein, thereby releasing the target polynucleotide. The protease can be any suitable protease, including but not limited to trypsin, chymotrypsin, intracellular protease Lys-C, endoprotease AspN, endoprotease GluC, elastase, proteinase K, or papain. The corresponding protease reaction conditions and incubation times are known in the art. In some embodiments, the extraction includes contacting the solid phase with a protein denaturant (e.g., a detergent and an organic solvent or a chaotropic agent). In some embodiments, the extraction includes raising the temperature of the solid phase (e.g., by warming the liquid in which the solid phase is located) to denature the Cas protein or Argonaute protein. In some embodiments, the extraction includes raising or lowering the pH of the liquid in which the solid phase is located.
[0151] In some embodiments, the normalization method includes performing the following steps in sequence: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution, the generating including combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adapter sequence of the target polynucleotide of the sample, and (c) dCas or dArgonaute at a predetermined concentration, the dCas or dArgonaute comprising an affinity tag, wherein the predetermined concentration of the dCas or dArgonaute corresponds to the guide polynucleotide; (iii) contacting the solution with a solid phase comprising an affinity tag binding molecule capable of binding to the affinity tag; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0152] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes, for each of the at least two samples: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence shown in any of SEQ ID NOs: 27-36; (ii) generating a solution, the generating including combining (a) the sample, (b) a Cas gRNA polynucleotide at a predetermined concentration, the Cas gRNA polynucleotide comprising a homologous region complementary to the adapter sequence, and (c) dCas9 at a predetermined concentration, the dCas9 comprising a His-tag shown in SEQ ID NO: 47 (e.g., a his-tag located at the n-terminus of dCas9); (iii) contacting the solution with a magnetic solid phase comprising Ni2+ ions; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0153] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes, for each of the at least two samples: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence shown in any of SEQ ID NOs: 27-36; (ii) generating a solution, the generating including combining (a) the sample, (b) an Argonaute guide polynucleotide (e.g., siDNA) at a predetermined concentration, the Argonaute guide polynucleotide comprising a targeting region complementary to the adapter sequence, and (c) an Argonaute protein at a predetermined concentration, the Argonaute protein comprising a His-tag shown in SEQ ID NO: 47; (iii) contacting the solution with a magnetic solid phase comprising Ni2+ ions; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0154] In some aspects, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples, each of the at least two samples comprising the target polynucleotide, wherein the target polynucleotide not captured during normalization is digested. In some embodiments, the method for each of the at least two samples comprises: (i) obtaining the sample (e.g., as described herein), wherein the target polynucleotide of the sample comprises an adaptor sequence; (ii) generating a solution, the generating comprising combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adaptor sequence of the target polynucleotide of the sample, and (c) dCas or dArgonaute at a predetermined concentration, the dCas or dArgonaute comprising an affinity tag; and (iii) contacting the solution with a nuclease (e.g., a catalytically active Cas or Argonaute protein 9) to normalize the concentration of the target polynucleotide between the two or more samples (e.g., digest the target polynucleotide not bound to the solid phase).
[0155] In some embodiments, the method comprises contacting the solution with a nuclease. In some embodiments, the nuclease is an exonuclease. In certain embodiments, the exonuclease is exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VI, exonuclease VII, or exonuclease VIII. In some embodiments, the nuclease is a catalytically active Cas RNP or Argonaute RNP. In some embodiments, the catalytically active Cas RNP or Argonaute RNP comprises a guide polynucleotide complementary to the adaptor sequence. In some embodiments, the Cas endonuclease is a Cas9 endonuclease. In some embodiments, the Cas endonuclease is a Cpf1, C2c1, C2c3, C2c2, CasX, or CasY endonuclease. In some embodiments, the catalytic activity is an Argonaute protein (e.g., an Argonaute endonuclease). In some embodiments, the Argonaute protein is a CbAgo endonuclease. In certain embodiments, the Argonaute protein is an LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo endonuclease.
[0156] In some aspects, two or more samples that are normalized (e.g., using the methods described herein) contain a target polynucleotide that includes a first adapter sequence and a second adapter sequence. In some embodiments, the guide polynucleotide targeting region is complementary to the first adapter sequence. In some embodiments, the method further comprises contacting the sample with a primer and a polymerase (e.g., a DNA polymerase) under conditions sufficient to effect primer extension, the primer encoding a nucleic acid sequence that is complementary to a portion of the second adapter sequence. As described in the "Kits" section, this primer can be used to extend the target polynucleotide such that the first adapter sequence includes a double-stranded PAM for Cas (e.g., dCas) binding. In some embodiments, the primer is complementary to the proximal portion of the adapter sequence (e.g., the second adapter sequence). In such embodiments, extending the primer to produce a double-stranded polynucleotide may not produce an extended strand that can be sequenced by next-generation sequencing because the extended strand does not include the distal portion of the second adapter sequence. This can be advantageous because extension (e.g., DNA amplification) may introduce artifacts (e.g., mutations) into the polynucleotide, which may bias or distort the sequencing results.
[0157] In some embodiments, the primer comprises a nucleic acid sequence that is complementary to any one of SEQ ID NOs: 28, 30, 32, and 34. In some embodiments, the second adapter sequence is a 3' adapter sequence. In some embodiments, the second adapter sequence is a 5' adapter sequence.
[0158] As used herein, unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" include plural referents. Thus, for example, reference to "an antibody" optionally includes a combination of two or more such molecules, and the like.
[0159] Unless otherwise required, any and all examples or exemplary language (e.g., "such as") provided herein are only intended to better illustrate the invention and do not limit the scope of the invention.
[0160] Unless the context clearly indicates otherwise, the terms "may," "may be," "can," and "can be" and related terms are intended to convey that the subject matter involved is optional (i.e., the subject matter exists in some instances and not in others), rather than a reference to the ability or probability of the subject matter.
[0161] The terms "optional" and "optionally" mean that the subsequently described event, circumstance, or material may or may not occur or exist, and the description includes both the case where the event, circumstance, or material occurs or exists and the case where the event, circumstance, or material does not occur or exist.
[0162] The use herein of the terms "including", "comprising", or "having" and variations thereof means covering the elements listed thereafter and their equivalents as well as additional elements. Embodiments recited as "including", "comprising", or "having" certain elements are also contemplated as "consisting essentially of those particular elements" and "consisting of those particular elements". As used herein, "and / or" means and encompasses any and all possible combinations of one or more of the associated listed items and the absence of combinations when interpreted in the alternative ("or").
[0163] Examples
[0164] Example 1 - Library normalization achieved by Cas9-guide-mediated capture of library molecules.
[0165] This example demonstrates that the target polynucleotide concentration between samples can be normalized using a predetermined amount of dCas protein and gRNA that binds to the adapter sequence of the target polynucleotide.
[0166] Normalization of polynucleotide libraries (also referred to as NGS libraries) prior to equimolar pooling of multiple different polynucleotide libraries is a challenge in next-generation sequencing workflows. While there are various normalization strategies, these strategies typically require additional steps such as qPCR or special primers or oligonucleotides and are often destructive to unused materials.
[0167] The normalization method described herein includes a sample comprising a target polynucleotide library, the concentration of which is variable and unknown, and reducing the concentration of the library to a predetermined value. This enables the target polynucleotide concentrations of several unknown target polynucleotide libraries to be normalized to more similar molar concentrations and then pooled by volume prior to loading onto a sequencer. The general method relies on using an adjustable amount of a nuclease with a non-cleaving nucleic acid guide (e.g., dCas9 or dArgonaute) linked to magnetic beads to capture a known number of library molecules ( Figure 1)。After separating the magnetic beads from the initial binding incubation and subsequent buffer washes, unbound material is removed by pipetting to remove excess library molecules. After washing, a large amount of beads from each sample are resuspended in the same volume of buffer. The captured DNA can then be obtained by denaturing the non-cleaving nuclease-guided nucleic acid. The resulting solution is then pooled in equal volumes with other libraries prepared in the same manner prior to sequencing.
[0168] In this example, a nuclease-guided nucleic acid (NAGN) in a magnetic bead-compatible form is pre-loaded with a Cas guide RNA specific for an adapter common to NGS library polynucleotides (e.g., ILLUMINA i5 / i7 sequences), and the nuclease-guided nucleic acid in the magnetic bead-compatible form is capable of specifically binding to a target determined by a guide nucleic acid (e.g., guide RNA or DNA) but not cleaving the target (e.g., dCas9 D10A / H840A double mutant). Guide RNA sequences AA (corresponding to SEQ ID NO:1) and AB (corresponding to SEQ ID NO:2) were tested, and AA showed higher efficiency. In this case, the guide RNA sequence targets the i5 region of the ILLUMINA Tru-Seq adapter.
[0169] Guide RNA sequence AA: / AltR1 / rArG rArUrC rGrGrA rArGrA rGrCrG rUrCrG rUrGrUrGrUrU rUrUrA rGrArG rCrUrA rUrGrC rU / AltR2 /
[0170] Guide RNA sequence AB: / AltR1 / rGrA rUrCrG rGrArA rGrArG rCrGrU rCrGrU rGrUrArGrUrU rUrUrA rGrArG rCrUrA rUrGrC rU / AltR2 /
[0171] DNA containing the i5 adapter sequence (target) was normalized using a dCas-nickel bead pull-down system having a guide RNA designed to target the i5 adapter sequence. Two samples (the first sample containing 88 femtomoles of target DNA and the second sample containing 44 femtomoles of target DNA) were normalized using 125 femtomoles of dCas enzyme ( Figure 2 )。After normalization, the first and second samples were normalized to 4.3 femtomoles of target DNA and 4.2 femtomoles of target DNA, respectively.
[0172] Normalization is further demonstrated as follows: Dilution series of an ILLUMINA DNA sequencing library containing 4 nM, 6 nM, 8 nM, and 10 nM input DNA can be normalized to 4.68 nM, 5.4 nM, 5.55 nM, and 5.89 nM respectively( Figure 5 ). These results further demonstrate that this method can be used to normalize the concentration of target polynucleotides in a sample, where the sample has a difference of more than 2-fold in starting concentration. The results also show that the capture of DNA molecules containing the target is specifically driven by the presence of dCas9 RNP( Figure 4 ). Additionally, the amount of target DNA in each sample extracted using the normalization method can be linearly regulated by increasing the dCas9 RNP (dCas9 + Cas gRNA) concentration( Figure 3 ). Thus, one may be able to predict the amount of target polynucleotides that will be extracted from a sample based on the amount of dCas9 RNP used.
[0173] Method
[0174] 1. Form a ribonucleoprotein (RNP) binding complex
[0175] Combine 1 μM guide RNA, 1 μM dCas9 enzyme (6His-tagged), and 1X Cas9 dilution buffer, and allow the reaction mixture to incubate at room temperature for 5 - 10 minutes. This produces an RNP that is specific for the i5 sequence of the library adaptor and allows precise targeting of library molecules with dCas9.
[0176] 2. Perform a DNA binding reaction.
[0177] Combine 10X Cas9 reaction buffer, the desired nM of DNA library, 1 μM RNP complex, and dilute the 10X Cas9 reaction buffer to 1X using milliQ water. Incubate the reaction mixture at 37 °C for 1 hour. This allows dCas9 to bind to library molecules in a tight and specific manner, thus allowing subsequent capture using nickel magnetic beads (e.g., HisPur (Thermo Scientific)).
[0178] 3. Bead pull-down
[0179] Add HisPur magnetic beads and 1X Cas9 reaction buffer to each DNA binding reaction and incubate for 15 minutes. Remove the supernatant and wash the sample twice with 1X Cas reaction buffer to remove unbound DNA.
[0180] 4. Proteinase K digestion (elution)
[0181] Incubate the sample with Proteinase K at 56 °C for 10 minutes to release (elute) the bound DNA from the dCas enzyme.
[0182] 5. Quantification
[0183] Quantify the flow-through and the final eluted library by qPCR using library-specific primers.
[0184] Step 1.1 PCR-free library normalization achieved by Cas9-guide-mediated capture of library molecules.
[0185] Optionally, the method may include an additional step (Step 1.1) for generating a double-stranded PAM site on a single-stranded target polynucleotide without the need for PCR amplification ( Figure 7A -D). PCR amplification is a common strategy for constructing NGS libraries, especially as it allows NGS sequencing of small amounts of sample.
[0186] However, PCR amplification can be problematic for many NGS applications. For example, PCR amplification can introduce GC bias, which can interfere with data analysis, such as the identification of novel single nucleotide polymorphisms (SNPs). To allow library normalization as described in Example 1 while avoiding PCR amplification, the workflow described in this example introduces steps of library molecule denaturation and partial extension. Step 1.1 is performed prior to combining the RNP binding complex with the target nucleotide.
[0187] This step starts with a sample having a target polynucleotide comprising 3' and 5' adapter sequences ( Figure 7A ). The reverse complement of the PAM site is encoded on the 5' adapter. A partial primer (which is complementary to a portion of the 3' adapter sequence such that when the partial primer is extended, it does produce the reverse complement of the entire 3' adapter sequence) is contacted with the target polynucleotide ( Figure 7B ) under conditions such that primer extension can occur ( Figure 7C ), e.g., the conditions include the presence of a DNA polymerase. This generates a double-stranded PAM site that can bind to the dCas protein used in the normalization method as described herein ( Figure 7D ). However, the extended primer does not contain the complete 3' adapter and thus is not sequenced during next-generation sequencing. Thus, the sequencing results are not biased by amplified DNA (e.g., the extended primer).
[0188] Method
[0189] To generate a partial second strand that will enable, for example, Illumina PCR-free library normalization, the following method can be used:
[0190] Combine the ligated and unamplified library with an excess of primers specific for the i7 adapter sequence, such as: 5'GTGACTGGAGTTCAGACGTGT'3 (SEQ ID NO:49). This primer binds to the proximal part of the i7 adapter and critically does not contain the flow cell binding sequence from the distal adapter part. After primer binding, the primer is extended using polymerase in the presence of dNTPs. The resulting extended strand will not be able to bind to the flow cell and cluster and is thus inert. After extension, subject the library to the normalization workflow as described in this specification.
[0191] Advantages of the method:
[0192] ● Only the original non-PCR molecules will be able to cluster.
[0193] ● This method ensures that only fully ligated library molecules (molecules with 5' and 3' adapters ligated to the insert) will be targeted by dCas9 (or other nuclease-guided nucleases) in subsequent normalization steps.
[0194] Detailed experimental plan.
[0195] 1. Combine 100 femtomoles of the ligated library with 500 femtomoles of the i7 complement primer (5'GTGACTGGAGTTCAGACGTGT'3 (SEQ ID NO:49)) in the presence of 1X PCR buffer (Tris-HCl pH 8: 20 mM, magnesium chloride (MgCl2): 2 mM, potassium chloride (KCl): 50 mM, dNTP (each): 200 μM and 1 unit of Taq polymerase). Any other polymerase such as Bst, KOD, etc. can be used here.
[0196] 2. Subject the sample to a single denaturation extension step: 95°C for 30 seconds, followed by 60°C for 1 minute.
[0197] 3. Optionally, purify the extended duplex using SPRI beads (Ampure XP or equivalent) and subject the extended library to normalization using the method described in the examples.
[0198] 4. Only the original ligated library molecules can cluster on the NGS sequencer because they contain both the 5' adapter and the 3' adapter.
[0199] Example 2 - Normalization of Samples for NGS Using Catalytically Inactive Cas9 Protein Binding and Exonuclease Digestion
[0200] An NGS library is provided. For the NGS library, the exact concentration of the nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The nucleic acids in the library have adapter sequences at the first and second ends of the nucleic acid (e.g., ILLUMINA P5 / P7 sequences and a shared Y-adapter sequence (e.g., 13 bp)). The library is combined with a predetermined amount (80 fmol) of a catalytically inactive Cas9 protein having D10A and H840A substitutions. The catalytically inactive Cas9 protein is pre-loaded with two different guide RNAs or guide DNAs, each of which is specific for one of the adapter sequences present on the nucleic acids in the NGS library (e.g., the P5 sequence, the P7 sequence, or the shared Y-adapter sequence). Alternatively, the catalytically inactive Cas9 protein is pre-loaded with a single guide RNA or guide DNA specific for the Y-adapter sequence present on the nucleic acids in the NGS library or a single guide RNA or guide DNA specific for the same other sequence on both ends of the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-directed) association of the catalytically inactive Cas9 protein with the adapter sequence. This incubation results in a specific number of library molecules (80 fmol) being bound by the catalytically inactive Cas9 protein at their ends. Under conditions that maintain specific binding of the catalytically inactive Cas9 protein to the adapter sequence but also permit non-targeted nuclease activity, a non-targeted nuclease (Exonuclease III) is added to the reaction mixture. Only a specific number (80 fmol) of the library molecules bound by the catalytically inactive Cas9 protein at their ends remain in the resulting NGS library, and the remaining unprotected molecules are partially or completely digested by the nuclease.
[0201] These steps are performed simultaneously or sequentially on one or more additional NGS libraries. For the one or more additional NGS libraries, the exact concentration of the nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The resulting NGS libraries (which now have more similar concentrations between the NGS libraries compared to the starting concentration) are pooled and purified to remove the catalytically inactive Cas9 protein from the nucleic acids. If the desired concentration is reached, the nucleic acids in these libraries are sequenced. If the nucleic acids in the library have not reached the desired concentration, the nucleic acids are amplified and then sequenced.
[0202] Example 3 - Normalization of Samples for NGS by Binding with Catalytically Inactive CbAgo Protein and Exonuclease Digestion
[0203] An NGS library is provided. For the NGS library, the exact concentration of the nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The nucleic acids in the library have adapter sequences at the first and second ends of the nucleic acid (e.g., ILLUMINA P5 / P7 sequences and a shared Y-adapter sequence (e.g., 13 bp)). The library is combined with a predetermined amount (80 fmol) of catalytically inactivated CbAgo protein. The catalytically inactivated CbAgo protein is pre-loaded with one or more guide DNAs specific for the adapter sequences present on the nucleic acids in the NGS library (e.g., P5 sequence, P7 sequence, or shared Y-adapter sequence). The reaction mixture is incubated under conditions that promote specific and strong (guide-guided) association of the catalytically inactivated CbAgo protein with the adapter sequences. This incubation causes a specific number of library molecules (80 fmol) to be bound by the catalytically inactivated CbAgo protein at their ends. Under conditions that maintain tight binding of the catalytically inactivated CbAgo protein to the adapter sequences but also allow non-targeted nuclease activity, a non-targeted nuclease (exonuclease III) is added to the reaction mixture. Only a specific number (80 fmol) of the library molecules bound by the catalytically inactivated CbAgo protein at their ends remain in the resulting NGS library.
[0204] These steps are performed simultaneously or sequentially on one or more additional NGS libraries. For the one or more additional NGS libraries, the exact concentration of the nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The resulting NGS libraries (which now have more similar concentrations between the NGS libraries compared to the starting concentration) are pooled and purified to remove the catalytically inactivated CbAgo protein from the nucleic acids. If the desired concentration is reached, the nucleic acids in these libraries are sequenced. If the nucleic acids in the library have not reached the desired concentration, the nucleic acids are amplified and then sequenced.
[0205] Example 4 - Alternative method for normalizing samples for NGS using other nucleases
[0206] The normalization method is performed as described in Example 2 or 3, except that the non-target nuclease (exonuclease III) is replaced with an exonuclease I, a restriction enzyme, or a sequence-specific nuclease. For example, the nuclease is a 5'-to-3' exonuclease or a 3'-to-5' exonuclease, a single-strand specific nuclease, a double-strand specific nuclease, or a mixture thereof. Alternatively, the non-target nuclease is replaced with a nuclease guided by a catalytically active nucleic acid (e.g., Cas9, CbAgo). In some embodiments of the method, the non-target nuclease is replaced with a nuclease-guided nickase (e.g., Cas9 nickase). In certain embodiments, the nuclease-guided nickase is a Cas9 nickase having a D10A mutation or an H840A mutation. In some embodiments, the non-target nuclease is replaced with a transcription activator-like effector nuclease (TALEN) or a TALE nickase. In certain embodiments, the TALEN or the TALE nickase targets the same adapter sequence as the nucleic acid-binding protein (e.g., dCas protein, dAgo, etc.) used in the method.
[0207] Example 5 - Normalization of Samples for NGS Using Binding of Catalytically Inactive Cas9 Protein and Digestion with Cas9 Nickase
[0208] An NGS library is provided for which the exact concentration of the nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The nucleic acids in the library have adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y-adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acid. The library is combined with a predetermined amount (80 fmol) of a nucleic acid-guided binding protein to produce a reaction mixture. The nucleic acid-guided binding protein is a catalytically inactive Cas9 protein having D10A and H840A substitutions. The catalytically inactive Cas9 protein is pre-loaded with a single guide RNA or guide DNA specific for the shared Y-adapter sequence present on the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-guided) association of the catalytically inactive Cas9 protein with the adapter sequence. This incubation results in a specific number of library molecules (80 fmol) being bound by the catalytically inactive Cas9 protein at their ends. Cas9 nickase is added to the reaction mixture under conditions that maintain specific binding of the catalytically inactive Cas9 protein to the adapter sequence but also permit targeted Cas9 nickase activity. Only a specific number (80 fmol) of the library molecules bound by the catalytically inactive Cas9 protein at their ends remain in the resulting NGS library, and the remaining unprotected molecules are digested, partially or completely, by Cas9 nickase.
[0209] These steps are performed simultaneously or sequentially on one or more additional NGS libraries, for which the exact concentration of nucleic acid is unknown, but the concentration is at least 20 nM (20 fmol / μL). The resulting NGS libraries, which now have more similar concentrations between the NGS libraries than the starting concentration, are pooled and purified to remove protein from the nucleic acid. If the desired concentration is reached, the nucleic acid in these libraries is sequenced. If the nucleic acid in the library has not reached the desired concentration, the nucleic acid can be amplified and then sequenced.
[0210] Example 6 - Alternative Method for Normalizing Samples for NGS That Includes Amplification Prior to Digestion
[0211] An NGS library is provided for which the exact concentration of nucleic acid is unknown. The nucleic acid in the library has adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y - adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acid. Amplification is performed with primers that bind to the adapter sequences, thereby generating an amplified library.
[0212] The amplified library is combined with a predetermined amount (80 fmol) of a nucleic acid - directed binding protein to produce a reaction mixture. The nucleic acid - directed binding protein is dCas9, a catalytically inactivated Cas9 protein with D10A and H840A substitutions. The catalytically inactivated Cas9 protein is pre - loaded with a guide RNA or guide DNA that is specific for an adapter sequence (e.g., P5 sequence, P7 sequence, or shared Y - adapter sequence) present on the nucleic acid in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide - directed) association of the catalytically inactivated Cas9 protein with the adapter sequence. This incubation causes a specific number of library molecules to be bound by the catalytically inactivated Cas9 protein at their ends. A non - target nuclease (exonuclease III) is added to the reaction mixture under conditions that maintain tight binding of the catalytically inactivated Cas9 protein to the adapter sequence but also allow non - target nuclease activity. Only a specific number of the library molecules that are bound by the catalytically inactivated Cas9 protein at their ends remain in the resulting NGS library.
[0213] These steps are performed simultaneously or sequentially on one or more additional NGS libraries. That is, dCas9 is combined with individual libraries, the protected libraries are pooled, and the pooled protected libraries are exposed to an excess of nucleases (e.g., exonuclease III, exonuclease I, restriction enzymes, sequence-specific nucleases, 5'-to-3' exonuclease, 3'-to-5' exonuclease, single-strand-specific nuclease, double-strand-specific nuclease, catalytically active nuclease-guided nucleases (e.g., Cas9, CbAgo), nuclease-guided nickases (e.g., Cas9 nickase), transcription activator-like effector nucleases (TALEN) or TALE nickases or mixtures thereof).
[0214] The resulting NGS libraries (which now have more similar concentrations between the NGS libraries than the starting concentration) are pooled, and the pooled NGS libraries are subjected to standard SPRI purification. The nucleic acids in the library are then sequenced.
[0215] Example 7 - Normalization of Sequencing Libraries after Amplification
[0216] Typical Illumina sequencer library preparation chemistry targets a final amplified library molarity between 200 femtomoles (fmol) and 1,000 fmol. To demonstrate the ability of library normalization across a range of inputs, an Illumina TruSeq adapter-ligated double-stranded DNA (dsDNA) library was amplified to produce a high concentration DNA input. DNA input amounts of 50 fmol, 500 fmol, 750 fmol, or 1,000 fmol of this DNA were used as inputs for two different normalization reactions containing 250 fmol or 300 fmol of ribonucleoprotein (RNP, e.g., dCas9-guide RNA complex). The resulting normalized libraries were quantified using an HSD1000 DNA tape on an Agilent Tapestation with a region of 100 - 1,000 base pairs. Both 250 fmol and 300 fmol of RNP normalized the range of DNA library inputs, where the resulting amount of the normalized output depends on the amount of RNP. Between a low DNA input of 50 fmol and a high DNA input of 1,000 fmol, there is a 20-fold difference in DNA amount. After normalization with 250 fmol RNP, there is only a 1.5-fold difference, and after normalization with 300 fmol RNP, there is only a 1.3-fold difference. Between all samples normalized with 250 fmol RNP, there is an average output of 37 fmol, while normalization with 300 fmol RNP produces an average output of 56 fmol, indicating that the amount of library retained during normalization can be regulated by the amount of RNP used ( Figure 8 ).
[0217] Example 8 - Normalization of Sequencing Library Pre-Targeted Capture
[0218] Targeted capture / enrichment groups typically require high-quality amplified libraries, usually in the range of 100 ng - 1,000 ng of library molecules. Although it depends on the size of the library, such a quality range generally corresponds to 300 - 2,000 fmol of DNA. To demonstrate the ability to normalize to achieve this final output range, an Illumina TruSeq adapter-ligated dsDNA library was amplified to produce a high concentration of DNA input, and 3,400 fmol of this DNA was used as the input. Normalization was performed with 2,000 fmole, 6,000 fmole, 10,000 fmole, or 14,000 fmole of RNP. The output DNA was analyzed on an Agilent Tapestation using the D1000 tape with a region setting of 180 - 1,000 base pairs. Increasing the RNP resulted in more DNA library being captured. By increasing the RNP to 6,000 fmol, approximately 2,000 fmol of the library was captured, supporting the upstream application of normalization upstream of target enrichment and similar applications( Figure 9 ).
[0219] Example 9. Normalization Using Single Guide RNA (sgRNA) in RNP
[0220] Using a single guide RNA (sgRNA) as opposed to separate CRISPR RNAs (crRNAs) and trans-activating CRISPR RNAs (tracrRNAs) provides a simpler format for normalization. To demonstrate performance similar to that of crRNA / tracrRNA RNPs, an Illumina TruSeq adapter-ligated dsDNA library was amplified to generate a high concentration of DNA input, and 500 fmol of this DNA was used. Normalization was performed with 250 fmol, 600 fmol, or 1,000 fmol of RNPs containing crRNA / tracrRNA (duplex guide) or sgRNA. The sgRNA sequence was: 5'-mA*mG*mA*rUrCrG rGrArA rGrArG rCrGrUrCrGrU rGrUrG rUrUrU rUrArG rArGrCrUrArG rArArA rUrArG rCrArA rGrUrU rArArArArUrA rArGrG rCrUrA rGrUrC rCrGrU rUrArU rCrArA rCrUrU rGrArA rArArA rGrUrGrGrCrA rCrCrG rArGrU rCrGrGrUrGrC mU*mU*mU*rU-3' (SEQ ID NO:48) (mA / mG / mU are 2'-O-methyl RNA bases; * indicates phosphorothioate bond).
[0221] The crRNA sequence was:
[0222] 5'- / AltR1 / rGrCrA rCrGrG rArGrA rCrGrG rArUrG rUrUrA rUrUrGrUrUrUrUrArG rArGrC rUrArUrGrCrU / AltR2 / -3' (SEQ ID NO:50) ( / AltR1 / and / AltrR2 / are proprietary IDT modifications).
[0223] Output DNA was analyzed using the D1000 tape on an Agilent Tapestation with a region setting of 180 - 1,000 base pairs. RNPs containing sgRNA were found to behave similarly to RNPs containing crRNA / tracrRNA. RNPs containing sgRNA retained approximately 20% less DNA than RNPs containing crRNA / tracrRNA. However, increasing the fmol of sgRNA RNP produced more captured DNA, thus allowing the range demonstrated with crRNA / tracrRNA to be achieved with sgRNA ( Figure 10 ).
[0224] These embodiments are provided to illustrate the disclosure and not to limit its scope. Other variations of the disclosure will be apparent to those of ordinary skill in the art and are covered by the appended claims.
Claims
1. A method for normalizing the concentration of a target polynucleotide between at least two samples, each of the at least two samples containing the target polynucleotide, and for each of the at least two samples, the method comprises: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adaptor sequence; (ii) generating a solution, the generating comprising combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adaptor sequence of the target polynucleotide of the sample; and (c) dCas or dArgonaute at a predetermined concentration, the dCas or the dArgonaute comprising an affinity tag, wherein the predetermined concentration of the dCas or the dArgonaute is homologous to the guide polynucleotide; (iii) contacting the solution with a solid phase comprising an affinity tag-binding molecule capable of binding to the affinity tag; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
2. The method according to claim 1, wherein the targeting region is complementary to a static region of the adaptor sequence.
3. The method according to claim 1 or claim 2, wherein the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
4. The method according to claim 2, wherein the static region is a static region of a next-generation sequencing adaptor sequence.
5. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 27-36.
6. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 27-28.
7. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 29-30.
8. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 31-32.
9. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to an adaptor sequence shown in any one of SEQ ID NOs: 33-34.
10. The method according to any one of claims 1 to 5, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
35.
11. The method according to any one of claims 1 to 4, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
36.
12. The method according to any one of claims 1 to 11, wherein the Cas gRNA polynucleotide targeting region is a homologous region.
13. The method according to claim 12, wherein the Cas gRNA polynucleotide comprises a homology region as shown in any one of SEQ ID NOs: 1-26.
14. The method according to claim 13, wherein the Cas gRNA polynucleotide comprises the homology region as shown in SEQ ID NO:
1.
15. The method according to claim 1, wherein the Argonaute guide polynucleotide is siRNA, miRNA, piRNA, shRNA or siDNA.
16. The method according to any one of claims 1 to 15, wherein the guide polynucleotide is a DNA polynucleotide.
17. The method according to any one of claims 1 to 16, wherein the guide polynucleotide is an RNA polynucleotide.
18. The method according to claim 16 or claim 17, wherein the guide polynucleotide comprises a modified nucleic acid.
19. The method according to claim 18, wherein the modified nucleic acid comprises 2'F RNA, 2'OMe RNA and / or phosphorothioate bond (PS).
20. The method according to claim 19, wherein the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not comprise a modified nucleic acid at a position interacting with a Cas protein homologous to the gRNA.
21. The method according to any one of claims 1 to 20, wherein the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX or CasY protein.
22. The method according to claim 21, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
23. The method according to any one of claims 1 to 11, wherein the dArgonaute is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo.
24. The method according to any one of claims 1 to 23, wherein the affinity tag binding molecule comprises Ni2+ and the affinity tag comprises a His tag.
25. The method according to any one of claims 1 to 23, wherein the affinity tag binding molecule comprises biotin and the affinity tag comprises avidin.
26. The method according to any one of claims 1 to 23, wherein the affinity tag binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag.
27. The method according to any one of claims 1 to 26, wherein the affinity tag and the corresponding affinity tag binding molecule are selected from Table 1.
28. The method according to any one of claims 1 to 27, wherein the solid phase comprises magnetic beads.
29. The method according to any one of claims 1 to 28, wherein separating the solution from the solid phase comprises immobilizing the solid phase and washing the solid phase.
30. The method according to any one of claims 1 to 29, wherein extracting the target polynucleotide from the solid phase comprises combining the solid phase and a protease.
31. The method according to any one of claims 1 to 29, wherein extracting the target polynucleotide from the solid phase comprises combining the solid phase and proteinase K in a solution sufficient to generate proteinase K activity.
32. The method according to claim 31, wherein proteinase K digests the dCas or dArgonaute bound to the solid phase, thereby extracting the target polynucleotide.
33. The method according to any one of claims 1 to 32, wherein steps (i)-(v) are carried out in sequence.
34. The method according to any one of claims 1 to 33, wherein normalization comprises bringing the concentration of the target polynucleotide between the at least two samples within 15% of each other after normalization.
35. A method for normalizing the concentration of a target polynucleotide between at least two samples, each of the at least two samples comprising a target polynucleotide, the method for each of the at least two samples comprises: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adaptor sequence; (ii) generating a solution, the generating comprising combining (a) the sample, (b) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adaptor sequence of the target polynucleotide of the sample, and (c) a dCas or dArgonaute at a predetermined concentration, the dCas or the dArgonaute comprising an affinity tag; and (iii) contacting the solution with a predetermined amount of catalytically active Cas or Argonaute protein to normalize the concentration of the target polynucleotide between the two or more samples.
36. The method according to claim 35, wherein the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY protein.
37. The method according to claim 36, wherein the dCas protein comprises an amino acid sequence having at least 95% identity with SEQ ID NO:
37.
38. The method according to claim 36, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
39. The method according to claim 35, wherein the dArgonaute is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo.
40. The method according to any one of claims 35 to 39, wherein (ii) comprises binding of the dCas or dArgonaute to the adaptor sequence of at least some of the target polynucleotides in the target polynucleotide.
41. The method according to claim 40, wherein the catalytically active Cas9 or Argonaute digests the target polynucleotide that is not bound by the dCas9 or the dArgonaute.
42. The method according to any one of claims 35 to 40, wherein the normalization comprises bringing the concentrations of the target polynucleotides between the at least two samples within 15% of each other after normalization.
43. The method according to any one of claims 35 to 40, wherein the normalization comprises bringing the concentrations of the target polynucleotides between the at least two samples within 10% of each other after normalization.
44. The method according to any one of claims 1 to 43, wherein the target polynucleotide comprises a first adaptor sequence and a second adaptor sequence, and wherein the guide polynucleotide targeting region is complementary to the first adaptor sequence.
45. The method according to claim 44, further comprising performing the following after (i) and before (ii): (a) contacting the sample with a primer encoding a nucleic acid sequence complementary to a portion of the second adaptor sequence, wherein the portion of the adaptor sequence is located in the proximal portion of the adaptor sequence; and (b) contacting the sample with a DNA polymerase under conditions sufficient to promote primer extension.
46. The method according to claim 45, wherein the portion of the adaptor sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34.
47. The method according to claim 45 or claim 46, wherein the second adaptor sequence is a 3' adaptor sequence.
48. The method according to claim 45 or claim 46, wherein the second adaptor sequence is a 5' adaptor sequence.
49. A guide polynucleotide comprising a targeting region that is complementary to the static region of an adaptor sequence shown in any one of SEQ ID NOs: 28 and 30, 32, 34 - 36.
50. The guide polynucleotide according to claim 49, wherein the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
51. The guide polynucleotide according to claim 49 or claim 50, wherein the static region is the static region of a next-generation sequencing adaptor sequence.
52. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
28.
53. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
30.
54. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
32.
55. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
34.
56. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
35.
57. The guide polynucleotide according to claim 51, wherein the targeting region is complementary to the adaptor sequence shown in SEQ ID NO:
36.
58. The guide polynucleotide according to any one of claims 51 to 57, wherein the Cas gRNA polynucleotide targeting region is a homologous region.
59. The guide polynucleotide according to claim 58, wherein the Cas gRNA polynucleotide comprises a homologous region shown in any one of SEQ ID NOs: 1-26.
60. The guide polynucleotide according to claim 59, wherein the Cas gRNA polynucleotide comprises the homologous region shown in SEQ ID NO:
1.
61. The guide polynucleotide according to any one of claims 50 to 57, wherein the Argonaute guide polynucleotide is siRNA, miRNA, piRNA, shRNA or siDNA.
62. The guide polynucleotide according to any one of claims 49 to 61, wherein the guide polynucleotide is a DNA polynucleotide.
63. The guide polynucleotide according to any one of claims 49 to 61, wherein the guide polynucleotide is an RNA polynucleotide.
64. The guide polynucleotide according to claim 62 or claim 63, wherein the guide polynucleotide comprises a modified nucleic acid.
65. The guide polynucleotide according to claim 64, wherein the modified nucleic acid comprises 2'F RNA, 2'Ome RNA and / or phosphorothioate bond (PS).
66. The guide polynucleotide according to claim 65, wherein the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not comprise a modified nucleic acid at the position interacting with the Cas protein, and the Cas protein and the gRNA are homologous.
67. A kit, comprising: (a) The guide polynucleotide according to any one of claims 49 to 66; and (b) A Cas protein or an Argonaute protein, or a polynucleotide sequence encoding a Cas protein or an Argonaute protein.
68. The kit according to claim 67, wherein the Cas protein is a catalytically inactivated Cas protein (dCas).
69. The kit according to claim 67 or claim 68, wherein the Cas protein is a catalytically inactivated Cas9 protein (dCas9).
70. The kit according to any one of claims 67 to 69, wherein the Cas9 protein comprises D10A mutation and H840A mutation.
71. The kit according to any one of claims 67 to 68, wherein the Cas protein is a catalytically inactivated Cpf1, C2c1, C2c3, C2c2, CasX or CasY protein.
72. The kit according to any one of claims 68 to 71, wherein the dCas protein comprises an amino acid sequence having at least 95% identity with SEQ ID NO:
37.
73. The kit according to any one of claims 68 to 71, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
74. The kit according to claim 67, wherein the Argonaute protein is catalytically inactive (dArgonaute).
75. The kit according to claim 67, wherein the Argonaute protein is catalytically inactive CbAgo protein.
76. The kit according to claim 67, wherein the Argonaute protein is catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo protein.
77. The kit according to any one of claims 74 to 76, wherein the dArgonaute protein comprises an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 39 - 45.
78. The kit according to any one of claims 74 to 76, wherein the dArgonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 39 - 45.
79. The kit according to any one of claims 67 to 78, further comprising a catalytically active Cas protein or a catalytically active Argonaute protein.
80. The kit according to claim 79, wherein the catalytically active Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY.
81. The kit according to claim 80, wherein the catalytically active Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo.
82. The kit according to any one of claims 67 to 81, wherein the guide polynucleotide is capable of binding to the Cas protein or the Argonaute protein to form a ribonucleoprotein complex, and the ribonucleoprotein complex is capable of binding to the adapter sequence.
83. The kit according to any one of claims 67 to 82, further comprising a primer complementary to a part of the adapter sequence.
84. The kit according to claim 83, wherein the part of the adapter sequence is located in the proximal part of the adapter sequence.
85. The kit according to claim 84, wherein the portion of the adapter sequence is located in the proximal portion of the adapter sequence, and the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NO: 28, 30, 32, and 34.
86. The kit according to any one of claims 83 to 85, wherein the adapter sequence is a 3'-adapter sequence.
87. The kit according to any one of claims 83 to 85, wherein the adapter sequence is a 5'-adapter sequence.
88. The kit according to any one of claims 83 to 87, wherein the primer is not complementary to the adapter sequence, and the adapter sequence is complementary to the guide polynucleotide targeting region.
89. A reaction mixture comprising: (i) a plurality of target polynucleotides, wherein the target polynucleotide comprises an adapter sequence shown in any one of SEQ ID NO: 28, 30, 32, and 34-36; (ii) a guide polynucleotide at a predetermined concentration, the guide polynucleotide comprising a targeting region complementary to the adapter sequence; and (iii) a dCas protein or a dArgonaute protein at a predetermined concentration.
90. The reaction mixture according to claim 89, wherein the predetermined concentration of the guide polynucleotide or the dCas protein or the dArgonaute protein is lower than the concentration of the target polynucleotide in the reaction mixture.
91. The reaction mixture according to claim 89 or claim 90, wherein the dArgonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo or a catalytically inactivated variant thereof, and the guide polynucleotide is homologous to the dArgonaute protein.
92. The reaction mixture according to claim 89 or claim 90, wherein the dCas protein comprises a catalytically inactivated Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY or a variant thereof, and the guide polynucleotide is homologous to the dCas protein.
93. The reaction mixture according to any one of claims 89 to 92, wherein the adapter sequence comprises the polynucleotide sequence of any one of SEQ ID NO: 28, 30, 32, and 34-36.
94. The reaction mixture according to any one of claims 89 to 93, further comprising a catalytically active Cas protein or a catalytically active Argonaute protein.
95. A ribonucleoprotein (RNP) complex comprising: (i) a guide polynucleotide according to any one of claims 49 to 66; (ii) a Cas protein or an Argonaute protein, the Cas protein or the Argonaute protein being homologous to the guide polynucleotide; and (iii) an adapter sequence The targeting region of the guide polynucleotide is complementary to the adaptor sequence, and the guide polynucleotide is homologous to the Cas protein or the Argonaute protein.
96. The RNP complex according to claim 95, wherein the Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo or a catalytically inactivated variant thereof.
97. The RNP complex according to claim 95, wherein the Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY or a catalytically inactivated variant thereof.
98. The RNP complex according to any one of claims 95 to 97, wherein the adaptor sequence comprises the polynucleotide sequence of any one of SEQ ID NO: 28, 30, 32 and 34-36.