Methods and compositions for sequencing library normalization
The use of a ribonucleoprotein complex to bind and normalize adapter sequences in polynucleotide libraries addresses the inaccuracies in existing methods, achieving consistent sequencing results by adjusting nucleic acid concentrations.
Patent Information
- Application Number
- JP2025522560
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-27
- Filing Date
- 2023-10-20
- Publication Date
- 2025-10-24
AI Technical Summary
Existing methods for normalizing nucleic acid concentrations in sequencing libraries are inaccurate, time-consuming, and expensive, leading to under- or over-sequencing of samples due to varying nucleic acid amounts.
A method involving the use of a predetermined amount of ribonucleoprotein, such as a catalytically inactive CRISPR-associated protein (dCas)-guide RNA (gRNA) complex, to bind to adapter sequences in polynucleotide libraries, allowing for the extraction and normalization of polynucleotide concentrations between samples.
This method effectively normalizes polynucleotide concentrations between samples, ensuring accurate sequencing by bringing concentrations within 15% of each other, thereby improving sequencing accuracy and efficiency.
Smart Images

Figure 2025535369000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 63 / 380,488, filed October 21, 2022, entitled "Methods for Sequencing Library Normalization," and U.S. Provisional Patent Application No. 63 / 516,033, filed July 27, 2023, entitled "Methods and Compositions for Sequencing Library Normalization," the entire contents of each of which are incorporated herein by reference.
[0002] Reference to electronic sequence listing The contents of the electronic sequence listing (W109470000WO00-SEQ-ARM.xml; size: 53,547 bytes; created on: October 20, 2023) are incorporated herein by reference in their entirety. [Background technology]
[0003] background Massively parallel, deep sequencing, or next-generation sequencing (NGS), enables large-scale DNA and RNA sequencing. These methods involve the parallel multiplex analysis of large numbers of nucleic acid sequences, allowing millions to billions of sequences from individual strands to be analyzed individually but simultaneously.
[0004] Multiplex analysis can be difficult because the amount of nucleic acid present in the samples analyzed often varies, which can lead to inaccuracies, such as low-concentration samples being under-sequenced and high-concentration samples being over-sequenced.Therefore, various methods have been utilized to normalize the nucleic acid concentration between samples.Spectrophotometry, electrophoresis, fluorometry, and quantitative PCR (qPCR) have been used to detect the nucleic acid concentration in samples to normalize the concentration between samples.Each of these methods has drawbacks, such as the inaccuracy introduced by manual adjustment, lack of sensitivity, and the large number of steps that require considerable time and / or expense. Summary of the Invention [Means for solving the problem]
[0005] Abstract In some aspects, the present disclosure describes methods and compositions for normalizing the polynucleotide concentration between two or more polynucleotide sequencing libraries.Preparing polynucleotides for sequencing (for example, next-generation sequencing (NGS)) often involves adding an adapter polynucleotide sequence to polynucleotides.Certain technologies, such as ILLUMINA sequencing, use adapter sequences to sequence polynucleotides.It is recognized herein that these adapter sequences can be utilized to normalize the polynucleotide concentration between different polynucleotide sequencing libraries.Specifically, it is recognized that polynucleotide sequencing library normalization can be achieved by contacting a polynucleotide sequencing library with a predetermined amount of ribonucleoprotein that binds to the adapter sequence of the polynucleotide library (for example, catalytically inactive CRISPR-associated protein (dCAS)-guide RNA (gRNA) complex), and extracting the polynucleotides that (1) contain adapters and (2) are bound by the ribonucleoprotein complex. By performing this method on a number of different polynucleotide sequencing libraries, it is demonstrated that the concentrations of polynucleotides between polynucleotide sequencing libraries can be normalized.
[0006] Accordingly, the present disclosure provides, in some aspects, a method for normalizing the concentration of a target polynucleotide between at least two samples, each comprising a target polynucleotide, the method comprising, for each sample of the at least two samples: (i) obtaining a sample, wherein a target polynucleotide of the sample comprises an adapter sequence; (ii) forming a solution comprising: (a) a sample; (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to an adapter sequence of a target polynucleotide of the sample; and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag, wherein the predetermined concentration of dCas or dArgonaute is cognate to the guide polynucleotide; (iii) contacting the solution with a solid phase comprising an affinity tag binding molecule capable of binding to the affinity tag; (iv) separating the solution from the solid phase; (v) extracting the target polynucleotide from the solid phase and normalizing the concentration of the target polynucleotide between the two or more samples; In some embodiments, the targeting region is complementary to a static region of the adapter sequence. In some embodiments, the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide. In some embodiments, the static region is a static region of a next-generation sequencing adapter sequence. In some embodiments, the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 27-36. In some embodiments, the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 27-28. In some embodiments, the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 29-30. In some embodiments, the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 31-32. In some embodiments, the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 33-34. In some embodiments, the targeting region is complementary to an adapter sequence of SEQ ID NO: 35. In some embodiments, the targeting region is complementary to an adapter sequence of SEQ ID NO: 36.
[0007] In some embodiments, the Cas gRNA polynucleotide targeting region is a homologous region. In some embodiments, the Cas gRNA polynucleotide comprises a homologous region of any one of SEQ ID NOs: 1-26. In some embodiments, the Cas gRNA polynucleotide comprises a homologous region of SEQ ID NO: 1.
[0008] In some embodiments, the Argonaute guide polynucleotide is an siRNA, miRNA, piRNA, shRNA, or siDNA.
[0009] In some embodiments, the guide polynucleotide is a DNA polynucleotide. In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide comprises a modified nucleic acid. In some embodiments, the modified nucleic acid comprises 2'F RNA, 2'OMe RNA, and / or phosphorothioate linkages (PS).
[0010] In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not contain a modified nucleic acid at a position that interacts with the gRNA's cognate Cas protein. In some embodiments, the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, or CasY protein. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO: 37.
[0011] In some embodiments, the dArgonaute is catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo.
[0012] In some embodiments, the affinity tag binding molecule comprises Ni2+ and the affinity tag comprises a His tag.
[0013] In some embodiments, the affinity tag binding molecule comprises biotin and the affinity tag comprises avidin. In some embodiments, the affinity tag binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag. In some embodiments, the affinity tag and the corresponding affinity tag binding molecule are selected from Table 1.
[0014] In some embodiments, the solid phase comprises magnetic beads. In some embodiments, separating the solution from the solid phase comprises immobilizing the solid phase and washing the solid phase. In some embodiments, extracting the target polynucleotide from the solid phase comprises combining the solid phase with a protease. In some embodiments, extracting the target polynucleotide from the solid phase comprises combining the solid phase with proteinase K in a solution sufficient for proteinase K activity. In some embodiments, proteinase K digests dCas or dArgonaute bound to the solid phase, thereby extracting the target polynucleotide.
[0015] In some embodiments, steps (i) through (v) are performed in sequential order.
[0016] In some embodiments, the normalizing step comprises having target polynucleotide concentrations between at least two samples within 15% of each other after normalization.
[0017] In some aspects, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples, each comprising a target polynucleotide, the method comprising: for each sample of the at least two samples: (i) obtaining a sample, wherein a target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution comprising combining (a) a sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region complementary to an adapter sequence of a target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag; (iii) contacting the solution with a predetermined amount of catalytically active Cas or Argonaute protein to normalize the concentration of the target polynucleotide between two or more samples; The present invention provides a method comprising:
[0018] In some embodiments, the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY protein. In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO: 37.
[0019] In some embodiments, the dArgonaute is catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo.
[0020] In some embodiments, step (ii) comprises binding of dCas or dArgonaute to at least a portion of the adaptor sequence of the target polynucleotide. In some embodiments, catalytically active Cas9 or Argonaute digests target polynucleotides not bound by dCas9 or dArgonaute. In some embodiments, the normalizing step comprises bringing the target polynucleotide concentrations between at least two samples within 15% of each other after normalization. In some embodiments, the normalizing step comprises bringing the target polynucleotide concentrations between at least two samples within 10% of each other after normalization. In some embodiments, the target polynucleotide comprises a first adaptor sequence and a second adaptor sequence, and the guide polynucleotide targeting region is complementary to the first adaptor sequence.
[0021] In some embodiments, the method further comprises, after step (i) and before step (ii): (a) contacting the sample with a primer encoding a nucleic acid sequence that is complementary to a portion of a second adapter sequence, wherein the portion of the adapter sequence is located proximal to the adapter sequence; and (b) contacting the sample with a DNA polymerase under conditions sufficient to promote primer extension; Further includes:
[0022] In some embodiments, a portion of the adaptor sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34.
[0023] In some embodiments, the second adapter sequence is a 3' adapter sequence. In some embodiments, the second adapter sequence is a 5' adapter sequence.
[0024] In some aspects, the disclosure provides a guide polynucleotide comprising a targeting region complementary to an adapter sequence of any one of SEQ ID NOs: 28 and 30, 32, 34-36. In some embodiments, the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
[0025] In some embodiments, the static region is a static region of a next-generation sequencing adapter sequence. In some embodiments, the targeting region is complementary to the adapter sequence of SEQ ID NO: 28. In some embodiments, the targeting region is complementary to the adapter sequence of SEQ ID NO: 30. In some embodiments, the targeting region is complementary to the adapter sequence of SEQ ID NO: 32. In some embodiments, the targeting region is complementary to the adapter sequence of SEQ ID NO: 34. In some embodiments, the targeting region is complementary to the adapter sequence of SEQ ID NO: 35. In some embodiments, the targeting region is complementary to any one of the adapter sequences of SEQ ID NO: 36.
[0026] In some embodiments, the Cas gRNA polynucleotide targeting region is a homologous region. In some embodiments, the Cas gRNA polynucleotide comprises a homologous region of any one of SEQ ID NOs: 1-26. In some embodiments, the Cas gRNA polynucleotide comprises a homologous region of SEQ ID NO: 1.
[0027] In some embodiments, the Argonaute guide polynucleotide is an siRNA, miRNA, piRNA, shRNA, or siDNA.
[0028] In some embodiments, the guide polynucleotide is a DNA polynucleotide. In some embodiments, the guide polynucleotide is an RNA polynucleotide. In some embodiments, the guide polynucleotide comprises a modified nucleic acid. In some embodiments, the modified nucleic acid comprises 2'F RNA, 2'Ome RNA, and / or phosphorothioate linkages (PS). In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide, and the Cas gRNA polynucleotide does not comprise a modified nucleic acid in a position that interacts with the cognate Cas protein for the gRNA.
[0029] In some embodiments, the present disclosure provides a kit comprising (a) a guide polynucleotide described herein and (b) a Cas protein or Argonaute protein, or a polynucleotide sequence encoding a Cas protein or Argonaute protein. In some embodiments, the Cas protein is a catalytically inactive Cas protein (dCas). In some embodiments, the Cas protein is a catalytically inactive Cas9 protein (dCas9). In some embodiments, the Cas9 protein comprises a D10A mutation and an H840A mutation. In some embodiments, the Cas protein is a catalytically inactive Cpfl, C2cl, C2c3, C2c2, CasX, or CasY protein. In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO: 37.
[0030] In some embodiments, the Argonaute protein is catalytically inactive (dArgonaute). In some embodiments, the Argonaute protein is a catalytically inactive CbAgo protein. In some embodiments, the Argonaute protein is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo protein. In some embodiments, the dArgonaute protein comprises an amino acid sequence having at least 95% identity to any one of SEQ ID NOs: 39-45. In some embodiments, the dArgonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 39-45.
[0031] In some embodiments, the kit further comprises a catalytically active Cas protein or a catalytically active Argonaute protein. In some embodiments, the catalytically active Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY. In some embodiments, the catalytically active Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo. In some embodiments, the guide polynucleotide can bind to the Cas protein or Argonaute protein to form a ribonucleoprotein complex, and the ribonucleoprotein complex can bind to an adapter sequence.
[0032] In some embodiments, the kit further comprises a primer complementary to a portion of the adapter sequence. In some embodiments, the portion of the adapter sequence is located proximal to the adapter sequence. In some embodiments, the portion of the adapter sequence is located proximal to the adapter sequence, and the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34. In some embodiments, the adapter sequence is a 3' adapter sequence. In some embodiments, the adapter sequence is a 5' adapter sequence. In some embodiments, the primer is not complementary to the adapter sequence complementary to the guide polynucleotide targeting region.
[0033] In some aspects, the present disclosure provides a reaction mixture comprising: (i) a plurality of target polynucleotides, wherein the target polynucleotides comprise an adapter sequence of any one of SEQ ID NOs: 28, 30, 32, and 34-36; (ii) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence; and (iii) a predetermined concentration of a dCas protein or a dArgonaute protein.
[0034] In some embodiments, the predetermined concentration of guide polynucleotide, or dCas protein or dArgonaute protein is less than the concentration of target polynucleotide in the reaction mixture.
[0035] In some embodiments, the d-Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo, or a catalytically inactive variant thereof, and the guide polynucleotide is cognate to the d-Argonaute protein. In some embodiments, the dCas protein comprises catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY, or a variant thereof, and the guide polynucleotide is cognate to the dCas protein. In some embodiments, the adapter sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 28, 30, 32, and 34-36. In some embodiments, the reaction mixture further comprises a catalytically active Cas protein or a catalytically active Argonaute protein.
[0036] In some embodiments, the present disclosure provides a ribonucleoprotein (RNP) complex comprising: (i) a guide polynucleotide described herein; (ii) a Cas protein or Argonaute protein cognate to the guide polynucleotide; and (iii) an adapter sequence, wherein a targeting region of the guide polynucleotide is complementary to the adapter sequence, and the guide polynucleotide is cognate to the Cas protein or Argonaute protein. In some embodiments, the Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo, or a catalytically inactive variant thereof. In some embodiments, the Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY, or a catalytically inactive variant thereof. In some embodiments, the adapter sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 28, 30, 32, and 34-36. [Brief explanation of the drawings]
[0037] [Figure 1] Figure 1 shows a representative schematic of the polynucleotide capture process by CAS-gRNA RNPs bound to metal beads.
[0038] [Figure 2] Figure 2 shows the normalization of two different samples, each containing an i5-adapter-containing polynucleotide. Two different samples of i5-containing DNA fragments were subjected to normalization using the same mass (125 femtomoles, fmol) of dCas9. After elution following proteinase digestion, both samples contained equivalent numbers of molecules, demonstrating the concept of normalization. The ratio of molecules between the samples varied from 2:1 to 1:1, as intended. When these samples were pooled using equivalent volumes and sequenced by NGS, both samples were expected to produce equivalent numbers of clusters.
[0039] [Figure 3] Figure 3 shows that the amount of polynucleotides extracted from a sample can be adjusted by titrating the amount of ribonucleoprotein (e.g., RNP) that binds to the adaptor of the polynucleotide added to the sample. RNP (dCas9-gRNA complex) titration and subsequent bead binding and elution were performed using samples containing a fixed amount of DNA (160 fmol). Increasing the amount of RNP resulted in a linear increase in the amount of DNA bound (and ultimately recovered). At a DNA input level of 160 fmol, there was a linear relationship in the range of 125 fmol to 1000 fmol of RNP. These data demonstrate that the amount of DNA library extracted can be adjusted by adjusting the amount of RNP added to the sample.
[0040] [Figure 4] Figure 4 shows the specific and stoichiometric targeting and retention of target DNA molecules driven by interaction with dCas9 RNPs and subsequent binding to HisPur beads via 6His-tagged dCas9. Known quantities of i5-sequence-containing DNA fragments were bound to HisPur beads after contact with RNPs (RNP+) or without RNPs (RNP-). The majority of the library remained bound to the beads only after exposure to RNPs and elution after proteinase K digestion (compare RNP+ flow-through with RNP+ elution). In the absence of RNPs, all of the DNA was found in the flow-through, with none bound to the beads (compare RNP- flow-through with RNP- elution).
[0041] [Figure 5]Figure 5 shows the normalization of four different samples of PCR-amplified ILLUMINA libraries using dCas9-gRNA complexes. Four different samples of PCR-amplified ILLUMINA libraries (input) were subjected to normalization using the same mass of dCas9. After elution following proteinase digestion (elution), all samples contained significantly more uniform numbers of molecules. The highest concentration sample was 25% more concentrated than the lowest after normalization (4.68 nM elution vs. 5.89 nM elution), compared to 150% before normalization (4 nM input vs. 10 nM input). In addition, the majority of unbound library was retained in the flow-through for all libraries.
[0042] [Figure 6] Figure 6 shows a schematic demonstrating an exemplary embodiment of the disclosed normalization method. In this exemplary embodiment, two samples containing different concentrations of end-labeled target nucleic acids (Library A with a concentration of 9 fmol and Library B with a concentration of 18 fmol) are provided. In this exemplary embodiment, a reaction mixture is generated by adding a predetermined concentration of 3 fmol of guide polynucleotides and catalytically inactive Cas9 protein (which may be pre-assembled) to Library A and Library B, allowing the protein to bind to the adapter sequence at one end of the end-labeled target nucleic acids in the sample, thereby protecting that end of the target nucleic acid from nuclease digestion. An exonuclease is added to the reaction mixture to digest unprotected nucleic acids. The resulting Library A and Library B each contain 3 fmol of target nucleic acid. In other embodiments, the exonuclease in this scheme can be replaced with a nucleic acid-guided nuclease (e.g., a Cas9 nuclease), a nucleic acid-guided nickase (e.g., a Cas9 nickase), a restriction enzyme nuclease or nickase, or a transcription activator-like effector nuclease (TALEN), or another similar nuclease or nickase.
[0043] [Figure 7]Figures 7A-D show a schematic diagram of library molecule denaturation and partial sequence extension to generate a double-stranded PAM site. In Figure 7A, the library molecule is ligated to two Y-shaped adapters at the 3' and 5' ends. In Figure 7B, the library molecule is denatured and a partial primer is attached to the 3' end. In Figure 7C, the attached primer is extended by polymerase, resulting in the creation of a complementary strand (dashed line) that does not span the entire 3' adapter but includes a double-stranded full-length 5' adapter, which contains the PAM site for dCas9 targeting. Figure 7D illustrates the attachment of a dCas9 molecule to the double-stranded PAM site.
[0044] [Figure 8] Figure 8 is a bar graph showing femtomoles of captured DNA (fmol out) after normalization and amplification for different amounts of starting DNA (fmol in) and two concentrations of ribonucleoprotein (RNP fmol).
[0045] [Figure 9] Figure 9 shows the fmol of DNA captured by various concentrations of ribonucleoprotein (RNP fmol) (fmol out). The input DNA concentration is 3,400 fmol.
[0046] [Figure 10] Figure 10 shows the fmol of DNA (fmol out) captured by various concentrations of ribonucleoprotein (RNP fmol) and either single guide (bottom line) or double-stranded guide (top line). DETAILED DESCRIPTION OF THE INVENTION
[0047] Detailed Description Guide polynucleotide In some aspects, the present disclosure describes a guide polynucleotide that includes a targeting region that is complementary to an adapter sequence.
[0048] "Polynucleotide" refers to a polymer of nucleotides (e.g., deoxyribonucleotides, ribonucleotides, and modified nucleotides). The polymer can be in either single-stranded or double-stranded form. Unless specifically limited, the term encompasses synthetic, naturally occurring, and non-naturally occurring polynucleotides containing known nucleotide analogs or modified backbone residues or linkages, which have similar binding properties to the reference nucleic acid, and are metabolized similarly to naturally occurring nucleotides. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, and peptide nucleic acids (PNAs).
[0049] A "guide polynucleotide" refers to a polynucleotide that (1) comprises a targeting region complementary to a target (e.g., an adapter sequence) and (2) can facilitate binding of a ribonucleoprotein complex comprising the guide polynucleotide to the target. In some embodiments, the guide polynucleotide is a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide. In some embodiments, a Cas gRNA polynucleotide refers to a two-polynucleotide system comprising a tracrRNA and a crRNA, e.g., as described in Karvelis T et al., RNA Biol. 2013 May;10(5):841-51. PMID: 23535272. The tracrRNA comprises a sequence encoding a stem-loop structure that associates with a Cas protein (e.g., dCas9). The crRNA comprises a homologous region complementary to the target (e.g., an adapter sequence) and a region complementary to the tracrRNA. The crRNA and tracrRNA can form a complex with a Cas protein, which then binds to a target DNA or RNA polynucleotide (depending on the type of Cas). In some embodiments, a Cas gRNA polynucleotide refers to a single guide RNA (sgRNA) polynucleotide, e.g., as described in Jinek M et al., Science. 2012 Aug 17;337(6096):816-21. doi: 10.1126 / science.1225829. The sgRNA contains a sequence encoding a stem-loop structure that associates with a Cas protein and a homologous region. Many Cas proteins bind to a polynucleotide (e.g., double-stranded DNA) at a position adjacent to a protospacer adjacent motif (PAM) site. In some embodiments, the homologous region is complementary to a portion of the adapter sequence adjacent to the PAM site (e.g., sufficiently adjacent so that Cas binding to the adapter sequence can occur).Methods for making and using gRNAs are known in the art, for example, as described in Mohr SE, et al., FEBS J. 2016 Sep;283(17):3232-8. PMCID: PMC5014588.
[0050] An Argonaute guide polynucleotide refers to a polynucleotide encoding a small interfering RNA (siRNA), a microRNA (miRNA), a P element-induced wimpy testis (PIWI)-interacting RNA (piRNA), and a small interfering DNA (siDNA), for example, as described in Wu J et al., J Adv Res. 2020 Apr 29;24:317-324, PMID: 32455006. In some embodiments, an Argonaute guide polynucleotide comprises a targeting region that is complementary to an adapter sequence described herein. In some embodiments, a guide polynucleotide is single-stranded (e.g., an RNA Cas gRNA polynucleotide). In some embodiments, a guide polynucleotide is double-stranded (e.g., a double-stranded DNA encoding a Cas gRNA polynucleotide or a dsRNA used in an Argonaute guide polynucleotide).
[0051] In some embodiments, the guide polynucleotide is an RNA molecule. In some embodiments, the guide polynucleotide is a DNA molecule. In some embodiments, the guide polynucleotide comprises modified nucleic acids (e.g., 2'F RNA, 2'OMe RNA, and / or phosphorothioate linkages (PS)). Nucleic acids in a guide polynucleotide (e.g., an RNA guide polynucleotide such as a Cas guide RNA) can be modified to increase the stability of the guide polynucleotide (e.g., reduce nuclease digestion). A Cas gRNA (e.g., a Cas9 gRNA) can comprise modified nucleic acids (e.g., 2'OMe) in a region of the gRNA that does not interact with a Cas protein (e.g., Cas9) but still maintains the ability to guide the Cas protein to a target, as described in Yin H et al., Nature Biotechnology 2017 Dec; 35(12): 1179-1187. In some embodiments, the Cas guide RNA polynucleotide comprises a modification of an e-sgRNA as described in Yin et al. 2017. In some embodiments, the Cas9 sgRNA does not comprise a modified nucleic acid (e.g., a 2'OH modification) at one or more of positions 22-27, 43-45, 47, 49, 51, 58, 59, 62, 63-65, 68-69, or 82, counting from the 5' end of the sgRNA (i.e., for an sgRNA of (GGGCGAGGAGCUGUUCACCGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU, SEQ ID NO: 46)).
[0052] In some embodiments, the sgRNA comprises a 2'-O-methyl RNA base. In a 2'-O-methyl RNA base, the 2'-hydroxyl (-OH) group of the ribose sugar in an RNA molecule is replaced with a methyl group (-CH3). This modification affects the ribose sugar; the nitrogenous bases (adenine, guanine, cytosine, uracil) themselves are typically not modified in this context. 2'-O-methyl RNA bases are 2'O-methyl adenosine, 2'O-methyl guanosine, 2'O-methyl cytidine, and 2'O-methyl uridine. In some embodiments, the sgRNA has the sequence: rGrArA rGrArG rCrGrU rCrGrU rGrUrG rUrUrU rUrArG rArGrCrUrArG rArArA rUrArG rCrArA rGrUrU rArArA rArUrA rArGrG rCrUrA rGrUrC rCrGrU rUrArU rCrArA rCrUrU rGrArA rArArA rGrUrG rGrCrA rCrCrG rArGrU rCrGrGrUrGrC mU*mU*mU*rU-3' (SEQ ID NO: 48), wherein mA and mG are 2'-O-methyl adenosine, 2'O-methyl guanosine RNA bases.
[0053] A "targeting region" of a guide polynucleotide refers to a region of the guide polynucleotide that is complementary to a target polynucleotide sequence (e.g., an adapter sequence). In some embodiments, the guide polynucleotide is a Cas gRNA polynucleotide that includes a homology region. A Cas gRNA polynucleotide homology region can include a series of contiguous amino acids that are complementary to a target polynucleotide sequence (e.g., an adapter sequence). In some embodiments, the homology region is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the homology region is 18-26 nucleotides in length. In some embodiments, the homology region is 17-24 nucleotides in length. In some embodiments, the homology region is 18-22 nucleotides in length. In some embodiments, the homology region is 20 nucleotides in length.
[0054] In some embodiments, the guide polynucleotide is an Argonaute guide polynucleotide. The Argonaute guide polynucleotide can be an RNA interference (RNAi) polynucleotide. In some embodiments, the Argonaute guide polynucleotide comprises a targeting region of small interfering RNA (siRNA), microRNA (miRNA), P-element-induced wimpy testis (PIWI)-interacting RNA (piRNA) and small interfering DNA (siDNA), for example, as described in Wu J et al., J Adv Res. 2020 Apr 29;24:317-324, PMID: 32455006.
[0055] In some embodiments, the targeting region is complementary to an adapter sequence described herein. In some embodiments, the targeting region is complementary to a static region of the adapter sequence (e.g., a region that does not generally vary between adapters). In some embodiments, the targeting region is complementary to a sequence that includes a region added to the sequence for the purpose of being complementary to the guide polynucleotide targeting region. For example, an additional sequence can be added to the adapter sequence for the purpose of being complementary to the guide polynucleotide targeting region (e.g., for use in library normalization). Such an additional sequence can be a sequence that is specifically bound by a Cas protein-gRNA complex or an Argonaute guide polynucleotide complex compared to other sequences in a polynucleotide sequencing library.
[0056] In some embodiments, the targeted region is complementary to a next-generation sequencing adapter sequence. In some embodiments, the static region is a static region of a next-generation sequencing adapter sequence. In some embodiments, the targeted region is complementary to a p5 adapter or a p7 adapter. In some embodiments, the p5 adapter and the p7 adapter are ILLUMINA sequencing adapters. In some embodiments, the targeted region is complementary to the i5 polynucleotide sequence of SEQ ID NO:27 or SEQ ID NO:28. In some embodiments, the targeted region is complementary to the i7 adapter sequence of SEQ ID NO:29 or SEQ ID NO:30. In some embodiments, the targeted region is complementary to a NEXTERA Read 1 Adapter (e.g., SEQ ID NO:31 or SEQ ID NO:32) or a NEXTERA Read 2 Adapter (e.g., SEQ ID NO:33 or SEQ ID NO:34). In some embodiments, the targeted region is complementary to an ION TORRENT A Adapter or an ION TORRENT P1 Adapter. In some embodiments, the targeted region is complementary to an ION TORRENT A Adapter of SEQ ID NO:35. In some embodiments, the targeting region is complementary to the ION TORRENT P1 Adapter of SEQ ID NO: 36. In some embodiments, the gRNA polynucleotide comprises a homologous region of any one of SEQ ID NOs: 1-26. In some embodiments, the gRNA polynucleotide comprises a homologous region of SEQ ID NO: 1.
[0057] The term "complementary," as used herein, refers to the expected degree of Watson-Crick base pairing between a first polynucleotide (e.g., a guide polynucleotide targeting region) and a second polynucleotide (e.g., an adapter sequence). Complementary nucleotides are generally A and T (or A and U) and G and C. In some embodiments, complementary refers to 100% complementarity between two polynucleotides (e.g., a guide polynucleotide targeting region and an adapter sequence). In some embodiments, complementary refers to 70%, 80%, 90%, 95%, or 99% complementarity between two polynucleotides. For example, a targeting region can be complementary to an adapter sequence even if one, two, or three nucleic acids in the targeting region are not complementary to the adapter sequence. In some embodiments, complementary refers to a sufficient degree of Watson-Crick base pairing between a guide polynucleotide targeting region (of an RNP (e.g., dCas9 bound to a guide RNA)) and an adapter sequence for the RNP to bind to the adapter sequence. In some embodiments, complementary refers to a sufficient degree of Watson-Crick base pairing between an Argonaute guide polynucleotide (e.g., miRNA, siRNA, pwRNA, shRNA, or siDNA) and a target sequence (e.g., an adapter sequence) to direct binding of an Argonaute protein containing the Argonaute guide polynucleotide to bind to the target.
[0058] "Adapter sequence" refers to a polynucleotide that is added (e.g., ligated) to the end (e.g., 3' and / or 5') of a target polynucleotide for use in sequencing the polynucleotide. In some embodiments, adapter sequence refers to the sequence of a nucleic acid encoding the adapter. In some embodiments, adapter sequences can be used for sequencing using a specific sequencing platform. For example, ILLUMINA i5 / p5 and i7 / p7 adapter sequences can be used for sequencing using an ILLUMINA sequencing platform. In some embodiments, an adapter sequence comprises a static region (e.g., a region that generally does not vary between adapters) and a dynamic region (e.g., a region that can vary between adapters). In some embodiments, an adapter sequence comprises a distal region (generally static), an index region (generally dynamic), and a proximal region (generally static). For example, an ILLUMINA adapter sequence can comprise a p5 region that is static, an index region that is dynamic, and an i5 region that is static. In some embodiments, the guide polynucleotide targeting region is complementary to the static region of the adapter sequence. In some embodiments, the guide polynucleotide targeting region is complementary to a dynamic region of the adapter sequence. In some embodiments, the guide polynucleotide comprises a targeting region that is complementary to a static region of the adapter sequence (e.g., a distal or proximal region of the adapter sequence). In some embodiments, the guide polynucleotide comprises a targeting region that is complementary to a dynamic region of the adapter sequence (e.g., an index). In some embodiments, at least some of the target polynucleotides comprise the same static adapter sequence (e.g., the same distal and / or proximal region). In some embodiments, the majority of the target polynucleotides comprise the same static adapter sequence. In some embodiments, all of the target polynucleotides comprise the same static adapter sequence.
[0059] In some embodiments, the adapter sequence is modified to include a Cas protein protospacer adjacent motif (PAM) or the reverse complement of a PAM. A "protospacer adjacent motif (PAM)" is a nucleotide motif (often three consecutive nucleotides in a polynucleotide) required for Cas protein binding to a target polynucleotide. The canonical Cas9 PAM sequence is 5'-NGG, where N is any nucleotide A, G, C, or T. In some embodiments, the adapter sequence is modified to include a PAM or the reverse complement of a PAM to facilitate binding of the Cas-gRNA complex to the target polynucleotide. For example, a PAM sequence can be added to the 3' or 5' end of the adapter sequence. In some embodiments, a distal sequence can be modified to include a protospacer adjacent motif (PAM) or its reverse complement.
[0060] In some embodiments, the adapter sequence is an adapter sequence used for ILLUMINA, ION TORRENT, PACBIO, ELEMENT, ULTIMA, OMNIOME, SINGULAR, or MGI sequencing. In some embodiments, the adapter sequence is an adapter sequence from any one of the following kits: WATCHMAKER DNA Library Prep Kit with Fragmentation or WATCHMAKER RNA Library Prep Kit with Polaris Depletion ILLUMINA TruSeq PCR-Free Library Preparation Kit, TruSeq Nano DNA Library Prep Kit, NEXTERA DNA Library Prep Kit, NEXTERA DNA XT Library Prep Kit, NEXTERA Rapid Capture Exome Kit, NEXTERA Rapid Capture Expanded Exome Kit, AmpliSeq for ILLUMINA Library Prep Kit, ILLUMINA RNA Prep with Enrichment Kit, ILLUMINA Stranded mRNA Prep Kit, TruSeq RNA Library Prep Kit, TruSeq Stranded Total RNA Kit, TruSeq Stranded mRNA Kit, or TruSeq Small RNA Kit.In some embodiments, the adapter sequence is an adapter sequence used in any one of the following kits: NEBNEXT Fast DNA for ION TORRENT, NEBNEXT Fast DNA Fragmentation & Library Prep Set for ION TORRENT, THERMOPHISHER Precision ID Library Kit, THERMOPHISHER Ion Xpress Plus Fragment Library Kit, THERMOPHISHER Ion Xpress Barcode Adapters 1-96 Kit, and / or THERMOPHISHER Ion AmpliSeq Transcriptome Human Gene Expression Kit. In some embodiments, the adapter sequence is an adapter sequence used in any one of the following kits: PACBIO SMRTbell Template Prep Kit 1.0, SMRTbell Barcoded Adapter Complete Prep Kit-96, or Barcoded Adapter Kit. In some embodiments, the targeted region may be complementary to an adapter sequence used in the QIAGEN QIAseq Stranded RNA Library Kit or QIAseq UPX 3' Transcriptome Kit, the Perkin Elmer NEXTFLEX Rapid Directional RNA-Seq Kit or NEXTFLEX Small RNA-Seq Kit, or the Takara Bio SMART-Seq mRNA Kit or SMART-Seq mRNA LP Kit.
[0061] In some embodiments, the adapter sequence comprises a shared Y adapter sequence. In certain embodiments, the shared Y adapter sequence is a 13 bp sequence. In other embodiments, the shared Y adapter sequence is a 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bp sequence.
[0062] In some embodiments, the adapter sequence may comprise any one of SEQ ID NOs: 27-36.
[0063] kit In some aspects, the present disclosure describes a kit that includes (a) a guide polynucleotide described herein; and (b) a Cas protein or an Argonaute protein, or a polynucleotide sequence encoding a Cas protein or an Argonaute protein (e.g., as described herein).
[0064] In some embodiments, the kit comprises a Cas protein. "Cas protein" refers to a clustered regularly interspaced short palindromic repeats (CRISPR)-associated protein. In some embodiments, the Cas protein has nuclease activity. In some embodiments, the Cas protein does not have nuclease activity (dCas). In some embodiments, the Cas protein has nicking activity (nickase).
[0065] In some embodiments, the Cas protein is a Cas9 protein. "Cas9" refers to a Cas9 protein or fragment thereof present in any bacterial species that encodes a type II CRISPR / Cas9 system. See, for example, Makarova et al., Nature Reviews, Microbiology, 9: 467-477 (2011), including supplementary information. Cas9 homologs are found in a wide variety of eubacteria, including, but not limited to, bacteria from the following taxonomic groups: Actinobacteria, Aquificae, Bacteroidetes-Chlorobi, Chlamydiae-Verrucomicrobia, Chlroflexi, Cyanobacteria, Firmicutes, Proteobacteria, Spirochaetes, and Thermotogae. An exemplary Cas9 protein is the Streptococcus pyogenes Cas9 protein. Additional Cas9 proteins and their homologs are described, for example, in Chylinksi et al., 2013, RNA Biol. 10(5): 726-37; Hou et al., 2013, Proc. Natl. Acad. Sci. USA 110(39): 15644-49; Sampson et al., 2013, 497(7448): 254-57; and Jinek et al., 2012, Science 337(6096): 816-21. Full-length Cas9 is an RNA-mediated endonuclease containing a recognition domain and two nuclease domains (HNH and RuvC, respectively). In terms of amino acid sequence, the HNH is linearly continuous, while the RuvC is divided into three regions: one on the left side of the recognition domain and two on the right side of the recognition domain, adjacent to the HNH domain. Cas9 from Streptococcus pyogenes is targeted to genomic sites within cells by interacting with a guide RNA that hybridizes to a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by Cas9.
[0066] In some embodiments, the Cas9 is a catalytically inactive or nuclease-dead Cas9 (dCas9) modified to inactivate Cas9 nuclease activity. Modifications include, but are not limited to, changing one or more amino acids to inactivate nuclease activity or the nuclease domain. For example, but not limited to, D10A, H840A, and / or R1335K mutations can be made in Cas9 from Streptococcus pyogenes to inactivate Cas9 nuclease activity. In some embodiments, the Cas9 protein contains the D10A and H840A mutations. Other modifications include removing all or a portion of the nuclease domain of Cas9 so that the Cas9 lacks the sequence that exhibits nuclease activity. Thus, a catalytically inactive Cas9 can contain a polypeptide sequence modified to inactivate nuclease activity or can include removal of a polypeptide sequence(s) to inactivate nuclease activity. Catalytically inactive Cas9 retains the ability to bind to DNA even if its nuclease activity has been inactivated. Thus, dCas9 contains the polypeptide sequence(s) necessary for DNA binding, but either contains a modified nuclease sequence or lacks the nuclease sequence responsible for nuclease activity.
[0067] In some embodiments, the catalytically inactive Cas9 protein is a full-length Cas9 sequence from S. pyogenes that lacks the polypeptide sequences of the RuvC and / or HNH nuclease domains and retains DNA-binding function. In other embodiments, the catalytically inactive Cas9 protein sequence has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to a Cas9 polypeptide sequence that lacks the RuvC and / or HNH nuclease domains and retains DNA-binding function. In some embodiments, the Cas protein is Alt-R™ SpdCas9 Protein V3 (IDT).
[0068] In some embodiments, the Cas protein is a Cpf1 (Cas12a), C2c1, C2c3, C2c2, CasX, or CasY protein. In some embodiments, the Cas protein is modified to inactivate Cas9 nuclease activity. Modifications include, but are not limited to, changing one or more amino acids to inactivate the nuclease activity or nuclease domain of the Cas protein. For example, D908, D832, E993, R1226, and / or D1235 mutations can be made in Cpf1 from Acidaminococcus sp. BV3L6, Lachnospiraceae ND2006, or Francisella tularensis subsp. novicida U112 to inactivate Cpf1 nuclease activity. Other modifications include removing all or part of the nuclease domain so that no sequence exhibiting nuclease activity is present. Thus, catalytically inactive Cas proteins may contain polypeptide sequences modified to inactivate nuclease activity or may contain removal of polypeptide sequence(s) to inactivate nuclease activity. Catalytically inactive Cas proteins retain the ability to bind to DNA even if their nuclease activity is inactivated. Thus, catalytically inactive Cas proteins contain the polypeptide sequence(s) necessary for DNA binding but contain modified nuclease sequences or lack the nuclease sequences responsible for nuclease activity. In some embodiments, catalytically inactive Cas proteins are full-length Cas sequences lacking the polypeptide sequence of the nuclease domain and retaining DNA-binding function. In other embodiments, catalytically inactive Cas protein sequences have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to a Cas polypeptide sequence lacking the nuclease domain and retaining DNA-binding function.
[0069] In some embodiments, the Cas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 37 or SEQ ID NO: 38. In some embodiments, the Cas protein comprises the amino acid sequence of SEQ ID NO: 37 or SEQ ID NO: 38.
[0070] In some embodiments, the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 37. In some embodiments, the dCas protein comprises the amino acid sequence of SEQ ID NO: 37.
[0071] In some embodiments, the Cas protein comprises a nuclease localization sequence.
[0072] In some embodiments, a kit includes an Argonaute protein. "Argonaute protein" refers to a protein that binds to small non-coding nucleic acids (e.g., Argonaute guide polynucleotides) and utilizes them for guided cleavage of complementary nucleic acid targets or for indirect gene silencing by recruiting additional factors. Catalytically active Argonaute proteins are capable of guided binding of nucleic acids to complementary nucleic acid targets (e.g., DNA) and cleavage of the nucleic acid targets. See, for example, Kaya et al., 2016, PNAS 113(5):4057-62. The prokaryotic Argonaute (AGO) gene family encodes several domains: an N-terminal (N), PAZ, MID, and a C-terminal PIWI domain. The MID and PAZ domains are involved in binding the 5' and 3' ends of the guide polynucleotide, respectively. In contrast to eukaryotic Agos, which exclusively use small RNA guides, the majority of characterized prokaryotic Agos bind to single-stranded DNA (ssDNA) guides. The Ago PIWI domain contains nuclease activity. In some embodiments, the Argonaute protein is catalytically inactive. In some embodiments, the Argonaute protein is a catalytically inactive CbAgo protein. In some embodiments, the Argonaute protein is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo protein. In certain embodiments, the Argonaute protein has been modified to inactivate its nuclease activity or has essentially reduced nuclease activity (see, for example, Kaya et al., 2016, PNAS 113(5):4057-62). Modifications may include, but are not limited to, changing one or more amino acids to inactivate the nuclease activity or nuclease domain of the Argonaute protein. For example, in some embodiments, the catalytically inactive Argonaute protein is a CbAgo protein that includes a D541A mutation and a D611A mutation.See, for example, Hegge et al., 2019, Nuc. Acids Res. 47(11):5809-21. Alternatively, all or a portion of the nuclease domain can be removed so that the sequence exhibiting nuclease activity is absent. Thus, a catalytically inactive Argonaute protein can contain a polypeptide sequence modified to inactivate nuclease activity, or can contain the removal of a polypeptide sequence(s) to inactivate nuclease activity. A catalytically inactive Argonaute protein retains the ability to bind to DNA, even if its nuclease activity is inactivated. Thus, a catalytically inactive Argonaute protein contains the polypeptide sequence(s) necessary for DNA binding, but contains a modified nuclease sequence or lacks the nuclease sequence responsible for nuclease activity. In some embodiments, a catalytically inactive Argonaute protein is a full-length Argonaute sequence that lacks the polypeptide sequence of the nuclease domain and retains DNA binding function. In other embodiments, the catalytically inactive Argonaute protein sequence has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to an Argonaute polypeptide sequence that lacks the nuclease domain and retains DNA binding function.
[0073] In some embodiments, the d-Argonaute protein comprises an amino acid sequence having at least 95% identity to any one of SEQ ID NOs: 40-45. In some embodiments, the d-Argonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 40-45.
[0074] In some embodiments, the kits include a catalytically inactive Cas protein (dCas protein) or a catalytically inactive Argonaute protein (e.g., dArgonaute) described herein. In some embodiments, the kits include a catalytically inactive Cas protein (dCas protein) and a catalytically active Cas protein described herein. In some embodiments, the kits include a catalytically inactive Argonaute protein (e.g., dArgonaute) or a catalytically active Argonaute protein described herein. In some embodiments, the kits include dCas and a cognate dCas guide RNA polynucleotide. In some embodiments, the kits include dCas9 and a cognate dCas9 guide RNA. In some embodiments, the kits include a dArgonaute and a cognate dArgonaute guide polynucleotide.
[0075] In some embodiments, the kit comprises a Cas protein comprising an affinity tag or an Argonaute protein comprising an affinity tag. In some embodiments, the affinity tag is a protein comprising an affinity tag, as described in Kimple ME et al., Curr Protoc Protein Sci. 2013 Sep 24;73:9.9.1-9.9.23. doi: Albumin-binding protein (ABP), alkaline phosphatase (AP), AU1 epitope, AU5 epitope, bacteriophage T7 epitope (T7-tag), bacteriophage V5 epitope (V5-tag), biotin-carboxy carrier protein (BCCP), bluetongue virus tag (B-tag), calmodulin-binding peptide (CBP), chloramphenicol acetyltransferase (CAT), cellulose-binding domain (CBP), chitin-binding domain (CBD), choline-binding domain (CBD), dihydrofolate reductase (DHFR), E2 epitope, FLAG epitope, as described in 10.1002 / 0471140864.ps0909s73. , galactose-binding protein (GBP), green fluorescent protein (GFP), Glu-Glu (EE-tag), Glu-Glu (EE-tag), human influenza virus hemagglutinin (HA), HaloTag®, histidine affinity tag (HAT), horseradish peroxidase (HRP), HSV epitope, ketosteroid isomerase (KSI), KT3 epitope, LacZ, luciferase, maltose-binding protein (MBP), Myc epitope, Nus, PDZ domain, PDZ ligand, polyarginine (Arg-tag), polyaspartate (Asp-tag), polyhistidine (His-tag), polyphenylalanine (Phe-tag), Profinity eXact, Protein C, S1-tag, S-tag, Streptavidin-binding peptide (SBP), Staphylococcal Protein A (Protein A), Staphylococcal Protein G (Protein G), Strep-tag, Streptavidin, Small Ubiquitin-like Modifier (SUMO), Tandem Affinity Purification (TAP), T7 epitope, Thioredoxin (Trx), TrpE, Ubiquitin, Universal, and VSV-G.
[0076] In some embodiments, the affinity tag and corresponding affinity tag binding molecule in the kit or for use in the method are selected from Table 1. Table 1: Affinity tags and affinity tag-binding molecules. [Table 1-1] [Table 1-2]
[0077] In some embodiments, the kit comprises a guide polynucleotide that can bind to a Cas protein (e.g., the guide polynucleotide is cognate to a Cas protein) or Argonaute protein of the kit to form a ribonucleoprotein complex, and the ribonucleoprotein complex can bind to an adapter sequence.
[0078] The term "cognate," when used in the context of a guide polynucleotide and a Cas protein or Argonaute protein, refers to a guide polynucleotide that is compatible with the Cas protein or Argonaute protein (e.g., the guide polynucleotide is capable of directing the Cas protein or Argonaute protein to a target polynucleotide). For example, a Cas9 sgRNA is cognate to Cas9 and dCas9.
[0079] The terms "bind" or "specifically bind" and similar terms refer to a molecule (e.g., a Cas-gRNA complex or an Argonaute-guide polynucleotide complex) that binds to a target nucleic acid with at least two-fold greater affinity than a non-target nucleic acid, e.g., with at least one of 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 25-fold, 50-fold, 100-fold, 1,000-fold, 10,000-fold, or greater affinity for the target nucleic acid compared to an unrelated nucleic acid, when assayed under the same binding affinity assay conditions. In some embodiments, the terms "bind" or "specifically bind" as used herein refer to, for example, a molecule that binds to a target nucleic acid with at least two-fold greater affinity than a non-target nucleic acid, e.g., with at least one of 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 25-fold, 50-fold, 100-fold, 1,000-fold, 10,000-fold, or greater affinity for the target nucleic acid compared to an unrelated nucleic acid. -2 M or smaller, e.g., 10 -3 M, 10 -4 M, 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M, 10 -9 M, 10 -10 M, 10 -11 M, or 10 -12 This can be exhibited by a molecule (e.g., a Cas-gRNA complex or an Argonaute-guide polynucleotide complex) having an equilibrium dissociation constant, KD, of M. In some embodiments, the antibody has a KD of less than 10 nM or less than 100 nM.
[0080] In some embodiments, the kit further includes a primer complementary to a portion of the adapter sequence (e.g., as described herein). In some embodiments, the primer is complementary to a proximal portion of the adapter sequence. This primer can be used to generate a double-stranded target polynucleotide containing a double-stranded protospacer adjacent motif (PAM) site that can be used for binding by a dCas protein (e.g., as described herein). In some embodiments, the primer binds to the proximal portion of the adapter sequence, such that the new polynucleotide (complementary to the target polynucleotide) does not contain the distal portion of the adapter sequence, and therefore, the polynucleotide cannot be sequenced using next-generation sequencing (e.g., ILLUMINA sequencing). Without being bound by theory, this method may be advantageous compared to PCR-based methods because errors (e.g., mutations) introduced during primer extension are not sequenced, but a double-stranded target polynucleotide can still be generated and bound to the dCas protein.
[0081] In some embodiments, the primer is a DNA primer. In some embodiments, the DNA primer comprises a nucleic acid modification (e.g., as described herein). In some embodiments, a portion of the adapter sequence is located in the proximal portion of the adapter sequence. In some embodiments, the proximal portion of the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34. In some embodiments, the primer is complementary to the 3' adapter sequence. In some embodiments, the primer is complementary to the 5' adapter sequence. In some embodiments, the primer is not complementary to the adapter sequence that is complementary to the guide polynucleotide targeting region. In some embodiments, the primer is bound to the adapter sequence and, when extended, does not extend the entire adapter sequence.
[0082] Reaction mixture In some aspects, the present disclosure describes a reaction mixture comprising: (i) a plurality of target polynucleotides, where the target polynucleotides comprise adapter sequences; (ii) a predetermined concentration of a guide polynucleotide comprising a targeting region complementary to the adapter sequence; and (iii) a predetermined concentration of a dCas protein or a dArgonaute protein. In some embodiments, the reaction mixture comprises an aqueous solution (e.g., a buffer suitable for carrying out the methods described herein). In some embodiments, the reaction mixture comprises a predetermined concentration of a dCas9 protein. In some embodiments, the reaction mixture comprises a predetermined concentration of a dArgonaute protein. In some embodiments, the reaction mixture comprises a target polynucleotide comprising an adapter sequence of any one of SEQ ID NOs: 27-36. In some embodiments, the reaction mixture comprises a predetermined concentration of a guide polynucleotide of any one of SEQ ID NOs: 27-36.
[0083] In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 50 femtomol (fmol) to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 200 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 300 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 500 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 750 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 2,000 fmol to 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 50 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 3,400 fmol. In some embodiments, the reaction mixture comprises a plurality of target polynucleotides comprising adapter sequences at a concentration of 500 fmol to 3,400 fmol.
[0084] In some embodiments, the reaction mixture comprises a predetermined concentration of dCas9 of SEQ ID NO: 38, a target polynucleotide comprising an adapter sequence of SEQ ID NO: 28, and a predetermined concentration of a guide polynucleotide (e.g., an RNA guide polynucleotide) comprising a homologous region that is complementary to the adapter sequence (e.g., the homologous region comprises SEQ ID NO: 37).
[0085] In some embodiments, the reaction mixture comprises a predetermined concentration of dArgonaute of any one of SEQ ID NOs: 40-45 (e.g., a catalytically inactive variant of any one of SEQ ID NOs: 40-45), a target polynucleotide comprising an adapter sequence of SEQ ID NO: 28, and a predetermined concentration of dArgonaute guide polynucleotide comprising a targeting region that is complementary to the adapter sequence.
[0086] In some embodiments, the reaction mixture further comprises a catalytically active nuclease as described herein.
[0087] Ribonucleoprotein (RNP) complexes In some aspects, the present disclosure provides a ribonucleoprotein (RNP) complex comprising: (i) a guide polynucleotide; (ii) a Cas protein (e.g., dCas) or an Argonaute protein (e.g., dArgonaute); and (iii) an adapter sequence, wherein a targeted region of the guide polynucleotide is complementary to the adapter sequence.
[0088] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide, (ii) a Cas protein described herein, and (iii) an adapter sequence, wherein the homologous region of the RNA Cas gRNA polynucleotide is complementary to the adapter sequence.
[0089] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide comprising a homologous region of any one of SEQ ID NOs: 1-26, (ii) a dCas protein of SEQ ID NO: 37, and (iii) an adapter sequence of any one of SEQ ID NOs: 27-36, wherein the homologous region of the RNA Cas gRNA polynucleotide is complementary to the adapter sequence.
[0090] In some embodiments, the RNP complex comprises (i) an RNA Cas gRNA polynucleotide comprising a homologous region of any one of SEQ ID NO: 1, (ii) a dCas protein of SEQ ID NO: 37, and (iii) an adapter sequence of any one of SEQ ID NO: 28, wherein the homologous region of the RNA Cas gRNA polynucleotide is complementary to the adapter sequence.
[0091] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide, (ii) an Argonaute protein, and (iii) an adapter sequence, wherein the targeted region of the Argonaute guide polynucleotide is complementary to the adapter sequence.
[0092] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide (e.g., siDNA), (ii) an Argonaute protein of any one of SEQ ID NOs: 39-45, and (iii) an adapter sequence of SEQ ID NO: 28, wherein the targeted region of the Argonaute guide polynucleotide is complementary to the adapter sequence.
[0093] In some embodiments, the RNP complex comprises (i) an Argonaute guide polynucleotide (e.g., siDNA), (ii) an Argonaute protein of any one of SEQ ID NOs: 40-45, and (iii) an adapter sequence of SEQ ID NO: 28, wherein the targeted region of the Argonaute guide polynucleotide is complementary to the adapter sequence.
[0094] In some embodiments of the RNP complexes provided herein, the adapter sequence is unmodified. In some embodiments, the adapter sequence is unmodified at the 5' end.
[0095] In some embodiments, the RNP complex is present at a concentration of 250 femtomolar (fmol) to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of 300 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of 2,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of 6,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of 10,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 250 fmol. In some embodiments, the RNP complex is present at a concentration of at least 300 fmol. In some embodiments, the RNP complex is present at a concentration of at least 2,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 6,000 fmol. In some embodiments, the RNP complex is present at a concentration of at least 10,000 fmol. In some embodiments, the RNP complex is present at a concentration of 250 fmol. In some embodiments, the RNP complex is present at a concentration of 300 fmol. In some embodiments, the RNP complex is present at a concentration of 2,000 fmol. In some embodiments, the RNP complex is present at a concentration of 6,000 fmol. In some embodiments, the RNP complex is present at a concentration of 10,000 fmol. In some embodiments, the RNP complex is present at a concentration of 14,000 fmol. array Table 2: Sequences [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8]
[0096] Methods for Polynucleotide Library Normalization In some aspects, the present disclosure provides methods for normalizing the concentration of a target polynucleotide between two or more samples (eg, normalizing between two or more polynucleotide libraries).
[0097] As used herein, "normalizing," "normalization," and similar terms refer to the process of generating subsequent samples having a desired concentration of target polynucleotides from an initial sample having a different starting concentration of target polynucleotides than the subsequent samples. In some embodiments, "normalization" of two or more samples results in two or more subsequent samples each having a target polynucleotide concentration that is more similar to each other after the normalization method than before the normalization method. For example, a first sample (e.g., a first polynucleotide library) may contain five times more target polynucleotides than a second sample (e.g., a second polynucleotide library). In this example, after normalization of the first and second samples, the difference in concentration between the first and second samples will be less than five-fold (e.g., less than four-fold, less than three-fold, or less than two-fold). In some embodiments, normalization of two or more samples results in two or more subsequent samples each having equimolar (i.e., 1:1) concentrations. In some embodiments, the normalization method results in a difference in target polynucleotide concentration between two or more samples that is less than 50% (e.g., less than 40%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, or less than 2.5%). In some embodiments, the normalization method results in a difference in target polynucleotide concentration between two or more samples that is 5% to 40%, 5% to 30%, 5% to 20%, 5% to 10%, 10% to 40%, 10% to 30%, or 10% to 20% in concentration. In some embodiments, the difference in starting target polynucleotide concentration between two or more samples may be relatively small (e.g., less than 5-10%) before performing the normalization method described herein. In such embodiments, the normalization method may be performed, but the two or more samples may not have detectably more similar concentrations after normalization than they were before normalization.
[0098] A "target polynucleotide" or "target polynucleotides" refers to a polynucleotide comprising a nucleic acid encoding an adapter sequence or a static region of an adapter sequence. In some embodiments, a target polynucleotide comprises a polynucleotide comprising: (1) a nucleic acid encoding an adapter sequence or a sequence (e.g., a zinc finger binding domain or a Talen binding domain) added to a polynucleotide for the purpose of normalization using the methods described herein; and (2) a nucleic acid encoding a sequence of interest. In some embodiments, a target polynucleotide comprises a nucleic acid encoding an adapter sequence described herein and a nucleic acid encoding a sequence of interest. The sequence of interest can be any sequence for which normalization is desired. For example, the sequence of interest can include, but is not limited to, DNA, genomic DNA, circulating tumor DNA, RNA, rRNA, mRNA, miRNA, or cDNA. The adapter sequence can be ligated to the 5' and / or 3' end of the sequence of interest. In some embodiments, the target polynucleotide comprises a first adapter sequence (e.g., a p5 adapter) at the 5' end of the sequence of interest and a second adapter sequence (e.g., a p7 adapter) at the 3' end of the target sequence.
[0099] A "sample" refers to a composition or solution containing one or more target polynucleotides. In some embodiments, a sample contains a plurality of target polynucleotides. "Plurality" refers to two or more. For example, a plurality of target polynucleotides includes two or more target polynucleotides. In some embodiments, a plurality of target polynucleotides includes multiple copies of a single target polynucleotide sequence (e.g., at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, or at least 1,000,000, or at least 10,000,000, or at least 100,000,000 target polynucleotides). In some embodiments, the plurality of target polynucleotides comprises different target polynucleotide sequences (e.g., at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, or at least 1,000,000, or at least 10,000,000, or at least 100,000,000 different target polynucleotides), but at least some of the target polynucleotides (e.g., all of the target polynucleotides) comprise the same adapter sequence. In some embodiments, the sample comprises target polynucleotides and other molecules (e.g., polynucleotides that are not target polynucleotides and do not include an adapter sequence or that include an adapter sequence different from that complementary to the guide polynucleotide). In some embodiments, the sample comprises target polynucleotides comprising nucleic acids encoding genomic DNA. In some embodiments, the sample comprises target polynucleotides comprising nucleic acids encoding RNA. In some embodiments, the sample comprises target polynucleotides comprising nucleic acids encoding cDNA (e.g., of mRNA or lncRNA). In some embodiments, the sample comprises target polynucleotides comprising nucleic acids encoding exons. In some embodiments, the sample comprises a target polynucleotide that comprises a nucleic acid encoding an intron.
[0100] In some embodiments, the plurality of target polynucleotides in the sample are present at a concentration of 50 femtomol (fmol) to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 200 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 300 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 500 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 750 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 1,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides are present at a concentration of 2,000 fmol to 3,400 fmol. In some embodiments, the plurality of target polynucleotides is present at a concentration of 50 fmol. In some embodiments, the plurality of target polynucleotides is present at a concentration of 3,400 fmol. In some embodiments, the plurality of target polynucleotides is present at a concentration of 500 fmol to 3,400 fmol.
[0101] In some embodiments, the normalization method is performed on at least two samples (e.g., at least three samples, at least four samples, at least five samples, at least six samples, at least seven samples, at least eight samples, at least nine samples, at least 10 samples, at least 25 samples, at least 50 samples, at least 75 samples, at least 100 samples, at least 150 samples, or at least 200 samples, at least 300 samples, at least 384 samples, at least 400 samples, at least 500 samples, at least 600 samples, at least 700 samples, at least 800 samples, at least 900 samples, at least 1000 samples, at least 1500 samples, at least 2000 samples, at least 3000 samples, or at least 5000 samples or more). In some embodiments, the normalization method is performed on 2 to 384 samples. In some embodiments, the normalization method is performed on 24 to 384 samples. In some embodiments, the normalization method is performed on between 24 and 500 samples. In some embodiments, the normalization method is performed on between 24 and 750 samples. In some embodiments, the normalization method is performed on between 24 and 1000 samples. In some embodiments, the normalization method is performed on between 24 and 2000 samples. In some embodiments, the normalization method is performed on between 24 and 3000 samples.
[0102] In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 70-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 60-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 50-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 40-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 30-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 20-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 10-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 5-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 2-fold of each other. In some embodiments, the sample comprises target polynucleotides at pre-normalized concentrations within 1.5-fold of each other.
[0103] In some embodiments, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples (e.g., two polynucleotide libraries), each comprising a target polynucleotide, the method comprising, for each of the at least two samples, (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution comprising combining (a) the sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to a static region of the adapter sequence of the target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute; (iii) contacting the solution with a solid phase comprising a dCas or dArgonaute binding molecule; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0104] In some embodiments, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples (e.g., two polynucleotide libraries), each comprising a target polynucleotide, the method comprising, for each of the at least two samples, (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution comprising combining (a) the sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence of the target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute; (iii) contacting the solution with a solid phase comprising a dCas or dArgonaute binding molecule; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between the two or more samples.
[0105] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes obtaining a sample containing the target polynucleotide. The sample can be obtained from any suitable source, including, but not limited to, cells (e.g., eukaryotic or prokaryotic cells), tissues (e.g., brain, heart, lung, liver, kidney, fat, skin, gallbladder, or breast), diseased tissues (e.g., tumors), bodily fluids (e.g., saliva, sweat, blood, urine, mucus, or cerebrospinal fluid), or from in vitro generation. In some embodiments, obtaining the sample includes obtaining a polynucleotide of interest (e.g., extracting the polynucleotide of interest from a biological sample) and modifying the polynucleotide of interest to include an adapter sequence (e.g., modifying the polynucleotide of interest to become a target polynucleotide). Modifying the polynucleotide of interest to include an adapter sequence can be performed using any suitable method, including, but not limited to, using a kit for attaching an adapter sequence to a polynucleotide described herein. In some embodiments, the sample contains 1 to 100 nM of the target polynucleotide. In some embodiments, the sample comprises between 1 and 1000 nM of the target polynucleotide.
[0106] In some embodiments, generating a solution, in the context of a normalization method, refers to combining a sample with a predetermined concentration of guide polynucleotide and a predetermined concentration of dCas or dArgonaute in an aqueous solution (e.g., a buffer suitable for the method). In some embodiments, generating a solution comprises combining a predetermined amount of dCas protein or dArgonaute protein with a predetermined amount of guide polynucleotide before combining with the sample. In some embodiments, generating a solution comprises combining a predetermined amount of guide RNA polynucleotide, which is RNA, with a predetermined amount of dCas9 protein or dArgonaute protein. In some embodiments, generating a solution comprises combining a predetermined amount of guide polynucleotide, which is DNA, with a predetermined amount of dArgonaute protein. In some embodiments, generating a solution comprises combining a predetermined amount of guide RNA comprising a homologous region of any one of SEQ ID NOS: 1-26 with a predetermined amount of dCas9 of SEQ ID NOS: 37.
[0107] A "predetermined concentration" refers to a concentration (e.g., nanomolar concentration) or amount (e.g., nanomolar) of a reagent (e.g., guide polynucleotide, dCas protein, or dArgonaute protein) selected to extract a specific or consistent amount of target polynucleotide from a sample. In some embodiments, the same or similar amount of dCas protein or dArgonaute protein at a predetermined concentration is added to each sample to be normalized. In some embodiments, the same or similar amount of guide RNA polynucleotide at a predetermined concentration is added to each sample to be normalized. In some embodiments, a predetermined concentration of dCas protein or dArgonaute protein is added to a sample, and the predetermined amount of guide RNA polynucleotide added to the sample exceeds the predetermined amount of dCas protein or dArgonaute protein added to the sample. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is between 1 and 2000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is between 60 and 2000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is between 60 and 1000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is between 125 and 1000 femtomoles. In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is at least 60 femtomoles (e.g., at least 60 femtomoles, at least 125 femtomoles, at least 250 femtomoles, at least 500 femtomoles, at least 750 femtomoles, or at least 1000 femtomoles). In some embodiments, the predetermined amount of dCas protein or dArgonaute protein is 50, 60, 125, 250, 750, 1000, 1500, or 2000 femtomoles. In some embodiments, the predetermined concentration of guide RNA polynucleotide is greater than the predetermined concentration of dCas9 protein or dArgonaute protein.
[0108] In some embodiments, the ribonucleoprotein (RNP) complex (e.g., guide RNA, dCas protein or dArgonaute protein, and adapter sequence) is present in the sample at a concentration of 250 femtomol (fmol) to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 300 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 2,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 6,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 10,000 fmol to 14,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 250 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 300 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 2,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 6,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of at least 10,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 250 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 300 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 2,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 6,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 10,000 fmol. In some embodiments, the RNP complex is present in the sample at a concentration of 14,000 fmol.
[0109] In some embodiments, the method comprises generating a solution comprising a predetermined concentration of a dCas protein comprising an affinity tag or a dArgonaute protein comprising an affinity tag as described herein.
[0110] In some embodiments, the affinity tag binding molecule comprises a metal ion (e.g., Ni2+) and the affinity tag comprises a His tag. In some embodiments, the affinity tag binding molecule comprises biotin and the affinity tag comprises avidin. In some embodiments, the affinity tag binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag. In some embodiments, the affinity tag binding molecule and the corresponding affinity tag are selected from Table 1.
[0111] In some embodiments, the dArgonaute or dCas of the method does not comprise an affinity tag. In some embodiments, the method includes contacting the solution with a solid phase comprising a molecule that binds to dCas or dArgonaute (e.g., an antibody that binds to dCas or dArgonaute).
[0112] In some aspects, the method includes contacting a solution with a solid phase. In some embodiments, the solid phase includes a molecule (e.g., an antibody or affinity tag-binding molecule) that binds to dArgonaute or dCas. In some embodiments, the solid phase includes beads (e.g., microparticles or nanoparticles). In some embodiments, the beads are metal beads, polymer beads, protein beads, or lipid beads. In some embodiments, the beads (e.g., metal beads) include metal ions (e.g., Ni2+, Co2+, Cu2+, and / or Zn2+ ions) on the surface of the beads. In some embodiments, the metal beads are magnetic (e.g., paramagnetic, diamagnetic, or ferromagnetic).
[0113] In some embodiments, the method includes incubating the solid phase and the solution. In some embodiments, the incubation is at room temperature. In some embodiments, the incubation is at 30-40°C. In some embodiments, the incubation is at 35-38°C. In some embodiments, the incubation is at or about 37°C. In some embodiments, the incubation is for at least 15 minutes (e.g., at least 30 minutes, at least 45 minutes, at least 60 minutes, or at least 2 hours). In some embodiments, the incubation is for at least 60 minutes. In some embodiments, the incubation is for 15 minutes to 2 hours.
[0114] In some embodiments, the method further includes a step of separating the solution from the solid phase. "Separating," as used in the context of this method, refers to removing the solution from the solid phase or removing the solid phase from the solution. In some embodiments, the separating is not a complete separation; for example, a small amount of solution (e.g., less than about 1 microliter) and the solid phase may still be in contact with each other after separation. In some embodiments, the purpose of the separating step is to separate unbound target polynucleotides in the solution from target polynucleotides bound to the solid phase, which is part of concentration normalization. This separation can be achieved by (1) capturing magnetic beads bound to the target polynucleotides (e.g., using a magnet), (2) removing the solution from the solid phase (e.g., using a pipette), and (3) repeatedly rinsing the solid phase with a buffer (e.g., a buffer that does not denature the Cas protein or Argonaute protein). In some embodiments, the method includes a step of washing the solid phase at least once (e.g., at least two, at least three, at least four, or at least five times, or more).
[0115] In some embodiments, the method further comprises extracting the target polynucleotide from the solid phase. In some embodiments, the extracting step comprises contacting the solid phase with a protease that proteolyzes the Cas protein or Argonaute protein, thereby liberating the target polynucleotide. The protease may be any suitable protease, including, but not limited to, trypsin, chymotrypsin, endoproteinase Lys-C, endoproteinase AspN, endoproteinase GluC, elastase, proteinase K, or papain. Corresponding protease reaction conditions and incubation times are known in the art. In some embodiments, the extracting step comprises contacting the solid phase with a protein denaturant (e.g., a surfactant, an organic solvent, or a chaotropic agent). In some embodiments, the extracting step comprises increasing the temperature of the solid phase (e.g., by warming the liquid containing the solid phase) to denature the Cas protein or Argonaute protein. In some embodiments, the extracting step comprises increasing or decreasing the pH of the liquid containing the solid phase.
[0116] In some embodiments, the normalization method includes, in sequential order: (i) obtaining a sample, wherein a target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution comprising combining (a) the sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence of the target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag, wherein the predetermined concentration of dCas or dArgonaute corresponds to the guide polynucleotide; (iii) contacting the solution with a solid phase comprising an affinity tag binding molecule capable of binding to the affinity tag; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentrations of the target polynucleotides between two or more samples.
[0117] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes, for each of the at least two samples, (i) obtaining a sample, wherein the target polynucleotide of the sample comprises an adapter sequence of any one of SEQ ID NOs: 27-36; (ii) generating a solution, the solution comprising combining (a) the sample, (b) a predetermined concentration of a Cas gRNA polynucleotide comprising a homologous region complementary to the adapter sequence, and (c) a predetermined concentration of dCas9 comprising a His tag of SEQ ID NO: 47 (e.g., a His tag at the N-terminus of dCas9); (iii) contacting the solution with a magnetic solid phase comprising Ni2+ ions; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between two or more samples.
[0118] In some embodiments, a method for normalizing the concentration of a target polynucleotide between at least two samples includes, for each of the at least two samples, (i) obtaining a sample, wherein the target polynucleotide of the sample comprises an adapter sequence of any one of SEQ ID NOs: 27-36; (ii) generating a solution comprising combining (a) the sample, (b) a predetermined concentration of an Argonaute guide polynucleotide (e.g., siDNA) comprising a targeting region complementary to the adapter sequence, and (c) a predetermined concentration of an Argonaute protein comprising a His-tag of SEQ ID NO: 47; (iii) contacting the solution with a magnetic solid phase comprising Ni2+ ions; (iv) separating the solution from the solid phase; and (v) extracting the target polynucleotide from the solid phase to normalize the concentration of the target polynucleotide between two or more samples.
[0119] In some aspects, the present disclosure provides a method for normalizing the concentration of a target polynucleotide between at least two samples, each containing a target polynucleotide, wherein uncaptured target polynucleotides are digested during normalization. In some embodiments, the method includes, for each of the at least two samples, (i) obtaining a sample (e.g., as described herein), wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution, the solution comprising combining (a) the sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region complementary to the adapter sequence of the target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag; and (iii) contacting the solution with a nuclease (e.g., catalytically active Cas or Argonaute protein 9) to normalize the concentration of the target polynucleotide between the two or more samples (e.g., digest target polynucleotides not bound to a solid phase).
[0120] In some embodiments, the method includes contacting the solution with a nuclease. In some embodiments, the nuclease is an exonuclease. In certain embodiments, the exonuclease is exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VI, exonuclease VII, or exonuclease VIII. In some embodiments, the nuclease is a catalytically active Cas RNP or Argonaute RNP. In some embodiments, the catalytically active Cas RNP or Argonaute RNP comprises a guide polynucleotide that is complementary to an adapter sequence. In some embodiments, the Cas endonuclease is a Cas9 endonuclease. In some embodiments, the Cas endonuclease is a Cpf1, C2c1, C2c3, C2c2, CasX, or CasY endonuclease. In some embodiments, the Argonaute protein (e.g., an Argonaute endonuclease) is catalytically active. In some embodiments, the Argonaute protein is a CbAgo endonuclease. In certain embodiments, the Argonaute protein is an LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo endonuclease.
[0121] In some aspects, the two other samples to be normalized (e.g., using the methods described herein) contain a target polynucleotide comprising a first adapter sequence and a second adapter sequence. In some embodiments, the guide polynucleotide targeting region is complementary to the first adapter sequence. In some embodiments, the method further includes contacting the sample with a primer encoding a nucleic acid sequence complementary to a portion of the second adapter sequence and a polymerase (e.g., a DNA polymerase) under conditions sufficient for primer extension. As described in the "Kits" section, this primer can be used to extend the target polynucleotide so that the first adapter sequence contains a double-stranded PAM for Cas (e.g., dCas) binding. In some embodiments, the primer is complementary to a proximal portion of the adapter sequence (e.g., the second adapter sequence). In such embodiments, extending the primer to generate a double-stranded polynucleotide cannot generate an extended strand that can be sequenced by next-generation sequencing because the extended strand does not contain the distal portion of the second adapter sequence. This can be advantageous. This is because extension (e.g., DNA amplification) can introduce artifacts (e.g., mutations) into the polynucleotide that can bias or distort the sequencing results.
[0122] In some embodiments, the primer comprises a nucleic acid sequence complementary to any one of SEQ ID NOs: 28, 30, 32, and 34. In some embodiments, the second adapter sequence is a 3' adapter sequence. In some embodiments, the second adapter sequence is a 5' adapter sequence.
[0123] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an antibody" includes combinations of two or more such molecules, as appropriate.
[0124] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illustrate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed.
[0125] The terms "may," "may be," "can," and "can be," and related terms, unless the context clearly dictates otherwise, are intended to convey that the accompanying subject matter is necessary (i.e., the subject matter is present in some instances and absent in other instances) and are not a reference to the subject matter's ability or probability.
[0126] The terms "optionally" and "as needed" mean that the subsequently described event, circumstance, or material may or may not occur or be present, and that the description includes instances in which the event, circumstance, or material occurs or is present, as well as instances in which it does not occur or be present.
[0127] Use herein of the terms "including," "comprising," or "having," and variations thereof, is intended to encompass the subsequently listed elements and equivalents thereof, as well as additional elements. Embodiments described as "including," "comprising," or "having" certain elements are also considered to "consist essentially of" and "consist of" those certain elements. As used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or"). [Example]
[0128] Example 1 Library normalization by Cas9-guided mediated capture of library molecules This example demonstrates that a predetermined amount of dCas protein and a gRNA that binds to the adapter sequence of a target polynucleotide can be used to normalize target polynucleotide concentrations between samples.
[0129] Polynucleotide library (also known as NGS library) normalization prior to equimolar pooling of many different polynucleotide libraries is a challenging task in next-generation sequencing workflows. While various strategies for normalization exist, these strategies often require additional steps, such as qPCR, or special primers or oligonucleotides, which often result in the destruction of unused material.
[0130] The normalization method described herein involves samples containing a target polynucleotide library with variable and unknown concentrations, resulting in the reduction of the library concentration to a predetermined value. This allows the target polynucleotide concentrations of several unknown target polynucleotide libraries to be normalized to a more similar molar concentration and then pooled by volume before loading onto the sequencer. A common method relies on capturing a known number of library molecules using an adjustable amount of a non-cleaving nucleic acid-guided nuclease (e.g., dCas9 or dArgonaute) bound to magnetic beads (Figure 1). Excess library molecules are removed by pipetting unbound material after magnetic bead separation from the initial binding incubation and subsequent buffer washes. After washing, the magnetic beads from each sample are resuspended in an equal volume of buffer. The non-cleaving nucleic acid-guided nuclease can then be denatured to obtain the captured DNA. The resulting solution is then pooled in equal volume with other similarly prepared libraries prior to sequencing.
[0131] In this example, a magnetic bead-compatible form of a nucleic acid-guided nuclease (NAGN) capable of specifically binding to a target determined by a guide nucleic acid (e.g., guide RNA or DNA) but unable to cleave the target (e.g., a dCas9 D10A / H840A double mutant) was preloaded with a Cas guide RNA specific for an adapter (e.g., an ILLUMINA i5 / i7 sequence) common to the NGS library polynucleotides. The guide RNA sequences AA (corresponding to SEQ ID NO: 1) and AB (corresponding to SEQ ID NO: 2) were tested, with AA demonstrating higher efficiency. In this case, the guide RNA sequence targets the i5 region of the ILLUMINA Tru-Seq adapter. Guide RNA sequence AA: / AltR1 / rArG rArUrC rGrGrA rArGrA rGrCrG rUrCrG rUrGrU rGrUrU rUrUrA rGrArG rCrUrA rUrGrC rU / AltR2 / Guide RNA sequence AB: / AltR1 / rGrA rUrCrG rGrArA rGrArG rCrGrU rCrGrU rGrUrA rGrUrU rUrUrA rGrArG rCrUrA rUrGrC rU / AltR2 /
[0132] DNA containing the i5 adapter sequence (target) was normalized using a dCas-nickel bead pull-down system with a guide RNA designed to target the i5 adapter sequence. Two samples, the first containing 88 femtomoles of target DNA and the second containing 44 femtomoles of target DNA, were normalized using 125 femtomoles of dCas enzyme (Figure 2). After normalization, the first and second samples were normalized to 4.3 femtomoles and 4.2 femtomoles of target DNA, respectively.
[0133] Normalization was further demonstrated by successfully normalizing a dilution series of an ILLUMINA DNA sequencing library containing 4 nM, 6 nM, 8 nM, and 10 nM input DNA to 4.68 nM, 5.4 nM, 5.55 nM, and 5.89 nM, respectively (Figure 5). These results further demonstrate that this method can be used to normalize the concentration of target polynucleotides in samples when the starting concentrations of the samples differ by more than two-fold. The results also showed that capture of target-containing DNA molecules was specifically driven by the presence of dCas9 RNP (Figure 4). Additionally, the amount of target DNA extracted per sample using the normalization method could be linearly adjusted by increasing the dCas9 RNP (dCas9 + Cas gRNA) concentration (Figure 3). Therefore, it may be possible to predict the amount of target polynucleotide that will be extracted from a sample based on the amount of dCas9 RNP used.
[0134] method 1. Forms ribonucleoprotein (RNP)-bound complexes Combine 1 uM guide RNA, 1 uM dCas9 enzyme (6His-tagged), and 1x Cas9 dilution buffer and incubate the reaction at room temperature for 5-10 minutes. This will generate RNPs specific to the i5 sequence of the library adapter, allowing for specific targeting of library molecules with the correct amount of dCas9.
[0135] 2. Perform the DNA binding reaction. Combine 10x Cas9 reaction buffer, the desired nM DNA library, and 1 uM RNP complex, and dilute the 10x Cas9 reaction buffer to 1x using milliQ water. Incubate the reaction for 1 hour at 37°C. This results in dCas9 that is tightly and specifically bound to the library molecules, allowing for subsequent capture using nickel magnetic beads, e.g., HisPur (Thermo Scientific).
[0136] 3. Bead pull-down To each DNA binding reaction, add HisPur magnetic beads and 1x Cas9 reaction buffer and incubate for 15 min. Remove the supernatant and wash the samples twice with 1x Cas9 reaction buffer to remove unbound DNA.
[0137] 4. Proteinase K Digestion (Elution) The sample is incubated with proteinase K for 10 minutes at 56°C to release (elute) the bound DNA from the dCas enzyme.
[0138] 5. Quantification The flow-through and final eluted libraries were quantified by qPCR using library-specific primers.
[0139] Step 1.1 PCR-free library normalization by Cas9-guided mediated capture of library molecules.
[0140] Optionally, this method can include an additional step (step 1.1) to generate double-stranded PAM sites on single-stranded target polynucleotides without the need for PCR amplification (Figure 7A-D). PCR amplification is a common strategy in NGS library construction, particularly because it enables NGS sequencing with small sample volumes.
[0141] However, PCR amplification can be problematic for some NGS applications. For example, PCR amplification can introduce GC bias, which can interfere with data analysis, such as the identification of novel single nucleotide polymorphisms (SNPs). To enable library normalization as described in Example 1 while avoiding PCR amplification, the workflow described in this example introduces library molecule denaturation and partial extension steps. Step 1.1 is performed before combining the RNP-bound complex with the target nucleotide.
[0142] This step begins with a sample containing a target polynucleotide containing 3' and 5' adapter sequences (Figure 7A). The reverse complement of the PAM site is encoded on the 5' adapter. A partial primer—complementary to a portion of the 3' adapter sequence and thus generating the reverse complement of the entire 3' adapter sequence when the partial primer is extended—is contacted with the target polynucleotide (Figure 7B) under conditions that allow primer extension to occur (Figure 7C) (e.g., conditions including the presence of DNA polymerase). This generates a double-stranded PAM site that can bind to the dCas protein used in the normalization method, as described herein (Figure 7D). However, the extended primer does not contain the complete 3' adapter and therefore is not sequenced during next-generation sequencing. Therefore, the sequencing results are not biased by amplified DNA (e.g., the extended primer).
[0143] method For example, the following approach can be used to generate partial second strands that allow for normalizing Illumina PCR-free libraries: The ligated, non-amplified library is combined with an excess of a primer specific to the i7 adapter sequence, such as 5'GTGACTGGAGTTCAGACGTGT'3 (SEQ ID NO: 49). This primer binds to the proximal portion of the i7 adapter and, importantly, does not contain the flow cell binding sequence from the distal adapter portion. After primer binding, the primer is extended using a polymerase in the presence of dNTPs. The resulting extended strand cannot bind to the flow cell and clusters and is therefore inactive. After extension, the library is subjected to the normalization workflow described herein.
[0144] Advantages of this method: Only the original PCR-free molecules can be clustered. This method ensures that only fully ligated library molecules (those with 5' and 3' adapters ligated to the insert) are targeted by dCas9 (or other nucleic acid-guided nucleases) in the subsequent normalization step.
[0145] Detailed protocol. 1. 100 femtomoles of the ligated library is combined with 500 femtomoles of i7 complementary primer (5'GTGACTGGAGTTCAGACGTGT'3 (SEQ ID NO: 49)) in the presence of 1x PCR buffer (Tris-HCl pH 8: 20 mM, magnesium chloride (MgCl2): 2 mM, potassium chloride (KCl): 50 mM, dNTP (each): 200 μM, and 1 unit of Taq polymerase). Any other polymerase, such as Bst, KOD, etc., can be used here. 2. The sample is subjected to a single denaturation-extension step: 95C for 30 seconds followed by 60C for 1 minute. 3. Optionally, purify the extended duplexes using SPRI beads (Ampure XP or equivalent) and subject the extended library to normalization using the methods described in the Examples. 4. Only the original ligated library molecules contain both 5' and 3' adapters and can be clustered on the NGS sequencer.
[0146] Example 2 Sample normalization for NGS by catalytically inactive Cas9 protein binding and exonuclease digestion An NGS library is provided in which the exact concentration of nucleic acids is unknown but is at least 20 nM (20 fmol / uL). The nucleic acids in the library have adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acids. The library is combined with a predetermined amount (80 fmol) of catalytically inactive Cas9 protein with D10A and H840A substitutions. The catalytically inactive Cas9 protein is preloaded with two different guide RNAs or guide DNAs, each specific to one of the adapter sequences (e.g., P5 sequence, P7 sequence, or shared Y adapter sequence) present on the nucleic acids in the NGS library. Alternatively, the catalytically inactive Cas9 protein is preloaded with a single guide RNA or guide DNA specific to the Y adapter sequence present on the nucleic acids in the NGS library, or a single guide RNA or guide DNA specific to another sequence on both ends of the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-induced) association of catalytically inactive Cas9 protein with the adapter sequence. This incubation results in a specific number of library molecules (80 fmol) with catalytically inactive Cas9 protein attached to their ends. A non-targeted nuclease, exonuclease III, is added to the reaction mixture under conditions that maintain the specific binding of catalytically inactive Cas9 protein with the adapter sequence but also allow the activity of the non-targeted exonuclease. Only the specific number of library molecules (80 fmol) with catalytically inactive Cas9 protein attached to their ends remain in the resulting NGS library, while the remaining unprotected molecules are partially or completely digested by the nuclease.
[0147] These steps are carried out simultaneously or sequentially for one or more additional NGS libraries, whose nucleic acid concentration is unknown but at least 20 nM (20 fmol / uL). The resulting NGS libraries (which now have concentrations more similar than the starting concentrations between NGS libraries) are pooled and purified to remove catalytically inactive Cas9 protein from the nucleic acid. If the nucleic acid in the library is at the desired concentration, it is sequenced. If the nucleic acid in the library is not yet at the desired concentration, it is amplified and then sequenced.
[0148] Example 3 Sample normalization for NGS by catalytically inactive CbAgo protein binding and exonuclease digestion An NGS library is provided, with an unknown but at least 20 nM (20 fmol / uL) concentration of nucleic acids. The nucleic acids in the library have adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acids. The library is combined with a predetermined amount (80 fmol) of catalytically inactive CbAgo protein. The catalytically inactive CbAgo protein is preloaded with one or more guide DNAs specific to the adapter sequences (e.g., P5, P7, or shared Y adapter sequences) present on the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-induced) association of the catalytically inactive CbAgo protein with the adapter sequences. This incubation results in a specific number of library molecules (80 fmol) with catalytically inactive CbAgo protein attached to their ends. An untargeted nuclease, exonuclease III, is added to the reaction mixture under conditions that maintain tight binding between the catalytically inactive CbAgo protein and the adapter sequence, but also allow the activity of the untargeted exonuclease. Only a certain number of library molecules (80 fmol) with catalytically inactive CbAgo protein attached to their ends remain in the resulting NGS library.
[0149] These steps are performed simultaneously or sequentially on one or more additional NGS libraries whose nucleic acid concentrations are unknown but at least 20 nM (20 fmol / uL). The resulting NGS libraries (which now have concentrations more similar than the starting concentrations between the NGS libraries) are pooled and purified to remove catalytically inactive CbAgo proteins from the nucleic acids. If the nucleic acids in the library are at the desired concentration, they are sequenced. If the nucleic acids in the library are not yet at the desired concentration, they are amplified and then sequenced.
[0150] Example 4 Alternative methods for normalizing samples for NGS using other nucleases The normalization method is performed as described in Example 2 or 3, except that exonuclease I, a restriction enzyme, or a sequence-specific nuclease is used instead of the untargeted nuclease exonuclease III. For example, the nuclease is a 5' to 3' exonuclease or a 3' to 5' exonuclease, a single-strand-specific nuclease, a double-strand-specific nuclease, or a mixture thereof. Alternatively, a catalytically active nucleic acid-guided nuclease (e.g., Cas9, CbAgo) is used instead of the untargeted nuclease. In some embodiments of the method, a nucleic acid-guided nickase (e.g., Cas9 nickase) is used instead of the untargeted nuclease. In certain embodiments, the nucleic acid-guided nickase is a Cas9 nickase with a D10A mutation or an H840A mutation. In some embodiments, a transcription activator-like effector nuclease (TALEN) or a TALE nickase is used instead of the untargeted nuclease. In certain embodiments, the TALEN or TALE nickase targets the same adapter sequence as the nucleic acid binding protein (e.g., dCas protein, dAgo, etc.) used in the method.
[0151] Example 5 Sample normalization for NGS by catalytically inactive Cas9 protein binding and Cas9 nickase digestion An NGS library is provided, with an unknown but at least 20 nM (20 fmol / uL) concentration of nucleic acids. The nucleic acids in the library have adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acids. The library is combined with a predetermined amount (80 fmol) of a nucleic acid-guided binding protein to create a reaction mixture. The nucleic acid-guided binding protein is a catalytically inactive Cas9 protein with D10A and H840A substitutions. The catalytically inactive Cas9 protein is preloaded with a single guide RNA or guide DNA specific for the shared Y adapter sequence present on the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-guided) association of the catalytically inactive Cas9 protein with the adapter sequence. This incubation results in a specific number of library molecules (80 fmol) with catalytically inactive Cas9 protein attached to their ends. Cas9 nickase is added to the reaction mixture under conditions that maintain specific binding of the catalytically inactive Cas9 protein to the adapter sequence but also allow targeted Cas9 nickase activity. Only a certain number of library molecules (80 fmol) with catalytically inactive Cas9 protein attached to their termini remain in the resulting NGS library, while the remaining unprotected molecules are partially or completely digested by the Cas9 nickase.
[0152] These steps are performed simultaneously or sequentially on one or more additional NGS libraries whose nucleic acid concentrations are unknown but are at least 20 nM (20 fmol / uL). The resulting NGS libraries (which now have concentrations more similar than the starting concentrations between NGS libraries) are pooled and purified to remove proteins from the nucleic acids. If the nucleic acids in the library are at the desired concentration, they can be sequenced. If the nucleic acids in the library are not yet at the desired concentration, they can be amplified and then sequenced.
[0153] Example 6 An alternative method for normalizing samples for NGS, involving amplification before digestion An NGS library is provided in which the exact concentration of nucleic acids is unknown. The nucleic acids in the library have adapter sequences (e.g., ILLUMINA P5 / P7 sequences and a shared Y adapter sequence (e.g., 13 bp)) at the first and second ends of the nucleic acids. Amplification is performed using primers that bind to the adapter sequences to generate an amplified library.
[0154] The amplified library is combined with a predetermined amount (80 fmol) of nucleic acid-guided binding protein to create a reaction mixture. The nucleic acid-guided binding protein is dCas9, a catalytically inactive Cas9 protein with D10A and H840A substitutions. The catalytically inactive Cas9 protein is preloaded with a guide RNA or guide DNA specific to an adapter sequence (e.g., a P5 sequence, a P7 sequence, or a shared Y adapter sequence) present on the nucleic acids in the NGS library. The reaction mixture is incubated under conditions that promote specific and strong (guide-induced) association between the catalytically inactive Cas9 protein and the adapter sequence. This incubation results in a specific number of library molecules with the catalytically inactive Cas9 protein attached to their ends. An untargeted nuclease, exonuclease III, is added to the reaction mixture under conditions that maintain tight association between the catalytically inactive Cas9 protein and the adapter sequence but also allow untargeted exonuclease activity. Only a certain number of library molecules with catalytically inactive Cas9 proteins attached to their ends remain in the resulting NGS library.
[0155] These steps are performed simultaneously or sequentially on one or more additional NGS libraries. Briefly, dCas9 is attached to each library, the protected libraries are pooled, and the pool of protected libraries is exposed to an excess amount of nuclease (e.g., exonuclease III, exonuclease I, restriction enzyme, sequence-specific nuclease, 5'→3' exonuclease, 3'→5' exonuclease, single-strand-specific nuclease, double-strand-specific nuclease, catalytically active nucleic acid-guided nuclease (e.g., Cas9, CbAgo), nucleic acid-guided nickase (e.g., Cas9 nickase), transcription activator-like effector nuclease (TALEN), or TALE nickase, or a mixture thereof). The resulting NGS libraries (now with concentrations more similar than the starting concentrations between the NGS libraries) are pooled, and standard SPRI purification is performed on the pooled NGS libraries. The nucleic acids in the libraries are then sequenced.
[0156] Example 7 Normalization of post-amplification sequencing libraries Typical Illumina sequencer library preparation chemistry targets a molar amount of the final amplified library between 200 and 1,000 femtomoles (fmol). To demonstrate library normalization capabilities across various input amounts, Illumina TruSeq adapter-ligated double-stranded DNA (dsDNA) libraries were amplified to generate high-concentration DNA inputs. DNA inputs of 50, 500, 750, or 1,000 fmol were used as input to two different normalization reactions containing either 250 fmol or 300 fmol of ribonucleoprotein (RNP, e.g., dCas9-guide RNA complex). The resulting normalized libraries were quantified using an HSD1000 DNA tape on an Agilent Tapestation over a range of 100 to 1,000 base pairs. Both 250 fmol and 300 fmol of RNP normalized this range of DNA library inputs, and the amount of resulting normalized output depended on the RNP amount. There was a 20-fold difference in DNA amount between the low DNA input of 50 fmol and the high DNA input of 1,000 fmol. After normalization with 250 fmol of RNP, there was only a 1.5-fold difference, and after normalization with 300 fmol of RNP, there was only a 1.3-fold difference. Across all samples normalized with 250 fmol of RNP, the average yield was 37 fmol, while normalization with 300 fmol of RNP yielded an average yield of 56 fmol. This demonstrates that the amount of library retained during normalization can be modulated by the amount of RNP used (Figure 8).
[0157] Example 8 Normalization of sequencing libraries prior to targeted capture Targeted capture / enrichment panels typically require high-mass amplified libraries, often ranging from 100 ng to 1,000 ng of library molecules. Depending on the library size, this mass range roughly corresponds to 300–2,000 fmol of DNA. To demonstrate the ability of normalization to achieve this final output range, an Illumina TruSeq adapter-ligated dsDNA library was amplified to generate a high-concentration DNA input, and 3,400 fmol of this DNA was used as input. Normalization was performed using 2,000, 6,000, 10,000, or 14,000 fmol of RNP. The output DNA was analyzed using a D1000 tape on an Agilent Tapestation with a 180–1,000 base pair region setting. Increasing the RNP resulted in greater capture of the DNA library. Increasing the RNP to 6,000 fmol captured approximately 2,000 fmol of library, supporting normalization for upstream use in target enrichment and similar applications (Figure 9).
[0158] Example 9 Normalization using single guide RNA (sgRNA) in RNPs The use of single-guide RNAs (sgRNAs), as opposed to separate CRISPR RNAs (crRNAs) and trans-activating CRISPR RNAs (tracrRNAs), presents a simpler format for performing normalization. To demonstrate similar performance to crRNA / tracrRNA RNPs, an Illumina TruSeq adapter-ligated dsDNA library was amplified to generate a highly concentrated DNA input, and 500 fmol of this DNA was used. Normalization was performed using 250, 600, or 1,000 fmol of RNPs containing either crRNA / tracrRNA (double-stranded guide) or sgRNA. The sgRNA sequences were as follows: 5'-mA*mG*mA* rUrCrG rGrArA rGrArG rCrGrU rCrGrU rGrUrG rUrUrU rUrArG rArGrCrUrArG rArArA rUrArG rCrArA rGrUrU rArArA rArUrA rArGrG rCrUrA rGrUrC rCrGrU rUrArU rCrArA rCrUrU rGrArA rArArA rGrUrG rGrCrA rCrCrG rArGrU rCrGrGrUrGrC mU*mU*mU*rU-3' (SEQ ID NO: 48) (mA / mG / mU are 2'-O-methyl RNA bases and * indicates a phosphorothioate linkage). The crRNA sequence was as follows: 5'- / AltR1 / rGrCrA rCrGrG rArGrA rCrGrG rArUrG rUrUrA rUrUrGrUrUrU rUrArG rArGrC rUrArUrGrCrU / AltR2 / -3' (SEQ ID NO: 50) ( / AltR1 / and / AltrR2 / are proprietary IDT modifications).
[0159] The resulting DNA was analyzed using a D1000 tape set on an Agilent Tapestation with a region setting of 180–1,000 base pairs. The sgRNA-containing RNPs were found to behave similarly to the crRNA / tracrRNA-containing RNPs. The sgRNA-containing RNPs retained approximately 20% less DNA than the crRNA / tracrRNA-containing RNPs. However, increasing the fmol of sgRNA RNP resulted in more DNA being captured, and thus, it was possible to achieve the range demonstrated with crRNA / tracrRNA with sgRNA (Figure 10).
[0160] These examples are provided to illustrate the present disclosure, but not to limit its scope. Other variations of the present disclosure will be readily apparent to those skilled in the art and are therefore encompassed by the appended claims.
Claims
1. 1. A method for normalizing the concentration of a target polynucleotide between at least two samples, each sample comprising a target polynucleotide, the method comprising: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) forming a solution comprising: (a) the sample; (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence of the target polynucleotide of the sample; and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag, wherein the dCas or dArgonaute at the predetermined concentration is cognate to the guide polynucleotide; (iii) contacting the solution with a solid phase comprising an affinity tag binding molecule capable of binding to the affinity tag; (iv) separating the solution from the solid phase; (v) extracting said target polynucleotide from said solid phase to normalize the concentration of target polynucleotide between two or more of said samples; A method comprising:
2. The method of claim 1 , wherein the targeting region is complementary to a static region of an adapter sequence.
3. 3. The method of Claim 1 or Claim 2, wherein the guide polynucleotide comprises a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
4. The method of claim 2, wherein the static region is a static region of a next-generation sequencing adapter sequence.
5. 5. The method of any one of claims 1 to 4, wherein the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 27 to 36.
6. 5. The method of claim 1, wherein the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 27-28.
7. 5. The method of any one of claims 1 to 4, wherein the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 29 to 30.
8. 5. The method of any one of claims 1 to 4, wherein the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 31 to 32.
9. 5. The method of claim 1, wherein the targeting region is complementary to an adapter sequence of any one of SEQ ID NOs: 33-34.
10. 6. The method of any one of claims 1 to 5, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
35.
11. 5. The method of any one of claims 1 to 4, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
36.
12. 12. The method of any one of claims 1 to 11, wherein the Cas gRNA polynucleotide targeting region is a homologous region.
13. 13. The method of claim 12, wherein the Cas gRNA polynucleotide comprises a homologous region of any one of SEQ ID NOs: 1-26.
14. 14. The method of Claim 13, wherein the Cas gRNA polynucleotide comprises a homologous region of SEQ ID NO:
1.
15. The method of claim 1, wherein the Argonaute guide polynucleotide is an siRNA, miRNA, piRNA, shRNA, or siDNA.
16. 16. The method of any one of claims 1 to 15, wherein the guide polynucleotide is a DNA polynucleotide.
17. 17. The method of any one of claims 1 to 16, wherein the guide polynucleotide is an RNA polynucleotide.
18. 18. The method of claim 16 or claim 17, wherein the guide polynucleotide comprises a modified nucleic acid.
19. 19. The method of claim 18, wherein the modified nucleic acid comprises 2'F RNA, 2'OMe RNA, and / or phosphorothioate linkages (PS).
20. 20. The method of Claim 19, wherein the guide polynucleotide is a Cas gRNA polynucleotide, and wherein the Cas gRNA polynucleotide does not comprise a modified nucleic acid at a position that interacts with a cognate Cas protein to the gRNA.
21. 21. The method of any one of claims 1 to 20, wherein the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, or CasY protein.
22. 22. The method of claim 21, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
23. 12. The method of any one of claims 1 to 11, wherein the d-argonaute is catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo.
24. 24. The method of any one of claims 1 to 23, wherein the affinity tag binding molecule comprises Ni2+ and the affinity tag comprises a His tag.
25. 24. The method of any one of claims 1 to 23, wherein the affinity tag binding molecule comprises biotin and the affinity tag comprises avidin.
26. 24. The method of any one of claims 1 to 23, wherein the affinity tag binding molecule comprises an anti-myc antibody and the affinity tag comprises a myc tag.
27. 27. The method of any one of claims 1 to 26, wherein the affinity tag and corresponding affinity tag binding molecule are selected from Table 1.
28. 28. The method of any one of claims 1 to 27, wherein the solid phase comprises magnetic beads.
29. 29. The method of any one of claims 1 to 28, wherein separating the solution from the solid phase comprises immobilizing the solid phase and washing the solid phase.
30. 30. The method of any one of claims 1 to 29, wherein the step of extracting the target polynucleotide from the solid phase comprises combining the solid phase with a protease.
31. 30. The method of any one of claims 1 to 29, wherein the step of extracting the target polynucleotide from the solid phase comprises combining the solid phase with proteinase K in a solution sufficient for proteinase K activity.
32. 32. The method of claim 31, wherein proteinase K digests the dCas or dArgonaute bound to the solid phase, thereby extracting the target polynucleotide.
33. 33. The method of any one of claims 1 to 32, wherein steps (i) to (v) are performed in sequential order.
34. 34. The method of any one of claims 1 to 33, wherein the normalizing step comprises bringing the concentrations of the target polynucleotide between at least two of the samples after normalization to within 15% of each other.
35. 1. A method for normalizing the concentration of a target polynucleotide between at least two samples, each sample comprising a target polynucleotide, the method comprising: (i) obtaining the sample, wherein the target polynucleotide of the sample comprises an adapter sequence; (ii) generating a solution comprising combining (a) the sample, (b) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence of the target polynucleotide of the sample, and (c) a predetermined concentration of dCas or dArgonaute comprising an affinity tag; (iii) contacting said solution with a predetermined amount of catalytically active Cas or Argonaute protein to normalize the concentration of target polynucleotide between two or more of said samples; A method comprising:
36. 36. The method of claim 35, wherein the dCas is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY protein.
37. 37. The method of Claim 36, wherein the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO:
37.
38. 37. The method of claim 36, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
39. 36. The method of claim 35, wherein the d-argonaute is catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo.
40. 40. The method of any one of claims 35 to 39, wherein (ii) comprises binding of dCas or dArgonaute to the adapter sequence of at least a portion of the target polynucleotide.
41. 41. The method of Claim 40, wherein the catalytically active Cas9 or Argonaute digests a target polynucleotide that is not bound by the dCas9 or dArgonaute.
42. 41. The method of any one of Claims 35 to 40, wherein the normalizing step comprises bringing the concentrations of the target polynucleotide between at least two of the samples after normalization to within 15% of each other.
43. 41. The method of any one of Claims 35 to 40, wherein the normalizing step comprises bringing the concentrations of the target polynucleotide between at least two of the samples after normalization to within 10% of each other.
44. 44. The method of any one of Claims 1 to 43, wherein the target polynucleotide comprises a first adapter sequence and a second adapter sequence, and the guide polynucleotide targeting region is complementary to the first adapter sequence.
45. After (i) and before (ii), (a) contacting the sample with a primer encoding a nucleic acid sequence that is complementary to a portion of the second adapter sequence, wherein the portion of the adapter sequence is located proximal to the adapter sequence; and (b) contacting the sample with a DNA polymerase under conditions sufficient to promote primer extension.
45. The method of claim 44, further comprising:
46. 46. The method of claim 45, wherein the portion of the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34.
47. 47. The method of claim 45 or claim 46, wherein the second adapter sequence is a 3' adapter sequence.
48. 47. The method of claim 45 or claim 46, wherein the second adapter sequence is a 5' adapter sequence.
49. A guide polynucleotide comprising a targeting region that is complementary to a static region of an adapter sequence of any one of SEQ ID NOs: 28 and 30, 32, 34-36.
50. 50. The guide polynucleotide of Claim 49, comprising a CRISPR-associated protein (Cas) guide RNA (gRNA) polynucleotide or an Argonaute guide polynucleotide.
51. 51. The guide polynucleotide of Claim 49 or Claim 50, wherein the static region is a static region of a next generation sequencing adaptor sequence.
52. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
28.
53. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
30.
54. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
32.
55. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
34.
56. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of SEQ ID NO:
35.
57. 52. The guide polynucleotide of Claim 51, wherein the targeting region is complementary to the adapter sequence of any one of SEQ ID NO:
36.
58. 58. The guide polynucleotide of any one of Claims 51-57, wherein the Cas gRNA polynucleotide targeting region is a homologous region.
59. 59. The guide polynucleotide of Claim 58, wherein the Cas gRNA polynucleotide comprises a region of homology to any one of SEQ ID NOs: 1-26.
60. 60. The guide polynucleotide of Claim 59, wherein said Cas gRNA polynucleotide comprises a homologous region of SEQ ID NO:
1.
61. 58. The guide polynucleotide of any one of claims 50 to 57, wherein the Argonaute guide polynucleotide is an siRNA, miRNA, piRNA, shRNA, or siDNA.
62. 62. The guide polynucleotide of any one of claims 49 to 61, which is a DNA polynucleotide.
63. 62. The guide polynucleotide of any one of claims 49 to 61, which is an RNA polynucleotide.
64. 64. The guide polynucleotide of Claim 62 or Claim 63, comprising a modified nucleic acid.
65. 65. The guide polynucleotide of Claim 64, wherein the modified nucleic acid comprises 2'F RNA, 2'Ome RNA, and / or phosphorothioate linkages (PS).
66. 66. The guide polynucleotide of Claim 65, wherein said guide polynucleotide is a Cas gRNA polynucleotide, and wherein said Cas gRNA polynucleotide does not comprise a modified nucleic acid at a position that interacts with a cognate Cas protein to said gRNA.
67. (a) a guide polynucleotide according to any one of claims 49 to 66; and (b) a Cas protein or an Argonaute protein, or a polynucleotide sequence encoding a Cas protein or an Argonaute protein; Includes a kit.
68. 68. The kit of claim 67, wherein the Cas protein is a catalytically inactive Cas protein (dCas).
69. 69. The kit of claim 67 or claim 68, wherein the Cas protein is a catalytically inactive Cas9 protein (dCas9).
70. 70. The kit of any one of Claims 67-69, wherein the Cas9 protein comprises a D10A mutation and a H840A mutation.
71. 69. The kit of any one of claims 67 to 68, wherein the Cas protein is a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, or CasY protein.
72. 72. The kit of any one of claims 68-71, wherein the dCas protein comprises an amino acid sequence having at least 95% identity to SEQ ID NO:
37.
73. 72. The kit of any one of claims 68 to 71, wherein the dCas protein comprises the amino acid sequence of SEQ ID NO:
37.
74. 68. The kit of claim 67, wherein the Argonaute protein is catalytically inactive (dArgonaute).
75. 68. The kit of claim 67, wherein the Argonaute protein is a catalytically inactive CbAgo protein.
76. 68. The kit of claim 67, wherein the Argonaute protein is a catalytically inactive LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo protein.
77. 77. The kit of any one of claims 74 to 76, wherein the dArgonaute protein comprises an amino acid sequence having at least 95% identity to any one of SEQ ID NOs: 39 to 45.
78. 77. The kit of any one of claims 74 to 76, wherein the dArgonaute protein comprises the amino acid sequence of any one of SEQ ID NOs: 39 to 45.
79. 79. The kit of any one of claims 67 to 78, further comprising a catalytically active Cas protein or a catalytically active Argonaute protein.
80. 80. The kit of Claim 79, wherein the catalytically active Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY.
81. 81. The kit of claim 80, wherein the catalytically active Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo.
82. 82. The kit of any one of Claims 67-81, wherein the guide polynucleotide is capable of binding to the Cas protein or the Argonaute protein to form a ribonucleoprotein complex, and the ribonucleoprotein complex is capable of binding to an adapter sequence.
83. 83. The kit of any one of claims 67 to 82, further comprising a primer that is complementary to a portion of the adapter sequence.
84. 84. The kit of claim 83, wherein the portion of the adapter sequence is located proximal to the adapter sequence.
85. 85. The kit of claim 84, wherein the portion of the adapter sequence is located proximal to the adapter sequence, and the adapter sequence comprises the nucleic acid sequence of any one of SEQ ID NOs: 28, 30, 32, and 34.
86. 86. The kit of any one of claims 83 to 85, wherein the adapter sequence is a 3' adapter sequence.
87. 86. The kit of any one of claims 83 to 85, wherein the adapter sequence is a 5' adapter sequence.
88. 88. The kit of any one of Claims 83 to 87, wherein the primer is not complementary to an adapter sequence that is complementary to a guide polynucleotide targeted region.
89. (i) a plurality of target polynucleotides, wherein the target polynucleotides comprise an adapter sequence of any one of SEQ ID NOs: 28, 30, 32, and 34-36; (ii) a predetermined concentration of a guide polynucleotide comprising a targeting region that is complementary to the adapter sequence; (iii) a predetermined concentration of dCas protein or dArgonaute protein; A reaction mixture comprising:
90. 90. The reaction mixture of Claim 89, wherein the predetermined concentration of the guide polynucleotide, or the dCas protein or the dArgonaute protein is less than the concentration of target polynucleotide in the reaction mixture.
91. 91. The reaction mixture of claim 89 or claim 90, wherein the d-Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo or SeAgo, or a catalytically inactive variant thereof, and the guide polynucleotide is cognate to the d-Argonaute protein.
92. 91. The reaction mixture of Claim 89 or Claim 90, wherein the dCas protein comprises a catalytically inactive Cpf1, C2c1, C2c3, C2c2, CasX, Cas9 or CasY, or a variant thereof, and the guide polynucleotide is cognate to the dCas protein.
93. 93. The reaction mixture of any one of claims 89 to 92, wherein the adapter sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 28, 30, 32, and 34-36.
94. 94. The reaction mixture of any one of claims 89 to 93, further comprising a catalytically active Cas protein or a catalytically active Argonaute protein.
95. (i) a guide polynucleotide according to any one of claims 49 to 66; and (ii) a Cas protein or an Argonaute protein that is cognate to the guide polynucleotide; and (iii) an adapter sequence; A ribonucleoprotein (RNP) complex comprising: A ribonucleoprotein (RNP) complex, wherein the targeted region of the guide polynucleotide is complementary to the adapter sequence, and the guide polynucleotide is cognate to the Cas protein or the Argonaute protein.
96. 96. The RNP complex of claim 95, wherein the Argonaute protein comprises LrAgo, PfAgo, TtAgo, AaAgo, AfAgo, MjAgo, MpAgo, NgAgo, RsAgo, CpAgo, IbAgo, KmAgo, or SeAgo, or a catalytically inactive variant thereof.
97. 96. The RNP complex of claim 95, wherein the Cas protein comprises Cpf1, C2c1, C2c3, C2c2, CasX, Cas9, or CasY, or a catalytically inactive variant thereof.
98. 98. The RNP complex of any one of claims 95 to 97, wherein the adapter sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 28, 30, 32, and 34 to 36.