Methods using guide RNAs with chemical modifications
Patent Information
- Application Number
- JP2024515702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-09
- Filing Date
- 2022-09-14
- Publication Date
- 2025-09-22
AI Technical Summary
Current CRISPR-based technologies face challenges in efficiency and stability, particularly in delivering guide RNAs (gRNAs) for gene editing and modulation due to nuclease degradation, which limits their effectiveness in vivo and ex vivo applications.
The use of chemically modified guide RNAs with phosphorothioate modifications at the 5' end and phosphonocarboxylate or thiophosphonocarboxylate modifications at the 3' end enhances the stability and editing efficiency of gRNAs, allowing for effective gene editing and modulation even under subsaturating conditions.
The modified gRNAs demonstrate significantly higher editing and modulation efficiency, achieving at least 10-5 times the yield of unmodified gRNAs, even when delivered in the presence of nucleases, thus improving the efficacy of CRISPR-based therapies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 243,985, filed September 14, 2021, and U.S. Provisional Patent Application No. 63 / 339,737, filed May 9, 2022, the entire contents of each of which are incorporated herein by reference in their entirety.
[0002] The present disclosure relates to the field of molecular biology. In particular, the present disclosure relates to cluster of regularly interspaced short palindromic repeats (CRISPR) technology. [Background technology]
[0003] Native prokaryotic CRISPR-Cas systems contain an array of short repeats (i.e., clusters of regularly interspaced short palindromic repeats, or "CRISPR") with constant length intervening variable sequences and CRISPR-associated ("Cas") proteins. The transcribed RNA of the CRISPR array is processed by a subset of Cas proteins into small guide RNAs, which generally have two components: There are at least six different systems: Type I, Type II, Type III, Type IV, Type V, and Type VI. These six systems differ in the enzymes involved in processing the RNA into mature crRNA. In native prokaryotic Type II systems, the guide RNA ("gRNA") contains two short non-coding RNAs, called CRISPR RNA ("crRNA") and trans-acting RNA ("tracrRNA"). In native V-type systems, the guide RNA contains enough crRNA to form an active complex with Cas12 (e.g., Cas12a is also known as Cpf1) protein, but does not have a tracrRNA segment. The gRNA forms a complex with Cas protein (ribonucleoprotein "RNP" complex). The gRNA:Cas protein complex binds to a target polynucleotide sequence that has a protospacer adjacent motif ("PAM") and a protospacer, where the protospacer contains a sequence complementary to a portion of the gRNA. Recognition and binding of the target polynucleotide by the gRNA:Cas protein complex induces cleavage of the target polynucleotide. The native CRISPR-Cas system functions as an immune system in prokaryotes, where the gRNA:Cas protein complex recognizes and silences exogenous genetic elements in a manner similar to RNAi in eukaryotes, thereby conferring resistance to exogenous genetic elements such as plasmids and phages.
[0004] Many improvements and refinements of CRISPR technology have been developed and continue to be developed. Early approaches include using the CRISPR-Cas system to cleave both strands of the target DNA, and editing is performed by homologous recombination or non-homologous end joining via double-strand breaks. Newer techniques include modulation of gene expression and other gene editing methods. For example, prime editing is a CRISPR-based technique for editing targeted sequences in DNA, which allows for various forms of base substitutions, such as transversions and pairwise mutations. It also allows for precise insertions and deletions, including large deletions up to about 700 bp in length. Notably, prime editing does not require an exogenous DNA repair template. Instead, a polymerized template containing the desired edit is contained within the guide RNA, which forms a complex with a Cas protein fused to a polymerase (such as reverse transcriptase). Upon binding to the target site, the Cas protein nicks the target site, and the polymerase can use the polymerized template to synthesize a new strand of DNA. Base editing is another gene editing technique in which a base editor enzyme, such as cytidine deaminase, is delivered together with a Cas protein and a guide RNA. The base editor enzyme is directed to the target site by the gRNA:Cas protein complex and catalyzes the deamination and thus the mutation of the cytidine residue at the target site. Modulation of gene expression can be achieved, for example, by fusing a transcription activator or inhibitor with a Cas protein that does not have cleavage activity but can form a complex with the gRNA and bind to the target site. As a result, the transcription activator or inhibitor can regulate gene expression at the target site. Therefore, the technology is called CRISPRa and CRISPRi, respectively, where "a" stands for activation and "i" stands for inhibition.
[0005] Despite these advances, there remains a need in the art to further improve CRISPR technology, in particular to improve the efficiency and stability of CRISPR-based systems, e.g., to support the adoption of CRISPR-based gene editing or modulation. [Brief description of the drawings]
[0006] [Figure 1] Figure 1 shows the results of a titration study in which increasing amounts of gRNA were mixed with a constant amount of Cas9 protein for transfection into 200,000 HepG2 cells targeting the HBB gene for the creation of indels at the target site. These results demonstrate the concept of saturating amounts of transfected components for editing, whereby with increasing amounts, editing activity reaches a plateau and further increases in amount do not increase editing yields for a fixed number of cells. [Diagram 2] FIG. 13 shows on-target and off-target editing of HBB in HepG2 cells transfected with subsaturating amounts of Cas mRNA and gRNA (0.0625 pmol Cas9 mRNA and 10 pmol gRNA for 0.2 million cells) after washing the cells with PBS buffer to remove residual serum. [Diagram 3] FIG. 13 shows on-target and off-target editing of HBB in HepG2 cells transfected with subsaturating amounts of Cas mRNA and gRNA (0.0625 pmol Cas9 mRNA and 10 pmol gRNA for 0.2 million cells) after washing the cells with PBS buffer to remove residual serum. [Figure 4] FIG. 13 shows on-target and off-target editing of HBB in HepG2 cells transfected with subsaturating amounts of Cas mRNA and gRNA (0.5 pmol Cas9 mRNA and 30 pmol gRNA for 0.2 million cells) when cells were not washed with PBS buffer to remove residual serum prior to transfection. [Diagram 5]FIG. 13 shows on-target and off-target editing of HBB in HepG2 cells transfected with subsaturating amounts of Cas protein and gRNA (12.5 pmol Cas9 protein and 30 pmol sgRNA for 0.2 million cells) when cells were not washed with buffer to remove serum prior to transfection. [Figure 6] Two exemplary gRNAs are depicted that incorporate 3xMS at the 5' and 3' ends (top), or 3xMS at the 5' end and 3xMP at the 3' end (bottom). [Figure 7] 1 shows the results of an experiment evaluating the relative levels of chemically modified gRNA over time in K562 cells. Cells were washed with PBS buffer to remove residual serum, and then the cells were transfected with gRNA in the absence of Cas protein. [Figure 8] FIG. 13 shows on-target and off-target editing of HBB in primary human T cells transfected with subsaturating amounts of Cas9 mRNA and gRNA (0.0625 pmol Cas9 mRNA and 5 pmol sgRNA for 0.2 million cells) after cells were washed with PBS buffer to remove residual serum. [Figure 9] FIG. 13 shows the results of cytidine base editing of HBB in K562 cells using chemically modified gRNAs with MS or MP at the 3' end compared to a control using unmodified gRNA. Cells were co-transfected with gRNA and mRNA encoding Cas9 nickase fused with cytidine deaminase. [Figure 10] 1 is a diagram depicting prime editing using an exemplary CRISPR-Cas system. [Figure 11] FIG. 1 shows the efficiency of prime editing of EMX1 in K562 cells using the first set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 12]
[0023] Figure 1 shows the efficiency of prime editing of EMX1 in Jurkat cells using the first set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 13] FIG. 13 shows the efficiency of prime editing of EMX1 in K562 cells using a second set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 14]
[0023] Figure 1 shows the efficiency of prime editing of EMX1 in Jurkat cells using a second set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 15] FIG. 13 shows the efficiency of prime editing of RUNX1 in K562 cells using the first set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 16]
[0023] Figure 1 shows the efficiency of prime editing of RUNX1 in Jurkat cells using the first set of chemically modified pegRNAs. Cells were co-transfected with pegRNAs and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 17] Illustrated are the chemical structures of two examples of chemically modified nucleotides that may be incorporated into the pegRNA disclosed herein, 2'-O-methyl-3'-phosphorothioate (MS) and 2'-O-methyl-3'-phosphonoacetate (MP). [Figure 18] Illustrates prime editing of EMX1 and RUNX1 using exemplary target sequences. [Figure 19]1 shows the results of an experiment to determine prime editing of EMX1 in K562 cells. In this case, prime editing was used to knock out the PAM in EMX1. Cells were co-transfected with pegRNA and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 20] 1 shows the results of an experiment to determine prime editing of EMX1 in Jurkat cells, where prime editing was used to knock out the PAM in EMX1. Cells were co-transfected with pegRNA and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 21] 1 shows the results of an experiment to determine prime editing of RUNX1 in K562 cells. In this case, prime editing was used to introduce a three-base insertion into RUNX1. Cells were co-transfected with pegRNA and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 22]
[0023] Figure 1 shows the results of an experiment to determine prime editing of RUNX1 in Jurkat cells. In this case, prime editing was used to introduce a three-base insertion into RUNX1. Cells were co-transfected with pegRNA and mRNA encoding Cas9 nickase fused to reverse transcriptase. [Figure 23] Graph showing the results of an experiment determining editing of the HBB sickle cell allele (and known intergenic off-target loci) in unwashed HepG2 cells co-transfected with sgRNA and mRNA encoding the Cas9 protein. [Figure 24] Graph showing the results of an experiment determining editing of the HBB sickle cell allele (and known intergenic off-target loci) in unwashed HepG2 cells transfected with a ribonucleoprotein (RNP) complex formed from a chemically modified sgRNA precomplexed with Cas9 protein. [Diagram 25]Graph showing the results of an experiment to determine editing of the HBB sickle cell allele (and known intergenic off-target loci) in unwashed HepG2 cells transfected with ribonucleoprotein (RNP) complexes formed by chemically modified 163-nt sgRNAs precomplexed with Cas9 protein. 163-nt sgRNAs were designed for the CRISPRa SAM system, but instead of using them for gene activation by CRISPRa, they were used with SpCas9 protein to generate indels. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] Provided herein are methods for CRISPR / Cas-based genome editing and / or modulation of gene expression in cells (e.g., primary cells for use in ex vivo therapy) or in vivo cells (e.g., cells in an organ or tissue of a subject, such as a human). In particular, the methods provided herein utilize chemically modified guide RNAs (gRNAs) that have higher activity or yield in gene editing or regulation compared to the corresponding unmodified gRNA. In some aspects, the disclosure provides methods for editing the sequence of a target nucleic acid in a cell or modulating the expression of a target nucleic acid by introducing a chemically modified gRNA that hybridizes with the target nucleic acid together with either a Cas protein, an mRNA encoding a Cas protein, or a recombinant expression vector that includes a nucleotide sequence encoding a Cas protein. In some aspects, the Cas protein can be a variant that lacks nuclease activity (e.g., dCas9) or has nickase activity. In some aspects, the Cas protein is a fusion protein that includes a Cas polypeptide and a reverse transcriptase polypeptide. In some aspects, the present disclosure provides methods for preventing or treating a genetic disease in a subject by administering a sufficient amount of a chemically modified gRNA to correct a genetic mutation associated with the disease (e.g., by editing the patient's genomic DNA or by modulating the expression of a gene associated with the disease).
[0008] The embodiments of the present disclosure utilize conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA that are within the skill of one of ordinary skill in the art. See Sambrook, Fritsch and Maniatis, Molecular Cloning: A Laboratory Manual, 2nd Edition (1989), Current Protocols in Molecular Biology (FMAusubel et al., eds., (1987)), the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (RI Freshney, eds. (1987)).
[0009] Non-commercially available oligonucleotides can be chemically synthesized, for example, by the solid-phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Lett. 22:1859-1862 (1981), using an automated synthesizer as described in Van Devanter et al., Nucleic Acids Res. 12:6159-6168 (1984). Purification of oligonucleotides is performed using any strategy accepted in the art, for example, native acrylamide gel electrophoresis or anion-exchange high performance liquid chromatography (HPLC) as described in Pearson and Reanier, J. Chrom. 255:137-149 (1983).
[0010] [Definitions and Abbreviations] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art.In addition, any method or material similar or equivalent to the method or material described herein can be used to carry out the method and prepare the composition described herein.For the purposes of this disclosure, the following terms are defined:
[0011] As used herein, the terms "a," "an," or "the" include not only embodiments with one member, but also embodiments with more than one member. For example, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to a "cell" includes a plurality of such cells, a reference to an "agent" includes reference to one or more agents known to those of skill in the art, and so forth.
[0012] The term "CRISPR-associated protein" or "Cas protein" or "Cas polypeptide" refers to a wild-type Cas protein, a fragment thereof, or a mutant or variant thereof. The term "Cas mutant" or "Cas variant" refers to a protein or polypeptide derivative of a wild-type Cas protein, such as a protein with one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. In certain embodiments, a "Cas mutant" or "Cas variant" substantially retains the nuclease activity of a Cas protein. In certain embodiments, a "Cas mutant" or "Cas variant" is mutated such that one or both nuclease domains are inactive (the proteins may be referred to as Cas nickases or dead Cas proteins, respectively). In certain embodiments, a "Cas mutant" or "Cas variant" has nuclease activity. In certain embodiments, a "Cas mutant" or "Cas variant" lacks some or all of the nuclease activity of its wild-type counterpart. The term "CRISPR-associated protein" or "Cas protein" also includes wild-type Cpf1 protein, also referred to as Cas12a (named after the clustered regularly interspaced short palindromic repeats 1 ribonucleoprotein or CRISPR / Cpf1 ribonucleoprotein from Prevotella and Francisella), fragments thereof, or mutants or variants thereof, of various species of prokaryotes. Cas proteins include any of the CRISPR-associated proteins, including, but not limited to, any one of six different CRISPR systems: Type I, Type II, Type III, Type IV, Type V, and Type VI.
[0013] The term "nuclease domain" of a Cas protein refers to a polypeptide sequence or domain within the protein that has catalytic activity for DNA cleavage. Cas9 typically catalyzes a double-stranded cleavage upstream of a PAM sequence. A nuclease domain may be contained in a single polypeptide chain, and cleavage activity may result from the association of two (or more) polypeptides. A single nuclease domain may consist of more than one isolated stretch of amino acids within a given polypeptide. Examples of these domains include the RuvC-like motif (amino acids 7-22, 759-766, and 982-989 in SEQ ID NO:1) and the HNH motif (amino acids 837-863); see Gasiunas et al. (2012) Proc. Natl. Acad. Set. USA 109:39, E2579-E2586 and WO / 2013176772.
[0014] A synthetic guide RNA ("gRNA") with "gRNA functionality" is one that has one or more of the functions of a naturally occurring guide RNA, such as associating with a Cas protein to form a ribonucleoprotein (RNP) complex or a function performed by a guide RNA associated with a Cas protein (i.e., a function of an RNP complex). In certain embodiments, the functionality includes binding to a target polynucleotide. In certain embodiments, the functionality includes targeting a target polynucleotide to a Cas protein or a gRNA:Cas protein complex. In certain embodiments, the functionality includes nicking a target polynucleotide. In certain embodiments, the functionality includes cleaving a target polynucleotide. In certain embodiments, the functionality includes associating with or binding to a Cas protein. For example, a Cas protein may be engineered to be a "dead" Cas protein (dCa) fused to one or more proteins or portions thereof, such as transcription factors enhancers or repressors, deaminase proteins, reverse transcriptases, polymerases, etc., such that the fused protein or portions thereof can exert its function at the target site. In certain embodiments, the functionality comprises base editing functionality. In other embodiments, the functionality comprises prime editing functionality. In certain embodiments, the functionality comprises activation, repression or interference of gene expression. In other embodiments, the functionality comprises epigenetic modification. In certain embodiments, the functionality is any other known function of guide RNA in a CRISPR-Cas system using Cas proteins, including artificial CRISPR-Cas systems using engineered Cas proteins. In certain embodiments, the functionality is any other function of natural guide RNA. Synthetic guide RNAs may have gRNA functionality to a greater or lesser extent than naturally occurring guide RNAs. In certain embodiments, synthetic guide RNAs may have greater activity for one function and less activity for another function compared to a similar naturally occurring guide RNA.
[0015] A Cas protein with single-stranded "nicking" activity refers to a Cas protein, including a Cas mutant or Cas variant, that has a reduced ability to cleave one of the two strands of dsDNA compared to a wild-type Cas protein. For example, in certain embodiments, a Cas protein with single-stranded nicking activity has a mutation (e.g., an amino acid substitution) that reduces the function of the RuvC domain (or HNH domain), resulting in a reduced ability to cleave one strand of target DNA. Examples of such variants include D10A, H839A / H840A, and / or N863A substitutions in S. pyogenes Cas9, as well as the same or similar substitutions at equivalent sites in other species of Cas9 enzymes.
[0016] A Cas protein having "binding" activity or "binding" to a target polynucleotide refers to a Cas protein that forms a complex with a guide RNA, in which the guide RNA hybridizes and base pairs with another polynucleotide, such as a target polynucleotide sequence, via hydrogen bonds between the bases of the guide RNA and the other polynucleotide. Hydrogen bonds can occur by Watson-Crick model base pairing or in any other sequence-specific manner. A hybrid can include two strands forming a duplex, three or more strands forming a multi-stranded triplex, or any combination thereof.
[0017] A "CRISPR system" is a system that utilizes at least one Cas protein and at least one gRNA to provide a function or effect, including but not limited to gene editing, DNA cleavage, DNA nicking, DNA binding, regulation of gene expression, CRISPR activation (CRISPRa), CRISPR interference (CRISPRi), and any other function that can be achieved by linking the Cas protein to another effector, thereby performing an effector function on a target sequence recognized by the Cas protein. For example, a nuclease-free Cas protein can be fused to a transcription factor, deaminase, methylase, reverse transcriptase, etc. The resulting fusion protein can be used to edit, regulate the transcription of, deaminate, or methylate a target in the presence of a guide RNA for the target. As another example in prime editing, a Cas protein is used with a reverse transcriptase or other polymerase (optionally as a fusion protein) to edit a target nucleic acid in the presence of a pegRNA.
[0018] A "fusion protein" is a protein that comprises at least two peptide sequences (i.e., amino acid sequences) covalently linked to each other, where the two peptide sequences are not covalently linked in nature. The two peptide sequences can be linked directly (by a bond between them) or indirectly (by a linker between them, which can include any chemical structure, including but not limited to a third peptide sequence).
[0019] A "prime editor" is a molecule or collection of molecules that has both Cas protein activity and reverse transcriptase activity. In some embodiments, the Cas protein is a nickase. In some embodiments, the prime editor is a fusion protein that includes both a Cas protein and a reverse transcriptase. As noted elsewhere in this disclosure, other polymerases can be used in place of reverse transcriptase for prime editing, so the prime editor may include a non-reverse transcriptase polymerase in place of RT. Various versions of the prime editor have been developed and are referred to as PEI, PE2, PE3, etc. For example, "PE2" refers to a PE complex that includes a fusion protein (PE2 protein) that includes a variant of Cas9(H840A) nickase and MMLV RT having the structure [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)] and the desired pegRNA. "PE3" refers to PE2 plus a second strand nicking guide RNA that complexes with the PE2 protein to prime the cell to repair the targeted region and introduces a nick into the unedited DNA strand, facilitating incorporation of the edit into the genome (see Anzalone et al. 2019; Liu, W02020191153). Prime editors use specialized gRNAs, referred to as prime editing gRNAs or "pegRNAs," which are described in detail elsewhere in this disclosure.
[0020] A "base editor" or "BE" is a molecule or collection of molecules that has both a Cas protein (or mutant protein) and deaminase or transglycosylation activity. Base editors (BEs) are typically fusions of a Cas domain with a nucleotide-modifying domain (e.g., naturally occurring or evolved deaminases such as cytidine deaminases, e.g., APOBEC1 ("apolipoprotein B mRNA editing enzyme catalytic polypeptide 1"), CDA ("cytidine deaminase"), and AID ("activation-induced cytidine deaminase"), or adenosine deaminases, e.g., TadA (bacterial tRNA-specific adenosine deaminase). To date, two classes of deaminase base editors have been generally described: cytosine base editors ("CBEs"), which convert targeted C:G base pairs to T:A base pairs, and adenosine base editors ("ABEs"), which convert A:T base pairs to G:C base pairs. Taken together, these two classes of base editors allow the targeted introduction of all possible pairwise mutations (C to T, G to A, A to G, T to C, C to U, and A to U). Gaudelli, NM et al., Programmable base editing of A:T to G:C in genomic DNA without DNA cleavage, incorporated herein by reference.See Nature 551, 464-471 (2017). Another nucleotide-modifying domain used for base editing is a transglycosylase domain, such as wild-type tRNA guanine transglycosylase (TGT), or a variant thereof, e.g., TGT that substitutes a first nucleobase (i.e., thymine) for a second nucleobase at the ribose-nucleobase glycosidic bond. Transglycosylase editors provide thymine-to-guanine or "TGBE" (or adenine-to-cytosine or "ACBE") transversion base editors. In some cases, base editors may also include proteins or domains that modify cellular DNA repair processes to increase the efficiency and / or stability of the resulting single nucleotide changes. In some embodiments, base editors include one or more NLSs (nuclear localization sequences) and may further include one or more uracil-DNA glycosylase inhibitor (UGI) domains, which can inhibit uracil-DNA glycosylase, thereby improving the base editing efficiency of C to T base editor proteins. In some embodiments, the Cas domain is a nickase (e.g., nCas9). In some embodiments, the Cas protein is a complete nuclease inactivated protein or dead Cas9 "dCas9". In some embodiments, the base editor is a fusion protein that includes both a Cas protein (or a portion thereof) and a deaminase (or a portion thereof). In some embodiments, the base editor is a fusion protein that includes both a Cas protein (or a portion thereof) and a transglycosylase (or a portion thereof). Different versions of base editors that show improvements over previous systems, for example, base editors with different or extended PAM compatibility (Kim, YB et al. Increasing the genome-targeting scope and precision of base editing with engineered Cas9-cytidine deaminase fusions. Nature biotechnology 35, 371-376 (2017); Hu, JHEvolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018); Li, X. et al. Base editing with a Cpf1-cytidine deaminase fusion. Nature biotechnology 36, 324-327 (2018)), high-fidelity base editors with reduced off-target activity (Hu, JH et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018); Rees, HA et al. Improving the DNA specificity and applicability of base editing through protein engineering and protein delivery. Nat Commun 8, 15790 (2017); Kleinstiver, BP, Pattanayak, V., Prew, MS & Nature, T.-SQHigh-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target (see Chen, JS et al. Enhanced proofreading governs CRISPR-Cas9 targeting accuracy. Nature 550, 407-410 (2017); Slaymaker, IM et al. Rationally engineered Cas9 nucleases with improved specificity. Science 351, 84-88 (2016)), base editors with narrower editing windows (usually about 5 nucleotides wide) (Kim, YBIncreasing the genome-targeting scope and precision of base editing with engineered Cas9-cytidine deaminase fusions.Nature biotechnology 35, 371-376 (2017)), and cytidine base editors with reduced by-products (BE4) (Komor, AC et al. Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity.Sci Adv 3, eaao4774 (2017)) have been developed. Different versions of base editors are referred to as BE1, BE2, BE3, BE4, etc. Unlike prime editors, base editors work in concert with a "traditional" gRNA (e.g., Cas9-style or Cpf1-style) that programs the Cas effector portion of the base editor to target nucleic acids at desired sequence locations. .
[0021] "Guide RNA" (or "gRNA") generally refers to an RNA molecule (or a group of RNA molecules collectively) that can bind to a Cas protein and help target the Cas protein to a specific location within a target polynucleotide (e.g., DNA). Thus, a guide RNA includes a guide sequence that can hybridize to a target sequence, and another portion of the guide RNA (the "scaffold") functions to bind to a Cas protein to form a ribonucleoprotein (RNP) complex of the guide RNA and the Cas protein. There are various styles of guide RNA, including but not limited to Cas9-style and Cpf1-style guide RNA. "Cas9-style" guide RNA includes a crRNA segment and a tracrRNA segment. As used herein, the term "crRNA" or "crRNA segment" refers to an RNA molecule or portion thereof that includes a polynucleotide-targeting guide sequence; a scaffold sequence that helps interact with a Cas protein; and, optionally, a 5'-overhang sequence. As used herein, the term "tracrRNA" or "tracrRNA segment" refers to an RNA molecule or portion thereof that contains a protein-binding segment that can interact with a CRISPR-associated protein, such as Cas9. In addition to Cas9, there are other Cas proteins that utilize Cas9-style guide RNAs, and the phrase "Cas9" is used simply to identify a representative member of the various Cas proteins that utilize this style in the term "Cas9-style". "Cpf1-style" is a single-molecule guide RNA that contains a scaffold that is 5' of the guide sequence. In the literature, Cpf1 guide RNAs are often described as having only crRNAs, not tracrRNAs. It should be noted that regardless of terminology, all guide RNAs have a guide sequence for binding to the target and a scaffold region that can interact with a Cas protein. Unlike prime editing, which uses specialized gRNAs (pegRNAs), base editing uses conventional gRNAs (i.e., Cas9-style and Cpf1-style).
[0022] The term "guide RNA" encompasses single guide RNAs ("sgRNAs") that contain all functional parts in one molecule. For example, in Cas9-style sgRNAs, the crRNA and tracrRNA segments are located in the same RNA molecule. As another example, Cpf1 guide RNA is originally a single guide RNA molecule. The term "guide RNA" also encompasses a group of two or more RNA molecules collectively; for example, the crRNA and tracrRNA segments may be located in separate RNA molecules. Furthermore, the term "gRNA" as used herein encompasses guide RNAs used in prime editing (pegRNA), base editing, and gene expression modulation and any other CRISPR technology that employs gRNA.
[0023] Optionally, a "guide RNA" may include one or more additional segments that perform one or more accessory functions upon recognition and binding by a cognate polypeptide or enzyme that performs a molecular function together with the function of the Cas protein associated with the gRNA. For example, a gRNA for prime editing (commonly referred to as a "pegRNA") may include a primer binding site and a template for reverse transcriptase. In another example, the gRNA may comprise one or more polynucleotide segments that form one or more aptamers (e.g., MS2 aptamers) that recognize and bind to an aptamer-binding polypeptide (the aptamer-binding polypeptide is optionally fused to other polypeptides (e.g., MS2-p65-HSF1) that perform auxiliary functions such as transcription activation alongside a Cas protein or Cas fusion protein (e.g., dCas9-VP64); these systems are known as synergistic activation mediator "SAM" systems; see S. Konermann et al., Genome-scale transcriptional activation an engineered CRISPR-Cas9 complex. Nature. 517, 583-588 (2015); MA Horlbeck et al., Compact and highly active next-generation libraries for CRISPR-mediated gene repression and activation. eLife. 5, e19760 (2016)).
[0024] Optionally, the "guide RNA" may comprise an additional polynucleotide segment (such as a 3' (or 5')-terminal polyuridine tail, a hairpin, a stem-loop, a toe-loop, etc.) that can increase the stability of the gRNA by preventing its degradation, as can occur by nucleases, e.g., endonucleases and / or exonucleases.
[0025] The term "guide sequence" refers to a contiguous sequence of nucleotides in a gRNA (or pegRNA) that has partial or complete complementarity with a target sequence in a target polynucleotide and can hybridize to the target sequence through base pairing promoted by a Cas protein. In some cases, the target sequence is adjacent to a PAM site (PAM sequence). In some cases, the target sequence may be located immediately upstream of the PAM sequence. The target sequence that hybridizes with the guide sequence may be immediately downstream of the complement of the PAM sequence. In other examples, such as Cpf1, the location of the target sequence that hybridizes with the guide sequence may be upstream of the complement of the PAM sequence.
[0026] Guide sequences can be as short as about 14 nucleotides or as long as about 30 nucleotides. Typical guide sequences are 15, 16, 17, 18, 19, 20, 21, 22, 23 and 24 nucleotides long. The length of the guide sequence varies between the two classes and six types of CRISPR-Cas systems mentioned above. Synthetic guide sequences for Cas9 are usually 20 nucleotides long, but can be longer or shorter. If the guide sequence is shorter than 20 nucleotides, this is typically a deletion from the 5' end compared to the 20 nucleotide guide sequence. As an example, the guide sequence can consist of 20 nucleotides complementary to the target sequence. In other words, the guide sequence is identical to the 20 nucleotides upstream of the PAM sequence except for the A / U difference between DNA and RNA. If this guide sequence is truncated by 3 nucleotides from the 5' end, nucleotide 4 of the 20 nucleotide guide sequence now becomes nucleotide 1, which is 17 bases long, nucleotide 5 of the 20 nucleotide guide sequence now becomes nucleotide 2, which is 17 bases long, and so on. The new position for the 17-mer guide sequence is the original position minus 3.
[0027] As used herein, the term "prime editing guide RNA" (or "pegRNA") refers to a guide RNA (gRNA) that includes a reverse transcriptase template sequence that encodes one or more edits to a target sequence of a nucleic acid and a primer binding site (also called a target site) that can bind to a sequence in the target region. For example, a pegRNA can include a reverse transcriptase template sequence that includes one or more nucleotide substitutions, insertions, or deletions to a sequence in the target region. The pegRNA functions to form a complex with a Cas protein and hybridize to a target sequence, typically within a target region in the genome of a cell, resulting in editing of a sequence in the target region. Without being bound by theory, in some embodiments, the pegRNA forms an RNP complex with the Cas protein and binds to the target sequence within the target region, the Cas protein nicks one strand of the target region resulting in a flap, the primer binding site of the pegRNA hybridizes to the flap, and reverse transcriptase uses the flap as a primer on the reverse transcriptase template of the hybridized pegRNA, which acts as a template to synthesize a new DNA sequence on the nicked end of the flap, which then contains the desired edit, and finally, this new DNA sequence replaces the original sequence within the target region resulting in editing of the target.
[0028] A "pegRNA" may contain a reverse transcriptase template and primer binding site near its 5' or 3' end. The "prime editing end" is one end of the pegRNA, either 5' or 3', that is closer to the reverse transcriptase template and primer binding site than the guide sequence. The other end of the pegRNA is the "distal end" that is closer to the guide sequence than the reverse transcriptase template or primer binding site. Thus, the order of these components, in the 5' or 3' direction, is: Prime editing end - (primer binding site and template for reverse transcriptase) - (guide sequence and scaffold) - distal end where the brackets indicate that the order of the two segments described therein can be switched between each other depending on the format of the pegRNA (e.g., Cas9 format or Cpf1 format) as well as the location of the prime editing end (i.e., 5' end or 3' end). It should be noted that if the pegRNA is not a single guide RNA but comprises more than one RNA molecule, the prime editing end refers to the end in the RNA molecule that contains the primer binding site and the template for reverse transcriptase that is closer to these components, whereas the opposite end of this RNA molecule is the distal end. The guide sequence may be in a different RNA molecule of the pegRNA, separate from the RNA molecule with the prime editing end and the distal end.
[0029] A "nicking guide RNA" or "nicking gRNA" is a guide RNA (not a pegRNA) that can be added to a prime edit on demand to cause nicking of the unedited strand within or near the target region. Such nicking helps prime the cell in which the prime edit takes place to repair the relevant region, i.e. the target region.
[0030] An "extended tail" is a nucleotide stretch of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that can be added to either the 5' or 3' end of a guide RNA, such as a pegRNA. A "poly(N) tail" is a homopolymeric extended tail containing 1-10 nucleotides with the same nucleobase, e.g., A, U, C, or T. A "polyuridine tail" or "poly U tail" is a poly(N) tail containing 1-10 uridines. Similarly, a "poly A tail" contains 1-10 adenosines.
[0031] The term "nucleic acid", "nucleotide", or "polynucleotide" refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) and polymers thereof in either single-, double-, or multiple-stranded form. The term includes, but is not limited to, single-, double-, or multiple-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and / or pyrimidine bases or other natural, chemically modified, biochemically modified, non-natural, synthetic, or derivatized nucleotide bases. In some embodiments, nucleic acids may include mixtures of DNA, RNA, and their analogs. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid. Unless otherwise indicated, a particular nucleic acid sequence also encompasses the sequence explicitly stated, as well as implicitly encompassing its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences. Specifically, substitution of degenerate codons may be accomplished by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid is used interchangeably with gene, cDNA, and mRNA encoded by a gene.
[0032] The term "nucleotide analog" or "modified nucleotide" refers to a nucleotide that contains one or more chemical modifications (e.g., substitutions) in or on the nitrogenous base of the nucleoside (e.g., cytosine (C), thymine (T) or uracil (U), adenine (A) or guanine (G)), in or on the sugar moiety of the nucleoside (e.g., ribose, deoxyribose, modified ribose, modified deoxyribose, six-membered sugar analogue, or open-ring sugar analogue), or to the phosphate.
[0033] The term "gene" or "nucleotide sequence encoding a polypeptide" refers to a segment of DNA involved in producing a polypeptide chain. The DNA segment may include intervening sequences (introns) between individual coding segments (exons) as well as regions preceding and following the coding region (leader and trailer) which are involved in the transcription / translation of the gene product and regulation of transcription / translation.
[0034] The terms "polypeptide", "peptide" and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. These terms apply to naturally occurring and non-naturally occurring amino acid polymers, as well as to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of corresponding naturally occurring amino acids. As used herein, these terms encompass any length of amino acid chain, including full-length proteins, in which amino acid residues are linked by covalent peptide bonds.
[0035] The term "nucleic acid", "polynucleotide" or "oligonucleotide" refers to a DNA molecule, an RNA molecule, or an analog thereof. As used herein, the terms "nucleic acid", "polynucleotide" and "oligonucleotide" include, but are not limited to, DNA molecules such as cDNA, genomic DNA or synthetic DNA, and RNA molecules such as guide RNA, messenger RNA or synthetic RNA. Moreover, as used herein, these terms include single-stranded and double-stranded forms.
[0036] The term "hybridization" or "hybridizing" refers to the process in which fully or partially complementary polynucleotide strands come together under suitable hybridization conditions to form a double-stranded structure or region in which the two constituent strands are connected by hydrogen bonds. As used herein, the term "partial hybridization" includes cases in which the double-stranded structure or region contains one or more bulges or mismatches. Hydrogen bonds typically form between adenine and thymine, or adenine and uracil (A and T, or A and U, respectively), or cytosine and guanine (C and G), although other atypical base pairs may form (see, e.g., Adams et al., The Biochemistry of the Nucleic Acids, 11th ed., 1992). It is believed that modified nucleotides may form hydrogen bonds that atypically enable or facilitate hybridization.
[0037] The term "complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either by traditional Watson-Crick or other non-traditional methods. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. "Substantially complementary," as used herein, refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.
[0038] As used herein, the terms "portion," "segment," "element," or "fragment" of a sequence refer to any portion of a sequence (e.g., a nucleotide subsequence or an amino acid subsequence) that is shorter than the complete sequence. A portion, segment, element, or fragment of a polynucleotide can be any length greater than one, for example, at least 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, 150, 200, 300, or 500 nucleotides in length or longer.
[0039] The term "oligonucleotide" as used herein means a multimer of nucleotides. For example, an oligonucleotide may have a length of about 2 to about 200 nucleotides, up to about 50 nucleotides, up to about 100 nucleotides, up to about 500 nucleotides, or any integer value between 2 and 500 nucleotides. In some embodiments, an oligonucleotide may range from 30 to 300 nucleotides or 30 to 400 nucleotides. An oligonucleotide may contain ribonucleotide monomers (i.e., may be an oligoribonucleotide) and / or deoxyribonucleotide monomers. An oligonucleotide may be 10 to 20, 21 to 30, 31 to 40, 41 to 50, 51 to 60, 61 to 70, 71 to 80, 80 to 100, 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, or 350 to 400 nucleotides in length, e.g., any integer value between these ranges.
[0040] A "recombinant expression vector" is a recombinantly or synthetically produced nucleic acid construct that has a set of specific nucleic acid elements that allow the transcription of a particular polynucleotide sequence in a host cell. An expression vector can be part of a plasmid, a viral genome, or a nucleic acid fragment. Typically, an expression vector contains a polynucleotide to be transcribed operably linked to a promoter. "Operably linked" in this context means two or more genetic elements, such as a polynucleotide coding sequence and a promoter, positioned relative to each other such that the proper biological function of the element, such as the promoter, directing the transcription of the coding sequence, is exerted. The term "promoter" is used herein to refer to an array of nucleic acid control sequences that direct the transcription of a nucleic acid. As used herein, a promoter includes necessary nucleic acid sequences near the transcription start site, such as a TATA element in the case of a polymerase II type promoter. A promoter also includes distal enhancer or repressor elements, which may be located as far away as several thousand base pairs from the transcription start site, as needed. Other elements that may be present in an expression vector include those that enhance transcription (e.g., enhancers) and terminate transcription (e.g., terminators), as well as those that confer a certain binding affinity or antigenicity to the recombinant protein produced from the expression vector.
[0041] "Recombinant" refers to a genetically modified polynucleotide, polypeptide, cell, tissue, or organism. For example, a recombinant polynucleotide (or a copy or complement of a recombinant polynucleotide) is one that has been engineered using well-known methods. A recombinant expression cassette that includes a promoter operably linked to a second polynucleotide (e.g., a coding sequence) can include a promoter that is heterologous to the second polynucleotide as a result of human manipulation (e.g., by the methods described in Sambrook et al., Molecular Cloning - A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, (1989) or Current Protocols in Molecular Biology, Vol. 1-3, John Wiley & Sons, Inc. (1994-1998)). A recombinant expression cassette (or expression vector) typically includes polynucleotides in combinations not found in nature. For example, human-engineered restriction sites or plasmid vector sequences may flank or separate the promoter from other sequences. Recombinant proteins are those expressed from recombinant polynucleotides, and recombinant cells, tissues, and organisms are those that contain recombinant sequences (polynucleotides and / or polypeptides).
[0042] "Editing" a nucleic acid target means making a change in the nucleotide sequence of the target. The change can be an insertion, deletion or substitution of a single nucleotide or multiple nucleotides, respectively. When multiple nucleotides are inserted, deleted or substituted, these nucleotides can be consecutive or non-consecutive. The change can also be a combination of any of the above. "Editing" includes "base editing" and "prime editing" techniques.
[0043] "Editing efficiency" is a measure of the Cas-induced editing achieved in one or more cells. The results of genome editing at the target and potential off-target sites can be measured using standard methods known in the art, such as genomic DNA sequencing, RNA sequencing, or deep sequencing of PCR amplicons of the target site and any off-target sites of interest. Indel mutations in genomic DNA can also be identified using the SURVEYOR® Mutation Detection Kit (Integrated DNA Technologies, Coralville, Iowa) or Guide-it™ Indel Identification Kit (Clontech, Mountain View, CA). In addition, techniques that measure the presence or absence of proteins, such as gel or capillary electrophoresis, western blotting, flow cytometry, or mass spectrometry techniques, can be used to quantify the efficiency of editing aimed at introducing or knocking out protein-coding genes. These techniques can be applied to cell populations in bulk preparations or at the single cell level. In some embodiments, efficiency is measured using the number of correct edits in a cell population measured in bulk or at the single cell level. In some embodiments, efficiency is measured as the percentage of correctly edited targets, or the number or percentage of cells exhibiting a corrected genotype or phenotype.
[0044] "Modulating expression of a gene" means altering (reducing or activating) the expression of a specific gene product. CRISPR activation or "CRISPRa" refers to the activation of a gene, whereas CRISPR interference or "CRISPRi" refers to the interference of gene expression. Both systems use a nuclease-deficient Cas protein (dCas9) fused or interacting in combination with a transcription effector (activator or repressor). CRISPRa may be performed with the SAM system (dCas9-VP64) as described above. When used in gene-specific CRISPRa, a gRNA containing the MS2 aptamer recruits the MS2-p65-HSF1 fusion to the transcription start site (TSS) of the targeted gene to initiate activation. CRISPRa and CRISPRi may both be performed and combined in a multiplex manner (e.g., targeting multiple genes). CRISPRoff is a programmable epigenetic memory writer consisting of a dead Cas9 fusion protein that establishes DNA methylation and repressive histone modifications that can heritably alter gene expression (Nunez et al., Genome-wide programmable transcriptional memory by CRISPR-based epigenome editing, Cell. (2021) 184(9):2503-2519).
[0045] "Gene expression modulation efficiency" can be measured, for example, by techniques that measure the relative or absolute levels of different RNAs, such as qRT-PCR or RNA sequencing, or by various methods that measure the relative or absolute levels of proteins, such as gel or capillary electrophoresis, Western blotting, flow cytometry, or mass spectrometry techniques. These techniques can be applied to a population of cells in bulk preparations or at the single cell level. In some embodiments, efficiency is measured using the amount of protein or RNA expressed from target genes in a cell population measured as bulk or at the single cell level.
[0046] The term "single nucleotide polymorphism" or "SNP" refers to a change in a single nucleotide with respect to a polynucleotide, including within an allele. This can include the exchange of one nucleotide for another, as well as the deletion or insertion of a single nucleotide. Most typically, SNPs are biallelic markers, although tri- and tetraallelic markers can exist. As a non-limiting example, a nucleic acid molecule containing SNP A\C can contain a C or an A at the polymorphic position.
[0047] "Nuclease" as used herein means an enzyme that can break the phosphodiester linkage between nucleotides of nucleic acid. Nucleases can variously cause both single-stranded and / or double-stranded breaks in DNA and / or RNA molecules. In living organisms, they are essential mechanisms for many aspects of DNA repair. As used herein, nuclease refers to both exonucleases and endonucleases, and includes not only ribonucleases but also deoxyribonucleases.
[0048] The term "primary cells" refers to cells that are directly isolated from a multicellular organism. Primary cells have typically undergone few population doublings and are therefore more representative of the main functional components of the tissue from which they are derived than continuous (tumor or artificially immortalized) cell lines. In some cases, primary cells are cells that are used immediately after isolation. In other cases, primary cells cannot divide indefinitely and therefore cannot be cultured in vitro for long periods of time.
[0049] The term "nuclease-containing liquid" is used herein to refer to any medium in which nucleases are present. For example, the medium may be a cell culture medium or a medium derived from a cell culture medium, meaning that the cells are transferred from the cell culture medium to a new medium without washing the cells or removing all components contained in the original medium, and therefore may still contain nucleases. For example, the cells may be transferred from the cell culture medium to the reaction medium without washing the cells and without removing substantially all components of the cell culture medium, and therefore nucleases may be present when the cells are contacted with the gRNA and Cas protein (RNP), or the mRNA or DNA vector encoding the gRNA and the editing Cas effector. The liquid may be serum, human serum, animal serum, bovine serum (BSA), fetal serum, cerebrospinal fluid (CSF) or another body fluid.
[0050] The terms "cultivate", "culturing", "growing", "growing", "maintaining", "maintaining", "expanding", "expanding" and the like when referring to the cell culture itself or the culturing process can be used interchangeably to mean that cells (e.g., primary cells) are maintained in controlled conditions outside their normal environment, e.g., conditions suitable for survival. Keeping cultured cells alive and culturing can result in cell growth, quiescence, differentiation or division. The terms do not imply that all cells in the culture survive, grow, or divide, as some cells may naturally die or senesce. Cells are typically cultured in a medium that can be replaced during the culturing process.
[0051] The terms "subject," "patient," and "individual" are used interchangeably herein to include humans or animals. For example, an animal subject can be a mammal, a primate (e.g., monkey), a livestock animal (e.g., horse, cow, sheep, pig, or goat), a companion animal (e.g., dog, cat), a laboratory test animal (e.g., mouse, rat, guinea pig, bird), an animal of veterinary significance, or an animal of economic significance.
[0052] As used herein, the term "administering" includes oral administration, topical contact, administration as a suppository, intravenous, intraperitoneal, intramuscular, intralesional, intrathecal, intranasal, or subcutaneous administration to a subject. Administration is by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, or transdermal). Parenteral administration includes, for example, intravenous, intramuscular, intraarteriolar, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, and the like.
[0053] The term "treating" refers to an approach to obtain beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. By therapeutic benefit, it is meant any therapeutically meaningful improvement or effect on one or more diseases, conditions, or symptoms being treated. For prophylactic benefit, the composition may be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more of the physiological symptoms of the disease, even though the disease, condition, or symptom may not yet be manifested.
[0054] The term "effective amount" or "sufficient amount" refers to an amount of an agent (e.g., Cas protein, modified gRNA / pegRNA, etc.) sufficient to cause a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the subject and condition being treated, the subject's weight and age, the severity of the condition, the mode of administration, etc., which can be readily determined by one of skill in the art. The particular amount may vary depending on one or more of the particular agent selected, the type of target cell, the location of the target cell within the subject, the dosing regimen to be followed, whether it is administered in combination with other agents, the timing of administration, and the physical delivery system by which it is delivered.
[0055] As disclosed herein, some ranges of values are provided. It is understood that each intervening value between the upper and lower limits of the range is also specifically contemplated. Each smaller range or intervening value encompassed by a stated range is also specifically contemplated. The term "about" generally refers to ±10% of the indicated number. For example, "about 10%" may indicate a range of 9% to 11%, and "about 20" may mean 18 to 22. Other meanings of "about," such as rounding, may be apparent from the context, so for example, "about 1" may also mean 0.5 to 1.4.
[0056] Some chemically modified nucleotides are described herein.It should be noted that MS, MP and MSP can each refer to the corresponding modification or the nucleotide containing the corresponding modification.The following abbreviations shall be used in the relevant context.
[0057] "PACE": phosphonoacetate
[0058] "MS": 2'-O-methyl-3'-phosphorothioate
[0059] "MP": 2'-O-methyl-3'-phosphonoacetate
[0060] "MSP": 2'-O-methyl-3'-thiophosphonoacetate
[0061] "2'-MOE": 2'-O-methoxyethyl
[0062] Other definitions of terms may appear throughout the specification.
[0063] The present invention demonstrates that specific modifications of guide RNA at specific positions make the guide RNA particularly resistant to degradation by nuclease. This is particularly important for in vivo delivery of guide RNA for CRISPR-mediated gene editing or gene expression modulation, since in vivo nuclease activity is high. For example, body fluids such as serum and cerebrospinal fluid (CSF) contain relatively abundant nucleases. In such a challenging environment, guide RNA tends to be degraded, and therefore its concentration does not reach a level that can achieve higher performance (i.e., below saturation). Therefore, any increase in guide RNA concentration, and therefore in gene editing opportunities and gene expression modulation, is important in this industry. The present invention provides a surprising discovery that certain guide RNAs, such as those with phosphorothioate modification at 5' end, as well as phosphonocarboxylate or thiophosphonocarboxylate modification at 3' end, have high CRISPR activity even in the presence of serum, compared to unmodified or other modification-containing counterparts (e.g., containing phosphorothioate instead of phosphonocarboxylate or thiophosphonocarboxylate, but otherwise the same).
[0064] Similarly, cells to be subjected to CRISPR-mediated editing / modulation for ex vivo therapy are usually in cell culture medium containing serum, or contain body fluids in the environment if they are freshly collected from a subject. As such, nucleases present in serum or body fluids will degrade the guide RNA delivered to cells, reducing the efficiency of CRISPR-mediated editing / modulation. Although cells can be washed to reduce the amount of serum or body fluids before CRISPR treatment, extensive washing can be harmful to cells. In addition, CRISPR-mediated editing / modulation does not occur immediately after guide RNA and other CRISPR effectors are added to cells, but rather cells need to be cultured for a period of time. Cultivation in the absence of serum is often harmful to cells, which is a risk factor for ex vivo therapy as cells are then delivered to patients. The modified guide RNA of the present invention, which is more resistant to nuclease degradation, is a significant improvement to solve these problems.
[0065] The modified guide RNAs of the present invention are useful when introduced "naked" into a cell and directly exposed to nucleases, e.g., co-transfected or otherwise delivered with DNA or mRNA encoding a Cas protein. However, the modifications described herein are similarly advantageous even if the guide RNA is not naked, e.g., present in a ribonucleoprotein (RNP) with a Cas protein, or in a nanoparticle with or without a Cas protein.
[0066] In one aspect of the disclosure, a method is provided for editing a target region in a nucleic acid in a cell or modulating the expression of a target gene in the target region. The method includes providing a cell with a) a CRISPR-associated ("Cas") protein, and b) a modified guide RNA that includes a 5'-end and a 3'-end, a guide sequence that can hybridize with a target sequence in the target region, and a scaffold region that interacts with the Cas protein. The modified guide RNA also includes one or more phosphorothioate modifications within the five nucleotides of the 5'-end, and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the five nucleotides of the 3'-end. The cell is present ex vivo in the presence of a nuclease-containing liquid, or is present in vivo. In the method, providing the cell with the Cas protein and the modified guide RNA results in the editing of the target region or the modulation of the expression of the target gene.
[0067] In some embodiments, the modified guide RNA comprises 2, 3, 4, or 5 phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal 5 nucleotides. At least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal 5 nucleotides of the gRNA may comprise at least 2, 3, 4, or 5 MP nucleotides, which may be arranged in any order, including 2 consecutive and 1 or 2 non-consecutive modified nucleotides, 3 consecutive and 1 non-consecutive modified nucleotides, 2 pairs of 2 consecutive modified nucleotides, or 5 of 2 consecutive modified nucleotides. In some embodiments, the modified guide RNA comprises at least 2, at least 3, at least 4, or 5 consecutive MP nucleotides within the 3'-terminal 5 nucleotides. In some embodiments, the modified guide RNA comprises 1, 2, 3, 4, or 5 phosphorothioate modifications within the 5'-terminal 5 nucleotides. The one or more phosphorothioate modifications within the 5 nucleotides at the 5' end of the gRNA may comprise at least one, at least two, at least three, at least four, or five MS nucleotides, which may be arranged in any order, including consecutive or non-consecutive. In some embodiments of the method, the modified guide RNA comprises at least two, at least three, at least four, or five consecutive MS nucleotides within the 5 nucleotides at the 5' end. The one or more modified nucleotides within the 5 nucleotides at the 3' or 5' end of the gRNA may be independently selected (e.g., the number and / or order of modified nucleotides may be different at the 5' and 3' ends of the gRNA).
[0068] In some embodiments, the modified guide RNA further comprises modified nucleotides located outside the 5 nucleotides in the 5'-end and 3'-end.The modified guide RNA may comprise one or more modifications in the guide sequence that enhance target specificity (e.g., as described in U.S. Patent No. 10,767,175).For example, the modified guide RNA may comprise modified nucleotides at position 5 or position 11 of the modified guide sequence.
[0069] In some embodiments, the modified guide RNA is a single-stranded guide RNA. In some embodiments, the guide RNA is exactly 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 1 10, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170 , 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200 nucleotides or at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73 , 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140,141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173 , 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200 nucleotides, and / or up to 18 0, 179, 178, 177, 176, 175, 174, 173, 172, 171, 170, 169, 168, 167, 166, 165, 164, 163, 162, 161, 159, 158, 157, 156, 155, 154, 153, 152, 151, 150, 149, 148, 147, 14 1, 2, 3, 4, 5, 6, 145, 144, 143, 142, 141, 140, 139, 138, 137, 136, 135, 134, 133, 132, 131, 130, 129, 128, 127, 126, 125, 124, 123, 122, 121, or 120 nucleotides. It is expressly contemplated that any of the foregoing minimums and maximums can be combined to form a range, so long as the minimum is less than the maximum.
[0070] In some embodiments, the Cas protein is provided to the cell as an mRNA encoding the Cas protein or a variant or fusion protein thereof. In some embodiments, the Cas protein is provided to the cell as a recombinant expression vector comprising a nucleotide sequence encoding the Cas protein or a variant or fusion protein thereof. The cell can be transfected with the mRNA or expression vector encoding the Cas protein separately or together with the modified guide RNA. In some embodiments, the cell is co-transfected with the modified guide RNA and the mRNA or expression vector encoding the Cas protein. When co-transfected, the modified guide RNA and the mRNA or expression vector encoding the Cas protein can be provided to the cell in separate delivery systems or in a single delivery system. Alternatively, the modified guide RNA may be transfected into the cell before or after transfection of the mRNA or expression vector encoding the Cas protein. The cell can be transfected by electroporation, microinjection, lipofection or exposure to nanoparticles or other delivery systems (described in more detail below). In some embodiments, the mRNA or expression vector encoding the Cas protein and / or the modified guide RNA is provided in a nanoparticle, e.g., a lipid nanoparticle.
[0071] In some embodiments, Cas protein and modified guide RNA are provided as ribonucleoprotein complex (RNP). Modified guide RNA can be complexed with Cas protein or its variant or fusion protein to form RNP for introduction into cells. RNP can be provided to cells in a delivery system, such as by electroporation, microinjection, virus-like particle, lipofection or exposure to nanoparticles or other delivery systems (described in more detail below). In some embodiments, Cas protein and / or modified guide RNA are provided in nanoparticles, for example, lipid nanoparticles.
[0072] In some embodiments, the cell to be edited or modulated is ex vivo. In other embodiments, the cell to be edited or modulated is in vivo. The method can be used to edit a target region in a nucleic acid in an ex vivo cell that has been previously cultured in a medium containing serum, or to modulate the expression of a target gene in a target region, where the cell has been incompletely separated from serum or one or more serum components. For example, the method can include transferring the cell from a cell culture medium to a reaction medium without washing the cell or without extensive washing of the cell. In some embodiments, the modified guide RNA and Cas protein are provided to the cell (e.g., an in vivo cell in blood, plasma or serum) in the presence of serum or one or more serum components.
[0073] In some embodiments, the cells are a population of cells each containing a target region. For example, the population of cells may be a cell culture or may be derived from a cell culture. The cells or population of cells may be in a cell culture medium or a nuclease-containing solution before the cells are provided with the modified guide RNA and Cas protein, and in some embodiments, the cells are washed but not completely free of cell culture medium or one or more components of the cell culture medium before the introduction of the editing components, so that the nuclease is still present. For example, the cells may be transferred from the cell culture medium to the reaction medium without washing the cells and without removing substantially all components of the cell culture medium. Alternatively, the cells or population of cells may be in the cell culture medium at the time of providing the modified guide RNA and Cas protein. In such an embodiment, the cell culture medium can function as a reaction medium for editing or modulating a target sequence in the cells. In some embodiments, the cells are in or transferred from a cell culture medium that includes serum or one or more other medium components, such as one or more natural proteins of human or animal origin. In some embodiments, the cells are in or transferred from cell culture medium that includes bovine serum albumin, horse serum, or fetal bovine serum.
[0074] In some embodiments, the editing or modulation of expression resulting from providing a cell with a modified guide RNA is at least 10%, at least 20%, at least 25%, or at least 50% more efficient than the editing or modulation caused by an unmodified guide RNA that is otherwise identical to the modified guide RNA. For example, the present invention has an average indel yield or an average editing yield that is at least 10%, at least 20%, at least 25%, or at least 50% higher than the yield obtained by a corresponding method employing an unmodified guide RNA that is otherwise identical to the modified guide RNA. In some embodiments, the editing or modulation resulting from providing a cell with a modified guide RNA is at least 2-fold, at least 3-fold, or at least 5-fold more efficient than the editing or modulation caused by an unmodified guide RNA that is otherwise identical to the modified guide RNA. For example, the present method has an average indel yield or an average editing yield that is at least 2-fold, at least 3-fold, or at least 5-fold higher than a corresponding method employing an unmodified guide RNA that is otherwise identical to the modified guide RNA.
[0075] In the present invention, multiplexing is contemplated by using multiple modified gRNAs for multiple target regions. In some embodiments, two modified guide RNAs of the present application are used to edit (or modulate) two different target regions in the same cell, preferably simultaneously. In some embodiments, a modified guide RNA is used to edit a first target region, and a second modified guide RNA is used to multiplex modulate the expression of a target region (which may be the same as or different from the first target region).
[0076] In recent years, CRISPR-based technology has emerged as a potentially revolutionary therapeutic approach (e.g., to correct genetic defects). However, the use of CRISPR systems has been limited due to practical issues. In particular, there is a need for methods to stabilize guide RNAs (gRNAs) for in vivo delivery of CRISPR-Cas components. Previous studies have investigated the use of gRNAs with chemically modified nucleotides. As described herein, the present disclosure is based in part on the surprising discovery that incorporation of specific modified nucleotides at the 3' end of gRNA can improve the yield of Cas-mediated editing or modulation of target nucleic acids, with significant improvement when gRNA and mRNA or DNA encoding Cas proteins are introduced into cells under harsh conditions (e.g., co-transfection).
[0077] In some aspects, the guide RNAs disclosed herein are useful in applications where the guide RNA is introduced into a cell in one or more challenging circumstances, e.g. i. the cells are in medium containing serum (e.g., fetal bovine serum); ii. the cells have been previously cultured in medium containing serum and serum is still present when the guide RNA is introduced; iii. the cells have been pre-cultured in medium containing one or more nucleases and the nucleases are still present when the guide RNA is introduced; iv. the cells have a relatively high level of nuclease activity, e.g., a relatively high expression of one or more nucleases; v. The cells have a relatively low level of nuclease inhibitor activity, e.g., a relatively low expression of a nuclease inhibitor; vi. the modified guide RNA is not complexed with a Cas protein prior to delivery into the cell; vii. the cell is present in vivo; and viii. Combinations thereof, where applicable This may be particularly advantageous in
[0078] The concept of saturation is well known in the art. When a substance is at its "saturation level", any further increase in the amount of substance will not result in higher activity. A "subsaturating level" is below the saturation level, and adding more of the substance in question may result in higher activity. The saturation threshold may be determined empirically using conventional assays. For example, Figure 1 shows the results of an assay in which Cas editing activity was evaluated after co-transfection of increasing amounts of synthetic gRNA and a constant amount of Cas protein. As shown by this figure, Cas-mediated editing activity reaches a plateau at 25-31.25 pmol of gRNA, when the level of gRNA reaches a saturation point for transfection of 200,000 cells.
[0079] In many cases, it is ideal to use saturating levels of components required by the chemical reaction. However, such conditions are not always feasible, especially in the case of therapies where saturating levels of one or more compounds may not be possible or safe to treat humans or animals. For CRISPR-based therapies, transfection efficiency is typically a bottleneck that limits the effectiveness of the therapy. For example, current CRISR-based therapies typically require co-transfection of gRNA and mRNA encoding Cas protein into one or more cells of a patient. If transfection efficiency is low, one or more components of the therapy may be delivered at a level below the effective amount required for therapeutic effect. The modified guide RNA constructs disclosed herein address this need in the art in that they typically exhibit high levels of Cas editing activity when transfected at subsaturating levels. Indeed, the incorporation of one or more phosphonocarboxylate modifications at the 3' end of synthetic gRNA is particularly advantageous for CRISPR-based methods involving co-transfection of Cas mRNA with synthetic gRNA.
[0080] As mentioned above, the present disclosure also provides modified pegRNA constructs and methods that retain high levels of prime editing activity in harsh conditions, for example when transfected in less than saturating amounts. This result is particularly surprising because the structure of traditional guide RNAs (gRNAs) is very different from that of prime editing gRNAs (pegRNAs), and before the present disclosure, it was unclear how chemical modifications of pegRNAs would affect their activity. In particular, pegRNAs contain additional sequences in their 3' portion compared to typical gRNAs (i.e., reverse transcriptase template and primer binding site sequences), and the 3' end of pegRNAs plays a different function in prime editing than the 3' end of typical gRNAs in other CRISPR-Cas systems. Thus, phosphoribose (or other chemical) modifications at the 3' end of pegRNAs have the potential to interfere with the role of primer binding site sequences. The primer binding site sequence hybridizes to the 3' end of the nicked strand of the DNA target site, such that reverse transcriptase recognizes the resulting RNA:DNA duplex as an acceptable substrate for primer extension of the nicked 3' end to achieve primed editing.
[0081] Based on this understanding, it is expected that some phosphoribose modifications, such as MS and MP, within the RNA segment of an RNA:DNA duplex may interfere with or reduce the affinity of reverse transcriptase for the duplex, and thus reduce prime editing activity. Moreover, it is expected that positions and / or combinations of positions at which the phosphoribose is modified (e.g., by MS or MP) will interfere with the function of reverse transcriptase in prime editing, and thus reduce prime editing activity. The published co-crystal structure of a complex between an RNA:DNA duplex and a portion of the duplex-complexed polypeptide fragment of reverse transcriptase from xenotropic murine leukemia virus-related virus (a close relative of Moloney murine leukemia virus (MMLV) whose reverse transcriptase is typically employed for prime editing) lacks the portion of the reverse transcriptase that interacts with the 3' end of the RNA strand within the RNA:DNA duplex (Nowak et al., Nucl. Acids Res. 2013, 3874-3887), and information remains lacking in the art regarding RNA-protein contacts that may be important at the 3' end of pegRNA in prime editing.
[0082] The present disclosure is based, in part, on the surprising discovery that modified gRNAs or pegRNAs comprising one or more MP modifications at the 3' end, optionally with one or more modifications at the 5' end, can enhance Cas-mediated editing activity, particularly when the modified guide RNAs are transfected into cells at subsaturating levels. As described in more detail below, chemically synthesized single-stranded guide RNAs of various designs, which can be about 100 nt in length, and typically longer pegRNAs, were co-transfected with Cas proteins or mRNAs encoding Cas proteins into cultured human cells, and enhanced activity was observed when MS or MP modifications were added to the phosphoribose at the 3' end of the gRNA / pegRNA.
[0083] To evaluate the effect of various 3' and / or 5' end modifications, a series of synthetic gRNAs targeting the HBB gene were created by systematically incorporating MS or MP phosphoribose modifications at the 3' end of the gRNA as listed in Table 1. The 5' and 3' end modifications are indicated in the name of each synthetic gRNA. For example, HBB-101-3xMS,3xMP refers to a guide RNA for the HBB gene with three MS modifications at the 5' end and three MP modifications at the 3' end of the gRNA. The exact positions of the modifications are underlined in the sequence shown in Figure 1. The name also indicates the length of the RNA. For example, HBB-101-etc. refers to an sgRNA strand targeting the HBB gene that is composed of 101 nucleotides. Similarly, HBB-99-etc. refers to an sgRNA strand composed of 99 nucleotides. The difference between these lengths and similar lengths of sequences is the number of uridines in the short poly-uridine (poly-U) tail at the 3' end of the sgRNA, as specified by the sequences defined in Table 1. In the examples given in Table 1, the 3' poly-U tail consists of 3, 4, 5, 6 or 7 consecutive uridines (as a rule of thumb, the 3' poly-U tail on the native tracrRNA generally consists of 7 consecutive uridines). Modifications in the guide sequence, if any, are also indicated after the name of the target gene and the length of the RNA. For example, HBB-102-11MP-3xMS,3xMP refers to a guide RNA for the HBB gene that consists of 102 nucleotides, has 3 MS modifications at the 5' end and 3 MP modifications at the 3' end, and includes an MP modification at position 11 in the guide sequence. The exact positions of the modifications are indicated by underlines in the sequences shown in Table 1 (as well as in Tables 2-4, where the MP modifications in the guide sequence are noted by bold underlines).
[0084] [Table 1] JPEG2024533448000003.jpg224170
[0085] The number and type of chemical modifications at the 3' end of gRNA can substantially increase their effectiveness for DNA editing in situations where subsaturating amounts of gRNA are delivered into cells (e.g., by nucleofection). This benefit is particularly evident in methods that use gRNA co-transfected with mRNA encoding Cas protein, as opposed to co-transfected in complex with Cas protein as a ribonucleoprotein (RNP) complex. The number and type of chemical modifications incorporated into gRNA can also improve the editing efficiency of Cas RNP complexes, as illustrated by the data provided herein for transfection of cells suspended in growth medium containing serum (known to contain nucleases). See, for example, Figures 4 and 5. The experimental data described herein also indicate that certain chemical modifications and certain sequence positions in the transfected gRNA sequence can be particularly advantageous for increasing editing yields, for example, by incorporating one or more MP modifications into consecutive 3'-terminal phosphoribose on the 3' end of gRNA.
[0086] Any of the 5' and 3' end modifications described herein may be combined with modifications in the guide sequence of a gRNA that enhance target specificity (e.g., as described in U.S. Patent No. 10,767,175). For example, as shown in Table 1 and examined in Figures 2-5, an MP modification at the 3' end (e.g., an MP at the second nucleotide from the 3' end, meaning that the first internucleotide linkage from the 3' end contains a phosphonoacetate) may be combined with an MP or other modification at positions 5 or 11 (counting from the 5' end of the guide sequence within the 20 nucleotide guide sequence) in the guide sequence portion of a gRNA or pegRNA.
[0087] Chemical modifications can be incorporated during chemical synthesis of gRNA by using chemically modified phosphoramidites in amidite coupling selection cycles to produce desired sequences. After synthesis, chemically modified gRNA is used for gene editing or regulation in the same way as unmodified gRNA. A preferred embodiment is to co-transfect chemically modified synthetic gRNA with mRNA or DNA encoding Cas protein. Chemical modifications enhance the activity of gRNA in transfected cells, including when delivered by electroporation, lipofection, or exposure of living cells or tissues to nanoparticles charged with gRNA and / or mRNA encoding Cas protein.
[0088] Exemplary synthetic pegRNAs are shown in Tables 2 and 3 below. These pegRNAs were modified by systematically incorporating MS or MP phosphoribose modifications at the 3' end. The 5' and 3' end modifications are indicated in the name of each synthetic pegRNA, which also indicates the target gene. For example, "EMX1-peg-3xMS,3xMP" refers to a pegRNA for the EMX1 gene, which has three MS modifications at the 5' end and three MP modifications at the 3' end of the pegRNA. The exact positions of the modifications are underlined in the sequence shown in Table 2. Some of the pegRNA designs have a short polyuridine tract (i.e., a poly-U tail) added to the 3' end as indicated by "+3'UU", "+3'UUU", or "+3'UUUU" in the pegRNA name.
[0089] [Table 2] JPEG2024533448000005.jpg40170
[0090] [Table 3]
[0091] [Table 4]
[0092] As demonstrated by the examples below, the use of chemical modifications at the 3' end of the pegRNA substantially improves the efficacy of synthetic pegRNAs with a prime editor (compared to pegRNAs with unmodified 3' ends). In contrast to the sustained editing activity when pegRNAs and prime editors are constitutively expressed in cells transfected with DNA vectors as originally reported in the literature (see Anzalone et al. 2019), the use of synthetic pegRNAs for prime editing may be preferred when aiming to limit the duration of editing activity. From the data, the present disclosure further demonstrates that certain chemical modifications and certain sequence positions within the pegRNA sequence may be particularly advantageous in some embodiments, such as incorporating two MP modifications at consecutive 3'-terminal phosphoribose (without adding a downstream poly-U tail at the 3' end) on a pegRNA strand that terminates in a primer binding segment at the 3' end.
[0093] A. Exemplary CRISPR / Cas Systems The CRISPR / Cas system of genome modification includes a Cas protein (e.g., Cas9 nuclease), a DNA-targeting RNA (e.g., modified gRNA) that contains a guide sequence that targets the Cas protein to the target DNA, and a scaffold region (e.g., tracrRNA) that interacts with the Cas protein. In some cases, variants of Cas proteins can be used, such as Cas9 mutants that contain one or more of the following mutations: D10A, H840A, D839A, and H863A. In other examples, fragments of Cas proteins (or variants thereof) with desired properties (e.g., can generate single-stranded or double-stranded breaks and / or modulate gene expression) can be used. For example, donor repair templates may be used in some CRISPR applications, which may include nucleotide sequences that code for reporter polypeptides, such as fluorescent proteins or antibiotic resistance markers, and homology arms that are homologous to the target DNA and flank the gene modification site. Alternatively, the donor repair template can be a single-stranded oligodeoxynucleotide (ssODN). In some embodiments, the CRISPR / CAS system may include a Cas protein that can act as a prime editor (e.g., a fusion protein that includes a Cas protein that exhibits nickase activity fused to a reverse transcriptase protein or domain thereof). Prime editors may be used with a pegRNA that incorporates a reverse transcriptase template that contains one or more edits to the sequence of a target nucleic acid to modify the sequence of the target nucleic acid by a process referred to as prime editing.
[0094] 1. Cas proteins and their variants The CRISPR (clustered regularly interspaced short palindromic repeats) / Cas (CRISPR-associated proteins) nuclease system was discovered in bacteria but is used in eukaryotic cells (e.g., mammals) for genome editing / modulation of gene expression. It is based in part on the adaptive immune response of many microbes and archaea. When a virus or plasmid invades such a microbe, a segment of the invader's DNA is integrated into a CRISPR locus (or "CRISPR array") in the microbial genome. Expression of the CRISPR locus produces a non-coding CRISPR RNA (crRNA). In type II CRISPR systems, the crRNA then associates with another type of RNA called tracrRNA through a region of partial complementarity to guide the Cas (e.g., Cas9) protein to a region homologous to the crRNA in the target DNA called the "protospacer." The Cas (e.g., Cas9) protein cleaves the DNA to generate a blunt end at the double-stranded break at the site specified by a 20-nucleotide guide sequence contained within the crRNA transcript. Cas (e.g., Cas9) proteins require both crRNA and tracrRNA for site-specific DNA recognition and cleavage. The system has been engineered to combine crRNA and tracrRNA into one molecule (single-stranded guide RNA or "sgRNA") (see, e.g., Jinek et al. (2012) Science, 337:816-821; Jinek et al. (2013) eLife, 2:e00471; Segal (2013) eLife, 2:e00563). Thus, the CRISPR / Cas system can be engineered to generate double-stranded breaks at desired targets in the cell genome and take advantage of cell-intrinsic mechanisms to repair the induced breaks by homology-directed repair (HDR) or non-homologous end joining (NHEJ).
[0095] In some embodiments, the Cas protein has DNA cleavage activity. The Cas protein can direct the cleavage of one or both strands to a location in the target DNA sequence. For example, the Cas protein can be a nickase with one or more inactivating catalytic domains that cleave a single strand of the target DNA sequence (e.g., as in the case of a prime editor Cas protein).
[0096] Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas11, Cas12, Cas13, Cas14, CasΦ, CasX, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2 , Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Cpf1, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, their variants, their fragments, their mutants, and their derivatives. There are at least six types of Cas proteins (types I-VI), and at least 33 subtypes (see, e.g., Makarova et al., Nat. Rev. Microbiol., 2020, 18:2, 67-83). Type II Cas proteins include Cas1, Cas2, Csn2, and Cas9. Cas protein is known to those skilled in the art.For example, the amino acid sequence of Streptococcus pyogenes wild-type Cas9 polypeptide is shown in, for example, NBCI reference sequence number NP_269215, and the amino acid sequence of Streptococcus thermophilus wild-type Cas9 polypeptide is shown in, for example, NBCI reference sequence number WP011681470.CRISPR-related endonucleases useful in the embodiments of the present disclosure are disclosed in, for example, U.S. Patent No. 9,267,135; U.S. Patent No. 9,745,610; and U.S. Patent No. 10,266,850.
[0097] Cas proteins, such as Cas9 polypeptides, are useful in the detection and treatment of bacterial infections in, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, and the like. pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synovie synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, LactobacillusLactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsuginis, Sphaerochaeta globus globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidates Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobactershibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis The bacteria may be derived from a variety of bacterial species including Bacillus subsp. meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, Proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.
[0098] "Cas9" refers to an RNA-guided double-stranded DNA-binding nuclease or nickase protein. Wild-type Cas9 nuclease has two functional domains, e.g., RuvC and HNH, that cleave different DNA strands. Cas9 can introduce double-stranded breaks in genomic DNA (target DNA) when both functional domains are active. The Cas9 enzyme is a nuclease that is active in Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, and other organisms. In some embodiments, the catalytic domains may include one or more catalytic domains of a Cas9 protein from bacteria belonging to the group consisting of: Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, and Campylobacter. In some embodiments, the two catalytic domains are from different bacterial species.
[0099] "Cas12" (including variants Cas12a (also known as Cpf1), Cas12b, c2c1, c2c3, CasX, and CasY) refers to an RNA-guided double-stranded DNA-binding nuclease protein that contains a mixed alpha / beta domain, RuvC-I, followed by a helical region, RuvC-II, and a zinc finger-like domain or nickase protein. Wild-type Cas12 nuclease creates staggered 5' overhangs at dsDNA target sequences and does not require tracrRNA. Cas12 and its variants recognize 5' AT-rich PAM sequences on target dsDNA. An insertion domain, called Nuc, of the Cas12a protein has been demonstrated to be responsible for cleavage of the target strand. Cas12 enzymes may contain one or more catalytic domains of Cas12 proteins derived from bacteria belonging to the group consisting of Francisella and Prevotella.
[0100] Useful variants of Cas9 protein may contain a single inactive catalytic domain, such as RuvC or HNH enzymes, both of which are nickases. Such Cas proteins are useful, for example, in the context of prime editing. Cas9 nickases have only one functional domain and can only cleave one strand of target DNA, thereby generating a single-strand break or nick. In some embodiments, the Cas protein is a mutant Cas9 nuclease with at least a D10A mutation, and a Cas9 nickase. In other embodiments, the Cas protein is a mutant Cas9 nuclease with at least a H840A mutation, and a Cas9 nickase. Other examples of mutations present in Cas9 nickases include, but are not limited to, N854A and N863A. Double-strand breaks can be introduced using Cas9 nickases when using at least two DNA-targeting RNAs that target opposite DNA strands. The staggered double nick-introduced double-strand break can be repaired by NHEJ or HDR (Ran et al., 2013, Cell, 154:1380-1389; Anzalone et al. Nature 576:7785, 2019, 149-15). This gene editing strategy favors HDR and reduces the frequency of indel mutations as a by-product. Non-limiting examples of Cas9 nucleases or nickases are described, for example, in U.S. Pat. Nos. 8,895,308; 8,889,418; 8,865,406; 9,267,135; and 9,738,908; and U.S. Patent Application Publication No. 2014 / 0186919. Cas9 nucleases or nickases can be codon-optimized for the target cell or organism.
[0101] In some embodiments, the Cas protein may be a Cas9 polypeptide containing two silencing mutations (D10A and H840A) in the RuvCl and HNH nuclease domains, referred to as dCas9 (Jinek et al., Science, 2012, 337:816-821; Qi et al., Cell, 152(5):1173-1183). In one embodiment, the dCas9 polypeptide from Streptococcus pyogenes contains at least one mutation at positions D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, A987, or any combination thereof. A description of such dCas9 polypeptides and variants thereof is provided, for example, in International Patent Publication No. WO2013 / 176772. The dCas9 enzyme may contain mutations at D10, E762, H983, or D986, as well as mutations at H840 or N863. In some cases, the dCas9 enzyme contains a D10A or D10N mutation. The dCas9 enzyme may also include H840A, H840Y, or H840N. In some embodiments, the dCas9 enzyme used in the aspects of the present disclosure includes D10A and H840A; D10A and H840Y; D10A and H840N; D10N and H840A; D10N and H840Y; or D10N and H840N substitutions. The substitutions may be conservative or non-conservative substitutions to render the Cas9 polypeptide catalytically inactive and capable of binding to target DNA.
[0102] dCas9 polypeptide is catalytically inactive and lacks nuclease activity.In some cases, dCas9 enzyme or its variant or fragment can block the transcription of target sequence, and in some cases block RNA polymerase.In other examples, dCas9 enzyme or its variant or fragment can activate the transcription of target sequence when fused with target sequence, such as transcription activator polypeptide.In some embodiments, Cas protein or protein variant comprises one or more NLS sequences.
[0103] In some embodiments, the Cas protein may be a fusion protein comprising one or more Cas nuclease domains fused, optionally with an intervening linker, to one or more heterologous functional domains of a second protein, where the linker does not interfere with the activity of the fusion protein. Heterologous in this context means that the functional domain is derived from a protein other than the Cas protein. In some embodiments, the heterologous functional domain comprises an enzymatic domain and / or a binding domain. In some embodiments, the heterologous enzymatic domain is a nuclease, nickase, recombinase, deaminase, methyltransferase, polymerase, reverse transcriptase, methylase, acetylase, acetyltransferase, transcriptional activator, or transcriptional repressor domain. In some embodiments, the heterologous enzyme domain comprises base editing activity, nucleotide deaminase activity, methylase activity, demethylase activity, translation activation activity, translation repression activity, transcription activation activity, transcription repression activity, transcription release factor activity, chromatin modifying or remodeling activity, histone modifying activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, nucleic acid binding activity, detectable activity, or any combination thereof.
[0104] In some embodiments, the Cas protein comprises a heterologous functional domain that is a base editor, such as a cytidine deaminase domain from the catalytic polypeptide-like (APOBEC) family of deaminases, including, for example, apolipoprotein B mRNA editing enzyme, APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4; an activation-induced cytidine deaminase (AID), e.g., activation-induced cytidine deaminase (AICDA); cytosine deaminase 1 (CDA1) or CDA2; or a cytosine deaminase acting on tRNA (CDAT). In some embodiments, the heterologous functional domain is a deaminase that modifies adenosine DNA base, for example, the deaminase is adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA1 (ADAR1), ADAR2, ADAR3; adenosine deaminase acting on tRNA1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA). In some embodiments, the heterologous functional domain is a biological tether. In some embodiments, the biological tether is MS2, Csy4 or lambda N protein. In some embodiments, the heterologous functional domain is FokI.
[0105] In some embodiments, the Cas protein comprises a heterologous functional domain that is an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways, e.g., uracil DNA glycosylase inhibitor (UGI), which inhibits uracil DNA glycosylase (also known as UDG, uracil N-glycosylase, or UNG)-mediated cleavage of uracil to initiate BER; or a DNA end-binding protein, such as Gam from bacteriophage Mu.
[0106] In some embodiments, the Cas protein comprises a heterologous functional domain that is a transcriptional activation domain, e.g., a VP64 domain, a p65 domain, a MyoD1 domain, or a HSF1 domain. In some embodiments, the Cas protein comprises a heterologous functional domain that is a transcriptional repression domain, e.g., a Krueppel-associated box (KRAB) domain, an ERF repressor domain (ERD), an mSin3A interaction domain (SID) domain, a SID4X domain, a NuE domain, or an NcoR domain. In some embodiments, the Cas protein comprises a heterologous functional domain that is a nuclease domain, e.g., a Fok1 domain. In some embodiments, the Cas protein comprises a transcriptional silencer domain, e.g., a heterochromatin protein 1 (HP1), e.g., HP1a or HP1D. In some embodiments, the heterologous functional domain of the Cas protein is an enzyme that modifies the methylation state of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase (DNMT) or a TET protein. In some embodiments, the TET protein is TET1. In some embodiments, the heterologous functional domain of the Cas protein is an enzyme that modifies a histone subunit. In some embodiments, the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), a histone deacetylase (HDAC), a histone methyltransferase (HMT), or a histone demethylase.
[0107] For gene regulation (e.g., modulating the transcription of target DNA), nuclease-deficient Cas proteins, such as, but not limited to, dCas9, can be used for transcription activation or transcription repression. Methods for inactivating gene expression using nuclease-null Cas proteins are described, for example, in Larson et al., Nat. Protoc., 2013, 8(11):2180-2196.
[0108] In some embodiments, the Cas protein comprises one or more nuclear localization signal (NLS) domains. The one or more NLS domains may be located at or near or adjacent to the terminus of the effector protein (e.g., C2c2), and when there is more than one NLS, each of the two may be located at or near or adjacent to the terminus of the effector protein (e.g., C2c2).
[0109] In some embodiments, the nucleotide sequence encoding the Cas protein is present in a recombinant expression vector. In some cases, the recombinant expression vector is a viral construct, such as a recombinant adeno-associated viral construct, a recombinant adenoviral construct, a recombinant lentiviral construct, and the like. For example, the viral vector can be based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, and the like. The retroviral vector can be based on murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, mammary tumor virus, and the like. Useful expression vectors are known to those skilled in the art, and many are commercially available. As examples for eukaryotic host cells, the following vectors are provided: pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40. However, any other vector may be used as long as it is compatible with the host cell.
[0110] Depending on the target cell / expression system used, any of several transcriptional and translational control elements, including promoters, transcriptional enhancers, transcriptional terminators, etc., may be used in the expression vector. Useful promoters may be derived from viruses or any organism, e.g., prokaryotes or eukaryotes. Suitable promoters include, but are not limited to, SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoter, e.g., CMV immediate early promoter region (CMVIE), Rous sarcoma virus (RSV) promoter, human U6 small nuclear promoter (U6), enhanced U6 promoter, human H1 promoter (H1), etc.
[0111] The Cas protein can be introduced into a cell (e.g., a cell such as a primary cell for ex vivo therapy or a cell in vivo, such as in a patient) as a Cas polypeptide, an mRNA encoding a Cas polypeptide, or a recombinant expression vector comprising a nucleotide sequence encoding a Cas polypeptide.
[0112] 2. Chemically modified guide RNA (gRNA) Modified gRNAs for use in the CRISPR / Cas system of genome modification typically include a guide sequence that is complementary to a target nucleic acid sequence and a scaffold region that interacts with a Cas protein.
[0113] The guide sequence of the modified guide RNA can be any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence (e.g., a target DNA sequence) to hybridize with the target sequence and direct sequence-specific binding of the CRISPR complex to the target sequence. In some embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the guide sequence of the modified guide RNA and its corresponding target sequence is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, or is greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more. Optimal alignment may be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows Wheeler aligner), Clustal W, Clustal X, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 nucleotides long or longer. In some cases, the guide sequence is about 20 nucleotides long. In other examples, the guide sequence is about 15 nucleotides long. In other examples, the guide sequence is about 25 nucleotides long. The ability of the guide sequence to direct the sequence-specific binding of the CRISPR complex to the target sequence may be evaluated by any suitable assay. Binding can be determined directly or indirectly, for example, by using editing or truncation as a proxy.For example, components of a CRISPR system sufficient to form a CRISPR complex comprising a test guide sequence may be provided to a host cell having a corresponding target sequence, such as by transfection of a vector encoding the components of the CRISPR sequence followed by assessment of editing or cleavage in the target sequence. Similarly, cleavage of a target polynucleotide sequence may be assessed in vitro by providing the target sequence and components of a CRISPR complex comprising the test guide sequence and a control guide sequence different from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test guide sequence reaction and the control guide sequence reaction.
[0114] The nucleotide sequence of the guide RNA can be selected using any of the web-based software mentioned above. Considerations for selecting a DNA-targeting RNA include the PAM sequence for the Cas protein (e.g., Cas9 polypeptide) to be used, and a strategy for minimizing off-target modification. Tools such as CRISPR design tools can provide sequences for preparing modified gRNAs, for determining target modification efficiency, and / or for determining cleavage at off-target sites. Another consideration for selecting the sequence of a modified guide RNA includes reducing the degree of secondary structure within the guide sequence. The secondary structure may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. Examples of suitable algorithms include mFold (Zuker and Stiegler, Nucleic Acids Res, 9 (1981), 133-148), the UNAFold package (Markham et al., Methods Mol Biol, 2008, 453:3-31) and RNAfold from the ViennaRNA package.
[0115] One or more nucleotides of the guide sequence of the modified guide RNA and / or one or more nucleotides of the scaffold region can be modified nucleotides.For example, a guide sequence of about 20 nucleotides in length can have one or more, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more modified nucleotides.In some cases, the guide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more modified nucleotides.In other examples, the guide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20 or more modified nucleotides.Modified nucleotides can be located at any nucleic acid position of the guide sequence. In other words, the modified nucleotides can be at or near the first and / or last nucleotide of the guide sequence, and / or at any position therebetween. For example, for a guide sequence that is 20 nucleotides long, one or more modified nucleotides can be located at nucleic acid position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, position 15, position 16, position 17, position 18, position 19, and / or position 20 of the guide sequence. In some cases, about 10% to about 30%, e.g., about 10% to about 25%, about 10% to about 20%, about 10% to about 15%, about 15% to about 30%, about 20% to about 30%, or about 25% to about 30% of the guide sequence can comprise modified nucleotides. In other examples, about 10% to about 30% of the guide sequences may comprise modified nucleotides, e.g., about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, or about 30%.
[0116] In some embodiments, the scaffold region of modified guide RNA contains one or more modified nucleotides.For example, the scaffold region of about 80 nucleotides can have one or more, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 76, 77, 78, 79, 80 or more modified nucleotides.In some cases, the scaffold region comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more modified nucleotides. In other examples, the scaffold region comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20 or more modified nucleotides. The modified nucleotides may be located at any nucleic acid position in the scaffold region. For example, the modified nucleotides may be located at or near the first and / or last nucleotide of the scaffold region, and / or at any position therebetween. For example, for a scaffold region that is about 80 nucleotides in length, the one or more modified nucleotides may be at nucleic acid position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, position 15, position 16, position 17, position 18, position 19, position 20, position 21, position 22, position 23, position 24, position 25, position 26, position 27, position 28, position 29, position 30, position 31, position 32, position 33, position 34, position 35, position 36, position 37, position 38, position 39, position 40, position 41, position 42, position 43, position 44, position 45, position 46, position 47, position 48, position 49, position 50, position 51, position 52, position 53, position 54, position 55, position 56, position 57, position 58, position 59, position 60, position 61, position 62, position 63, position 64, position 65, position 66, position 67, position 68, position 69, position 70, position 71, position 72, position 73, position 74, position 75, position 76, position 77, position 78, position 79, position 80, position 81, position 82, position 83, position 84, position 85, position 86, position 87, position 88, position 89, position 90, position 91, position 92, position 93, position 94, position 95, position 96, position 97, position position 38, position 39, position 40, position 41, position 42, position 43, position 44, position 45, position 46, position 47, position 48, position 49, position 50, position 51, position 52, position 53, position 54, position 55, position 56, position 57, position 58, position 59, position 60, position 61, position 62, position 63, position 64, position 65, position 66, position 67, position 68, position 69, position 70, position 71, position 72, position 73, position 74, position 75, position 76, position 77, position 78, position 79, and / or position 80.In some cases, about 1% to about 10%, e.g., about 1% to about 8%, about 1% to about 5%, about 5% to about 10%, or about 3% to about 7% of the scaffold region may comprise modified nucleotides. In other examples, about 1% to about 10%, e.g., about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the scaffold region may comprise modified nucleotides.
[0117] Modified nucleotides of the guide RNA may include modifications at the ribose (e.g., sugar) group, the phosphate group, the nucleobase, or any combination thereof. In some embodiments, the modification at the ribose group includes a modification at the 2' position of the ribose.
[0118] In some embodiments, the modified nucleotides include 2'fluoro-arabino nucleic acid, tricycle-DNA (tc-DNA), peptide nucleic acid, cyclohexene nucleic acid (CeNA), locked nucleic acid (LNA), ethylene-bridged nucleic acid (ENA), xenonucleic acid (XNA), phosphodiamidate morpholino, or combinations thereof.
[0119] Modified nucleotides or nucleotide analogs may include sugar-modified and / or backbone-modified ribonucleotides (i.e., including modifications to the phosphate-sugar backbone). For example, the phosphodiester linkages of native or natural RNA may be modified to include at least one of nitrogen or sulfur heteroatoms. In some backbone-modified ribonucleotides, the phosphoester group connecting adjacent ribonucleotides may be replaced by a modified group, for example, a phosphorothioate group. In preferred sugar-modified ribonucleotides, the 2' portion is a group selected from H, OR, R, halo, SH, SR, NH2, NHR, NR2 or ON, where R is C1-C6 alkyl, alkenyl or alkynyl, and halo is F, Cl, Br or I.
[0120] In some embodiments, the modified nucleotide contains a sugar modification. Non-limiting examples of sugar modifications include 2'-deoxy-2'-fluoro-oligoribonucleotides (2'-fluoro-2'-deoxycytidine-5'-triphosphate, 2'-fluoro-2'-deoxyuridine-5'-triphosphate), 2'-deoxy-2'-deamine oligoribonucleotides (2'-amino-2'-deoxycytidine-5'-triphosphate, 2'-amino-2'-deoxyuridine-5'-triphosphate), 2'-O-alkyl oligoribonucleotides, 2'-deoxy- Included are 2'-C-alkyl oligoribonucleotides (2'-O-methylcytidine-5'-triphosphate, 2'-methyluridine-5'-triphosphate), 2'-C-alkyl oligoribonucleotides, and their isomers (2'-aracytidine-5'-triphosphate, 2'-arauidine-5'-triphosphate), azidotriphosphates (2'-azido-2'-deoxycytidine-5'-triphosphate, 2'-azido-2'-deoxyuridine-5'-triphosphate), and combinations thereof.
[0121] In some embodiments, the modified guide RNA contains one or more 2'-fluoro, 2'-amino and / or 2'-thio modifications. In some cases, the modifications are 2'-fluoro-cytidine, 2'-fluoro-uridine, 2'-fluoro-adenosine, 2'-fluoro-guanosine, 2'-amino-cytidine, 2'-amino-uridine, 2'-amino-adenosine, 2'-amino-guanosine, 2,6-diaminopurine, 4-thio-uridine, 5-amino-allyl-uridine, 5-bromo-uridine, 5-iodo-uridine, 5-methyl-cytidine, ribo-thymidine, 2-aminopurine, 2'-amino-butyryl-pyrene-uridine, 5-fluoro-cytidine, and / or 5-fluoro-uridine.
[0122] More than 96 naturally occurring nucleoside modifications are found in mammalian RNA.See, for example, Limbach et al., Nucleic Acids Research, 22(12):2183-2196 (1994).The preparation of nucleotides and modified nucleotides and nucleosides is well known in the art and is described, for example, in U.S. Patent Nos. 4,373,071, 4,458,066, 4,500,707, 4,668,777, 4,973,679, 5,047,524, 5,132,418, 5,153,319, 5,262,530 and 5,700,642.Many modified nucleosides and modified nucleotides suitable for use as described herein are commercially available. The nucleoside can be an analog of a naturally occurring nucleoside. In some cases, the analog is dihydrouridine, methyladenosine, methylcytidine, methyluridine, methylpseudouridine, thiouridine, deoxycytodine, and deoxyuridine.
[0123] In some cases, the modified guide RNAs described herein include nucleobase-modified ribonucleotides, i.e., ribonucleotides that contain at least one non-naturally occurring nucleobase in place of a naturally occurring nucleobase.Non-limiting examples of modified nucleobases that can be incorporated into modified nucleosides and modified nucleotides include m5C (5-methylcytidine), m5U (5-methyluridine), m6A (N6-methyladenosine), s2U (2-thiouridine), Um (2'-O-methyluridine), m1A (1-methyladenosine), m2A (2-methyladenosine), Am (2-1-O-methyladenosine), ms2m6A (2-methylthio-N6-methyladenosine), i6A (N6-isopentenyl adenosine), ms2i6A (2-methylthio-N6-isopentenyladenosine), io6A (N6-(cis-hydroxyisopentenyl)adenosine), ms2io6A (2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine), g6A (N6-glycinylcarbamoyladenosine), t6A (N6-threonylcarbamoyladenosine), ms2t6A (2-methylthio-N6-threonylcarbamoyladenosine), m6t6A (N6-methyl-N 6-threonylcarbamoyl adenosine), hn6A (N6-hydroxynorvalylcarbamoyl adenosine), ms2hn6A (2-methylthio-N6-hydroxynorvalylcarbamoyl adenosine), Ar(p) (2-O-ribosyladenosine (phosphate)), I (inosine), m1I (1-methylinosine), m'Im (1,2'-O-dimethylinosine), m3C (3-methylcytidine), Cm (2T-O-methylcytidine), s2C (2-thiocytidine), ac4C (N 4-acetylcytidine), f5C (5-formylcytidine), m5Cm (5,2-O-dimethylcytidine), ac4Cm (N4 acetyl-2O-methylcytidine), k2C (lysidine), m1G (1-methylguanosine), m2G (N2-methylguanosine), m7G (7-methylguanosine), Gm (2'-O-methylguanosine), m22G (N2,N2-dimethylguanosine), m2Gm (N2,2'-O-dimethylguanosine), m22Gm (N2,N2,2'-O-trimethylguanosine), Gr(p)(2'-O-ribosylguanosine (phosphate)), yW(wybutosine), o2yW(peroxywybutosine), OHYW(hydroxywybutosine), OHYW*(hypomodified hydroxywybutosine), imG(wybutosine), mimG(methylguanosine), Q(queuosine), oQ(epoxyqueuosine), galQ(galactosyl-queuosine), manQ(mannosyl-queuosine), preQo(7-cyano-7-deazaguanosine), preQi (7-aminomethyl-7-deazaguanosine), G (archeosine), D (dihydrouridine), m5Um (5,2'-O-dimethyluridine), s4U (4-thiouridine), m5s2U (5-methyl-2-thiouridine), s2Um (2-thio-2'-O-methyluridine), acp3U (3-(3-amino-3-carboxypropyl)uridine), ho5U (5-hydroxyuridine), mo5U (5-methoxyuridine), cmo5U (uridine 5-oxyacetic acid), mcmo5U (uridine 5-oxyacetic acid methyl ester), chm5U (5-(carboxyhydroxymethyl)uridine), mchm5U (5-(carboxyhydroxymethyl)uridine methyl ester), mcm5U (5-methoxycarbonylmethyluridine), mcm5Um (S-methoxycarbonylmethyl-2-O-methyluridine), mcm5s2U (5-methoxycarbonylmethyl-2-thiouridine), nm5s2U (5-aminomethyl-2-thiouridine), mnm5U (5-methylaminomethyluridine), mnm5s2U (5-methylaminomethyl-2-thiouridine), mnm5se2U (5 -methylaminomethyl-2-selenouridine), ncm5U (5-carbamoylmethyluridine), ncm5Um (5-carbamoylmethyl-2'-O-methyluridine), cmnm5U (5-carboxymethylaminomethyluridine), cnmm5Um (5-carboxymethylaminomethyl-2-LO-methyluridine), cmnm5s2U (5-carboxymethylaminomethyl-2-thiouridine), m62A (N6,N6-dimethyladenosine), Tm (2'-O-methylinosine), m4C (N4-methylcytidine), m4Cm (N4,2-O-dimethylcytidine), hm5C (5-hydroxymethylcytidine), m3U (3-methyluridine), cm5U (5-carboxymethyluridine), m6Am (N6,TO-dimethyladenosine), m62Am (N6,N6,O-2-trimethyladenosine), m2'7G (N2,7-dimethylguanosine), m2'2'7G (N2,N2,7-trimethylguanosine), m 3Um (3,2T-O-dimethyluridine), m5D (5-methyldihydrouridine), f5Cm (5-formyl-2'-O-methylcytidine), m1Gm (1,2'-O-dimethylguanosine), m'Am (1,2-O-dimethyladenosine), tm5s2U (S-taurinomethyl-2-thiouridine), imG-14 (4-demethylguanosine), im G2 (isoguanosine), or ac6A (N6-acetyladenosine), hypoxanthine, inosine, 8-oxo-adenine, its 7-substituted derivatives, dihydrouracil, pseudouracil, 2-thiouracil, 4-thiouracil, 5-aminouracil, 5-(C1-C6)-alkyluracil, 5-methyluracil, 5-(C2-C6)-alkenyluracil, 5-(C2-C6) -Alkynyluracil, 5-(hydroxymethyl)uracil, 5-chlorouracil, 5-fluorouracil, 5-bromouracil, 5-hydroxycytosine, 5-(C1-C6)-alkylcytosine, 5-methylcytosine, 5-(C2-C6)-alkenylcytosine, 5-(C2-C6)-alkynylcytosine, 5-chlorocytosine, 5-fluorocytosine, 5-bromocytosine, N, 2 -dimethylguanine, 7-deazaguanine, 8-azaguanine, 7-deaza-7-substituted guanine, 7-deaza-7-(C2-C6)alkynylguanine, 7-deaza-8-substituted guanine, 8-hydroxyguanine, 6-thioguanine, 8-oxoguanine, 2-aminopurine, 2-amino-6-chloropurine, 2,4-diaminopurine, 2,6-diaminopurine, 8-azapurine, substituted 7-deazapurines, 7-deaza-7-substituted purines, 7-deaza-8-substituted purines, and combinations thereof.
[0124] In some embodiments, the phosphate backbone of the guide RNA is altered. Modified gRNAs may include one or more phosphorothioate, phosphoramidate (e.g., N3'-P5'-phosphoramidate (NP)), 2'-O-methoxy-ethyl (2'MOE), 2'-O-methyl-ethyl (2'ME), and / or methylphosphonate linkages.
[0125] In certain embodiments, one or more of the modified nucleotides of the guide sequence of the guide RNA and / or one or more of the modified nucleotides of the scaffold region of the guide RNA include 2'-O-methyl (M) nucleotides, 2'-O-methyl-3'-phosphorothioate (MS) nucleotides, 2'-O-methyl-3'-phosphonoacetate (MP) nucleotides, 2'-O-methyl-3'thioPACE (MSP) nucleotides, or combinations thereof. In some cases, the guide RNA includes one or more MS nucleotides. In other examples, the guide RNA includes one or more MP / MSP nucleotides. In yet other examples, the guide RNA includes one or more MS nucleotides and one or more MP / MSP nucleotides. In further examples, the guide RNA does not include an M nucleotide. In some cases, the guide RNA includes one or more MS nucleotides and / or one or more MP / MSP nucleotides, and further includes one or more M nucleotides. In certain other examples, the MS nucleotides and / or MP / MSP nucleotides are the only modified nucleotides present in the guide RNA.
[0126] In some embodiments, the modified guide RNAs described herein and the Cas proteins (or the mRNAs encoding same) may be present in a composition (e.g., a CRISPR / Cas reaction mixture) in specific amounts, ratios, or ranges. For example, a reaction mixture may include: a) 1-200 pmol guide RNA; b) 1-100 pmol Cas protein, or 0.01-3.0 pmol DNA or mRNA encoding a Cas protein; c) a molar ratio of 0.1:1 to 3:1 guide RNA to Cas protein; and / or d) a molar ratio of 1:1 to 200:1 guide RNA to DNA or mRNA encoding a Cas protein. For example, in some embodiments, a reaction mixture comprises a plurality of cells; i) 1-100 pmol guide RNA (or pegRNA) per 100,000 cells, and / or ii) 1-50 pmol Cas protein or 0.01-3.0 pmol DNA or mRNA encoding a Cas protein per 100,000 cells. Similarly, in some embodiments, a reaction mixture comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 pmol of DNA or mRNA encoding a Cas protein per pmol of DNA or mRNA encoding a Cas protein. It may comprise 0, 180, 190, or 200 pmol, or up to 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 pmol of guide RNA, or an amount within a range bounded by any combination of said values.In some embodiments, the molar ratio of guide RNA to DNA or mRNA encoding a Cas protein is at least about 200:1, 190:1, 180:1, 170:1, 160:1, 150:1, 140:1, 130:1, 120:1, 110:1, 100:1, 90:1, 80:1, 70:1, 60:1, 50:1, 40:1, 30:1, 20:1, or 10:1. 0:1, 110:1, 100:1, 90:1, 80:1, 70:1, 60:1, 50:1, 40:1, 30:1, 20:1, or 10:1, or in a range bounded by up to 200:1, 190:1, 180:1, 170:1, 160:1, 150:1, 140:1, 130:1, 120:1, 110:1, 100:1, 90:1, 80:1, 70:1, 60:1, 50:1, 40:1, 30:1, 20:1, or 10:1, or any combination of said ratios. In some embodiments, a reaction mixture according to the present disclosure comprises at least about 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, or 3.0 pmol of Cas protein per pmol of Cas protein. , 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9 or 3.0 pmol, or up to 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9 or 3.0 pmol of guide RNA, or an amount within a range bounded by any combination of said values.
[0127] It should be noted that any of the modifications described herein can be combined and incorporated within the guide sequence and / or scaffold region of a modified gRNA.
[0128] In some cases, the guide RNA also comprises a structural modification such as a stem loop, e.g., an MS2 stem loop or a tetraloop.
[0129] Guide RNA can be synthesized by any method known to those skilled in the art. Modified gRNA can be synthesized using 2'-O-thionocarbamate protected nucleoside phosphoramidite. Methods are described, for example, in Dellinger et al., J.American Chemical Society 133, 11540-11556 (2011); Threlfall et al., Organic & Biomolecular Chemistry 10, 746-754 (2012); and Dellinger et al., J.American Chemical Society 125, 940-950 (2003).
[0130] Chemically modified gRNA or pegRNA can be used with any CRISPR-related technology, such as RNA guide technology. As described herein, guide RNA can serve as a guide for any Cas protein or variant or fragment thereof, including any engineered or artificial Cas9 polypeptide. Modified gRNA or pegRNA can target DNA and / or RNA molecules in isolated primary cells for ex vivo therapy or in vivo (e.g., in animals). The methods disclosed herein can be applied to genome editing, gene regulation, imaging, and any other CRISPR-based applications.
[0131] 3. Donor Repair Template In some embodiments, the present disclosure provides a recombinant donor repair template that includes two homology arms that are homologous to a portion of a target DNA sequence (e.g., a target gene or locus) on either side of a Cas protein (e.g., Cas9 nuclease) cleavage site. In some cases, the recombinant donor repair template includes a reporter cassette that includes a nucleotide sequence that encodes a reporter polypeptide (e.g., a detectable polypeptide, a fluorescent polypeptide, or a selection marker), and two homology arms that flank the reporter cassette and are homologous to a portion of the target DNA on either side of the Cas protein cleavage site. The reporter cassette may further include a sequence that encodes a self-cleaving peptide, one or more nuclear localization signals, and / or a fluorescent polypeptide, such as superfolder GFP (sfGFP).
[0132] In some embodiments, the homology arms are the same length. In other embodiments, the homology arms are different lengths. The homology arms are at least about 10 base pairs (bp), e.g., at least about 10bp, 15bp, 20bp, 25bp, 30bp, 35bp, 45bp, 55bp, 65bp, 75bp, 85bp, 95bp, 100bp, 150bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 10 In some embodiments, the nucleic acid sequence may be 00 bp, 1.1 kilobases (kb), 1.2 kb, 1.3 kb, 1.4 kb, 1.5 kb, 1.6 kb, 1.7 kb, 1.8 kb, 1.9 kb, 2.0 kb, 2.1 kb, 2.2 kb, 2.3 kb, 2.4 kb, 2.5 kb, 2.6 kb, 2.7 kb, 2.8 kb, 2.9 kb, 3.0 kb, 3.1 kb, 3.2 kb, 3.3 kb, 3.4 kb, 3.5 kb, 3.6 kb, 3.7 kb, 3.8 kb, 3.9 kb, 4.0 kb, or longer. The homology arm may be from about 10 bp to about 4 kb, for example, from about 10 bp to about 20 bp, from about 10 bp to about 50 bp, from about 10 bp to about 100 bp, from about 10 bp to about 200 bp, from about 10 bp to about 500 bp, from about 10 bp to about 1 kb, from about 10 bp to about 2 kb, from about 10 bp to about 4 kb, from about 100 bp to about 200 bp, from about 100 bp to about 500 bp, from about 100 bp to about 1 kb, from about 100 bp to about 2 kb, from about 100 bp to about 4 kb, from about 500 bp to about 1 kb, from about 500 bp to about 2 kb, from about 500 bp to about 4 kb, from about 1 kb to about 2 kb, from about 1 kb to about 2 kb, from about 1 kb to about 4 kb, or from about 2 kb to about 4 kb.
[0133] The donor repair template can be cloned into an expression vector. Conventional viral and non-viral based expression vectors known to those of skill in the art can be used.
[0134] Instead of recombinant donor repair template, single-stranded oligodeoxynucleotide (ssODN) donor template can be used for homologous recombination mediated repair. ssODN is useful for introducing short modifications into target DNA. For example, ssODN is suitable for precisely correcting genetic mutations such as SNPs. ssODN may contain two adjacent homologous sequences on each side of the target site of Cas protein cleavage, and may be oriented in sense or antisense direction relative to target DNA. Each flanking sequence can be at least about 10 base pairs (bp), e.g., at least about 10bp, 15bp, 20bp, 25bp, 30bp, 35bp, 40bp, 45bp, 50bp, 55bp, 60bp, 65bp, 70bp, 75bp, 80bp, 85bp, 90bp, 95bp, 100bp, 150bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1kb, 2kb, 4kb, or more. In some embodiments, each homology arm is from about 10 bp to about 4 kb, e.g., from about 10 bp to about 20 bp, from about 10 bp to about 50 bp, from about 10 bp to about 100 bp, from about 10 bp to about 200 bp, from about 10 bp to about 500 bp, from about 10 bp to about 1 kb, from about 10 bp to about 2 kb, from about 10 bp to about 4 kb, from about 100 bp to about 200 bp, from about 100 bp to about 500 bp, from about 100 bp to about 1 kb, from about 100 bp to about 2 kb, from about 100 bp to about 4 kb, from about 500 bp to about 1 kb, from about 500 bp to about 2 kb, from about 500 bp to about 4 kb, from about 1 kb to about 2 kb, from about 1 kb to about 2 kb, from about 1 kb to about 4 kb, or from about 2 kb to about 4 kb. The ssODN can be at least about 25 nucleotides (nt) in length, e.g., at least about 25nt, 30nt, 35nt, 40nt, 45nt, 50nt, 55nt, 60nt, 65nt, 70nt, 75nt, 80nt, 85nt, 90nt, 95nt, 100nt, 150nt, 200nt, 250nt, 300nt, or longer.In some embodiments, the ssODN is about 25 to about 50; about 50 to about 100; about 100 to about 150; about 150 to about 200; about 200 to about 250; about 250 to about 300; or about 25 nt to about 300 nt in length.
[0135] In some embodiments, the ssODN template comprises at least one, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or more modified nucleotides as described herein. In some cases, at least 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 99% of the sequence of the ssODN comprises modified nucleotides. In some embodiments, the modified nucleotides are located at one or both ends of the ssODN. The modified nucleotides can be the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th nucleotides from the end, or any combination thereof. For example, the modified nucleotides can be at the three terminal nucleotides at both ends of the ssODN template. Additionally, the modified nucleotides may be located internally relative to the termini.
[0136] In some embodiments, such as prime editing, no exogenous DNA repair template is required. For example, the modified pegRNA described herein comprises a reverse transcriptase sequence (e.g., at the 3' end adjacent to the primer binding site sequence) that contains one or more edits to the target nucleic acid, which is used as a template by the prime editor Cas protein when performing prime editing of the target nucleic acid.
[0137] 4.Target DNA In the CRISPR / Cas system, the target DNA sequence may be immediately followed by a protospacer adjacent motif (PAM) sequence. The target DNA site may be located directly 5' of the PAM sequence specific to the bacterial species of the Cas protein Cas9 used. For example, the PAM sequence of Cas9 from Streptococcus pyogenes is NGG; the PAM sequence of Cas9 from Neisseria meningitidis is NNNNGATT; the PAM sequence of Cas9 from Streptococcus thermophilus is NNAGAA; the PAM sequence of Cas9 from Treponema denticola is NAAAAC. In some embodiments, the PAM sequence may be 5'-NGG (N is any nucleotide); 5'-NRG (N is any nucleotide and R is a purine); or 5'-NNGRR (N is any nucleotide and R is a purine). For the S. pyogenes system, the selected target DNA sequence should immediately precede (e.g., located 5') a 5' NGG PAM (where N is any nucleotide) such that the guide sequence of the DNA-targeting RNA (e.g., modified gRNA) base-pairs with the opposite strand to mediate cleavage approximately 3 base pairs upstream of the PAM sequence.
[0138] In some embodiments, the degree of complementarity between a guide sequence of a DNA-targeting RNA (e.g., a guide RNA) and its corresponding target DNA sequence, when optimally aligned using a suitable alignment algorithm, is about or greater than about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more. Optimal alignment may be determined using any algorithm suitable for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, Selangor, Malaysia), and ELAND (Illumina, San Diego, Calif.).
[0139] Target DNA sites can be selected from within predefined genomic sequences (genes) using web-based software such as ZiFiT Targeter software (Sander et al., 2007, Nucleic Acids Res, 35:599-605; Sander et al., 2010, Nucleic Acids Res, 38:462-468), E-CRISP (Heigwer et al., 2014, Nat Methods, 11:122-123), RGEN Tools (Bae et al., 2014, Bioinformatics, 30(10):1473-1475), CasFinder (Aach et al., 2014, bioRxiv), DNA2.0 gNRA Design Tool (DNA2.0, Menlo Park, Calif.), and CRISPick Design Tool (Broad Institute, Cambridge, Mass.). Such tools analyze genomic sequences (e.g., genes or loci of interest) and identify target sites suitable for gene editing. To determine off-target gene modifications for each DNA-targeting RNA (e.g., modified gRNA), computational prediction of off-target sites is performed based on quantitative specificity analysis of the identity, location and distribution of base-pairing mismatches.
[0140] 5. Modulate gene expression CRISPR / Cas system may be used to regulate gene expression, such as inhibiting gene expression or activating gene expression. As a non-limiting example, a complex comprising a Cas9 variant or fragment and a gRNA capable of binding to a target DNA sequence can block or hinder the transcription initiation and / or elongation by RNA polymerase. This in turn can inhibit or suppress gene expression of the target DNA. Alternatively, a complex comprising a different Cas9 variant or fragment and a gRNA capable of binding to a target DNA sequence can induce or activate gene expression of the target DNA.
[0141] Detailed descriptions of methods for performing CRISPR interference (CRISPRi) to inactivate or reduce gene expression can be found, for example, in Larson et al., Nature Protocols, 2013, 8(11):2180-2196, and Qi et al., Cell, 152, 2013, 1173-1183. In CRISPRi, the gRNA-Cas9 variant complex can bind to the non-template strand of a protein-coding region and block transcription elongation. In some cases, when the gRNA-Cas9 variant complex binds to the promoter region of a gene, the complex prevents or interferes with transcription initiation.
[0142] Detailed descriptions of methods for performing CRISPR activation to increase gene expression can be found, for example, in Cheng et al., Cell Research, 2013, 23:1163-1171, Konerman et al., Nature, 2015, 517:583-588, and U.S. Patent No. 8,697,359.
[0143] For CRISPR-based control of gene expression, catalytically inactive variants of Cas proteins (e.g., Cas9 polypeptides) that lack endonucleolytic activity can be used. In some embodiments, the Cas protein is a Cas9 variant that contains at least two point mutations in the RuvC-like and HNH nuclease domains. In some embodiments, the Cas9 variant has D10A and H840A amino acid substitutions, referred to as dCas9 (Jinek et al., Science, 2012, 337:816-821; Qi et al., Cell, 152(5):1173-1183). In some cases, the dCas9 polypeptide from Streptococcus pyogenes contains at least one mutation at positions D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, A987, or any combination thereof. Descriptions of such dCas9 polypeptides and their variants are provided, for example, in International Patent Application Publication No. WO2013 / 176772. dCas9 enzymes can contain mutations at D10, E762, H983 or D986, as well as mutations at H840 or N863. In some cases, dCas9 enzymes contain D10A or D10N mutations. Also, dCas9 enzymes can include H840A, H840Y, or H840N. In some cases, dCas9 enzymes include D10A and H840A; D10A and H840Y; D10A and H840N; D10N and H840A; D10N and H840Y; or D10N and H840N substitutions. Substitutions can be conservative or non-conservative substitutions to render Cas9 polypeptide catalytically inactive and capable of binding to target DNA.
[0144] In certain embodiments, dCas9 polypeptide is catalytically inactive, e.g., lacks nuclease activity.In some cases, dCas9 enzyme or its variant or fragment can block the transcription of target sequence, and in some cases block RNA polymerase.In other examples, dCas9 enzyme or its variant or fragment can activate the transcription of target sequence.
[0145] In certain embodiments, a Cas9 variant lacking endonucleolytic activity (e.g., dCas9) can be fused to a transcriptional repression domain, e.g., a Kruppel-associated box (KRAB) domain, or a transcriptional activation domain, e.g., a VP16 transactivation domain. In some embodiments, the Cas9 variant is a fusion polypeptide comprising dCas9 and a transcription factor, e.g., RNA polymerase omega factor, heat shock factor 1, or a fragment thereof. In other embodiments, the Cas9 variant is a fusion polypeptide comprising dCas9 and a DNA methylase, histone acetylase, or a fragment thereof.
[0146] For CRISPR-based control of gene expression mediated by RNA binding and / or RNA cleavage, suitable Cas protein (e.g., Cas9 polypeptide) variants with endoribonuclease activity can be used, for example, as described in O'Connell et al., Nature, 2014, 516:263-266. Other useful Cas protein (e.g., Cas9) variants are described, for example, in U.S. Pat. No. 9,745,610. Other CRISPR-associated enzymes that can cleave RNA include Csy4 endoribonuclease, CRISPR-associated Cas6 enzymes, Cas5 family member enzymes, Cas6 family member enzymes, type I CRISPR system endoribonucleases, type II CRISPR system endoribonucleases, type III CRISPR system endoribonucleases, and variants thereof.
[0147] In some embodiments of CRISPR-based RNA cleavage, a DNA oligonucleotide containing a PAM sequence (e.g., a PAMmer) is used with the modified gRNA and Cas protein (e.g., Cas9) variants described herein to bind and cleave single-stranded RNA transcripts. A detailed description of suitable PAMmer sequences can be found, for example, in O'Connell et al., Nature, 2014, 516:263-266.
[0148] In some embodiments, multiple modified gRNAs and / or pegRNAs are used to target different regions of a target gene to regulate gene expression of the target gene. Multiple modified gRNAs and / or pegRNAs can provide synergistic modulation (e.g., inhibition or activation) of gene expression of a single target gene compared to each modified gRNA alone. In other embodiments, multiple modified gRNAs / pegRNAs are used to regulate gene expression of at least two different target genes.
[0149] B. Ex vivo cells In some aspects of the method, the target sequence is in a cell.The method can be used to edit, modulate, cleave, nick, or bind to the target sequence in the nucleic acid in any cell of interest, including primary cells, immortalized cells, cells from cell lines, cells from cell cultures, and others.In some embodiments, the cell is a cell type that has one or more severe conditions, for example, a cell that has high nuclease (e.g., ribonuclease, exonuclease, exoribonuclease) expression, concentration, and / or activity, for example, a cell type that has high specific nuclease.
[0150] The compositions and methods disclosed herein can be used to edit or regulate the expression of target nucleic acid in primary cells of interest. Primary cells can be cells isolated from any multicellular organism, such as plant cells (e.g., rice cells, wheat cells, tomato cells, Arabidopsis thaliana cells, Zea mays cells, etc.), cells from multicellular protists, cells from multicellular fungi, cells from invertebrates (e.g., Drosophila, cnidarians, echinoderms, nematodes, etc.) or cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals, etc.), cells from humans, cells from healthy humans, cells from human patients, cells from cancer patients, etc. In some cases, primary cells with genome editing or induced gene regulation can be transplanted into a subject (e.g., a patient). For example, primary cells can be derived from the subject (e.g., a patient) to be treated.
[0151] Any type of primary cell can be a cell of interest, for example, stem cells, for example, embryonic stem cells, induced pluripotent stem cells, adult stem cells (e.g., mesenchymal stem cells, neural stem cells, hematopoietic stem cells, organ stem cells), progenitor cells, somatic cells (e.g., fibroblasts, hepatocytes, cardiac cells, liver cells, pancreatic cells, muscle cells, skin cells, blood cells, nerve cells, immune cells), and any other cells of the body, for example, the human body. Primary cells are typically derived from a subject, for example, an animal subject or a human subject, and grown in vitro for a limited number of passages. In some embodiments, the cell is a diseased cell or is derived from a subject with a disease. For example, the cell can be a cancer cell or a tumor cell.
[0152] Primary cells can be collected from a subject by any standard method. For example, cells from tissues such as skin, muscle, bone marrow, spleen, liver, kidney, pancreas, lung, intestine, stomach, etc. can be collected by tissue biopsy or fine needle aspiration. Blood cells and / or immune cells can be isolated from whole blood, plasma or serum. In some cases, suitable primary cells include peripheral blood mononuclear cells (PBMCs), peripheral blood lymphocytes (PBLs), and other blood cell subsets, such as but not limited to T cells, natural killer cells, monocytes, natural killer T cells, monocyte progenitor cells, hematopoietic stem and progenitor cells (HSPCs), such as CD34+ HSPCs, or non-pluripotent stem cells. In some cases, the cells can be any immune cell, including but not limited to any T cell, such as tumor infiltrating cells (TILs), CD3+ T cells, CD4+ T cells, CD8+ T cells, or any other type of T cell. T cells can also include memory T cells, memory stem T cells, or effector T cells. T cells can also be biased to a particular population and phenotype. For example, T cells can be biased to a phenotype, including CD45RO(-), CCR7(+), CD45RA(+), CD62L(+), CD27(+), CD28(+) and / or IL-7Ra(+). Appropriate cells can be selected that include one or more markers selected from the list including CD45RO(-), CCR7(+), CD45RA(+), CD62L(+), CD27(+), CD28(+) and / or IL-7Ra(+). Induced pluripotent stem cells may be generated from differentiated cells by standard protocols described, for example, in U.S. Patent Nos. 7,682,828, 8,058,065, 8,530,238, 8,871,504, 8,900,871 and 8,791,248.
[0153] C. Ex Vivo Therapy The methods described herein can be used for ex vivo therapy. Ex vivo therapy can include administering a composition (e.g., cells) that is generated or modified outside of an organism to a subject (e.g., a patient). In some embodiments, a composition (e.g., including cells) can be generated or modified by the methods disclosed herein. For example, ex vivo therapy can include administering a primary cell that is generated or modified outside of an organism to a subject (e.g., a patient), where the primary cell has been cultured in vitro and edited / modulated by the methods disclosed herein, including contacting a target nucleic acid in the primary cell with one or more modified gRNAs described herein, and a Cas protein (e.g., a Cas9 polypeptide) or a variant or fragment thereof, an mRNA encoding a Cas protein (e.g., a Cas9 polypeptide) or a variant or fragment thereof, or a recombinant expression vector comprising a nucleotide sequence encoding a Cas protein (e.g., a Cas9 polypeptide) or a variant or fragment thereof.
[0154] In some embodiments, the composition (e.g., cells) may be derived from a subject (e.g., a patient) to be treated by ex vivo therapy. In some embodiments, the ex vivo therapy may include a cell-based therapy, such as adoptive immunotherapy.
[0155] In some embodiments, the composition used for ex vivo therapy can be a cell. The cell can be a primary cell, including but not limited to peripheral blood mononuclear cells (PBMC), peripheral blood lymphocytes (PBL), and other blood cell subsets. The primary cell can be an immune cell. The primary cell can be a T cell (e.g., CD3+ T cell, CD4+ T cell, and / or CD8+ T cell), a natural killer cell, a monocyte, a natural killer T cell, a monocyte progenitor cell, a hematopoietic stem cell or a non-pluripotent stem cell, a stem cell, or a progenitor cell. The primary cell can be a hematopoietic stem or progenitor cell (HSPC), such as a CD34+ HSPC. The primary cell can be a human cell. The primary cell can be isolated, selected, and / or cultured. The primary cell can be expanded ex vivo. The primary cell can be expanded in vivo. The primary cells can be CD45RO(-), CCR7(+), CD45RA(+), CD62L(+), CD27(+), CD28(+), and / or IL-7Ra(+). The primary cells can be autologous to the subject receiving the cells. Or the primary cells can be non-autologous to the subject. The primary cells can be a reagent that meets Good Manufacturing Practice (GMP) standards. The primary cells can be part of a combination therapy for treating diseases, including cancer, infectious diseases, autoimmune disorders, or graft-versus-host disease (GVHD), in subjects with or at risk for such diseases.
[0156] As a non-limiting example of ex vivo therapy, a primary cell can be isolated from a multicellular organism (e.g., a plant, a multicellular protist, a multicellular fungus, an invertebrate, a vertebrate, such as a human, etc.) before contacting the target nucleic acid in the primary cell with the Cas protein and the modified gRNA. After contacting the target nucleic acid with the Cas protein and the guide RNA, the primary cell or its progeny (e.g., a cell derived from the primary cell) can be returned to the multicellular organism.
[0157] In some embodiments, the Cas protein and guide RNA are introduced into a living organism, such as by introduction into a serum-containing fluid (e.g., whole blood, plasma or serum) in or derived from the living organism.
[0158] D. Methods for Introducing Nucleic Acids and / or Polypeptides into Target Cells Methods for introducing polypeptides and nucleic acids into target cells (host cells) are known in the art and can be employed in the present methods to introduce nucleic acids (e.g., nucleotide sequences encoding Cas proteins, modified guide RNAs, donor repair templates for homology directed repair (HDR), etc.), polypeptides (Cas proteins, polymerases, deaminases, etc.), or RNPs (e.g., gRNA / Cas protein complexes) into cells, e.g., primary cells such as stem cells, progenitor cells, or differentiated cells. Non-limiting examples of suitable methods include electroporation, viral or bacteriophage infection, transfection, microinjection, conjugation, protoplast fusion, lipofection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated delivery (e.g., lipid nanoparticle-mediated delivery, polymeric nanoparticle-mediated delivery, hybrid lipid-polymer nanoparticle-mediated delivery), and others.
[0159] In some embodiments, the components of the CRISPR system can be introduced into cells using a delivery system. In some cases, the delivery system includes nanoparticles, microparticles (e.g., polymeric micropolymers), liposomes, micelles, virosomes, virus particles, virus-like particles (VLPs), nucleic acid complexes, transfection agents, electroporation agents (e.g., using the NEON transfection system), nucleofection agents, lipofection agents, and / or buffer systems that contain the components to be delivered. For example, the components can be mixed with lipofection agents to be encapsulated or packaged in cationic submicron oil-in-water emulsions. Alternatively, the components can be delivered without a delivery system, for example as an aqueous solution.
[0160] Methods for preparing liposomes and encapsulating polypeptides and nucleic acids in liposomes are described, for example, in Methods and Protocols, Vol. 1: Pharmaceutical Nanocarriers: Methods and Protocols, (ed. Weissig). Humana Press, 2009, and Heyes et al. (2005) J Controlled Release 107:276-87. Methods for preparing microparticles and encapsulating polypeptides and nucleic acids are described, for example, in Functional Polymer Colloids and Microparticles Vol. 4 (Microspheres, microcapsules & liposomes), (eds. Arshady and Guyot). Citus Books, 2002, and Microparticulate Systems for the Delivery of Proteins and Vaccines, (eds. Cohen and Bernstein). CRC Press, 1996. For a review of the preparation of nanoparticles, such as lipid, polymer or hybrid lipid-polymer nanoparticles, see Advanced Drug Delivery Reviews 2021, Vol. 168.
[0161] E. Methods for determining the efficiency of genome editing To functionally test the presence of correct genome editing modification, target DNA can be analyzed by standard methods known to those skilled in the art.For example, indel mutations can be identified by sequencing using SURVEYOR® Mutation Detection Kit (Integrated DNA Technologies, Coralville, Iowa) or Guide-it™ Indel Identification Kit (Clontech, Mountain View, Calif.). Homologous recombination repair (HDR), base editing or prime editing mediated editing can be detected by PCR-based methods and in combination with sequencing or RFLP analysis.Non-limiting examples of PCR-based kits include Guide-it Mutation Detection Kit (Clontech) and GeneArt® Genome Break Detection Kit (Life Technologies, Carlsbad, Calif.).Deep sequencing can also be used, especially for a large number of samples or potential target / off-target sites.
[0162] In certain embodiments, the efficiency (e.g., specificity) of genome editing corresponds to the number or percentage of on-target genome editing events relative to the number or percentage of all genome editing events, including on-target and off-target events. In some embodiments, the editing efficiency of a target region corresponds to the expected number of edits of that target region at the level of a single cell or cell population.
[0163] In some embodiments, the modified gRNAs described herein can enhance genome editing of a target DNA sequence in a cell, such as a primary cell, compared to the corresponding unmodified gRNA. Genome editing can include homology-directed repair (HDR) (e.g., insertion, deletion, or point mutation), prime editing, base editing, or non-homologous end joining (NHEJ).
[0164] In certain embodiments, the efficiency of nuclease-mediated genome editing of a target DNA sequence in a cell is increased by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, 1x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 2x, 2.5x, 3x, 3.5x, 4x, 4.5x, 5x, 5.5x, 6x, 6.5x, 7x, 7.5x, 8x, 8.5x, 9x, 9.5x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, or more in the presence of a guide RNA described herein compared to a corresponding unmodified gRNA sequence. In some other embodiments, the efficiency is compared to the corresponding gRNA with different modifications to achieve the above-mentioned enhanced level. For example, a gRNA with 1x, 2x, or 3x MS at the 5' end, as well as 2x, 3x, or 4x MP or MSP at the 3' end may be compared to a gRNA with the same number of MS instead of MP / MSP (i.e., 1x, 2x, or 3x MS at the 5' end, as well as 2x, 3x, or 4x MS at the 3' end).
[0165] F. Methods of Preventing or Treating a Genetic Disease in a Subject Modified gRNAs can be applied to targeted nuclease-based therapy of genetic diseases. Current approaches to precisely correct genetic mutations in the genome of primary patient cells can be very efficient (sometimes precisely editing less than 1% of cells). Modified gRNAs described herein can enhance the activity of genome editing and increase the efficacy of genome editing-based therapy. In certain embodiments, modified gRNAs may be used for in vivo gene editing of genes in subjects with genetic diseases. Modified gRNAs can be administered to subjects via any suitable administration route in a dose or amount sufficient to enhance the effect of nuclease-based therapy (e.g., improve genome editing efficiency).
[0166] Provided herein is a method for preventing or treating a genetic disease in a subject in need thereof by correcting a genetic mutation associated with the disease. The method comprises administering to the subject a modified guide RNA as described herein in an amount sufficient to correct the mutation. Also provided herein is the use of a modified guide RNA as described herein in the manufacture of a medicament for preventing or treating a genetic disease in a subject in need thereof by correcting a genetic mutation associated with the disease. The modified guide RNA may be contained in a composition that also includes a Cas protein (e.g., a Cas polypeptide), an mRNA encoding a Cas protein, or a recombinant expression vector comprising a nucleotide sequence encoding a Cas protein. In some cases, the modified guide RNA is included in the delivery system described above.
[0167] Genetic diseases that may be corrected by the methods include, but are not limited to, X-linked severe combined immunodeficiency, sickle cell anemia, thalassemia, hemophilia, neoplasia, cancer, age-related macular degeneration, schizophrenia, trinucleotide repeat disorders, fragile X syndrome, prion-related disorders, amyotrophic lateral sclerosis, drug addiction, autism, Alzheimer's disease, Parkinson's disease, cystic fibrosis, blood and coagulation diseases or disorders, inflammation, immune-related diseases or disorders, metabolic diseases, liver diseases and disorders, kidney diseases and disorders, musculoskeletal diseases and disorders (e.g., muscular dystrophies, Duchenne muscular dystrophy), neurological and neuronal diseases and disorders, cardiovascular diseases and disorders, pulmonary diseases and disorders, ophthalmic diseases and disorders, viral infections (e.g., HIV infection), and the like. EXAMPLES
[0168] Aspects of the present teachings can be further understood in light of the following examples, which should not be construed as limiting the scope of the present teachings in any way.
[0169] Although it is understood that variations and substitutions in preparation, testing and other details may be employed in accordance with the teachings herein, various general methods and reagents were used in the following examples and are described below to facilitate an understanding of the examples.
[0170] Preparation of gRNA and mRNA. RNA oligomers were synthesized on controlled pore glass (LGC) using 2'-O-thionocarbamate protected nucleoside phosphoramidites (Sigma-Aldrich and Hongene) using Dr.Oligo 48 and 96 synthesizers (Biolytic Lab Performance Inc.) as previously described. 2'-O-methyl-3'-O-(diisopropylamino)-phosphinoacetic acid-1,1-dimethylcyanoethyl ester-5'-O-dimethoxytrityl nucleosides used to synthesize MP-modified RNA were purchased from Glen Research and Hongene. For phosphorothioate-containing oligomers, the iodine oxidation step after the coupling reaction was replaced by a 6 min sulfurization step using a 0.05 M solution of 3-((N,N-dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-5-thione in a pyridine-acetonitrile (3:2) mixture. Unless otherwise stated, reagents for solid-phase RNA synthesis were purchased from Glen Research and Honeywell. Phosphonoacetate modifications incorporated into MP-modified gRNAs were synthesized by using commercially available protected nucleoside phosphinoamidite monomers as described above using protocols adapted from previous publications (e.g., Dellinger et al., 2003 and Threlfall et al., 2012, see above). All oligonucleotides were purified using reversed-phase high-performance liquid chromatography (RP-HPLC) and analyzed by liquid chromatography-mass spectrometry (LC-MS) using an Agilent 1290 Infinity Series LC system coupled with an Agilent 6545 Q-TOF (time-of-flight) mass spectrometer. In all cases, the mass determined by deconvolution of a series of peaks containing multiple charge states in the mass spectrum of purified gRNA matched the expected mass within the error of the calibrated instrument (the standard for quality assurance used in this assay is that the observed mass of purified gRNA is within 0.01% of the calculated mass), thus confirming the composition of each synthetic gRNA.
[0171] CleanCap Cas9 mRNA fully substituted with 5-methoxyuridine was purchased from TriLink (L-7206). BE4-Gam mRNA and PE2 mRNA encoding the BE4-Gam and PE2 proteins, respectively, were custom-ordered and purchased from TriLink by TriLink adding their own proprietary 5' and 3' UTRs to the coding sequence provided. The custom mRNAs were fully substituted with 5-methylcytidine and pseudouridine, and had a CleanCap AG cap and polyA tail.
[0172] Cell culture and nucleofection. Human K562 cells were obtained from ATCC and cultured in RPMI 1640 + GlutaMax medium (gibco) supplemented with 10% fetal bovine serum (gibco). Nucleofection was performed on K562 cells (within passages 4-14) using 200,000 cells per transfection in 20 μL of SF buffer utilizing the Lonza SF Cell Line Kit (V4SC-2960) combined with 125 pmol gRNA and 6 μL of 1.87 pmol BE4-Gam mRNA in PBS buffer for cytidine base editing, or 8 μL of 125 pmol pegRNA with 100 pmol nicking gRNA and 1.35 pmol PE2 mRNA in PBS buffer for prime editing, using the Lonza 4D-Nucleofector (96-well shuttle device, program FF-120) according to the manufacturer's instructions. Cells were cultured at 37°C in ambient oxygen and 5% carbon dioxide, and cells were harvested 48 hours after transfection.
[0173] Human Jurkat clone E6-1 cells were obtained from ATCC and cultured in RPMI 1640 + GlutaMax medium supplemented with 10% fetal bovine serum. Nucleofection was performed on Jurkat cells (within passage 7-20) using the Lonza SE Cell Line kit (V4SC-1960) with 200,000 cells in 20 μL of SE buffer combined with 125 pmol of pegRNA, 100 pmol of nicking gRNA and 8 μL of 1.35 pmol of PE2 mRNA in PBS buffer (program CL-120). Cultured cells were harvested 72 hours after transfection.
[0174] Human HepG2 cells were obtained from ATCC and cultured in Dulbecco's Modified Eagle's Medium (DMEM) + L-glutamine + 4.5 g / L D-glucose medium (Gibco) supplemented with 10% fetal bovine serum. HepG2 cells (within passages 4-13) were spun down from the medium and spun down again with or without a PBS rinse. Cells were nucleofected using the Lonza SF Cell Line Kit (V4SC-2960) with 200,000 cells in 20 μL SF buffer combined with 10 pmol gRNA and 3 μL 0.0625 pmol Cas9 mRNA in PBS buffer (program EH-100) or nucleofected in the presence of residual serum by combining 200,000 cells in 20 μL SF buffer with 30 pmol gRNA and 0.5 pmol Cas9 mRNA or 5 μL 12.5 pmol Streptococcus pyogenes Cas9 (SpCas9) protein (Aldeveron) in PBS buffer. For the 163 residue gRNA, these were similarly nucleofected in the presence of residual serum and SF buffer by combining them with 125 pmol 163 residue gRNA and 5 μL 50 pmol SpCas9 protein in PBS buffer. For all RNP transfections, gRNA was precomplexed with SpCas9 protein (Aldevron) in PBS buffer by combining and incubating at room temperature for approximately 20 min before combining with cells in SF buffer for nucleofection. For mRNA transfections, gRNA was similarly combined with Cas9 mRNA (TriLink) in PBS buffer and kept on ice for approximately 20 min until combining with cells in SF buffer for nucleofection. Cultured HepG2 cells were harvested approximately 72 h after transfection.
[0175] Human primary T cells (LP, CR, CD3+, NS) were obtained from AllCells (Alameda, CA) and cultured in RPMI 1640+GlutaMax medium supplemented with 10% fetal bovine serum, 5ng / mL human IL-7 and 5ng / mL human IL-15 (gibco). Primary T cells were activated for 48 hours using anti-human CD3 / CD28 magnetic Dynabeads (Thermo Fisher) at a bead:cell concentration of 3:1. Nucleofection was performed on bead-depleted primary T cells using the Lonza P3 Primary Cell Kit (V4SP-3960) with 200,000 cells in 20μL of P3 buffer combined with 5pmol of gRNA and 2.7μL of 0.0625pmol of Cas9 mRNA in PBS buffer (program EO-115). Cultured cells were harvested 7 days after transfection. T cells were maintained at a density of approximately 1 million cells / mL of medium throughout the culture period, and additional medium was added every 2 days after electroporation.
[0176] qRT-PCR assay. Human K562 cells were cultured as described above and 200,000 cells per replicate were nucleofected with 125 pmol of gRNA (without Cas9 mRNA or protein) as described. At each time point, cells were collected in 1.7 mL Eppendorf tubes, rinsed with PBS, then resuspended in 750 μL Qiazol and kept at room temperature for 5 min before transferring to a -20°C freezer. Total RNA was isolated in PBS from Qiazol+chloroform extracts on a QiaCube HT using the miRNeasy kit (Qiagen) and immediately reverse transcribed using the Protoscript II first strand cDNA synthesis kit (NEB). qRT-PCR was performed on an Applied Biosystems QuantStudio 6 Flex instrument using TaqPath ProAmp master mix (Thermo Fisher) with two TaqMan MGB probes, one against FAM-labeled gRNA and the other against VIC-labeled U6 snRNA for normalization to the amount of total RNA isolated, calculated as ΔCt. ΔCt values for triplicate samples were averaged and normalized to the lowest observed mean ΔCt value to calculate ΔΔCt values. Relative gRNA levels were calculated as 2 -ΔΔCt It was calculated as:
[0177] PCR-targeted deep sequencing and quantification of targeted genomic modifications. Purification of genomic DNA and construction of PCR-targeted deep sequencing libraries were performed as described above. Library concentrations were determined using Qubit dsDNA BR Assay Kit (Thermo Fisher). Paired-end 2 × 220 bp reads were sequenced on 0.8 ng / μL PCR amplified libraries using MiSeq (Illumina) with 20.5% PhiX.
[0178] Paired-end reads were merged using FLASH version 1.2.11 software and then mapped against the human genome using BWA-MEM software (bwa-0.7.10) set to default parameters. Reads were scored as having or not having indels depending on whether an insertion or deletion was found within 10 bp of the Cas9 cleavage site. For prime editing analysis, reads were scored as having an edit if the desired edit was identified within the read. For cytidine base editing analysis, reads were scored as having a base edited if a cytidine was edited within a window 10-20 bp upstream of the PAM site. For each replicate in each experiment, mapped reads were separated according to the amplicon locus to which they were mapped and binned by the presence or absence of indels or edits. Counting of reads per bin was used to calculate the % indels or % edits produced at each locus. For plots, indel or editing yields and standard deviations were calculated by logit transformation of % indels or % editing, transformed as ln(r / (1-r)) to approximate a normal distribution, where r is the % indels or % editing per specific locus. Triplicate mock transfections gave average mock controls (or negative controls), while triplicate samples showing average indel yields or average editing yields significantly higher (p<0.05 by t-test) than the corresponding negative controls were considered to be higher than background.
[0179] [Example 1] This example assessed the stability of guide RNAs with 2'-O-methyl-3'-phosphonoacetate (MP) and 2'-O-methyl-3'-phosphorothioate (MS) modifications at the 3' end. To assess the relative longevity of single guide RNAs with MS or MP modifications at the 3' end in transfected cells, guide RNAs were synthesized with MS modifications at the first three internucleotide linkages at the 5' end and either MS modifications at the last three internucleotide linkages at the 3' end (denoted as 3xMS,3xMS) or two, three or four consecutive MP modifications at the terminal internucleotide linkages at the 3' end (denoted as 3xMS,2xMP; 3xMS,3xMP; and 3xMS,4xMP, respectively). Each modified gRNA was individually transfected into human K562 cells in the absence of Cas9, and qRT-PCR was used to measure the relative amount of sgRNA remaining in cells harvested at a series of time points from 1 to 96 hours after transfection.
[0180] As shown in Figure 7, a more rapid drop was observed in the relative levels of 3xMS,3xMS gRNA detected over 1, 6, and 24 h post-transfection compared to gRNAs whose 3' ends were modified with either MPs (either 2, 3, or 4 consecutive MPs). Specifically, after 1 h of transfection, the relative amounts of transfected gRNAs differed only by 2.6-fold, with error bars largely overlapping between all four variants of 3'-end protection, whereas a much larger difference was observed after 6 h of transfection, with the remaining amount of 3xMS,3xMS-protected gRNA dropping to a relative level (0.039) that was approximately 1 / 10th that of the 3xMS,3xMP- and 3xMS,4xMP-protected gRNAs (0.341-0.351). The differences were even greater at the 24-hour time point, when they varied in a logical order along with the 3'-end protection level, from having 3xMS at the 3'-end to having 2xMP, 3xMP, 4xMP, resulting in residual gRNA levels ranging approximately 250-fold, consistent with the level of 3'-end protection. Thus, it was found that incorporating MP modifications at the 3'-end of uncomplexed gRNAs can significantly enhance their stability in transfected cells compared to MS modifications, specifically by one to two orders of magnitude for three different MP-modified gRNAs tested in parallel with a gRNA modified with MS alone. Designs with three or four consecutive MPs at the 3'-end can extend the lifespan of free gRNAs for longer time points (72 and 96 hours after transfection).
[0181] [Example 2] It has been demonstrated that phosphonate modifications can be stably incorporated into DNA and RNA oligonucleotides and increase their nuclease resistance compared to phosphorothioates. Previous reports exploring the use of MPs to increase gRNA specificity by incorporating MPs into the 20 nt guide sequence portion have found that MPs at specific sequence positions such as positions 5 or 11 (counting from the 5' end of the 20 nucleotides) can significantly reduce off-target editing while maintaining high on-target editing, as described, for example, in Ryan et al., Nucleic Acids Research 46, 792-803 (2018). However, it has also been reported that incorporating MP modifications within the first 1, 2 or 3 nucleotides at the 5' end of the gRNA can reduce their on-target cleavage activity and / or increase their off-target activity, thus reducing specificity, for some guide sequences (see, for example, Ryan et al., 2018).
[0182] To further explore the potential utility of phosphonate modifications in guide RNAs, the performance of gRNAs containing different numbers of consecutive 2'-O-methyl-3'-phosphonoacetate (2'-O-methyl-3'-PACE, or "MP") modifications at their 3' ends was evaluated compared to guide RNAs with 2'-O-methyl-3'-phosphorothioate (or "MS") modifications at their ends. The results of this study are further described in Ryan et al. "Phosphonoacetate Modifications Enhance the Stability and Editing Yields of Guide RNAs for Cas9 Editors." Biochemistry (2022) doi.org / 10.1021 / acs.biochem.lc00768.
[0183] This experiment was designed to evaluate Cas activity following co-transfection of HepG2 cells with relatively low (subsaturating) amounts of chemically modified guide RNA and mRNA encoding a Cas protein, using HBB as the target gene. Such subsaturating amounts constitute a challenging situation for editing a target region of cells.
[0184] For three sets of samples, mRNA encoding Cas9 was co-transfected into human hepatocytes (HepG2 cells) with modified gRNA targeting HBB (see Table 1 above). For the fourth set of HepG2 cells, modified gRNA targeting the same site in HBB was precomplexed with purified recombinant Cas9 protein to form RNPs, which were then transfected into the cells. Each transfection was performed in triplicate samples of cells cultured separately. Genomic DNA was harvested, and HBB target and off-target sequences were amplified using primers specific for the HBB gene and intergenic off-target sites, respectively, to produce amplicons that were sequenced, and the extent of editing at the target and off-target sites ("% indels") was determined from the sequencing results. ON and OFF indicate on-target and off-target sequences, respectively. Intergenic off-target loci were monitored, as they are known to undergo high side activity when targeted against selected target sequences in the HBB gene. The editing yields for the modified gRNAs listed in Table 1 are plotted as bar graphs in Figures 2-5.
[0185] As demonstrated by Figures 2-5, co-transfection of subsaturating levels of modified guide RNA targeting HBB with mRNA (or RNP complex of modified guide RNA) encoding Cas proteins resulted in higher levels of editing yield compared to samples co-transfected with equivalent amounts of unmodified gRNA. Furthermore, the addition of 2, 3, or 4 MP modifications at the 3' end of the modified gRNA resulted in progressively higher increases in editing yield in addition to a substantial increase compared to modified gRNAs containing 3 MS modifications at the 3' end of the modified gRNA. See, e.g., Figures 2 and 4. As shown by Figures 3 and 5, inclusion of MP modifications at positions 5 or 11 (counting from the 5' end of the 20 nt guide sequence in the gRNA) also reduced off-target activity. Notably, inclusion of MP at position 5 in the gRNA had minimal impact on editing yield, whereas substantially reduced off-target activity.
[0186] It was further observed that MP modifications at the 3' end significantly enhanced the editing yield in HepG2 cells (Figure 3). For example, designs with 2, 3 or 4 consecutive MP modifications at the 3' end gave at least 2-fold more Cas9-mediated indels than comparable designs with 3xMS at the 3' end (81%-83% at on-target sites for 2xMP, 3xMP and 4xMP modifications at the 3' end vs. 38% for 3xMS). A similar trend was observed for the same gRNAs transfected into primary human T cells, although the enhancement was more modest, in that 2xMP, 3xMP and 4xMP modifications gave 1.3-fold higher levels of on-target indels than when using 3xMS at the 3' end (Figure 8). Incorporation of an additional MP at position 5 in the 20 nt guide sequence portion of the gRNA significantly reduced editing at the OFF1 site in both cell types while maintaining high on-target editing efficiency as previously reported as a means to increase specificity (see, e.g., Ryan et al., 2018). Indeed, indels at the OFF1 site were reduced 7- to 10-fold in HepG2 cells by incorporation of an MP at position 5, as well as 6- to 7-fold in primary T cells.
[0187] As shown by FIG. 9, the use of chemically modified gRNAs in combination with base editors was also evaluated. Base editors are a class of alternative genome editing systems built around Cas9 nickase (nCas9) or dead Cas9 (dCas9) fused to one of a variety of deaminases that allow editing of genomic DNA in cells without creating double-stranded breaks. Both cytidine base editors (CBEs) and adenosine base editors (ABEs) have been reported, which have undergone several modifications for base editing. The potential benefit of using MP modifications as opposed to MS modifications at the 3' end of such gRNAs was tested in the context of CBEs, i.e., BE4-Gam mRNA. A 1.4-fold higher level of cytidine editing was observed by using CBE mRNA in K562 cells co-transfected with gRNAs modified with MP at the 3' end, compared to alternative designs with MS at the 3' end.
[0188] [Example 3] This example evaluated the use of 2'-O-methyl-3'-phosphonoacetate (MP) and 2'-O-methyl-3'-phosphorothioate (MS) modifications at the 3' end of chemically synthesized pegRNA. Experiments were performed to explore two approaches for prime editing adopted from the literature: knocking out the PAM in EMX1 or introducing a three-base insertion into RUNX1. Both of these approaches utilize pegRNA with a primer binding sequence containing 15 nucleotides. The specific sequence editing evaluated in this experiment is shown in Figure 18. mRNA encoding the prime editor (in this case a fusion protein containing Cas9 nickase and MMLV-derived reverse transcriptase) was introduced into K562 or Jurkat cells along with pegRNA targeting the EMX1 gene. Each transfection was performed in triplicate cell samples cultured separately. Genomic DNA was harvested, the EMX1 target sequence was amplified using primers specific for EMX1 to produce amplicons, the amplicons were sequenced, and the degree of prime editing ("% edited") was determined from the sequencing results. The degree of unwanted indel formation ("% indel") at the nickase site within the EMX1 target sequence was also determined from the sequencing results. Such indels are known by-products of prime editing and are generally considered undesirable (see Anzalone et al. 2019). The yield of prime editing and the yield of indel by-products per pegRNA are plotted as bar graphs in Figures 11-16. The sequences used for this assay were selected from the sequences shown in Table 2. The data in Figures 11-12 were obtained using a first batch synthesis of pegRNA targeting EMX1, while the data in Figures 13-14 were obtained using a second batch synthesis of pegRNA targeting EMX1. Note that some of the same sequences were synthesized again in the second batch synthesis. Conversely, the data in Figures 15-16 were obtained using pegRNA targeting RUNX1 (i.e., using the sequences listed in Table 3).
[0189] As illustrated by the results shown in Figures 11-16, inclusion of MS and MP nucleotides as chemical modifications at the 5' and 3' ends, respectively, of pegRNA increases prime editing activity. The high activity of constructs with modified nucleotides at the 3' end of pegRNA is particularly surprising given the fact that the 3' end of pegRNA contains additional functional sites (e.g., primer binding site and template sequence for reverse transcriptase). As discussed above, prior to the present disclosure, inclusion of chemically modified nucleotides (e.g., MS and / or MP) at this site was expected to interfere with functionality provided by these other 3' end components of pegRNA.
[0190] [Example 4] This example evaluated the incorporation of MP or MS modifications at the 3' end of chemically synthesized pegRNA. The method used in this experiment is consistent with the above method. Briefly, a prime editing approach was adopted to knock out the PAM in EMX1 or introduce a 3-base insertion into RUNX1. For editing EMX1 or RUNX1, K562 cells were co-transfected with prime editor mRNA (in this case a fusion protein containing Cas9 nickase and MMLV-derived reverse transcriptase) and synthetic pegRNA modified at the 5' end by 3xMS and at the 3' end by various modification schemes (as indicated). For editing EMX1 or RUNX1, the same pegRNA was used to similarly transfect Jurkat cells. Editing efficiency was measured by deep sequencing of PCR amplicons of the target locus for both the desired editing (%Edited) and any contaminating indel by-products (%By-indel). Bars in the accompanying figures represent the mean with standard deviation (n=3).
[0191] As shown by Figures 19-22, this experiment compared pegRNAs with 3xMS at the 3' end to alternative designs with one, two or three consecutive MPs at the 3' end for both targets. Each pegRNA was co-transfected with PE2 mRNA in K562 or Jurkat cells, and the results show that pegRNAs with MP modifications at the 3' end performed well and could achieve similar, or in some cases somewhat higher, editing yields than 3xMS. For the two pegRNA sequences tested here, designs with 2xMP and / or 3xMP at the 3' end consistently performed better (specifically 1.2-1.4 times better) than designs with 1xMP at the 3' end.
[0192] [Example 5] In this example, we demonstrated that the use of MP modification at the 3' end of chemically synthesized gRNA helps maximize editing yields in the presence of serum. To simulate the harsher cellular environment that CRISPR-Cas components may encounter when delivered in vivo (by nanocarriers or other cell-permeable formulations), we performed experiments in which Cas9 mRNA was co-transfected with gRNA into cells that were isolated from the culture medium but were not rinsed with PBS buffer to remove residual serum. Serum is known to contain nucleases.
[0193] The method used in this experiment is consistent with the above method.However, it is noted that under this experimental condition, more gRNA and Cas9 mRNA are required to achieve substantial editing levels.Specifically, 3 times more gRNA and 8 times more Cas9 are used per transfection for the experiment, resulting in the data shown in Figure 23.This is compared with the experiment that results in the data shown in Figure 3, in which the same number of cells per transfection are washed with buffer before introducing CRISPR-Cas components.
[0194] Based on the results of this study, it appears that extracellular nucleases in serum that were not rinsed from the cells degraded the transfected RNA. We found that when unwashed HepG2 cells were co-transfected with Cas9 mRNA, gRNAs with MP modifications at the 3' end gave substantially higher (one order of magnitude or more) editing yields compared to gRNAs with MS modifications at the 3' end (Figure 23). Specifically, we observed 15%-44% editing yields for gRNAs with one or more MPs at the 3' end compared to less than 2% for those with 3xMS at the 3' end.
[0195] In parallel experiments, RNP versions of each gRNA were prepared by precomplexing with Cas9 protein in PBS buffer and transfected into aliquots of unwashed HepG2 cells. As expected, unmodified and 3xMS modified gRNAs in RNP formulations gave higher indel yields than when they were co-transfected with Cas9 mRNA, because precomplexing of gRNA with Cas9 protein in the RNP state helps to shield the gRNA from nucleolytic degradation (compare results shown in Figure 24 and Figure 23). Although the improvement in editing efficiency between RNPs incorporating gRNAs with MP and MS modifications at the 3' end was not as dramatic as when these modifications were used for co-transfection with Cas9 mRNA, designs with MP at the 3' end gave significantly higher Cas9-mediated indels than similar designs with 3xMS at the 3' end (70%-73% indels for 2xMP, 3xMP, and 4xMP modifications at the 3' end compared to 52% for 3xMS at the 3' end at the ON target site, a difference of about 1.3 fold) (see Figure 24). Similar outcomes were observed for a different set of synthetic 163-residue gRNAs designed for the CRISPRa SAM system but used with SpCas9 protein in an RNP formulation to produce indels instead of using them for gene activation by CRISPRa (Figure 25).
[0196] Exemplary embodiments <Modifications Section A> Embodiment A1. A method for editing a target region in a nucleic acid under one or more challenging conditions, comprising: To the cells, a) CRISPR-associated ("Cas") proteins, and b) a modified guide RNA comprising a guide sequence capable of hybridizing to a target region and a scaffold that interacts with a Cas protein, the modified guide RNA comprising a 5' end and a 3' end, the modified guide RNA further comprising one or more modified nucleotides within five nucleotides of the 3' end, the one or more modified nucleotides comprising at least one nucleotide having a 2' modification and an internucleotide linkage modification, the 2' modification being selected from 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl (2'-MOE) and 2'-deoxy, and the internucleotide linkage modification being a phosphonocarboxylate or a thiophosphonocarboxylate. providing a The one or more difficult circumstances are: i. the target region or cells containing the target region are in a medium containing serum (e.g., fetal bovine serum); ii. the cells containing the target region have been previously cultured in medium containing serum, and the cells have been incompletely separated from the serum; iii. the cells containing the target region were previously cultured in medium containing one or more exoribonucleases, and the cells were incompletely separated from the one or more exoribonucleases; iv. cells containing the target region have a relatively high level of exoribonuclease activity, such as a relatively high expression of one or more exoribonucleases; v. cells containing the target region have a relatively low level of ribonuclease inhibitor activity, such as a relatively low expression of a ribonuclease inhibitor; vi. the modified guide RNA is not complexed with a Cas protein prior to delivery to a cell containing the target region; and vii. Applicable combinations thereof selected from the group consisting of; A method in which a Cas protein and a modified guide RNA form a complex that results in editing of the target region. Embodiment A2. The method of embodiment A1, wherein the internucleotide linkage modification is a phosphonocarboxylate. Embodiment A3. The method of embodiment A2, wherein the phosphonocarboxylate is a phosphonoacetate. Embodiment A4. The method of embodiment A1 wherein the thiophosphonocarboxylate is a thiophosphonoacetate. Embodiment A5. The method of any of embodiments A1-4, wherein the Cas protein is introduced as an mRNA encoding the Cas protein. Embodiment A6. The method of any of embodiments A1-4, wherein the Cas protein is introduced as an expression vector encoding the Cas protein. Embodiment A7. The method of embodiment A5 or A6, wherein the mRNA or expression vector encoding the Cas protein is contained in a nanoparticle when introduced into the target area. Embodiment A8. The method of any one of embodiments A1 to A4, wherein the Cas protein and the guide RNA are introduced as a ribonucleoprotein (RNP) complex. Embodiment A9. The method of any preceding embodiment, wherein the 2' modification is 2'-O-methyl. Embodiment A10. The method of any one of embodiments A1-A8, wherein the 2' modification is 2'-fluoro. Embodiment A11. The method of any one of embodiments A1-A8, wherein the 2' modification is 2'-MOE. Embodiment A12. The method of any one of embodiments A1-A8, wherein the 2' modification is 2'-deoxy. Embodiment A13. The method of any of the previous embodiments, wherein the one or more edits comprise one or more single nucleotide changes, one or more insertions of nucleotides, and / or one or more deletions of nucleotides. Embodiment A14. The method of any preceding embodiment, wherein the target region is present in a cell-free assay. Embodiment A15. The method of embodiment A14, further comprising extracting nucleic acid from the cells, such as by lysing the cells, forming an assay mixture comprising the extracted nucleic acid and one or more other cellular components, such as an exoribonuclease or other enzyme, and introducing a guide RNA into the assay mixture. Embodiment A16. The method of any of the previous embodiments, wherein the target region is in a cell with high ribonuclease expression, concentration and / or activity, e.g., a cell type with high specific nuclease. Embodiment A17. The method of embodiment A16, wherein the cells comprise primary cells. The method of embodiment A17, wherein the cells are ex vivo and the method further comprises one or more steps for isolating the cells from a living organism. The cells can be separated into the reaction mixture, or the separated cells can be transferred into the reaction mixture. Embodiment A19. The method of any of embodiments A16-A18, wherein the cell is isolated from a multicellular organism prior to introducing the modified guide RNA and Cas protein into the target region within the cell. Embodiment A20. The method of embodiment A19, wherein the cell or its progeny is returned to the multicellular organism after introducing the modified guide RNA and Cas protein into the target region within the cell. Embodiment A21. The method of any of embodiments A16-A20, wherein the cells are primary cells. Embodiment A22. The method of embodiment A21, wherein the primary cells are stem cells or immune cells. Embodiment A23. The method of embodiment A22, wherein the stem cells are hematopoietic stem and progenitor cells (HSPCs), mesenchymal stem cells, neural stem cells, or organ stem cells. Embodiment A24. The method of embodiment A22, in which the immune cell is a T cell, a natural killer cell, a monocyte, a peripheral blood mononuclear cell (PBMC), or a peripheral blood lymphocyte (PBL). Embodiment A25. The method of embodiment A24, wherein the cell is a T cell. Embodiment A26. The method of any of embodiments A16-A20, wherein the cells are hepatocytes. Embodiment A27. The method of any of embodiments A16-A26, wherein the cells are a population of cells each comprising a target region. Embodiment A28. The method of any of embodiments A16-A27, wherein the cells are in cell culture and the cells are in a cell culture medium that includes serum or one or more other medium components. Embodiment A29. The method of embodiment A28, wherein the cells are separated from the cell culture medium before the Cas protein and the modified guide RNA are introduced. Embodiment A30. The method of any of embodiments A1-13 and A16-A29, wherein the Cas protein and the modified guide RNA are introduced into a living organism. The method of embodiment A30, wherein the Cas protein and the modified guide RNA are introduced into a living organism or into a serum-containing fluid from a living organism. Embodiment A32. The method of any preceding embodiment, wherein the edit is a prime edit and the modified guide RNA further comprises a region containing the desired edit. Embodiment A33. The method of any of the preceding embodiments, wherein the editing comprises homology directed repair (HDR), non-homologous end joining (NHEJ), prime editing, or base editing. Embodiment A34. The method of any of the preceding embodiments, wherein the Cas protein is a Cas9 or Cas12 protein. Embodiment A35. The method of any preceding embodiment, wherein the Cas protein is a Cas nickase capable of nicking a single strand of DNA. Embodiment A36. The method of any preceding embodiment, wherein the Cas protein is a fusion protein comprising a Cas domain and a heterologous functional domain, and the heterologous functional domain comprises base editing activity, nucleotide deaminase activity, transglycosylase activity, methylase activity, demethylase activity, reverse transcriptase activity, polymerase activity, translation activating activity, translation repressing activity, transcription activating activity, transcription repressing activity, transcription release factor activity, chromatin modifying or remodeling activity, histone modifying activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, nucleic acid binding activity, detectable activity, or any combination thereof. Embodiment A37. The method of embodiment A36, wherein the fusion protein comprises a Cas nickase domain and a nucleotide deaminase. Embodiment A38. The method of embodiment A36, wherein the nucleotide deaminase is an adenosine deaminase or a cytidine deaminase. Embodiment A39. The method of embodiment A36, wherein the fusion protein comprises one or more nucleic acid modifying domains. Embodiment A40. The method of embodiment A36, wherein the nucleic acid-modifying domain is a DNA polymerase domain, a recombinase domain, a ribonucleotide reductase domain, a methyltransferase domain, a diadenosine tetraphosphate hydrolase domain, a DNA helicase domain, or an RNA helicase domain. Embodiment A41. The method of embodiment A36, wherein the fusion protein comprises a Cas nickase domain and a reverse transcriptase domain. Embodiment A42. The method of any preceding embodiment, wherein the guide RNA is a single guide RNA. Embodiment A43. The modified guide RNA is at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93 , 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 1 17, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, or 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 1 72, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200 nucleotides, and / or Max 180, 179, 178, 177, 176, 175, 174, 173, 172, 171, 170, 169, 168, 167, 166, 165, 164, 163, 162, 161, 159, 158, 157, 156, 155, 154, 153, 152, 151, 150, 149, 148, 147, 146, 145 , 144, 143, 142, 141, 140, 139, 138, 137, 136, 135, 134, 133, 132, 131, 130, 129, 128, 127, 126, 125, 124, 123, 122, 121, or 120 nucleotides. Embodiment A44. The method of any preceding embodiment, wherein the guide RNA further comprises one or more modified nucleotides within the 5'-terminal 5 nucleotides or within the 5'-terminal 3 nucleotides. Embodiment A45. The method of embodiment A44, wherein the one or more modified nucleotides at the 5' terminus comprise at least one nucleotide having a 2' modification and an internucleotide linkage modification, wherein the 2' modification is selected from 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl (2'-MOE) and 2'-deoxy, and the internucleotide linkage modification is selected from phosphonocarboxylate, thiophosphonocarboxylate, and phosphorothioate. Embodiment A46. The method of any preceding embodiment, wherein the guide RNA further comprises one or more modified nucleotides at one or more positions other than at least 5 nucleotides from both the 5' and 3' ends of the guide RNA. Embodiment A47. A method of modulating expression of a target gene in a target region in a nucleic acid in a cell under one or more stringent conditions, comprising administering to the cell a) a CRISPR-associated ("Cas") protein, or DNA or mRNA encoding a Cas protein, and b) a modified guide RNA comprising a guide sequence capable of hybridizing to a target region and a region that interacts with a Cas protein, the modified guide RNA comprising a 5' end and a 3' end, the modified guide RNA further comprising one or more modified nucleotides within the 5 nucleotides of the 3' end, the one or more modified nucleotides comprising at least one nucleotide having a 2' modification and an internucleotide linkage modification, the 2' modification being selected from 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl (2'-MOE) and 2'-deoxy, and the internucleotide linkage modification being a phosphonocarboxylate or a thiophosphonocarboxylate; including providing; A method in which the Cas protein and the modified guide RNA form a complex that results in modulation of expression of the target region. Embodiment A48. The method of embodiment A47, wherein the Cas protein or modified guide RNA further comprises an epigenetic modifier, or a transcriptional or translational activation or repression signal. Embodiment A49. The method of embodiment A47, wherein the Cas protein is a fusion protein comprising an inactive Cas nuclease domain and a heterologous functional domain selected from a transcriptional activation domain and a transcriptional repression domain. Embodiment A50. The method of embodiment A49, wherein the heterologous functional domain is a transcription activation domain. Embodiment A51. The method of embodiment A50, wherein the transcription activation domain is a VP64 domain, a p65 domain, a MyoD1 domain, or an HSF1 domain. Embodiment A52. The method of embodiment A49, wherein the heterologous functional domain is a transcriptional repression domain. Embodiment A53. The method of embodiment A52, wherein the transcriptional repression domain is a KRAB domain, a SID domain, a SID4X domain, a NuE domain, or an NcoR domain. Embodiment A54. A method for prime editing a target region in a nucleic acid in one or more stringent conditions, comprising: a) In a cell, a Cas protein that can nick a single strand of nucleic acid; Reverse transcriptase and i) a guide sequence capable of hybridizing to a target region; ii) a region that interacts with a Cas protein; iii) a reverse transcriptase template sequence containing one or more edits to the sequence of the nucleic acid; and iv) a primer binding site sequence capable of binding to the complement of the target region; a modified prime-edited guide RNA ("pegRNA") comprising providing a the modified pegylated RNA comprises a 5' end and a 3' end, the modified pegylated RNA further comprises one or more modified nucleotides within the 5 nucleotides of the 3' end, the one or more modified nucleotides comprising at least one nucleotide having a 2' modification and an internucleotide linkage modification, the 2' modification being selected from 2'-O-methyl, 2'-fluoro, 2'-O-methoxyethyl (2'-MOE) and 2'-deoxy, and the internucleotide linkage modification is a phosphonocarboxylate or a thiophosphonocarboxylate; A method in which the Cas protein and the modified guide RNA form a complex that results in editing of the target region. Embodiment A55. The method of embodiment A54, wherein the Cas protein and the reverse transcriptase are connected by a linker to form a fusion protein. Embodiment A56. The method of any one of the preceding embodiments, wherein the guide RNA comprises at least one phosphorothioate internucleotide linkage within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate internucleotide linkages within the 3'-terminal five nucleotides. Embodiment A57. The method of any one of the preceding embodiments, wherein the guide RNA comprises at least one phosphorothioate internucleotide linkage within the 5'-terminal five nucleotides and at least two consecutive phosphonoacetate or thiophosphonoacetate internucleotide linkages within the 3'-terminal five nucleotides. Embodiment A58. The method of any one of the preceding embodiments, wherein the guide RNA comprises at least one MS within the 5 nucleotides of the 5' end and at least two consecutive MPs or MSPs within the 5 nucleotides of the 3' end. Embodiment A59. The method of any one of the preceding embodiments, wherein the guide RNA comprises three MSs within the 5 nucleotides of the 5' end and three MPs or MSPs within the 5 nucleotides of the 3' end. Embodiment A60. The method of any one of the preceding embodiments, wherein editing and / or modulation of target gene expression is performed in a multiplex manner (i.e., for at least two target genes or at least two target regions).
[0197] <Section B> Embodiment B1. A method of editing a target region in a nucleic acid in a cell, comprising administering to the cell: a) CRISPR-associated ("Cas") proteins, and b) with a 5' end and a 3' end; a guide sequence capable of hybridizing to a target sequence within the target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; Modified guide RNA containing providing a the cell is present ex vivo in the presence of a nuclease-containing solution or is present in vivo; The method, wherein the providing step results in editing of the target region. Embodiment B1.1. A method of editing a target region in a nucleic acid in a cell, comprising administering to the cell: a) CRISPR-associated ("Cas") proteins, and b) a modified guide RNA that is a prime edited guide RNA (pegRNA) comprising a 5' end and a 3' end, one of which is a prime edited end and the other is a distal end, a guide sequence capable of hybridizing to a target sequence within the target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the five nucleotides of the distal end and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the five nucleotides of the prime edited end; A modified guide RNA further comprising providing a The cell is present ex vivo in the presence of a nuclease-containing solution or is present in vivo; The method, wherein the providing step results in editing of the target region. Embodiment B2. The method of embodiment B1 or B1.1, wherein the editing occurs more efficiently than with an unmodified gRNA that is otherwise identical to the modified guide RNA. Embodiment B3. A method of modulating expression of a target gene within a target region in a nucleic acid in a cell, comprising administering to the cell: a) CRISPR-associated ("Cas") proteins, and b) with a 5' end and a 3' end; a guide sequence capable of hybridizing to a target sequence within the target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; providing a modified guide RNA comprising: the cell is present ex vivo in the presence of a nuclease-containing solution or is present in vivo; The method, wherein said providing step results in modulation of expression of a target gene. Embodiment B4. The method of embodiment B3, wherein modulation occurs more efficiently than with an unmodified gRNA that is otherwise identical to the modified guide RNA. Embodiment B5. The method of any one of the preceding B embodiments, wherein the cell is in vivo. Embodiment B6. The method of any one of the preceding B embodiments, wherein the cells are ex vivo in the presence of a nuclease-containing solution. Embodiment B7. The method of any one of the preceding B embodiments, wherein the modified guide RNA comprises at least two consecutive 2'-O-methyl-3'-phosphorothioates (MS) within the five nucleotides of the 5' end (Exception: "distal end" instead of "5' end" if this embodiment is subject to embodiment B1.1). Embodiment B8. The method of any one of the preceding B embodiments, wherein the phosphonocarboxylate is a phosphonoacetate and the thiophosphonocarboxylate is a thiophosphonoacetate. Embodiment B9. The method of any one of the preceding B embodiments, wherein the modified guide RNA comprises at least two consecutive 2'-O-methyl-3'-phosphonoacetates (MP) or 2'-O-methyl-3'-thiophosphonoacetates (MSP) within the 5 nucleotides of the 3' end (Exception: "prime stop end" instead of "5' end" if this embodiment is subject to embodiment B1.1). Embodiment B10. The method of any one of the preceding B embodiments, wherein the modified guide RNA further comprises modified nucleotides located outside of 5 nucleotides within the 5' and 3' ends. Embodiment B11. The method of any one of the preceding B embodiments, wherein the modified guide RNA is a single-stranded guide RNA. Embodiment B12. The method of any one of the preceding B embodiments, wherein the Cas protein is provided as an mRNA encoding the Cas protein. Embodiment B13. The method of any one of embodiments B1-B11, wherein the Cas protein is provided as DNA encoding the Cas protein. Embodiment B14. The method of embodiment B13, wherein the DNA is a viral expression vector. Embodiment B15. The method of any one of embodiments B1 to B11, wherein the Cas protein and the modified guide RNA are provided as a ribonucleoprotein complex (RNP). Embodiment B16. The method of any one of embodiments B1 to B11, wherein the Cas protein and / or the modified guide RNA is provided in a nanoparticle. Embodiment B17. The method of any one of the preceding B embodiments, wherein the efficiency is at least 5% higher. Embodiment B18. The method of any one of the preceding B embodiments, wherein the efficiency is at least 10% higher. Embodiment B19. The method of any one of the preceding B embodiments, wherein the efficiency is at least 15% higher. Embodiment B20. The method of any one of the preceding B embodiments, wherein the efficiency is at least 20% higher. Embodiment B21. The method of any one of the preceding B embodiments, wherein the efficiency is at least 25% greater. Embodiment B22. The method of any one of the preceding B embodiments, wherein the efficiency is at least 30% greater. Embodiment B23. The method of any one of the preceding B embodiments, wherein the efficiency is at least 35% greater. Embodiment B24. The method of any one of the preceding B embodiments, wherein the efficiency is at least 40% greater. Embodiment B25. The method of any one of the preceding B embodiments, wherein the efficiency is at least 45% greater. Embodiment B26. The method of any one of the preceding B embodiments, wherein the efficiency is at least 50% greater. Embodiment B27. The method of any one of the preceding B embodiments, wherein the Cas protein is capable of cleaving both strands of DNA. Embodiment B28. The method of any one of embodiments B1 to B26, wherein the Cas protein is a nickase. Embodiment B29. The method of any one of embodiments B1 to B26, wherein the Cas protein does not have nuclease activity. Embodiment B30. The method of any one of the preceding B embodiments, wherein the Cas protein is part of a fusion protein further comprising a heterologous protein. Embodiment B31. The method of any one of the preceding B embodiments, wherein the Cas protein is a type II Cas protein. Embodiment B32. The method of any one of the preceding B embodiments, wherein the Cas protein is a Cas9 protein, or a variant or fragment thereof. Embodiment B33. The method of embodiment B32, wherein the Cas9 protein is derived from Streptococcus pyogenes. Embodiment B34. The method of any one of embodiments B1 to B32, wherein the Cas protein is a Cpf1 protein, or a variant or fragment thereof. Embodiment B35. The method of any one of the preceding B embodiments, wherein the Cas protein is a hybrid protein having sequences from at least two different wild-type Cas proteins. Embodiment B36. The method of any one of the preceding B embodiments, wherein the modified guide RNA is 40-70 nucleotides in length. Embodiment B37. The method of any one of the preceding B embodiments, wherein the modified guide RNA is 40-100 nucleotides in length. Embodiment B38. The method of any one of embodiments B1 to B35, wherein the modified guide RNA is 90 to 110 nucleotides in length. Embodiment B39. The method of any one of embodiments B1-B35, wherein the modified guide RNA is 90-130 nucleotides in length. Embodiment B40. The method of any one of embodiments B1-B35, wherein the modified guide RNA is 130-160 nucleotides in length. Embodiment B41. The method of any one of embodiments B1-B35, wherein the modified guide RNA is 160-200 nucleotides in length. Embodiment B42. The method of any one of the preceding B embodiments, wherein the modified guide RNA is a pegRNA. Embodiment B43. The method of any one of the preceding B embodiments, wherein a phosphorothioate, phosphonocarboxylate or thiophosphonocarboxylate modification is present in a nucleotide that also contains a 2'-O-methyl modification, respectively. Embodiment B44. A 5' end and a 3' end, a guide sequence capable of hybridizing to a second target sequence within the second target region; one or more phosphorothioate modifications within the 5'-terminal five nucleotides (except that this is the distal end if this embodiment is dependent on B1.1), and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides (except that this is the prime edited end if this embodiment is dependent on B1.1); The method of any one of the preceding B embodiments, further comprising editing a second target region in the cell using a second modified guide RNA comprising: Embodiment B45. A 5' end and a 3' end, a guide sequence capable of hybridizing to a third target sequence within the third target region; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; The method of any one of the preceding B embodiments, further comprising modulating expression of a third target gene at a third target region in the cell using a third modified guide RNA comprising: Embodiment B46. The method of any one of the preceding B embodiments, wherein the nuclease is an exonuclease. Embodiment B47. The method of any one of the preceding B embodiments, wherein the nuclease is a ribonuclease.
[0198] <Section C> Embodiment C1. A method of editing two or more nucleic acid target regions, including a first target region and a second target region, in a cell, comprising: a) CRISPR-associated ("Cas") proteins; b) a 5' end and a 3' end; a first guide sequence capable of hybridizing to a first target sequence within a first target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; a first modified guide RNA comprising: c) a 5' end and a 3' end; a second guide sequence capable of hybridizing to a second target sequence within the second target region; and A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; a second modified guide RNA comprising providing a the cell is present ex vivo in the presence of a nuclease-containing solution or is present in vivo; The method, wherein the providing step results in editing of the first and second target regions. Embodiment C2. A method of modulating expression of at least a first target gene of a first target region and a second target gene in a second target region in a cell, comprising administering to the cell: a) CRISPR-associated ("Cas") proteins; b) a 5' end and a 3' end; a first guide sequence capable of hybridizing to a first target sequence within a first target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; a first modified guide RNA comprising: c) a second guide sequence having a 5' end and a 3' end and capable of hybridizing to a second target sequence within a second target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides and at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; a second modified guide RNA comprising providing a the cell is present ex vivo in the presence of a nuclease-containing solution or is present in vivo; The method, wherein said providing step results in modulation of expression of the first and second target genes. Embodiment C3. The method of embodiment C1 or C2, wherein editing of the first target region or modulation of the first target gene has a first efficiency that is greater than that of an otherwise identical unmodified guide RNA to the first modified guide RNA. Embodiment C4. The method of embodiment C3, wherein editing of the second target region or modulation of the second target gene has a second efficiency that is greater than that of an otherwise identical unmodified guide RNA with the second modified guide RNA. Embodiment C5. The method of any one of the preceding C embodiments, wherein the cell is in vivo. Embodiment C6. The method of any one of embodiments C1-C4, wherein the cells are ex vivo in the presence of a nuclease-containing solution. Embodiment C7. The method of any one of the preceding C embodiments, further comprising the additional limitations applicable from each of the A or B embodiments.
[0199] The foregoing description of exemplary or preferred embodiments should be considered as illustrative, rather than limiting, of the present disclosure as defined by the claims. As will be readily recognized, numerous variations and combinations of the features set forth above can be utilized without departing from the present disclosure as set forth in the claims. Such variations are not considered as a departure from the scope of the present disclosure, and all such variations are intended to be included within the scope of the following claims. All references cited herein are incorporated by reference in their entirety.
[0200] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
Claims
1. 1. A method of editing a target region in a nucleic acid in a cell, or modulating expression of a target gene within a target region, comprising administering to the cell: a) a CRISPR-associated ("Cas") protein; and b) a 5' end and a 3' end; a guide sequence capable of hybridizing to the target sequence within the target region; A scaffold region that interacts with the Cas protein; one or more phosphorothioate modifications within the 5'-terminal five nucleotides; at least two consecutive phosphonocarboxylate or thiophosphonocarboxylate modifications within the 3'-terminal five nucleotides; providing a modified guide RNA comprising: the cells are present ex vivo in the presence of a nuclease-containing solution or are present in vivo; wherein said providing step results in editing of said target region or modulation of expression of said target gene.
2. 2. The method of claim 1, wherein the phosphorothioate, phosphonocarboxylate, and thiophosphonocarboxylate modifications are each present in a nucleotide that also contains a 2'-O-methyl.
3. 2. The method of claim 1, wherein the modified guide RNA comprises at least two consecutive 2'-O-methyl-3'-phosphorothioates (MS) within the 5 nucleotides of the 5' end.
4. 2. The method of claim 1, wherein the phosphonocarboxylate is a phosphonoacetate and the thiophosphonocarboxylate is a thiophosphonoacetate.
5. 4. The method of claim 3, wherein the modified guide RNA comprises at least two consecutive 2'-O-methyl-3'-phosphonoacetate (MP) or 2'-O-methyl-3'-thiophosphonoacetate (MSP) residues within the 3'-terminal five nucleotides.
6. 2. The method of claim 1, wherein the modified guide RNA further comprises modified nucleotides positioned outside of 5 nucleotides within the 5' and 3' ends.
7. 2. The method of claim 1, wherein the modified guide RNA is a single-stranded guide RNA.
8. 2. The method of claim 1, wherein the Cas protein is provided as an mRNA encoding the Cas protein.
9. 10. The method of claim 1, wherein the Cas protein and the modified guide RNA are provided as a ribonucleoprotein complex (RNP).
10. 10. The method of claim 1, wherein the Cas protein and / or modified guide RNA is provided in a nanoparticle.
11. 2. The method of claim 1, wherein the editing or modulation occurs with greater efficiency than with an unmodified gRNA that is otherwise identical to the modified guide RNA.
12. The method of claim 11 , wherein the efficiency is at least 10% higher.
13. The method of claim 1 , wherein the nuclease-containing liquid is serum.
14. 2. The method of claim 1, wherein the nuclease-containing fluid is cerebrospinal fluid (CSF).
15. The method of claim 1 , wherein the nuclease-containing liquid is a cell culture medium.
16. The method of claim 1 , wherein the nuclease-containing fluid is a body fluid.
17. The method of claim 1 , wherein the cell is in vivo.
18. The method of claim 1 , wherein the cells are present ex vivo in the presence of a nuclease-containing solution.