Engineered casphi2 (cas12j-2) proteins

Engineered CasPhi2 variants with specific mutations act as NTS nickases, addressing the size and delivery challenges of SpCas9, enabling efficient base and prime editing in human cells.

WO2025171041A1PCT designated stage Publication Date: 2025-08-14THE GENERAL HOSPITAL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014634
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2025-02-05
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing CRISPR-Cas9 systems, particularly SpCas9 variants, are large in size, making them difficult to deliver using size-limited vectors like adeno-associated virus (AAV) and challenging for producing larger mRNAs, and it is complex to develop CasPhi2 variants as TS or NTS nickases for base editing or prime editing due to the single RuvC domain active site.

Method used

Engineered CasPhi2 variants with specific mutations, such as T498K and S502K, are developed to function as NTS nickases, enabling efficient base and prime editing by reducing nuclease-induced gene editing efficiencies, and fusion proteins are created to enhance delivery and activity in human cells.

Benefits of technology

The engineered CasPhi2 variants demonstrate reduced nuclease-induced gene editing efficiencies, allowing for effective base and prime editing with smaller sizes suitable for viral and mRNA delivery methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000028_0001
    Figure IMGF000028_0001
  • Figure IMGF000019_0001
    Figure IMGF000019_0001
  • Figure IMGF000020_0001
    Figure IMGF000020_0001
Patent Text Reader

Abstract

Described herein are variants of CasPhi2 polypeptides with enhanced editing capabilities and methods of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Engineered CasPhi2 (Casl2j-2) Proteins

[0002] CLAIM OF PRIORITY

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 549,915, filed on February 5, 2024. The entire contents of the foregoing are hereby incorporated by reference.

[0004] FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0005] This invention was made with Government support under Grant Nos. GM118158 and HG009490 awarded by the National Institutes of Health. The Government has certain rights in the invention.

[0006] TECHNICAL FIELD

[0007] The present disclosure provides CasPhi2 polypeptides and fusion proteins harboring these proteins that can function for gene editing, e.g., as nickases, nucleases, as base editors or prime editors, and so on. The present disclosure provides systems, methods, and kits comprising such CasPhi2 polypeptides and fusion proteins.

[0008] BACKGROUND

[0009] CRISPR-Cas nucleases have provided a robust and simple-to-use platform for performing targeted gene editing in a wide variety of different organisms. Recently, nextgeneration CRISPR-based gene editor platforms have been described including CRISPR base editors (that induce targeted base substitutions, e.g., C-to-T or A-to-G) and CRISPR prime editors (that allow the installation of any base substitution(s) and also small insertions and deletions of user-specified lengthx). Base editors and prime editors both require the use of fusion proteins that harbor an engineered CRISPR-Cas nuclease variant that nicks one strand of DNA rather than inducing double-stranded cuts as is the case with the wild-type nucleases. Base editor activities are enhanced by nicking of the “target strand” (TS) of the DNA (i.e., the DNA strand in the Cas nuclease-induced R-loop that is complementary to the guide RNA) whereas prime editors work most efficiently with nicking of the “non-target strand” (NTS) of the DNA (i.e., the DNA strand in the R-loop i that is not complementary to the guide RNA), although recent work has suggested that prime editors that nick the target strand can also be used to introduce targeted alterations2. The most commonly used versions of base editors and prime editors utilize the CRISPR-Cas9 from Streptococcus pyogenes (SpCas9), which harbors two nuclease domains (RuvC and HNH domains), each of which nicks one of the two DNA strands in a target site in a specified fashion. For example, SpCas9 variants with mutations within the RuvC domain (e.g., at position DIO) function as TS nickases whereas those with mutations within the HNH domain (e.g., at position H840) function as NTS nickases.

[0010] SUMMARY

[0011] Provided herein are isolated CasPhi2 proteins, comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence of SEQ ID NO:1, and mutations at one, two, three, four, five or more of the following positions: A36, S106, D134, L149, E159, S160, S164, D167, El 68, P277, T355, T357, T518, L571, S616, D679, and / or Q684, and further comprising one or more mutations at R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E503, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, R678. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: L149, E159, S160, S164, D167, and E168. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: A36R, S106R, D134R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: A36R, S106R, D134R, L149R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: A36K, S106K, D134K, P277K, D337K, T355R, T357K, V531R, T539A, A543K, L571K, S616K, D679K, and T691K. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: A36K, S106K, D134K, P277K, D337K, T355R, T357K, V531R, T539A, A543K, L571K, S616K, D679K, Q684R, and T691K. In some embodiments, the isolated CasPhi2 protein comprises the following mutations: A36R, S106R, D134R, L149R, E159A, S160A, S164A, D167K, El 68 A, P277R, T357K, T518R, L571K, S616R, Q684R, T355R, and D679K.

[0012] In some embodiments, the isolated CasPhi2 protein further comprises a mutation at F517, wherein the mutation at position F517 is F517L, F517A or F517W.

[0013] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following positions: S496, N497, T498, T499, S502, E503, L506, S509, and / or S511. In some instances, the mutations are S496K, N497K, T498K, T499K, S502K, E503K, L506K, S509K, and / or S511K.

[0014] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following positions: K522, K523, K526, K527, and / or K528. In some instances, the positions are mutated to either a glutamate (E) or glycine (G).

[0015] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following position: R678. In some instances, the positions are mutated to either an alanine (A) or glycine (G).

[0016] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following mutations: T498K, S502K, R678A, and / or R678G.

[0017] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following positions: R437, W440, D441, R442, E444, E445, E446, R448, R450. In some instances, the CasPhi2 protein comprises the following mutations: R437A, D441A, and / or E444A.

[0018] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following positions: K618, R619, K630, and / or K631. In some instances, the positions are mutated to either an alanine (A) or glycine (G). In some instances, the CasPhi2 protein comprises one of the following sets of mutations: (1) K618G / R619G or (2) K630G / K631G.

[0019] In some embodiments, the isolated CasPhi2 protein further comprises one or more mutations at the following positions: T498, S502, K618 / R619, or K630 / K631. In some instances, the isolated CasPhi2 protein comprises one of the following sets of mutations: T498K, S502K, K618G / R619G, or K630G / K631G.

[0020] Also provided herein are fusion proteins comprising any one or more of the isolated CasPhi2 proteins as described above, fused to at least one heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

[0021] In some embodiments, the heterologous functional domain is a biological tether. In some instances, the biological tether is MS2, Csy4 or lambda N protein.

[0022] In some embodiments, the heterologous functional domain is Fokl.

[0023] In some embodiments, the heterologous functional domain is a deaminase. In some instances, the heterologous functional domain is a cytidine deaminase. In some instances, the cytidine deaminase is selected from the group consisting of APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, AP0BEC3H, APOBEC4, activation-induced cytidine deaminase (AID), cytosine deaminase 1 (CDA1), pmCDAl, CDA2, and cytosine deaminase acting on tRNA (CDAT). In some instances, the heterologous functional domain is an adenosine deaminase. In some instances, the adenosine deaminase is selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA).

[0024] In some embodiments, the heterologous functional domain is a reverse transcriptase (e.g., MMLV-RT pentamutant ARNaseH variant).

[0025] In some embodiments, the fusion proteins comprise at least two heterologous functional domains, wherein the additional heterologous functional domain comprises an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways. In some instances, the additional heterologous functional domain is a uracil DNA glycosylase inhibitor (UGI) that inhibits uracil DNA glycosylase (UDG, also known as uracil N-glycosylase, or UNG); or Gam from the bacteriophage Mu.

[0026] Also provided herein are isolated nucleic acids encoding the isolated CasPhi2 proteins as described above, or the fusion proteins as described above. Also provided are vectors comprising these isolated nucleic acids. Also provided are isolated host cells comprising these nucleic acids. In some instances, the host cell is a mammalian host cell. Also provided herein are complexes for prime editing comprising: (a) any of the isolated CasPhi2 proteins as described above, (b) a domain comprising an RNA- dependent DNA polymerase activity; and (b) a pegRNA. In some instances, parts (a) and (b) are expressed together as a single fusion protein from a single plasmid. In other instances, parts (a) and (b) are expressed from two separate plasmids. In some instances, the domain comprising an RNA-dependent DNA polymerase activity is a reverse transcriptase (e.g., MMLV-RT pentamutant ARNaseH variant). In some instances, the pegRNA comprises a primer binding sequence (PBS), a reverse transcriptase template (RTT), one or more crRNAs, and a modified U6 stem loop sequence appended to the 5’ end of the pegRNA. In some instances, the pegRNA further comprises a linker sequence between the PBS+RTT and the one or more crRNAs.

[0027] Also provided are methods for prime editing, the methods comprising: expressing in a cell any of the complexes as described above.

[0028] Also provided are methods of altering a genome of a cell, the methods comprising expressing in the cell, or contacting the cell with, any of the isolated CasPhi2 proteins as described above or any of the fusion proteins as described above, and one or more crRNAs, wherein the one or more crRNAs direct any of the isolated CasPhi2 proteins as described above or any of the fusion proteins as described above to one or more target genomic sequences. In some instances, the cell is a stem cell. In some instances, the stem cell is an embryonic stem cell, a mesenchymal stem cell, or an induced pluripotent stem cell; is in a living animal; or is in or is an embryo.

[0029] Also provided are methods of altering a double stranded DNA (dsDNA) molecule, the methods comprising contacting the dsDNA with any of the isolated CasPhi2 proteins as described above or any of the fusion proteins as described above, and one or more crRNAs, wherein the one or more crRNAs direct any of the isolated CasPhi2 proteins as described above or any of the fusion proteins as described above to one or more target genomic sequences. In some instances, the dsDNA molecule is in vitro.

[0030] Also described herein are CasPhi2 variants comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence of SEQ ID NO: 1 and comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E503, L506, S509, S51 1, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1.

[0031] In some embodiments, the CasPhi2 variant comprises the mutation at a position selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, T498, T499, S502, E503, K618, R619, K630, K631, and R678.

[0032] In some embodiments, the CasPhi2 variant comprises the mutation at a position selected from the group consisting of: F517, R437, D441, E444, T498, S502, K618, R619, K630, K631 and R678.

[0033] In some embodiments, the CasPhi2 variant comprises the mutation at position F517. In some embodiments, the F517 mutation is F517L, F517A or F517W. In some embodiments, the F517 mutation is F517A.

[0034] In some embodiments, the CasPhi2 variant comprises the mutation at a position selected from the group consisting of: T498, S502, and R678. In some embodiments, the mutations are selected from T498K, S502K, R678A, or R678G.

[0035] In some embodiments, the CasPhi2 variant comprises a mutation at a position selected from R437, D441, or E444. In some embodiments, the mutations are R437A, D441A, or E444A.

[0036] In some embodiments, the CasPhi2 variant comprises at least two mutations at positions selected from K618, R619, K630, or K631. In some embodiments, the at least two mutations comprise K618G and R619G. In some embodiments, the at least two mutations comprise K630G and K631G.

[0037] Also described herein are fusion proteins comprising the CasPhi2 variant comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E503, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1, fused to at least one heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

[0038] In some embodiments, the at least one heterologous functional domain comprises: a biological tether, a FokI, a deaminase, or a reverse transcriptase.

[0039] In some embodiments, the biological tether MS2, Csy4 or lambda N protein. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the cytidine deaminase is selected from the group consisting of APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced cytidine deaminase (AID), cytosine deaminase 1 (CDA1), pmCDAl, CDA2, and cytosine deaminase acting on tRNA (CD AT).

[0040] In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the adenosine deaminase is selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (ADAR1), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA).

[0041] In some embodiments, the reverse transcriptase is a MMLV-RT pentamutant ARNaseH variant.

[0042] In some embodiments, the fusion proteins comprise at least two heterologous functional domains, wherein an additional heterologous functional domain comprises an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways.

[0043] In some embodiments, the the additional heterologous functional domain is a uracil DNA glycosylase inhibitor (UGI) or Gam from the bacteriophage Mu.

[0044] Also described herein are isolated nucleic acids encoding the CasPhi2 variant described herein or the fusion proteins described herein. Additionally described herein are vector comprising the isolated nucleic acid. Also described are isolated host cells comprising the nucleic acid. In some embodiments, the host cell is a mammalian host cell.

[0045] Also described are complexes for prime editing comprising: (a) the CasPhi2 variant comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E503, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1, (b) a domain comprising an RNA- dependent DNA polymerase activity; and (c) a pegRNA. In some embodiments, parts (a) and (b) expressed together as a single fusion protein from a single plasmid.

[0046] In some embodiments, parts (a) and (b) are expressed from two separate plasmids.

[0047] In some embodiments, the domain is a reverse transcriptase. In some embodiments, the reverse transcriptase is a MMLV-RT pentamutant ARNaseH variant.

[0048] In some embodiments, the pegRNA comprises a primer binding sequence (PBS), a reverse transcriptase template (RTT), one or more crRNAs, and a modified U6 stem loop sequence appended to the 5’ end of the pegRNA. In some embodiments, the PegRNA further comprises a linker sequence between the PBS+RTT and the one or more crRNAs.

[0049] Also described herein are methods of altering a genome of a cell, the method comprising expressing in the cell, or contacting the cell with, the CasPhi2 variant comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E503, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1, and one or more crRNAs, wherein the one or more crRNAs direct the CasPhi2 variant to one or more target genomic sequences, wherein the CasPhi2 variant nicks a strand of the one or more target genomic sequences.

[0050] A method of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA with the CasPhi2 variant comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E5O3, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1, or the fusion protein comprising the CasPhi variant, and one or more crRNAs, wherein the one or more crRNAs direct the CasPhi2 variant or the fusion protein to one or more target genomic sequences, wherein the CasPhi2 variant nicks a strand of the one or more target genomic sequences. In some embodiments, the CasPhi2 variant is a nickase. In some embodiments, the strand is a non-target strand. In some embodiments, the strand is a target strand.

[0051] In some instances of any one of the preceding embodiments, the method is performed in vivo. In some instances of any one of the preceding embodiments, the method is performed in vitro.

[0052] In some instances of any one of the preceding embodiments, the method is performed ex vivo.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0054] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.

[0055] DESCRIPTION OF DRAWINGS

[0056] Fig 1A-1C. Testing the indel frequencies induced by CasPhi2-17AA and the potential NTS nickase mutants CasPhi2-17AAF517X (X=L, A, or W) using crRNA with longer spacers (20, 24, 25, 26, or 27nts) in HEK293T cells (n=l).’ No treatment’ sample was used as a negative control. Indel frequencies were determined by targeted amplicon sequencing at (A) IL2RA site 31 (B) TRAC site 19, and (C) VEGFA site 3.1. Mutants with potential nickase activity will show lower indel frequencies since only the NTS strand is cut. CasPhi2-17AA mutants with low indel frequencies are labeled in bold.

[0057] Fig 2A-2D. Close-up view of CasPhi2 ternary complex structure (PDB: 7LYT) (A) showing capping residue (F517), residues at a-helix 17 (S496, N497, T498, T499, S502, E503, L506, S509, S511), lysine-rich protein loop (K522, K523, K526, K527, K528), and (B) near active site (R678) that was mutated to prepare potential NTS nickase variants of CasPhi-17AA. (C) Close-up view of residues located near the crRNA spacer region (K618, R619, K630, and K631) that were mutated to prepare potential TS nickase variants of CasPhi2-17AA. (D) Heat map showing indel frequencies of CasPhi2-17AA variants screened for NTS nickase activity at three genomic target sites in HEK293T cells (n=l). Indel frequencies were determined by targeted amplicon sequencing and ‘no treatment’ was used as a negative control. Mutants with potential nickase activity will show lower indel frequencies since only the NTS strand is cut. CasPhi2-17AA mutants with low indel frequencies are labeled with an asterisk (*).

[0058] Fig 3. Testing gene editing activity of CasPhi2-17AA mutants where 9 residues within a-helix 14 were individually mutated to alanine at 3 genomic loci in HEK293T cells (n=l). CasPhi2-17AAand ‘no treatment’ were used as controls. Indel frequencies were determined by targeted amplicon sequencing of the genomic loci. CasPhi2-17AA mutants with low indel frequencies are labeled with an asterisk (*).

[0059] Fig 4A-4B. Testing prime editing activity of the split-PEs based on potential CasPhi2-17AA NTS nickase variants and MMLV-RT A RNase H, along with pegRNAs harboring a modified U6 stem-loop at the target site VEGFA site 3.1 in HEK293T cells (n=l). Split PEs based on CasPhi2-17AA and potential TS nickase variants (CasPhi2- 17AA K618G / R619G and K630G / K631G) and ‘no treatment’ were included as controls. (A) Predicted secondary structures of U6 stem-loop and modified U6 stem-loop added to pegRNAs at the 5’ end (Top left). Diagrams (Top right) and example schematic maps showing pUC 19-based U6 expression vector of pegRNAs used in the experiment (Bottom). We tested two different pegRNAs with 15 nt spacer designed to introduce the intended ATG’ insertion at position 12 or 13 (P12 or P13). SpCas9 tracrRNA sequence was utilized as a linker. Sequences shown in FIG. 4A: U6 promoter: 5’- GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGT TAGAGAGATAATTAGAATTAATTTGACTGTAAACACAAAGATATTAGTACAA AATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATT ATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATT TCTTGGCTTTATATATCTTGTGGAAAGGACGAAAC ACC-3’ (SEQ ID NO: 18); Modified U6 stem loop: 5’-GTGCTGCTTCGGCAGCAC-3’ (SEQ ID NO: 19); ATG insertion at P12: 5’-ACCCCTGGCCCATTTCTCCCCGCT-3’ (SEQ ID NO:20); ATG insertion at P13: 5’-GACCCCTGGCCATCTTCTCCCCGC-3’ (SEQ ID NO:21); SpCas9 tracrRNA Scaffold: 5’-GTTTTAGAGCTAGAAATAGCAAGTTAAAAT AAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3 ’ (SEQ ID NO:22); CasPhi2 DR: 5’-CAACGATTGCCCCTCACGAGGGGAC-3’ (SEQ ID NO 23); and Spacer 15nt: 5’-GAGCGGGGAGAAGGC-3’ (SEQ ID NO:23). (B) Prime editing efficiencies and indel frequencies observed from targeted amplicon sequencing of VEGFA site 3.1.

[0060] DETAILED DESCRIPTION

[0061] One disadvantage of SpCas9 and SpCas9 variants is their large size (1368 amino acids), which makes it difficult to encode these proteins in size-limited delivery vectors such as adeno associated virus (AAV) (which has a cargo size limit of 4.7 kb) and also poses challenges for the production of larger size mRNAs. Recent work has described the existence of smaller CRISPR-Cas nucleases, which can potentially overcome the disadvantages of the larger size SpCas9 nuclease. One recent example of such a hypercompact CRISPR-Cas nuclease is the bacteriophage CasPhi2 (Casl2j-2), which is substantially smaller (757 amino acids) than the SpCas9 protein (1368 amino acids). In contrast to previously published reports3 4, CasPhi2 showed very little to no nuclease- induced gene-editing activity in human cells (see WO / 2024 / 086845). To overcome this, CasPhi2 variants were evolved bearing up to seventeen amino acid mutations that exhibited robust and highly efficient nuclease-induced gene editing activities in human cells (see WO / 2024 / 086845).

[0062] Base editors and prime editors built using hypercompact CRISPR-Cas nuclease variants such as CasPhi2 would be smaller in size, which as noted above has potential advantages for viral and mRNA delivery methodologies. However, unlike SpCas9 which is a CRISPR type II nuclease harboring two nuclease domains that uses two different active sites to cut the TS and NTS, respectively, CasPhi2 is a type V nuclease that uses a single RuvC domain active site to cut the NTS followed by TS3. Therefore, it is more challenging and less straightforward to develop a CasPhi2 TS or NTS nickase that can be used for base editing or prime editing since simply mutating any of the RuvC active site residues would be expected to hamper cleavage of both strands.

[0063] Here, we report the rational structure-guided and other engineering of CasPhi2- 17AA (a CasPhi2 variant bearing 17 amino acid mutations that we recently evolved to have high nuclease-induced gene editing activity in human cells) with the intention of creating TS or NTS nickase versions of this enzyme that could be used for base editing and prime editing in human cells. In a human cell-based screen, we identified two mutants of CasPhi2-17AA (T498K and S502K) engineered to have NTS nicking activity that showed reduced nuclease-induced gene editing efficiencies, consistent with reduced cleavage of one of the two DNA strands. Indeed, when we built prime editors using these two mutants, we observed prime editing activity with substantially reduced nuclease- induced indels relative to matched prime editors harboring CasPhi2-17AA nuclease suggesting that the two variants are behaving as NTS nickases. These experiments provide an important proof-of-principle for creating CasPhi2 nickase variants that can be used to create nickase-based gene editors that are active in human cells.

[0064] Engineered CasPhi2 Variants

[0065] Provided herein are CasPhi2 variants. The CasPhi2 wild type sequence is as follows (GenBank Accession No. 7LYS A; Pausch P, Soczek KM, Herbst DA, Tsuchida CA, Al-Shayeb B, Banfield JF, Nogales E, Doudna JA. DNA interference states of the hypercompact CRISPR-Cas effector. Nat Struct Mol Biol. 2021 Aug;28(8):652-661):

[0066] 1 MPKPAVESEF SKVLKKHFPG ERFRSSYMKR GGKI LAAQGE EAVVAYLQGK SEEE PPNFQP 61 PAKCHVVTKS RDFAEWPIMK ASEAIQRYIY ALSTTERAAC KPGKSSESHA AWFAATGVSN

[0067] 121 HGYSHVQGLN LI FDHTLGRY DGVLKKVQLR NEKARARLES INASRADEGL PE IKAEEEEV 181 ATNETGHLLQ PPGTNPS FYV YQTI SPQAYR PRDEIVLPPE YAGYVRDPNA PI PLGVVRNR 241 CDIQKGCPGY I PEWQREAGT AIS PKTGKAV TVPGLS PKKN KRMRRYWRSE KEKAQDALLV 301 TVRI GTDWVV IDVRGLLRNA RWRTIAPKDI SLNALLDLFT GDPVIDVRRN IVTFTYTLDA 361 CGTYARKWTL KGKQTKATLD KLTATQTVAL VAIDLGQTNP I SAGI SRVTQ ENGALQCE PL

[0068] 421 DRFTLPDDLL KDI SAYRIAW DRNEEELRAR SVEALPEAQQ AEVRALDGVS KETARTQLCA 481 DFGLDPKRLP WDKMSSNTTF I SEALLSNSV SRDQVFFTPA PKKGAKKKAP VEVMRKDRTW 541 ARAYKPRLSV EAQKLKNEAL WALKRTS PEY LKLSRRKEEL CRRS INYVIE KTRRRTQCQI 601 VI PVIEDLNV RFFHGSGKRL PGWDNFFTAK KENRWFIQGL HKAFSDLRTH RS FYVFEVRP 661 ERTS ITCPKC GHCEVGNRDG EAFQCLSCGK TCNADLDVAT HNLTQVALTG KTMPKREE PR 721 DAQGTAPARK TKKASKSKAP PAEREDQTPA QE PSQTS ( SEQ ID NO : 1 )

[0069] The CasPhi2 variants described herein can include mutations at one or more of the following positions: T355 and / or D679 (or at positions analogous thereto). In some embodiments, the CasPhi2 variants described herein can include a mutation at T355. In some embodiments, the CasPhi2 variants described herein can include a mutation at D679. In some embodiments, the CasPhi2 variants described herein can include mutations at T355 and D679. In some embodiments, the mutation at T335 is T355R or T355K. In some embodiments, the mutation at D679 is D679R, D679K, D679H, or D679T.

[0070] In some embodiments, the CasPhi2 variants include mutations at all of the following positions: A36, S106, D134, L149, E159, S160, S164, D167, E168, P277, T355, T357, T518, L571, S616, D679, and Q684. In some instances, the mutations are: A36R, S106R, D134R, L149R, E159A, S160A, S164A, D167K, E168A, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R (referred to herein after as the “17 amino acid” or “17AA” CasPhi2 variant).

[0071] CasPhi proteins are described in WO2022159822; the CasPhi proteins can include one or more mutations or modifications; in some embodiments, the CasPhi proteins comprise the following mutations: A36R, S106R, D134R, L149R, E159A, S160A, S164A, D167K, E168A, P277R, T357K, T518R, L571K, S616R, Q684R, T355R, and D679K. In some embodiments, the CasPhi proteins comprise the following mutations: A36R, S106R, D134R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R, optionally further comprising a mutation at one or more of the following positions: Si l, S25, G138, T203, A261, D337, N497, L506, S507, N508, S509, D513, Q514, A520, G524, A525, K527, P530, V531, R538, T539, R542, A543, E569, E578, T628, T649, E674, and / or T691, optionally further comprising the following mutations: F23S and S26R, or optionally further comprising the following mutations: T340G, D341R, and D342G. In some embodiments, the CasPhi proteins comprise the following mutations: A36R, S106R, D134R, L149R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R. In some embodiments, the CasPhi proteins comprise the following mutations: A36K, S106K, D134K, P277K, D337K, T355R, T357K, V531R, T539A, A543K, L571K, S616K, D679K, and T691K, optionally further comprising the following mutation: Q684R. In some embodiments, the CasPhi proteins comprise a mutation that catalytically inactivates nuclease activity, wherein the mutation is D394A or E606Q. The numbering is relative to the CasPhi2 wild type sequence (GenBank Accession No. 7LYS_A; Pausch P, Soczek KM, Herbst DA, Tsuchida CA, Al-Shayeb B, Banfield JF, Nogales E, Doudna JA. DNA interference states of the hypercompact CRISPR-Cas effector. Nat Struct Mol Biol. 2021 Aug;28(8):652-661).

[0072] The CasPhi proteins need not be active nucleases or nickases, but can be part of a fusion protein, e.g., a base editor or prime editor.

[0073] In some embodiments, the CasPhi2 variant is a target strand (TS) nickase. In some embodiments, the TS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes a mutation at position F517. In some instances, the mutation at position F517 is F517L, F517A or F517W. In some instances, the mutation at position F517 is F517A.

[0074] In some embodiments, the CasPhi2 variant is a non-target strand (NTS) nickase. In some embodiments, the NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following positions: S496, N497, T498, T499, S502, E503, L506, S509, and / or S511. In some instances, the mutation at one or more of these positions is a lysine (K). In some embodiments, the NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following positions: K522, K523, K526, K527, and / or K528. In some instances, the mutation at one or more of these positions is a glutamate (E) or glycine (G). In some embodiments, the NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following position: R678. In some instances, the mutation at one or more of these positions is alanine (A) or glycine (G)). In some embodiments, the NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more of the following mutations: T498K, S502K, R678A, and / or R678G.

[0075] In some embodiments, the CasPhi2 variant is a kinetic non-target strand (kNTS) nickase. In some embodiments, the kinetic NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following positions: R437, W440, D441, R442, E444, E445, E446, R448, R450. In some instances, the mutation at one or more of these positions is alanine (A). In some embodiments, the kinetic NTS nickase comprises the mutations of the 17AA CasPhi2 variant and further includes one or more of the following mutations: R437A, D441A, and / or E444A. In some embodiments, the CasPhi2 variant is a target strand (TS) nickase that can be used to construct more efficient base editors as well as prime editor proteins that function by nicking the target strand. In some embodiments, the CasPhi2 variant comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following positions: K618, R619, K630, and / or K631. In some instances, the mutation at one or more of these positions is glycine (G). In some instances, the CasPhi2 variant comprises the mutations of the 17AA CasPhi2 variant and further includes one of the following sets of mutations: (1) K618G / R619G or (2) K630G / K631G.

[0076] In some embodiments, the CasPhi2 variant is a NTS nickase that can be used to construct more efficient base editors as well as prime editor proteins that function by nicking the non-target strand. In some embodiments, the CasPhi2 variant comprises the mutations of the 17AA CasPhi2 variant and further includes one or more mutations at the following positions or sets of positions: T498, S502, K618 / R619, or K630 / K631. In some instances, the CasPhi2 variant comprises the mutations of the 17AA CasPhi2 variant and further includes one or more of the following mutations or sets of mutations: T498K, S502K, K618G / R619G, or K630G / K631G.

[0077] In some embodiments, the CasPhi2 variants are at least 70%, e.g., at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the amino acid sequence of SEQ ID NO: 1, e g., have differences at up to 5%, 10%, 15%, 20%, 25%, or 30% of the amino acid residues of SEQ ID NO: 1 replaced, e.g., with conservative mutations, in addition to mutations described herein. In preferred embodiments, the variant retains or has improved desired activity of the parent, e.g., the nickase activity.

[0078] To determine the percent identity of two nucleic acid sequences, the sequences are aligned for optimal comparison purposes (e.g., gaps can be introduced in one or both of a first and a second amino acid or nucleic acid sequence for optimal alignment and non-homologous sequences can be disregarded for comparison purposes). The length of a reference sequence aligned for comparison purposes is at least 80% of the length of the reference sequence, and in some embodiments is at least 90% or 100%. The nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, then the molecules are identical at that position (as used herein nucleic acid “identity” is equivalent to nucleic acid “homology”). The percent identity between the two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps, and the length of each gap, which need to be introduced for optimal alignment of the two sequences. Percent identity between two polypeptides or nucleic acid sequences is determined in various ways that are within the skill in the art, for instance, using publicly available computer software such as Smith Waterman Alignment (Smith, T. F. and M. S. Waterman (1981) J Mol Biol 147: 195-7); “BestFit” (Smith and Waterman, Advances in Applied Mathematics, 482-489 (1981)) as incorporated into GeneMatcher Plus™, Schwarz and Dayhof (1979) Atlas of Protein Sequence and Structure, Dayhof, M.O., Ed, pp 353-358; BLAST program (Basic Local Alignment Search Tool; (Altschul, S. F., W. Gish, et al. (1990) J Mol Biol 215: 403-10), BLAST-2, BLAST-P, BLAST-N, BLAST-X, WU- BLAST-2, ALIGN, ALIGN-2, CLUSTAL, or Megalign (DNASTAR) software. In addition, those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared. In general, for proteins or nucleic acids, the length of comparison can be any length, up to and including full length (e.g., 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%). For purposes of the present compositions and methods, at least 80% of the full length of the sequence is aligned using the BLAST algorithm and the default parameters.

[0079] For purposes of the present invention, the comparison of sequences and determination of percent identity between two sequences can be accomplished using a Blossum 62 scoring matrix with a gap penalty of 12, a gap extend penalty of 4, and a frameshift gap penalty of 5.

[0080] Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. Fusions including CasPhi2 variants

[0081] In addition, the variants described herein can be used in fusion proteins in place of the wild-type CasPhi2 or other CasPhi2 mutants as known in the art, e.g., a fusion protein with a heterologous functional domains as described in US 8,993,233; US 20140186958; US 9,023,649; WO / 2014 / 099744; WO 2014 / 089290; WO2014 / 144592; WO144288; WO2014 / 204578; WO2014 / 152432; WO2115 / 099850; US8,697,359; US2010 / 0076057; US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989;

[0082] WO / 2010 / 054108; WO / 2012 / 164565; WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US 20150071899 and WO 2014 / 124284.

[0083] For example, the CasPhi2 variants, can be fused to a heterologous functional domain on the N- terminus or C- terminus. In some embodiments, the CasPhi2 variant can have a heterologous functional domain that is inlaid within the protein (i.e., internally inserted).

[0084] In some embodiments, the heterologous functional domain is a base editor, e.g., a deaminase that modifies cytosine DNA bases, e.g., a cytidine deaminase from the apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like (APOBEC) family of deaminases, including APOBEC 1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D / E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4 (see, e.g, Yang et al, J Genet Genomics. 2017 Sep 20;44(9):423-437); activation-induced cytidine deaminase (AID), e.g, activation induced cytidine deaminase (AICDA), cytosine deaminase 1 (CDA1), and CDA2, and cytosine deaminase acting on tRNA (CDAT). The following table provides exemplary sequences; other sequences can also be used.

[0085] Table 2: Exemplary Sequences of Base Editors

[0086] * from Saccharomyces cerevisiae S288C

[0087] * from sea lamprey (Petromyzon marinus) In some embodiments, the heterologous functional domain is a deaminase that modifies adenosine DNA bases, e.g., the deaminase is an adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3 (see, e g., Savva et al., Genome Biol. 2012 Dec 28; 13(12):252); adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3 (see Keegan et al., RNA. 2017 Sep;23(9): 1317-1328 and Schaub and Keller, Biochimie. 2002 Aug;84(8):791-803); and naturally occurring or engineered tRNA-specific adenosine deaminase (TadA) (see, e.g., Gaudelli et al., Nature. 2017 Nov 23 ;551 (7681):464-471) (NP_417054.2 (Escherichia coli str. K-12 substr. MG1655); See, e.g., Wolf et al., EMBO J. 2002 Jul 15;21 (14):3841 - 51. The following table provides exemplary sequences; other sequences can also be used.

[0088] Table 4: Exemplary Sequences of Deaminases In some embodiments, the heterologous functional domain is an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways, e.g., thymine DNA glycosylase (TDG; GenBank Acc Nos. NM_003211.4 (nucleic acid) and NP_003202.3 (protein)) or uracil DNA glycosylase (UDG, also known as uracil N-glycosylase, or UNG; GenBank Acc Nos. NM 003362.3 (nucleic acid) and NP_003353.1 (protein)) or uracil DNA lycosylase inhibitor (UGI) that inhibits UNG mediated excision of uracil to initiate BER (see, e.g., Mol et al., Cell 82, 701-708 (1995); Komor et al., Nature. 2016 May 19;533(7603)); or DNA endbinding proteins such as Gam, which is a protein from the bacteriophage Mu that binds free DNA ends, inhibiting DNA repair enzymes and leading to more precise editing (less unintended base edits; Komor et al., Sci Adv. 2017 Aug 30;3(8):eaao4774).

[0089] In some embodiments, all or part of the protein, e.g., at least a catalytic domain that retains the intended function of the enzyme, can be used.

[0090] In some embodiments, the heterologous functional domain is a biological tether, and comprises all or part of (e.g., DNA binding domain from) the MS2 coat protein, endoribonuclease Csy4, or the lambda N protein. These proteins can be used to recruit RNA molecules containing a specific stem-loop structure to a locale specified by the CasPhi2 variant gRNA targeting sequences. For example, a CasPhi2 variant fused to MS2 coat protein, endoribonuclease Csy4, or lambda N can be used to recruit a long noncoding RNA(lncRNA) such as XIST or HOTAIR; see, e g., Keryer-Bibens et al., Biol. Cell 100: 125-138 (2008), that is linked to the Csy4, MS2 or lambda N binding sequence. Alternatively, the Csy4, MS2 or lambda N protein binding sequence can be linked to another protein, e.g., as described in Keryer-Bibens et al., supra, and the protein can be targeted to the CasPhi2 variant binding site using the methods and compositions described herein. In some embodiments, the Csy4 is catalytically inactive. In some embodiments, the CasPhi2 variant, preferably a CasPhi2 variant, is fused to FokI as described in US 8,993,233; US 20140186958; US 9,023,649; WO / 2014 / 099744; WO 2014 / 089290; WO2014 / 144592; WO144288; WO2014 / 204578; WO2014 / 152432; WO2115 / 099850; US8,697,359; US2010 / 0076057; US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565;

[0091] WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US 20150071899 and WO 2014 / 204578.

[0092] In some embodiments, the fusion proteins include a linker between the CasPhi2 variant and the heterologous functional domains. Linkers that can be used in these fusion proteins (or between fusion proteins in a concatenated structure) can include any sequence that does not interfere with the function of the fusion proteins. In preferred embodiments, the linkers are short, e.g., 2-40 amino acids, and are typically flexible (i.e., comprising amino acids with a high degree of freedom such as glycine, alanine, and serine). In some embodiments, the linker comprises one or more units consisting of GGGS (SEQ ID NO:2) or GGGGS (SEQ ID NOG), e.g., two, three, four, or more repeats of the GGGS (SEQ ID NOG) or GGGGS (SEQ ID NOG) unit. In some embodiments, the linker comprises an XTEN linker (e.g., a 32 amino acid modified XTEN linker (flanked with extended GlySer linkers on both sides)). Other linker sequences can also be used (see Table 4).

[0093] Table 4: Different linkers used to fuse dCasPhi2-17AA to deaminase domains Prime Editors

[0094] Prime editors are another example, which include a Cas protein fused to an engineered reverse transcriptase that is paired with a prime editing guide RNA (pegRNA) that specifies the target site and encodes the desired edit (Anzalone et al., Nature volume 576, pagesl49-157 (2019)). For example, in some instances any of the CasPhi variants described herein are fused to an engineered reverse transcriptase. In other instances, the reverse transcriptase and CasPhi are expressed from two separate plasmids. Examples of reverse transcriptases include those comprising or derived from any of the following: Eubacterium rectale RT (aka Marathon-RT), Human endogenous retrovirus K consensus (HERV-Kcon) RT, Geobacillus stearotherniophilus GsI-IIC RT (WT), Geobacillus stearothermophilus GsI-IIC intron RT (GsI-IIC RT), MMLV-RT (MMLV-RT pentamutant ARNaseH variant). The present compositions and methods can make use of variants as known in the art and as provided herein, e.g., MarathonRT, GsI-IIC RT, and MMLV-RT 20 variants, e.g., PE2 MMLV RT (with D200N, T306K, W313F, T330P, L603W mutations), or MMLV or PE2 MMLV RT truncations (truncations 2, 5, and 6; Griinewald et al. Nat Biotechnol. 2023 Mar;41(3):337-343), as well as RT HFV, HERV, LtrA, HERV-Kcon, Tel4c, Marathon, GsI-IIC, Ma-Int5, engineered Marathon (optionally with D14R, N26R, D74R, N116K, or N197R mutations), etc.

[0095] GPPs

[0096] In some embodiments, the variant protein includes a cell-penetrating peptide sequence that facilitates delivery to the intracellular space, e.g., HIV-derived TAT peptide, penetratins, transportans, or hCT derived cell-penetrating peptides, see, e.g., Caron et al., (2001) Mol Ther. 3(3):310-8; Langel, Cell-Penetrating Peptides: Processes and Applications (CRC Press, Boca Raton FL 2002); El-Andaloussi et al., (2005) Curr Pharm Des. 11(28):3597-611 ; and Deshayes et al., (2005) Cell Mol Life Sci. 62(16): 1839-49.

[0097] Cell penetrating peptides (CPPs) are short peptides that facilitate the movement of a wide range of biomolecules across the cell membrane into the cytoplasm or other organelles, e.g., the mitochondria and the nucleus. Examples of molecules that can be delivered by CPPs include therapeutic drugs, plasmid DNA, oligonucleotides, siRNA, peptide-nucleic acid (PNA), proteins, peptides, nanoparticles, and liposomes. CPPs are generally 30 amino acids or less, are derived from naturally or non-naturally occurring protein or chimeric sequences, and contain either a high relative abundance of positively charged amino acids, e.g., lysine or arginine, or an alternating pattern of polar and nonpolar amino acids. CPPs that are commonly used in the art include Tat (Frankel et al., (1988) Cell. 55: 1189-1193, Vives et al., (1997) J. Biol. Chem. 272: 16010-16017), penetratin (Derossi et al., (1994) J. Biol. Chem. 269: 10444-10450), polyarginine peptide sequences (Wender et al., (2000) Proc. Natl. Acad. Sci. USA 97: 13003-13008, Futaki et al., (2001) J. Biol. Chem. 276:5836-5840), and transportan (Pooga et al., (1998) Nat. Biotechnol. 16:857-861).

[0098] CPPs can be linked with their cargo through covalent or non-covalent strategies. Methods for covalently joining a CPP and its cargo are known in the art, e.g., chemical cross-linking (Stetsenko et al., (2000) J. Org. Chem. 65:4900-4909, Gait et al. (2003) Cell. Mol. Life. Sci. 60:844-853) or cloning a fusion protein (Nagahara et al., (1998) Nat. Med. 4:1449-1453). Non-covalent coupling between the cargo and short amphipathic CPPs comprising polar and non-polar domains is established through electrostatic and hydrophobic interactions.

[0099] CPPs have been utilized in the art to deliver potentially therapeutic biomolecules into cells. Examples include cyclosporine linked to polyarginine for immunosuppression (Rothbard et al., (2000) Nature Medicine 6(11): 1253-1257), siRNA against cyclin Bl linked to a CPP called MPG for inhibiting tumorigenesis (Crombez et al., (2007) Biochem Soc. Trans. 35:44-46), tumor suppressor p53 peptides linked to CPPs to reduce cancer cell growth (Takenobu et al., (2002) Mol. Cancer Then 1(12): 1043-1049, Snyder et al., (2004) PLoS Biol. 2:E36), and dominant negative forms of Ras or phosphoinositol 3 kinase (PI3K) fused to Tat to treat asthma (Myou et al., (2003) J. Immunol. 171 :4399- 4405).

[0100] CPPs have been utilized in the art to transport contrast agents into cells for imaging and biosensing applications. For example, green fluorescent protein (GFP) attached to Tat has been used to label cancer cells (Shokolenko et al., (2005) DNA Repair 4(4):511-518). Tat conjugated to quantum dots have been used to successfully cross the blood-brain barrier for visualization of the rat brain (Santra et al., (2005) Chem. Commun. 3144-3146). CPPs have also been combined with magnetic resonance imaging techniques for cell imaging (Liu et al., (2006) Biochem. And Biophys. Res. Comm. 347(1): 133-140). See also Ramsey and Flynn, Pharmacol Ther. 2015 Jul 22. Pii: S0163- 7258(15)00141-2.

[0101] In some embodiments, alternatively or in addition, the variant proteins can include a nuclear localization sequence, e.g., SV40 large T antigen NLS (PKKKRRV (SEQ ID NO: 11)) and nucleoplasmin NLS (KRPAATKKAGQAKKKK (SEQ ID NO: 12)). Other NLSs are known in the art; see, e.g., Cokol et al., EMBO Rep. 2000 Nov 15; 1(5): 411-415; Freitas and Cunha, Curr Genomics. 2009 Dec; 10(8): 550-557.

[0102] In some embodiments, the variants include a moiety that has a high affinity for a ligand, for example GST, FLAG or hexahistidine sequences. Such affinity tags can facilitate the purification of recombinant variant proteins.

[0103] For methods in which the variant proteins are delivered to cells, the proteins can be produced using any method known in the art, e.g., by in vitro translation, or expression in a suitable host cell from nucleic acid encoding the variant protein; a number of methods are known in the art for producing proteins. For example, the proteins can be produced in and purified from yeast, E. coli, insect cell lines, plants, transgenic animals, or cultured mammalian cells; see, e.g., Palomares et al., “Production of Recombinant Proteins: Challenges and Solutions,” Methods Mol Biol. 2004;267: 15-52. In addition, the variant proteins can be linked to a moiety that facilitates transfer into a cell, e.g., a lipid nanoparticle, optionally with a linker that is cleaved once the protein is inside the cell. See, e.g., LaFountaine et al., Int J Pharm. 2015 Aug 13;494(1): 180-194.

[0104] Methods of Use

[0105] The variants described herein can be used for altering the genome of a cell; the methods generally include expressing the variant proteins in the cells, along with a guide RNA (or crRNA) having a region complementary to a selected portion of the genome of the cell. Methods for selectively altering the genome of a cell are known in the art, see, e.g., US8,697,359; US2010 / 0076057; US2011 / 0189776; US2011 / 0223638; US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565;

[0106] WO / 2013 / 098244; WO / 2013 / 176772; US20150050699; US20150045546; US20150031134; US20150024500; US20140377868; US20140357530; US20140349400; US20140335620; US20140335063; US20140315985; US20140310830; US20140310828; US20140309487; US20140304853; US20140298547; US20140295556; US20140294773; US20140287938; US20140273234; US20140273232; US20140273231; US20140273230; US20140271987; US20140256046; US20140248702; US20140242702; US20140242700; US20140242699; US20140242664; US20140234972; US20140227787; US20140212869; US20140201857; US20140199767; US20140189896; US20140186958; US20140186919; US20140186843; US20140179770; US20140179006; US20140170753; Makarova et al., “Evolution and classification of the CRISPR-Cas systems” 9(6) Nature Reviews Microbiology 467-477 (1-23) (Jun. 2011); Wiedenheft et al., “RNA-guided genetic silencing systems in bacteria and archaea” 482 Nature 331-338 (Feb. 16, 2012); Gasiunas et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria” 109(39) Proceedings of the National Academy of Sciences USAE2579-E2586 (Sep. 4, 2012); Jinek et al., “A Programmable Dual-RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity” 337 Science 816-821 (Aug. 17, 2012); Carroll, “A CRISPR Approach to Gene Targeting” 20(9) Molecular Therapy 1658-1660 (Sep. 2012); U.S. Appl. No. 61 / 652,086, filed May 25, 2012; Al-Attar et al., Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs): The Hallmark of an Ingenious Antiviral Defense Mechanism in Prokaryotes, Biol Chem. (2011) vol. 392, Issue 4, pp. 277-289; Hale et al., Essential Features and Rational Design of CRISPR RNAs That Function With the Cas RAMP Module Complex to Cleave RNAs, Molecular Cell, (2012) vol. 45, Issue 3, 292-302.

[0107] The variant proteins described herein can be used in place of the corresponding proteins (e.g., nickases, prime editors, base editors, etc.) described in the foregoing references or in combination with analogous mutations described therein, with a guide RNA (or crRNA) appropriate for the selected CasPhi2. The methods can include delivering either nucleic acids encoding one or both of the CasPhi2 and / or guide RNA, or a ribonucleoprotein (RNP) complex comprising CasPhi2 protein and guide RNA, to the cell. In addition, the methods can be used in vitro, e.g., on non-genomic target DNA. The proteins described herein can be used with their corresponding proteins to edit or modify target DNA, e.g., to induce single or double stranded breaks (e.g., nucleases or nickases), or to alter base sequence (e.g., cytosine or adenine base editors), insert sequences (e.g., prime editors), alter transcriptional regulation (e.g., fusions with a transcriptional activator or repressor), alter histone methylation or acetylation modifiers (e.g., fusion with a histone acetyltransferase (HAT), histone deacetylase (HD AC), histone methyltransferase (HMT), or histone demethylase), or alter DNA methylation. See, e g., WO 2014 / 152432. Such methods can include contacting the target DNA (e.g., in vitro in a dish, or ex vivo in isolated living cells (e g., in culture), or in vivo in a living organism) with the modified gRNAs described herein and their corresponding Cas protein. This can be achieved, e.g., by contacting the target DNA with a ribonucleoprotein (RNP) complex comprising the Cas protein and the modified gRNA, or by expressing the gRNA and Cas protein in a cell comprising the target DNA, or by expressing the Cas protein and contacting the cell with the modified gRNA.

[0108] Nucleic Acids

[0109] Also provided herein are isolated nucleic acids encoding the CasPhi2 variants, vectors comprising the isolated nucleic acids, optionally operably linked to one or more regulatory domains for expressing the variant proteins, and host cells, e.g., mammalian host cells, comprising the nucleic acids, and optionally expressing the variant proteins.

[0110] Guide RNAs (gRNAs) / CRISPR RNAs (crRNAs) for CasPhi2 and variants

[0111] In contrast to Cas9 guide RNAs, which can consist of separate CRISPR RNAs (crRNAs) and tracrRNAs that function together to guide cleavage or chimeric fused crRNA-tracrRNAs (referred to as a single guide RNA or sgRNA, see also Jinek et al., Science 2012; 337:816-821), CasPhi nucleases (and CasPhi2 in particular) are guided to their target sites by a crRNAthat contains a 5’ direct repeat and a 3’ spacer sequence (the latter being complementary to the target DNA sequence), without the need for a tracrRNA.

[0112] In some embodiments, vectors (e.g., plasmids) encoding one or more CasPhi2 crRNA are used, e g., plasmids encoding, 2, 3, 4, 5, or more crRNAs directed to different sites in the same region of a target gene or to different target genes.

[0113] CasPhi2 variants as described herein can be guided to specific genomic targets using a crRNA consisting of a 25-40 nt repeat sequence (in some instances, including the pre-crRNA) a at its 5’ end and a 14-30 nt (e.g., 24, 25, 26 or 27 nts) spacer sequence (also referred to herein as “spacer region,” “crRNA spacer,” or the like) at its 3’ end that is complementary to the “target strand” or the “non-target strand” of the target DNA site. In this application, we refer to the CasPhi2 crRNAs as “crRNAs”, “guide RNAs” or “gRNAs” and use these terms interchangeably.

[0114] The CasPhi2 gRNAs / crRNAs can include on the 5’ and / or 3’ ends additional XN sequences, which can be any sequence (X is any nucleotide), wherein N (in the RNA) can be 1-200, e.g., 1-100, 1-50, or 1-20, that does not interfere with the binding of the ribonucleic acid to CasPhi2.

[0115] In some embodiments, the gRNA / crRNA includes one or more Adenine (A) or Uracil (U) nucleotides on the 3’ end. In some embodiments the RNA includes zero or more U, e g., 0 to 8 or more Us (e g., U, UU, UUU, UUUU, UUUUU, UUUUUU, UUUUUUU, UUUUUUUU) at the 3’ end of the molecule, as a result of the optional presence of one or more Ts used as a termination signal to terminate RNA PolIII transcription of these RNAs from DNA expression vectors.

[0116] In some embodiments, the gRNA / crRNA is targeted to a site that is at least three or more mismatches different from any sequence in the rest of the genome in order to minimize off-target effects. In some embodiments, the guide RNA includes one or more Guanine (G) nucleotides at the 5’ end for enhanced expression from a U6 promoter from DNA expression vectors in mammalian cells. In some embodiments, the guide RNA includes one or more Guanine (G) nucleotides (e.g., one G or two G’s at the 5’ end, preferably two Gs, i.e. 5’GG) at the 5’ end for enhanced expression from a T7 promoter for in vitro transcription (IVT) of the gRNA.

[0117] Modified RNA oligonucleotides such as locked nucleic acids (LNAs) have been demonstrated to increase the specificity of RNA-DNA hybridization by locking the modified oligonucleotides in a more favorable (stable) conformation. For example, 2’-O- methyl RNA is a modified base where there is an additional covalent linkage between the 2’ oxygen and 4’ carbon which when incorporated into oligonucleotides can improve overall thermal stability and selectivity (Formula I).

[0118] Formula I - Locked Nucleic Acid Thus in some embodiments, the gRNAs / crRNAs disclosed herein may comprise one or more modified RNA oligonucleotides. For example, the gRNA / crRNA molecules described herein can have one, some or all of the 17-18 or 17-19 nts 5’ region of the gRNA / crRNA spacer that is complementary to the target strand of the target sequence is / are modified, e.g., locked (2’-O-4’-C methylene bridge), 5 ’-methylcytidine, 2’-O- methyl-pseudouridine, or in which the ribose phosphate backbone has been replaced by a polyamide chain (peptide nucleic acid), e.g., a synthetic ribonucleic acid.

[0119] In other embodiments, one, some or all of the nucleotides of the gRNA / crRNA sequence may be modified, e.g., locked (2’-O-4’-C methylene bridge), 5 ’-methylcytidine, 2’-O-methyl-pseudouridine, or in which the ribose phosphate backbone has been replaced by a polyamide chain (peptide nucleic acid), e.g., a synthetic ribonucleic acid.

[0120] In some embodiments, the gRNAs and / or crRNAs can include one or more Adenine (A) or Uracil (U) nucleotides on the 3’ end.

[0121] Existing Cas9-based RNA-guided nucleases use gRNA-DNA heteroduplex formation to guide targeting to genomic sites of interest. However, RNA-DNA heteroduplexes can form a more promiscuous range of structures than their DNA-DNA counterparts. In effect, DNA-DNA duplexes are more sensitive to mismatches, suggesting that a DNA-guided nuclease may not bind as readily to off-target sequences, making them comparatively more specific than RNA-guided nucleases. Thus, the gRNA / crRNAs usable in the methods described herein can be hybrids, i.e., wherein one or more deoxyribonucleotides, e.g., a short DNA oligonucleotide, replaces all or part of the gRNA, e.g., all or part of the complementarity region of a gRNA. This DNA-based molecule could replace either all or part of the gRNA / crRNA. Such a system that incorporates DNA into the spacer complementarity region should more reliably target the intended genomic DNA sequences due to the general intolerance of DNA-DNA duplexes to mismatching compared to RNA-DNA duplexes. Methods for making such duplexes are known in the art, See, e.g., Barker et al., BMC Genomics. 2005 Apr 22;6:57; and Sugimoto et al., Biochemistry. 2000 Sep 19;39(37): 11270-81.

[0122] In a cellular context, complexes of CasPhi2 with these synthetic gRNAs / crRNAs could be used to improve the genome-wide specificity of the CRISPR / Cas9 nuclease system. The methods described can include expressing in a cell, or contacting the cell with, a CasPhi2 gRNA / crRNA plus a fusion protein as described herein.

[0123] Specific examples of CasPhi crRNAs and Proteins include the following.

[0124] CasPhi crRNAs are described, e.g., in WO2022159822. In some embodiments, the crRNA includes a protein-binding region that binds the CasPhi protein and a targeting region that is complementary to 14-24 nucleotides of a respective target genomic sequence or sequences. In some embodiments, the crRNA comprises one of the following sequences:

[0125] 5’-CAACGAUUGCCCCUCACGAGGGGAC-Ni2-24-Uo-8, SEQ ID NO: 13, or

[0126] 5’-GCAACGAUUGCCCCUCACGAGGGGAC-Ni2-24-Uo-8, SEQ ID NO: 14, or pre-crRNAs, e.g.,

[0127] 5’-GUCGGAACGCUCAACGAUUGCCCCUCACGAGGGGAC-Ni2-24-Uo-8, SEQ ID NO: 15,

[0128] 5’-GGUCGGAACGCUCAACGAUUGCCCCUCACGAGGGGAC-Ni2-24-Uo-8, SEQ ID NO: 16,

[0129] 5’-GGCAACGAUUGCCCCUCACGAGGGGAC-N12-24-Uo-8, SEQ ID NO: 17, or 5’-GGGUCGGAACGCUCAACGAUUGCCCCUCACGAGGGGAC-Ni2-24-Uo-8, SEQ-ID N: 18, wherein N is any nucleotide (N12-24 represents the targeting sequence that is complementary to the target genomic sequence). With CasPhi2, we generally observed higher editing with 5’ modified crRNA and 5’ modified pegRNA.

[0130] Expression Systems

[0131] To use the CasPhi2 variants described herein, it may be desirable to express them from a nucleic acid that encodes them. This can be performed in a variety of ways. For example, the nucleic acid encoding the CasPhi2 variant can be cloned into an intermediate vector for transformation into prokaryotic or eukaryotic cells for replication and / or expression. Intermediate vectors are typically prokaryote vectors, e.g., plasmids, or shuttle vectors, or insect vectors, for storage or manipulation of the nucleic acid encoding the CasPhi2 variant for production of the CasPhi2 variant. The nucleic acid encoding the CasPhi2 variant can also be cloned into an expression vector, for administration to a plant cell, animal cell, preferably a mammalian cell or a human cell, fungal cell, bacterial cell, or protozoan cell.

[0132] To obtain expression, a sequence encoding a CasPhi2 variant is typically subcloned into an expression vector that contains a promoter to direct transcription. Suitable bacterial and eukaryotic promoters are well known in the art and described, e.g., in Sambrook et al., Molecular Cloning, A Laboratory Manual (3d ed. 2001); Kriegler, Gene Transfer and Expression: ALaboratory Manual (1990); and Current Protocols in Molecular Biology (Ausubel et al., eds., 2010). Bacterial expression systems for expressing the engineered protein are available in, e.g., E. coli, Bacillus sp., and Salmonella (Palva et al., 1983, Gene 22:229-235). Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available.

[0133] The promoter used to direct expression of a nucleic acid depends on the particular application. For example, a strong constitutive promoter is typically used for expression and purification of fusion proteins. In contrast, when the CasPhi2 variant is to be administered in vivo for gene regulation, either a constitutive or an inducible promoter can be used, depending on the particular use of the CasPhi2 variant. In addition, a preferred promoter for administration of the CasPhi2 variant can be a weak promoter, such as HSV TK or a promoter having similar activity. The promoter can also include elements that are responsive to transactivation, e.g., hypoxia response elements, Gal4 response elements, lac repressor response element, and small molecule control systems such as tetracycline-regulated systems and the RU-486 system (see, e.g., Gossen & Bujard, 1992, Proc. Natl. Acad. Sci. USA, 89:5547; Oligino et al., 1998, Gene Then, 5:491-496; Wang et al., 1997, Gene Then, 4:432-441; Neering et al., 1996, Blood, 88: 1147-55; and Rendahl et al., 1998, Nat. Biotechnol., 16:757-761).

[0134] In addition to the promoter, the expression vector typically contains a transcription unit or expression cassette that contains all the additional elements required for the expression of the nucleic acid in host cells, either prokaryotic or eukaryotic. A typical expression cassette thus contains a promoter operably linked, e.g., to the nucleic acid sequence encoding the CasPhi2 variant, and any signals required, e.g., for efficient polyadenylation of the transcript, transcriptional termination, ribosome binding sites, or translation termination. Additional elements of the cassette may include, e.g., enhancers, and heterologous spliced intronic signals.

[0135] The particular expression vector used to transport the genetic information into the cell is selected with regard to the intended use of the CasPhi2 variant, e.g., expression in plants, animals, bacteria, fungus, protozoa, etc. Standard bacterial expression vectors include plasmids such as pBR322 based plasmids, pSKF, pET23D, and commercially available tag-fusion expression systems such as GST and LacZ.

[0136] For delivery of CasPhi2 and episomal expression of CasPhi2 and / or (pre)crRNAs in mammalian cells ex vivo or in vivo, adeno associated virus (AAV)-based vector systems or integration-deficient lentiviruses (IDLV) can be used. For ex vivo integration of CasPhi2 sequences in the cellular genome, lentiviruses or gammaretroviruses could be used as vector systems.

[0137] Expression vectors containing regulatory elements from eukaryotic viruses are often used in eukaryotic expression vectors, e.g., SV40 vectors, papilloma virus vectors, and vectors derived from Epstein-Barr virus. Other exemplary eukaryotic vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV40 early promoter, SV40 late promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0138] The vectors for expressing the CasPhi2 variants can include RNAPol III promoters to drive expression of the crRNAs or pre-crRNAs, e.g., the Hl, U6 or 7SK promoters. These promoters allow for expression of the crRNAs or pre-crRNAs in mammalian cells following plasmid transfection.

[0139] Some expression systems have markers for selection of stably transfected cell lines such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase. High yield expression systems are also suitable, such as using a baculovirus vector in insect cells, with the CasPhi2 variant and the crRNA or pre-crRNA encoding sequence under the direction of the polyhedrin promoter or other strong baculovirus promoters. The elements that are typically included in expression vectors also include a replicon that functions in E. coli, a gene encoding antibiotic resistance to permit selection of bacteria that harbor recombinant plasmids, and unique restriction sites in nonessential regions of the plasmid to allow insertion of recombinant sequences.

[0140] Standard transfection methods are used to produce bacterial, mammalian, yeast or insect cell lines that express large quantities of protein, which are then purified using standard techniques (see, e.g., Colley et al., 1989, J. Biol. Chem., 264: 17619-22; Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990)). Transformation of eukaryotic and prokaryotic cells are performed according to standard techniques (see, e.g., Morrison, 1977, J. Bacteriol. 132:349-351; Clark-Curtiss & Curtiss, Methods in Enzymology 101 :347-362 (Wu et al., eds, 1983).

[0141] Any of the known procedures for introducing foreign nucleotide sequences into host cells may be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, nucleofection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors, both episomal and integrative, and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA or other foreign genetic material into a host cell (see, e.g., Sambrook et al., supra). It is only necessary that the particular genetic engineering procedure used be capable of successfully introducing at least one gene into the host cell capable of expressing the CasPhi2 variant.

[0142] The present invention also includes the vectors and cells comprising the vectors.

[0143] Also provided herein are compositions and kits comprising the variants described herein. In some embodiments, the kits include the fusion proteins and a cognate guide RNA (i.e., a guide RNA that binds to the protein and directs it to a target sequence appropriate for that protein). In some embodiments, the kits also include labeled detector DNA, e.g., for use in a method of detecting a target ssDNA or dsDNA. Labeled detector DNAs are known in the art, e.g., as described in US20170362644; East-Seletsky et al., Nature. 2016 Oct 13; 538(7624): 270-273; Gootenberg et al., Science. 2017 Apr 28; 356(6336): 438-442, and WO2017219027 Al, and can include labeled detector DNAs comprising a fluorescence resonance energy transfer (FRET) pair or a quencher / fluor pair, or both. The kits can also include one or more additional reagents, e.g., additional enzymes (such as RNA polymerases) and buffers, e.g., for use in a method described herein.

[0144] EXAMPLES

[0145] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

[0146] Methods

[0147] The following materials and methods were used in the Examples below.

[0148] Molecular cloning. All crRNAs and pegRNAs used in this study were cloned into a pUC19-U6 mammalian expression vector which was digested with BsmBI-v2 and Hindlll-HF. DNA fragments that contain the direct repeat sequence / scaffold and the spacer sequence were prepared by overlap extension PCR with oligos with overlapping sequences using Phusion high-fidelity DNA polymerase. The PCR fragments were separated by 1% agarose gel electrophoresis and purified using QIAquick PCR purification kit. The purified PCR fragments were inserted into the digested vector prepared as above by Gibson Assembly using 2x Gibson master mix at 50°C for 1.5 h and the reaction mixture was then used to transform competent Escherichia coli XLl-Blue cells. The plasmids were purified from transformed cells by QIAgen Miniprep or Plus Midi kits. The plasmids encoding the enzymes used in this study (CasPhi2-17AA (ref to patent?) and MML-RT variant without RNase H domain3) were prepared previously by Gibson assembly using pCMV mammalian expression vector. To generate mutants of CasPhi2-17AA used in this study, DNA fragments containing the mutations were prepared by PCR amplification using primers carrying the mutation and Phusion high- fidelity DNA polymerase and inserted into pCMV mammalian expression vector digested with Agel-HF and Notl-HF. The mutants were cloned and plasmids prepared as for the RNAs detailed above.

[0149] Cell culture. HEK293T cells were cultured in Dulbecco’s modified Eagle medium (Gibco) containing 2 mM L-glutamine, supplemented with 10% FBS, 50 units / ml penicillin and 50 pg / ml streptomycin, and 1% GlutaMAX (Gibco). Cells were grown at 37 °C with 5% CO2 and passaged when reaching 80% confluency. Cell culture supernatants were tested for mycoplasma contamination every 4 weeks using MycoAlert PLUS mycoplasma detection kit (Lonza). Transfection. HEK293T cells were seeded at 1 .25 x 104cells in 92 pL growth medium / well. After 18-24 h incubation, the cells were transfected with 30 ng CasPhi2- 17AA or CasPhi2-17AA mutant and 10 ng crRNA using 0.3 pL TransIT-X2 lipofection reagent (Minis) and 9 pL of Opti-MEM (Gibco) per well. For split prime editing, the cells were transfected with 30 ng CasPhi2-17AA or CasPhi2-17AA mutant, 15 ng MMLV-RT ARNase H, and 10 ng of pegRNA with 5’ modified U6 stem-loop with 15 nt spacer sequence using 0.3 pL TransIT-X2 lipofection reagent (Minis) and 9 pL of Opti- MEM (Gibco) per well. After transfection, the cells were incubated at 37 °C with 5% CO2 for 72 h before extraction of genomic DNA.

[0150] DNA extraction. HEK293T cells were washed with 1 * PBS (Corning) and treated with 43.5 pl of gDNA lysis buffer (100 mM Tris-HCl at pH 8, 200 mM NaCl, 5 mM EDTA, 0.05% SDS) supplemented with 5.25 pl of 20 mg ml-1 Proteinase K (NEB) and 1.25 pL of 1 M DTT (Sigma) per well in 96-well plates. Cells were lysed overnight by shaking at 500 rpm at 55 °C. Subsequently, gDNA was extracted from lysates using 2x paramagnetic beads, washed twice with 70% ethanol, and eluted with 30pl of O. lx EB.

[0151] Library preparation for targeted amplicon sequencing. The gDNA concentrations were determined using a Qubit fluorometer and dsDNA HS Assay Kit (Thermo Fisher). The amplicon library for Miseq was generated by a 2-PCR process. For PCR1, the sequence of interest was amplified from 30-100ng gDNA using primers containing Illumina adapter sequences. The amplicons were purified with 0.7x paramagnetic beads and eluted in 30 pl nuclease-free water. In PCR2, 20-100ng of the PCR1 product was used in a PCR to add Illumina-compatible barcodes which was subsequently purified with 0.7x paramagnetic beads, eluted in 20 pl nuclease-free water, and quantified using the Quantifluor system (Promega). The PCR2 products were pooled based on the concentrations to ensure that all samples were represented equally in the final library which was then sequenced using an Illumina Miseq kit (Miseq Reagent Kit v.2; 300 cycles, 2 x 150 bp, paired-end). The FASTQ files were downloaded from BaseSpace (Illumina) for sequencing data analysis.

[0152] Next-generation sequencing analysis. Amplicon sequencing data were analyzed using CRISPResso2 using Base Editor Output mode. The CRISPResso2 output table ‘CRISPRessoBatch_quantification_of_editing_frequency.txt.’ was utilized to calculate indel frequencies reported around the cut site using the window parameters (-wc -1 -w 6) using the formula: ((‘insertions’+’deletions’-‘insertions and deletions’) / ’ reads aligned’) * 100.

[0153] Example 1: Rational design and human cell-based testing of potential CasPhi2-l 7AA NTS variants by mutation of the F517 capping residue

[0154] The published ternary structure of CasPhi2 with a DNA substrate harboring phosphorothioate backbone modification clearly showed that the NTS is accessible at the active site for cleavage prior to cleavage of the TS (PDB: 7LYT4). Since the TS is not accessible at the active site until after the NTS is cleaved, we hypothesized that we could develop an NTS nickase mutant of CasPhi2 by reducing or limiting the access of the TS at the active site. The published structure of CasPhi2 also revealed that amino acid F517 of CasPhi2 acts as a capping residue that limits the size of the R-loop between the TS and crRNA by intercalating in the R-loop4(Fig. 2A). Hence, we envisioned creating a NTS nickase by making mutations to F517, which would be predicted to permit the formation of an extended length R-loop between the TS and the CasPhi2 guide RNA (which is referred to as a CRISPR RNA or crRNA), and thereby potentially limiting access of the TS at the nuclease active site.

[0155] To test this hypothesis, we constructed CasPhi2-17AA variants harboring F517L, F517A or F517W mutations and crRNAs with standard (20 nts) and extended-length spacer sequences (24, 25, 26 or 27 nts) that might be annealed to longer stretches of the TS to facilitate the formation of the desired extended-length R-loop. We used a human cell-based screening assay to compare nuclease-induced gene editing activities of all combinations of these CasPhi2-17AA variants and extended-length crRNAs with those of CasPhi2-17AA and standard-length crRNAs with 20 nt spacers, reasoning that variants with desired nickase activity would show reduced indel mutations at the target sites. The results of these experiments performed at three different endogenous gene target loci in HEK293T cells demonstrated that the F517A mutation resulted in lower indel frequencies at all three target sites tested (Fig. 1A-1C). Furthermore, although reductions in indels relative to the CasPhi2-17AA were observed across all lengths of crRNAs, in general greater reductions in indel frequencies were observed with the more extended length crRNAs (25, 26, and 27 nts) (Fig. 1A-1C). These reductions in gene editing activities of the CasPhi2-17AA (F517A) variant suggest that it is a nickase due to prevention of R-loop capping and formation of a longer-length R-loop leading to reduced TS cleavage activity as we hypothesized.

[0156] Example 2: Additional rational design and human cell-based testing of potential CasPhi2-l 7AA NTS nickase variants

[0157] We also pursued three additional structure-guided strategies to rationally engineer CasPhi2-17AA variants with potential NTS nickase activities:

[0158] (1) We focused on mutating a-helix 17 of CasPhi2-17AA, which previously published structural studies show is located far from the nuclease active site but is close enough to the TS to interact with it. We reasoned that if we could create a new interaction between this alpha-helix and the TS that doing so would perhaps hold this strand away from the active site, thereby reducing or preventing its cleavage. To accomplish this, we introduced positively charged lysine (K) amino acid substitutions at a few residues near and within a-helix 17 (S496, N497, T498, T499, S502, E503, L506, S509, S511; Fig. 2A) designed with the goal to create an interaction between each of these positions and the TS.

[0159] (2) We also introduced mutations in CasPhi2-17AAin a lysine-rich protein loop that is proximal to the TS (based on a published cryo-EM structure) and that we hypothesized may interact with the TS and play a role in positioning this DNA strand with the nuclease active site. To disrupt this interaction with the TS and thereby reduce cleavage of this DNA strand, we mutated five lysine residues in this loop (K522, K523, K526, K527, K528; Fig. 2A) to either glutamate (E) or glycine (G).

[0160] (3) We mutated R678 residue in CasPhi2-17AAnear the active site to either alanine (A) or glycine (G) (Fig. 2B). The rationale for choosing to mutate these particular residues was based on analogy to a previous study with AsCasl2a (another type V nuclease which also harbors only a single DNA nuclease domain) which reported that an R1226A mutation at a protein loop near the active site led to preferential single-strand cleavage over double-strand cleavage activity6.

[0161] We again used a human cell-based screening assay to compare nuclease-induced gene editing activities of these various CasPhi2-17AA variants with the parental CasPhi2-17AA nuclease, reasoning that variants with desired nickase activity would show reduced indel mutations at the target sites. We performed these screens for three different endogenous genome loci in human HEK293T cells and found that at least four of the 31 variants we tested (T498K, S502K, R678A, and R678G) showed significant reductions in indel frequencies (Fig. 2D), consistent with these four variants being potential NTS nickases.

[0162] Example 3: Construction of and human-cell testing of alanine scanning variants of CasPhi2-l 7AA bearing mutations in a-helix 14 to create potential kinetic NTS nickases

[0163] The bridge helix of the type V CRISPR-Casl2a nuclease plays an important role in DNA cleavage through structural change of the enzyme upon DNA binding7. A previously published study has shown that introduction of a W890A within the LbCasl2a bridge helix leads to a substantially slowed rate of TS cleavage relative to NTS cleavage7. Based on this observation, we sought to develop an analogous kinetic NTS nickase of CasPhi2-17AA. To do this, we performed a structure-based comparison between different Cast 2a enzymes and CasPhi2 and noted that a-helix 14 of CasPhi2 appears similar to the bridge helix of Casl2a. Based on this, we conducted an alanine scanning mutagenesis of residues within a-helix 14 of CasPhi2-17AA (R437, W440, D441, R442, E444, E445, E446, R448, R450), replacing each residue individually with alanine. We screened these nine CasPhi2-17AA variants and the parental CasPhi2-17AA nuclease for their abilities to induce targeted indels at three different endogenous gene loci in HEK293T cells. We found that three variants (R437A, D441A, and E444A) each consistently induced lower indel frequencies than CasPhi2-17AA across all three target sites (Fig. 3), suggesting that these three variants may be potential kinetic NTS nickases. Example 4: Rational design and human cell-based testing of potential CasPhi2-l 7AA TS nickase mutants

[0164] In addition to generating NTS nickase mutants, we also envisioned developing a CasPhi2-17AA TS nickase, which could be used to construct more efficient CasPhi2- based base editors. To do this, we started from a published finding showing that FnCasl2a can be converted to a TS nickase by introducing two mutations simultaneously: K1013G and R1014G8. Based on a structural comparison we performed between FnCasl2a and CasPhi2, we found a region of the latter nuclease that appeared structurally similar, is enriched for positively charged amino acid residues (lysine and arginine), and is located near the crRNA spacer region. Based on this analysis, we mutated four positively charged residues in this region (K618, R619, K630, and K631; Fig. 2C), creating single and double-amino acid substitutions to glycine at these positions in CasPhi2-17AA. We screened six such variants and the parental CasPhi2-17AA nuclease for their abilities to induce targeted indels at three different endogenous gene loci in HEK293T cells, again reasoning that TS nickases should show reduced gene editing activities in this assay (Fig. 2D). The results of these experiments showed that the two double mutants of CasPhi2-17AA (K618G / R619G and K630G / K631G) exhibited lowered activities relative to CasPhi2-17AA (Fig. 2D), as would be expected for nickases. These potential TS nickases could potentially be used to construct more efficient base editors as well as prime editor proteins that function by nicking the target strand as recently described2.

[0165] Example 5: Construction and testing of prime editors made using CasPhi2-l 7AA NTS nickase variants

[0166] We investigated whether some of the potential CasPhi2 NTS nickases we generated might be used to perform prime editing. To do this, we transfected plasmid expressing CasPhi2-17AA nuclease, a potential CasPhi2-17AA NTS nickase (bearing T498K or S502K mutations), or a potential CasPhi2-17AA TS nickases (bearing K618G / R619G or K630G / K631G mutations) into HEK293T cells together with a plasmid expressing an MMLV-RT pentamutant ARNaseH variant1,5and a plasmid expressing one of two pegRNAs. These two pegRNAs harbored PBS+RTT sequences designed to introduce an ATG insertion at positions 12 or 13 (Pl 2 or Pl 3) of a CasPhi2 crRNA spacer in the VEGFA 3.1 target site present in the human endogenous VEGFA gene (Fig. 4A). These pegRNAs also contained a linker sequence (derived from an SpCas9 tracrRNA sequence) between the PBS+RTT and the CasPhi2 crRNA (consisting of a direct repeat (DR) and spacer targeting sequence) and a modified U6 stem loop sequence appended to the 5’ end (added to increase the activity of the pegRNA) (Fig. 4A). The full sequences of these two pegRNA are shown in Fig. 4A. We isolated genomic DNA from the transfected cells and performed targeted amplicon sequencing to assess the frequencies of desired prime editing (insertion of the ATG) and of indel mutations at the intended VEGFA 3.1 target site. This experiment demonstrated that the desired prime editing event (insertion of an ATG) was observed with the potential CasPhi2-17AA NTS nickases but not with the potential CasPhi2-17AA TS nickases or the CasPhi2-17AA nuclease, consistent with the possibility that our potential NTS nickases are nicking the NTS as expected (Fig. 4B). Importantly, indel mutations were only observed with the CasPhi2-17AA nuclease and not with the potential NTS or TS nickases (Fig. 4B), consistent with the idea that are variants no longer possess nuclease activity and may only cleave one of the DNA strands. Although we performed prime editing in this example using an MMLV-RT variant that is expressed in trans to the CasPhi2 nickase, it should also be possible (by analogy to SpCas9-based prime editor architectures (Grunewald et al., Nat Biotechn 2023 (PMID: 36163548)) to perform prime editing with the MMLV-RT variant expressed in cis as a fusion to the CasPhi2 nickase.

[0167] Taken together, the results of our prime editing experiments suggest that CasPhi2- 17AA T498K and S502K variants may be NTS nickases and are consistent with the possibility that the CasPhi2-17AA K618G / R619G or K630G / K631G variants may be TS nickases that could be used to build base editors or prime editors that require nicking of the TS. Furthermore, it seems likely that combinations of various mutations that may convert CasPhi2-17AA into NTS nickases might possess even more preferential cutting of the NTS relative to the TS. Similarly, it also seems likely that combinations of various mutations that may convert CasPhi2-17AA into TS nickases might possess even more preferential cutting of the TS relative to the NTS. References:

[0168] 1. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019).

[0169] 2. Bill Kim, Y. et al. A novel mechanistic framework for precise sequence replacement using reverse transcriptase and diverse CRISPR-Cas systems. http : / / biorxiv. org / lookup / doi / 10.1101 / 2022.12.13.520319 (2022) doi: 10.1101 / 2022.12.13.520319.

[0170] 3. Pausch, P. et al. CRISPR-CasO from huge phages is a hypercompact genome editor. Science 369, 333-337 (2020).

[0171] 4. Pausch, P. et al. DNA interference states of the hypercompact CRISPR Cas<b effector. Nat. Struct. Mol. Biol. 28, 652-661 (2021).

[0172] 5. Griinewald, J. et al. Engineered CRISPR prime editors with compact, untethered reverse transcriptases. Nat. Biotechnol. (2022) doi: 10.1038 / s41587-022-01473-l.

[0173] 6. Yamano, T. et al. Crystal Structure of Cpfl in Complex with Guide RNA and Target DNA. Cell 165, 949-962 (2016).

[0174] 7. Ma, E. et al. Improved genome editing by an engineered CRISPR-Cas 12a. Nucleic Acids Res. gkacl 192 (2022) doi: 10.1093 / nar / gkacl 192.

[0175] 8. Paul, B., Chaubet, L., Verver, D. E. & Montoya, G. Mechanics of CRISPR-Casl2a and engineered variants on / -DNA. Nucleic Acids Res. 50, 5208-5225 (2022).

[0176] OTHER EMBODIMENTS

[0177] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. An isolated CasPhi2 protein, comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence of SEQ ID NO: 1, and mutations at one, two, three, four, five or more of the following positions: A36, S106, D134, L149, E159, S160, S164, D167, E168, P277, T355, T357, T518, L571, S616, D679, and / or Q684, and further comprising one or more mutations at R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E5O3, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, R678.

2. The isolated CasPhi2 protein of claim 1, comprising the following mutations: L149, E159, S160, S164, D167, and E168.

3. The isolated CasPhi2 protein of claim 1, comprising the following mutations: A36R, S106R, D134R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R.

4. The isolated CasPhi2 protein of claim 1 , comprising the following mutations: A36R, S106R, D134R, L149R, P277R, T355R, T357K, T518R, L571K, S616R, D679K, and Q684R.

5. The isolated CasPhi2 protein of claim 1, comprising the following mutations: A36K, S106K, D134K, P277K, D337K, T355R, T357K, V531R, T539A, A543K, L571K, S616K, D679K, and T691K.

6. The isolated CasPhi2 protein of claim 1, comprising the following mutations: A36K, S106K, D134K, P277K, D337K, T355R, T357K, V531R, T539A, A543K, L571K, S616K, D679K, Q684R, and T691K.

7. The isolated CasPhi2 protein of claim 1 , comprising the following mutations: A36R, S106R, D134R, L149R, E159A, S160A, S164A, D167K, E168A, P277R, T357K, T518R, L571K, S616R, Q684R, T355R, and D679K.

8. The isolated CasPhi2 protein of any one of claims 1-7, further comprising a mutation at F517, wherein the mutation at position F517 is F517L, F517A or F517W.

9. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following positions: S496, N497, T498, T499, S502, E503, L506, S509, and / or S511.

10. The isolated CasPhi2 protein of claim 9, wherein the mutations are S496K, N497K, T498K, T499K, S502K, E5O3K, L506K, S509K, and / or S511K.

11. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following positions: K522, K523, K526, K527, and / or K528.

12. The isolated CasPhi2 protein of claim 11, wherein the positions are mutated to either a glutamate (E) or glycine (G).

13. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following position: R678.

14. The isolated CasPhi2 protein of claim 13, wherein the positions are mutated to either an alanine (A) or glycine (G).

15. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following mutations: T498K, S502K, R678A, and / or R678G.

16. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following positions: R437, W440, D441, R442, E444, E445, E446, R448, R450.

17. The isolated CasPhi2 protein of claim 16, comprising the following mutations: R437A, D441A, and / or E444A.

18. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following positions: K618, R619, K630, and / or K631.

19. The isolated CasPhi2 protein of claim 18, wherein the positions are mutated to either an alanine (A) or glycine (G).

20. The isolated CasPhi2 protein of claim 19, comprising one of the following sets of mutations: (1) K618G / R619G or (2) K630G / K631G.

21. The isolated CasPhi2 protein of any one of claims 1-7, further comprising one or more mutations at the following positions: T498, S502, K618 / R619, or K630 / K631.

22. The isolated CasPhi2 protein of claim 21, comprising one of the following sets of mutations: T498K, S502K, K618G / R619G, or K630G / K631G.

23. A fusion protein comprising the isolated CasPhi2 protein of any one of claims 1-22, fused to at least one heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

24. The fusion protein of claim 23, wherein the heterologous functional domain is a biological tether.

25. The fusion protein of claim 24, wherein the biological tether is MS2, Csy4 or lambda N protein.

26. The fusion protein of claim 23, wherein the heterologous functional domain is Fokl.

27. The fusion protein of claim 23, wherein the heterologous functional domain is a deaminase.

28. The fusion protein of claim 27, wherein the heterologous functional domain is a cytidine deaminase.

29. The fusion protein of claim 28, wherein the cytidine deaminase is selected from the group consisting of AP0BEC1, AP0BEC2, AP0BEC3A, AP0BEC3B, APOBEC3C, AP0BEC3D / E, APOBEC3F, AP0BEC3G, AP0BEC3H, AP0BEC4, activation-induced cytidine deaminase (AID), cytosine deaminase 1 (CDA1), pmCDAl, CDA2, and cytosine deaminase acting on tRNA (CDAT).

30. The fusion protein of claim 27, wherein the heterologous functional domain is an adenosine deaminase.

31. The fusion protein of claim 30, wherein the adenosine deaminase is selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (AD ARI), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (ADAT1), ADAT2, ADAT3; and naturally occurring or engineered tRNA- specific adenosine deaminase (TadA).

32. The fusion protein of claim 23, wherein the heterologous functional domain is a reverse transcriptase (optionally MMLV-RT pentamutant ARNaseH variant).

33. The fusion protein of any one of claims 23 or 27 to 31 , comprising at least two heterologous functional domains, wherein the additional heterologous functional domain comprises an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways.

34. The fusion protein of claim 33, wherein the additional heterologous functional domain is a uracil DNA glycosylase inhibitor (UGI) that inhibits uracil DNA glycosylase (UDG, also known as uracil N-glycosylase, or UNG); or Gam from the bacteriophage Mu.

35. An isolated nucleic acid encoding the isolated CasPhi2 protein of any one of claims 1-22 or the fusion protein of claims 23-34.

36. A vector comprising the isolated nucleic acid of claim 35.

37. An isolated host cell comprising the nucleic acid of claim 36.

38. The isolated host cell of claim 37, wherein the host cell is a mammalian host cell.

39. A complex for prime editing comprising:(a) the isolated CasPhi2 protein of any one of claims 1-22,(b) a domain comprising an RNA-dependent DNA polymerase activity; and(b) a pegRNA.

40. The complex of claim 39, wherein parts (a) and (b) are expressed together as a single fusion protein from a single plasmid.

41. The complex of claim 39, wherein parts (a) and (b) are expressed from two separate plasmids.

42. The complex of any one of claims 39-41, wherein the domain comprising an RNA-dependent DNA polymerase activity is a reverse transcriptase (optionally MMLV-RT pentamutant ARNaseH variant).

43. The complex of any one of claims 39-42, wherein the pegRNA comprises a primer binding sequence (PBS), a reverse transcriptase template (RTT), one or more crRNAs, and a modified U6 stem loop sequence appended to the 5’ end of the pegRNA.

44. The complex of claim 43, wherein the pegRNA further comprises a linker sequence between the PBS+RTT and the one or more crRNAs.

45. A method for prime editing, the method comprising: expressing in a cell the complex of any one of claims 39-44.

46. A method of altering a genome of a cell, the method comprising expressing in the cell, or contacting the cell with, the isolated CasPhi2 protein of any one of claims 1-22 or the fusion protein of any one of claims 23-34, and one or more crRNAs, wherein the one or more crRNAs direct the isolated CasPhi2 protein of any one of claims 1-22 or the fusion protein of any one of claims 23-34 to one or more target genomic sequences.

47. The method of claim 46, wherein the cell is a stem cell.

48. The method of claim 47, wherein the stem cell is an embryonic stem cell, a mesenchymal stem cell, or an induced pluripotent stem cell; is in a living animal; or is in or is an embryo.

49. A method of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA with the isolated CasPhi2 protein of any one of claims 1-22 or the fusion protein of any one of claims 23-34, and one or more crRNAs, wherein the one or more crRNAs direct the isolated CasPhi2 protein of any one of claims1 -22 or the fusion protein of any one of claims 23-34 to one or more target genomic sequences.

50. The method of claim 49, wherein the dsDNA molecule is in vitro.

51. A CasPhi2 variant comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence of SEQ ID NO: 1 and comprising a mutation at one or more positions selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, N497, T498, T499, S502, E5O3, L506, S509, S511, K522, K523, K526, K527, K528, K618, R619, K630, K631, and R678 relative to SEQ ID NO: 1.

52. The CasPhi2 variant of claim 51, comprising the mutation at a position selected from the group consisting of: R437, W440, D441, R442, E444, E445, E446, R448, R450, F517, S496, T498, T499, S502, E503, K618, R619, K630, K631, and R678.

53. The CasPhi2 variant of claim 51 or 52, comprising the mutation at a position selected from the group consisting of: F517, R437, D441, E444, T498, S502, K618, R619, K630, K631 and R678.

54. The CasPhi2 variant of any one of claims 51-53, comprising the mutation at position F517.

55. The CasPhi2 variant of claim 54, wherein the F517 mutation is F517L, F517A or F517W.

56. The CasPhi2 variant of claim 55, wherein the F517 mutation is F517A.

57. The CasPhi2 variant of any one of claims 51-53, comprising the mutation at a position selected from the group consisting of: T498, S502, and R678.

58. The CasPhi2 variant of claim 57, wherein the mutations are selected from T498K, S502K, R678A, or R678G.

59. The CasPhi2 variant of any one of claims 51-53, comprising a mutation at a position selected from R437, D441, or E444.

60. The CasPhi2 variant of claim 59, wherein the mutations are R437A, D441A, or E444A.

61. The CasPhi2 variant of any one of claims 51-53, comprising at least two mutations at positions selected from K618, R619, K630, or K631.

62. The CasPhi2 variant of claim 61, wherein the at least two mutations comprise K618G and R619G.

63. The CasPhi2 variant of claim 61, wherein the at least two mutations comprise K630G and K631G.

64. A fusion protein comprising the CasPhi2 variant of any one of claims 51- 63, fused to at least one heterologous functional domain, with an optional intervening linker, wherein the linker does not interfere with activity of the fusion protein.

65. The fusion protein of claim 64, wherein the at least one heterologous functional domain comprises: a biological tether, a FokI, a deaminase, or a reverse transcriptase.

66. The fusion protein of claim 65, wherein the biological tether MS2, Csy4 or lambda N protein.

67. The fusion protein of claim 65, wherein the deaminase is a cytidine deaminase.

68. The fusion protein of claim 67, wherein the cytidine deaminase is selected from the group consisting of AP0BEC1, AP0BEC2, AP0BEC3A, AP0BEC3B, APOBEC3C, AP0BEC3D / E, APOBEC3F, AP0BEC3G, AP0BEC3H, AP0BEC4, activation-induced cytidine deaminase (AID), cytosine deaminase 1 (CDA1), pmCDAl, CDA2, and cytosine deaminase acting on tRNA (CDAT).

69. The fusion protein of claim 65, wherein the deaminase is an adenosine deaminase.

70. The fusion protein of claim 69, wherein the adenosine deaminase is selected from the group consisting of adenosine deaminase 1 (ADA1), ADA2; adenosine deaminase acting on RNA 1 (ADAR1), ADAR2, ADAR3; adenosine deaminase acting on tRNA 1 (AD ATI), ADAT2, ADAT3; and naturally occurring or engineered tRNA- specific adenosine deaminase (TadA).7E The fusion protein of claim 65, wherein the reverse transcriptase is a MMLV-RT pentamutant ARNaseH variant.

72. The fusion protein of any one of claims 65-71, comprising at least two heterologous functional domains, wherein an additional heterologous functional domain comprises an enzyme, domain, or peptide that inhibits or enhances endogenous DNA repair or base excision repair (BER) pathways.

73. The fusion protein of claim 72, wherein the additional heterologous functional domain is a uracil DNA glycosylase inhibitor (UGI) or Gam from the bacteriophage Mu.

74. An isolated nucleic acid encoding the CasPhi2 variant of any one of claims 51-63 or the fusion protein of claims 64-73.

75. A vector comprising the isolated nucleic acid of claim 74.

76. An isolated host cell comprising the nucleic acid of claim 74.

77. The isolated host cell of claim 76, wherein the host cell is a mammalian host cell.

78. A complex for prime editing comprising:(a) the CasPhi2 variant of any one of claims 51-63,(b) a domain comprising an RNA-dependent DNA polymerase activity; and(b) a pegRNA.

79. The complex of claim 78, wherein parts (a) and (b) expressed together as a single fusion protein from a single plasmid.

80. The complex of claim 78, wherein parts (a) and (b) are expressed from two separate plasmids.

81. The complex of any one of claims 78-80, wherein the domain is a reverse transcriptase.

82. The complex of claim 81, wherein the reverse transcriptase is a MMLV- RT pentamutant ARNaseH variant.

83. The complex of any one of claims 78-82, wherein the pegRNA comprises a primer binding sequence (PBS), a reverse transcriptase template (RTT), one or more crRNAs, and a modified U6 stem loop sequence appended to the 5’ end of the pegRNA.

84. The complex of claim 83, wherein the PegRNA further comprises a linker sequence between the PBS+RTT and the one or more crRNAs.

85. A method of altering a genome of a cell, the method comprising expressing in the cell, or contacting the cell with, the CasPhi2 variant of any one of claims 51-63, and one or more crRNAs, wherein the one or more crRNAs direct the CasPhi2 variant of any one of claims 51-63 to one or more target genomic sequences, wherein the CasPhi2 variant nicks a strand of the one or more target genomic sequences.

86. A method of altering a double stranded DNA (dsDNA) molecule, the method comprising contacting the dsDNA with the CasPhi2 variant of any one of claims 51-63 or the fusion protein of any one of claims 64-73, and one or more crRNAs, wherein the one or more crRNAs direct the CasPhi2 variant of any one of claims 51-63 or the fusion protein of any one of claims 64-73 to one or more target genomic sequences, wherein the CasPhi2 variant nicks a strand of the one or more target genomic sequences.

87. The method of claim 86, wherein the CasPhi2 variant is a nickase.

88. The method of claim 86, wherein the strand is a non-target strand.

89. The method of claim 86, wherein the strand is a target strand.

90. The method of any one of the preceding claims, wherein the method is performed in vivo.

91. The method of any one of the preceding claims, wherein the method is performed in vitro.

92. The method of any one of the preceding claims, wherein the method is performed ex vivo.

Citation Information

Patent Citations

  • Crispr-CAS effector polypeptides and methods of use thereof

    WO2020181101A1

  • Engineering immune orthoganol AAV and immune stealth crispr-cas

    WO2022120103A1

  • Engineered casphi2 nucleases

    WO2024086845A2