Enhanced CAS9 variants and uses thereof
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INTELLIA THERAPEUTICS INC
- Filing Date
- 2026-01-31
- Publication Date
- 2026-08-06
Smart Images

Figure IMGF000084_0001 
Figure IMGF000085_0001 
Figure IMGF000137_0001
Abstract
Description
ENHANCED CAS9 VARIANTS AND USES THEREOFCROSS REERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to US Provisional Application No.63 / 752,200, filed date January 31, 2025. The foregoing application is incorporated herein by reference in its entirety.SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted in xml format and is hereby incorporated by reference in its entirety. Said xml copy, created on January 22, 2026, is named 121468_1010WO_00348_SL.xml and is 5,517 kilobytes in size.INTRODUCTION
[0003] The present disclosure relates to polypeptides, polynucleotides, compositions, and methods for genome editing involving RNA-guided DNA binding agents such as CRISPR-Cas systems and subunits thereof.
[0004] RNA-guided DNA binding agents such as CRISPR-Cas systems can be used for targeted genome editing, including in eukaryotic cells and in vivo. Such editing has been shown to be capable of inactivating certain deleterious alleles or correcting certain deleterious point mutations. The CRISPR / Cas9 system relies on a nuclease, termed CRISPR- associated protein 9 (Cas9), which induces site-specific breaks in DNA. Cas9 is guided to specific DNA sequences by small RNA molecules termed guide RNA (gRNA).
[0005] CRISPR / Cas9 systems are found in various bacterial species, each possessing distinct properties such as varying levels of sequence specificity and editing activity. The wild-type Neisseria meningitidis Cas9 (WT NmeCas9) is notable for its low off-target cleavage rate compared to the Streptococcus pyogenes Cas9 (SpyCas9). However, NmeCas9 can often display lower potency than SpyCas9 for a given target. Therefore, to improve editing efficiency with NmeCas9, there is a need in the art for NmeCas9 nucleases with enhanced on-target editing capabilities compared to WT NmeCas9.SUMMARY
[0006] Provided herein are engineered NmeCas9 polypeptides that can provide improved editing efficiency or other benefits, relative to wild-type NmeCas9. The providedvariants of NmeCas9 have enhanced on-target editing efficiency while maintaining minimal off-target effects.
[0007] In one aspect, provided herein is a Neisseria meningitidis Cas9 (NmeCas9) polypeptide comprising an amino acid sequence with at least 90% identity to the amino acid sequence of SEQ ID NO: 1, wherein the amino acid sequence of the NmeCas9 polypeptide comprises one or more substitution(s) at a position(s) selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0008] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises one or more substitution(s) at a position(s) selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
[0009] In certain embodiments, the substitution is at a position selected from the group consisting of N1026, K266, Q422, and D418 of SEQ ID NO: 1
[0010] In one embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises an N1026 substitution relative to SEQ ID NO: 1. In one embodiment, the N1026 substitution is aN1026R substitution relative to SEQ ID NO: 1. In another embodiment, the N1026 substitution is aN1026K substitution relative to SEQ ID NO: 1.
[0011] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a K266 substitution relative to SEQ ID NO: 1. In one embodiment, the K266 substitution is a K266R substitution relative to SEQ ID NO: 1.
[0012] In certain embodiments, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a Q422 substitution relative to SEQ ID NO: 1. In one embodiment, the substitution is a Q422K or a Q422R substitution relative to SEQ ID NO: 1.
[0013] In another embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises a D418 substitution relative to SEQ ID NO: 1. In one embodiment, the substitution is a D418K or a D418R substitution relative to SEQ ID NO: 1.
[0014] In certain aspects, the amino acid sequence of the NmeCas9 polypeptide comprises an E932 substitution relative to SEQ ID NO: 1. In one embodiment, the E932 substitution is a substitution selected from the group consisting of E932K, E932N, E932Q, E932M, E932R, E932H, E932A, E932S, and E932T substitution relative to SEQ ID NO: 1.
[0015] In other aspects, the amino acid sequence of the NmeCas9 polypeptide comprises an E868 substitution relative to SEQ ID NO: 1. In one embodiment, the E868 substitution is an E868R or a E868K substitution relative to SEQ ID NO: 1.
[0016] In one embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises a K929 substitution relative to SEQ ID NO: 1. In one embodiment, the K929 substitution is a K929R substitution relative to SEQ ID NO: 1.
[0017] In another embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution selected from the group consisting of N1026R, N1026K, K266R, E932K, E932N, E932Q, E932M, E932H, E932A, E932S, E932T, E868, Q422, D418, E932, K929, and K1044 relative to SEQ ID NO: 1.
[0018] In other aspects, the amino acid sequence of the NmeCas9 polypeptide further comprises a K333 substitution relative to SEQ ID NO: 1
[0019] In still other embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a single substitution at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0020] In certain aspects, the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution. In one embodiment, the nickase comprises single substitution relative to SEQ ID NO: 1 at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
[0021] In other aspects, the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution. In one embodiment, the dCas9 comprises single substitution relative to SEQ ID NO: 1 at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
[0022] In another embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises two substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
[0023] In other embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of: a) E932 and any one of N1026, K266, E868, Q422, D418, or K929; b) N1026 and any one of E932, K266, E868, Q422, D418, or K929; c) K266 and any one of E932, N1026, E868, Q422, D418, or K929; d) E868 and any one of E932, N1026, K266, Q422, D418, or K929; e) Q422 and any one of E932, N1026, K266, E868, D418, or K929;f) D418 and any one of E932, N1026,K266, E868, Q422, or K929; and g) K929 and any one of E932, N1026, K266, E868, Q422, or D418 relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0024] In one embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of: a) E932 and N1026; b) E932 and E868; c) E932 and K266; d) E932 and D418; e) E932 and Q422; f) K266 and N1026; g) Q422 and N1026; h) D418 and N1026; and i) E868 and N1026. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0025] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises two amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of: N1026R and K266R; N1026R and D418K; N1026R and E868R; N1026R and E868K; N1026R and E932N; N1026R and E932Q; N1026R and E932M;N1026R and E932R; N1026R and E932H; N1026R and E932A; N1026R and E932S;N1026R and E932T; N1026K and E932N; N1026K and E932Q; N1026K and E932M;N1026K and E932R; N1026K and E932H; N1026K and E932A; N1026K and E932S;N1026K and E932T; E932R and K266R; E932R and D418K; E932R and Q422K; E932R and E868R; E932R and N1026R; E868R and N1026K; E932N and N1026R; E932M and N1026R; and E932A and N1026R, relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0026] In certain embodiments, the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: l.the nickase comprises two substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
[0027] In one embodiment, the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution. In one embodiment, the dCas9 comprises two substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
[0028] In certain aspects, the amino acid sequence of the NmeCas9 polypeptide comprises three substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of: E932 and any two of N1026, K266, E868, Q422, D418, and K929; N1026 and any two of E932, K266, E868, Q422, D418, and K929; K266 and any two of E932, N1026, E868, Q422, D418, and K929; E868 and any two of E932, N1026, K266, Q422, D418, and K929; Q422 and any two of E932, N1026, K266, E868, D418, or K929; D418 and any two of E932, N1026, K266, E868, Q422, and K929; and K929 and any two of E932, N1026, K266, E868, Q422, and D418, relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0029] In certain aspects, the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of: E932, K266, and D418; K266, Q422, andN1026; K266, Q422, and E868; Q422, E868, and N1026; K266, E868, and N1026; E932, D418, and Q422; E868, E932, and N1026; K266, E868, and N1026; and D418, E868, and N1026, relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0030] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of: E932R, K266R, and Q422K; E932R, K266R, and D418K; E932R, K266R, and D422K; K266R, Q422K, and E868R; E932R, D418K, and Q422K; E868R, E932R, and N1026K; K266R, E868R, and N1026R; and D418K, E868R, and N1026R, relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0031] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0032] In certain embodiments, the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution. In one embodiment, the nickase comprises three substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
[0033] In some embodiments, the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution. In one embodiment, the dCas9 comprises three substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0034] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises four substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of: E932 and any three of N1026, K266, E868, Q422, D418, and K929; N1026 and any three of E932, K266, E868, Q422, D418, and K929; K266 and any three of E932, N1026, E868, Q422, D418, and K929; E868 and any three of E932, N1026, K266, Q422, D418, and K929 Q422 and any three of E932, N1026, K266, E868, D418, or K929; D418 and any three of E932, N1026, K266, E868, Q422, and K929; and K929 and any three of E932, N1026, K266, E868, Q422, and D418, relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0035] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises four substitutions at positions relative to SEQ ID NO: 1 at positions selected from the group consisting of: K266, Q422, E868, and N1026; E932, K266, D418, and Q422; E932, K266, K333, D418, and Q422; E932, K266, K333, D418, Q422, E508, K517, and K549; and E932, K266, D418, Q422, E508, K517, and K549 relative to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
[0036] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:bK266R, Q422K, E868R, andN1026R; E932, K266R, D418K, and Q422K; E932, K266R, K333R, D418K, and Q422K; and E932R, K266R, K333R, D418K, Q422K, E508K, K517R, and K549R; and E932R, K266R, D418K, Q422K, E508K, K517R, K549R, relative to SEQ ID NO: 1.
[0037] In one embodiment, the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution. In one embodiment, the nickase comprises four substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
[0038] In certain embodiments, the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution. In one embodiment, the dCas9 comprises four substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
[0039] In certain embodiments, the polypeptide described herein comprises an amino acid sequence with at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1.
[0040] In certain embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 32, 65, 87, 98, 307, or 318.
[0041] In certain embodiments, the NmeCas9 polypeptide is an Nme2Cas9 polypeptide, an NmelCas9 polypeptide, or an Nme3Cas9 polypeptide.
[0042] In certain embodiments, the NmeCas9 polypeptide is a Nme2 Cas9 polypeptide.
[0043] In other embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise an E932D relative to SEQ ID NO: 1.
[0044] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise an E932R substitution relative to SEQ ID NO: 1.
[0045] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises neither an E932D substitution nor an E932R substitution relative to SEQ ID NO: 1.
[0046] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise an E868K relative to SEQ ID NO: 1.
[0047] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise a K929R substitution relative to SEQ ID NO: 1
[0048] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise a substitution at K1044 relative to SEQ ID NO: 1.
[0049] Also provided herein is a Neisseria meningitidis Cas9 (NmeCas9) polypeptide comprising an amino acid sequence with at least 90% identity to the amino acid sequence of SEQ ID NO: 2 or 3, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a substitution at S1022 of SEQ ID NO: 2 or G1022 of SEQ ID NO:3. In one embodiment, the NmeCas9 polypeptide comprises a substitution at S1022 relative to SEQ ID NO: 2 and comprises an amino acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 2. In another embodiment, the NmeCas9 polypeptide comprises a G1022 substitution relative to SEQ ID NO: 3 and comprises an amino acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 3.
[0050] In certain aspects, the NmeCas9 polypeptide further comprises a nuclear localization signal (NLS). In one embodiment, the NLS is selected from a c-Myc NLS, an SV40 NLS, a nucleoplasmin NLS, or a snurportin-1 importin-P NLS.
[0051] In certain aspects, the NmeCas9 polypeptide further comprises one or more additional heterologous functional domains. In one embodiment, the one or more additional heterologous functional domains is selected from an HMGB1 domain, a deaminase, a uracil glycosylase inhibitor (UGI), and a polymerase. In one embodiment, the one or more additional heterologous functional domains is a deaminase. In one embodiment, the deaminase is a cytidine deaminase or an adenine deaminase. On another embodiment, the cytidine deaminase is an apolipoprotein B mRNA editing enzyme (APOBEC) deaminase.
[0052] In certain aspects, the NmeCas9 polypeptide has cleavase activity. In one embodiment, the NmeCas9 polypeptide has increased editing activity as compared to a corresponding wild type NmeCas9 polypeptide.
[0053] In other aspects, provided is a polynucleotide comprising an open reading frame (ORF) that encodes any NmeCas9 polypeptide disclosed herein.
[0054] In one embodiment, the ORF has been codon optimized for increased translation of the mRNA in a mammal.
[0055] In one embodiment, the ORF has been codon optimized for increased translation of the mRNA in a human. In another embodiment, the increased translation is relative to the extent of translation of a wild type sequence of the ORF, or relative to an ORF having a codon distribution matching the codon distribution of the organism from which the ORF was derived.
[0056] In certain embodiments, the polynucleotide is an mRNA.
[0057] In further aspects, provided is a vector comprising a polynucleotide disclosed herein.
[0058] In one embodiment, the vector is a viral vector. In certain embodiments, the viral vector is an adeno-associated virus (AAV) vector.
[0059] In one embodiment, the vector further encodes one or more gRNAs. In certain embodiments, the one or more gRNAs is a single guide RNA (sgRNA). In other embodiments, the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA. In still other embodiments, the one or more gRNAs comprises a scaffold region that binds the NmeCas9 polypeptide, and a targeting region that hybridizes with a target genomic sequence in one or more cells of interest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide. In one embodiment, wherein the PAM sequence recognized by the NmeCas9 polypeptide is N4CC.
[0060] In additional aspects, provided is a cell comprising an Nme polypeptide described herein, a polynucleotide described herein, or a vector described herein.
[0061] In yet a further aspect, provided herein is a lipid nanoparticle (LNP) comprising a polynucleotide disclosed herein. In one embodiment, the LNP comprises an ionizable lipid.
[0062] Further provided is a composition comprising one or more guide RNAs (gRNAs); and an NmeCas9 polypeptide disclosed herein; a polynucleotide disclosed herein; a vector disclosed herein; or an LNP disclosed herein. In one embodiment, the one or more gRNAs is a single gRNA (sgRNA). In yet another embodiment, the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA. In one embodiment, the full-length sgRNA is a full-length Nme sgRNA as set forth in SEQ ID NO: 475. In one embodiment, the shortened gRNA is 101 nucleotides in length, 102 nucleotides in length, 103 nucleotides in length, 104 nucleotides in length, 105 nucleotides in length, or 115 nucleotides in length. In one embodiment, the shortened gRNA is 115 nucleotides in length. In yet another embodiment, the shortened sgRNA comprises a scaffold region that binds a NmeCas9 polypeptide, wherein the scaffold region comprises a polynucleotide sequence having at least 90% identity to any one of SEQ ID NOs 572-592.
[0063] In one embodiment, the scaffold region comprises a polynucleotide sequence having at least 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs 572-592. In one embodiment, the scaffold region comprises the polynucleotide sequence of any one of SEQ ID NOs: 572-592. In another embodiment, the scaffold region comprises atleast 90% identity to SEQ ID NO: 576. In still another embodiment, the scaffold region comprises at least 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 576. In a further embodiment, the scaffold region comprises SEQ ID NO: 576.
[0064] In certain aspects, the one or more gRNAs is a chemically modified gRNA.
[0065] In other aspects, the one or more gRNAs binds the NmeCas9 polypeptide and hybridizes with a target sequence in one or more cells of interest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide. In one embodiment, the PAM sequence recognized by the NmeCas9 polypeptide is N4CC.
[0066] In still another embodiment, the composition disclosed herein further comprises a template for a polymerase
[0067] In certain aspects, provide is a pharmaceutical composition comprising a NmeCas9 polypeptide disclosed herein; a polynucleotide disclosed herein; a vector disclosed herein; an LNP disclosed herein; or a composition disclosed herein; and a pharmaceutically acceptable carrier.
[0068] Also provided is a method for binding a target sequence in a cell comprising delivering to the cell: one or more guide RNAs (gRNAs); and a NmeCas9 polypeptide disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, an LNP disclosed herein, or a pharmaceutical composition disclosed herein; thereby binding the target sequence with the NmeCas9 polypeptide of the composition.
[0069] Further provided is a method for cleaving a target sequence in a cell comprising delivering to the cell: one or more guide RNAs (gRNAs); and an NmeCas9 polypeptide disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, an LNP disclosed herein, or a pharmaceutical composition disclosed herein; wherein the NmeCas9 polypeptide cleaves the target sequence.
[0070] In other aspects, provided is a method for modifying a target sequence in a cell comprising delivering to the cell: one or more guide RNAs (gRNAs); and an NmeCas9 polypeptide disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, an LNP disclosed herein, or a pharmaceutical composition disclosed herein, wherein the NmeCas9 polypeptide modifies the target sequence.
[0071] In one embodiment of a method disclosed herein, the one or more gRNAs is a single gRNA (sgRNA).
[0072] In one embodiment of a method disclosed herein, the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA. In one embodiment, the full-length sgRNA is a full-length Nme sgRNA as set forth in SEQ ID NO: 475.
[0073] In other embodiment of a method disclosed herein, the gRNA comprises a scaffold region that binds the NmeCas9 polypeptide and a targeting region that hybridizes with a target sequence in one or more cells of interest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide. In one embodiment, the PAM sequence recognized by the NmeCas9 polypeptide isN4CC.
[0074] In other aspects, provided is a method for binding a target sequence in a cell comprising delivering to the cell a vector disclosed herein or a composition disclosed herein, thereby binding the target sequence with the NmeCas9 polypeptide of the composition.
[0075] In still other aspects, provided is a method for cleaving a target sequence in a cell comprising delivering to the cell a vector disclosed herein or a composition disclosed herein, wherein the NmeCas9 polypeptide cleaves the target sequence.
[0076] In other embodiments, provided is a method for modifying a target sequence in a cell comprising delivering to the cell a vector disclosed herein or a composition disclosed herein, wherein the NmeCas9 polypeptide modifies the target sequence.
[0077] In certain embodiments, the method is an ex vivo method.
[0078] In other embodiments, the method is an in vivo method.
[0079] In certain aspects of a method disclosed herein, the target sequence is a target gene in a genome of the cell.
[0080] In one embodiment, the cell is a mammalian cell.
[0081] In another embodiment, the cell is a human cell.
[0082] In still another embodiment, the NmeCas9 of the invention is Nme2Cas9.
[0083] Further embodiments are provided throughout and described in the claims and Figures.BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Fig. 1A shows the crystal structure of wild-type Nme2Cas9 complexed with a sgRNA and a dsDNA, with exemplary amino acid positions of the Nme2Cas9 indicated. “NTS”: non-target strand of genomic DNA duplex. “TS”: target strand of genomic DNA duplex. Fig. IB shows the interactions of exemplary Nme2Cas9 amino acids with (i) asgRNA comprising a spacer, and (ii) a genomic dsDNA site comprising a target strand (“TS”) to which the spacer binds, and a non-target strand (“NTS”). Fig. 1C shows Nme2Cas9 domains; numbers indicate amino acid positions in the wild-type (WT) Nme2Cas9 protein (SEQ ID NO: 1). Bold residues refer to residues from the genomic locus that are complementary to the spacer of the sgRNA.
[0085] Figs. 2A-2D show editing at the PCSK9 locus in PMH treated with a mRNA encoding a Nme2Cas9 variant, along with a gRNA targeting the PCSK9 locus. Fig. 2A depicts in vitro editing by Nme2Cas9 variants including at least one of fusion with an HMGB1 domain or one of the following amino acid substitutions (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932K / K1044R, K145R, K266R, K333R, K390R, D418K, K419R, Q422K, Q422R, E508K, E508R, K517R, K549R, or K555R. Fig. 2B depicts in vitro editing by Nme2Cas9 variants including at least one of fusion with an HMGB1 domain or one of the following amino acid substitutions (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932K / K1044R, K53R, G197K, K351R, Q679R, L846K, L846Y, E868R, K870R, K929R, T930K, K962R, K963R, K965R, Q967R, K989R, K1005R, N1026K, S1051K, Q1053K, N1054K, orN1054R. Fig. 2C depicts in vitro editing by Nme2Cas9 variants including at least one of fusion with an HMGB1 domain or one of the following amino acid substitutions (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932K / K1044R, E932K, E932N, E932Q, E932M, E932R, E932H, E932A, E932D, E932S, and E932T. Fig. 2D depicts in vitro editing by Nme2Cas9 variants including at least one of fusion with an HMGB 1 domain or one of the following amino acid substitutions (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932K / K1044R, K1044R, K1044Q,K1044D, K1044A.
[0086] Fig. 3 shows dose response editing at the PCSK9 locus in PMH treated with a mRNA encoding a Nme2Cas9 N1026K, Nme2Cas9 N1026R, or WT Nme2Cas9, along with a gRNA targeting the PCSK9 locus.
[0087] Fig. 4 shows dose response editing at the PCSK9 locus in PMH treated with LNPs comprising a mRNA encoding a Nme2Cas9 N1026R variant or a wild-type Nme2Cas9, along with a gRNA targeting the PCSK9 locus.
[0088] Fig. 5 shows editing efficiencies at the PCSK9 locus in mouse liver, in animals treated with the 0.1 or 0.3 mpk (mg per kg of animal body weight) of LNPs comprising a gRNA targeting the PCSK9 locus, and a mRNA encoding a Nme2Cas9 N1026R variant or a wild-type Nme2Cas9.
[0089] Figs. 6A-6D show an alignment of the amino acid sequences for Nme2Cas9 (SEQ ID NO: 1), NmelCas9 (SEQ ID NO: 2), or Nme3Cas9 (SEQ ID NO: 3). The equivalent amino acid residues in NmelCas9 and Nme3Cas9 that correspond to N1026, K266, Q422, and D418 of Nme2Cas9 are notated by boxes.
[0090] Figs. 7A-7D show editing efficiency at the PCSK9 locus with a mRNA encoding a Nme2Cas9 and a gRNA targeting the PCSK9 locus. Fig. 7A shows dose response editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): at the positions indicated. Fig. 7B shows editing efficiency in vivo in mouse liver with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932R, E932N, E932M, or N1026R; or WT Nme2Cas9. Fig. 7C shows dose response editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932R, E932N, E932M, E932K, N1026R, enNme2.CC.NR (K104T, D152A, F260L, A263T, A303S, D451V, E932K, K1044R, Q1047R, V1056A), or iNme2 (K929R, K870R, E868K, D844G, E932K, D873A, D911G); or WT Nme2Cas9. Fig. 7D shows editing efficiency in vivo in mouse liver with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932K, N1026R, enNme2.CC.NR (K104T, D152A, F260L, A263T, A303S, D451V, E932K, K1044R, Q1047R, V1056A), or iNme2 (K929R, K870R, E868K, D844G, E932K, D873A, D911G); or WT Nme2Cas9.
[0091] Figs. 8A-8D show editing efficiency at the PCSK9 locus with a mRNA encoding a Nme2Cas9 and a gRNA targeting the PCSK9 locus. Fig. 8A shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932R either alone or in combination with an additional substitution; or WT Nme2Cas9. Fig. 8B shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932E either alone or in combination with two or more additional substitutions; or WT Nme2Cas9. Fig. 8C shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): E868R, E868K, E932R, N1026K, N1026R, either alone or in combination; or WT Nme2Cas9. Fig. 8D shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to thewild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): E932R either alone or in combination with an additional substitution; or WT Nme2Cas9.
[0092] Figs. 9A-9C show editing efficiency at the PCSK9 locus with a mRNA encoding a Nme2Cas9 and a gRNA targeting the PCSK9 locus. Fig. 9A shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): N1026R either alone or in combination with various E932 substitutions. Fig. 9B shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): N1026R either alone or in combination with one or more additional substitutions. Fig. 9C shows editing efficiency in PMH with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): N1026R either alone or in combination with one or more additional substitutions
[0093] Figs. 10A-10B show editing efficiency at the PCSK9 locus with a mRNA encoding a Nme2Cas9 and a gRNA targeting the PCSK9 locus in vivo in mouse liver and serum PCSK9 protein levels, respectively, in mice treated with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): N1026R or N1026R / K266R; or WT Nme2Cas9.
[0094] Fig. 11 shows editing efficiency at the PCSK9 locus in PMH with a mRNA encoding a Nme2Cas9 and a gRNA targeting the PCSK9 locus with an Nme2Cas9 with an amino acid substitution (i.e., relative to the wildtype Nme2Cas9 sequence set forth in SEQ ID NO: 1): atN1026R or N1026R / K266R.
[0095] Figs. 12A-12B show editing efficiencies obtained using a Nme2Cas9 variant with an amnio acid substitution (i.e., relative to the wild-type Nme2Cas9 sequence set forth in SEQ ID NO: 1): N1026R; or WT Nme2Cas9 with each of the 47 guides in Table 22A targeting 16 different loci, in primary human hepatocytes (PHH). Fig. 12A shows the aggregate editing efficiency of all guides tested. Fig. 12B shows the change in editing efficiency for each of the guides between Nme2Cas9 N1026R and WT Nme2Cas9.DETAILED DESCRIPTION
[0096] The present disclosure provides for variants of Neisseria meningitidis Cas9 (NmeCas9) that have enhanced editing activity, as well as related methods and compositions for using the enhanced NmeCas9 variants in genome editing.I. Definitions
[0097] Unless stated otherwise, the following terms and phrases as used herein are intended to have the following meanings:
[0098] The term “or combinations thereof’ as used herein refers to all permutations and combinations of the listed terms preceding the term. For example, “A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AAB, BBC, AAABC, CBBA, BABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
[0099] As used herein, the term “kit” refers to a packaged set of related components, such as one or more polynucleotides or compositions and one or more related materials such as delivery devices (e.g., syringes), solvents, solutions, buffers, instructions, or desiccants.
[0100] “Or” is used in the inclusive sense, i.e., equivalent to “and / or,” unless the context requires otherwise.
[0101] “Polynucleotide” and “nucleic acid” are used herein to refer to a multimeric compound comprising nucleosides or nucleoside analogs which have nitrogenous heterocyclic bases or base analogs linked together along a backbone, including conventional RNA, DNA, mixed RNA-DNA, and polymers that are analogs thereof. A nucleic acid “backbone” can be made up of a variety of linkages, including one or more of sugar phosphodiester linkages, peptide-nucleic acid bonds (“peptide nucleic acids” or PNA; PCT No. WO 95 / 32305), phosphorothioate linkages, methylphosphonate linkages, or combinations thereof. Sugar moieties of a nucleic acid can be ribose, deoxyribose, or similar compounds with substitutions, e.g., 2’ methoxy or 2’ halide substitutions. Nitrogenous bases can be conventional bases (A, G, C, T, U), analogs thereof (e.g., modified uridines such as 5-methoxyuridine, pseudouridine, orNl -methylpseudouridine, or others); inosine; derivatives of purines or pyrimidines (e.g., N4-methyl deoxy guanosine, deaza- or aza-purines, deaza- or aza-pyrimidines, pyrimidine bases with substituent groups at the 5 or 6 position (e.g., 5-methylcytosine), purine bases with a substituent at the 2, 6, or 8 positions, 2-amino-6-methylaminopurine, O6-methylguanine, 4-thio-pyrimidines, 4-amino-pyrimidines, 4-dimethylhydrazine-pyrimidines, and O4-alkyl-pyrimidines; US Pat. No. 5,378,825 and PCTNo. WO 93 / 13121). For general discussion see The Biochemistry of the Nucleic Acids 5-36, Adams et al., ed., 11th ed., 1992). Nucleic acids can include one or more “abasic” residues where the backbone includes no nitrogenous base for position(s) of the polymer (US Pat. No.5,585,481). A nucleic acid can comprise only conventional RNA or DNA sugars, bases and linkages, or can include both conventional components and substitutions (e.g., conventional bases with 2’ methoxy linkages, or polymers containing both conventional bases and one or more base analogs). Nucleic acid includes “locked nucleic acid” (LNA), an analogue containing one or more LNA nucleotide monomers with a bicyclic furanose unit locked in an RNA mimicking sugar conformation, which enhance hybridization affinity toward complementary RNA and DNA sequences (Vester and Wengel, 2004, Biochemistry 43(42): 13233-41). RNA and DNA have different sugar moieties and can differ by the presence of uracil or analogs thereof in RNA and thymine or analogs thereof in DNA.
[0102] “Polypeptide” as used herein refers to a multimeric compound comprising amino acid residues that can adopt a three-dimensional conformation. Polypeptides include but are not limited to enzymes, enzyme precursor proteins, regulatory proteins, structural proteins, receptors, nucleic acid binding proteins, antibodies, etc. Polypeptides may, but do not necessarily, comprise post-translational modifications, non-natural amino acids, prosthetic groups, and the like. In certain embodiments, polypeptides include only natural amino acids, which may include post-translational modifications. In some embodiments, a polypeptide may comprise an amino acid sequence provided by a SEQ ID NO listed herein (e.g., in Table 23) that includes an N-terminal methionine. In some embodiments, such a polypeptide may comprise the SEQ ID NO comprising the N-terminal methionine. In other embodiments, the polypeptide may comprise the SEQ ID NO: as listed, but without the N-terminal methionine. For example, a SEQ ID NO provided herein may, in some embodiments, be operably linked to another moiety at the N-terminus (e.g., a nuclear localization signal, a second polypeptide, etc.). In such embodiments, a person of skill in the art will recognize that the N-terminal methionine is optional. Accordingly, in SEQ ID NOs provided herein that contain an N-terminal methionine residue, the N-terminal methionine residue should be considered an optional feature. Likewise, for nucleic acid sequences provided herein which encode a protein or polypeptide that comprises an N-terminal methionine residue, the nucleotides encoding the N-terminal methionine residue should be considered an optional feature.
[0103] As used herein, “substitution” in a polypeptide sequence is understood as a change in the identity of an amino acid at a specific position without insertion or deletion of an amino acid. As used herein, a deletion or change in the identity of the initiator methionine in a polypeptide sequence, e.g., in SEQ ID NO: 1, 2, or 3, is not considered a mutation or a substitution.
[0104] “Cas nuclease”, also called “Cas protein”, as used herein, encompasses Cas cleavases (i.e., with double stranded DNA cleavage activity), Cas nickases, and dCas DNA binding agents (e.g., in which cleavase / nickase activity is inactivated).
[0105] As used herein, “NmeCas9” or “Nme Cas9” is generic and encompasses any type of NmeCas9, including, NmelCas9 (e.g., SEQ ID NO: 2), Nme2Cas9 (e.g., SEQ ID NO: 1), and Nme3Cas9 (e.g., SEQ ID NO: 3). Several Cas9 orthologs have been obtained from N. meningitidis (Esvelt et al., NAT. METHODS, vol. 10, 2013, 1116 - 1121; Hou et al., PNAS, vol. 110, 2013, pages 15644 - 15649) (NmelCas9, Nme2Cas9, and Nme3Cas9). The Nme2Cas9 ortholog functions efficiently in mammalian cells, recognizes an N4CC PAM, and can be used for in vivo editing (Ran et al., NATURE, vol. 520, 2015, pages 186 - 191; Kim et al., NAT. COMMUN., vol. 8, 2017, pages 14500). Nme2Cas9 has been shown to be naturally resistant to off-target editing (Lee et al., MOL. THER., vol. 24, 2016, pages 645 -654; Kim et al., 2017). See also e.g., WO / 2020081568 (e.g., pages 28 and 42), describing an Nme2Cas9 D16A nickase, the contents of which are hereby incorporated by reference in its entirety. Further, NmeCas9 variants are known in the art, see, e.g., Huang et al., Nature Biotech. 2022, doi.org / 10.1038 / s41587-022-01410-2, which describes Cas9 variants targeting single- nucleotide-pyrimidine PAMs. The sequence for wildtype NmelCas9 is set forth in SEQ ID NO: 2; the sequence for wildtype Nme2Cas9 is set forth in SEQ ID NO: 1; and the sequence for wildtype Nme3Cas9 is set forth in SEQ ID NO: 3.
[0106] As used herein, a “nickase” is an enzyme that creates a single-strand break (also known as a “nick”) in double strand DNA, i.e., cuts one strand but not the other of the DNA double helix. As used herein, an “RNA-guided nickase” means a polypeptide or complex of polypeptides having DNA nickase activity, wherein the DNA nickase activity is sequence-specific and depends on the sequence of the RNA. Exemplary RNA-guided nickases include Cas nickases. Cas nickases include, but are not limited to, nickase forms of a Csm or Cmr complex of a type III CRISPR system, the CaslO, Csml, or Cmr2 subunit thereof, a Cascade complex of a type I CRISPR system, the Cas3 subunit thereof, and Class 2 Cas nucleases. Class 2 Cas nickases include Class 2 Cas nuclease variants in which only oneof the two catalytic domains is inactivated, which have RNA-guided DNA nickase activity. Wild type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. Class 2 Cas nickases include, for example, NmeCas9 (e.g., with a D16A substitution in the RuvC domain or H588A substitution in the HNH domain of NmelCas9, Nme2Cas9, or Nme3Cas9), SpyCas9 (e.g., H840A, D10A, or N863 A variants of SpyCas9), Cpfl, C2cl, C2c2, C2c3, HF Cas9 (e.g., N497A, R661 A, Q695A, Q926A variants), HypaCas9 (e.g, N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g, K810A, KI 003 A, R1060A variants), and eSPCas9(l.l) (e.g, K848A, K1003A, R1060A variants) proteins and modifications thereof. Cpfl protein, Zetsche et al., Cell, 163: 1-13 (2015), is homologous to Cas9, and contains a RuvC-like protein domain. Cpfl sequences of Zetsche are incorporated by reference in their entirety. See, e.g., Zetsche, Tables SI and S3. “Cas9” encompasses 5. pyogenes (Spy) Cas9, the variants of Cas9 listed herein, and equivalents thereof. See, e.g., Makarova et al., Nat Rev Microbiol, 13(11): 722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015).
[0107] As used herein, the term “fusion protein” refers to a hybrid polypeptide which comprises polypeptides from at least two different proteins or sources. One polypeptide may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxyterminal (C- terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy -terminal fusion protein,” respectively. It is understood that, if one or more of the polypeptides includes an N-terminal methionine corresponding to a start codon, the methionine may be omitted from the fusion protein. For example, a polypeptide that is not located at the N-terminal portion in some cases may lack an N-terminal methionine and its coding sequence may similarly lack a start codon in its open reading frame that ordinarily would encode an N-terminal methionine, e.g. in the context of a wild-type or non-fusion protein. Such an omission is not considered a mutation or substitution as compared to a reference sequence, e.g., SEQ ID NO: 1, 2, or 3. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0108] As used herein, the term “operably linked” is understood as a juxtaposition of at least two components, e.g., polypeptide chain components or polynucleotide chain components, in a manner to allow one component to exert an effect on the other, or to allow the components when operably linked to have an effect that the two components could not exert individually or as a mixture of the components. In certain embodiments, two components may be operably linked by a covalent linkage, e.g., a peptide bond between two polypeptide chains encoded by a single open reading frame. In certain embodiments, the linkage can be mediated by a linker, e.g., a peptide linker. In certain embodiments, the linker is covalently attached to at least one of the components. In certain embodiments, the linker is covalently attached to both components. In certain embodiments, e.g., polynucleotides, the components can be joined by one or more phosphodiester or phosphorothioate bonds, e.g., a single PO or PS bond, or a polynucleotide linker sequence, or by other types of linkages, or any combination thereof.
[0109] The term “High Mobility Group Box 1” or “HMGB1,” as used herein in the context of the HMGB1 protein, refers to a non-histone, nuclear DNA-binding protein belonging to the High Mobility Group-Box superfamily, or a portion thereof. The wildtype HMGB1 protein is composed of 215 amino acids in three structural domains: a HMGB1 Box A DNA-binding domain (alternatively referred elsewhere herein as a HMGB1 Box A domain), a Box B DNA-binding domain (alternatively referred elsewhere herein as a HMGB1 Box B domain), and an acidic tail domain. The wildtype HMGB1 protein comprising SEQ ID NO: 928 may be understood to comprise, from N-terminus to C-terminus, a HMGB1 Box A DNA-binding domain comprising the sequence of amino acid residues 2-78 of SEQ ID NO: 928, a HMGB1 Box B DNA-binding domain comprising the sequence of amino acid residues 89-162 of SEQ ID NO: 928, a cryptic nuclear localization signal (NLS) comprising the sequence of amino acid residues 179-185 of SEQ ID NO: 928, and an acidic tail corresponding to amino acid residues 186-215 of SEQ ID NO: 928. The term “HMGB1 polypeptide” as used herein refers to a polypeptide comprising the amino acid sequence of HMGB1, or a portion or fragment thereof. For example, in some embodiments, an HMGB1 polypeptide can comprise one or more HMGB1 domains (such as an HMGB1 Box B domain, an HMGB1 Box A domain, and an acidic tail domain). It is understood that, when an HMGB1 polypeptide is present as part of a fusion protein, the HMGB1 polypeptide may be referred as “HMGB1” (e.g., “a fusion protein comprising, from N-terminus to C-terminus, deaminase-first linker-DNA-binding domain-heterologous NLS-second linker-HMGB1”). In some embodiments, the acidic tail may serve as a transcription stimulatory domain. “HMGB1” as used herein in the context of nucleic acids refers to a nucleic acid (e.g., DNA or mRNA) encoding an HMGB1 polypeptide. The human HMGB1 gene has accession number NC_000013.11 (30456704.30617597). It is understood that, when an HMGB1 polypeptide (e.g., an HMGB1 protein of SEQ ID NO: 928 is present as part of a fusion protein, the N-terminal methionine residue of HMGB1 may be omitted.
[0110] As used herein, in the context of the wildtype HMGB 1 protein, the term “Box A” refers to an amino acid sequence comprising or consisting of SEQ ID NO: 929In some embodiments, the HMGB1 Box A domain may comprise or consist of amino acid residues 2-78 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box A domain may have DNA binding activity.
[0111] As used herein, in the context of the wildtype HMGB 1 protein, the term “Box B” refers to an amino acid sequence comprising or consisting of SEQ ID NO: 930In some embodiments, the HMGB1 Box B domain may comprise or consist of amino acid residues 89-162 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box B domain may have DNA binding activity.
[0112] As used herein, in the context of the wildtype HMGB 1 protein, the term “cryptic nuclear localization signal” or “cryptic NLS” refers to an amino acid sequence comprising or consisting of EKSKKKK (SEQ ID NO: 931) or a variant thereof. In some embodiments, the cryptic NLS may comprise amino acid residues 179-185 of SEQ ID NO: 928. In some embodiments, the cryptic NLS may comprise additional amino acid residues. In some embodiments, the cryptic NLS may comprise a part of or all of amino acid residues 166-185 of SEQ ID NO: 928. In some embodiments, the cryptic NLS disclosed herein comprises a variant of EKSKKKK (SEQ ID NO: 931) wherein one amino acid residue of SEQ ID NO: 931 comprises a conservative substitution thereof. As used herein, a cryptic NLS, as provided by SEQ ID NO: 931, or as a variant thereof, is not a heterologous NLS described herein.
[0113] In some embodiments, the cryptic NLS may have constitutive NLS activity (that is, act as a signal fragment that mediates the nuclear import of the HMGB1 protein). In some embodiments, the cryptic NLS may not have constitutive NLS activity.
[0114] As used herein, in the context of the wildtype HMGB 1 protein, the term “receptor for advanced glycation end-products (RAGE) binding domain” or “RAGE binding domain” refers to an amino acid sequence comprising amino acid residues 150-183 of SEQID NO: 928. In some embodiments, the HMGB1 polypeptide described herein comprises a RAGE binding domain, or portion thereof, C-terminal to the HMGB1 Box B domain. In some embodiments, the RAGE binding domain comprises the amino acid sequence of SEQ ID NO: 932 C terminal to the HMGB1 Box B domain.
[0115] As used herein, in the context of the wildtype HMGB1 protein, the term “acidic tail domain” of an amino acid sequence comprising or consisting of SEQ ID NO: 934. In some embodiments, the acidic tail domain may comprise or consist of amino acid residues 186-215 of SEQ ID NO: 928. In some embodiments, the acidic tail domain may have transcriptional stimulatory function.
[0116] In certain embodiments, the HMGB 1 polypeptide is joined to the Cas9Nme by an intervening polypeptide sequence, e.g., a linker, an NLS.
[0117] As used herein, a “cytidine deaminase” means a polypeptide or complex of polypeptides that is capable of cytidine deaminase activity, that is catalyzing the hydrolytic deamination of cytidine or deoxy cytidine, typically resulting in uridine or deoxyuridine. Cytidine deaminases encompass enzymes in the cytidine deaminase superfamily, and in particular, enzymes of the APOBEC family (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005;Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274: 18470-6, 1999); Carrington et al., Cells 9:1690 (2020)). In some embodiments, variants of any known cytidine deaminase or APOBEC protein are encompassed. Variants include proteins having a sequence that differs from wild-type protein by one or several mutations (i.e., substitutions, deletions, insertions), such as one or several single point substitutions. For instance, a shortened sequence could be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids, preferably one to four amino acids at the C-terminus of the sequence. As used herein, the term “variant” refers to allelic variants, splicing variants, and natural or artificial mutants, which are homologous to a reference sequence. The variant is “functional” in that it shows a catalytic activity of DNA editing.
[0118] As used herein, the term “APOBEC3A” refers to a cytidine deaminase such as the protein expressed by the human A3 A gene. The APOBEC3 A may have catalytic DNA editing activity. An amino acid sequence of APOBEC3 A has been described (UniPROT accession ID: p31941) and is included herein as SEQ ID NO: 370. In some embodiments, the APOBEC3 A protein is a human APOBEC3 A protein or a wild-type protein. Variants includeproteins having a sequence that differs from wild-type APOBEC3 A protein by one or several mutations (i.e., substitutions, deletions, insertions), such as one or several single point substitutions. For instance, a shortened AP0BEC3A sequence could be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids, preferably one to four amino acids at the C-terminus of the sequence. As used herein, the term “variant” refers to allelic variants, splicing variants, and natural or artificial mutants, which are homologous to an APOBEC3 A reference sequence. The variant is “functional” in that it shows a catalytic activity of DNA editing. In some embodiments, an APOBEC3 A (such as a human APOBEC3 A) has a wild-type amino acid position 57 (as numbered in the wild-type sequence). In some embodiments, an APOBEC3 A (such as a human APOBEC3 A) has an asparagine at amino acid position 57 (as numbered in the wild-type sequence).
[0119] As used herein, the term “uracil glycosylase inhibitor”, “uracil-DNA glycosylase inhibitor” or “UGI” refers to a protein that is capable of inhibiting a uracil-DNA glycosylase (UDG) base-excision repair enzyme (e.g., UniProt ID: P14739; SEQ ID NO: 369).
[0120] The term “linker,” as used herein, refers to a chemical group or a molecule linking two adjacent molecules or moi eties. Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). Exemplary peptide linkers are disclosed elsewhere herein. In some embodiments, the linker is a peptide linker comprising an amino acid or a plurality of amino acids (e.g., a peptide or protein) such as a 16-amino acid residue “XTEN” linker, or a variant thereof (See, e.g., the Examples; and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009) or a GS linker. As used herein, the term “GS linker” refers to a linker sequence that is rich in glycine and serine, e.g., at least 50%, 60%, 70%, 80%, 90%, or 100% glycine or serine amino acid residues. In some embodiments, the linker comprises one or more sequences provided herein.
[0121] “Modified uridine” is used herein to refer to a nucleoside other than thymidine with the same hydrogen bond acceptors as uridine and one or more structural differences from uridine. In some embodiments, a modified uridine is a substituted uridine, i.e., a uridine in which one or more non-proton substituents (e.g., alkoxy, such as methoxy) takes the place of a proton. In some embodiments, a modified uridine is pseudouridine. Insome embodiments, a modified uridine is a substituted pseudouridine, i.e., a pseudouridine in which one or more non-proton substituents (e.g., alkyl, such as methyl) takes the place of a proton. In some embodiments, a modified uridine is any of a substituted uridine, pseudouridine, or a substituted pseudouridine, e.g., Nl-methyl-psuedouridine.
[0122] “Uridine position” as used herein refers to a position in a polynucleotide occupied by a uridine or a modified uridine. Thus, for example, a polynucleotide in which “100% of the uridine positions are modified uridines” contains a modified uridine at every position that would be a uridine in a conventional RNA (where all bases are standard A, U, C, or G bases) of the same sequence. Unless otherwise indicated, a U in a polynucleotide sequence of a sequence table or sequence listing in or accompanying this disclosure can be a uridine or a modified uridine.
[0123] As used herein, a first sequence is considered to “comprise a sequence with at least X% identity to” a second sequence if an alignment of the first sequence to the second sequence shows that X% or more of the positions of the second sequence in its entirety are matched by the first sequence. For example, the sequence AAGA comprises a sequence with 100% identity to the sequence AAG because an alignment would give 100% identity in that there are matches to all three positions of the second sequence. The differences between RNA and DNA (generally the exchange of uridine for thymidine or vice versa) and the presence of nucleoside analogs such as modified uridines do not contribute to differences in identity or complementarity among polynucleotides as long as the relevant nucleotides (such as thymidine, uridine, or modified uridine) have the same complement (e.g., adenosine for all of thymidine, uridine, or modified uridine; another example is cytosine and 5-methylcytosine, both of which have guanosine as a complement). Thus, for example, the sequence 5’-AXG where X is any modified uridine, such as pseudouridine, N1 -methyl pseudouridine, or 5-methoxy uridine, is considered 100% identical to AUG in that both are perfectly complementary to the same sequence (5’-CAU). Exemplary alignment algorithms are the Smith- Waterman and Needleman-Wunsch algorithms, which are well-known in the art. One skilled in the art will understand what choice of algorithm and parameter settings are appropriate for a given pair of sequences to be aligned; for sequences of generally similar length and expected identity >50% for amino acids or >75% for nucleotides, the Needleman-Wunsch algorithm with default settings of the Needleman-Wunsch algorithm interface provided by the EBI at the www.ebi.ac.uk web server is generally appropriate.
[0124] “mRNA” is used herein to refer to a polynucleotide that is RNA or modified RNA and comprises an open reading frame that can be translated into a polypeptide (i.e., can serve as a substrate for translation by a ribosome and amino-acylated tRNAs). mRNA can comprise a phosphate-sugar backbone including ribose residues or analogs thereof, e.g., 2’-methoxy ribose residues. In some embodiments, the sugars of an mRNA phosphate-sugar backbone consist essentially of ribose residues, 2’-methoxy ribose residues, or a combination thereof. In general, mRNAs do not contain a substantial quantity of thymidine residues (e.g., 0 residues or fewer than 30, 20, 10, 5, 4, 3, or 2 thymidine residues; or less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, or 0.1% thymidine content). An mRNA can contain modified uridines at some or all of its uridine positions.
[0125] As used herein, an “RNA-guided DNA binding agent” means a polypeptide or complex of polypeptides having RNA and DNA binding activity, or a DNA-binding subunit of such a complex, wherein the DNA binding activity is sequence-specific and depends on the sequence of the RNA. Exemplary RNA-guided DNA binding agents include Cas cleavases / nickases and inactivated forms thereof (“dCas DNA binding agents”). “Cas nuclease”, also called “Cas protein”, as used herein, encompasses Cas cleavases, Cas nickases, and dCas DNA binding agents. The dCas DNA binding agent may be a dead nuclease comprising non-functional nuclease domains (RuvC or HNH domain). In some embodiments the Cas cleavase or Cas nickase encompasses a dCas DNA binding agent modified to permit DNA cleavage, e.g., via fusion with a FokI domain. Exemplary nucleotide and polypeptide sequences of Cas9 molecules are provided below. Methods for identifying alternate nucleotide sequences encoding Cas9 polypeptide sequences, including alternate naturally occurring variants, are known in the art. Sequences with at least 75%, 80%, 85%, preferably 90%, 95%, 96%, 97%, 98%, or 99% identity to any of the Cas9 nucleic acid sequences, amino acid sequences, or nucleic acid sequences encoding the amino acid sequences provided herein are also contemplated.
[0126] As used herein, the “minimal uridine codon(s)” for a given amino acid is the codon(s) with the fewest uridines (usually 0 or 1 except for a codon for phenylalanine, where the minimal uridine codon has 2 uridines). Modified uridine residues are considered equivalent to uridines for the purpose of evaluating uridine content.
[0127] As used herein, the “uridine dinucleotide (UU) content” of an ORF can be expressed in absolute terms as the enumeration of UU dinucleotides in an ORF or on a rate basis as the percentage of positions occupied by the uridines of uridine dinucleotides (forexample, AUUAU would have a uridine dinucleotide content of 40% because 2 of 5 positions are occupied by the uridines of a uridine dinucleotide). Modified uridine residues are considered equivalent to uridines for the purpose of evaluating uridine dinucleotide content.
[0128] As used herein, the “minimal adenine codon(s)” for a given amino acid is the codon(s) with the fewest adenines (usually 0 or 1 except for a codon for lysine and asparagine, where the minimal adenine codon has 2 adenines). Modified adenine residues are considered equivalent to adenines for the purpose of evaluating adenine content.
[0129] As used herein, the “adenine dinucleotide content” of an ORF can be expressed in absolute terms as the enumeration of AA dinucleotides in an ORF or on a rate basis as the percentage of positions occupied by the adenines of adenine dinucleotides (for example, UAAUA would have an adenine dinucleotide content of 40% because 2 of 5 positions are occupied by the adenines of an adenine dinucleotide). Modified adenine residues are considered equivalent to adenines for the purpose of evaluating adenine dinucleotide content.
[0130] As used herein, the “minimum repeat content” of a given open reading frame (ORF) is the minimum possible sum of occurrences of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF that encodes the same amino acid sequence as the given ORF. The repeat content can be expressed in absolute terms as the enumeration of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF or on a rate basis as the enumeration of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF divided by the length in nucleotides of the ORF (for example, UAAUA would have a repeat content of 20% because one repeat occurs in a sequence of 5 nucleotides). Modified adenine, guanine, cytosine, thymine, and uracil residues are considered equivalent to adenine, guanine, cytosine, thymine, and uracil residues for the purpose of evaluating minimum repeat content.
[0131] As used herein, “open reading frame” or “ORF” of a gene refers to a sequence consisting of a series of codons that specify the amino acid sequence of the protein that the gene codes for. The ORF generally begins with a start codon (e.g., ATG in DNA or AUG in RNA) and ends with a stop codon, e.g., TAA, TAG or TGA in DNA or UAA, UAG, or UGA in RNA. The ORF sequences described herein may or may not include a start codon encoding an N-terminal methionine. Thus, in some cases, an ORF described in a SEQ ID NO herein comprises a start codon and the ORF includes the start codon. In other cases, an ORF comprises the sequence of the listed SEQ ID NO but without the start codon provided in the sequence of the listed SEQ ID NO. In some cases, an ORF described in a SEQ ID NO hereincomprises a stop codon and the ORF includes the stop codon. In other cases, an ORF comprises the sequence of the listed SEQ ID NO but without the stop codon provided in the sequence of the listed SEQ ID NO.
[0132] “Guide RNA”, “gRNA”, and “guide” are used herein interchangeably to refer to either a crRNA (also known as CRISPR RNA), or the combination of a crRNA and a trRNA (also known as tracrRNA). The crRNA and trRNA may be associated as a single RNA molecule (single guide RNA, sgRNA) or in two separate RNA molecules (dual guide RNA, dgRNA). “Guide RNA” or “gRNA” refers to each type. The trRNA may be a naturally occurring sequence, or a trRNA sequence with modifications or variations compared to naturally occurring sequences. Guide RNAs can include modified RNAs as described herein. Unless otherwise clear from the context, guide RNAs described herein are suitable for use with an NmeCas9, e.g., an Nmel, Nme2, or Nme3 Cas9.
[0133] As used herein, a “guide sequence” or “guide region” or “spacer” or “spacer sequence” or “spacer region” and the like refers to a sequence within a guide RNA that is complementary to a target sequence and functions to direct a guide RNA to a target sequence for binding or modification (e.g., cleavage) by NmeCas9. A guide sequence can be 20-25 nucleotides in length, e.g., in the case of NmeCas9 and related Cas9 homologs / orthologs. Shorter or longer sequences can also be used as guides, e.g., 20-, 21-, 22-, 23-, 24-, or 25 -nucleotides in length. A guide sequence can be at least 22-, 23-, 24-, or 25-nucleotides in length in the case of NmeCas9. A guide sequence can form a 22-, 23-, 24, or 25-continuous base pair duplex, e.g., a 24-continuous base pair duplex, with its target sequence in the case of NmeCas9. In certain embodiments, a sgRNA can comprise a spacer region or guide region, and a “structural region” or “scaffold region” positioned 3’ of the spacer region. In some embodiments, the structural or scaffold region can comprise a repeat / anti-repeat region, a hairpin 1 region, a hairpin2 region, and (optionally) a tail region. This structural or scaffold region can interact with a Cas9 nuclease (e.g., NmeCas9). Accordingly, the spacer or guide region of a sgRNA can perform certain functions of a crRNA, and the structural or scaffold region of a sgRNA can perform certain functions of a trRNA.
[0134] Target sequences for Cas proteins include both the positive and negative strands of genomic DNA (i.e., the sequence given and the sequence’s reverse compliment), as a nucleic acid substrate for a Cas protein is a double stranded nucleic acid. Accordingly, where a guide sequence is said to be “complementary to a target sequence”, it is to be understood that the guide sequence may direct a guide RNA to bind to the reversecomplement of a target sequence. Thus, in some embodiments, where the guide sequence binds the reverse complement of a target sequence, the guide sequence is identical to certain nucleotides of the target sequence (e.g., the target sequence not including the PAM) except for the substitution of U for T in the guide sequence.
[0135] As used herein, a “shortened” region in a gRNA is a conserved region of a gRNA that lacks at least 1 nucleotide compared to the corresponding conserved region shown in Table 1. Similarly, “shortened” with respect to an sgRNA means that its conserved region comprises fewer nucleotides than the sgRNA conserved region shown in Table 2. Under no circumstances does “shortened” imply any particular limitation on a process or manner of production of the gRNA.
[0136] The term a “conserved region” in the context of a shortened sgRNA refers to a conserved region of an N. meningitidis Cas9 (“NmeCas9”) gRNA as shown in Table 2. The first row shows the numbering of the nucleotides; the second row shows an exemplary sequence (e.g., SEQ ID NO: 475); and the third and fourth rows show the regions. Shortened conserved regions lack at least one nucleotide shown in Table 2, as discussed in detail below.
[0137] As used herein, a “gene editing” or “genetic modification” is a change at the DNA level or changes in DNA expression, e.g., induced by a gRNA / Cas complex. Gene editing or genetic modification may comprise an insertion, deletion, or substitution (base substitution, e.g., C-to-T, or point mutation), typically within a defined sequence or genomic locus. A genetic modification changes the nucleic acid sequence of the DNA. A genetic modification may be at a single nucleotide position. A genetic modification may be at multiple nucleotides, e.g., 2, 3, 4, 5 or more nucleotides, typically in close proximity to each other, e.g., contiguous nucleotides. In some embodiments, the method or use results in gene editing. In some embodiments, the method or use results in a double-stranded break within the target gene. In some embodiments, the method or use results in formation of indel mutations during non- homologous end joining of the DSB. In some embodiments, the method or use results in an insertion or deletion of nucleotides in a target gene. In some embodiments, the insertion or deletion of nucleotides in a target gene leads to a frameshift mutation or premature stop codon that results in a non-functional protein. In some embodiments, the insertion or deletion of nucleotides in a target gene leads to a knockdown or elimination of target gene expression. In some embodiments, the method or use comprises homology directed repair of a DSB. In some embodiments, the method or use further comprises delivering to the cell a template, wherein at least a part of the templateincorporates into a target DNA at or near a double strand break site induced by the nuclease. In some embodiments, the method or use results in a single strand break within the target gene. In some embodiments, the method or use results in a base change, e.g., by deamination, within the target gene. The gene editing typically occurs within or adjacent to the portion of the target gene with which the spacer sequence forms a duplex.
[0138] The terms “cleave” or “cleavage” refer to the hydrolysis of at least one phosphodiester bond within the backbone of one or both strands of a double- stranded target sequence (e.g., target DNA sequence) that can result in either single-stranded or doublestranded breaks within the target sequence.
[0139] “Editing efficiency,” “genome editing activity,” “editing percentage,” or “percent editing” as used herein is the total number of sequence reads with insertions, deletions, or base changes of nucleotides into the target region of interest over the total number of sequence reads following genetic modification by a Cas RNP. In the context of catalytically inactive Cas9 variants (dCas9 variants), genome editing activity can alternatively be measured by alterations in gene expression relative to a control condition.
[0140] An “enhancement” or “increase” in genome editing activity, editing efficiency, or editing potency refers to an increase in on-target genome editing activity relative to a reference level. In the context of NmeCas9 variants comprising an amino acid substitution, an increase in genome editing activity may be measured relative to wildtype NmeCas9. The skilled person will understand that such a comparison is to be made with appropriate controls, e.g., by providing an NmeCas9 variant and a wildtype NmeCas9 in equivalent amounts and under the same or highly similar conditions. Amino acid substitutions described herein that enhance the genome editing activity of NmeCas9 relative to wildtype NmeCas9 are alternatively referred to as “enhancing” amino acid substitutions (e.g., one or more amino acid substitutions described in Table 3, Table 4, or Table 5) and the resulting NmeCas9 variants are alternatively referred to as “enhanced NmeCas9” variants.
[0141] Comparison of double-stranded cleavase activity, e.g., efficiency, between two NmeCas9 polypeptides can be performed, for example, using a dilution series and assay method in primary hepatocytes as in Example 7 in view of Example 1. The species of hepatocytes is selected based on the guide RNA used. The upper and lower end of the dilution series are selected to provide an appropriate curve to allow for determination of EC50 of the double-stranded cleavase using NGS analysis using methods known in the art. The activity of an Nme2Cas9 variant is compared to its corresponding wild-type Nme2Cas9.Such considerations related to selection of an appropriate control are well understood in the art.
[0142] In certain embodiments, the test NmeCas9 polypeptide has increased editing efficiency as compared to the wildtype NmeCas9 by at least 20%. In certain embodiments, the testNmeCas9 polypeptide has increased editing efficiency of at least 50%, preferably at least 100%. In certain embodiments, the editing efficiency of the testNmeCas9 polypeptide is increased by at least 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, or 10-fold as compared to a wildtype NmeCas9 polypeptide.
[0143] As used herein, a “transition” is a substitution mutation in which a purine nucleotide is replaced with a different purine nucleotide or a pyrimidine nucleotide is replaced with a different pyrimidine nucleotide. Exemplary transition mutations include A-to-G, G-to-A, C-to-T, and T-to-C.
[0144] As used herein, a “transversion” is a substitution mutation in which a purine nucleotide is replaced with a pyrimidine nucleotide or a pyrimidine nucleotide is replaced with a purine nucleotide. Exemplary transversion mutations include A-to-C, C-to-A, G-to-T, T-to-G, A-to-T, T-to-A, G-to-C, and C-to-G.
[0145] As used herein, “indel” refers to an insertion or deletion mutation consisting of a number of nucleotides that are either inserted, deleted, or inserted and deleted, e.g., at the site of a nick or double-strand break (DSB), in a target nucleic acid. As used herein, when indel formation results in an insertion, the insertion is a random insertion at the site of a nick or DSB and is not directed by or based on a template sequence.
[0146] As used herein, “base editing” is understood as a change in a nucleotide sequence promoted by a deamination reaction close to the site of a nick in a double stranded DNA molecule. That is, the change in the nucleotide sequence resulting from base editing is template independent and is not dependent on a polymerase. Deaminase reactions are catalyzed, for example, by a cytosine deaminase or an adenosine deaminase. The editing method is limited by both the nucleotide sequence present in the genome and the base changes that can be made through deamination, e.g., C to T and A to G. The compositions and methods provided herein do not include a deaminase.
[0147] As used herein, “knockdown” refers to a decrease in expression of a particular gene product (e.g., protein, mRNA, or both). Knockdown of a protein can be measured either by detecting protein secreted by tissue or population of cells (e.g., in serum or cell media) or by detecting total cellular amount of the protein from a tissue or cell population of interest.Methods for measuring knockdown of mRNA are known and include sequencing of mRNA isolated from a tissue or cell population of interest. In some embodiments, “knockdown” may refer to some loss of expression of a particular gene product, for example a decrease in the amount of mRNA transcribed or a decrease in the amount of protein expressed or secreted by a population of cells (including in vivo populations such as those found in tissues).
[0148] As used herein, “knockout” refers to a loss of expression of a particular protein in a cell. Knockout can be measured either by detecting the amount of protein secretion from a tissue or population of cells (e.g., in serum or cell media) or by detecting total cellular amount of a protein a tissue or a population of cells. In some embodiments, the methods of the disclosure “knockout” a target protein one or more cells (e.g., in a population of cells including in vivo populations such as those found in tissues). In some embodiments, a knockout is not the formation of mutant of the target protein, for example, created by indels, but rather the complete loss of expression of the target protein in a cell.
[0149] As used herein, “ribonucleoprotein” (RNP) or “RNP complex” refers to a guide RNA together with an RNA-guided DNA binding agent, such as a Cas cleavase, nickase, or dCas DNA binding agent (e.g., Cas9). In some embodiments, the guide RNA guides the RNA-guided DNA binding agent such as Cas9 to a target sequence, and the guide RNA hybridizes with and the agent binds to the target sequence; in cases where the agent is a cleavase or nickase, binding can be followed by cleaving or nicking.
[0150] As used herein, a “target sequence” refers to a sequence of nucleic acid in a target gene that has complementarity to the guide sequence of the gRNA. The interaction of the target sequence and the guide sequence directs an RNA-guided DNA binding agent to bind, and potentially nick or cleave (depending on the activity of the agent), within the target sequence.
[0151] In some embodiments, the target sequence may be adjacent to a PAM. In some embodiments, the PAM may be adjacent to or within 1, 2, 3, or 4, nucleotides of the 3' end of the target sequence. The length and the sequence of the PAM may depend on the Cas protein used. For example, the PAM may be selected from a consensus or a particular PAM sequence for a specific NmeCas9 protein or NmeCas9 ortholog (Edraki et al., 2019). In some embodiments, the PAM may comprise 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length.Nonlimiting exemplary PAM sequences include NCC, N4GAYW, N4GYTT, N4GTCT, NNNNCC(a), NNNNCAAA (wherein N is defined as any nucleotide, W is defined as eitherA or T, and R is defined as either A or G; and (a) is a preferred, but not required, A after the second C)). In some embodiments, the PAM sequence may be NCC.
[0152] As used herein, “treatment” refers to any administration or application of a therapeutic for disease or disorder in a subject, and includes slowing or arresting disease development or progression, relieving one or more signs or symptoms of the disease, curing the disease, or preventing reoccurrence of one or more symptoms of the disease.
[0153] As used herein, the term “lipid nanoparticle” (LNP) refers to a particle that comprises a plurality of (i.e., more than one) lipid molecules physically associated with each other by intermolecular forces. The LNPs may include a lipid component comprising (i) an ionizable lipid for encapsulation and for endosomal escape, and one or more (i.e., one, two, or three) of the following: (ii) a neutral lipid for stabilization (e.g., a phospholipid), (iii) a helper lipid, also for stabilization (e.g. cholesterol), and (iv) a structural PEG lipid for reducing particle aggregation and controlling particle size. The LNPs may encapsulate a cargo composition, including, but not limited to nucleic acids (RNA or DNA), small molecules, biologies, and other chemical species. Any LNP known to those of skill in the art to be capable of delivering nucleotides to subjects may be utilized with the guide RNAs and the nucleic acid encoding an RNA-guided DNA binding agent described herein.
[0154] As used herein, the terms “nuclear localization signal” (NLS) or “nuclear localization sequence” refers to an amino acid sequence which induces transport of molecules comprising such sequences or linked to such sequences into the nucleus of eukaryotic cells. The nuclear localization signal may form part of the molecule to be transported. In some embodiments, the NLS may be linked to the remaining parts of the molecule by covalent bonds, hydrogen bonds or ionic interactions.
[0155] As used herein, “delivering” and “administering” are used interchangeably, and include ex vivo and in vivo applications.
[0156] Co-admini strati on, as used herein, means that a plurality of substances are administered sufficiently close together in time so that the agents act together.Coadministration encompasses administering substances together in a single formulation and administering substances in separate formulations close enough in time so that the agents act together.
[0157] As used herein, the phrase “pharmaceutically acceptable” means that which is useful in preparing a pharmaceutical composition that is generally non-toxic and is notbiologically undesirable and that are not otherwise unacceptable for pharmaceutical use. Pharmaceutically acceptable generally refers to substances that are non-pyrogenic.
[0158] Pharmaceutically acceptable can refer to substances that are sterile, especially for pharmaceutical substances that are for injection or infusion.
[0159] Reference will now be made in detail to certain embodiments of the invention, examples of which are illustrated in the accompanying drawings. While the invention will be described in conjunction with the illustrated embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the invention as defined by the appended claims.
[0160] Before describing the present teachings in detail, it is to be understood that the disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a conjugate” includes a plurality of conjugates and reference to “a cell” includes a plurality of cells and the like.
[0161] Numeric ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with the measurement. Accordingly, unless indicated to the contrary, the numerical parameters set forth in the following specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0162] The term “about” is used herein to mean within the typical ranges of tolerances in the art. For example, “about” can be understood as about 2 standard deviations from the mean. In certain embodiments, about means +10%. In certain embodiments, about means +5%, +2%, or +1%. When about is present before a series of numbers or a range, it is understood that “about” can modify each of the numbers in the series or range. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0163] The use of “comprise”, “comprises”, “comprising”, “contain”, “contains”, “containing”, “include”, “includes”, and “including” are not intended to be limiting. It is to be understood that both the foregoing general description and detailed description are exemplary and explanatory only and are not restrictive of the teachings.
[0164] The term “at least” prior to a number or series of numbers is understood to include the number adjacent to the term “at least”, and all subsequent numbers or integers that could logically be included, as clear from context. For example, the number of nucleotides in a nucleic acid molecule must be an integer. For example, “at least 17 nucleotides of a 20 nucleotide nucleic acid molecule” means that 17, 18, 19, or 20 nucleotides have the indicated property. When at least is present before a series of numbers or a range, it is understood that “at least” can modify each of the numbers in the series or range.
[0165] As used herein, “no more than” or “less than” is understood as the value adjacent to the phrase and logical lower values or integers, as logical from context, to zero. For example, a duplex region of “no more than 2 nucleotide base pairs” has 2, 1, or 0 nucleotide base pairs. When “no more than” or “less than” is present before a series of numbers or a range, it is understood that each of the numbers in the series or range is modified.
[0166] As used herein, ranges include both the upper and lower limits.
[0167] As used herein, it is understood that when the maximum amount of a value is represented by 100% (e.g., 100% inhibition) that the value is interpreted in light of the method of detection. For example, 100% inhibition, and the like, is understood as inhibition to a level below the level of detection of the assay.
[0168] Unless specifically noted in the above specification, embodiments in the specification that recite “comprising” various components are also contemplated as “consisting of or “consisting essentially of the recited components; embodiments in the specification that recite “consisting of various components are also contemplated as “comprising” or “consisting essentially of the recited components; and embodiments in the specification that recite “consisting essentially of various components are also contemplated as “consisting of or “comprising” the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0169] The section headings used herein are for organizational purposes only and are not to be construed as limiting the desired subject matter in any way. In the event that any literature incorporated by reference contradicts the express content of this specification,including but not limited to a definition, the express content of this specification controls. While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.II. NmeCas9 Polypeptides
[0170] Provided herein are variants of Neisseria meningitidis Cas9 (NmeCas9) polypeptides that have enhanced genome editing activity relative to wildtype NmeCas9, as well as related methods and compositions for using the enhanced NmeCas9 variants in genome editing in a cell (e.g., ex vivo or in vivo).II.A. RNA-guided DNA binding agent; NmeCas9
[0171] The NmeCas9 polypeptides described herein comprise amino acid substitutions relative to the amino acid sequence of a wildtype NmeCas9 polypeptide. In certain embodiments, the wildtype NmeCas9 is Nme2Cas9 (SEQ ID NO: 1), NmelCas9 (SEQ ID NO: 2), or Nme3Cas9 (SEQ ID NO: 3).
[0172] In certain embodiments, the NmeCas9 polypeptide comprises one or more amino acid substitutions that enhance an activity of the NmeCas9 polypeptide (e.g., DNA binding, cleaving, or nicking) relative to wild-type NmeCas9.
[0173] In some embodiments, the NmeCas9 polypeptide is an Nme2Cas9 polypeptide and comprises at least one amino acid substitution. In certain embodiments, the Nme2Cas9 polypeptide (e.g., an enhanced Nme2Cas9 variant) comprises an enhancing amino substitution described herein and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprise one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, L846, E868, K870, K929, T930, E932, K962, K963, K965, Q967, K989, K1005, K1026, K1044, S1051, Q1053, or N1054 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide isan Nme2Cas9 polypeptide and comprises at least one amino acid substitution described in Table 3 (i.e, K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K / R, E508K / R, K517R, K549R, K555R, Q679R, L846K / Y, E868K / R, K870R, K929R, T930K, E932K / R, K962R, K963R, K965R, Q967R, K989R, K1005R, N1026K / R, K1044R, S1051K, Q1053K, or N1054K / R relative to SEQ ID NO: 1).
[0174] In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the Nme2Cas9 polypeptide comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the Nme2Cas9 polypeptide comprises one, two, three, four, five, six, seven, eight, nine, ten, or more than ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide).
[0175] In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than three substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than two substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide contains only one substitution relative to SEQ ID NO: 1. In some embodiments, the Nme2Cas9 polypeptide may comprise additional mutations at positions other than those listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1.
[0176] In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418,E932, E868, K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5, S6, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0177] In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 wherein the substitutions comprises at least one position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, and also comprises at least one substitution selected from the group consisting of S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5, S6, and E520. In certain embodiments, two substitutions are at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0178] In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to SEQ ID NO: 1 wherein the substitutions comprised at least one positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, and at least one position selected from the group consisting of S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5, S6, and E520 relative to SEQ ID NO: 1. . Optionally, the two substitutions are at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0179] In some embodiments, the Nme2Cas9 polypeptide is a nickase comprising a RuvC substitution or an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 2 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the nickase substitution, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional substitutions relative to SEQ ID NO: 1, does not contain more than two additional substitutions relative to SEQ ID NO: 1, contains only one additional substitution relative to SEQ ID NO: 1.
[0180] In some embodiments, the Nme2Cas9 polypeptide is a dCas9 comprising a RuvC substitution and an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 2 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, S6 , G33, E47, R63,V68, K104, Al 16,T123, DI 52, El 54, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the dCas9 substitutions, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional substitutions relative to SEQ ID NO: 1, does not contain more than two additional substitutions relative to SEQ ID NO: 1, contains only one additional substitution relative to SEQ ID NO: 1.
[0181] In some embodiments, the Nme2Cas9 polypeptide comprises an amino acid sequence having at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 32, 65, 87, 98, 307, or 318 (as shown in Table 23).
[0182] In some embodiments, the NmeCas9 polypeptide is an NmelCas9 polypeptide and comprises at least one amino acid substitution. In certain embodiments, the NmelCas9 polypeptide (e.g., an enhanced NmelCas9 variant) comprises an enhancing amino substitution described herein and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype Nmel Cas9 polypeptide). In some embodiments, the amino acid sequence of the NmelCas9 polypeptide comprise one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, S197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, P1003, S1022, K1040, G1051, K1053, T1054 of SEQ ID NO: 2 (i.e., a wildtype NmelCas9polypeptide). In some such embodiments, the NmeCas9 polypeptide is an NmelCas9 polypeptide and comprises at least one amino acid substitution described in Table 4 (i.e., K53R, K145R, S197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, P1003K / R, S1022R, K1040R, G1051K, K1053R, or T1054K / R relative to SEQ ID NO: 2). In certain embodiments, theNmelCas9 polypeptide comprises an amino acid substitution at position S1022 of SEQ ID NO: 2 (e.g., an S1022R substitution). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide has four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some embodiments, the amino acid sequence of the NmelCas9 polypeptide has four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 4 relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some embodiments, the Nme2Cas9 polypeptide may comprise additional substitutions at positions other than those listed in Table 4 relative to the amino acid sequence of SEQ ID NO: 2. In some embodiments, the NmelCas9 polypeptide may comprise additional substitutions at positions other than those listed in Table 4 relative to the amino acid sequence of SEQ ID NO: 2.
[0183] In some embodiments, the amino acid sequence of the NmelCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some embodiments, the NmelCas9 polypeptide comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some embodiments, theNmelCas9 polypeptide comprises one, two, three, four, five, six, seven, eight, nine, ten, or more than ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype Nmel Cas9 polypeptide).
[0184] In some embodiments, the NmeCas9 polypeptide is an Nme3Cas9 polypeptide and comprises at least one amino acid substitution. In certain embodiments, the Nme3Cas9 polypeptide (e.g., an enhanced Nme3Cas9 variant) comprises an enhancing amino substitution described herein and comprises at least 70%, at least 75%, at least 80%, at least85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide comprise one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, S1003, G1022, S1040, G1050, K1052, or T1053 of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an Nme3Cas9 polypeptide and comprises at least one amino acid substitution described in Table 5 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, S1003K / R, G1022R, S1040K / R, G1050K, K1052R, or T1053K / R substitution relative to SEQ ID NO: 3). In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide has four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide has four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 5 relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, theNme3Cas9 polypeptide may comprise additional substitutions at positions other than those listed in Table 5 relative to the amino acid sequence of SEQ ID NO: 3.
[0185] In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the Nme3Cas9 polypeptide comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, theNmelCas9 polypeptide comprises one, two, three, four, five, six, seven, eight, nine, ten, or more than ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide).
[0186] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprise one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 of SEQ ID NO: 1.
[0187] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions N1026, K266, Q422, and D418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises one substitution at a position selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises two substitutions at positions selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises three substitutions at positions selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions at positions N1026, K266, Q422, and D418 of SEQ ID NO: 1.
[0188] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprise one or more (e.g., one or more, two or more, three or more, or four or more) amino acid substitutions at positions N1026, E932, E868, K929, K266, Q422, and D418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises one substitution at a position selected from N1026, E932, E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises two substitutions at positions selected from N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises three substitutions at positions selected from N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises four amino acid substitutions at positions N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the one or more amino acid substitutions comprises a substitution at position N1026 of SEQ ID NO: 1. In some embodiments, the one or more amino acid substitutions comprises a substitution at position E932 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9polypeptide has four or fewer amino acid substitutions relative SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide has four or fewer amino acid substitutions listed in Table 3 in SEQ ID NO: 1. In some embodiments, the Nme2Cas9 polypeptide may comprise additional mutations at positions other than those listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1.
[0189] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position N1026 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an N1026X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an N1026X substitution relative to SEQ ID NO: 1, where X is a basic amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a N1026K or N1026R substitution relative to SEQ ID NO: 1.
[0190] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises aN1026K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution N1026K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 307. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:307 (Nme2Cas9(N1026K)). In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a N1026R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution N1026R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 318. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:318 (Nme2Cas9(N1026R)).
[0191] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position K266 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a K266X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a K266X substitution relative to SEQ ID NO: 1, where X is a basic amino acid.
[0192] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a K266R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution K266R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 32. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:32 (Nme2Cas9(K266R)).
[0193] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position Q422 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an Q422X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a Q422X substitution relative to SEQ ID NO: 1, where X is a basic amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution selected from Q422K or Q422R relative to SEQ ID NO: 1.
[0194] In certain embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution Q422K relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution Q422K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 87. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO: 87 (Nme2Cas9(Q422K)).
[0195] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution Q422R relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution Q422R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 98. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:98 (Nme2Cas9(Q422R)).
[0196] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position D418 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an D418X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, theamino acid sequence of the NmeCas9 polypeptide comprises an D418X substitution relative to SEQ ID NO: 1, where X is a basic amino acid. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution selected from D418K or D418R relative to SEQ ID NO: 1.
[0197] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution D418K relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution D418K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 65. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:65 (Nme2Cas9(D418K)).In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution at E932 relative to SEQ ID NO: 1, e.g., to K, N, Q, M, R, H, A, S, or T; optionally to K, N, M, or T. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution at E932 and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid sequence comprises amino acids 40-1120 of SEQ ID NOS: 654, 659, 664, 669, 674, 679, 684, 694, or 699.
[0198] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution E868 relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution E868K or E868R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 1. the amino acid sequence comprises amino acids 40-1120 of SEQ ID NOS 714 or 774.
[0199] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution K929R relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises the amino acid substitution K929R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 704. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO: 704 (Nme2Cas9(K929R)). Insome embodiments, calculation of identity is relative to amino acids 40-1120 of SEQ ID NO: 704 and does not consider amino acids 1-39 of SEQ ID NO: 704 which includes heterologous sequences, e.g., NLS and linker sequences.
[0200] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266X, where X is any amino acid;N1026X and Q422X, where X is any amino acid;N1026X and D418X, where X is any amino acid;K266X and Q422X, where X is any amino acid;K266X and D418X, where X is any amino acid;Q422X and D418X, where X is any amino acid;E932Y and N1026X, where Y is any amino acid except D, and where X is any amino acid;E932Y and E868X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K266X, where Y is any amino acid except D, and where X is any amino acid;E932Y and D418X, where Y is any amino acid except D, and where X is any amino acid;E932Y and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K333X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K1044X, where Y is any amino acid except D, and where X is any amino acid; orE868X and N1026X, where Y is any amino acid except D, and where X is any amino acid.
[0201] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266X, where X is a basic amino acid;N1026X and Q422X, where X is a basic amino acid;N1026X and D418X, where X is a basic amino acid;K266X and Q422X, where X is a basic amino acid;K266X and D418X, where X is a basic amino acid; orQ422X and D418X, where X is a basic amino acid;E932Y and N1026X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and E868X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K266X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and D418X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K333X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K1044X, where Y is any amino acid except D, and where X is a basic amino acid; orE868X and N1026X, where Y is any amino acid except D, and where X is a basic amino acid.
[0202] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266X, where X is K or R;N1026X and Q422X, where X is K or R;N1026X and D418X, where X is K or R;K266X and Q422X, where X is K or R;K266X and D418X, where X is K or R;Q422X and D418X, where X is K or R;E932Y and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and E868X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and K266X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and D418X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y and Q422X, where Y is N, Q, M, H, A, S, T, or K and where X is K or R; E932Y and K333X, where Y is N, Q, M, H, A, S, T, or K and where X is K or R; E932Y and K1044X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; or E868X and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R.
[0203] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X, K266X, and Q422X, where X is any amino acid;N1026X, K266X, and D418X, where X is any amino acid;N1026X, Q422X, and D418X, where X is any amino acid;K266X, Q422X, and D418X, where X is any amino acid;E932Y, K266X, and D418X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;K266X, Q422X, and E868X, where X is any amino acid;Q422X, E868X, and N1026X, where X is any amino acid;K266X, E868X, and N1026X, where X is any amino acid;E868X, E932Y, N1026X, where Y is any amino acid except D, and where X is any amino acid;E932Y, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E868X, E932Y, and N1026X, where Y is any amino acid except D, and where X is any amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;D418X, E868S, and N1026X, where X is any amino acid;E932Y, K266X, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, K333X, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is any amino acid; orK266X, Q422X, E868X, and N1026X, where X is any amino acid.
[0204] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X, K266X, and Q422X, where X is a basic amino acid;N1026X, K266X, and D418X, where X is a basic amino acid;N1026X, Q422X, and D418X, where X is a basic amino acid;K266X, Q422X, and D418X, where X is a basic amino acid;E932Y, K266X, and D418X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, and Q422X, where Y is any amino acid except D, and where X is any basic amino acid;K266X, Q422X, and E868X, where X is a basic amino acid;Q422X, E868X, and N1026X, where X is a basic amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;E868X, E932Y, N1026X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E868X, E932Y, and N1026X, where Y is any amino acid except D, and where X is a basic amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;D418X, E868S, and N1026X, where X is a basic amino acid;E932Y, K266X, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, K333X, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is a basic amino acid; orK266X, Q422X, E868X, and N1026X, where X is a basic amino acid.
[0205] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X, K266X, and Q422X, where X is K or R;N1026X, K266X, and D418X, where X is K or R;N1026X, Q422X, and D418X, where X is K or R;K266X, Q422X, and D418X, where X is K or R;E932Y, K266X, and D418X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;K266X, Q422X, and E868X, where X is K or R;Q422X, E868X, and N1026X, where X is K or R;K266X, E868X, and N1026X, where X is K or R;E868X, E932Y, N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E868X, E932Y, N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E868X, E932Y, and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;K266X, E868X, and N1026X, where X is K or R;D418X, E868S, and N1026X, where X is K or R;E932Y, K266X, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, K333X, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; orK266X, Q422X, E868X, and N1026X, where X is K or R.
[0206] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions at positions N1026X / K266X / Q422X / D418X (e.g., relative to SEQ ID NO: 1):
[0207] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266R, where X is R or K;N1026X and Q422X, where X is R or K;N1026X and D418X, where X is R or K;K266R and Q422X; where X is R or K;K266R and D418X, where X is R or K; orQ422X and D418X, where X is R or K.
[0208] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X, K266R, and Q422X, where X is R or K;N1026X, K266R, and D418X, where X is R or K;N1026X, Q422X, and D418X, where X is R or K;K266R, Q422X, and D418X, where X is R or K.
[0209] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions at positions N1026X, K266R, Q422X, D418X, where X is R or K (e.g., relative to SEQ ID NO: 1).
[0210] In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at positions 1026, 266, 422, or 418 of SEQ ID NO: 1 and further comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more) additional amino acid substitutions. In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 of SEQ ID NO: 1.
[0211] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions at positions E932Y, where Y is any amino acid except D, optionally where Y is N, Q, M, H, A, S, T, or K, in combination with one, two, or three of K266X, Q422X, and K1044X, where each X is independently R or K (e.g., relative to SEQ ID NO: 1).
[0212] In some embodiments, the NmeCas9 polypeptide comprises a combination of amino acid substitutions at positions E932Y where Y is any amino acid except D, optionally where Y is N, Q, M, H, A, S, T, or K, in combination with one, two, or three of K266X, Q422X, and K1044X, where each X is independently R or K (e.g., relative to SEQ ID NO: 1) and further comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more) additional amino acid substitutions. In some embodiments, the NmeCas9 polypeptide has no more than 5 substitutions relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide has no more than 4 substitutions relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide has no more than 3 substitutions relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide has no more than 2 substitutions relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide has only a single substitution relative to SEQ ID NO: 1.
[0213] In some embodiments, the one or more additional substitutions is at position E932 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution E932N, E932Q, E932H, E932T, E932S, E932M, E932A, E932R, or E932K relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a substitution selected from E932N, E932Q, E932M, E932R, E932H, E932A, E932S, E932T, and E932K relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises a substitution selected from E932N, E932M, E932R, E932T, and E932K relative to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide does not comprise the substitution E932D or E932R relative to SEQ ID NO: 1.
[0214] In some embodiments, the one or more additional substitutions is at position E868 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide further comprises a substitution at position E868 of SEQ ID NO: 1. In some embodiments, the amino acid sequence of the NmeCas9 polypeptide comprises the substitution E868R or E868K relative to SEQ ID NO: 1.
[0215] In some embodiments, the NmeCas9 polypeptide does not comprise the substitution E932K relative to SEQ ID NO: 1.
[0216] In some embodiments, the NmeCas9 polypeptide does not comprise a substitution at amino acid K1044 relative to SEQ ID NO: 1.
[0217] It is to be understood that, if the sequence includes a HiBiT tag (e.g., VSGWRLFKKIS; SEQ ID NO: 571), the tag is not included in the amino acid sequence within the scope of the claimed invention.
[0218] RNA-guided DNA binding agents described herein encompass Neisseria meningitidis Cas9 (NmeCas9) and modified and variants thereof. In some embodiments, the NmeCas9 is Nme2 Cas9. In some embodiments, the NmeCas9 is Nmel Cas9. In some embodiments, the NmeCas9 is Nme3 Cas9.
[0219] Modified versions having one catalytic domain, either RuvC or HNH, that is inactive are termed “nickases.” Nickases cut only one strand on the target DNA, thus creating a single-strand break. A single-strand break may also be known as a “nick.” In some embodiments, the compositions and methods comprise nickases. In some embodiments, the compositions and methods comprise a nickase RNA-guided DNA binding agent, such as a nickase Cas, e.g., a nickase Cas9, that induces a nick rather than a double strand break in the target DNA.
[0220] Accordingly, in another aspect, provided are NmeCas9 nickases that comprise one or more amino acid substitutions that enhance the genome editing activity of the NmeCas9 nickase.
[0221] In some embodiments, the NmeCas9 nuclease (e.g., an enhanced NmeCas9 nickase variant described herein) may be modified to contain only one functional nuclease domain. For example, the RNA-guided DNA binding agent may be modified such that one of the nuclease domains is mutated or fully or partially deleted to reduce its nucleic acid cleavage activity.
[0222] In some embodiments, a NmeCas9 nickase (e.g., comprising one or more enhancing amino acid substitutions described herein) is used having a RuvC domain with reduced activity. In some embodiments, aNmeCas9 nickase is used having an inactive RuvC domain. In some embodiments, a NmeCas9 nickase is used having an HNH domain with reduced activity. In some embodiments, aNmeCas9 nickase is used having an inactive HNH domain.
[0223] In some embodiments, a conserved amino acid within a NmeCas9 nuclease domain is substituted to reduce or alter nuclease activity. Wild type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 nuclease comprises more than one RuvC domain or more than one HNH domain. In someembodiments, the Cas9 nuclease is a wild type Cas9. In some embodiments, the Cas9 is capable of inducing a double strand break in target DNA, i.e., is a cleavase. In certain embodiments, the Cas nuclease may cleave dsDNA, it may cleave one strand of dsDNA, or it may not have DNA cleavase or nickase activity. In some embodiments, a NmeCas9 may comprise an amino acid substitution in the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include D16A (based on the A. meningitidis Cas9 protein). In some embodiments, the Cas protein may comprise an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include H588A (based on the NmeCas9 protein). These substitutions can be used in combination with one or more of the enhancing amino acid substitutions described herein to generate an enhanced NmeCas9 nickase. Specific examples of NmeCas9 nickases that include a D16A or H588A amino acid substitution in combination with an amino acid substitution that enhances that genome editing activity of NmeCas9 are further described herein.
[0224] In some embodiments, chimeric Cas proteins are used, where one domain or region of the protein is replaced by a portion of a different protein. In some embodiments, a NmeCas9 nuclease domain may be replaced with a domain from a different nuclease such as Fokl. In some embodiments, a NmeCas9 protein may be a modified NmeCas9 nuclease.
[0225] In some embodiments, the nuclease may be modified to induce a point mutation or base change, e.g., a deamination.
[0226] In some embodiments, the Cas protein comprises a fusion protein comprising a Cas nuclease (e.g., NmeCas9), which is a nickase or is catalytically inactive, linked to a heterologous functional domain. In some embodiments, the Cas protein comprises a fusion protein comprising a catalytically inactive Cas nuclease (e.g., NmeCas9) linked to a heterologous functional domain (see, e.g., WO2014152432). In some embodiments, the catalytically inactive Cas9 is from the A meningitidis Cas9. In some embodiments, the catalytically inactive Cas comprises mutations that inactivate the Cas.
[0227] In some embodiments, the heterologous functional domain is a domain that modifies gene expression, histones, or DNA. In some embodiments, the heterologous functional domain is a transcriptional activation domain or a transcriptional repressor domain. In some embodiments, the nuclease is a catalytically inactive Cas nuclease, such as dCas9.
[0228] In some embodiments, the heterologous functional domain is a deaminase, such as a cytidine deaminase or an adenine deaminase. In certain embodiments, theheterologous functional domain is a C to T base converter (cytidine deaminase), such as an apolipoprotein B mRNA editing enzyme (APOBEC) deaminase. A heterologous functional domain such as a deaminase may be part of a fusion protein with a Cas nuclease having nickase activity or a Cas nuclease that is catalytically inactive discussed further below.
[0229] In some embodiments, the NmeCas9 has double stranded endonuclease activity.
[0230] In some embodiments, the NmeCas9 has nickase activity.
[0231] In some embodiments, the NmeCas9 comprises a dCas9 DNA binding domain.
[0232] In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%.II.B. Nuclear localization signals (NLS)
[0233] In some embodiments, the RNA-guided DNA binding agent (e.g., NmeCas9 polypeptide) comprises a heterologous functional domain that facilitates transport of the RNA-guided DNA binding agent into the nucleus of a cell. In some embodiments, the heterologous functional domain is a nuclear localization signal (NLS).
[0234] In some embodiments, a NmeCas9 polypeptide disclosed herein is fused with one to five NLS(s). In some embodiments, the NmeCas9 polypeptide is fused with 1, 2, 3, or 4 NLS(s). In some embodiments, the NmeCas9 polypeptide may is with one NLS. In some embodiments, the NmeCas9 polypeptide is fused with two NLS(s). In some embodiments, the NmeCas9 polypeptide is fused with three NLSs. In some embodiments, the NmeCas9 polypeptide is fused with four NLSs. The fusion comprising the one or more NLSs may further comprises one or more linkers, such as a peptide linker that operably links the one or more NLSs to the NmeCas9 polypeptide.
[0235] Where one NLS is used, the NLS may be linked at the N-terminus or the C-terminus of the NmeCas9 polypeptide. In some embodiments, the NLS is not linked to the C-terminus. It may also be inserted within the NmeCas9 polypeptide. In certain circumstances, at least the two NLSs are the same (e.g., two SV40 NLSs). In certain embodiments, at least two different NLSs are present the NmeCas9 polypeptide. In some embodiments, the NmeCas9 polypeptide is fused to two SV40 NLS sequences linked at the C-terminus. In some embodiments, the NmeCas9 polypeptide is fused with two NLSs, one linked at the N-terminus and one at the C-terminus. In some embodiments, the NmeCas9 polypeptide isfused with three NLSs. In some embodiments, the NmeCas9 polypeptide is fused with no NLS.
[0236] The nuclear localization signal (NLS) disclosed herein may facilitate transport of the RNA-guided DNA-binding agent into the nucleus of a cell. The first NLS and, when present, the second NLS disclosed herein may be linked at the N-terminus to the RNA-guided DNA-binding agent sequence, i.e., the RNA-guided DNA binding agent is the C-terminal domain in the encoded polypeptide. The first NLS and, when present, the second NLS disclosed herein may be linked at the N-terminus to the NmeCas9 coding sequence.Additional NLS may be linked at the N-terminus of the NmeCas9 coding sequence. In some embodiments, the encoded polypeptide comprises three NLSs at the N-terminus to the NmeCas9 coding sequence. In some embodiments, at least one NLS is provided at the C-terminus of the RNA-guided DNA-binding agent sequence (e.g., with or without an intervening spacer or linker between the NLS and the preceding domain). In some embodiments, a first NLS and a second NLS are provided at the C-terminus of the RNA-guided DNA-binding agent sequence (e.g., with or without an intervening spacer or linker between the NLS and the preceding domain).
[0237] In some embodiments, the NLS may be a monopartite sequence, such as, e.g., the SV40 NLS, PKKKRKV (SEQ ID NO: 441) or PKKKRRV (SEQ ID NO: 473). In some embodiments, the NLS may be a bipartite sequence, such as the NLS of nucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 474). In some embodiments, the NLS sequence may comprise LAAKRSRTT (SEQ ID NO: 462), QAAKRSRTT (SEQ ID NO: 463), PAPAKRERTT (SEQ ID NO: 464), QAAKRPRTT (SEQ ID NO: 465), RAAKRPRTT (SEQ ID NO: 466), AAAKRSWSMAA (SEQ ID NO: 467), AAAKRVWSMAF (SEQ ID NO: 468), AAAKRSWSMAF (SEQ ID NO: 469), AAAKRKYFAA (SEQ ID NO: 470), RAAKRKAFAA (SEQ ID NO: 471), or RAAKRKYFAV (SEQ ID NO: 472). The NLS may be a snurportin-1 importin- (IBB domain, e.g., an SPNl-imp sequence. See Huber et al., 2002, J. Cell Bio., 156, 467-479. In a specific embodiment, a single PKKKRKV (SEQ ID NO: 441). In some embodiments, the first and second NLS are independently selected from an SV40 NLS, a nucleoplasmin NLS, a bipartite NLS, a c-myc like NLS, and an NLS comprising the sequence KTRAD (SEQ ID NO: 593). In certain embodiments, the first and second NLSs may be the same (e.g., two SV40 NLSs). In certain embodiments, the first and second NLSs may be different.
[0238] In some embodiments, the first NLS is a SV40NLS and the second NLS is a nucleoplasmin NLS.
[0239] In some embodiments, the SV40 NLS comprises a sequence of PKKKRKVE (SEQ ID NO: 436) or KKKRKVE (SEQ ID NO: 437). In some embodiments, the nucleoplasmin NLS comprises a sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 474). In some embodiments, the bipartite NLS comprises a sequence of KRTADGSEFESPKKKRKVE (SEQ ID NO: 438). In some embodiments, the c-myc like NLS comprises a sequence of PAAKKKKLD (SEQ ID NO: 439).
[0240] In some embodiments, one or more NLS(s) according to any of the foregoing embodiments are present in the RNA-guided DNA-binding agent in combination with one or more additional heterologous functional domains, such as any of the heterologous functional domains described below.II. C. Other Heterologous functional domains
[0241] In some embodiments, the RNA-guided DNA binding agent (e.g., NmeCas9 polypeptide) comprises one or more additional heterologous functional domains (e.g., is or comprises a fusion polypeptide).
[0242] In some embodiments, the heterologous functional domain may be capable of modifying the intracellular half-life of the RNA-guided DNA binding agent. In some embodiments, the half-life of the RNA-guided DNA binding agent may be increased. In some embodiments, the half-life of the RNA-guided DNA-binding agent may be reduced. In some embodiments, the heterologous functional domain may be capable of increasing the stability of the RNA-guided DNA-binding agent. In some embodiments, the heterologous functional domain may be capable of reducing the stability of the RNA-guided DNA-binding agent. In some embodiments, the heterologous functional domain may act as a signal peptide for protein degradation. In some embodiments, the protein degradation may be mediated by proteolytic enzymes, such as, for example, proteasomes, lysosomal proteases, or calpain proteases. In some embodiments, the heterologous functional domain may comprise a PEST sequence. In some embodiments, the RNA-guided DNA-binding agent may be modified by addition of ubiquitin or a polyubiquitin chain. In some embodiments, the ubiquitin may be a ubiquitin-like protein (UBL). Non-limiting examples of ubiquitin-like proteins include small ubiquitin-like modifier (SUMO), ubiquitin cross-reactive protein (UCRP, also known as interferon-stimulated gene-15 (ISG15)), ubiquitin-related modifier-1 (URM1), neuronal-precursor-cell-expressed developmentally downregulated protein-8 (NEDD8, also called Rubl in 5. cerevisiae), human leukocyte antigen F-associated (FAT10), autophagy-8 (ATG8) and -12 (ATG12), Fau ubiquitin-like protein (FUB1), membrane-anchored UBL (MUB), ubiquitin fold-modifier- 1 (UFM1), and ubiquitin-like protein-5 (UBL5).
[0243] In some embodiments, the heterologous functional domain may be a marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and reporter gene sequences. In some embodiments, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, sfGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreenl ), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire,), cyan fluorescent proteins (e.g, ECFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (e.g, mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFPl, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato) or any other suitable fluorescent protein. In other embodiments, the marker domain may be a purification tag or an epitope tag. Non-limiting exemplary tags include glutathione-S -transferase (GST), chitin binding protein (CBP), maltose binding protein (MBP), thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, SI, T7, V5, VSV-G, 6xHis (SEQ ID NO: 594), 8xHis (SEQ ID NO: 595), biotin carboxyl carrier protein (BCCP), poly -His, calmodulin, and HiBiT. Non-limiting exemplary reporter genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, or fluorescent proteins.
[0244] In additional embodiments, the heterologous functional domain may target the RNA-guided DNA-binding agent to a specific organelle, cell type, tissue, or organ.
[0245] In further embodiments, the heterologous functional domain may be an effector domain. When the RNA-guided DNA-binding agent is directed to its target sequence, e.g., when a Cas nuclease is directed to a target sequence by a gRNA, the effector domain may modify or affect the target sequence. In some embodiments, the effector domain may be chosen from a nucleic acid binding domain, a nuclease domain (e.g., a non-Casnuclease domain), an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain. In some embodiments, the heterologous functional domain is a nuclease, such as a Fokl nuclease. See, e.g., US Pat. No. 9,023,649. In some embodiments, the heterologous functional domain is a transcriptional activator or repressor. See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression,” Cell 152: 1173-83 (2013); Perez-Pinera et al., “RNA-guided gene activation by CRISPR-Cas9-based transcription factors,” Nat. Methods 10:973-6 (2013); Mali et al., “CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering,” Nat. Biotechnol. 31:833-8 (2013); Gilbert et al., “CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes,” Cell 154:442-51 (2013). As such, the RNA-guided DNA-binding agent essentially becomes a transcription factor that can be directed to bind a desired target sequence using a guide RNA. In certain embodiments, the DNA modification domain is a methylation domain, such as a demethylation or methyltransferase domain. In certain embodiments, the effector domain is a DNA modification domain, such as a base-editing domain. In particular embodiments, the DNA modification domain is a nucleic acid editing domain that introduces a specific modification into the DNA, such as a deaminase domain, which are further discussed below.II.D. Linkers
[0246] In some embodiments, the RNA-guided DNA binding agent (e.g., NmeCas9 polypeptide) further comprises a linker that connects the RNA-guided DNA binding agent to a heterologous functional domain. For example, in some embodiments, the RNA-guided DNA binding agent polypeptide comprising a deaminase and a RNA-guided nickase described herein further comprises a linker that connects the deaminase and the RNA-guided nickase. In some embodiments, the RNA-guided DNA binding agent polypeptide comprising a RNA-guided DNA binding agent and a nuclear localization signal (NLS) further comprises a linker that connects the NLS and the RNA-guided DNA binding agent. In some embodiments, the linker is a peptide linker.
[0247] In some embodiments, the peptide linker is any stretch of amino acids having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or more amino acids.
[0248] In some embodiments, the peptide linker is the 16 residue “XTEN” linker, or a variant thereof (See, e.g, the Examples; and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises the sequence SGSETPGTSESATPES (SEQ ID NO: 371), SGSETPGTSESA (SEQ ID NO: 372), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 373).
[0249] In some embodiments, the peptide linker comprises a (GGGGS)n (SEQ ID NO: 375), a (G)n, an (EAAAK)n(SEQ ID NO: 376), a (GGS)n, or an SGSETPGTSESATPES (SEQ ID NO: 371) motif (see, e g, Guilinger J P, Thompson D B, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; the entire contents are incorporated herein by reference), or an (XP)n motif, or a combination of any of these, wherein n is independently an integer between 1 and 30. See, W02015089406, e.g., paragraph
[0012] , the entire content of which is incorporated herein by reference.
[0250] Additional examples of peptide linkers are provided in SEQ ID NOs: 374-434 and further described in WO2023081689, which is hereby incorporated by reference.III. NmeCas9 Polynucleotides
[0251] Further provided herein are polynucleotides comprising open reading frames (ORFs) that encode variants of Neisseria meningitidis Cas9 (NmeCas9) that have enhanced genome editing activity relative to wildtype NmeCas9 (e.g., enhanced variants of NmeCas9 polypeptide comprising one or more enhancing amino acid substitutions described herein (e.g., see Table 3, Table 4, or Table 5)).
[0252] The polynucleotides described herein comprise an ORF encoding an NmeCas9 polypeptide that has amino acid substitutions relative to the amino acid sequence of a wildtype NmeCas9. In certain embodiments, the wildtype NmeCas9 polypeptide is Nme2Cas9 (SEQ ID NO: 1), NmelCas9 (SEQ ID NO: 2), or Nme3Cas9 (SEQ ID NO: 3).
[0253] In certain embodiments, the ORF encodes an NmeCas9 polypeptide comprising one or more amino acid substitutions that enhance an activity of the NmeCas9 polypeptide (e.g., DNA binding, cleaving, or nicking) relative to wild-type NmeCas9.
[0254] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide that comprises at least one amino acid substitution. In certain embodiments, the Nme2Cas9 polypeptide (e.g., an enhanced Nme2Cas9 variant) comprises an enhancing aminosubstitution described herein and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the ORF encodes a Nme2Cas9 polypeptide comprising one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, L846, E868, K870, K929, T930, E932, K962, K963, K965, Q967, K989, K1005, K1026, K1044, S1051, Q1053, orN1054 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some such embodiments, the ORF encodes a NmeCas9 polypeptide that comprises at least one amino acid substitution described in Table 3 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K / R, E508K / R, K517R, K549R, K555R, Q679R, L846K / Y, E868K / R, K870R, K929R, T930K, E932K / R, K962R, K963R, K965R, Q967R, K989R, K1005R, N1026K / R, K1044R, S1051K, Q1053K, or N1054K / R relative to SEQ ID NO: 1).
[0255] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprising one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the ORF encodes a Nme2Cas9 polypeptide comprising at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the ORF encodes a Nme2Cas9 polypeptide comprising one, two, three, four, five, six, seven, eight, nine, ten, or more than ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide).
[0256] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wild-type Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wild-type Nme2Cas9 polypeptide). In some embodiments, the aminoacid sequence of the Nme2Cas9 polypeptide does not contain more than four substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than three substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than two substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide contains only one substitution relative to SEQ ID NO: 1.
[0257] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5S6, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0258] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprises amino acid substitutions wherein the substitutions comprise at least one position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, and at least one position selected from the group consisting of S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, VI 056, N1064, LI 075, D844, D870, D873, D911, D5, S6, and E520 relative to SEQ ID NO: 1. Optionally, the two substitutions are at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0259] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to SEQ ID NO: 1 wherein the substitutions comprised at least one positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, and at least one position selected from the group consisting of S6 , G33, E47,R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5, S6, and E520 relative to SEQ ID NO: 1. . Optionally, the two substitutions are at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0260] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide that is a nickase comprising a RuvC substitution or an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the nickase substitution, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional substitutions relative to SEQ ID NO: 1, does not contain more than two substitutions relative to SEQ ID NO: 1, contains only one additional substitution relative to SEQ ID NO: 1.
[0261] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide that is a dCas9 comprising a RuvC substitution and an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the dCas9 substitutions, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional substitutions relative to SEQ ID NO: 1, does not contain more than two additional substitutions relative to SEQ ID NO: 1, contains only one additional substitution relative to SEQ ID NO: 1.
[0262] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprising an amino acid sequence having at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 31, 64, 86, 97, 306, or 317 (as shown in Table 23). In some embodiments, the ORF comprises aNme2Cas9 polynucleotide comprising a nucleic acid sequence having at least 91%, 92%, 93%, 94%, 95%, 96%, 97%,98%, or 99% identity to any one of SEQ ID NOs: 31, 64, 86, 97, 306, or 317. In some embodiments, the ORF comprises a Nme2Cas9 polynucleotide comprising the nucleic acid sequence of any one of SEQ ID NOs: 31, 64, 86, 97, 306, or 317.
[0263] In some embodiments, the ORF encodes an NmelCas9 polypeptide that comprises at least one amino acid substitution. In certain embodiments, the ORF encodes a NmelCas9 polypeptide (e.g., an enhanced NmelCas9 variant) that comprises an enhancing amino substitution described herein and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some embodiments, the ORF encodes a NmelCas9 polypeptide comprising one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, S197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, P1003, S1022, K1040, G1051, K1053, T1054 of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some such embodiments, the ORF encodes a NmelCas9 polypeptide that comprises at least one amino acid substitution described in Table 4 (i.e., K53R, K145R, S197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, P1003K / R, S1022R, K1040R, G1051K, K1053R, or T1054K / R relative to SEQ ID NO: 2). In certain embodiments, the NmelCas9 polypeptide comprises an amino acid substitution at position S1022 of SEQ ID NO: 2 (e.g., an S1022R substitution).
[0264] In some embodiments, the ORF encodes an Nme2Cas9 polypeptide comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wild-type Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the NmelCas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 4 relative to the amino acid sequence of SEQ ID NO: 2 (i.e., a wild-type Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four substitutions relative to SEQ ID NO: 2. Nme2Cas9 polypeptide does not contain more than three substitutions relative to SEQ ID NO: 2. Nme2Cas9 polypeptide does not contain more thantwo substitutions relative to SEQ ID NO: 2. Nme2Cas9 polypeptide contains only one substitution relative to SEQ ID NO: 2.
[0265] In some embodiments, the ORF encodes a Nme3Cas9 polypeptide that comprises at least one amino acid substitution. In certain embodiments, the ORF encodes a Nme3Cas9 polypeptide (e.g., an enhanced Nme3Cas9 variant) that comprises an enhancing amino substitution described herein and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the ORF encodes a Nme3Cas9 polypeptide comprising one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, or more than six) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, S1003, G1022, S1040, G1050, K1052, or T1053 of SEQ ID NO: 3 (i.e., a wildtype NmelCas9 polypeptide). In some such embodiments, the ORF encodes aNmelCas9 polypeptide comprising at least one amino acid substitution described in Table 5 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, S1003K / R, G1022R, S1040K / R, G1050K, K1052R, or T1053K / R substitution relative to SEQ ID NO: 3).
[0266] In some embodiments, the ORF encodes a Nme3Cas9 polypeptide that comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more) amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the ORF encodes aNme3Cas9 polypeptide comprising at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some embodiments, the ORF encodes aNmelCas9 polypeptide comprising one, two, three, four, five, six, seven, eight, nine, ten, or more than ten amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide).
[0267] In some embodiments, the ORF encodes an Nme3Cas9 polypeptide comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acidsubstitutions relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wild-type Nme3Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 5 relative to the amino acid sequence of SEQ ID NO: 3 (i.e., a wild-type Nme3Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme3Cas9 polypeptide does not contain more than four substitutions relative to SEQ ID NO: 3. Nme3Cas9 polypeptide does not contain more than three substitutions relative to SEQ ID NO: 3. Nme3Cas9 polypeptide does not contain more than two substitutions relative to SEQ ID NO: 3. Nme3Cas9 polypeptide contains only one substitution relative to SEQ ID NO: 3.
[0268] In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929.
[0269] In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions N1026, K266, Q422, and D418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises one substitution at a position selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. 1 In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises two substitutions at positions selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises three substitutions at positions selected from N1026, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises amino acid substitutions at positions N1026, K266, Q422, and D418 of SEQ ID NO: 1.
[0270] In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises one or more (e.g., one or more, two or more, three or more, or four or more) amino acid substitutions at positions N1026, E932, E868, K929, K266, Q422, and D418 of SEQ ID NO: 1 (i.e., a wild- type Nme2 Cas9 polypeptide). In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequencethat comprises one substitution at a position selected from N1026, E932, E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises two substitutions at positions selected from N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises three substitutions at positions selected from N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises amino acid substitutions at positions N1026, E932 , E868, K929, K266, Q422, and D418 of SEQ ID NO: 1. In some embodiments, the one or more amino acid substitutions comprises a substitution at positions N1026 of SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that comprises a substitution at positions E932 of SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that has four or fewer amino acid substitutions relative SEQ ID NO: 1. In some embodiments, ORF encodes an NmeCas9 polypeptide having an amino acid sequence that has four or fewer amino acid substitutions listed in Table 3 in SEQ ID NO: 1.
[0271] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises an amino acid substitution at position N1026 of SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises an N1026X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the ORF encodes aNmeCas9 polypeptide that comprises an N1026X substitution relative to SEQ ID NO: 1, where X is a basic amino acid. In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises a N1026K and N1026R substitution relative to SEQ ID NO: 1.
[0272] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises aN1026K substitution relative to SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid substitution N1026K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 307. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:307(Nme2Cas9(N1026K)).
[0273] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution N1026K and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 306. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 306 (Nme2Cas9(N1026K)), including the start or stop codon. In other embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO:306 (Nme2Cas9(N1026K)), excluding the start or stop codon.
[0274] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises aN1026R substitution relative to SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid substitution N1026R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 318. In other embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid sequence of SEQ ID NO:318 (Nme2Cas9(N1026R)).
[0275] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution N1026R and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 317. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 317 (Nme2Cas9(N1026R)), including the start or stop codon. In other embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO:317 (Nme2Cas9(N1026R)), excluding the start or stop codon.
[0276] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises an amino acid substitution at position K266 of SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a K266X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises a K266X substitution relative to SEQ ID NO: 1, where X is a basic amino acid.
[0277] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a K266R substitution relative to SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid substitution K266R andcomprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 32. In other embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid sequence of SEQ ID NO:32 (Nme2Cas9(K266R)).
[0278] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution K266R and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 31. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 31 (Nme2Cas9(K266R)), including the start or stop codon. In other embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO:31 (Nme2Cas9(K266R)), excluding the start or stop codon.
[0279] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises an amino acid substitution at position Q422 of SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises an Q422X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the ORF encodes aNmeCas9 polypeptide that comprises a Q422X substitution relative to SEQ ID NO: 1, where X is a basic amino acid.
[0280] In certain embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the substitution Q422K relative to SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid substitution Q422K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 87. In other embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid sequence of SEQ ID NO:87 (Nme2Cas9(Q422K)).
[0281] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution Q422K and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 86. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 86 (Nme2Cas9(Q422K)), including the start or stop codon. In otherembodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 86 (Nme2Cas9(Q422K)), excluding the start or stop codon.
[0282] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises the substitution Q422R relative to SEQ ID NO: 1. In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises the amino acid substitution D418R and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 98. In other embodiments, the NmeCas9 polypeptide comprises the amino acid sequence of SEQ ID NO:98 (Nme2Cas9(Q422R)).
[0283] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution Q422R and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 97. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 97 (Nme2Cas9(Q422R)), including the start or stop codon. In other embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO:97(Nme2Cas9(Q422R)), excluding the start or stop codon.
[0284] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises an amino acid substitution at position D418 of SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises an D418X substitution relative to SEQ ID NO: 1, where X is any amino acid. In some embodiments, the ORF encodes aNmeCas9 polypeptide that comprises an D418X substitution relative to SEQ ID NO: 1, where X is a basic amino acid. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises an amino acid substitution selected from D418K or D418R relative to SEQ ID NO: 1.
[0285] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the substitution D418K relative to SEQ ID NO: 1. In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises the amino acid substitution D418K and comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 65. In other embodiments,the ORF encodes a NmeCas9 polypeptide that comprises the amino acid sequence of SEQ ID NO: 65 (Nme2Cas9(D418K)).
[0286] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution D418K and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 64. In some embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 64 (Nme2Cas9(D418K)), including the start or stop codon. In other embodiments, the ORF comprises the nucleotide sequence of SEQ ID NO: 64 (Nme2Cas9(D418K)), excluding the start or stop codon.
[0287] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution at E932 and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 4. In some embodiments, the ORF encodes an NmeCas9 polypeptide having an amino acid substitution selected from E932N, E932Q, E932M, E932R, E932H, E932A, E932S, E932T, and E932K relative to SEQ ID NO: 1, excluding the start or stop codon.
[0288] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution E868R and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 4.
[0289] In some embodiments, the ORF encodes a NmeCas9 polypeptide having the amino acid substitution K929R and the ORF comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the nucleotide sequence of SEQ ID NO: 4.
[0290] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266X, where X is any amino acid;N1026X and Q422X, where X is any amino acid;N1026X and D418X, where X is any amino acid;K266X and Q422X, where X is any amino acid;K266X and D418X, where X is any amino acid;Q422X and D418X, where X is any amino acid;E932Y and N1026X, where Y is any amino acid except D, and where X is any amino acid;E932Y and E868X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K266X, where Y is any amino acid except D, and where X is any amino acid;E932Y and D418X, where Y is any amino acid except D, and where X is any amino acid;E932Y and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K333X, where Y is any amino acid except D, and where X is any amino acid;E932Y and K1044X, where Y is any amino acid except D, and where X is any amino acid; orE868X and N1026X, where Y is any amino acid except D, and where X is any amino acid.
[0291] In some embodiments, the ORF encodes aNmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X and K266X, where X is a basic amino acid;N1026X and Q422X, where X is a basic amino acid;N1026X and D418X, where X is a basic amino acid;K266X and Q422X, where X is a basic amino acid;K266X and D418X, where X is a basic amino acid;Q422X and D418X, where X is a basic amino acid;E932Y and N1026X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and E868X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K266X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and D418X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K333X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y and K1044X, where Y is any amino acid except D, and where X is a basic amino acid; orE868X and N1026X, where Y is any amino acid except D, and where X is a basic amino acid.
[0292] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X and K266X, where X is K or R;N1026X and Q422X, where X is K or R;N1026X and D418X, where X is K or R;K266X and Q422X, where X is K or R;K266X and D418X, where X is K or R;Q422X and D418X, where X is K or R;E932Y and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and E868X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and K266X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and D418X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E932Y and Q422X, where Y is N, Q, M, H, A, S, T, or K and where X is K or R; E932Y and K333X, where Y is N, Q, M, H, A, S, T, or K and where X is K or R; E932Y and K1044X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; or E868X and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R.
[0293] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X, K266X, and Q422X, where X is any amino acid;N1026X, K266X, and D418X, where X is any amino acid;N1026X, Q422X, and D418X, where X is any amino acid; orK266X, Q422X, and D418X, where X is any amino acidE932Y, K266X, and D418X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;K266X, Q422X, and E868X, where X is any amino acid;Q422X, E868X, and N1026X, where X is any amino acid;K266X, E868X, and N1026X, where X is any amino acid;E868X, E932Y, N1026X, where Y is any amino acid except D, and where X is any amino acid;E932Y, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E868X, E932Y, and N1026X, where Y is any amino acid except D, and where X is any amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;D418X, E868S, and N1026X, where X is any amino acid;E932Y, K266X, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, K333X, D418X, and Q422X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is any amino acid;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is any amino acid; orK266X, Q422X, E868X, and N1026X, where X is any amino acid.
[0294] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X, K266X, and Q422X, where X is a basic amino acid;N1026X, K266X, and D418X, where X is a basic amino acid;N1026X, Q422X, and D418X, where X is a basic amino acid;K266X, Q422X, and D418X, where X is a basic amino acid;E932Y, K266X, and D418X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, and Q422X, where Y is any amino acid except D, and where X is any basic amino acid;K266X, Q422X, and E868X, where X is a basic amino acid;Q422X, E868X, and N1026X, where X is a basic amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;E868X, E932Y, N1026X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E868X, E932Y, and N1026X, where Y is any amino acid except D, and where X is a basic amino acid;K266X, E868X, and N1026X, where X is a basic amino acid;D418X, E868S, and N1026X, where X is a basic amino acid;E932Y, K266X, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, K333X, D418X, and Q422X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is a basic amino acid;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is any amino acid except D, and where X is a basic amino acid; orK266X, Q422X, E868X, and N1026X, where X is a basic amino acid.
[0295] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X, K266X, and Q422X, where X is K or R;N1026X, K266X, and D418X, where X is K or R;N1026X, Q422X, and D418X, where X is K or R;K266X, Q422X, and D418X, where X is K or R;E932Y, K266X, and D418X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;K266X, Q422X, and E868X, where X is K or R;Q422X, E868X, and N1026X, where X is K or R;K266X, E868X, and N1026X, where X is K or R;E868X, E932Y, N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; E868X, E932Y, N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E868X, E932Y, and N1026X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;K266X, E868X, and N1026X, where X is K or R;D418X, E868S, and N1026X, where X is K or R;E932Y, K266X, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, K333X, D418X, and Q422X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, K333X, D418X, Q422X, E508X, K517X, and K549X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R;E932Y, K266X, D418X, Q422X, E508X, K517X, and K549X, where Y is N, Q, M, H, A, S, T, or K, and where X is K or R; orK266X, Q422X, E868X, and N1026X, where X is K or R.
[0296] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions at positions N1026X / K266X / Q422X / D418X (e.g., relative to SEQ ID NO: 1):
[0297] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ ID NO: 1):N1026X and K266R, where X is R or K;N1026X and Q422K, where X is R or K;N1026X and D418X, where X is R or K;K266R and Q422K;K266R and D418X, where X is R or K; orQ422K and D418X, where X is R or K.
[0298] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions selected from the following positions (e.g., relative to SEQ IDNO: 1):N1026X, K266R, and Q422K, where X is R or K;N1026X, K266R, and D418X, where X is R or K;N1026X, Q422K, and D418X, where X is R or K;K266R, Q422K, and D418X, where X is R or K;E932Y, K266X, and D418X, where Y is N, Q, H, T, S, A, or M, and where X is K or R;E932Y, K266X, and Q422X, where Y is N, Q, H, T, S, A, or M, and where X is K or R;K266X, Q422X, and E868X, where X is K or R;Q422X, E868X, and N1026X, where X is K or R;K266X, E868X, and N1026X, where X is K or R;E868X, E932Y, N1026X, where Y is N, Q, H, T, S, A, or M, and where X is K or R; orK266X, Q422X, E868X, and N1026X, where X is K or R.
[0299] In some embodiments, the ORF encodes a NmeCas9 polypeptide that comprises a combination of amino acid substitutions at positions N1026X, K266R, Q422K, D418X, where X is R or K (e.g., relative to SEQ ID NO: 1):
[0300] In some embodiments, the ORF encodes an NmeCas9 polypeptide that comprises an amino acid substitution at positions 1026, 266, 422, or 418 of SEQ ID NO: 1 and further comprises one or more (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more) additional amino acid substitutions. In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929. In some embodiments, the one or more additional substitutions is at position E932 of SEQ ID NO: 1. 1 In some embodiments, the ORF encodes an NmeCas9 polypeptide that further comprises the substitution E932N, E932Q, E932H, E932T, E932S, E932M, E932A, E932R, or E932K relative to SEQ ID NO: 1. In some embodiments, the ORF encodes an NmeCas9 polypeptide that further comprises a substitution selected from E932N, E932Q, E932M, E932R, E932H, E932A, E932S, E932T,and E932K relative to SEQ ID NO: 1, optionally E932N, E932M, E932R, E932T, and E932K.
[0301] In some embodiments, the one or more additional substitutions is at position E868 of SEQ ID NO: 1. In some embodiments, the ORF encodes an NmeCas9 polypeptide that further comprises a substitution at position E868 of SEQ ID NO: 1. In some embodiments, In some embodiments, the ORF encodes an NmeCas9 polypeptide that further comprises the substitution E868R or E868K relative to SEQ ID NO: 1.
[0302] In some embodiments, the ORF comprises codons that increase translation of the mRNA in a mammal. In some embodiments, the ORF comprises codons that increase translation of the mRNA in a human.
[0303] In some embodiments, the polynucleotide is DNA or RNA. In some embodiments, the polynucleotide is an mRNA.
[0304] It is to be understood that, if the sequence includes a HiBiT tag, the tag is not included in the amino acid sequence within the scope of the claimed invention.III. A. Exemplary Coding Sequences
[0305] In any of the embodiments set forth herein, the polynucleotide may be a mRNA comprising an ORF encoding an RNA-guided DNA binding agent (e.g., NmeCas9 polypeptide) disclosed above. In some embodiments, the polynucleotide is a mRNA comprising an ORF encoding an NmeCas9. In some embodiments, the polynucleotide is an expression construct comprising a promoter operably linked to an ORF encoding an RNA-guided DNA binding agent (e.g., NmeCas9).
[0306] Certain ORFs are translated in vivo more efficiently than others in terms of polypeptide molecules produced per mRNA molecule. The codon pair usage of such efficiently translated ORFs may contribute to translation efficiency. Further description of improvement of ORF coding sequence, codon pair usage, codon repeat contents are disclosed in WO 2019 / 0067910 and WO 2020 / 198641, the contents of each of which are hereby incorporated by reference in their entirety.
[0307] For example, in some embodiments, at least 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons of the ORF are minimal adenine codons or minimal uridine codons. In some embodiments, the ORF comprises or consists of codons that increase translation of the mRNA in a mammal. In some embodiments, the ORF comprises or consists of codons that increase translation of the mRNA in a human. An increase in translation in amammal, cell type, organ of a mammal, human, organ of a human, etc., can be determined relative to the extent of translation wild-type sequence of the ORF, or relative to an ORF having a codon distribution matching the codon distribution of the organism from which the ORF was derived or the organism that contains the most similar ORF at the amino acid level.
[0308] In some embodiments, the GC content of the ORF is greater than or equal to 56%. In some embodiments, the GC content of the ORF is greater than or equal to 56.5%. In some embodiments, the GC content of the ORF is greater than or equal to 57%. In some embodiments, the GC content of the ORF is greater than or equal to 57.5%. In some embodiments, the GC content of the ORF is greater than or equal to 58%. In some embodiments, the GC content of the ORF is greater than or equal to 58.5%. In some embodiments, the GC content of the ORF is greater than or equal to 59%. In some embodiments, the GC content of the ORF is less than or equal to 63%. In some embodiments, the GC content of the ORF is less than or equal to 62.6%. In some embodiments, the GC content of the ORF is less than or equal to 62.1%. In some embodiments, the GC content of the ORF is less than or equal to 61.6%. In some embodiments, the GC content of the ORF is less than or equal to 61.1%. In some embodiments, the GC content of the ORF is less than or equal to 60.6%. In some embodiments, the GC content of the ORF is less than or equal to 60.1%.
[0309] In some embodiments, the ORF consists of a set of codons of which at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 1.
[0310] Further examples of ORF encoding NmeCas9 polypeptides, including low adenine and / or low uridine contents, are further described in International Patent Publication No. WO2023081689A2, which is hereby incorporated by reference in its entirety.Table 1. Low A / U CodonsAmino Acid Low A / UGly GGCGlu GAGAsp GACVai GTGAla GCCArg CGGSer AGCLys AAGAsn AACMet ATGHe ATCThr ACCTrp TGGCys TGCTyr TACLeu CTGPhe TTCGin CAGHis CAC
[0311] In some embodiments, the polypeptide (e.g., NmeCas9 polypeptide) encoded by the ORF described herein comprises one or more additional heterologous functional domains (e.g., is or comprises a fusion polypeptide). In some embodiments, the ORF further comprises a nucleotide sequence encoding one or more additional heterologous functional domains, such as those further described herein.
[0312] In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for one or more nuclear localization signals (NLSs). In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for one nuclear localization signal (NLS). In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for two nuclear localization signals (NLSs). In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for the NLS located N-terminal to the NmeCas9 polypeptide. In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for the NLS located C-terminal to the NmeCas9 polypeptide. In some embodiments, the ORF encoding the polypeptide disclosed herein comprises a coding sequence for the first NLS and a coding sequence for the second NLS such that the encoded first NLS and second NLS are located to N-terminal to the NmeCas9 polypeptide. In some embodiments, the ORF further comprises a coding sequence for a third NLS C-terminal to the ORF encoding the NmeCas9.III.B. Codons that increase translation or that correspond to highly expressed tRNAs; exemplary codon sets
[0313] In some embodiments, the ORF has codons that increase translation in a mammal, such as a human. In further embodiments, the ORF has codons that increase translation in an organ, such as the liver, of the mammal, e.g., a human. In further embodiments, the ORF has codons that increase translation in a cell type, such as ahepatocyte, of the mammal, e.g., a human. An increase in translation in a mammal, cell type, organ of a mammal, human, organ of a human, etc., can be determined relative to the extent of translation wild-type sequence of the ORF, or relative to an ORF having a codon distribution matching the codon distribution of the organism from which the ORF was derived or the organism that contains the most similar ORF at the amino acid level.
[0314] In some embodiments, the polypeptide encoded by the ORF is a Cas9 nuclease derived from prokaryotes described below, and an increase in translation in a mammal, cell type, organ of a mammal, human, organ of a human, etc., can be determined relative to the extent of translation wild-type sequence of the ORF, or relative to an ORF of interest, such as an ORF encoding a human protein or transgene for expression in a human cell. For example, the ORF may be an ORF having a codon distribution matching the codon distribution of the organism from which the ORF was derived or the organism that contains the most similar ORF at the amino acid level, such as N. meningitidis with all else equal, including any applicable point mutations, heterologous domains, and the like. Codons useful for increasing expression in a human, including the human liver and human hepatocytes, can be codons corresponding to highly expressed tRNAs in the human liver / hepatocytes, which are discussed in Dittmar KA, PLOS Genetics 2(12): e221 (2006). In some embodiments, at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons corresponding to highly expressed tRNAs (e.g., the highest-expressed tRNA for each amino acid) in a mammal, such as a human. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons corresponding to highly expressed tRNAs (e.g., the highest- expressed tRNA for each amino acid) in a mammalian organ, such as a human organ. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons corresponding to highly expressed tRNAs (e.g., the highest- expressed tRNA for each amino acid) in a mammalian liver, such as a human liver. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons corresponding to highly expressed tRNAs (e.g., the highest- expressed tRNA for each amino acid) in a mammalian hepatocyte, such as a human hepatocyte.
[0315] Alternatively, codons corresponding to highly expressed tRNAs in an organism (e.g., human) in general may be used.
[0316] Any of the foregoing approaches to codon selection can be combined with selecting codon that contribute to lower repeat content; or using a codon set of Table 1, asshown above; using the minimal uridine or adenine codons as described in described in International Patent Publication No. WO2023081689A2, and then where more than one option is available, using the codon that corresponds to a more highly-expressed tRNA, either in the organism (e.g., human) in general, or in an organ or cell type of interest, such as the liver or hepatocytes (e.g., human liver or human hepatocytes).III. C. UTRs; Kozak sequences
[0317] In some embodiments, the polynucleotide comprises at least one UTR from Hydroxy steroid 17-Beta Dehydrogenase 4 (HSD17B4 or HSD), e.g., a 5’ UTR from HSD. In some embodiments, the polynucleotide comprises at least one UTR from a globin mRNA, for example, human alpha globin (HBA) mRNA, human beta globin (HBB) mRNA, or Xenopus laevis beta globin (XBG) mRNA. In some embodiments, the polynucleotide comprises a 5’ UTR, 3’ UTR, or 5’ and 3’ UTRs from a globin mRNA, such as HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises a 5’ UTR from bovine growth hormone, cytomegalovirus (CMV), mouse Hba-al, HSD, an albumin gene, HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises a 3’ UTR from bovine growth hormone, cytomegalovirus, mouse Hba-al, HSD, an albumin gene, HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises 5’ and 3’ UTRs from bovine growth hormone, cytomegalovirus, mouse Hba-al, HSD, an albumin gene, HBA, HBB, XBG, heat shock protein 90 (Hsp90), glyceraldehyde 3 -phosphate dehydrogenase (GAPDH), beta-actin, alphatubulin, tumor protein (p53), or epidermal growth factor receptor (EGFR).
[0318] In some embodiments, the polynucleotide comprises 5’ and 3’ UTRs that are from the same source, e.g., a constitutively expressed mRNA such as actin, albumin, or a globin such as HBA, HBB, or XBG.
[0319] In some embodiments, the polynucleotide disclosed herein comprises a 5’ UTR with at least 90% identity to any one of SEQ ID NOs: 444-451. In some embodiments, the polynucleotide disclosed herein comprises a 3’ UTR with at least 90% identity to any one of SEQ ID NOs: 452-459. In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, an mRNA disclosed herein comprises a 5’ UTR having the sequence of any one of SEQ ID NOs: 444-451. In some embodiments, the polynucleotide disclosed herein comprises a 3’ UTR having the sequence of any one of SEQ ID NOs: 452-459.
[0320] In some embodiments, the polynucleotide does not comprise a 5’ UTR, e.g., there are no additional nucleotides between the 5’ cap and the start codon. In some embodiments, the mRNA comprises a Kozak sequence (described below) between the 5’ cap and the start codon, but does not have any additional 5’ UTR. In some embodiments, the mRNA does not comprise a 3’ UTR, e.g., there are no additional nucleotides between the stop codon and the poly- A tail.
[0321] In some embodiments, the mRNA comprises a Kozak sequence. The Kozak sequence can affect translation initiation and the overall yield of a polypeptide translated from an mRNA. A Kozak sequence includes a methionine codon that can function as the start codon. A minimal Kozak sequence is NNNRUGN wherein at least one of the following is true: the first N is A or G and the second N is G. In the context of a nucleotide sequence, R means a purine (A or G). In some embodiments, the Kozak sequence is RNNRUGN, NNNRUGG, RNNRUGG, RNNAUGN, NNNAUGG, or RNNAUGG. In some embodiments, the Kozak sequence is RCCRUGG with zero mismatches or with up to one or two mismatches to positions in lowercase. In some embodiments, the Kozak sequence is RCCAUGG with zero mismatches or with up to one or two mismatches to positions in lowercase. In some embodiments, the Kozak sequence is GCCRCCAUGG (SEQ ID NO: 460) (nucleotides 4-13 of SEQ ID NO: 461; SEQ ID NO: 462) with zero mismatches or with up to one, two, or three mismatches to positions in lowercase. In some embodiments, the Kozak sequence is GCCACCAUG with zero mismatches or with up to one, two, three, or four mismatches to positions in lowercase. In some embodiments, the Kozak sequence is GCCACCAUG. In some embodiments, the Kozak sequence is GCCGCCRCCAUGG (SEQ ID NO: 461) with zero mismatches or with up to one, two, three, or four mismatches to positions in lowercase.III.D. 5 ’ cap
[0322] In some embodiments, the polynucleotide (e.g., mRNA) disclosed herein comprises a 5’ cap, such as a CapO, Capl, or Cap2.
[0323] A 5’ cap is generally a 7-methylguanine ribonucleotide (which may be further modified, as discussed below e.g., with respect to ARCA) linked through a 5 ’-triphosphate to the 5’ position of the first nucleotide of the 5’-to-3’ chain of the nucleic acid, i.e., the first cap-proximal nucleotide. In CapO, the riboses of the first and second cap-proximal nucleotides of the mRNA both comprise a 2’-hydroxyl. In Capl, the riboses of the first andsecond transcribed nucleotides of the mRNA comprise a 2’ -methoxy and a 2’ -hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the mRNA both comprise a 2’-methoxy. See, e.g., Katibah et al. (2014) Proc Natl Acad Sci USA 111(33): 12025-30; Abbas et al. (2017) Proc Natl Aca d Sci USA 114(ll):E2106-E2115. Most endogenous higher eukaryotic nucleic acids, including mammalian nucleic acids such as human nucleic acids, comprise Capl or Cap2. CapO and other cap structures differing from Capl and Cap2 may be immunogenic in mammals, such as humans, due to recognition as “non-self’ by components of the innate immune system such as IFIT-1 and IFIT-5, which can result in elevated cytokine levels including type I interferon. Components of the innate immune system such as IFIT-1 and IFIT-5 may also compete with eIF4E for binding of a nucleic acids with a cap other than Capl or Cap2, potentially inhibiting translation of the nucleic acid.
[0324] A cap can be included co-transcriptionally. For example, ARCA (anti-reverse cap analog; Thermo Fisher Scientific Cat. No. AM8045) is a cap analog comprising a 7-methylguanine 3 ’-m ethoxy-5’ -triphosphate linked to the 5’ position of a guanine ribonucleotide which can be incorporated in vitro into a transcript at initiation. ARCA results in a CapO cap or a CapO-like cap in which the 2’ position of the first cap-proximal nucleotide is hydroxyl. See, e.g., Stepinski et al., (2001) “Synthesis and properties of mRNAs containing the novel ‘anti -reverse’ cap analogs 7-methyl(3'-O-methyl)GpppG and 7-methyl(3'deoxy)GpppG,” RNA 7: 1486-1495. The ARCA structure is shown below.
[0325] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnologies Cat. No. N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG; TriLink Biotechnologies Cat. No. N-7133) can be used to provide a Capl structure co-transcriptionally. 3’-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also available from TriLink Biotechnologies as Cat. Nos. N-7413 andN-7433, respectively. The CleanCap™ AG structure is shown below. CleanCap™ structures are sometimes referred to herein using the last three digits of the catalog numbers listed above (e.g., “CleanCap™ 113” for TriLink Biotechnologies Cat. No. N-7113). / "“Ibn
[0326] Alternatively, a cap can be added to an RNA post-transcriptionally. For example, Vaccinia capping enzyme is commercially available (New England Biolabs Cat. No. M2080S) and has RNA triphosphatase and guanylyltransferase activities, provided by its DI subunit, and guanine methyltransferase, provided by its D 12 subunit. As such, it can add a 7-methylguanine to an RNA, so as to give CapO, in the presence of S-adenosyl methionine and GTP. See, e.g., Guo, P. and Moss, B. (1990) Proc. Natl. Acad. Sci. USA 87, 4023-4027; Mao, X. and Shuman, S. (1994) J. Biol. Chem. 269, 24472-24479. For additional discussion of caps and capping approaches, see, e.g., WO2017 / 053297 and Ishikawa et al., Nucl. Acids. Symp. Ser. (2009) No. 53, 129-130.III.E. Poly -A tail
[0327] In some embodiments, the polynucleotide is a mRNA that encodes a polypeptide disclosed herein comprising an ORF, and the mRNA further comprises a polyadenylated (poly- A) tail.
[0328] In some embodiments, the polynucleotide disclosed herein further comprises a poly-A tail sequence or a polyadenylation signal sequence. In some embodiments, the poly-Atail sequence comprises 100-400 (SEQ ID NO: 939) nucleotides. In some embodiments, the poly-A tail comprises at least 20, 30, 40, 50, 60, 70, 80, 90, or 100 adenines (SEQ ID NO: 940), optionally up to 300 adenines. In some embodiments, the poly-A tail comprises 95, 96, 97, 98, 99, or 100 adenine nucleotides (SEQ ID NO: 941). In some embodiments, the poly-A tail includes non-adenine nucleotides, z.e., is an interrupted poly-A tail. In certain embodiments, the poly-A tail is interrupted by a non-adenine nucleotide about every 40, 50, 60, 70, 80, or 90 nucleotides. In certain embodiments, the poly-A tail is interrupted by a non-adenine nucleotide about every 50 nucleotides.
[0329] In some embodiments, the poly-A sequence comprises non-adenine nucleotides. In some instances, the poly-A tail is “interrupted” with one or more non-adenine nucleotide “anchors” at one or more locations within the poly-A tail. The poly-A tails may comprise at least 8 consecutive adenine nucleotides, but also comprise one or more non-adenine nucleotide. As used herein, “non-adenine nucleotides” refer to any natural or nonnatural nucleotides that do not comprise adenine. Guanine, thymine, and cytosine nucleotides are exemplary non-adenine nucleotides. Thus, the poly-A tails on the mRNA described herein may comprise consecutive adenine nucleotides located 3’ to nucleotides encoding a polypeptide disclosed herein. In some instances, the poly-A tails on mRNA comprise non- consecutive adenine nucleotides located 3’ to nucleotides encoding an RNA-guided DNA- binding agent or a sequence of interest, wherein non-adenine nucleotides interrupt the adenine nucleotides at regular or irregularly spaced intervals.
[0330] In some embodiments, the poly-A tail is encoded in the plasmid used for in vitro transcription of mRNA and becomes part of the transcript. The poly-A sequence encoded in the plasmid, i.e., the number of consecutive adenine nucleotides in the poly-A sequence, may not be exact, e.g., a 100 poly-A sequence in the plasmid may not result in a precisely 100 poly-A sequence in the transcribed mRNA. In some embodiments, the poly-A tail is not encoded in the plasmid, and is added by PCR tailing or enzymatic tailing, e.g., using E. coli poly(A) polymerase.
[0331] In some embodiments, the one or more non-adenine nucleotides are positioned to interrupt the consecutive adenine nucleotides so that a poly(A) binding protein can bind to a stretch of consecutive adenine nucleotides. In some embodiments, one or more non-adenine nucleotide(s) is located after at least 8, 9, 10, 11, or 12 consecutive adenine nucleotides (SEQ ID NO: 596). In some embodiments, the one or more non-adenine nucleotide is located after at least 8-50 consecutive adenine nucleotides (SEQ ID NO: 597). In some embodiments, the one or more non-adenine nucleotide is located after at least 8-100 consecutive adenine nucleotides (SEQ ID NO: 598). In some embodiments, the non-adenine nucleotide is after one, two, three, four, five, six, or seven adenine nucleotides and is followed by at least 8 consecutive adenine nucleotides.
[0332] The poly-A tail of the present disclosure may comprise one sequence of consecutive adenine nucleotides followed by one or more non-adenine nucleotides, optionally followed by additional adenine nucleotides.
[0333] In some embodiments, the poly-A tail comprises or contains one non-adenine nucleotide or one consecutive stretch of 2-10 non-adenine nucleotides. In some embodiments, the non-adenine nucleotide(s) is located after at least 8, 9, 10, 11, or 12 consecutive adenine nucleotides (SEQ ID NO: 596). In some instances, the one or more non-adenine nucleotides are located after at least 8-50 consecutive adenine nucleotides (SEQ ID NO: 597). In some embodiments, the one or more non- adenine nucleotides are located after at least 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 consecutive adenine nucleotides (SEQ ID NO: 597).In some embodiments, the non-adenine nucleotide is guanine, cytosine, or thymine. In some instances, the non-adenine nucleotide is a guanine nucleotide. In some embodiments, the non-adenine nucleotide is a cytosine nucleotide. In some embodiments, the non-adenine nucleotide is a thymine nucleotide. In some instances, where more than one non-adenine nucleotide is present, the non-adenine nucleotide may be selected from: a) guanine and thymine nucleotides; b) guanine and cytosine nucleotides; c) thymine and cytosine nucleotides; or d) guanine, thymine and cytosine nucleotides.III.F. DNA Molecules, Expression Constructs, Host Cells, and Production Methods
[0334] In certain embodiments, the present disclosure provides a DNA molecule comprising an open reading frame (ORF) encoding an NmeCas9 polypeptide disclosed herein. In some embodiments, in addition to the ORF sequence, the DNA molecule further comprises nucleic acids that do not encode the polypeptide. Nucleic acids that do not encode the polypeptide include, but are not limited to, promoters, enhancers, regulatory sequences, and nucleic acids encoding a guide RNA.
[0335] In some embodiments, the DNA molecule further comprises a nucleotide sequence encoding a crRNA, a trRNA, or a crRNA and trRNA. In some embodiments, the nucleotide sequence encoding the crRNA, trRNA, or crRNA and trRNA comprises or consists of a guide sequence flanked by all or a portion of a repeat sequence from a naturally occurring CRISPR / Cas system. The nucleic acid comprising or consisting of the crRNA, trRNA, or crRNA and trRNA may further comprise a vector sequence wherein the vector sequence comprises or consists of nucleic acids that are not naturally found together with the crRNA, trRNA, or crRNA and trRNA. In some embodiments, the crRNA and the trRNA are encoded by non-contiguous nucleic acids within one vector. In other embodiments, thecrRNA and the trRNA may be encoded by a contiguous nucleic acid. In some embodiments, the crRNA and the trRNA are encoded by opposite strands of a single nucleic acid. In other embodiments, the crRNA and the trRNA are encoded by the same strand of a single nucleic acid.
[0336] In some embodiments, the DNA molecule further comprises a promoter operably linked to the sequence encoding any of the ORF encoding a polypeptide disclosed herein. In some embodiments, the DNA molecule is an expression construct suitable for expression in a mammalian cell, e.g., a human cell or a mouse cell, such as a human hepatocyte or a rodent (e.g., mouse) hepatocyte. In some embodiments, the DNA molecule is an expression construct suitable for expression in a cell of a mammalian organ, e.g., a human liver or a rodent (e.g., mouse) liver. In some embodiments, the DNA molecule is a plasmid or an episome. In some embodiments, the DNA molecule is contained in a host cell, such as a bacterium or a cultured eukaryotic cell. Exemplary bacteria include proteobacteria such as E. coli. Exemplary cultured eukaryotic cells include primary hepatocytes, including hepatocytes of rodent (e.g., mouse) or human origin; hepatocyte cell lines, including hepatocytes of rodent (e.g., mouse) or human origin; human cell lines; rodent (e.g., mouse) cell lines; CHO cells; microbial fungi, such as fission or budding yeasts, e.g., Saccharomyces, such as 5. cerevisiae,' and insect cells.
[0337] In some embodiments, a method of producing an mRNA disclosed herein is provided. In some embodiments, such a method comprises contacting a DNA molecule described herein with an RNA polymerase under conditions permissive for transcription. In some embodiments, the contacting is performed in vitro, e.g., in a cell-free system. In some embodiments, the RNA polymerase is an RNA polymerase of bacteriophage origin, such as T7 RNA polymerase. In some embodiments, NTPs are provided that include at least one modified nucleotide as discussed above. In some embodiments, the NTPs include at least one modified nucleotide as discussed above and do not comprise UTP.
[0338] In some embodiments, a method of producing a polynucleotide disclosed herein is provided. In some embodiments, such a method comprises contacting an expression construct disclosed herein with an RNA polymerase and NTPs that comprise at least one modified nucleotide. In some embodiments, the modified nucleotide comprises a modified uridine. In further embodiments, at least 80% of the uridine positions are modified uridines. In further embodiments, at least 90% of the uridine positions are modified uridines. In further embodiments, 100% of the uridine positions are modified uridines. In further embodiments,the modified uridine comprises or is a substituted uridine, pseudouridine, or a substituted pseudouridine. In further embodiments, the modified uridine comprises or is Nl-methyl-psuedouridine. In some embodiments, the expression construct comprises an encoded poly-A tail sequence.IV. Exemplary compositions comprising a NmeCas9 fusion protein
[0339] An enhanced NmeCas9 variant described herein (e.g., comprising one or more amino acid substitutions that enhance genome editing, e.g., see Table 3, Table 4, or Table 5) may further comprise one or more additional heterologous functional domains, thereby forming a NmeCas9 fusion protein. Non-limiting examples of heterologous functional domains that may be included in the NmeCas9 fusion protein include deaminases (e.g., a cytidine deaminase, such as APOB EC), a uracil glycosylase inhibitor (UGI)), an HMGB1 domain, or a polymerase, or a nuclear localization signal. In certain embodiments, when a heterologous functional domain is fused to the N-terminus of the NmeCas9, the initiator methionine is removed from the NmeCas9, e.g., from SEQ ID NO: 1, 2, or 3. Deletion of the initiator methionine from the N-terminus of SEQ ID NO: 1, 2, or 3 is not considered a substitution in a polypeptide as compared to the reference SEQ ID NO in the embodiments provided herein.
[0340] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or four or more) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, L846, E868, K870, K929, T930, E932, K962, K963, K965, Q967, K989, K1005, K1026, K1044, S1051, Q1053, orN1054 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an Nme2Cas9 polypeptide and comprises at least one amino acid substitution described in Table 3 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K / R, E508K / R, K517R, K549R, K555R, Q679R, L846K / Y, E868K / R, K870R, K929R, T930K, E932K / R, K962R, K963R, K965R, Q967R, K989R, K1005R, N1026K / R, K1044R, S1051K, Q1053K, or N1054K / R relative to SEQ ID NO: 1).
[0341] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acidsubstitutions relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1 (i.e., a wildtype Nme2Cas9 polypeptide). In some embodiments, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than three substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide does not contain more than two substitutions relative to SEQ ID NO: 1. Nme2Cas9 polypeptide contains only one substitution relative to SEQ ID NO: 1.
[0342] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the Nme2Cas9 polypeptide comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844, D870, D873, D911, D5S6, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO: 1.
[0343] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the Nme2Cas9 polypeptide is a nickase comprising a RuvC substitution or an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions listed in Table 2 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427,D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the nickase substitution, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional amino acid substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional amino acid substitutions relative to SEQ ID NO: 1, does not contain more than two additional amino acid substitutions relative to SEQ ID NO: 1, contains only one additional amino acid substitution relative to SEQ ID NO: 1.
[0344] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the Nme2Cas9 polypeptide is a dCas9 comprising a RuvC substitution and an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer,two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In certain embodiments, the dCas9 comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the dCas9 substitutions, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional amino acid substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional amino acid substitutions relative to SEQ ID NO: 1, does not contain more than two additional amino acid substitutions relative to SEQ ID NO: 1, contains only one additional amino acid substitution relative to SEQ ID NO: 1.
[0345] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions K53, K145, S197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, P1003, S1022, K1040, G1051, K1053, T1054 of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an NmelCas9 polypeptide and comprises at least one amino acid substitution described in Table 4 (i.e., K53R, K145R, S197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, P1003K / R, S1022R, K1040R, G1051K, K1053R, or T1054K / R relative to SEQ ID NO: 2). In certain embodiments, theNmelCas9 polypeptide comprises an amino acid substitution at position S1022 of SEQ ID NO: 2 (e.g., an S1022R substitution).
[0346] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, S1003, G1022, S1040, G1050, K1052, or T1053 of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an Nme3Cas9 polypeptide and comprises at least one amino acid substitution described in Table 5 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R,K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, S1003K / R, G1022R, S1040K / R, G1050K, K1052R, or T1053K / R substitution relative to SEQ ID NO: 3.
[0347] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0348] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 1026 and one or more (e.g., one or more, two or more, or all three) amino acid substitutions at positions 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0349] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0350] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0351] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 1026 (e.g., N1026) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9polypeptide comprises an N1026K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an N1026R substitution relative to SEQ ID NO: 1.
[0352] In some embodiments, the NmeCas9 polypeptide further comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 266 (e.g., K266) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a K266R substitution relative to SEQ ID NO: 1.
[0353] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 422 (e.g., Q422) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a Q422K substitution relative to SEQ ID NO: 1.
[0354] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 418 (e.g., D418) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a D418K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a D418R substitution relative to SEQ ID NO: 1.
[0355] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 932 (e.g., E932) of SEQ ID NO: 1 (i.e., a wild-type Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a E932 substitution to K, N, Q, M, R, H, A, S, or T relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a heterologous functionaldomain and the amino acid sequence of the NmeCas9 polypeptide comprises an E932 substitution to K, N, Q, M, R, H, A, S, or T relative to SEQ ID NO: 1.
[0356] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 868 (e.g., E868) of SEQ ID NO: 1 (i.e., a wild-type Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a E868R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an E868R substitution relative to SEQ ID NO: 1.
[0357] In some embodiments, the NmeCas9 polypeptide may further comprise a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 929 (e.g., K929) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a K929R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a heterologous functional domain and the amino acid sequence of the NmeCas9 polypeptide comprises a K929R substitution relative to SEQ ID NO: 1.
[0358] Further provided herein are polynucleotides comprising an open reading frame encoding the above NmeCas9 polypeptides that comprise a heterologous functional domain.
[0359] In some embodiments, methods of modifying a target gene are provided comprising contacting a cell with a NmeCas9 polypeptide that further comprises a heterologous functional domain, as further described herein.
[0360] IV.A. High Mobility Group Box l(HMGBl) Domain Fusions
[0361] Chromatin structure may impact the efficacy of various genome editing techniques, and target sequences located in chromatin-dense regions, in particular, may be less amenable to genome editing. For example, the CRISPR-associated nuclease Cas9 can exhibit insufficient editing of certain targets whose protospacer adjacent motifs (PAMs) are located within nucleosomes. Chromatin modulating peptides fused with Cas9 nucleases have been reported to improve editing efficiency at refractory target sites. See, e.g., Ding X,Seebeck T, Feng Y, Jiang Y, Davis GD, Chen F. Improving CRISPR-Cas9 Genome Editing Efficiency by Fusion with Chromatin-Modulating Peptides. CRISPR J. 2019 Feb;2:51-63.
[0362] The wildtype (high mobility group box 1) HMGB 1 protein is a non-sequencespecific DNA-binding protein that interacts with DNA through its DNA-binding domains. HMGB1 binding to DNA introduces local DNA distortions, which may affect the structure and stability of nearby DNA-protein complexes. In particular, HMGB1 binding within the linker region of nucleosomal DNA may destabilize the nucleosome structure, thereby increasing the accessibility of the surrounding chromatin. See, e.g., Starkova TY, Polyanichko AM, Artamonova TO, Tsimokha AS, Tomilin AN, Chikhirzhina EV. Structural Characteristics of High-Mobility Group Proteins HMGB1 and HMGB2 and Their Interaction with DNA. Int J Mol Sci. 2023 Feb 10;24(4):3577 and Lange SS, Vasquez KM. HMGB1: the jack-of-all-trades protein is a master DNA repair mechanic. Mol Carcinog. 2009 Jul;48(7):571-80.
[0363] The wildtype HMGB1 protein comprises, from N-terminus to C-terminus, a Box A DNA-binding domain, a Box B DNA-binding domain, a cryptic nuclear localization signal (NLS), and an acidic tail. The wildtype HMGB1 protein further comprises a receptor for advanced glycation end-products (RAGE) binding domain.
[0364] In the context of the wildtype protein, the acidic tail interacts with the Box A and Box B domains, causing the HMGB1 protein to adopt an auto-inhibited conformation that is not competent for DNA binding. The acidic tail therefore suppresses the chromatin remodeling activity of the wildtype HMGB1 protein. In contrast, a truncated HMGB1 protein lacking the acidic tail may remain in the active conformation and thereby exhibit increased chromatin remodeling activity.
[0365] While certain truncated HMGB 1 proteins lacking an acidic tail may have increased chromatin remodeling activity relative to the wildtype HMGB1 protein, other elements within the HMGB1 protein may also promote chromatin remodeling. For example, truncated HMGB1 proteins comprising a cryptic NLS, which may or may not have constitutive NLS activity, may exhibit increased chromatin remodeling activity compared to truncated HMGB1 proteins that lack the cryptic NLS.
[0366] The present disclosure provides for an HMGB1 polypeptide comprising certain domains of HMGB 1, or nucleic acids encoding the same. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box B domain. In some embodiments, theHMGB1 polypeptide comprises at least 7 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain.
[0367] In some embodiments, the HMGB1 polypeptide lacks all or part of an HMGB1 acidic tail domain. In some embodiments, the HMGB1 polypeptide lacks all of an acidic tail domain. In some embodiments, the HMGB 1 polypeptide lacks part of an acidic tail domain. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain.
[0368] In some embodiments, the HMGB1 polypeptide comprises a cryptic nuclear localization signal (NLS). In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 931) or a variant thereof. In some embodiments, the HMGB1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain. In some embodiments, the RAGE binding domain comprises the sequence of SEQ ID NO: 932. In some embodiments, the HMGB1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-20 amino acid residues immediately C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 933.
[0369] In some embodiments, the HMGB1 polypeptide lacks an acidic tail domain. In some embodiments, the HMGB1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 934. In some embodiments, the HMGB1 polypeptide lacks the sequence of SEQ ID NO: 934. In some embodiments, the HMGB1 polypeptide lacks transcriptional stimulatory function.
[0370] In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain, optionally wherein the cryptic NLS is comprised within the at least 7 contiguous amino acid residues. In some embodiments, the cryptic NLS immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain. In some embodiments, the cryptic NLS immediately precedes the N-terminal end of the HMGB1 Box B domain.
[0371] In some embodiments, the HMGB1 polypeptide lacks an HMGB1 Box A domain. In some embodiments, the HMGB1 polypeptide comprises two HMGB1 Box B domains. In some embodiments, the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domains. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box A domain. In some embodiments, the HMGB1 Box B domain is located C-terminal to the HMGB1 Box A domain.
[0372] In some embodiments, the HMGB1 polypeptide comprises, from N-terminus to C-terminus: (1) HMGB1 Box A domain-HMGBl Box B domain-cryptic NLS; (2) HMGB1 Box B domain-HMGBl Box B domain-cryptic NLS; or (3) HMGB1 Box B domain-cryptic NLS. In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOS: 929-938, or is encoded by a nucleic acid encoding the polypeptide. In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOS: 929-938. In some embodiments, the HMGB1 polypeptide comprises a sequence that is encoded by a nucleic acid encoding the polypeptide.
[0373] In some embodiments, the HMGB1 Box A domain comprises amino acid residues 2-78 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box A domain comprises amino acid residues 2-84 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box A domain comprises amino acid residues 2-88 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box B domain comprises amino acid residues 89-162 of SEQ ID NO: 928. In some embodiments, the HMGB1 Box B domain comprises amino acid residues 85-185 of SEQ ID NO: 928, including the cryptic nuclear localization signal. In some embodiments, the HMGB1 Box B domain consists of amino acid residues 85-185, including the cryptic nuclear localization signal.
[0374] In some embodiments, the present disclosure provides for an HMGB 1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7 amino acid residues C-terminal to the HMGB1 Box B domain, and wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain. In some embodiments, the acidic tail domain comprises amino acid residues 186-215 relative to SEQ ID NO: 928. In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 931) or a variant thereof. In some embodiments, the present disclosure provides for an HMGB1 polypeptide comprising an HMGB1 Box B domain and lacking an acidic tail domain, further wherein the HMGB1 polypeptide comprises a cryptic NLS comprising aminoacid residues 179-185 of SEQ ID NO: 928 or a variant thereof. In some embodiments, the present disclosure provides for an HMGB1 polypeptide comprising an HMGB1 polypeptide comprising an HMGB1 Box B domain and lacking an acidic tail domain, further wherein the HMGB1 polypeptide comprises amino acid residues 166-185 of SEQ ID NO: 928, which comprises the cryptic NLS (SEQ ID NO: 931), or a variant of amino acid residues 166-185 of SEQ ID NO: 928.
[0375] In some embodiments, the cryptic NLS is located C-terminal to the HMGB 1 Box B domain. In some embodiments, the cryptic NLS immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the cryptic NLS is located N-terminal to the HMGB1 Box B domain. In some embodiments, the cryptic NLS immediately precedes the C-terminal end of the HMGB1 Box B domain.
[0376] In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 931) is located C-terminal to the HMGB1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 931) immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 931) is located N-terminal to the HMGB 1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 931) immediately precedes the C-terminal end of the HMGB1 Box B domain.
[0377] In some embodiments, the HMGB1 polypeptide comprises the sequence of amino acid residues 166-185 relative to SEQ ID NO: 928, which comprises the cryptic NLS (SEQ ID NO: 931). In some embodiments, the HMGB1 polypeptide of amino acid residues 166-185 relative to SEQ ID NO: 928 has 2, 1, or 0 amino acid substitutions. In some embodiments, the amino acid substitutions are conservative amino acid substitutions.
[0378] In some embodiments, the HMGB1 polypeptide lacks an HMGB1 Box A domain. In some embodiments, the HMGB1 polypeptide comprises two HMGB1 Box B domains. In some embodiments, wherein the HMGB1 polypeptide comprises two HMGB1 Box B domains, the cryptic NLS is located C-terminal to the most C-terminal HMGB1 Box B domain. In some embodiments, wherein the HMGB1 polypeptide comprises two HMGB1 Box B domains, the sequence of EKSKKKK (SEQ ID NO: 931) is located C-terminal to the most C-terminal Box B domain. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box A domain. In some embodiments, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain, the HMGB1 Box B domain is located C-terminal to the HMGB1 Box A domain.
[0379] In some embodiments, wherein the HMGB1 polypeptide of amino acid residues 166-185 relative to SEQ ID NO: 928 is present, the HMGB1 polypeptide is located C-terminal to the most C-terminal Box B domain. When only one Box B domain is present, it is understood to be the most C-terminal Box B domain.
[0380] In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOS: 929-938 or that is encoded by a nucleic acid encoding the polypeptide.In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 935 or that is encoded by a nucleic acid encoding the polypeptide.IV.B. Base Editing Domain Fusions
[0381] An enhanced NmeCas9 variant described herein may further comprise a baseediting domain (e.g., a deaminase domain) that introduces a specific modification into a target nucleic acid. Accordingly, provided herein are enhanced NmeCas9 polypeptides that have been modified to act as a nickase (e.g., NmeCas9 polypeptides comprising a D16A substitution or a H588A substitution; and one or more enhancing amino acid substitutions described herein) and that further comprise a deaminase domain. Additionally provided are polynucleotides encoding such polypeptides. The NmeCas9 nickase and deaminase fusion may optionally further comprise one or more nuclear localization signals.
[0382] In some embodiments, the enhanced NmeCas9 polypeptide further comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, L846, E868, K870, K929, T930, E932, K962, K963, K965, Q967, K989, K1005, K1026, K1044, S1051, Q1053, orN1054 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an Nme2Cas9 polypeptide and comprises at least one amino acid substitution described in Table 3 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K / R, E508K / R, K517R, K549R, K555R, Q679R, L846K / Y, E868K / R, K870R, K929R, T930K, E932K / R, K962R, K963R, K965R, Q967R, K989R, K1005R, N1026K / R, K1044R, S1051K, Q1053K, or N1054K / R relative to SEQ ID NO: 1).
[0383] In some embodiments, the Nme2Cas9 polypeptide is a nickase comprising a RuvC substitution or an HNH substitution and further comprising four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions listed in Table 3 relative to the amino acid sequence of SEQ ID NO: 1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, S6 , G33, E47, R63,V68, K104, Al 16, T123, D152, E154, E221, F260, A263, A303, T396, H413, A427, D451, H452, E460, A484, A520, S629, S646, N674, V696, R711, D720, A724, V758, V765, Y767, K769, H771, S816, V821, D844, 1859, W865, K940, M951, K1005, D1028, S1029, N1031, R1033, K1044, Q1047, R1049, V1056, N1064, L1075, D844 , D870, D873, D911, D56, and E520 relative to SEQ ID NO: 1, optionally from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, K1005, and Q1053 relative to SEQ ID NO:1. In certain embodiments, the nickase comprises four or fewer (e.g., four or fewer, three or fewer, two or fewer, or one) additional amino acid substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1. In some embodiments, in addition to the nickase substitution, the amino acid sequence of the Nme2Cas9 polypeptide does not contain more than four additional amino acid substitutions relative to SEQ ID NO: 1, e.g., does not contain more than three additional amino acid substitutions relative to SEQ ID NO: 1, does not contain more than two additional amino acid substitutions relative to SEQ ID NO: 1, contains only one additional amino acid substitution relative to SEQ ID NO: 1.
[0384] In some embodiments, the NmeCas9 polypeptide further comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions K53, K145, S197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, P1003, S1022, K1040, G1051, K1053, T1054 of SEQ ID NO: 2 (i.e., a wildtype NmelCas9 polypeptide). In somesuch embodiments, the NmeCas9 polypeptide is an NmelCas9 polypeptide and comprises at least one amino acid substitution described in Table 4 (i.e., K53R, K145R, S197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, P1003K / R, S1022R, K1040R, G1051K, K1053R, or T1054K / R relative to SEQ ID NO: 2). In certain embodiments, the NmelCas9 polypeptide comprises an amino acid substitution at position S1022 of SEQ ID NO: 2 (e.g., an S1022R substitution).
[0385] In some embodiments, the NmeCas9 polypeptide further comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions K53, K145, G197, K266, K333, K351, K390, D418, K419, Q422, E508, K517, K549, K555, Q679, V847, Q866, K868, Q927, V928, K930, K958, S1003, G1022, S1040, G1050, K1052, or T1053 of SEQ ID NO: 3 (i.e., a wildtype Nme3Cas9 polypeptide). In some such embodiments, the NmeCas9 polypeptide is an Nme3Cas9 polypeptide and comprises at least one amino acid substitution described in Table 5 (i.e., K53R, K145R, G197K, K266R, K333R, K351R, K390R, D418K, K419R, Q422K, E508K / R, K517R, K549R, K555R, Q679R, V847K, Q866K / R, K868R, Q927K / R, V928K, K930R, K958R, S1003K / R, G1022R, S1040K / R, G1050K, K1052R, or T1053K / R substitution relative to SEQ ID NO: 3.
[0386] In certain embodiments, the enhanced NmeCas9 polypeptide further comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises any of the one or more of the exemplary amino acid substitutions provided herein. In certain embodiments, the one or more of the exemplary amino acid substitutions disclosed herein.
[0387] In some embodiments, the NmeCas9 polypeptide may further comprise a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0388] In some embodiments, the NmeCas9 polypeptide may further comprise a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 1026 and one or more (e.g., one or more, two or more, or all three) amino acid substitutions at positions 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, amino acid substitutions are at positions 1026 and one or more positions selected from 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1. In some embodiments, amino acid substitutions are at positions 1026 and no more than three positions selected from 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0389] In some embodiments, the NmeCas9 polypeptide may further comprise a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an additional amino acid substitution at position 932 and one or more (e.g., one or more, two or more, or three or more) amino acid substitutions at positions 1026, 266, 422, 418, 868, and 929 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, additional amino acid substitutions are at positions 932 and no more than three positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0390] In some embodiments, the NmeCas9 polypeptide may further comprise a cytidine deaminase (e.g, APOBEC3 A, alternatively referred to herein as A3 A), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0391] In some embodiments, the NmeCas9 polypeptide may further comprise an adenine deaminase, wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9polypeptide comprises one or more (e.g., one or more, two or more, three or more, or all four) amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0392] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 1026 (e.g., N1026) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises an N1026K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises an N1026R substitution relative to SEQ ID NO: 1.
[0393] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and N1026R substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 320. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and N1026R substitution) comprises the amino acid sequence of SEQ ID NO: 320.
[0394] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a N1026R and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 322. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a N1026R and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 322.
[0395] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and N1026K substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO:309. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and N1026K substitution) comprises the amino acid sequence of SEQ ID NO: 309.
[0396] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a N1026K and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 311. In some embodiments, theNmeCas9 polypeptide (e.g., comprising aN1026K and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 311.
[0397] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 932 (e.g., E932) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises an E932 substitution relative to SEQ ID NO: 1 to any amino acid except D, optionally to K, N, Q, M, R, H, A, S, or T. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises an E932 substitution and a D16A substitution relative to SEQ ID NO: 1.
[0398] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and E932 substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 1.
[0399] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 868 (e.g., E868) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickaseand the amino acid sequence of the NmeCas9 polypeptide comprises an E868R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises an E932R substitution and a D16A substitution relative to SEQ ID NO: 1.
[0400] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 929 (e.g., K929) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a K929R substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a K929R substitution and a D16A substitution relative to SEQ ID NO: 1.
[0401] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and E932 substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 1.
[0402] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 266 (e.g., K266) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a K266R substitution relative to SEQ ID NO: 1.
[0403] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and K266R substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, atleast 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 34. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and K266R substitution) comprises the amino acid sequence of SEQ ID NO: 34.
[0404] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a K266R and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 36. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a K266R and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 36.
[0405] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 422 (e.g., Q422) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a Q422K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a Q422R substitution relative to SEQ ID NO: 1.
[0406] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and Q422K substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 89. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and Q422K substitution) comprises the amino acid sequence of SEQ ID NO: 89.
[0407] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a Q422K and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO:91. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a Q422K and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 91.
[0408] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and Q422R substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 100. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and Q422R substitution) comprises the amino acid sequence of SEQ ID NO: 100.
[0409] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a Q422R and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 102. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a Q422R and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 102.
[0410] In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase (e.g., comprising a D16A substitution or a H588A substitution) and the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution at position 418 (e.g., D418) of SEQ ID NO: 1 (i.e., a wildtype Nme2 Cas9 polypeptide). In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a D418K substitution relative to SEQ ID NO: 1. In some embodiments, the NmeCas9 polypeptide comprises a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), wherein the NmeCas9 polypeptide is a nickase and the amino acid sequence of the NmeCas9 polypeptide comprises a D418R substitution relative to SEQ ID NO: 1.
[0411] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and D418K substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 67. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D16A and D418K substitution) comprises the amino acid sequence of SEQ ID NO: 67.
[0412] In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D418K and H588A substitution) comprises at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence of SEQ ID NO: 69. In some embodiments, the NmeCas9 polypeptide (e.g., comprising a D418K and H588A substitution) comprises the amino acid sequence of SEQ ID NO: 69.
[0413] Further provided herein are polynucleotides comprising an open reading frame encoding the above NmeCas9 polypeptides that comprise a deaminase domain. In some embodiments, the polypeptide comprising the deaminase and an RNA-guided nickase further comprises a uracil glycosylase inhibitor (UGI). In certain embodiments the polynucleotides comprising an open reading frame encoding one of the above NmeCas9 polypeptides that comprise a deaminase domain is provided with a polynucleotide providing a second open reading frame encoding a UGI. In certain embodiments, the polypeptide containing the open reading frame encoding one of the above NmeCas9 polypeptide, a first open reading frame, is in a separate polypeptide from the second open reading frame, i.e., the UGI is provided in trans.
[0414] In some embodiments, the polynucleotide comprises an open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g, APOBEC3 A, alternatively referred to herein as A3 A) and a C-terminal enhanced NmeCas9 nickase (e.g., comprising a D16A substitution or a H588A substitution and one or more enhancing amino acid substitutions described herein), and a first nuclear localization signal (NLS), wherein the polypeptide further comprises a uracil glycosylase inhibitor (UGI). In certain embodiments, the UGI is provided in a separate polypeptide from the NmeCas9 nickase further comprising the deaminase domain, i.e. the UGI is provided in trans.
[0415] In some embodiments, the polynucleotide comprises an open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3 A, alternatively referred to herein as A3 A) and a C-terminal enhanced NmeCas9 nickase (e.g., comprising a D16A substitution or a H588A substitution and one or more enhancing amino acid substitutions described herein), and a first nuclear localization signal (NLS), wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI). In some embodiments, the polypeptide is in combination with a polypeptide encoding UGI in trans.
[0416] In some embodiments, a second NLS is N-terminal to the NmeCas9 nickase. In some embodiments, the deaminase is N-terminal to an NLS (i.e., the first NLS or thesecond NLS). In some embodiments, the deaminase is N-terminal to all NLS in the polypeptide.
[0417] In some embodiments, the polynucleotide is DNA or RNA. In some embodiments, the polynucleotide is mRNA. In some embodiments, a polypeptide encoded by the mRNA is provided.
[0418] In some embodiments, the polypeptide comprises, from N to C terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A NmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and aD16ANme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, first and second NLSs, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A NmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, first and second NLSs, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A Nme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, A first NLS, a cytidine deaminase (e.g., APOBEC3A), a second NLS, an optional linker, and a D16ANmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, A first NLS, a cytidine deaminase (e.g., APOBEC3A), a second NLS, an optional linker, and a D16A Nme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein.
[0419] In some embodiments, the polypeptide comprises, from N to C terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3 A), an optional linker, and a H588A NmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and aH588ANme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, first and second NLSs, a cytidine deaminase (e.g., APOBEC3 A), an optional linker, and a H588A NmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, first and second NLSs, acytidine deaminase (e.g., AP0BEC3A), an optional linker, and a H588A Nme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, A first NLS, a cytidine deaminase (e.g., APOBEC3 A), a second NLS, an optional linker, and a H588A NmeCas9 nickase comprising one or more enhancing amino acid substitutions described herein. In some embodiments, the polypeptide comprises, from N to C terminus, A first NLS, a cytidine deaminase (e.g., APOBEC3 A), a second NLS, an optional linker, and a H588A Nme2Cas9 nickase comprising one or more enhancing amino acid substitutions described herein.
[0420] In some embodiments, the polypeptide comprising A3 A and an RNA-guided nickase further comprises a uracil glycosylase inhibitor (UGI).
[0421] In some embodiments, the polypeptide comprising A3 A and an RNA-guided nickase does not comprise a uracil glycosylase inhibitor (UGI). In certain embodiments, a UGI is provided in trans.
[0422] In some embodiments, a composition is provided comprising a first polypeptide, or an mRNA encoding a first polypeptide, comprising a cytidine deaminase, which is optionally an APOBEC3 A deaminase (A3 A); a C-terminal NmeCas9 nickase; a first nuclear localization signal (NLS); and, optionally, a second NLS; wherein the first NLS and, when present, the second NLS are located to N-terminal to the sequence encoding the NmeCas9 nickase, wherein the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI); and a second polypeptide, or an mRNA encoding a second polypeptide, comprising a uracil glycosylase inhibitor (UGI), wherein the second polypeptide is different from the first polypeptide.
[0423] In some embodiments, methods of modifying a target gene are provided comprising contacting a cell with the compositions described herein.
[0424] In some embodiments, the method comprises delivering to a cell a nucleic acid comprising an open reading frame encoding a polypeptide comprising a cytidine deaminase, which is optionally an APOBEC3 A deaminase (A3 A), and a NmeCas9 nickase (e.g., comprising a D16A substitution or a H588A substitution) having one or more amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1, wherein the polypeptide further comprises a uracil glycosylase inhibitor (UGI). In some embodiments, the one or more amino acid substitutions are at positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1.
[0425] In some embodiments, the method comprises delivering to a cell a nucleic acid comprising an open reading frame encoding a polypeptide comprising a cytidine deaminase, which is optionally an APOBEC3A deaminase (A3 A); a C-terminal NmeCas9 nickase (e.g., comprising a D16A substitution or a H588A substitution) having one or more amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1, optionally at one or more positions selected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1, wherein the polypeptide further comprises a uracil glycosylase inhibitor (UGI). In some embodiments, the polypeptide comprises a first nuclear localization signal (NLS); and, optionally, a second NLS; wherein the first NLS and, when present, the second NLS are located to N-terminal to the sequence encoding the NmeCas9 nickase.
[0426] In some embodiments, the method comprises delivering to a cell a first nucleic acid comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase, which is optionally an APOBEC3 A deaminase (A3 A); a C-terminal NmeCas9 nickase, wherein the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI), and a second nucleic acid comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), wherein the second nucleic acid is different from the first nucleic acid. In some embodiments, the first polypeptide comprises a first nuclear localization signal (NLS); and, optionally, a second NLS; wherein the first NLS and, when present, the second NLS are located to N-terminal to the sequence encoding the NmeCas9 nickase.
[0427] In some embodiments, the methods comprise delivering to a cell a polypeptide comprising a deaminase, which is optionally an APOBEC3 A deaminase (A3 A); a C-terminal NmeCas9 nickase, wherein the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding the polypeptide, and delivering to the cell a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding the UGI. In some embodiments, the first polypeptide comprises a first nuclear localization signal (NLS); and, optionally, a second NLS; wherein the first NLS and, when present, the second NLS are located to N-terminal to the sequence encoding the NmeCas9 nickase.
[0428] In some embodiments where the UGI is delivered in trans relative to the nickase, a molar ratio of the mRNA encoding UGI to the mRNA encoding the APOBEC3 A deaminase (A3 A) and an RNA-guided nickase is from about 1:35 to from about 30:1. In some embodiments, the molar ratio is from about 1:25 to about 25:1. In some embodiments, the molar ratio is from about 1 :20 to about 25 : 1. In some embodiments, the molar ratio is from about 1 : 10 to about 22: 1. In some embodiments, the molar ratio is from about 1:5 toabout 25 : 1. In some embodiments, the molar ratio is from about 1 : 1 to about 30: 1. In some embodiments, the molar ratio is from about 2:1 to about 10:1. In some embodiments, the molar ratio is from about 5:1 to about 20:1. In some embodiments, the molar ratio is from about 1 : 1 to about 25: 1. In some embodiments, the molar ratio may be about 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14: 1, 15:1, 16: 1, 17:1, 18:1, 19:1, 20:1, 21: 1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, or 30:1. In some embodiments, the molar ratio is equal to or larger than about 1 : 1. In some embodiments the molar ratio is about 1 : 1. In some embodiments the molar ratio is about 2: 1. In some embodiments the molar ratio is about 3 : 1. In some embodiments the molar ratio is about 4:1. In some embodiments the molar ratio is about 5: 1. In some embodiments the molar ratio is about 6: 1. In some embodiments the molar ratio is about 7: 1. In some embodiments the molar ratio is about 8: 1. In some embodiments the molar ratio is about 9: 1. In some embodiments the molar ratio is about 10:1. In some embodiments the molar ratio is about 11 : 1. In some embodiments the molar ratio is about 12: 1. In some embodiments the molar ratio is about 13 : 1. In some embodiments the molar ratio is about 14:1. In some embodiments the molar ratio is about 15:1. In some embodiments the molar ratio is about 16:1. In some embodiments the molar ratio is about 17:1. In some embodiments the molar ratio is about 18:1. In some embodiments the molar ratio is about 19:1. In some embodiments the molar ratio is about 20: 1. In some embodiments the molar ratio is about 21 : 1. In some embodiments the molar ratio is about 22: 1. In some embodiments the molar ratio is about 23 : 1. In some embodiments the molar ratio is about 24: 1. In some embodiments the molar ratio is about 25:1.
[0429] Similarly, in some embodiments, the molar ratio discussed above for the mRNA encoding the UGI protein to the mRNA encoding the APOBEC3 A deaminase (A3 A) and an RNA-guided nickase are similar if delivering protein.
[0430] In some embodiments, the composition described herein further comprises at least one gRNA. In some embodiments, a composition is provided that comprises an mRNA described herein and at least one gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a dual guide RNA (dgRNA).
[0431] In some embodiments, the composition is capable of effecting genome editing upon contacting the cell.IV. C. Cytidine deaminase Fusions; APOBEC 3 A Deaminase
[0432] Cytidine deaminases encompass enzymes in the cytidine deaminase superfamily, and in particular, enzymes of the APOBEC family (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274:18470-6, 1999); and Carrington et al., Cells 9:1690 (2020)).
[0433] In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC family. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC 1, APOBEC2, APOBEC4, and APOBEC3 subgroups. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC3 subgroup. In some embodiments, the cytidine deaminase disclosed herein is an APOBEC3 A deaminase (A3 A). In some embodiments, the deaminase comprises an APOBEC3 A deaminase.
[0434] In some embodiments, an APOBEC3 A deaminase (A3 A) disclosed herein is a human A3 A. In some embodiments, an APOBEC3 A deaminase (A3 A) disclosed herein is a human A3 A. In some embodiments, the A3 A is a wild-type A3 A.
[0435] In some embodiment, the A3 A is an A3 A variant. A3 A variants share homology to wild-type A3 A, or a fragment thereof. In some embodiments, a A3 A variant has at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to a wild type A3 A. In some embodiments, the A3A variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a wild type A3 A. In some embodiments, the A3 A variant comprises a fragment of an A3 A, such that the fragment has at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to the corresponding fragment of a wild-type A3 A.
[0436] In some embodiments, an A3A variant is a protein having a sequence that differs from a wild-type A3 A protein by one or several mutations, such as substitutions, deletions, insertions, one or several single point substitutions. In some embodiments, ashortened A3 A sequence could be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids. In some embodiments, a shortened A3 A sequence is used where one to four amino acids at the C-terminus of the sequence is deleted. In some embodiments, an APOBEC3 A (such as a human APOBEC3 A) has a wild-type amino acid position 57 (as numbered in the wild-type sequence). In some embodiments, an APOBEC3 A (such as a human APOBEC3 A) has an asparagine at amino acid position 57 (as numbered in the wildtype sequence).
[0437] In some embodiments, the wild-type A3A is a human A3A (UniProt accession ID: p319411, SEQ IDNO: 370).
[0438] In some embodiments, the A3 A disclosed herein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 370. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the A3 A comprises an amino acid sequence having at least 87% identity to SEQ ID NO: 370. In some embodiments, the A3A comprises an amino acid sequence with at least 90% identity to SEQ ID NO: 370. In some embodiments, the A3A comprises an amino acid sequence with at least 95% identity to SEQ ID NO: 370. In some embodiments, the A3 A comprises an amino acid sequence with at least 98% identity to SEQ ID NO: 370. In some embodiments, the A3A comprises an amino acid sequence with at least 99% identity to A3A ID NO: 370. In some embodiments, the A3A comprises the amino acid sequence of SEQ ID NO: 370.
[0439] In some embodiments, the cytidine deaminase comprises the amino acid sequence of any one of the cytidine deaminases described in International Patent Publication No. WO2023081689, which is hereby incorporated by reference (e.g., see SEQ ID NOs: 151-216 of WO2023081689A1), or an amino acid sequence having at least 80%, at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity thereto.IV. C. UGI
[0440] Without being bound by any theory, providing a UGI together with a polypeptide comprising a deaminase may be helpful in the methods described herein by inhibiting cellular DNA repair machinery (e.g., UDG and downstream repair effectors) that recognize a uracil in DNA as a form of DNA damage or otherwise would excise or modify the uracil or surrounding nucleotides. It should be understood that the use of a UGI may increase the editing efficiency of an enzyme that is capable of deaminating C residues.
[0441] Suitable UGI protein and nucleotide sequences are provided herein and additional suitable UGI sequences are known to those in the art, and include, for example, those published in Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264: 1163-1171(1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419(1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887(1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346(1999), the entire contents of each are incorporated herein by reference. It should be appreciated that any proteins that are capable of inhibiting a uracil-DNA glycosylase base-excision repair enzyme are within the scope of the present disclosure. Additionally, any proteins that block or inhibit base-excision repair as also within the scope of this disclosure. In some embodiments, a uracil glycosylase inhibitor is a protein that binds uracil. In some embodiments, a uracil glycosylase inhibitor is a protein that binds uracil in DNA. In some embodiments, a uracil glycosylase inhibitor is a single-stranded binding protein. In some embodiments, a uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein. In some embodiments, a uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein that does not excise uracil from the DNA. In some embodiments, a uracil glycosylase inhibitor is a catalytically inactive UDG.
[0442] In some embodiments, a uracil glycosylase inhibitor (UGI) disclosed herein comprises an amino acid sequence with at least 80% to SEQ ID NO: 369. In some embodiments, any of the foregoing levels of identity is at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the UGI comprises an amino acid sequence with at least 90% identity to SEQ ID NO: 369. In some embodiments, the UGI comprises an amino acid sequence with at least 95% identity to SEQ ID NO: 369. In some embodiments, the UGI comprises an amino acid sequence with at least 98% identity to SEQ ID NO: 369. In some embodiments, the UGI comprises an amino acid sequence with at least 99% identity to SEQ ID NO: 369. In some embodiments, the UGI comprises the amino acid sequence of SEQ ID NO: 369.IV.D. Compositions comprising an AP0BEC3A deaminase and a NmeCas9
[0443] In some embodiments, provided herein is an mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., A3 A) and an RNA-guided nickase (e.g., an NmeCas9 polypeptide comprising one or more enhancing amino acid substitutions described herein (e.g., see Table 3, Table 4, or Table 5). In some embodiments, the polypeptide comprises a human deaminase (e.g., A3 A) and a C-terminal RNA-guided nickase; and a nucleotide sequence encoding a first NLS and optionally a second NLS. In certain embodiments, the deaminase is N-terminal to an NLS. In certain embodiments, the deaminase is N-terminal to all NLS.
[0444] In some embodiments, provided herein is an mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., A3 A) and an RNA-guided nickase (e.g., an enhanced NmeCas9 nickase comprising a D16A substitution or a H588A substitution; and one or more enhancing amino acid substitutions described herein). In some embodiments, the polypeptide comprises a human deaminase (e.g., A3 A) and a C-terminal RNA-guided nickase; and a nucleotide sequence encoding a first NLS and optionally a second NLS. In certain embodiments, the deaminase is N-terminal to an NLS. In certain embodiments, the deaminase is N-terminal to all NLS.
[0445] In some embodiments, the polypeptide comprises a wild-type deaminase (e.g., A3 A) and a C-terminal RNA-guided nickase. In some embodiments, the polypeptide comprises an A3 A variant and an RNA-guided nickase. In some embodiments, the polypeptide comprises a deaminase (e.g., A3 A) and a Cas9 nickase. In some embodiments, the polypeptide comprises a deaminase (e.g., A3 A) and aD16ANmeCas9 nickase. In some embodiments, the polypeptide comprises a human deaminase (e.g., A3 A) and a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises an A3 A variant and a D16ANmeCas9 nickase. In some embodiments, the polypeptide lacks a UGI. In some embodiments, the deaminase (e.g., A3 A) and the RNA-guided nickase are linked via a linker. In some embodiments, the polypeptide further comprises one or more additional heterologous functional domains. In some embodiments, the polypeptide further comprises a nuclear localization sequence (NLS) (described herein).
[0446] In some embodiments, the polypeptide comprises a human deaminase (e.g., A3 A) and a C-terminal D16ANmeCas9 nickase (e.g., comprising one or more amino acid substitutions at positions 1026, 266, 422, and 418 of SEQ ID NO: 1), optionally at positionsselected from 1026, 266, 422, 418, 932, 868, and 929 relative to SEQ ID NO: 1, wherein the human deaminase (e.g., A3 A) and the D16ANmeCas9 are fused via a linker. In some embodiments, the polypeptide comprises a human A3 A and a C-terminal D16A NmeCas9 nickase, and a NLS at the N- terminus of the fused polypeptide. In some embodiments, the polypeptide comprises a human A3 A and a C-terminal D16ANmeCas9 nickase, wherein the human A3 A and the D16A NmeCas9 are fused via a linker, and a NLS fused to the N-terminus of the human A3 A, optionally via a linker.
[0447] The polypeptide may be organized in any number of ways to form a single chain. The first NLS and, when present, the second NLS are located to N-terminal to the sequence encoding the Cas9 nickase. Additional NLS can be N-terminal to the Cas9 nickase. The A3A can be N- or C-terminal as compared an NLS. In some embodiments, the polypeptide comprises, from N to C terminus, a first NLS, an optional second NLS, a deaminase, an optional linker, an RNA-guided nickase, and an optional NLS. In some embodiments, linkers are independently present between the first and second NLS, and an NLS and a deaminase. In some embodiments, the polypeptide comprises, from N to C terminus, a deaminase, a first NLS, an optional second NLS, a C-terminal RNA-guided nickase. In some embodiments, linkers are independently present between a deaminase and a first NLS, between a first NLS and a second NLS, and between an NLS and a C-terminal nickase.IV.E. Combinations comprising a polymerase and a NmeCas9
[0448] In certain embodiments, the NmeCas9 polypeptides provided herein are used in combination with a polymerase. In certain embodiments, the NmeCas9 nickase polypeptides provided herein are used in combination with a polymerase. The NmeCas9 nickase preferably comprises a substitution in the HNH nuclease domain (e.g., a H588A substitution in Nme2Cas9). Once bound to the target region, the NmeCas9 nickase may generate a nick in a non-target strand of the DNA duplex target nucleic acid, i.e., the strand (positive or negative) of the DNA duplex target nucleic acid that comprises a sequence that is complementary to the target sequence of the gRNA.
[0449] In certain embodiments, the polymerase is a reverse transcriptase. In certain embodiments, the polymerase is not a reverse transcriptase. In certain embodiments, the polymerase is a DNA-dependent DNA polymerase. As used herein, a “DNA-dependent DNA polymerase” or “DNA-templated DNA polymerase” (DDP) synthesizes DNA opposite a nucleic acid template preferentially comprised of DNA over RNA. In some embodiments, theextension activity of a DDP on DNA templates consisting only of DNA nucleotides is at least 4-fold higher than the extension activity of the DDP on RNA templates consisting only of RNA nucleotides, as measured, for example, in an assay using a template guide RNA.
[0450] In certain embodiments, DNA-dependent DNA polymerases are prokaryotic DNA polymerases. In certain embodiments, DNA-dependent DNA polymerases are viral DNA polymerases. In certain embodiments, DNA-dependent DNA polymerases are eukaryotic DNA polymerases.
[0451] In certain embodiments, DNA-dependent DNA polymerases are wild-type DNA polymerases. In certain embodiments, DNA-dependent DNA polymerases are variant DNA polymerases, i.e., not wild-type DNA polymerases. In certain embodiments, a variant DNA polymerase has altered activity as compared to the wild-type polymerase. For example, in certain embodiments, the processivity of the variant polymerase may be increased as compared to the wild-type polymerase. In certain embodiments, an exonuclease or certain exonuclease activity of the variant polymerase may be decreased or abolished as compared to the wild-type polymerase.In certain embodiments, the DNA polymerase is selected from DNA polymerase kappa (polK), phage T5 DNA polymerase (T5 pol), DNA polymerase 0, E. coli DNA polymerase I, or DNA polymerase N, optionally wherein the DNA-dependent DNA polymerase is polK or T5 DNA polymerase.
[0452] In certain embodiments, the polymerase is provided as a fusion protein with the NmeCas9 polypeptide. In certain embodiments, the polymerase is provided in trans, i.e., is not a fusion protein. In certain embodiments, the polymerase is provided in an ORF encoding the polymerase and the NmeCas9. Exemplary amino acid and nucleotide sequences for fusion with the NmeCas9 sequences herein are provided, for example, inWO2025101994.V. Guide RNA
[0453] In some embodiments, at least one guide RNA is provided in combination with an enhanced NmeCas9 polypeptide disclosed herein (e.g., comprising one or more enhancing amino acid substitutions described herein, e.g., see Table 3, Table 4, or Table 5) or polynucleotide comprising an ORF encoding the enhanced NmeCas9 polypeptide. In some embodiments, a guide RNA is provided as a separate molecule from the NmeCas9polypeptide or polynucleotide. In some embodiments, a guide RNA is provided as a part, such as a part of a UTR, of a polynucleotide disclosed herein.
[0454] In some embodiments, a composition comprising the polynucleot...
Claims
CLAIMSWhat is claimed is:
1. A Neisseria meningitidis Cas9 (NmeCas9) polypeptide comprising an amino acid sequence with at least 90% identity to the amino acid sequence of SEQ ID NO: 1, wherein the amino acid sequence of the NmeCas9 polypeptide comprises one or more substitution(s) at a position(s) selected from the group consisting of N1026, K266, Q422, D418, E932, E868, K929, Q422, E508, K517, K549, K555, K53, G197, Q679, L846, T930, K962, K965, Q967, KI 005, and Q1053 relative to SEQ ID NO: 1.
2. The NmeCas9 polypeptide of claim 1, wherein the amino acid sequence of the NmeCas9 polypeptide comprises one or more substitution(s) at a position(s) selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
3. The NmeCas9 polypeptide of claim 1, wherein the substitution is at a position selected from the group consisting of N1026, K266, Q422, and D418 of SEQ ID NO:
14. The NmeCas9 polypeptide of any one of claims 1-3, wherein the amino acid sequence of the NmeCas9 polypeptide comprises an N1026 substitution relative to SEQ ID NO: 1.
5. The NmeCas9 polypeptide of claim 4, wherein the N1026 substitution is a N1026R substitution relative to SEQ ID NO: 1.
6. The NmeCas9 polypeptide of claim 4, wherein the N1026 substitution is a N1026K substitution relative to SEQ ID NO: 1.
7. The NmeCas9 polypeptide of any one of claims 1-6, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a K266 substitution relative to SEQ ID NO: 1.
8. The NmeCas9 polypeptide of claim 7, wherein the K266 substitution is a K266R substitution relative to SEQ ID NO: 1.
9. The NmeCas9 polypeptide of any one of claims 1-8, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a Q422 substitution relative to SEQ ID NO: 1.
10. The NmeCas9 polypeptide of claim 9, wherein the substitution is a Q422K or a Q422R substitution relative to SEQ ID NO: 1.
11. The NmeCas9 polypeptide of any one of claims 1-10, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a D418 substitution relative to SEQ ID NO: 1.
12. The NmeCas9 polypeptide of claim 11 , wherein the substitution is a D418K or a D418R substitution relative to SEQ ID NO: 1.
13. The NmeCas9 polypeptide of any one of claims 1-12, wherein the amino acid sequence of the NmeCas9 polypeptide comprises an E932 substitution relative to SEQ ID NO: 1.
14. The NmeCas9 polypeptide of claim 13, wherein the E932 substitution is a substitution selected from the group consisting of E932K, E932N, E932Q, E932M, E932R, E932H, E932A, E932S, and E932T substitution relative to SEQ ID NO: 1.
15. The NmeCas9 polypeptide of any one of claims 1-14, wherein the amino acid sequence of the NmeCas9 polypeptide comprises an E868 substitution relative to SEQ ID NO: 1.
16. The NmeCas9 polypeptide of claim 15, wherein the E868 substitution is an E868R or a E868K substitution relative to SEQ ID NO: 1.
17. The NmeCas9 polypeptide of any one of claims 1-16, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a K929 substitution relative to SEQ ID NO: 1.
18. The NmeCas9 polypeptide of claim 17, wherein the K929 substitution is a K929R substitution relative to SEQ ID NO: 1.
20. The NmeCas9 polypeptide of any one of claims 1-3, wherein the amino acid sequence of the NmeCas9 polypeptide comprises an amino acid substitution selected from the group consisting of N1026R, N1026K, K266R, E932K, E932N, E932Q, E932M, E932H, E932A, E932S, E932T, E868, Q422, D418, E932, K929, and KI 044 relative to SEQ ID NO: 1.
21. The NmeCas9 polypeptide of any one of claims 1-20, wherein the amino acid sequence of the NmeCas9 polypeptide further comprises a K333 substitution relative to SEQ ID NO:
122. The NmeCas9 polypeptide of any one of claims 1-21, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a single substitution at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
23. The NmeCas9 polypeptide of claim 22, wherein the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
24. The NmeCas9 polypeptide of any one of claims 1 -22, wherein the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution.
25. The NmeCas9 polypeptide of claim 24, wherein the nickase comprises single substitution relative to SEQ ID NO: 1 at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
26. The NmeCas9 polypeptide of any one of claims 1 -22, wherein the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution.
27. The NmeCas9 polypeptide of claim 26, wherein the dCas9 comprises single substitution relative to SEQ ID NO: 1 at a position selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
28. The NmeCas9 polypeptide of any one of claims 1-21, wherein the amino acid sequence of the NmeCas9 polypeptide comprises two substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
29. The NmeCas9 polypeptide of claim 28, wherein the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of:a) E932 and any one of N1026, K266, E868, Q422, D418, or K929;b) N1026 and any one of E932, K266, E868, Q422, D418, or K929;c) K266 and any one of E932, N1026, E868, Q422, D418, or K929;d) E868 and any one of E932, N1026, K266, Q422, D418, or K929;e) Q422 and any one of E932, N1026, K266, E868, D418, or K929;f) D418 and any one of E932, N1026, K266, E868, Q422, or K929; andg) K929 and any one of E932, N1026, K266, E868, Q422, or D418 relative to SEQ ID NO: 1.
30. The NmeCas9 polypeptide of claim 28 or 29, wherein the amino acid sequence of the NmeCas9 polypeptide comprises substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) E932 and N1026;b) E932 and E868;c) E932 and K266;d) E932 and D418;e) E932 and Q422;f) K266 and N1026;g) Q422 and N1026;h) D418 and N1026; andi) E868 and N1026.
31. The NmeCas9 polypeptide of any one of claims 28-30, wherein the amino acid sequence of the NmeCas9 polypeptide comprises two amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) N1026R and K266R;b) N1026Rand D418K;c) N1026Rand E868R;d) N1026Rand E868K;e) N1026Rand E932N;f) N1026Rand E932Q;g) N1026Rand E932M;h) N1026Rand E932R;i) N1026Rand E932H;j) N1026Rand E932A;k) N1026Rand E932S;l) N1026Rand E932T;m) N1026K and E932N;n) N1026K and E932Q;o) N1026K and E932M;p) N1026K and E932R;q) N1026K and E932H;r) N1026K and E932A;s) N1026K and E932S;t) N1026K and E932T;u) E932R and K266R;v) E932RandD418K;w) E932R and Q422K;x) E932R and E868R;y) E932R and N1026R;z) E868R and N1026K;aa) E932N and N1026R;bb) E932M and N1026R; andcc) E932A and N1026R relative to SEQ ID NO: 1.
32. The NmeCas9 polypeptide of any one of claims 28-31, wherein the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
33. The NmeCas9 polypeptide of any one of claims 28-31, wherein the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution.
34. The NmeCas9 polypeptide of claim 33, wherein the nickase comprises two substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
35. The NmeCas9 polypeptide of any one of claims 28-31, wherein the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution.
36. The NmeCas9 polypeptide of claim 35, wherein the dCas9 comprises two substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
37. The NmeCas9 polypeptide of any one of claims 1-21, wherein the amino acid sequence of the NmeCas9 polypeptide comprises three substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
38. The NmeCas9 polypeptide of claim 37, wherein the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of:a) E932 and any two of N1026, K266, E868, Q422, D418, and K929;b) N1026 and any two of E932, K266, E868, Q422, D418, and K929;c) K266 and any two of E932, N1026, E868, Q422, D418, and K929;d) E868 and any two of E932, N1026, K266, Q422, D418, and K929;e) Q422 and any two of E932, N1026, K266, E868, D418, or K929;f) D418 and any two of E932, N1026, K266, E868, Q422, and K929; andg) K929 and any two of E932, N1026, K266, E868, Q422, and D418 relative to SEQ ID NO: 1.
39. The NmeCas9 polypeptide of claim 37 or 38, wherein the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) E932, K266, and D418;b) K266, Q422, and N1026;c) K266, Q422, and E868;d) Q422, E868, and N1026;e) K266, E868, and N1026;f)E932, D418, and Q422;g) E868, E932, and N1026;h) K266, E868, and N1026; andi) D418, E868, and N1026 relative to SEQ ID NO: 1.
40. The NmeCas9 polypeptide of any one of claims 37-39, wherein the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) E932R, K266R, and Q422K;b) E932R, K266R, and D418K;c) E932R, K266R, and D422K;d) K266R, Q422K, and E868R;f) E932R, D418K, and Q422K;g) E868R, E932R, and N1026K;h) K266R, E868R, and N1026R; andi) D418K, E868R, and N1026R relative to SEQ ID NO: 1.
41. The NmeCas9 polypeptide of any one of claims 37-40, wherein the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
42. The NmeCas9 polypeptide of any one of claims 37-40, wherein the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution.
43. The NmeCas9 polypeptide of claim 42, wherein the nickase comprises three substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
44. The NmeCas9 polypeptide of any one of claims 37-40, wherein the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution.
45. The NmeCas9 polypeptide of claim 44, wherein the dCas9 comprises three substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
46. The NmeCas9 polypeptide of any one of claims 1-21, wherein the amino acid sequence of the NmeCas9 polypeptide comprises four substitutions at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929 relative to SEQ ID NO: 1.
47. The NmeCas9 polypeptide of claim 46, wherein the amino acid sequence of the NmeCas9 polypeptide comprises substitutions at positions selected from the group consisting of:a) E932 and any three of N1026, K266, E868, Q422, D418, and K929;b) N1026 and any three of E932, K266, E868, Q422, D418, and K929;c) K266 and any three of E932, N1026, E868, Q422, D418, and K929;d) E868 and any three of E932, N1026, K266, Q422, D418, and K929;e) Q422 and any three of E932, N1026, K266, E868, D418, or K929;f) D418 and any three of E932, N1026, K266, E868, Q422, and K929; andg) K929 and any three of E932, N1026, K266, E868, Q422, and D418 relative to SEQ ID NO: 1.
48. The NmeCas9 polypeptide of claim 46 or 47, wherein the amino acid sequence of the NmeCas9 polypeptide comprises four substitutions at positions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) K266, Q422, E868, and N1026;b) E932, K266, D418, and Q422;c) E932, K266, K333, D418, and Q422;d) E932, K266, K333, D418, Q422, E508, K517, and K549; ande) E932, K266, D418, Q422, E508, K517, and K549 relative to SEQ ID NO: 1.
49. The NmeCas9 polypeptide of any one of claims 46-48, wherein the amino acid sequence of the NmeCas9 polypeptide comprises amino acid substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of:a) K266R, Q422K, E868R, and N1026R;b) E932, K266R, D418K, and Q422K;c) E932, K266R, K333R, D418K, and Q422K; andd) E932R, K266R, K333R, D418K, Q422K, E508K, K517R, and K549R; and e) E932R, K266R, D418K, Q422K, E508K, K517R, K549R relative to SEQ ID NO: 1.
50. The NmeCas9 polypeptide of any one of claims 46-49, wherein the amino acid sequence of the NmeCas9 polypeptide does not contain any additional mutations relative to SEQ ID NO: 1.
51. The NmeCas9 polypeptide of any one of claims 46-49, wherein the NmeCas9 is a nickase comprising a RuvC domain substitution or an HNH domain substitution.
52. The NmeCas9 polypeptide of claim 51, wherein the nickase comprises four substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the nickase does not contain any additional mutations relative to SEQ ID NO: 1.
53. The NmeCas9 polypeptide of any one of claims 46-49, wherein the NmeCas9 is a dCas9 comprising a RuvC domain substitution and an HNH domain substitution.
54. The NmeCas9 polypeptide of claim 53, wherein the dCas9 comprises four substitutions relative to SEQ ID NO: 1 at positions selected from the group consisting of N1026, K266, Q422, D418, E932, E868, and K929, wherein the dCas9 does not contain any additional mutations relative to SEQ ID NO: 1.
55. The NmeCas9 polypeptide of any one of claims 1-54 comprising an amino acid sequence with at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1.
56. The NmeCas9 polypeptide of claim 1, wherein the NmeCas9 polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 32, 65, 87, 98, 307, or 318.
57. The NmeCas9 polypeptide of any one of claims 1-56, wherein the NmeCas9 polypeptide is an Nme2Cas9 polypeptide, an NmelCas9 polypeptide, or an Nme3Cas9 polypeptide.
58. The NmeCas9 polypeptide of claim 57, wherein the NmeCas9 polypeptide is a Nme2 Cas9 polypeptide.
59. The NmeCas9 polypeptide of any one of claims 1-58 wherein the amino acid sequence of the NmeCas9 polypeptide does not comprise an E932D relative to SEQ ID NO: 1.
60. The NmeCas9 polypeptide of any one of claims 1-59, wherein the amino acid sequence of the NmeCas9 polypeptide does not comprise an E932R substitution relative to SEQ ID NO: 1.
61. The NmeCas9 polypeptide of any one of claims 1-60, wherein the amino acid sequence of the NmeCas9 polypeptide comprises neither an E932D substitution nor an E932R substitution relative to SEQ ID NO: 1.
62. The NmeCas9 polypeptide of any one of claims 1-61, wherein the amino acid sequence of the NmeCas9 polypeptide does not comprise an E868K relative to SEQ ID NO: 1.
63. The NmeCas9 polypeptide of any one of claims 1-62, wherein the amino acid sequence of the NmeCas9 polypeptide does not comprise a K929R substitution relative to SEQ ID NO: 1.
64. The NmeCas9 polypeptide of any one of claims 1-63, wherein the amino acid sequence of the NmeCas9 polypeptide does not comprise a substitution at KI 044 relative to SEQ ID NO: 1.
65. A Neisseria meningitidis Cas9 (NmeCas9) polypeptide comprising an amino acid sequence with at least 90% identity to the amino acid sequence of SEQ ID NO: 2 or 3, wherein the amino acid sequence of the NmeCas9 polypeptide comprises a substitution at SI 022 of SEQ ID NO: 2 or G1022 of SEQ ID NO:3.
66. The NmeCas9 polypeptide of claim 65, wherein the NmeCas9 polypeptide comprises a substitution at SI 022 relative to SEQ ID NO: 2 and comprises an amino acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 2.
67. The NmeCas9 polypeptide of claim 65, wherein the NmeCas9 polypeptide comprises a G1022 substitution relative to SEQ ID NO: 3 and comprises an amino acid sequence that is at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 3.
68. The NmeCas9 polypeptide of any one of claims 1-67, wherein the NmeCas9 polypeptide further comprises a nuclear localization signal (NLS).
69. The NmeCas9 polypeptide of claim 68, wherein the NLS is selected from a c-Myc NLS, an SV40 NLS, a nucleoplasmin NLS, or a snurportin-1 importin-P NLS.
70. The NmeCas9 polypeptide of any one of claims 1-69, further comprising one or more additional heterologous functional domains.
71. The NmeCas9 polypeptide of claim 70, wherein the one or more additional heterologous functional domains is selected from an HMGB1 domain, a deaminase, a uracil glycosylase inhibitor (UGI), and a polymerase.
72. The NmeCas9 polypeptide of claim 70, wherein the one or more additional heterologous functional domains is a deaminase.
73. The NmeCas9 polypeptide of claim 72, wherein the deaminase is a cytidine deaminase or an adenine deaminase.
74. The NmeCas9 polypeptide of claim 73, wherein the cytidine deaminase is an apolipoprotein B mRNA editing enzyme (APOBEC) deaminase.
75. The NmeCas9 polypeptide of any one of claims 1-74, wherein the NmeCas9 has cleavase activity.
76. The NmeCas9 polypeptide of claim 75, wherein the polypeptide has increased editing activity as compared to a corresponding wild type NmeCas9 polypeptide.
77. A polynucleotide comprising an open reading frame (ORF) that encodes the NmeCas9 polypeptide of any one of claims 1-76.
78. The polynucleotide of claim 77, wherein the ORF has been codon optimized for increased translation of the mRNA in a mammal.
79. The polynucleotide of claim 78, wherein the ORF has been codon optimized for increased translation of the mRNA in a human.
80. The polynucleotide of claim 77 or 78, wherein the increased translation is relative to the extent of translation of a wild type sequence of the ORF, or relative to an ORF having a codon distribution matching the codon distribution of the organism from which the ORF was derived.
81. The polynucleotide of any one of claims 77-80, wherein the polynucleotide is an mRNA.
82. A vector comprising the polynucleotide of any one of claims 77-80.
83. The vector of claim 82, wherein the vector is a viral vector.
84. The vector of claim 83, wherein the viral vector is an adeno-associated virus (AAV) vector.
85. The vector of any one of claims 82-84, wherein the vector further encodes one or more gRNAs.
86. The vector of claim 85, wherein the one or more gRNAs is a single guide RNA (sgRNA).
87. The vector of claim 85 or 86, wherein the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA.
88. The vector of any one of claims 85-87, wherein the one or more gRNAs comprises a scaffold region that binds the NmeCas9 polypeptide, and a targeting region that hybridizes with a target genomic sequence in one or more cells of interest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide.
89. The vector of claim 88, wherein the PAM sequence recognized by the NmeCas9 polypeptide is N4CC.
90. A cell comprising the Nme polypeptide of any one of claims 1-76, the polynucleotide of any one of claims 77-81, or the vector of any one of claims 82-89.
91. A lipid nanoparticle (LNP) comprising the polynucleotide of any one of claims 77-81.
92. The LNP of claim 91, wherein the LNP comprises an ionizable lipid.
93. A composition comprisingone or more guide RNAs (gRNAs); andthe NmeCas9 polypeptide of any one of claims 1-76, the polynucleotide of any one of claims 77-81; the vector of any one of claims 82-89; or the LNP of claim 91 or 92.
94. The composition of claim 93, wherein the one or more gRNAs is a single gRNA (sgRNA).
95. The composition of claim 93 or 94, wherein the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA.
96. The composition of claim 95, wherein the full-length sgRNA is a full-length Nme sgRNA as set forth in SEQ ID NO: 475.
97. The composition of claim 96, wherein the shortened gRNA is 101 nucleotides in length, 102 nucleotides in length, 103 nucleotides in length, 104 nucleotides in length, 105 nucleotides in length, or 115 nucleotides in length.
98. The composition of claim 97, wherein the shortened gRNA is 115 nucleotides in length.
99. The composition of claim 95, wherein the shortened sgRNA comprises a scaffold region that binds a NmeCas9 polypeptide, wherein the scaffold region comprises a polynucleotide sequence having at least 90% identity to any one of SEQ ID NOs 572-592.
100. The composition of claim 99, wherein the scaffold region comprises a polynucleotide sequence having at least 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs 572-592.
101. The composition of claim 99, wherein the scaffold region comprises the polynucleotide sequence of any one of SEQ ID NOs: 572-592.
102. The composition of claim 95, wherein the scaffold region comprises at least 90% identity to SEQ ID NO: 576.
103. The composition of claim 102, wherein the scaffold region comprises at least 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 576.
104. The composition of claim 102, wherein the scaffold region comprises SEQ ID NO: 576.
105. The composition of any one of claims 93-104, wherein the one or more gRNAs is a chemically modified gRNA.
106. The composition of any one of claims 93-105, wherein the one or more gRNAs binds the NmeCas9 polypeptide and hybridizes with a target sequence in one or more cells ofinterest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide.
107. The composition of claim 106, wherein the PAM sequence recognized by the NmeCas9 polypeptide is N4CC.
108. The composition of any one of claims 93-107, wherein the composition further comprises a template for a polymerase.
109. A pharmaceutical composition comprisingthe NmeCas9 polypeptide of any one of claims 1-76; the polynucleotide of any one of claims 77-81; the vector of any one of claims 82-89; or the LNP of claim 91 or 92; or the composition of any one of claims 93-108; anda pharmaceutically acceptable carrier.
110. A method for binding a target sequence in a cell comprising delivering to the cell: one or more guide RNAs (gRNAs); andthe NmeCas9 polypeptide of any one of claims 1-76; the polynucleotide of any one of claims 77-81; the vector of any one of claims 82-89; or the LNP of claim 91 or 92; or the pharmaceutical composition of claim 109;thereby binding the target sequence with the NmeCas9 polypeptide of the composition.
111. A method for cleaving a target sequence in a cell comprising delivering to the cell:one or more guide RNAs (gRNAs); andthe NmeCas9 polypeptide of any one of claims 1-76; the polynucleotide of any one of claims 77-81; the vector of any one of claims 82-89; or the LNP of claim 91 or 92; or the pharmaceutical composition of claim 109;wherein the NmeCas9 polypeptide cleaves the target sequence.
112. A method for modifying a target sequence in a cell comprising delivering to the cell:one or more guide RNAs (gRNAs); andthe NmeCas9 polypeptide of any one of claims 1-76; the polynucleotide of any one of claims 77-81; the vector of any one of claims 82-89; or the LNP of claim 91 or 92; or the pharmaceutical composition of claim 109;wherein the NmeCas9 polypeptide modifies the target sequence.
113. The method of any one of claims 110-112, wherein the one or more gRNAs is a single gRNA (sgRNA).
114. The method of any one of claims 110-112, wherein the one or more gRNAs is a shortened sgRNA relative to a full-length sgRNA.
115. The method of claim 114, wherein the full-length sgRNA is a full-length Nme sgRNA as set forth in SEQ ID NO: 475.
116. The method of any one of claims 110-115, wherein the gRNA comprises a scaffold region that binds the NmeCas9 polypeptide and a targeting region that hybridizes with a target sequence in one or more cells of interest, wherein the target sequence is located upstream of a Protospacer Adjacent Motif (PAM) sequence that is recognized by the NmeCas9 polypeptide.
117. The method of any one of claims 110-115, wherein the PAM sequence recognized by the NmeCas9 polypeptide is N4CC.
118. A method for binding a target sequence in a cell comprising delivering to the cell the vector of any one of claims 82-89 or the composition of any one of claims 93-108, thereby binding the target sequence with the NmeCas9 polypeptide of the composition.
119. A method for cleaving a target sequence in a cell comprising delivering to the cell the vector of any one of claims 82-89 or the composition of any one of claims 93-108, wherein the NmeCas9 polypeptide cleaves the target sequence.
120. A method for modifying a target sequence in a cell comprising delivering to the cell the vector of any one of claims 82-89 or the composition of any one of claims 93-108, wherein the NmeCas9 polypeptide modifies the target sequence.
121. The method of any one of claims 110-120, wherein the method is an ex vivo method.
122. The method of any one of claims 110-120, wherein the method is an in vivo method.
123. The method of any one of claims 110-122, wherein the target sequence is a target gene in a genome of the cell.
124. The method of any one of claims 110-123, wherein the cell is a mammalian cell.
125. The method of any one of claims 110-124, wherein the cell is a human cell.
126. The method of any one of claims 110-125, wherein the NmeCas9 is Nme2Cas9.