Improved CAS9 proteins
Amino acid-modified Cas9 proteins with enhanced activity and fidelity address the size and delivery limitations of traditional Cas9, facilitating efficient genome editing and therapeutic use.
Patent Information
- Application Number
- PCT/EP2025/069621
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-10
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-15
AI Technical Summary
Existing Cas9 proteins are too large for efficient delivery via methods like AAV and have low activity, limiting their suitability for therapeutic genome editing applications.
Development of a Cas9 protein with specific amino acid modifications, enhancing its endonuclease activity and fidelity, and incorporating nuclear localization signals, allowing for smaller size suitable for AAV delivery.
The modified Cas9 proteins exhibit significantly higher activity and fidelity, enabling effective genome editing and therapeutic applications, particularly in human cells.
Smart Images

Figure EP2025069621_15012026_PF_FP_ABST
Abstract
Description
IMPROVED CAS9 PROTEINSFIELD OF THE INVENTION
[0001] The present disclosure generally relates to improved Cas9 proteins, CRISPR-Cas9 systems and methods for providing site-specific modification of a target sequence in a eukaryotic cell using improved Cas9 proteins.BACKGROUND
[0002] Genome editing has undergone significant evolution, with systems like zinc fingers and TALENs showing limitations in precision and ease of use. In contrast, CRISPR- Cas9 systems offer greater ease of design and implementation. The CRISPR-Cas9 systems originate as a natural bacterial immune system, where they function as a defence mechanism against viral infections. These systems use guide RNA to target specific DNA sequences with a nuclease protein (Cas9), making them versatile and easy to implement across various applications, including genome editing.
[0003] The CRISPR-Cas9 gene editing system has been used successfully in a wide range of organisms and cell lines, both in order to induce double stranded break (DSB) formation in DNA using the wild type Cas9 protein or to nick a single DNA strand using a mutant protein termed Cas9n / Cas9 D10A (see, e.g., Mali etal., Science, 339 (6121): 823-826 (2013) and Sander and Joung, Nature Biotechnology 32(4): 347- 355 (2014), each of which is incorporated by reference herein in its entirety).
[0004] In addition to genome editing, the CRISPR system has a multitude of other applications, including regulating gene expression, genetic circuit construction, and functional genomics, amongst others (reviewed in Sander and Joung, 2014).
[0005] While some Cas9 proteins have been shown to be effective in a wide variety of in vivo and in vitro applications, their relatively large size is not suitable for fitting into delivery methods (e.g., AVV) for therapeutic applications.
[0006] Cas9 proteins have been isolated from different bacteria, including S. aureus and S. thermophilus. Cas9 proteins from Campylobacter lanienae have been previously described by WO2019099943 Al (which is incorporated by reference herein in its entirety), however the wild-type enzyme displays a very low activity, which is incompatible with genome editing for therapeutic applications.
[0007] There is a need for a highly effective Cas9 enzyme, which has also potentially a compatible size to be packaged into delivery methods, for instance AAV, and thus which is suitable for genome editing therapy in humans.SUMMARY OF THE INVENTION
[0008] The present disclosure is directed to a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises an amino acid modification at one or more positions relative to SEQ ID NO: 1 : 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229. In one embodiment, said polypeptide sequence consists of an amino acid modification at one or more positions relative to SEQ ID NO: 1: 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122,1160, 1161, 1164, 1206, 1226 and / or 1229.
[0009] The amino acid modification may be an amino acid substitution.
[0010] According to one embodiment, said Cas9 protein is an engineered and / or a non- naturally occurring Cas9 protein.
[0011] According to another embodiment, the amino acid modification is an amino acid substitution at one or more positions relative to SEQ ID NO: 1 : 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160,1161, 1164, 1206, 1226 and / or 1229.
[0012] According to one embodiment, said Cas9 protein comprises an enhanced endonuclease activity compared to a wild type Cas9 protein, such as a wild type CllCas9 protein, optionally wherein said endonuclease activity is measured by a gene editing reporter assay described herein and / or as described in Degtev et al., NatCommun 15, 9173 (2024) (which is incorporated by reference herein in its entirety) using the engineered Xential HEK293T cell line (see, for instance, Example 1 and / or Example 2). According to one embodiment, said endonuclease activity is measured as defined in Example 1 and / or Example 2. Said wild type CllCas9 protein may be the protein as set forth in SEQ ID NO: 1.
[0013] According to one embodiment, said Cas9 protein comprises an enhanced endonuclease activity compared to a wild type Cas9, such as SpCas9 protein; optionally wherein said endonuclease activity is measured by using the a gene editing reporter assay described herein and / or as described in Degtev et al., Nat Commun 15, 9173 (2024) (which is incorporated by reference herein in its entirety) using the engineered Xential HEK293T cell line (see, for instance, Example 1 and / or Example 2). According to one embodiment, said endonuclease activity is measured as defined in Example 1 and / Example 2. Said SpCas9 protein is the wild type SpCas9 (Cas9 from Streptococcus pyogenes) known in the art. Said SpCas9 protein may be as defined in SEQ ID NO: 238.
[0014] According to one embodiment, said Cas9 protein comprises an enhanced endonuclease activity and / or an enhanced fidelity compared to SEQ ID NO 211 (eSpotON), optionally wherein said endonuclease activity and / or fidelity is measured by using the editing reporter and amplicon sequencing methods described in Degtev et al., et al., Nat Commun 15, 9173 (2024) (which is incorporated by reference herein in its entirety) using the engineered Xential HEK293T cell line or wild type HEK293T cells, respectively (see, for instance, Example 1 and / or Example 2). According to one embodiment, said endonuclease activity is measured as defined in Example 1 and / Example 2.
[0015] According to one embodiment, the polypeptide sequence comprises or consists of one or more amino acid modifications selected from D1073A, A1160R, D1073G, D1091G, D1091N, D1091R, D1091S, D1122R, D1206G, D1206N, D1206R, D1206S, D37G, D37N, D37R, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D911K, D911R, D911S, D912G, D912N, D912R, D912S, E1164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, G39R, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, Q556R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
[0016] According to one embodiment, the polypeptide sequence comprises one or more amino acid modifications selected from D1073A, A1160R, D1073G, D1091G, D1091N, D1091R, D1091S, DI 122R, D1206G, D1206N, D1206R, D1206S, D37G, D37N, D37R, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D911S, D912G, D912N, D912R, D912S, E1164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, G39R, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, Q556R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
[0017] According to one embodiment, the polypeptide sequence comprises a first amino acid modification selected from D37R, D1091R, D912R, G39R, I978R, N1161R, and S1229R.
[0018] According to one embodiment, wherein the polypeptide sequence further comprises a second amino acid modification selected from D1091R, D912R, G39R, N1161R, and S1229R, optionally wherein the first amino acid modification is D37R and the second amino acid modification is selected from D1091R, N1161R, D912R, S1229R, and G39R, optionally wherein the first amino acid modification is D1091R and the second amino acid modification is selected from N1161R, S1229R, and D912R.
[0019] According to one embodiment, the Cas9 protein further comprises a third amino acid modification selected from A1160R, DI 073 A, D1073G, D1091G, DI 09 IN, D1091R, D1091S, D1122R, D1206G, D1206R, D1206S, D37G, D37N, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D911S, D912G, D912N, D912R, D912S, E1164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
[0020] According to one embodiment, the Cas9 protein further comprises a fourth amino acid modification selected from D1206S, D59S, D912N, D912R, D912S, E1164R, E1226G, E1226S, F94Y, I46R, Q556R, S1229R, S566R, and S57R.
[0021] According to one embodiment, the Cas9 protein further comprises a fifth amino acid modification selected from D1206N, D1206S, D912N, D912R, D912S, G39R, I46R, and S1229R, optionally a sixth amino acid modification selected fromS1229R, S566R, and optionally a seventh amino acid modification at position D912N.
[0022] According to one embodiment, said Cas9 protein comprises one or more mutations as defined in Table 2, or wherein said Cas9 protein comprises or consists of an amino acid sequence as defined in any one of SEQ ID NOs: 9-144, SEQ ID NOs: 14-144, SEQ ID NOs: 30-144, SEQ ID NOs: 60-144, SEQ ID NOs: 112-144, and / or SEQ ID NOs: 140-144.
[0023] According to another embodiment, said Cas9 protein further comprises at least one nuclear localization signal (NLS), such as a first nuclear localization signal attached to the N-terminus of the Cas9 protein and / or a second nuclear localization signal attached to the C-terminus of the Cas9 protein. In some embodiments, said at least one nuclear localization signal (NLS) is as defined in Table 3, such as VC, 53BP1, bpSV40, BRCAl l, BRCA1 2, EWS, FUS, HIV-I TAT, hnRNP-D, hnRNP-M, HsNPM2, HTLV-1, mpSV40, MYC, NLP, NPM2, PTHrP and / or REV HV1H2.
[0024] In some embodiments, said Cas9 protein further comprises at least one nuclear localization signal (NLS) attached to the N-terminus of the Cas9 protein, such as 53BP1, bpSV40, BRCAl l, BRCA1 2, FUS, HIV-I TAT, hnRNP-D, HsNPM2, HTLV-1, mpSV40, MYC, NLP, NPM2, PTHrP, and / or REV HV1H2.
[0025] In some embodiments, said Cas9 protein further comprises at least one nuclear localization signal (NLS) attached to the C-terminus of the Cas9 protein, such as 53bpl, 53BP1, bpSV40, BRCAl l, BRCA1 2, EWS, FUS, HIV-I TAT, hnRNP- D, hnRNP-M, HsNPM2, HTLV-1, mpSV40, MYC, NLP, PTHrP and / or REV HV1H2.
[0026] In some embodiments, said Cas9 protein further comprises at least one nuclear localization signal (NLS) attached to the N-terminus of the Cas9 protein, such as bpSV40, NLP, or NPM2, and said Cas9 protein further comprises at least one nuclear localization signal (NLS) attached to the C-terminus of the Cas9 protein, such as NLP, bpSV40, 53bpl, MYC, or mpSV40.
[0027] Any one or more NLS as defined in Table 3 may be combined with any one of the Cas9 proteins of SEQ ID NO: 9-144, such as SEQ ID NOs: 14-144, SEQ ID NOs:30-144, SEQ ID NOs: 60-144, SEQ ID NOs: 112-144, and / or SEQ ID NOs: 140- 144. Said NLS may advantageously increase the activity of the Cas9 proteins as defined herein.
[0028] In some embodiments, the Cas9 protein further comprises one or more NLS and is as set forth in any one or more of SEQ ID NOs: 161-194, SEQ ID NOs: 167-194, SEQ ID NOs: 178-194, SEQ ID NOs: 182-194, SEQ ID NOs: 189-194 or SEQ ID NOs: 192-194. In another embodiment, the Cas9 protein is as set forth in SEQ ID NO: 193 or 194, such as SEQ ID NO: 194.
[0029] According to one embodiment, said Cas9 protein comprises a NLS attached to the N-terminus which is selected from bpSV40, NLP, 53BP1 and NPM2, and / or said Cas9 protein comprises a NLS attached to the C-terminus which is selected from PTHrP, NLP, bpSV40, HTLV-1, 53bpl, MYC and mpSV40.
[0030] According to one embodiment, the Cas9 protein comprises a NLS at the N-terminus and / or at the C-terminus, wherein said NLS is any one NLS as set forth in SEQ ID NO: 239-244. For instance, the N-terminus NLS may be any one of SEQ ID NO: 239-244, and the C-terminus NLS may be any one of SEQ ID NO: 239-244.
[0031] Another aspect of the disclosure relates to a CRISPR-Cas system comprising a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises or consists of an amino acid modification at one or more positions relative to SEQ ID NO: 1: 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229, or a Cas9 protein as defined elsewhere herein; and a guide polynucleotide comprising a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell. In some embodiments, the eukaryotic cell may be an animal or a human cell.
[0032] Another aspect of the disclosure relates to a CRISPR-Cas system comprising a nucleic acid sequence encoding a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises or consists of an amino acid modification at one or more positions relative to SEQ ID NO: 1: 37, 39, 46, 47, 54, 57, 59, 61, 94,315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229, or a Cas9 protein as defined elsewhere herein; and a nucleic acid sequence encoding a guide polynucleotide that comprises a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell. In some embodiments, the eukaryotic cell may be an animal or a human cell.
[0033] According to one embodiment, the guide polynucleotide is an RNA and / or the nucleic acid sequence is a DNA.
[0034] According to one embodiment, the guide polynucleotide has from 80 to 140 bases in length, such as from 90 to 130 bases in length.
[0035] According to one embodiment, the guide polynucleotide comprises or consists of a polynucleotide sequence as defined in any one of SEQ ID NO: 195-209. In some embodiments, the spacer sequence comprises or consists of SEQ ID NO: 210.
[0036] Another aspect of the disclosure relates to a eukaryotic cell or a delivery particle comprising the Cas9 protein as defined elsewhere herein, or the CRISPR-Cas system as defined elsewhere herein. The delivery particle may be a vesicle or a viral vector.
[0037] According to another embodiment, the delivery particle is a vesicle or a viral vector. In some embodiments, the delivery particle is a lipid-based system, a liposome, a micelle, a microvesicle, an exosome or a gene gun. In some embodiments, the delivery particle is a lipid envelope. In some embodiments, the delivery particle is a sugar-based particle, for example, GalNAc. In some embodiments, the delivery particle is a nanoparticle. Preparation of delivery particles is further described in U.S. Patent Publication Nos. 2011 / 0293703, 2012 / 0251560, and 2013 / 0302401; and U.S. Patent Nos. 5,543 ,158, 5,855,913, 5,895,309, 6,007,845, and 8,709,843, each of which is incorporated herein by reference in its entirety.
[0038] Another aspect of the disclosure relates to a method for providing site-specific modification of a target sequence in a eukaryotic cell, the method comprising:a) introducing into the eukaryotic cell a nucleotide sequence encoding a guide polynucleotide that forms a complex with a Cas9 protein , wherein said guide polynucleotide comprises a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in host polynucleotide; and wherein the eukaryotic cell comprises the Cas9 protein as defined elsewhere herein; b) generating cohesive ends in the host polynucleotide with the Cas9 protein and the guide polynucleotide; and c) ligating i) the cohesive ends of (b) together, or ii) a 3' end of a polynucleotide sequence of interest to a first cohesive end of the host polynucleotide, and a 5’ end of the polynucleotide sequence of interest to a second cohesive end of the host polynucleotide; thereby modifying the target sequence.
[0039] According to another embodiment, step a) further comprises i) introducing into the eukaryotic cell a nucleotide sequence encoding the Cas9 protein as defined elsewhere herein, and ii) expressing the Cas9 protein by the eukaryotic cell.
[0040] According to another embodiment, the guide polynucleotide is an RNA.
[0041] According to another embodiment, the guide polynucleotide has from 80 to 140 bases in length, such as from 90 to 130 bases in length.
[0042] According to one embodiment, the guide polynucleotide further comprises a tracrRNA sequence. A "tracrRNA," or trans-activating CRISPR-RNA, forms an RNA duplex with a pre-crRNA, or pre-CRISPR-RNA, and is then cleaved by the RNA specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide RNA activates the Cas9 protein.
[0043] According to another embodiment, the guide polynucleotide comprises or consists of a polynucleotide sequence as defined in any one of SEQ ID NO: 195-209.
[0044] Another aspect of the disclosure relates to a method of modifying a target DNA, the method comprising: - contacting said target DNA with a Cas9 protein, said Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises an amino acid modification at one or more positions relative to SEQ ID NO: 1: 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229.
[0045] According to one embodiment, the method further comprises adding a guide polynucleotide for hybridizing with the target DNA.
[0046] According to one embodiment, said Cas9 protein comprises or consists of an amino acid sequence as defined in any one of SEQ ID NOs: 9-144, SEQ ID NOs: 14-144, SEQ ID NOs: 30-144, SEQ ID NOs: 60-144, SEQ ID NOs: 112-144, and / or SEQ ID NOs: 140-144.
[0047] Another aspect of the disclosure relates a method of treating a genetic disorder or a genetic disease in a subject comprising administering to said subject an effective amount of the CRISPR-Cas system as defined herein, the eukaryotic cell as defined herein, or the delivery particle as defined herein.
[0048] According to one embodiment, the genetic disorder or the genetic disease is caused by a DNA mutation, such as a point mutation, a deletion, an insertion, a duplication, or a repeat, in relation to a DNA sequence of a healthy subject.BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The following drawings form part of the present specification and are included to further demonstrate exemplary embodiments of certain aspects of the present invention.
[0050] Fig. 1 provides a diagram representation of the reporter system used for comparing nucleases in terms of editing activity. Nanoluciferase release into the medium isproportional to the targeted cleavage of the described reporter cassette integrated into the HBEGF locus.
[0051] Fig. 2 - First generation of engineered CllpCas9 enzymes compared in terms of editing reporter activation. Representative previous generation CllpCas9 (namely wild type or wt) shown for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0052] Fig. 3 - Second generation of engineered CllpCas9 enzymes compared in terms of editing reporter activation. The enzymes ePsCas9 and SpCas9 are included as controls. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0053] Fig. 4 - Third generation (Part 1) of engineered CllpCas9 enzymes compared in terms of editing reporter activation. All samples were run side-by-side with the ones presented in Figures 5 and 6, wtCHCas9 and previous generation CllpCas9 variants are shown on all figures for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0054] Fig. 5 - Third generation (Part 2) of engineered CllpCas9 enzymes compared in terms of editing reporter activation. All samples were run side-by-side with the ones presented in Figures 4 and 6, wtCHCas9 and previous generation CllpCas9 variants are shown on all figures for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0055] Fig 6. Third generation (Part 3) of engineered CllpCas9 enzymes compared in terms of editing reporter activation. All samples were run side-by-side with the ones presented in Figures 4 and 5, wtCHCas9 and previous generation CllpCas9 variants are shown on all figures for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0056] Fig. 7 - Fourth generation of engineered CllpCas9 enzymes compared in terms of editing reporter activation. Representative previous generation CllpCas9 variants shown for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0057] Fig. 8 - Fifth generation of engineered CllpCas9 enzymes compared in terms of editing reporter activation. Representative previous generation CllpCas9 variants shown for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0058] Fig. 9 - Sixth generation of engineered CllpCas9 enzymes compared in terms of editing reporter activation. Representative previous generation CllpCas9 variants shown for reference. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0059] Fig. 10 - Editing of the PCSK9 gene by representative CllpCas9 variants evaluated by next-generation sequencing. Well characterized ePsCas9 and SpCas9 with their corresponding guide RNAs are also shown.
[0060] Fig. 11-15 - Editing reporter activation by v4 CllpCas9 fused to different NLS combinations. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9 with no NLS (nlO).
[0061] Fig. 16 - Editing reporter activation by representative CllpCas9 variants fused to different NLS combinations (nl401, nl416, nl417) andnot fused to NLS (v004, v376, v392).
[0062] Fig. 17 - Editing reporter activation by wtCllpCas9 and engineered guide RNAs. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0063] Fig. 18 - Editing reporter activation by v376 CllpCas9 and engineered guide RNAs. A) Raw luminescence signal. B) Fold change of signal (editing) compared to wtCHCas9.
[0064] Fig. 19(A-C)- Editing activity of representative CllpCas9 variants compared to well characterized SpCas9 and PsCas9 enzymes on different targets. A) Editing of the T cell receptor alpha constant (TRAC) gene. B) Editing of the “safe harbor” AAVS1 locus. C) Editing of the dolichol-phosphate mannose synthase subunit 2 (DPM2) gene.
[0065] Fig. 20 - Off target activity of representative CllpCas9 variants compared to well characterized SpCas9 and PsCas9 enzymes using a known promiscuous spacer sequence in their respective guide RNAs. A) Editing of the commonly used target (HEK4) for benchmarking on- and off-target effects of genome editing tools. B) Editing of the known HEK4 off-target site 1. C) Editing of the known HEK4 off- target site 1.
[0066] Figure 21 - Sequence listing table.DETAILED DESCRIPTION OF THE INVENTION
[0067] The present disclosure provides a surprisingly highly active and improved Cas9 enzyme. Furthermore, said improved Cas9 enzyme is significantly smaller than Cas9 enzymes typically used in the field, such as SpCas9 or ePsCas9, which advantageously provides a Cas9 enzyme suitable for packaging and delivering to a cell, for instance by AAV delivery.
[0068] The inventors have surprisingly found that by mutating specific amino acid residue(s) within CllCas9 protein sequence, a substantially inactive wild-type enzyme may be transformed into an extremely active nuclease. Besides having an enhanced nuclease activity, it has also been surprisingly found that said Cas9 protein also comprises an improved fidelity, which may further improve therapeutic uses in a subject, such as a human patient. Furthermore, the highly active CllCas9 proteins are significantly smaller than traditional Cas9 proteins, thereby advantageously enabling AAV delivery.
[0069] Unless otherwise defined herein, scientific and technical terms used in the present disclosure shall have the meanings that are commonly understood by one of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. As used herein, "a" or "an" may mean one or more. As used herein, when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more than one. As used herein, "another" or "a further" may mean at least a second or more.
[0070] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include codedand non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.
[0071] The terms “CllCas9” and “CllpCas9” are used interchangeably and refer to a Cas9 protein expressed by Campylobacter lanienae.
[0072] The term “amino acid modification” refers to any amino acid modification, including amino acid substitution(s).
[0073] Percent identity of polynucleotides or polypeptides can be determined when the polynucleotide or polypeptide sequences are aligned over a specified comparison window. In some embodiments, only specific portions of two or more sequences are aligned to determine sequence identity. In some embodiments, only specific domains of two or more sequences are aligned to determine sequence similarity. A comparison window can be a segment of at least 10 to over 1000 residues, at least 20 to about 1000 residues, or at least 50 to 500 residues in which the sequences can be aligned and compared. Methods of alignment for determination of sequence identity are well-known and can be performed using publicly available databases such as BLAST. For example, in some embodiments, "percent identity" of two amino acid sequences is determined using the algorithm of Karlin and Altschul, Proc Nat Acad Sci USA 87:2264-2268 (1990), modified as in Karlin and Altschul, Proc Nat Acad Sci USA 90:5873-5877 (1993). Such algorithms are incorporated into BLAST programs, e.g., BLAST+ ortheNBLAST andXBLAST programs described in Altschul et al., J Mo! Biol, 215: 403-410 (1990). BLAST protein searches can be performed with programs such as, e.g., the XBLAST program, score=50, wordlength=3 to obtain amino acid sequences homologous to the protein molecules of the disclosure. Where gaps exist between two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Res 25(17): 3389-3402 (1997). When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used.
[0074] As used herein, the terms "comprising" (and any variant or form of comprising, such as "comprise" and "comprises"), "having" (and any variant or form of having, such as "have" and "has"), "including" (and any variant or form of including, such as "includes" and "include") or "containing" (and any variant or form of containing,such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited, elements or method steps.
[0075] As used herein, a "Cas protein" encompasses both Cas nucleases and Cas nickases. Cas proteins are part of the CRISPR / Cas system described herein. CRISPR / Cas systems, which include a Cas protein and a polynucleotide (also referred to as a "guide polynucleotide"), can be utilized for site-specific genome modifications.
[0076] As used herein, the term “enhanced endonuclease activity” refers to a Cas9 protein having a substantially higher endonuclease activity compared to a Cas9 protein in an unmodified state. Said endonuclease activity may be measured using the methods as described herein.
[0077] As used herein, the term “enhanced fidelity” refers to a Cas9 protein showing substantially reduced off-target activity compared to a reference Cas9 protein, such as CHCas9, SpCas9, or PsCas9. Said enhanced fidelity may refer to a Cas9 enzyme comprising one or more of the following properties: lower off-target cutting of DNA, less tolerance to DNA mismatches.
[0078] As used herein, the term "modification" of a target sequence encompasses singlenucleotide substitutions, multiple-nucleotide substitutions, insertions (i.e., knock-in) and deletions (i. e., knock-out) of a nucleic acid, frameshift mutations, and other nucleic acid modifications.
[0079] Following initial publications around the CRISPR-Cas9 system (Type II system), Cas9 variants have been identified in a range of bacterial species and a number have been functionally characterized. See, e.g., Chylinski et al., "Classification and evolution of type II CRISPR-Cas systems", Nucleic Acids Research 42(10): 6091- 6105 (2014), Ran et al., "In vivo genome editing using Staphylococcus aureus Cas9", Nature 520(7546): 186-91 (2015), and Esvelt et al., "Orthogonal Cas9 proteins for RNA-guided gene regulation and editing", Nature Methods 10(11): 1116-1121 (2013), each of which is incorporated by reference herein in its entirety.
[0080] The present document encompasses novel Cas9 proteins having an enhanced activity and / or fidelity.
[0081] According to one embodiment, the wild-type CllCas9 protein is as defined in Table 1 or SEQ ID NO: 1.
[0082] According to another embodiment, said Cas9 protein further comprises at least one nuclear localization signal (NLS), such as VC, 53BP1, bpSV40, BRCAl l, BRCA1 2, EWS, FUS, HIV-I TAT, hnRNP-D, hnRNP-M, HsNPM2, HTLV-1, mpSV40, MYC, NLP, NPM2, PTHrP and / or REV HV1H2. Advantageously, the present document shows that an enhanced Cas9 enzyme may have its activity significantly improved by addition of at least one, such as two, NLS at the N and / or C-terminal portion of the enzyme.
[0083] Another aspect of the disclosure relates to a CRISPR-Cas system comprising a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises or consists of an amino acid modification at one or more positions of 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229, or a Cas9 protein as defined elsewhere herein; and a guide polynucleotide comprising a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell. In some embodiments, said Cas9 protein is as defined in Table 2 or said Cas9 protein comprises or consists of an amino acid sequence as defined in any one of SEQ ID NOs: 9-144, SEQ ID NOs: 14-144, SEQ ID NOs: 30-144, SEQ ID NOs: 60-144, SEQ ID NOs: 112-144, and / or SEQ ID NOs: 140-144.
[0084] According to one embodiment, the guide polynucleotide comprises or consists of a polynucleotide sequence as defined in any one of SEQ ID NO: 195-209. Advantageously, the present document shows that a CRISPR-Cas system comprising an enhanced Cas9 enzyme may have its activity significantly improved by combining the enzyme with mutated gRNA sequences, such as the gRNA sequences as defined in any one of SEQ ID NO: 195-209.Table 1 - Wild-type CllCas9 protein sequence
[0085] Another aspect of the disclosure relates to a method for providing site-specific modification of a target sequence in a eukaryotic cell, the method comprising a) introducing into the eukaryotic cell a nucleotide sequence encoding a guide polynucleotide that forms a complex with a Cas9 protein and comprises a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in host polynucleotide; and wherein the eukaryotic cell comprises the Cas9 protein as defined above; b) generating cohesive ends in the host polynucleotide with the Cas9 protein and the guide polynucleotide; and c) ligating i) the cohesive ends of (b) together, or ii) a 3' end of a polynucleotide sequence of interest to a first cohesive end of the host polynucleotide, and a 5’ end of the polynucleotide sequence of interest to a second cohesive end of the host polynucleotide; thereby modifying the target sequence.
[0086] A "modification" of a target sequence encompasses single-nucleotide substitutions, multiple-nucleotide substitutions, insertions (i.e., knock-in) and deletions (i. e., knock-out) of a nucleic acid, frameshift mutations, and other nucleic acid modifications.
[0087] In embodiments of the method, the eukaryotic cell is an animal or human cell. In embodiments, the eukaryotic cell is an animal cell as described herein. In embodiments, the eukaryotic cell is a human cell. In embodiments, the eukaryotic cell is a human cell as described herein. In embodiments, the eukaryotic cell is a plant cell.
[0088] In embodiments of the method, the modification is deletion of at least part of the target sequence. In embodiments, the modification is mutation of the target sequence. In embodiments, the modification is inserting a sequence of interest into the target sequence. In embodiments, the modification is a modification as described herein.EXAMPLES
[0089] Example 1 - HEK293T Cell Genome Editing Reporter Assay
[0090] The editing reporter system was generated by using the Xential method as per Li S et al., Nat Commun 12, 497 (2021) (which is incorporated by reference herein in its entirety) on wild type HEK293T cells. Briefly, a reporter cassette coding for a constitutively expressed gene was integrated into the HBGEF locus. The expressed gene terminates with an array of stop codons flanked by short repeats and a downstream non-expressed NanoLuc® Luciferase gene. The sequence of the reporter cassette integrated into the HBEGF locus is as defined in SEQ ID NO: 212. The editing reporter system is described in Degtev et al., Nat Commun 15, 9173 (2024) (which is incorporated by reference herein in its entirety).
[0091] For reporter activation assays, the Xential reporter cell line was transfected with plasmids expressing Cas9 variants fused with designated nuclear localization sequences. Transfections were conducted by introducing a 5 pL mix containing Opti-MEM™ (Thermo), 50 ng plasmid and 0.3 pL FuGene (Promega) to each well containing 0.3 x 10A6 cells in a 96 well format, following the FuGene reagent use guidelines. After 24h, a designated amount of guide RNA targeting the stop codon array within the reporter cassette was introduced by using Lipofectamine™ RNAiMAX Transfection Reagent (Thermo) according to manufacturer’s instructions. The use of the editing reporter system is described in Degtev et al., Nat Commun 15, 9173 (2024) (which is incorporated by reference herein in its entirety).
[0092] Unless stated otherwise, guide RNA sequences (spacer underlined) used for targeting the integrated reporter cassette were CGAUUACCCUGUUAUCCCUACUGUUUCAAACCUUCGGCGUUUGAAAU CACCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCG GUUA (SEQ ID NO: 213), GUGCGCGCCGAGUAGGGAUAACGUUUCAGUUGAAAGAAGUCACUCU UAAAGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACU GACUCUGUCAUCCGGGU (SEQ ID NO: 214), AUUACCCUGUUAUCCCUACUGUUUUAGAGCUAGAAAUAGCAAGUUA AAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCG GUGCUUUU (SEQ ID NO: 215) for CllpCas9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O- methylated bases and phosphorothioated bonds in first and last 3 positions.
[0093] Introduced guide RNA causes Cas9 mediated cleavage of the stop codon array within the reporter cassette, which in case of microhomology mediated end joining repair leads to the expression of NanoLuc® that is secreted into the medium. 72 hours post guide RNA transfection, media was exchanged to a fresh one followed by 3-6h incubation. Equal amounts of media from treated reporter cells and Nano- Glo® Luciferase substrate (Promega) diluted at 1 : 1000 in PBS were mixed and were immediately subjected to a luminescence readout on PheraStar FSX plate reader (BMG Labtech). This method allows for a luminescence represented readout of Cas9 mediated gene editing in cells.
[0094] Editing activity of CllpCas9 variants was compared to wtCHCas9 using this system. All variants including the wtCHCas9 were fused to N-terminal bpSV40 and C- terminal mpSV40 nuclear localization sequences unless specifically designated otherwise. This system was also used to assess the functionality of different CllpCas9 guide designs.
[0095] The editing activity is shown in Table 2. On each generation (GEN), a number of mutations are screened, based on the CllCas9 proteins obtained in the previous generation. Generation “0” is the wild type CHCas9 protein. It is clear from theresults that the mutations that generated an improved enzyme were not predictable. In fact, some mutations may turn an improved enzyme of a previous generation into a non-active enzyme in the next generation. As shown in Figures 2-10, modified versions of the CllCas9 enzyme are extremely more active (from dozens to more than a hundred times more active) than the wild-type Cll Cas9 enzyme. Furthermore, as shown in Figures 3, 9 and 10, some CllCas9 variants are at least as active as a highly optimized enzyme known in the art, namely ePsCas9 (also known as eSpotON). Obtaining an enzyme which is at least as active (or more active) as ePsCas9 and which is also significantly smaller is particularly advantageous, for example in the context of gene therapy, where it is easier and more effective to package and deliver a smaller Cas9 enzyme into the cell, for instance using AAV delivery.
[0096] As shown in Table 3 and Figures 11-16, the editing activity of the Cas9 enzyme is significantly improved by the selection of one or more NLS to be attached to the N- and / or C-terminal end of the Cas9 enzyme. Some NLS and / or combinations thereof are particularly advantageous, for instance the ones selected from SEQ ID NO: 148- 194, SEQ ID NO: 161-194, SEQ ID NO: 168-194, SEQ ID NO: 173-194, SEQ ID NO: 178-194, SEQ ID NO: 182-194, SEQ ID NO: 185-194, SEQ ID NO: 189-194, and / or SEQ ID NO: 192-194. Variants 4, 376 and 392 of the CllCas9 have been used merely as an exemplary proof of concept, and the skilled person would expect that a similar improvement is achieved when combining the above NLS with any other CllCas9 variant disclosed herein, as shown on representative variants in Figures 11-16.
[0097] The amino acid sequences of the NLS are shown in Table 7.
[0098] In some embodiments, guide RNAs have been modified to improve the activity of the CRISPR-Cas system. Methods for modifying gRNAs are known in the art, for example in Riesenberg et al (Nature communications, 2021), which is incorporated by reference in its entirety. Improved gRNAs are shown in Table 4, Figures 17-18 and SEQ ID NO: 195-209. SEQ ID NO: 210 refers to the spacer sequence comprised in each of the gRNAs. Said gRNA are capable of significantly increasing the activity of the CllCas9, in particular the gRNA as defined in SEQ ID NO: 195, 201, 203 and / or 209. Wild-type enzyme and variant 376 of the CHCas9 have been usedmerely as an exemplary proof of concept, and the skilled person would expect that a similar improvement is achieved when combining the above gRNAs with any other CllCas9 variant disclosed herein.
[0099] As shown in Table 4, the first 22 nucleic acids from each one of the guide RNAs as set forth in SEQ ID NOs: 195-209 are the spacer sequence as defined in SEQ ID NO: 210. Therefore, the remaining nucleic acid sequences of each one of SEQ ID NO: 195-209 - from nucleic acid 23 onwards - are the scaffold sequences specific for each protein used (e.g. SpCas9, ePsCas9 or CllpCas9 variants).Table 2 - CllCas9 variantsTable 3 - CllCas9 and NLSTable 4 - Guide RNAsTable 5 - eSpotON protein sequence
[0100] Example 2 - Editing Efficiency and Off-target Analysis by Amplicon Sequencing
[0101] Wild-type HEK293T cells were treated identical as the reporter strain (described above) with the exception of using guide RNA targeting the designated gene.
[0102] Guide RNA sequences (spacer underlined) used for targeting the PCSK9 gene were UCCAGGUUCCACGGGAUGCUCUGUUUCAAACCUUCGGCGUUUGAAAUCA CCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCGGUUA (SEQ ID NO: 216), UCCAGGUUCCACGGGAUGCUCUGUUUCAGUUGAAAGAAGUCACUCUUAA AGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACUGACUC UGUCAUCCGGGU (SEQ ID NO: 217) and CAGGUUCCACGGGAUGCUCUGUUUUAGAGCUAGAAAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC UUUU (SEQ ID NO: 218) for CllpCas9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O-methylated bases and phosphorothioated bonds in first and last 3 positions.
[0103] Guide RNA sequences (spacer underlined) used for targeting the TRAC gene were GGUCAGGGUUCUGGAUAUCUGUGUUUCAAACCUUCGGCGUUUGAAAUC ACCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCGGUU A (SEQ ID NO: 219), GGUCAGGGUUCUGGAUAUCUGUGUUUCAGUUGAAAGAAGUCACUCUUA AAGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACUGACU CUGUCAUCCGGGU (SEQ ID NO: 220); and UCAGGGUUCUGGAUAUCUGUGUUUUAGAGCUAGAAAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC UUUU (SEQ ID NO: 221) for CllpAs9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O-methylated bases and phosphorothioated bonds in first and last 3 positions.
[0104] Guide RNA sequences (spacer underlined) used for targeting the DPM2 gene were AGAAUCACCCAGGCGGUGUAGUGUUUCAAACCUUCGGCGUUUGAAAUCA CCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCGGUUA (SEQ ID NO: 222);AGAAUCACCCAGGCGGUGUAGUGUUUCAGUUGAAAGAAGUCACUCUUA AAGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACUGACU CUGUCAUCCGGGU (SEQ ID NO:223); and AAUCACCCAGGCGGUGUAGUGUUUUAGAGCUAGAAAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC UUUU (SEQ ID NO: 224) for CllpAs9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O-methylated bases and phosphorothioated bonds in first and last 3 positions.
[0105] Guide RNA sequences (spacer underlined) used for targeting the AAVS1 gene were UAGUGGCCCCACUGUGGGGUGGGUUUCAAACCUUCGGCGUUUGAAAUCA CCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCGGUUA (SEQ ID NO: 225);UAGUGGCCCCACUGUGGGGUGGGUUUCAGUUGAAAGAAGUCACUCUUA AAGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACUGACU CUGUCAUCCGGGU (SEQ ID NO: 226); and GUGGCCCCACUGUGGGGUGGGUUUUAGAGCUAGAAAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC UUUU (SEQ ID NO: 227) for CllpCas9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O-methylated bases and phosphorothioated bonds in first and last 3 positions.
[0106] Guide RNA sequences (spacer underlined) used for targeting the HEK4 gene were GUGGCACUGCGGCUGGAGGUGGGUUUCAAACCUUCGGCGUUUGAAAUC ACCACAAGGUGUCUUCAUAUCGGGGCAACUGCCUUUGGCACCCCCGGUU A (SEQ ID NO: 228);GUGGCACUGCGGCUGGAGGUGGGUUUCAGUUGAAAGAAGUCACUCUUA AAGUGAGCUGAAAUCACUAAAAAUUAAGAUUGAACCCGGCUACUGACUCUGUCAUCCGGGU (SEQ ID NO:229); and GGCACUGCGGCUGGAGGUGGGUUUUAGAGCUAGAAAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC UUUU (SEQ ID NO: 230) for CllpCas9s, ePsCas9 and SpCas9, respectively. All guides were synthesized as end protected by having 2'-O-methylated bases and phosphorothioated bonds in first and last 3 positions.
[0107] Post 72h incubation step, genomic DNA was isolated using the QuickExtract DNA Extraction Solution (Lucigen) as per manufacturer’s instructions in final volume of 50 pL.
[0108] Primary amplicons were generated using 0.25 pM target-specific primers (IDT), 1.5uL genomic DNA and the Phusion Flash High-Fidelity PCR Master Mix (Thermo) as per manufacturers guidelines. PCR reaction conditions used were 98°C for 1 min followed by 32 cycles of 98°C for 10 s, 60-65°C for 10 s, and 72°C for 10 s. PCR products were cleaned up with Ampure XP beads (Beckman) and analyzed using the Fragment Analyzer (Agilent). Indexing was performed by secondary PCR using 1 ng of the primary amplicon, 0.5 pM indexing primers (IDT) and KAPA HiFi HotStart Ready Mix (Roche) used as per manufacturer’s instructions. Secondary PCR reaction conditions used were 72°C for 3 min; 98°C for 30 s; followed by 10 cycles of 98°C for 10 s, 63°C for 30 s, and 72°C for 3 min; followed by the final cycle of 72°C for 5 min. Cleanup and quality control were done identically as for the primary amplification.Table 6 - Reference sequences of each gene. Amplicons of corresponding sites used for NGS analysis.
[0109] Sequence library quantification was performed by using the Qubit 4 Fluorometer (Thermo). For sequencing the Illumina NextSeq platform was employed, as per manufacturer’s guidelines. Data demultiplexing was performed by using bcl2fastq software. Editing efficiencies were obtained by applying the RIMA method as described in Taheri-Ghahfarokhi A., et al., Nucleic Acids Res 46(16):8417-8434 (2018), which is incorporated herein by reference in its entirety.
[0110] The editing results of the next-generation sequencing are shown in Figures 10, 19 and 20. These results demonstrate that CllpCas9 variants are very effective nucleases acrossseveral assayed targets, exhibiting activity levels at least comparable to established nucleases such as SpCas9 and PsCas9.
[0111] Off-targeting activity was investigated on a previously reported representative site as described in Tsai, et al., Nat Biotechnol 33, 187-197 (2015) (which is incorporated by reference herein in its entirety). From each edited HEK293T cell sample using a guide RNA and its appropriate counterpart Cas9 protein targeting the HEK4 site three amplicons were derived, one for the HEK4 on-target site and two for known off-target sites as described in Table 6. Editing efficiency on these sites was quantified by amplicon sequencing analysis as described above. Data presented in Figure 20 strongly suggests that CllpCas9 is a high-fidelity nuclease showing very low levels of off- targeting activity on the HEK4 site. The off-target profile of CllpCas9 variants tested here is more similar to the high-fidelity PsCas9, rather than the more promiscuous SpCas9.Table 7: Nuclear Localization Sequences - amino acid residuesTable 8: Overview of the sequence listing
[0112] All references cited herein, including patents, patent applications, papers, textbooks and the like, and the references cited therein, to the extent that they are not already, are hereby incorporated herein by reference in their entirety.
[0113] It should it be understood that, in general, where the disclosure, or aspects of the disclosure, is / are referred to as comprising particular elements and / or features, certain embodiments of the disclosure or aspects of the disclosure consist, or consist essentially of, such elements and / or features. For purposes of simplicity, those embodiments have not been specifically set forth in haec verba herein.
[0114] All patents and publications mentioned in this specification are herein incorporated by reference to the same extent as if each independent patent and publication wasspecifically and individually indicated to be incorporated by reference. Citation or identification of any reference in any section of this application shall not be construed as an admission that such reference is available as prior art to the present invention.
Claims
CLAIMS1. A Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises an amino acid modification at one or more positions relative to SEQ ID NO: 1 : 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229.
2. The Cas9 protein according to claim 1, wherein said Cas9 protein comprises an enhanced endonuclease activity compared to a wild type Cas9 protein, such as a wild type CllCas9 protein, and optionally wherein said endonuclease activity is measured by a gene editing reporter assay.
3. The Cas9 protein according to any one of the preceding claims, wherein the polypeptide sequence comprises one or more amino acid modifications selected from D1073A, A1160R, D1073G, D1091G, D1091N, D1091R, D1091 S, D1122R, D1206G, D1206N, D1206R, D1206S, D37G, D37N, D37R, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D91 IK, D911R, D91 IS, D912G, D912N, D912R, D912S, El 164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, G39R, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, Q556R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
4. The Cas9 protein according to any one of the preceding claims, wherein the polypeptide sequence comprises one or more amino acid modifications selected from D1073A, A1160R, D1073G, D1091G, D1091N, D1091R, D1091 S, D1122R, D1206G, D1206N, D1206R, D1206S, D37G, D37N, D37R, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D91 IS, D912G, D912N, D912R, D912S, El 164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, G39R, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, Q556R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
5. The Cas9 protein according to any one of the preceding claims, wherein the polypeptide sequence comprises a first amino acid modification selected from D37R, D1091R, D912R, G39R, I978R, N1161R, and S1229R.
6. The Cas9 protein according to claim 5, wherein the polypeptide sequence further comprises a second amino acid modification selected from D1091R, D912R, G39R, N1161R or S1229R., optionally wherein:(i) the first amino acid modification is D37R and the second amino acid modification is selected from D1091R, N1161R, D912R, S1229R or G39R: or(ii) wherein the first amino acid modification is D1091R and the second amino acid modification is selected from N1161R, S1229R, and D912R.
7. The Cas9 protein according to claim 6, further comprising a third amino acid modification selected from A1160R, D1073A, D1073G, D1091G, D1091N, D1091R, D1091S, D1122R, D1206G, D1206R, D1206S, D37G, D37N, D47G, D47L, D47R, D59G, D59R, D59S, D61G, D61L, D61R, D911G, D911S, D912G, D912N, D912R, D912S, E1164R, E1226G, E1226R, E1226S, E508G, E508R, F94G, F94Y, I46A, I46R, I978R, K980R, N1161R, N501F, N501R, N501W, N501Y, Q315K, Q315R, S1229R, S566K, S566R, S57A, S57L, S57R, and T54R.
8. The Cas9 protein according to claim 7, further comprising a fourth amino acid modification selected from D1206S, D59S, D912N, D912R, D912S, E1164R, E1226G, E1226S, F94Y, I46R, Q556R, S1229R, S566R, and S57R, optionally a fifth amino acid modification selected from D1206N, D1206S, D912N, D912R, D912S, G39R, I46R, and S1229R, optionally a sixth amino acid modification selected from S1229R, S566R, and optionally a seventh amino acid modification at position D912N.
9. The Cas9 protein according to any one of the preceding claims, wherein said Cas9 protein comprises one or more mutations as defined in Table 2, or wherein said Cas9 protein comprises or consists of an amino acid sequence as defined in any one of SEQ ID NOs: 9-144, SEQ ID NOs: 14-144, SEQ ID NOs: 30-144, SEQ ID NOs: 60-144, SEQ ID NOs: 112-144, and / or SEQ ID NOs: 140-144.
10. The Cas9 protein according to any one of the preceding claims, wherein said Cas9 protein further comprises at least one nuclear localization signal (NLS), such as VC, 53BP1, bpSV40, BRCAl l, BRCA1 2, EWS, FUS, HIV-I TAT, hnRNP-D, hnRNP-M, HsNPM2, HTLV-1, mpSV40, MYC, NLP, NPM2, PTHrP and / or REV HV1H2, or wherein the Cas9 protein further comprises one or more NLS and is as set forth in any one or more of SEQ ID NOs: 161-194, SEQ ID NOs: 167-194, SEQ ID NOs: 178-194, SEQ ID NOs: 182-194, SEQ ID NOs: 189-194 or SEQ ID NOs: 192-194.
11. A CRISPR-Cas system comprising: a) a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises an amino acid modification at one or more positions relative to SEQ ID NO: 1 : 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229, or a Cas9 protein as defined in any one of the preceding claims; and b) a guide polynucleotide comprising a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell.
12. A CRISPR-Cas system comprising: a) a nucleic acid sequence encoding a Cas9 protein comprising a polypeptide sequence having at least 95% identity, such as 97%, 98% or 99%, to SEQ ID NO: 1, wherein said polypeptide sequence comprises an amino acid modification at one or more positions relative to SEQ ID NO: 1 : 37, 39, 46, 47, 54, 57, 59, 61, 94, 315, 501, 508, 556, 566, 911, 912, 978, 980, 1073, 1091, 1122, 1160, 1161, 1164, 1206, 1226 and / or 1229, or a Cas9 protein as defined in any one of claims 1-10 ; and b) a nucleic acid sequence encoding a guide polynucleotide that comprises a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in a eukaryotic cell.
13. The CRISPR-Cas system according to claim 11 or 12, wherein the guide polynucleotide comprises or consists of a polynucleotide sequence as defined in any one of SEQ ID NO: 195- 209.
14. A eukaryotic cell or a delivery particle comprising the Cas9 protein as defined in any one of claims 1 to 10 or the CRISPR-Cas system as defined in any one of claims 11-13.
15. A method for providing site-specific modification of a target sequence in a eukaryotic cell, the method comprising: a) introducing into the eukaryotic cell a nucleotide sequence encoding a guide polynucleotide that forms a complex with a Cas9 protein, wherein said guide polynucleotide comprises a guide sequence, wherein the guide sequence is capable of hybridizing with a target sequence in host polynucleotide; and wherein the eukaryotic cell comprises the Cas9 protein as defined in any one of claims 1-10, b) generating cohesive ends in the host polynucleotide with the Cas9 protein and the guide polynucleotide; and c) ligating i) the cohesive ends of (b) together, or ii) a 3' end of a polynucleotide sequence of interest to a first cohesive end of the host polynucleotide, and a 5’ end of the polynucleotide sequence of interest to a second cohesive end of the host polynucleotide; thereby modifying the target sequence.
16. The method according to claim 15, wherein step a) further comprises i) introducing into the eukaryotic cell a nucleotide sequence encoding the Cas9 protein as defined in any one of claims 1-10, and ii) expressing the Cas9 protein by the eukaryotic cell.