Nucleobase editors having reduced non-target deamination and assays for characterizing nucleobase editors
Patent Information
- Application Number
- JP2025034660
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-31
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-31
AI Technical Summary
Existing nucleobase editors, such as CRISPR-Cas proteins, suffer from spurious deamination and off-target editing, leading to unwanted mutations and bystander mutations in the genome, which are not effectively addressed by current technologies.
A fusion protein is developed with a deaminase inserted into a flexible loop of the Cas9 polypeptide, structured as NH2 - [N-terminal fragment of Cas9] - [Deaminase] - [C-terminal fragment of Cas9] - COOH, designed to reduce off-target deamination by positioning the deaminase proximal to the target nucleobase, thereby minimizing unwanted editing.
The fusion protein effectively reduces off-target deamination events, ensuring precise and targeted nucleobase editing with lower instances of spurious mutations, improving the accuracy and specificity of genetic modifications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background Art
[0001] Cross-reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 799,702, filed on Jan. 31, 2019 and its content is incorporated herein by reference in its entirety.
[0002] Deaminases combined with the precise targeting of CRISPR-Cas proteins are called nucleobase editors and have the ability to introduce site-specific point mutations into target polynucleotides . Nucleobase editors induce base changes without introducing double-stranded DNA breaks and include adenosine base editors that convert target A·T to G·C and cytidine base editors that convert target C·G to T·A. However, the introduction of nucleobase editors into cells can result in unwanted base editor-related editing, including spurious deamination, bystander mutations, and off-target editing throughout the genome . Spurious deamination events can occur throughout the genome and are catalyzed by the base editor deamination domain acting independently of target-directed base editing via the main programming of CRISPR-Cas by guide RNA. Without being bound by theory, spurious deamination events throughout the genome may occur where single-stranded DNA substrates are formed, such as by "DNA breathing" or at DNA replication forks . Off-target editing occurs outside of the on-target sequence but is a base editing event within approximately 200 bp upstream or downstream of the target region. Bystander mutations are guided by Cas9 / sgRNA Occurs within the edited base editing window of the target, but is not the desired target nucleic acid base Is a mutation. Bystander mutations can result in either silent mutations (no amino acid change) or non-synonymous Mutations (amino acid changes). Therefore, there is a need for a base editor with reduced off-target deamination SUMMARY OF THE INVENTION
[0003] As described below, the present invention provides compositions and methods of nucleic acid base editors, as well as An assay for characterizing nucleic acid base editors as having reduced off-target Deamination (e.g., as compared to programmed on-target deamination).
[0004] The compositions and articles defined by the present invention are isolated or otherwise produced in connection with the examples provided below Other features and advantages of the present invention will be apparent from the detailed description and Claims
[0005] In one aspect, provided herein is a fusion protein comprising a deaminase inserted into a flexible loop of a Cas9 polypeptide, wherein the fusion protein Has the following structure: NH2 - [N-terminal fragment of Cas9] - [Deaminase] - [C-terminal fragment of Cas9] - COOH Wherein each instance of "] - [" is an optional linker
[0006] In one aspect, provided herein is a fusion protein comprising a deaminase adjacent to an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide, wherein the C-terminus of the N-terminal fragment And The N-terminal fragment of the C-terminus contains a part of the flexible loop of the Cas9 polypeptide.
[0007] In some embodiments, the deaminase of the fusion protein deaminates the target nucleobase in the target polynucleotide sequence. In some embodiments, the flexible loop contains amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. In some embodiments, the flexible loop contains a part of the α-helix structure of the Cas9 polypeptide. In some embodiments, the target nucleobase is deaminated with lower off-target deamination compared to an end-terminal fusion protein containing a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the target nucleobase is deaminated with lower off-target deamination compared to an end-terminal fusion protein containing a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the target nucleobase is deaminated with lower off-target deamination compared to an end-terminal fusion protein containing a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the target nucleobase is deaminated with lower off-target deamination compared to an end-terminal fusion protein containing a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1.
[0008] In some embodiments, the target nucleobase is 1 to 20 nucleobases away from the protospacer adjacent motif (PAM) sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is 1 to 20 nucleobases away from the protospacer adjacent motif (PAM) sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is upstream of the PAM sequence by 2 to 12 nucleobases. In some embodiments, the flexible loop contains a region selected from the group consisting of amino acid residues at positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in SEQ ID NO: 1, or a corresponding region. In some embodiments, the flexible loop contains a region selected from the group consisting of amino acid residues at positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in SEQ ID NO: 1, or a corresponding region. In some embodiments, the flexible loop contains a region selected from the group consisting of amino acid residues at positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in SEQ ID NO: 1, or a corresponding region. In some embodiments, the deaminase is at amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, in SEQ ID NO: 1, numbered as such. In some embodiments, the deaminase is at amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, in SEQ ID NO: 1, numbered as such. Between 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or inserted between 1248-1249, or between corresponding amino acid positions.
[0009] In some embodiments, the deaminase is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1040-1041, 1068-1069, or 1247-1248 in the numbering of SEQ ID NO: 1, or between corresponding amino acid positions. In some embodiments, the deaminase is inserted between amino acid positions 1016-1017, 1023-1024, 1029-1030, 1040-1041, 1069-1070 or 1247-1248 in the numbering of SEQ ID NO: 1, or between corresponding amino acid positions.
[0010] In some embodiments, the N-terminal fragment comprises amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, and / or 1248-1297 of the Cas9 polypeptide in the numbering of SEQ ID NO: 1, or corresponding residues. In some embodiments, the C- terminal fragment comprises amino acid residues 1301-1368, 1248-1297, 1078-1231, 1026-1051, 948-1001, 692-942, 580-685, and / or 538-568 of the Cas9 polypeptide in the numbering of SEQ ID NO: 1, or corresponding residues. In some embodiments, the N-terminal fragment or C-terminal fragment of the Cas9 polypeptide binds to a target polynucleotide sequence.
[0011] In some embodiments, the N-terminal fragment or C-terminal fragment of the Cas9 polypeptide binds to DNA contains a domain. In certain embodiments, the N-terminal fragment or C-terminal fragment contains the RuvC domain In some embodiments, the N-terminal fragment or C-terminal fragment contains the HNH domain In some embodiments, neither the N-terminal fragment nor the C-terminal fragment contains the HNH domain In some embodiments, neither the N-terminal fragment nor the C-terminal fragment contains the RuvC domain In certain embodiments, the Cas9 polypeptide contains partial or complete deletions in one or more structural domains In certain embodiments, the deaminase is inserted at the position of the partial or complete deletion of the Cas9 polypeptide
[0012] In some embodiments, the deletion is within the RuvC domain. In some embodiments the deletion is within the HNH domain. In some embodiments, the deletion bridges the RuvC domain and the C-terminal domain, the L-I domain and the HNH domain, or the RuvC domain and the L-I domain In some embodiments, the Cas9 polypeptide contains a deletion of amino acids 1017-1069 or corresponding amino acids numbered in SEQ ID NO: 1 In some embodiments, the Cas9 polypeptide contains a deletion of amino acids 792-872 or corresponding amino acids numbered in SEQ ID NO: 1 In some embodiments, the Cas9 polypeptide contains a deletion of amino acids 792-906 or corresponding amino acids numbered in SEQ ID NO: 1 and SEQ ID NO: 1 In some embodiments, the Cas9 polypeptide contains a deletion of amino acids 792-906 or corresponding amino acids numbered in SEQ ID NO: 1
[0013] In one aspect, a fusion protein containing a deaminase inserted into the Cas9 polypeptide A protein is provided herein, wherein the fusion protein has the following structure: NH2 - [N-terminal fragment of Cas9] - [Deaminase] - [C-terminal fragment of Cas9] - COOH wherein each instance of "] - [" is an optional linker, the Cas9 polypeptide comprises a complete deletion of the HNH domain and the deaminase is inserted at the deletion position.
[0014] In some embodiments, the C-terminal amino acid of the N-terminal fragment is amino acid 791 numbered as in SEQ ID NO: 1. In some embodiments, the N-terminal amino acid of the C-terminal fragment is amino acid 907 numbered as in SEQ ID NO: 1. In some embodiments, the N-terminal amino acid of the C-terminal fragment is amino acid 873 numbered as in SEQ ID NO: 1.
[0015] In one aspect provided herein, a fusion protein comprising a deaminase inserted within a Cas9 polypeptide is provided, wherein the fusion protein has the following structure: NH2 - [N-terminal fragment of Cas9] - [Deaminase] - [C-terminal fragment of Cas9] - COOH wherein each instance of "] - [" is an optional linker, Cas9 comprises a complete deletion of the RuvC domain and the deaminase is inserted at the deletion position.
[0016] In certain embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase. In certain embodiments, the cytidine deaminase is an APOBEC cytidine deaminase, activation induced cytidine deaminase (AID), or CDA In certain embodiments, the APOBEC deaminase is APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, or APO EC4. In some embodiments, the APOBEC deaminase is rAPOBEC1. In a certain embodiment, any one of the fusion proteins of the above aspects further comprises a UGI domain.
[0017] In certain embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a modified TadA. In some embodiments, the TadA deaminase is TadA7.10. In a certain embodiment, the adenosine deaminase is a TadA dimer. In some embodiments, the TadA dimer comprises TadA7.10 and wild-type TadA. In some embodiments, any linker comprises (SGGS)n, (GGGS)n, (GGGGS)n, (G)n, (EAAAK)n, (GGS)n, SGSETPGTSESATPES, or (XP)n motif or a combination thereof, where n is independently an integer between 1 and 30.
[0018] In a certain embodiment, the N-terminal fragment of the Cas9 polypeptide is fused to the deaminase without a linker. In some embodiments, the C-terminal fragment of Cas9 is fused to the deaminase without a linker. In a certain embodiment, any one of the fusion proteins of the above aspects further comprises an additional catalytic domain.
[0019] In a certain embodiment, the additional catalytic domain is a second deaminase. In a certain embodiment, In a form, the second deaminase is fused to the N-terminus or C-terminus of the fusion protein . In certain embodiments, the deaminase is a cytidine deaminase or an adenosine de aminase. In certain embodiments, the fusion protein of any one of the above aspects is further comprising a nuclear localization signal. In certain embodiments, the nuclear localization signal is a bipartite ( bipartite) nuclear localization signal. In certain embodiments, the Cas9 polypeptide is Str eptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Strept ococcus thermophilus 1 Cas9 (St1Cas9), or a variant thereof. In certain embodiments , the Cas9 polypeptide is a modified Cas9 and has specificity for a modified PAM . In certain embodiments, the Cas9 polypeptide is a nickase. In certain embodiments , the Cas9 polypeptide is nuclease-inactive. In some embodiments , the fusion protein of any one of the above aspects forms a complex with a guide nucleic acid sequence and results in deamination of the target nucleobase. In certain embodiments, the fusion protein is further complexed with a target polynucleotide.
[0020] As used herein, a polynucleotide encoding the fusion protein of any one of the above aspects is provided .
[0021] As used herein, an expression vector comprising the above polynucleotide is provided
[0022] In certain embodiments, the expression vector is a mammalian expression vector. In certain embodiments In this regard, the vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenovirus vector, lentiviral vector, Sendai virus vector, and herpes virus vector. In certain embodiments the vector comprises a promoter.
[0023] As used herein, cells are provided that comprise any one of the fusion proteins of the above aspects, the polynucleotide as described above, or the vector as described above.
[0024] In certain embodiments, the cell is a bacterial cell, a plant cell, an insect cell, a human cell, or a mammalian cell.
[0025] As used herein, kits are provided that comprise any one of the fusion proteins of the above aspects, the polynucleotide as described above, or the vector as described above.
[0026] As used herein, methods for base editing are provided that comprise contacting a polynucleotide sequence with a fusion protein of any one of the above aspects, wherein the deaminase of the fusion protein deaminates a nucleic acid base in the polynucleotide, thereby editing the polynucleotide sequence.
[0027] In some embodiments, the method further comprises contacting a target polynucleotide sequence with a guide nucleic acid sequence to effect deamination of the target nucleic acid base.
[0028] In one aspect, methods are provided herein for editing a target nucleic acid base in a target polynucleotide sequence, the method comprising contacting the target polynucleotide sequence with a Cas9 poly A fusion protein comprising a deaminase adjacent to an N-terminal fragment and a C-terminal fragment of a peptide, and contacting, wherein the deaminase of the fusion protein deaminates a target nucleobase in a target polynucleotide sequence, and the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a part of a flexible loop of the Cas9 polypeptide.
[0029] Disclosed herein is a method for editing a target nucleobase in a target polynucleotide sequence, the method comprising contacting the target polynucleotide sequence with a fusion protein comprising a deaminase inserted within a flexible loop of a Cas9 polypeptide, wherein the fusion protein has the structure NH2-[N-terminal fragment of Cas9]-[deaminase]-[C-terminal fragment of Cas9]-COOH, where each instance of “]-[” is an optional linker, and the deaminase of the fusion protein deaminates a target nucleobase in the target polynucleotide sequence. NH2-[N-terminal fragment of Cas9]-[deaminase]-[C-terminal fragment of Cas9]-COOH where each instance of “]-[” is an optional linker, and the deaminase of the fusion protein deaminates a target nucleobase in the target polynucleotide sequence. In some embodiments, the method further comprises contacting the target polynucleotide sequence with a guide nucleic acid sequence to effect deamination of the target nucleobase. In some embodiments, the guide nucleic acid sequence comprises a spacer sequence complementary to a protospacer sequence of the target polynucleotide sequence, thereby forming an R-loop. In some embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase
[0030] In some embodiments, the method further comprises contacting the target polynucleotide sequence with a guide nucleic acid sequence to effect deamination of the target nucleobase. In some embodiments, the guide nucleic acid sequence comprises a spacer sequence complementary to a protospacer sequence of the target polynucleotide sequence, thereby forming an R-loop. In some embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase is deaminated with lower off-target deamination as compared to an end-terminal method that involves a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the deaminase of the fusion protein does not deaminate more than two nucleobases within the extent of the R-loop. In certain embodiments, the target nucleobase and is separated from the PAM sequence in the target polynucleotide sequence by 1 to 20 nucleobases. In some embodiments, the target nucleobase is upstream of the PAM sequence by 2 to 12 nucleobases.
[0031] In some embodiments, the flexible loop contains amino acids that are proximal to the target nucleobase when the deaminase of the fusion protein deaminates the target nucleobase. In some embodiments, the flexible loop contains a region selected from the group consisting of amino acid residues at positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in the numbering of SEQ ID NO: 1, or a corresponding region. In some embodiments, the deaminase is inserted between amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249 in the numbering of SEQ ID NO: 1, or between corresponding amino acid positions. In some embodiments, the deaminase is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1040-1041, 1068-1069, or 1247-1248 in the numbering of SEQ ID NO: 1, or between corresponding amino acid positions. In some embodiments, the deaminase is inserted between amino acid positions 1016-1017, 1023-1024, 1029-1030, 1040-1041, 1069- 1070 in the numbering of SEQ ID NO: 1, or between corresponding amino acid positions. It is inserted between 1070 and 1247 - 1248, or between the corresponding amino acid positions.
[0032] In some embodiments, the N - terminal fragment comprises amino acid residues 1 - 529, 538 - 568, 580 - 685, 692 - 942, 948 - 1001, 1026 - 1051, 1078 - 1231, and / or 1248 - 1297 of the Cas9 polypeptide as numbered in SEQ ID NO: 1, or the corresponding residues. In some embodiments the C - terminal fragment comprises amino acid residues 1301 - 1368, 1248 - 1297, 1078 - 1231, 1026 - 1051, 948 - 1001, 692 - 942, 580 - 685, and / or 538 - 568 of the Cas9 polypeptide as numbered in SEQ ID NO: 1, or the corresponding residues. In some embodiments, the N - terminal fragment or the C - terminal fragment of the Cas9 polypeptide binds to a target polynucleotide sequence. In certain embodiments, the N - terminal fragment or the C - terminal fragment comprises an RuvC domain. In some embodiments, the N - terminal fragment or the C - terminal fragment comprises an HNH domain. In some embodiments neither the N - terminal fragment nor the C - terminal fragment comprises an HNH domain. In some embodiments neither the N - terminal fragment nor the C - terminal fragment comprises an RuvC domain. In certain embodiments, the Cas9 polypeptide comprises a partial or complete deletion in one or more structural domains. In certain embodiments, the deaminase is inserted at the position of the partial or complete deletion of the Cas9 polypeptide.
[0033] In certain embodiments, the deletion is within the RuvC domain. In some embodiments, the deletion is within the HNH domain. In certain embodiments, the deletion is within the HNH domain. In some embodiments, the deletion is within the RuvC domain. In some embodiments, the deletion is within the HNH domain. Yes. In some embodiments, the deletion cross-links the RuvC domain and the C-terminal domain, the L-I domain and the HNH domain, or the RuvC domain and the L-I domain. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 1017-1069 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-872 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-906 as numbered in SEQ ID NO: 1 or corresponding amino acids. In one embodiment, the deaminase is a cytidine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In one embodiment, the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM). In one embodiment, the Cas9 polypeptide is a nickase. In one embodiment, the Cas9 polypeptide is nuclease-inactive. In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell.
[0034] In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-872 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-906 as numbered in SEQ ID NO: 1 or corresponding amino acids. In one embodiment, the deaminase is a cytidine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In one embodiment, the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM). In one embodiment, the Cas9 polypeptide is a nickase. In one embodiment, the Cas9 polypeptide is nuclease-inactive. In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 1017-1069 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-872 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-906 as numbered in SEQ ID NO: 1 or corresponding amino acids. In one embodiment, the deaminase is a cytidine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In one embodiment, the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM). In one embodiment, the Cas9 polypeptide is a nickase. In one embodiment, the Cas9 polypeptide is nuclease-inactive. In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 1017-1069 as numbered in SEQ ID NO: 1 or corresponding amino acids.
[0035] In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-872 as numbered in SEQ ID NO: 1 or corresponding amino acids. In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids 792-906 as numbered in SEQ ID NO: 1 or corresponding amino acids. In one embodiment, the deaminase is a cytidine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In one embodiment, the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM). In one embodiment, the Cas9 polypeptide is a nickase. In one embodiment, the Cas9 polypeptide is nuclease-inactive. In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell. In some embodiments, contacting is performed intracellularly. In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment, the cell is a pluripotent cell. In one embodiment, the cell is in vivo or ex vivo. In some embodiments, contacting is performed in a population of cells. In one embodiment, the population of cells is a mammalian cell or a human cell.
[0036] In one aspect, a method for treating a genetic condition in a subject is provided herein, the method comprising administering to the subject a fusion protein comprising a deaminase flanked by an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide, or a polynucleotide encoding the fusion protein, and a guide nucleic acid sequence or a polynucleotide encoding the guide nucleic acid sequence, wherein the guide nucleic acid sequence causes the deamination of a target nucleobase in the target polynucleotide sequence of the subject, thereby treating the genetic condition. A method for treating a genetic condition in a subject is provided herein, the method comprising administering to the subject a fusion protein comprising a deaminase inserted within a flexible loop of a Cas9 polypeptide, wherein the fusion protein has the following structure: NH2-[N-terminal fragment of Cas9]-[deaminase]-[C-terminal fragment of Cas9]-COOH wherein each instance of “]-[” is an optional linker, and the deaminase of the fusion protein deaminates a target nucleobase in the target polynucleotide sequence of the subject, thereby treating the genetic condition. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a part of the flexible loop of the Cas9 polypeptide. In some embodiments, the method further comprises administering a guide nucleic acid sequence to the subject to effect deamination of the target nucleobase. In certain embodiments, the target nucleobase comprises a mutation associated with the genetic condition. In certain embodiments, deamination of the target nucleobase substitutes the target nucleobase with a wild-type nucleobase. In some embodiments,
[0037]
[0038] In some embodiments, deamination of the target nucleobase replaces the target nucleobase with a non-wild-type nucleobase, and this deamination of the target nucleobase ameliorates the symptoms of the genetic condition.
[0039] In some embodiments, the target polynucleotide sequence contains a mutation associated with the genetic condition at a nucleobase other than the target nucleobase. In certain embodiments, deamination of the target nucleobase ameliorates the symptoms of the genetic condition. In certain embodiments, the target nucleobase is 1 to 20 nucleobases away from the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is 2 to 12 nucleobases upstream of the PAM sequence. In some embodiments, the flexible loop contains an amino acid that is proximal to the target nucleobase when the deaminase of the fusion protein deaminates the target nucleobase.
[0040] In some embodiments, the flexible loop contains a region selected from the group consisting of amino acid residues at numbering positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in SEQ ID NO: 1, or a corresponding region.
[0041] In some embodiments, the deaminase is at amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249 in SEQ ID NO: 1. is inserted between, or corresponding amino acid positions thereof. In some embodiments the deaminase is inserted between amino acid positions 768 - 769, 792 - 793, 102 2 - 1023, 1026 - 1027, 1040 - 1041, 1068 - 1069, or 1247 - 1248 in the numbering of SEQ ID NO: 1, or corresponding amino acid positions thereof. In some embodiments, the deaminase is inserted between amino acid positions 1016 - 1017, 1023 - 1024, 1029 - 1030, 1040 - 1041 , 1069 - 1070 or 1247 - 1248 in the numbering of SEQ ID NO: 1, or corresponding amino acid positions thereof. is inserted between, or corresponding amino acid positions thereof.
[0042] In some embodiments, the N - terminal fragment comprises amino acid residues 1 - 529, 538 - 568, 580 - 685, 692 - 942, 948 - 1001, 1026 - 1051, 1078 - 1231, and / or 1248 - 1297 of the Cas9 polypeptide in the numbering of SEQ ID NO: 1, or corresponding residues. In some embodiments the C - terminal fragment comprises amino acid residues 1301 - 1368, 1248 - 1297, 1078 - 1231, 1026 - 1051, 948 - 1001, 692 - 942, 580 - 685, and / or 538 - 568 of the Cas9 polypeptide in the numbering of SEQ ID NO: 1, or corresponding residues. In some embodiments, the N - terminal fragment or C - terminal fragment of the Cas 9 polypeptide binds to a target polynucleotide sequence. In certain embodiments, the N - terminal fragment or C - terminal fragment comprises an RuvC domain. In some embodiments, the N - terminal fragment or C - terminal fragment comprises an HNH domain.
[0043] In some embodiments, both the N-terminal fragment and the C-terminal fragment contain an HNH domain. In some embodiments, neither the N-terminal fragment nor the C-terminal fragment contains a RuvC domain. In one embodiment, the Cas9 polypeptide does not include one or more structural domains. In one embodiment, the deaminase comprises a partial or complete deletion in the domain. In some embodiments, the Cas9 polypeptide is inserted into the region of the partial or complete deletion. In some embodiments, the deletion is in the RuvC domain. In some embodiments, the deletion is in the RuvC domain and the C-terminal domain. , bridging the LI domain and the HNH domain, or the RuvC domain and the LI domain. In some embodiments, the Cas9 polypeptide comprises amino acid 101, as numbered in SEQ ID NO:1. Contains a deletion of amino acids 7 to 1069 or the corresponding amino acids.
[0044] In some embodiments, the Cas9 polypeptide is represented by the numbering in SEQ ID NO:1. In some embodiments, the deletion comprises amino acids 792 to 872 or the corresponding amino acids. The Cas9 polypeptide comprises amino acids 792 to 906 or therein, as numbered in SEQ ID NO:1. In one embodiment, the deaminase is a cytidine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In one embodiment, the Cas9 polypeptide is a modified Cas9 and a modified PA. In one embodiment, the Cas9 polypeptide is a nickase. In some embodiments, the Cas9 polypeptide is nuclease inactive. In embodiments, the subject is a mammal. In certain embodiments, the subject is a human.
[0045] A protein library for optimized base editing, comprising a plurality of fusion proteins is provided herein, wherein each of the plurality of fusion proteins comprises a deaminase flanked by an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide and the N-terminal fragment of each fusion protein is different from the N-terminal fragment of the remainder of the plurality of fusion proteins, or the C-terminal fragment of each fusion protein is different from the C-terminal fragment of the remainder of the plurality of fusion proteins, and the deaminase of each fusion protein deaminates a target nucleobase proximal to a protospacer adjacent motif (PAM) sequence in a target polynucleotide sequence, and the aforementioned N-terminal fragment or the aforementioned C-terminal fragment binds to the target polynucleotide sequence. In some embodiments, for each nucleobase 1 to 20 nucleobases away from the PAM sequence, at least one of the plurality of fusion proteins deaminates the nucleobase. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment of the Cas9 polypeptide of each of the plurality of fusion proteins comprises a part of a flexible loop of the Cas9 polypeptide. In some embodiments, at least one of the plurality of fusion proteins deaminates a target nucleobase with lower off-target deamination compared to an end-terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, at least one of the plurality of fusion proteins is 2 to 12 nucleotides of the PAM sequence In embodiments, the mammalian subject is a human subject.
[0046] In some embodiments, for each nucleobase 1 to 20 nucleobases away from the PAM sequence, at least one of the plurality of fusion proteins deaminates the nucleobase. In some embodiments, at least one of the plurality of fusion proteins deaminates a target nucleobase with lower off-target deamination compared to an end-terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment of the Cas9 polypeptide of each of the plurality of fusion proteins comprises a part of a flexible loop of the Cas9 polypeptide. In some embodiments, at least one of the plurality of fusion proteins is 2 to 12 nucleotides of the PAM sequence In some embodiments, the mammalian subject is a human subject. In some embodiments, for each nucleobase 1 to 20 nucleobases away from the PAM sequence, at least one of the plurality of fusion proteins deaminates the nucleobase. In some embodiments, at least one of the plurality of fusion proteins deaminates a target nucleobase with lower off-target deamination compared to an end-terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of SEQ ID NO: 1. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment of the Cas9 polypeptide of each of the plurality of fusion proteins comprises a part of a flexible loop of the Cas9 polypeptide. In some embodiments, at least one of the plurality of fusion proteins is 2 to 12 nucleotides of the PAM sequence Deaminates the target nucleobase upstream of the base. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment of the plurality of fusion proteins contains amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment of the plurality of fusion proteins contains amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. contains amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase.
[0047] In some embodiments, at least one deaminase of the plurality of fusion proteins is between amino acid positions 768 - 769, 791 - 792, 792 - 793, 1015 - 1016, 1022 - 1023, 1026 - 1027, 1029 - 1030, 1040 - 1041, 1052 - 1053, 1054 - 1055, 1067 - 1068, 1 068 - 1069, 1247 - 1248, or 1248 - 1249 or corresponding amino acid positions as numbered in SEQ ID NO: 1. In some embodiments, at least one deaminase of the fusion protein is between amino acid positions 768 - 769, 792 - 793, 1022 - 1023, 1026 - 1027, 104 0 - 1041, 1068 - 1069, or 1247 - 1248 or corresponding amino acid positions as numbered in SEQ ID NO: 1. In some embodiments, at least one deaminase of the fusion protein is between amino acid positions 1016 - 1017, 1023 - 1024, 1029 - 1030, 1040 - 1041, 1 069 - 1070 or 1247 - 1248 or corresponding amino acid positions as numbered in SEQ ID NO: 1. In certain embodiments, the deaminase is adenosine deaminase. In certain embodiments, the deaminase is cytidine deaminase. the deaminase is cytidine deaminase.
[0048] In certain embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9) 、Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1C as9), or variants thereof. In certain embodiments, the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM). In certain aspects, the Cas9 polypeptide is a nickase. In certain aspects, the Cas9 polypeptide is nuclease-inactive.
[0049] [Definitions] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with many general definitions of terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994 ); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Gl ossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings set forth below unless otherwise specified.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with many general definitions of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994 ); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Gl ossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meaning as defined below unless otherwise specified. "Adenosine deaminase" means a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In certain embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In certain embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered adenosine deaminases)
[0051] Nase (evolved adenosine deaminase) may be derived from any organism such as bacteria. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism. In some embodiments the deaminase or deaminase domain is not naturally occurring. For example in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90 %, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is derived from bacteria such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. cre scentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA (ecTa dA) deaminase or a fragment thereof. For example, truncated ecTadA may lack one or more N-terminal amino acids compared to full-length ecTadA. In some embodiments, truncated ecTadA lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues compared to full-length ecTadA. In some embodiments, the truncated
[0052] For example, truncated ecTadA may lack one or more N-terminal amino acids compared to full-length ecTadA. In some embodiments, truncated ecTadA lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues compared to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues. In some embodiments, the truncated The truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 4, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. In some embodiments, the TadA deaminase is TadA with a truncated N-terminus. In certain embodiments, TadA is any of the TadAs described in PCT / US2017 / 045381 (which is hereby incorporated by reference in its entirety).
[0053] In certain embodiments, the adenosine deaminase comprises the following amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPT AHAEIMALRQGGLVMQNYRLIDAT LYVTLEPCVMCAGAMIHSRIGRVVFGARDAKT GAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKA QKKAQSSTD This is referred to as the "TadA reference sequence".
[0054] In certain embodiments, the TadA deaminase is the full-length E. coli TadA deaminase. For example, in certain embodiments, the adenosine deaminase comprises the following amino acid sequence: including. MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VMQNYRLIDATLYVTLEPCVMCAGAMIHSRIG RVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSD FFRMRRQEI KAQKKAQSSTD
[0055] However, further adenosine deaminases useful in the present application will be apparent to those skilled in the art and are intended to be within the scope of the present disclosure. For example, the adenosine deaminase can be a homolog of adenosine deaminase acting on tRNA (ADAT). Exemplary ADAT homologs include, but are not limited to, the following. For example, the adenosine deaminase can be a homolog of adenosine deaminase acting on tRNA (ADAT). For example, the adenosine deaminase can be a homolog of adenosine deaminase acting on tRNA (ADAT). Exemplary ADAT homologs include, but are not limited to, the following. For example, the adenosine deaminase can be a homolog of adenosine deaminase acting on tRNA (ADAT). Exemplary ADAT homologs include, but are not limited to, the following.
[0056] Staphylococcus aureus TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCS GS LMNLLQQS NFNHRAIVDKG VLKE AC S TLLTTFFKNL RANKKS TN
[0057] Bacillus subtilis TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGC S GTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAAR KNLSE
[0058] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV
[0059] Shewanella putrefaciens (S. putrefaciens) TadA: MDE YWMQVAMQM AEKAEAAGE VPVGA VLVKDGQQIATGYNLS IS QHDPT AHAEI LCLRSAGKKLENYRLLDA TLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKK ALKLAQRAQQGIE
[0060] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQS DPT ΑΗ AEIIALRNG AKNI QN YRLLNS TLY VTLEPCTMC AG AILHS RIKRLVFG AS D YK TGAIGSRFHFFDDYKMNHTLEITSGVLAEE CSQKLSTFFQKRREEKKIEKALLKSLSD K
[0061] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAAKLGNYRL TDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRK AKI
[0062] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP
[0063] TadA7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD
[0064] Exemplary sequences containing TadA7.10 or a TadA7.10 variant include the following. are as follows. GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVG AVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVE ITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD
[0065] TadA7.10 CP65 TAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVF GVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQ VFNAQKKAQSSTDGSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTL AKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDP
[0066] TadA7.10 CP83 YRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLH YPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDGSSGS ETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVL NNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQN
[0067] TadA7.10 CP136 MNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDGSSGSETP GTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNR VIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVM CAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPG
[0068] TadA7.10 C-truncate GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVG AVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVE ITEGILADECAALLCYFFRMPRQVFN
[0069] TadA7.10 C-truncate 2 GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVG AVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVE ITEGILADECAALLCYFFRMPRQ
[0070] TadA7.10 delta59-66+C-truncate GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVG AVLVLNNRVIGEGWNRAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVM CAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILAD ECAALLCYFFRMPRQVFN
[0071] TadA7.10 delta 59-66 GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVG AVLVLNNRVIGEGWNRAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVM CAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILAD ECAALLCYFFRMPRQVFNAQKKAQSSTD
[0072] "Agent" means any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or a fragment thereof. thereof.
[0073] "Modifying a mutation" means
[0074] "Modification" means a change in the structure, expression level, or activity of a gene or polypeptide that is detected by standard methods known in the art as described herein. When used herein, a modification (e.g., an increase or decrease) includes a 10% change, a 25% change, a 40% change, and a change of 50% or more in the expression level. thereof. thereof. thereof.
[0075] "Analog" means a molecule that has similar functional or structural characteristics but is not identical. For example, a polynucleotide analog has certain modifications that enhance its analog function compared to a natural polynucleotide while retaining the biological activity of the corresponding natural polynucleotide. Such modifications can increase the affinity of the polynucleotide for DNA, its half-life, and / or its nuclease resistance. An analog can contain non-natural nucleotides or amino acids. thereof. thereof. thereof. thereof. thereof.
[0076] In the present disclosure, "comprises", "comprising", "containing", "having", etc. have the meanings defined in the United States Patent Law and can mean "includes", "including", etc., "consisting essentially of" thereof. thereof. "consisting essentially of" or "consists essentially y)" also has the meaning defined in the US Patent Law, and the term is open-ended such that other entities are tolerated as long as the basic or novel characteristics of the recited entity are not changed by other entities from what is recited, except for embodiments of the prior art that are excluded
[0077] "Base Editor (BE)" or "Nucleic Acid Base Editor (NBE)" means an agent that binds to a polynucleotide and has nucleic acid base modification activity. In one embodiment, the agent is a fusion protein comprising a domain having base editing activity, i.e., a domain capable of modifying a base (e.g., A, T, C, G, U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the domain having base editing activity is capable of deaminating a base within a nucleic acid molecule. In one embodiment, the base editor is capable of deaminating a base within a DNA molecule. In one embodiment, the base editor is capable of deaminating cytosine (C) or adenosine within DNA. In one embodiment, the base editor is a Cytidine Base Editor (CBE). In some embodiments, the base editor is an Adenosine Base Editor (ABE). In some embodiments, the base editor is an Adenosine Base Editor (ABE) and a Cytidine Base Editor (C BE). In one embodiment, the base editor is a nuclease-inactivated Cas9 (dCas9) fused to adenosine deaminase. In some embodiments, Cas9 is It is a circular permutant of Cas9 (such as spCas9 or saCas9). Circular permutant Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. In some embodiments, the base editor is fused to an inhibitor of base excision repair, such as the UGI domain. In certain embodiments, the fusion protein comprises a deaminase and a Cas9 nickase fused to an inhibitor of base excision repair such as the UGI domain. In other embodiments, the base editor is an abasic base editor.
[0078] In certain embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide-programmable DNA-binding domain is a CRISPR-associated (such as Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically dead Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the deaminase domain is an N-terminal or C-terminal fragment of the polynucleotide-programmable DNA-binding domain. In some embodiments, the deaminase is flanked by an N-terminal and a C-terminal fragment of the polynucleotide-programmable DNA-binding domain. In some embodiments, the deaminase domain is inserted into a site of the polynucleotide-programmable DNA-binding domain. In some embodiments, In this case, this base editor is fused to an inhibitor of base excision repair (BER). In certain embodiments, the inhibitor of base excision repair is uracil DNA glycosylase inhibitor (U GI). In certain embodiments, the inhibitor of base excision repair is an inosine base excision repair inhibitor. Details of the base editor are described in International PCT Application PCT / 2017 / 045381 (WO2018 / 02707 8) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is hereby incorporated by reference in its entirety into this specification. Also hereby incorporated by reference in its entirety into this specification is Komor, A.C., et al., “Programmable editing of a target base in genomic DN A without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N .M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A b ase editors with higher efficiency and product purity” Science Advances 3:eaao4 774 (2017), and Rees, H.A., et al., “Base editing: precision chemistry on the “genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770 -788. doi: 10.1038 / s41576-018-0059-1. See also
[0079] In some embodiments, the deaminase domain is inserted into a region of a polynucleotide-programmable DNA-binding domain. In some embodiments, the insertion site is determined by structural analysis of napDNAbp. In some embodiments, the insertion site is a flexible loop. In some embodiments, the deaminase domain is inserted into a site within a polynucleotide-programmable DNA-binding domain, where the site is selected from at least one from the group of amino acid positions consisting of 1029, 1026, 1054, 1022, 1015, 1068, 1247, 1040, 1248, and 768. In certain embodiments, the deaminase domain is inserted in place of a domain of a polynucleotide-programmable DNA-binding domain. In some embodiments, the domain is selected from the group consisting of RuvC, Rec1, Rec2, and HNH. In some embodiments, the deaminase domain is inserted in place of a range of amino acid residues in a polynucleotide-programmable DNA-binding domain, where the range of amino acid residues is selected from the group consisting of residues 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in Cas9 or corresponding positions in SEQ ID NO: 1. In some embodiments, the domain is selected from the group consisting of RuvC, Rec1, Rec2, and HNH. In some embodiments, the deaminase domain is inserted in place of a range of amino acid residues in a polynucleotide-programmable DNA-binding domain, where the range of amino acid residues is selected from the group consisting of residues 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in Cas9 or corresponding positions in SEQ ID NO: 1. In some embodiments, the deaminase domain is inserted in place of a range of amino acid residues in a polynucleotide-programmable DNA-binding domain, where the range of amino acid residues is selected from the group consisting of residues 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 in Cas9 or corresponding positions in SEQ ID NO: 1. 1052-1077, 1232-1248, and 1298-1300 or corresponding positions. is performed. By comparing the Cas9 amino acid sequences, how to identify the homologous regions in the different polynucleotide programmable DNA binding domains will be apparent to those skilled in the art. In some embodiments, the base editor comprises two or more deaminase domains inserted at two or more sites of the polynucleotide programmable DNA binding domain, and these sites are as described above.
[0080] In some embodiments, the base editor is generated by cloning an adenosine deaminase variant (e.g., TadA*7.10) into a scaffold comprising a circularly permuted Cas9 (e.g., spCAS9) and a bipartite nuclear localization sequence. Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circularly permuted sequences are described below, where the bold sequences indicate sequences derived from Cas9, the italicized sequences indicate linker sequences, and the underlined sequences indicate bipartite nuclear localization sequences.
[0081] JPEG2025102784000002.jpg187163
[0082] The nucleobase component of the base editor system and the polynucleotide programmable nucleotide binding component can be covalently or non-covalently linked to each other. For example, in some embodiments, the deaminase domain is targeted to the target nucleotide sequence by a polynucleotide programmable nucleotide binding domain. In certain embodiments, the polynucleotide programmable nucleotide binding domain The domain can be fused or linked to a deaminase domain. In some embodiments the polynucleotide programmable nucleotide binding domain targets the deaminase domain to the target nucleotide sequence by non-covalently interacting or binding with the deaminase domain. For example, in some embodiments the nucleic acid base editing component, such as the deaminase component, can interact, associate, or form a complex with an additional heterologous moiety or domain that is part of the polynucleotide programmable nucleotide binding domain. In some embodiments, the additional heterologous moiety can bind, interact, associate, or form a complex with a polypeptide. In some embodiments, the additional heterologous moiety can bind, interact, associate, or form a complex with a polynucleotide. In some embodiments, the additional heterologous moiety can bind to a guide polynucleotide. In some embodiments, the additional heterologous moiety can bind to a polypeptide linker. In some embodiments, the additional heterologous moiety can bind to a polynucleotide linker. The additional heterologous moiety can be a protein domain. In some embodiments, the additional heterologous moiety can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA
[0083] The base editor system can further include a guide polynucleotide component. The components of the base editor system are understood to be covalently, non-covalently bonded to each other, or any combination of those bonds and interactions. In certain embodiments, the deaminase domain can be targeted to the target nucleotide sequence by a guide polynucleotide. For example, in some embodiments , a nucleic acid base editing component of the base editor system, such as a deaminase component, can interact with, bind to, or form a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif), and can further include heterologous moieties or domains (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein). In some embodiments , an additional heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein) can be fused or linked to the deaminase domain. In some embodiments , an additional heterologous moiety can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments , an additional heterologous moiety can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments , an additional heterologous moiety can bind to a guide polynucleotide. In some The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.
[0084] In certain embodiments, the base editor system can further comprise an inhibitor component of base excision repair (BER). The components of the base editor system can be covalently bound, non-covalently interact, or any combination of their binding and interactions to each other. It should be understood that the inhibitor of the BER component can include a base excision repair inhibitor. In certain embodiments, the inhibitor of base excision repair can be a uracil DNA glycosylase inhibitor (UGI). In certain embodiments, the inhibitor of base excision repair can be an inosine base excision repair inhibitor. In certain embodiments, the inhibitor of base excision repair can be targeted to the target nucleotide sequence by a polynucleotide-programmable nucleotide binding domain. In certain embodiments, the polynucleotide-programmable nucleotide binding domain can be fused or linked to the inhibitor of base excision repair. In certain embodiments, the polynucleotide-programmable nucleotide binding domain can be fused or linked to a deaminase domain and the inhibitor of base excision The mingable nucleotide binding domain can non-covalently interact with an inhibitor of base excision repair or target the inhibitor of base excision repair to a target nucleotide sequence by associating with the inhibitor of base excision repair. For example, in some embodiments, the inhibitor of the base excision repair component may interact, associate, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide programmable nucleotide binding domain. In certain embodiments, the inhibitor of base excision repair can be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the inhibitor of base excision repair may include a further heterologous moiety or domain that can interact, associate, or form a complex with a part or segment of the guide polynucleotide (e.g., a polynucleotide motif), such as a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein) that can interact, associate, or form a complex with a part or segment of the guide polynucleotide. In some embodiments, the further heterologous moiety or domain of the guide polynucleotide (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein) can be fused or linked to the inhibitor of base excision repair. In some embodiments, an additional heterologous moiety can bind, interact, associate, or form a complex with a polynucleotide. In some embodiments, an additional heterologous moiety can bind to a guide polynucleotide. In some embodiments, an additional heterologous moiety can bind to a polypeptide linker. In some embodiments, an additional heterologous moiety can bind to a polynucleotide linker. The additional heterologous moiety can be a ta It may be a protein domain. In some embodiments, the additional heterologous moiety is K homologous (KH) domain, MS2 coat protein domain, PP7 coat protein domain, SfM u Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.
[0085] "Base editing activity" means acting to chemically change a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example, the activity of converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine deaminase activity, for example, the activity of converting A·T to G·C. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example, the activity of converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine deaminase activity, for example, the activity of converting A·T to G·C. In one embodiment, the base editing activity is cytidine deaminase activity, for example, the activity of converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine deaminase activity, for example, the activity of converting A·T to G·C. In another embodiment, the base editing activity is adenosine deaminase activity, for example, the activity of converting A·T to G·C. activity, for example, the activity of converting A·T to G·C.
[0086] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., an active, inactive, or partially active DNA cleavage domain of Cas9, and / or a protein comprising the gRNA-binding domain of Cas9). Cas9 nuclease is sometimes referred to as c asn1 nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids). A CRISPR cluster contains spacers, sequences complementary to the preceding mobile element, and a target invading nucleic acid. CRIS is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids). A CRISPR cluster contains spacers, sequences complementary to the preceding mobile element, and a target invading nucleic acid. CRIS The PR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein. TracrRNA guides the processing of pre-crRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require both a protein and both RNAs. However, a single guide RNA (“sgRNA” or simply “gRNA”) can be created that incorporates both sides of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are well known to those of skill in the art (e.g., “Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al In the tem, correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. TracrRNA serves as a guide for the processing of pre-crRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRN A endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then exo nucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require both a protein and both RNAs. However, it is possible to create a single guide RNA (the “sgRNA” or simply the “gRNA”) that incorporates both sides of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self. The sequence and structure of the Cas9 nuclease are well known to those skilled in the art (e.g., “Comp lete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al by incorporating both sides of crRNA and tracrRNA into a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self. The sequence and structure of the Cas9 nuclease are well known to those skilled in the art (e.g., “Comp lete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al in the CRISPR repeat sequence to help distinguish self from non-self. The sequence and structure of the Cas9 nuclease are well known to those of skill in the art (e.g., “Comp lete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al lete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al , J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C, Sez ate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F. Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughl in R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658 - 4663(2001); “CRISPR RNA maturati on by trans - encoded small RNA and host factor RNase III.” Deltcheva E., Chylins ki K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Ch arpentier E., Nature 471:602 - 607(2011); and “A programmable dual - RNA - guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonf ara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816 - 821(2012). See also the entire contents of which are incorporated herein by reference.). Cas9 orthologs have been described in various species including, but not limited to, S. pyogenes and S. thermophilus. Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure. See also the entire contents of which are incorporated herein by reference.). Cas9 orthologs have been described in various species including, but not limited to, S. pyogenes and S. thermophilus. Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure. See also the entire contents of which are incorporated herein by reference.). Cas9 orthologs have been described in various species including, but not limited to, S. pyogenes and S. thermophilus. Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure. resulting in such Cas9 nucleases and sequences being from organisms and loci disclosed in Chylinski, Rhun, and Charpentier , “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (201 3) RNA Biology 10:5, 726-737, and including Cas9 sequences from such organisms and loci. The entire content is incorporated herein by reference.
[0087] The nuclease-inactivated Cas9 protein can be interchangeably referred to as the “dCas9” protein (meaning nuclease-“ dead” Cas9) or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) having an inactive DNA cleavage domain are known (see, for example, Jinek et al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression”(2013) Cell. 28; 152 (5): 1173-83, each of which is incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, the mutations D10A and H840A suppress the nuclease activity of S. pyogenes Cas9 Completely inactivate the nuclease activity (Jinek et al, Science. 337:816-821(2012); Qi et a l, Cell. 28;152(5): 1173-83 (2013)). In certain embodiments, the Cas9 nuclease has an inactiv e (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase called the "nCas9" prote in (meaning "nickase" Cas9). In certain embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of the following two Cas9 domains: (1) the gRNA binding domain of Cas9; (2) the C as9 DNA cleavage domain. In certain embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a " Cas9 variant". Cas9 variants share homology with Cas9 or fragments thereof. For example, Cas9 variants have at least about 70% identity, at least about 8 0% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity with wild-type Cas9. In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 2 9, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 4 9, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 va riant... ... riant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 ba The variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), wherein the fragment has at least about 70% identity with the corresponding fragment of wild-type Cas9, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, and also has at least about 99.9% identity. In certain embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the corresponding amino acid length of wild-type Cas 9.
[0088] In certain embodiments, the fragment is at least 100 amino acids in length. In certain embodiments the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 60 0, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length. In certain embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcu s pyogenes (NCBI reference sequence: NC_17053.1, nucleotide sequence and amino acid sequence are as follows).
[0089] ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025102784000003.jpg172166(Underlined once: HNH domain; Underlined twice: RuvC domain)
[0090] In certain embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences: : ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAACCTGATCGCACAATTACCCGGAGAGAAGAAAAATGGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGACAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAAACGGGTACGCAGGTTATATTGACGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAAGATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGAACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCAAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025102784000004.jpg172165(Underlined once: HNH domain; Underlined twice: RuvC domain)
[0091] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NC BI reference sequence: NC_002737.2 (the following nucleotide sequence) and Uniprot reference sequence: Q99ZW2 (the following amino acid sequence). ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AAATTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025102784000005.jpg168163(SEQ ID NO: 1. Single underline: HNH domain; double underline: RuvC domain)
[0092] In certain embodiments, Cas9 is from Corynebacterium ulcerans (NCBI Refs: NC_015683.1 , NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcu s iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psyc hroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref : YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni ( NCBI Ref: YP_002344900.1) or Neisseria. meningitidis (NCBI Ref: YP_00234210 0.1), or Cas9 from any other organism.
[0093] In some embodiments, dCas9 corresponds to a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity, or comprises a part or all thereof. For example, in some embodiments, the dCas9 domain has D10A and H840A mutations or another Ca It includes the corresponding mutations in s9. In some embodiments, dCas9 includes the amino acid sequence of dCas9 (D10A and H840A): JPEG2025102784000006.jpg171165 (single underline: HNH domain; double underline: RuvC domain)
[0094] In some embodiments, the Cas9 domain includes the D10A mutation, while the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any of the amino acid sequences provided herein, remains histidine.
[0095] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which result in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain ). In one aspect, variants or homologs of dCas9 are provided that have at least about 70% identity, at least about 80% identity, at least about 90% identity , at least about 95% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity. In one aspect, variants of dCas9 are provided that have an amino acid sequence that is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids shorter or longer. It is.
[0096] In some embodiments, the Cas9 fusion protein provided herein comprises the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence and comprises only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and further suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art. In some embodiments, Cas9 is from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia
[0097] (NCBI Ref: NC_017861.1); Spiroplasma taiwanense(NCBI Ref: NC_021846.1); Strepto coccus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jej coccus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB I Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jej uni (NCBI Ref: YP_002344900.1); or Cas9 derived from Neisseria. meningitidis (NCBI Ref: YP_0023 42100.1).
[0098] Additional Cas9 proteins (e.g., nuclease - dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease - active Cas9) are to be understood to be within the scope of the present disclosure, including their variants and homologs . Exemplary Cas9 proteins include, but are not limited to, those provided below. In certain embodiments, the Cas9 protein is nuclease - dead Cas9 (dCas9). In certain embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease - active Cas9.
[0099] Exemplary catalytically - inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0100] Exemplary catalytic Cas9 nickase (nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0101] Exemplary catalytically active Cas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0102] In certain embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaea) that constitute the domain and kingdom of unicellular prokaryotic microorganisms. In certain embodiments, the Cas9 protein refers to, for example, CasX or CasY as described in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res . 2017 Feb 21. doi: 10.1038 / cr.2017.21, and its The entire content is incorporated herein by reference. Using genomic resolution metagenomics, many CRISPR-Cas systems have been identified, including Cas9, which was first reported in the archaeal domain of life. This divergent Cas9 protein was discovered as part of an active CRISPR-Cas system in Nanoarchaea, which have been little studied. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, and they fall into the most compact systems discovered to date. In some embodiments, Cas9 represents CasX or a variant of CasX. In some embodiments, Cas9 represents CasY or a variant of CasY. Other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp: nucleic acid programmable DNA binding protein), and it should be understood that they are within the scope of the present
[0103] disclosure. In some embodiments, any nucleic acid programmable DNA-binding protein (napDNAbp) of the fusion proteins provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least comprises an amino acid sequence having at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity. In some embodiments, napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, napDNAbp is at least 85%, at least at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the CasX or CasY proteins described herein. CasX and CasY from other bacterial species are also understood to be usable in accordance with the present disclosure.
[0104] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus isla ndicus (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG
[0105] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus is landicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG
[0106] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacterium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI
[0107] The term "cytidine deaminase" refers to a polypeptide or a fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine . PmCDA1 (Petromyzon marinus cytosine deaminase 1, "PmCDA1") derived from Petromyzon marinus, AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases . The nucleotide sequences and amino acid sequences of PmCDA1, and the nucleotide sequences and amino acid sequences of the CDS of human AID are shown below .
[0108] PmCDA1's nucleotide sequence and amino acid sequence, as well as the nucleotide sequence and amino acid sequence of the CDS of human AID are shown below .
[0109] >tr|A5H718|A5H718_PETMA Cytosine deaminase OS=Petromyzon marinus OX=7757 PE=2 SV =1 MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSG TERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLK IWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKT LKRAEKRRSELSIMIQVKILHTTKSPAV
[0110] >EF094822.1 Petromyzon marinus isolate PmCDA.21 cytosine deaminase mRNA, complet e cds TGACACGACACAGCCGTGTATATGAGGAAGGGTAGCTGGATGGGGGGGGGGGGAATACGTTCAGAGAGGA CATTAGCGAGCGTCTTGTTGGTGGCCTTGAGTCTAGACACCTGCAGACATGACCGACGCTGAGTACGTGA GAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACAACAAAAAATCCGTGTCGCA TAGATGCTACGTTCTCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTTGGGGCTATGCTGTG AATAAACCACAGAGCGGGACAGAACGTGGAATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAAT ACCTGCGCGACAACCCCGGACAATTCACGATAAATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGC TGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGC AAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGT TGAATGTAATGGTAAGTGAACACTACCAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATT GAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATG ATTCAGGTAAAAATACTCCACACCACTAAGAGTCCTGCTGTTTAAGAGGCTATGCGGATGGTTTTC
[0111] >tr|Q6QJ80|Q6QJ80_HUMAN Activation-induced cytidine deaminase OS=Homo sapiens OX =9606 GN=AICDA PE=2 SV=1 MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELL FLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRK AEPEGLRRLHRAGVQIAIMTFKAPV
[0112] >NG_011588.1:5001-15681 Homo sapiens activation induced cytidine deaminase (AICD A), RefSeqGene (LRG_17) on chromosome 12 AGAGAACCATCATTAATTGAAGTGAGATTTTTCTGGCCTGAGACTTGCAGGGAGGCAAGAAGACACTCTG GACACCACTATGGACAGGTAAAGAGGCAGTCTTCTCGTGGGTGATTGCACTGGCCTTCCTCTCAGAGCAA ATCTGAGTAATGAGACTGGTAGCTATCCCTTTCTCTCATGTAACTGTCTGACTGATAAGATCAGCTTGAT CAATATGCATATATATTTTTTGATCTGTCTCCTTTTCTTCTATTCAGATCTTATACGCTGTCAGCCCAAT TCTTTCTGTTTCAGACTTCTCTTGATTTCCCTCTTTTTCATGTGGCAAAAGAAGTAGTGCGTACAATGTA CTGATTCGTCCTGAGATTTGTACCATGGTTGAAACTAATTTATGGTAATAATATTAACATAGCAAATCTT TAGAGACTCAAATCATGAAAAGGTAATAGCAGTACTGTACTAAAAACGGTAGTGCTAATTTTCGTAATAA TTTTGTAAATATTCAACAGTAAAACAACTTGAAGACACACTTTCCTAGGGAGGCGTTACTGAAATAATTT AGCTATAGTAAGAAAATTTGTAATTTTAGAAATGCCAAGCATTCTAAATTAATTGCTTGAAAGTCACTAT GATTGTGTCCATTATAAGGAGACAAATTCATTCAAGCAAGTTATTTAATGTTAAAGGCCCAATTGTTAGG CAGTTAATGGCACTTTTACTATTAACTAATCTTTCCATTTGTTCAGACGTAGCTTAACTTACCTCTTAGG TGTGAATTTGGTTAAGGTCCTCATAATGTCTTTATGTGCAGTTTTTGATAGGTTATTGTCATAGAACTTA TTCTATTCCTACATTTATGATTACTATGGATGTATGAGAATAACACCTAATCCTTATACTTTACCTCAAT TTAACTCCTTTATAAAGAACTTACATTACAGAATAAAGATTTTTTAAAAATATATTTTTTTGTAGAGACA GGGTCTTAGCCCAGCCGAGGCTGGTCTCTAAGTCCTGGCCCAAGCGATCCTCCTGCCTGGGCCTCCTAAA GTGCTGGAATTATAGACATGAGCCATCACATCCAATATACAGAATAAAGATTTTTAATGGAGGATTTAAT GTTCTTCAGAAAATTTTCTTGAGGTCAGACAATGTCAAATGTCTCCTCAGTTTACACTGAGATTTTGAAA ACAAGTCTGAGCTATAGGTCCTTGTGAAGGGTCCATTGGAAATACTTGTTCAAAGTAAAATGGAAAGCAA AGGTAAAATCAGCAGTTGAAATTCAGAGAAAGACAGAAAAGGAGAAAAGATGAAATTCAACAGGACAGAA GGGAAATATATTATCATTAAGGAGGACAGTATCTGTAGAGCTCATTAGTGATGGCAAAATGACTTGGTCA GGATTATTTTTAACCCGCTTGTTTCTGGTTTGCACGGCTGGGGATGCAGCTAGGGTTCTGCCTCAGGGAG CACAGCTGTCCAGAGCAGCTGTCAGCCTGCAAGCCTGAAACACTCCCTCGGTAAAGTCCTTCCTACTCAG GACAGAAATGACGAGAACAGGGAGCTGGAAACAGGCCCCTAACCAGAGAAGGGAAGTAATGGATCAACAA AGTTAACTAGCAGGTCAGGATCACGCAATTCATTTCACTCTGACTGGTAACATGTGACAGAAACAGTGTA GGCTTATTGTATTTTCATGTAGAGTAGGACCCAAAAATCCACCCAAAGTCCTTTATCTATGCCACATCCT TCTTATCTATACTTCCAGGACACTTTTTCTTCCTTATGATAAGGCTCTCTCTCTCTCCACACACACACAC ACACACACACACACACACACACACACACACACAAACACACACCCCGCCAACCAAGGTGCATGTAAAAAGA TGTAGATTCCTCTGCCTTTCTCATCTACACAGCCCAGGAGGGTAAGTTAATATAAGAGGGATTTATTGGT AAGAGATGATGCTTAATCTGTTTAACACTGGGCCTCAAAGAGAGAATTTCTTTTCTTCTGTACTTATTAA GCACCTATTATGTGTTGAGCTTATATATACAAAGGGTTATTATATGCTAATATAGTAATAGTAATGGTGG TTGGTACTATGGTAATTACCATAAAAATTATTATCCTTTTAAAATAAAGCTAATTATTATTGGATCTTTT TTAGTATTCATTTTATGTTTTTTATGTTTTTGATTTTTTAAAAGACAATCTCACCCTGTTACCCAGGCTG GAGTGCAGTGGTGCAATCATAGCTTTCTGCAGTCTTGAACTCCTGGGCTCAAGCAATCCTCCTGCCTTGG CCTCCCAAAGTGTTGGGATACAGTCATGAGCCACTGCATCTGGCCTAGGATCCATTTAGATTAAAATATG CATTTTAAATTTTAAAATAATATGGCTAATTTTTACCTTATGTAATGTGTATACTGGCAATAAATCTAGT TTGCTGCCTAAAGTTTAAAGTGCTTTCCAGTAAGCTTCATGTACGTGAGGGGAGACATTTAAAGTGAAAC AGACAGCCAGGTGTGGTGGCTCACGCCTGTAATCCCAGCACTCTGGGAGGCTGAGGTGGGTGGATCGCTT GAGCCCTGGAGTTCAAGACCAGCCTGAGCAACATGGCAAAACGCTGTTTCTATAACAAAAATTAGCCGGG CATGGTGGCATGTGCCTGTGGTCCCAGCTACTAGGGGGCTGAGGCAGGAGAATCGTTGGAGCCCAGGAGG TCAAGGCTGCACTGAGCAGTGCTTGCGCCACTGCACTCCAGCCTGGGTGACAGGACCAGACCTTGCCTCA AAAAAATAAGAAGAAAAATTAAAAATAAATGGAAACAACTACAAAGAGCTGTTGTCCTAGATGAGCTACT TAGTTAGGCTGATATTTTGGTATTTAACTTTTAAAGTCAGGGTCTGTCACCTGCACTACATTATTAAAAT ATCAATTCTCAATGTATATCCACACAAAGACTGGTACGTGAATGTTCATAGTACCTTTATTCACAAAACC CCAAAGTAGAGACTATCCAAATATCCATCAACAAGTGAACAAATAAACAAAATGTGCTATATCCATGCAA TGGAATACCACCCTGCAGTACAAAGAAGCTACTTGGGGATGAATCCCAAAGTCATGACGCTAAATGAAAG AGTCAGACATGAAGGAGGAGATAATGTATGCCATACGAAATTCTAGAAAATGAAAGTAACTTATAGTTAC AGAAAGCAAATCAGGGCAGGCATAGAGGCTCACACCTGTAATCCCAGCACTTTGAGAGGCCACGTGGGAA GATTGCTAGAACTCAGGAGTTCAAGACCAGCCTGGGCAACACAGTGAAACTCCATTCTCCACAAAAATGG GAAAAAAAGAAAGCAAATCAGTGGTTGTCCTGTGGGGAGGGGAAGGACTGCAAAGAGGGAAGAAGCTCTG GTGGGGTGAGGGTGGTGATTCAGGTTCTGTATCCTGACTGTGGTAGCAGTTTGGGGTGTTTACATCCAAA AATATTCGTAGAATTATGCATCTTAAATGGGTGGAGTTTACTGTATGTAAATTATACCTCAATGTAAGAA AAAATAATGTGTAAGAAAACTTTCAATTCTCTTGCCAGCAAACGTTATTCAAATTCCTGAGCCCTTTACT TCGCAAATTCTCTGCACTTCTGCCCCGTACCATTAGGTGACAGCACTAGCTCCACAAATTGGATAAATGC ATTTCTGGAAAAGACTAGGGACAAAATCCAGGCATCACTTGTGCTTTCATATCAACCATGCTGTACAGCT TGTGTTGCTGTCTGCAGCTGCAATGGGGACTCTTGATTTCTTTAAGGAAACTTGGGTTACCAGAGTATTT CCACAAATGCTATTCAAATTAGTGCTTATGATATGCAAGACACTGTGCTAGGAGCCAGAAAACAAAGAGG AGGAGAAATCAGTCATTATGTGGGAACAACATAGCAAGATATTTAGATCATTTTGACTAGTTAAAAAAGC AGCAGAGTACAAAATCACACATGCAATCAGTATAATCCAAATCATGTAAATATGTGCCTGTAGAAAGACT AGAGGAATAAACACAAGAATCTTAACAGTCATTGTCATTAGACACTAAGTCTAATTATTATTATTAGACA CTATGATATTTGAGATTTAAAAAATCTTTAATATTTTAAAATTTAGAGCTCTTCTATTTTTCCATAGTAT TCAAGTTTGACAATGATCAAGTATTACTCTTTCTTTTTTTTTTTTTTTTTTTTTTTTTGAGATGGAGTTT TGGTCTTGTTGCCCATGCTGGAGTGGAATGGCATGACCATAGCTCACTGCAACCTCCACCTCCTGGGTTC AAGCAAAGCTGTCGCCTCAGCCTCCCGGGTAGATGGGATTACAGGCGCCCACCACCACACTCGGCTAATG TTTGTATTTTTAGTAGAGATGGGGTTTCACCATGTTGGCCAGGCTGGTCTCAAACTCCTGACCTCAGAGG ATCCACCTGCCTCAGCCTCCCAAAGTGCTGGGATTACAGATGTAGGCCACTGCGCCCGGCCAAGTATTGC TCTTATACATTAAAAAACAGGTGTGAGCCACTGCGCCCAGCCAGGTATTGCTCTTATACATTAAAAAATA GGCCGGTGCAGTGGCTCACGCCTGTAATCCCAGCACTTTGGGAAGCCAAGGCGGGCAGAACACCCGAGGT CAGGAGTCCAAGGCCAGCCTGGCCAAGATGGTGAAACCCCGTCTCTATTAAAAATACAAACATTACCTGG GCATGATGGTGGGCGCCTGTAATCCCAGCTACTCAGGAGGCTGAGGCAGGAGGATCCGCGGAGCCTGGCA GATCTGCCTGAGCCTGGGAGGTTGAGGCTACAGTAAGCCAAGATCATGCCAGTATACTTCAGCCTGGGCG ACAAAGTGAGACCGTAACAAAAAAAAAAAAATTTAAAAAAAGAAATTTAGATCAAGATCCAACTGTAAAA AGTGGCCTAAACACCACATTAAAGAGTTTGGAGTTTATTCTGCAGGCAGAAGAGAACCATCAGGGGGTCT TCAGCATGGGAATGGCATGGTGCACCTGGTTTTTGTGAGATCATGGTGGTGACAGTGTGGGGAATGTTAT TTTGGAGGGACTGGAGGCAGACAGACCGGTTAAAAGGCCAGCACAACAGATAAGGAGGAAGAAGATGAGG GCTTGGACCGAAGCAGAGAAGAGCAAACAGGGAAGGTACAAATTCAAGAAATATTGGGGGGTTTGAATCA ACACATTTAGATGATTAATTAAATATGAGGACTGAGGAATAAGAAATGAGTCAAGGATGGTTCCAGGCTG CTAGGCTGCTTACCTGAGGTGGCAAAGTCGGGAGGAGTGGCAGTTTAGGACAGGGGGCAGTTGAGGAATA TTGTTTTGATCATTTTGAGTTTGAGGTACAAGTTGGACACTTAGGTAAAGACTGGAGGGGAAATCTGAAT ATACAATTATGGGACTGAGGAACAAGTTTATTTTATTTTTTGTTTCGTTTTCTTGTTGAAGAACAAATTT AATTGTAATCCCAAGTCATCAGCATCTAGAAGACAGTGGCAGGAGGTGACTGTCTTGTGGGTAAGGGTTT GGGGTCCTTGATGAGTATCTCTCAATTGGCCTTAAATATAAGCAGGAAAAGGAGTTTATGATGGATTCCA GGCTCAGCAGGGCTCAGGAGGGCTCAGGCAGCCAGCAGAGGAAGTCAGAGCATCTTCTTTGGTTTAGCCC AAGTAATGACTTCCTTAAAAAGCTGAAGGAAAATCCAGAGTGACCAGATTATAAACTGTACTCTTGCATT TTCTCTCCCTCCTCTCACCCACAGCCTCTTGATGAACCGGAGGAAGTTTCTTTACCAATTCAAAAATGTC CGCTGGGCTAAGGGTCGGCGTGAGACCTACCTGTGCTACGTAGTGAAGAGGCGTGACAGTGCTACATCCT TTTCACTGGACTTTGGTTATCTTCGCAATAAGGTATCAATTAAAGTCGGCTTTGCAAGCAGTTTAATGGT CAACTGTGAGTGCTTTTAGAGCCACCTGCTGATGGTATTACTTCCATCCTTTTTTGGCATTTGTGTCTCT ATCACATTCCTCAAATCCTTTTTTTTATTTCTTTTTCCATGTCCATGCACCCATATTAGACATGGCCCAA AATATGTGATTTAATTCCTCCCCAGTAATGCTGGGCACCCTAATACCACTCCTTCCTTCAGTGCCAAGAA CAACTGCTCCCAAACTGTTTACCAGCTTTCCTCAGCATCTGAATTGCCTTTGAGATTAATTAAGCTAAAA GCATTTTTATATGGGAGAATATTATCAGCTTGTCCAAGCAAAAATTTTAAATGTGAAAAACAAATTGTGT CTTAAGCATTTTTGAAAATTAAGGAAGAAGAATTTGGGAAAAAATTAACGGTGGCTCAATTCTGTCTTCC AAATGATTTCTTTTCCCTCCTACTCACATGGGTCGTAGGCCAGTGAATACATTCAACATGGTGATCCCCA GAAAACTCAGAGAAGCCTCGGCTGATGATTAATTAAATTGATCTTTCGGCTACCCGAGAGAATTACATTT CCAAGAGACTTCTTCACCAAAATCCAGATGGGTTTACATAAACTTCTGCCCACGGGTATCTCCTCTCTCC TAACACGCTGTGACGTCTGGGCTTGGTGGAATCTCAGGGAAGCATCCGTGGGGTGGAAGGTCATCGTCTG GCTCGTTGTTTGATGGTTATATTACCATGCAATTTTCTTTGCCTACATTTGTATTGAATACATCCCAATC TCCTTCCTATTCGGTGACATGACACATTCTATTTCAGAAGGCTTTGATTTTATCAAGCACTTTCATTTAC TTCTCATGGCAGTGCCTATTACTTCTCTTACAATACCCATCTGTCTGCTTTACCAAAATCTATTTCCCCT TTTCAGATCCTCCCAAATGGTCCTCATAAACTGTCCTGCCTCCACCTAGTGGTCCAGGTATATTTCCACA ATGTTACATCAACAGGCACTTCTAGCCATTTTCCTTCTCAAAAGGTGCAAAAAGCAACTTCATAAACACA AATTAAATCTTCGGTGAGGTAGTGTGATGCTGCTTCCTCCCAACTCAGCGCACTTCGTCTTCCTCATTCC ACAAAAACCCATAGCCTTCCTTCACTCTGCAGGACTAGTGCTGCCAAGGGTTCAGCTCTACCTACTGGTG TGCTCTTTTGAGCAAGTTGCTTAGCCTCTCTGTAACACAAGGACAATAGCTGCAAGCATCCCCAAAGATC ATTGCAGGAGACAATGACTAAGGCTACCAGAGCCGCAATAAAAGTCAGTGAATTTTAGCGTGGTCCTCTC TGTCTCTCCAGAACGGCTGCCACGTGGAATTGCTCTTCCTCCGCTACATCTCGGACTGGGACCTAGACCC TGGCCGCTGCTACCGCGTCACCTGGTTCACCTCCTGGAGCCCCTGCTACGACTGTGCCCGACATGTGGCC GACTTTCTGCGAGGGAACCCCAACCTCAGTCTGAGGATCTTCACCGCGCGCCTCTACTTCTGTGAGGACC GCAAGGCTGAGCCCGAGGGGCTGCGGCGGCTGCACCGCGCCGGGGTGCAAATAGCCATCATGACCTTCAA AGGTGCGAAAGGGCCTTCCGCGCAGGCGCAGTGCAGCAGCCCGCATTCGGGATTGCGATGCGGAATGAAT GAGTTAGTGGGGAAGCTCGAGGGGAAGAAGTGGGCGGGGATTCTGGTTCACCTCTGGAGCCGAAATTAAA GATTAGAAGCAGAGAAAAGAGTGAATGGCTCAGAGACAAGGCCCCGAGGAAATGAGAAAATGGGGCCAGG GTTGCTTCTTTCCCCTCGATTTGGAACCTGAACTGTCTTCTACCCCCATATCCCCGCCTTTTTTTCCTTT TTTTTTTTTTGAAGATTATTTTTACTGCTGGAATACTTTTGTAGAAAACCACGAAAGAACTTTCAAAGCC TGGGAAGGGCTGCATGAAAATTCAGTTCGTCTCTCCAGACAGCTTCGGCGCATCCTTTTGGTAAGGGGCT TCCTCGCTTTTTAAATTTTCTTTCTTTCTCTACAGTCTTTTTTGGAGTTTCGTATATTTCTTATATTTTC TTATTGTTCAATCACTCTCAGTTTTCATCTGATGAAAACTTTATTTCTCCTCCACATCAGCTTTTTCTTC TGCTGTTTCACCATTCAGAGCCCTCTGCTAAGGTTCCTTTTCCCTCCCTTTTCTTTCTTTTGTTGTTTCA CATCTTTAAATTTCTGTCTCTCCCCAGGGTTGCGTTTCCTTCCTGGTCAGAATTCTTTTCTCCTTTTTTT TTTTTTTTTTTTTTTTTTTTAAACAAACAAACAAAAAACCCAAAAAAACTCTTTCCCAATTTACTTTCTT CCAACATGTTACAAAGCCATCCACTCAGTTTAGAAGACTCTCCGGCCCCACCGACCCCCAACCTCGTTTT GAAGCCATTCACTCAATTTGCTTCTCTCTTTCTCTACAGCCCCTGTATGAGGTTGATGACTTACGAGACG CATTTCGTACTTTGGGACTTTGATAGCAACTTCCAGGAATGTCACACACGATGAAATATCTCTGCTGAAG ACAGTGGATAAAAAACAGTCCTTCAAGTCTTCTCTGTTTTTATTCTTCAACTCTCACTTTCTTAGAGTTT ACAGAAAAAATATTTATATACGACTCTTTAAAAAGATCTATGTCTTGAAAATAGAGAAGGAACACAGGTC TGGCCAGGGACGTGCTGCAATTGGTGCAGTTTTGAATGCAACATTGTCCCCTACTGGGAATAACAGAACT GCAGGACCTGGGAGCATCCTAAAGTGTCAACGTTTTTCTATGACTTTTAGGTAGGATGAGAGCAGAAGGT AGATCCTAAAAAGCATGGTGAGAGGATCAAATGTTTTTATATCAACATCCTTTATTATTTGATTCATTTG AGTTAACAGTGGTGTTAGTGATAGATTTTTCTATTCTTTTCCCTTGACGTTTACTTTCAAGTAACACAAA CTCTTCCATCAGGCCATGATCTATAGGACCTCCTAATGAGAGTATCTGGGTGATTGTGACCCCAAACCAT CTCTCCAAAGCATTAATATCCAATCATGCGCTGTATGTTTTAATCAGCAGAAGCATGTTTTTATGTTTGT ACAAAAGAAGATTGTTATGGGTGGGGATGGAGGTATAGACCATGCATGGTCACCTTCAAGCTACTTTAAT AAAGGATCTTAAAATGGGCAGGAGGACTGTGAACAAGACACCCTAATAATGGGTTGATGTCTGAAGTAGC AAATCTTCTGGAAACGCAAACTCTTTTAAGGAAGTCCCTAATTTAGAAACACCCACAAACTTCACATATC ATAATTAGCAAACAATTGGAAGGAAGTTGCTTGAATGTTGGGGAGAGGAAAATCTATTGGCTCTCGTGGG TCTCTTCATCTCAGAAATGCCAATCAGGTCAAGGTTTGCTACATTTTGTATGTGTGTGATGCTTCTCCCA AAGGTATATTAACTATATAAGAGAGTTGTGACAAAACAGAATGATAAAGCTGCGAACCGTGGCACACGCT CATAGTTCTAGCTGCTTGGGAGGTTGAGGAGGGAGGATGGCTTGAACACAGGTGTTCAAGGCCAGCCTGG GCAACATAACAAGATCCTGTCTCTCAAAAAAAAAAAAAAAAAAAAGAAAGAGAGAGGGCCGGGCGTGGTG GCTCACGCCTGTAATCCCAGCACTTTGGGAGGCCGAGCCGGGCGGATCACCTGTGGTCAGGAGTTTGAGA CCAGCCTGGCCAACATGGCAAAACCCCGTCTGTACTCAAAATGCAAAAATTAGCCAGGCGTGGTAGCAGG CACCTGTAATCCCAGCTACTTGGGAGGCTGAGGCAGGAGAATCGCTTGAACCCAGGAGGTGGAGGTTGCA GTAAGCTGAGATCGTGCCGTTGCACTCCAGCCTGGGCGACAAGAGCAAGACTCTGTCTCAGAAAAAAAAA AAAAAAAGAGAGAGAGAGAGAAAGAGAACAATATTTGGGAGAGAAGGATGGGGAAGCATTGCAAGGAAAT TGTGCTTTATCCAACAAAATGTAAGGAGCCAATAAGGGATCCCTATTTGTCTCTTTTGGTGTCTATTTGT CCCTAACAACTGTCTTTGACAGTGAGAAAAATATTCAGAATAACCATATCCCTGTGCCGTTATTACCTAG CAACCCTTGCAATGAAGATGAGCAGATCCACAGGAAAACTTGAATGCACAACTGTCTTATTTTAATCTTA TTGTACATAAGTTTGTAAAAGAGTTAAAAATTGTTACTTCATGTATTCATTTATATTTTATATTATTTTG CGTCTAATGATTTTTTATTAACATGATTTCCTTTTCTGATATATTGAAATGGAGTCTCAAAGCTTCATAA ATTTATAACTTTAGAAATGATTCTAATAACAACGTATGTAATTGTAACATTGCAGTAATGGTGCTACGAA GCCATTTCTCTTGATTTTTAGTAAACTTTTATGACAGCAAATTTGCTTCTGGCTCACTTTCAATCAGTTA AATAAATGATAAATAATTTTGGAAGCTGTGAAGATAAAATACCAAATAAAATAATATAAAAGTGATTTAT ATGAAGTTAAAATAAAAAATCAGTATGATGGAATAAACTTG
[0113] Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like (APOBEC) is an evolutionarily conserved cytidine deaminase family. Members of this family are editing enzymes that convert C to U 。The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain 。More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain 。 and is important for cytidine deamination. APOBEC family members include AP OBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced cytidine deaminase. Without limitation, a number of modified cytidine deaminases including SaBE3, SaKKH-BE3, VQR-BE3, EQR-B E3, VRER-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3 are commercially available and can be obtained from Addgene (plasmids 85169, 85170, 85171, 85172,
[0114] Other exemplary deaminases that can be fused to Cas9 according to aspects of the present disclosure are provided below. In some embodiments, it should be understood that domains without the active domain of each sequence, for example, localization signals (nuclear localization sequence, nuclear export signal, cytoplasmic
[0115] Human AID: JPEG2025102784000007.jpg33165 (underline: nuclear localization sequence; double
[0116] underline: nuclear export signal) Mouse AID:
[0117] Dog AID: JPEG2025102784000009.jpg33165(Underline: Nuclear localization sequence; Double underline: Nuclear export signal)
[0118] Bovine AID: JPEG2025102784000010.jpg34165(Underline: Nuclear localization sequence; Double underline: Nuclear export signal)
[0119] Rat AID: JPEG2025102784000011.jpg35165(Underline: Nuclear localization sequence; Double underline: Nuclear export signal)
[0120] Mouse APOBEC-3 JPEG2025102784000012.jpg53164(Italic: Nucleic acid editing domain)
[0121] Rat APOBEC-3 JPEG2025102784000013.jpg52166(Italic: Nucleic acid editing domain)
[0122] Rhesus macaque APOBEC-3 G: JPEG2025102784000014.jpg47166(Italic: Nucleic acid editing domain; Underline: Cytoplasmic localization signal)
[0123] Chimpanzee APOBEC-3 G: JPEG2025102784000015.jpg44168(Italic: Nucleic acid editing domain; Underline: Cytoplasmic localization signal)
[0124] Green monkey APOBEC-3G: JPEG2025102784000016.jpg46168(Italic: Nucleic acid editing domain; Underline: Cytoplasmic localization signal)
[0125] Human APOBEC-3G: JPEG2025102784000017.jpg47168(Italic: Nucleic acid editing domain; Underline: Cytoplasmic localization signal)
[0126] Human APOBEC-3F: JPEG2025102784000018.jpg44168 (Italic: nucleic acid editing domain)
[0127] Human APOBEC-3B: JPEG2025102784000019.jpg45165 (Italic: nucleic acid editing domain)
[0128] Rat APOBEC-3B: MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRK QGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCR LIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRY YRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILR IYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL
[0129] Bovine APOBEC-3B: DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQL KQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHW IKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLQRILTAPI
[0130] Chimpanzee APOBEC-3B: MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQL SAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENF VYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCN EAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYK EALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPPQSPGP CLPLCSEPPLGSLLPTGRPAPSLPFLLTASFSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWA SVSKEGRDLG
[0131] Human APOBEC-3C: JPEG2025102784000020.jpg29165
[0132] Gorilla APOBEC3C: JPEG2025102784000021.jpg25165 (Italic: Nucleic acid editing domain)
[0133] Human APOBEC-3A: JPEG2025102784000022.jpg24165 (Italic: Nucleic acid editing domain)
[0134] Rhesus macaque APOBEC-3 A: JPEG2025102784000023.jpg25165 (Italic: Nucleic acid editing domain)
[0135] Bovine APOBEC-3 A: JPEG2025102784000024.jpg25168 (Italic: Nucleic acid editing domain)
[0136] Human APOBEC-3H: JPEG2025102784000025.jpg24168 (Italic: Nucleic acid editing domain)
[0137] Rhesus macaque APOBEC-3H: MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYL TWSPCPSCAGELVDFIKAHRHLNLRIFASRLYYHWRPNYQEGLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPS EKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR
[0138] Human APOBEC-3D JPEG2025102784000026.jpg44168 (Italic: Nucleic acid editing domain)
[0139] Human APOBEC-1: MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSM SCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR
[0140] Mouse APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNT RCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYP PSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK
[0141] Rat APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK
[0142] Human APOBEC-2: MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKG GQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEP EIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK
[0143] Mouse APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKG GQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK
[0144] Rat APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK
[0145] Bovine APOBEC-2: MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEP EIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK
[0146] Petromyzon marinus CDA1 (pmCDAl) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLR DNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCR KIFIQSSHNQLNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV
[0147] Human APOBEC3G D316R D317R MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKL HRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKFNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGF LCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTC FTSWSPCFSCAQEMAKFISKKHVSLCIFTARIYRRQGRC QEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN
[0148] Human APOBEC3G chain A MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDY RVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISF TYSEFKHCWDTFVDHQGCP FQPWDGLD EHSQDLSGRLRAILQ
[0149] Human APOBEC3G chain A, D120R, D121R MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQD YRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGC PFQPWDGLDEHSQDLSGRLRAILQ
[0150] The term "deaminase" or "deaminase domain" refers to a protein or a fragment thereof that catalyzes a deamination reaction. The term "detect" refers to identifying the presence, absence, or amount of an analyte to be detected.
[0151] In one embodiment, sequence changes in a polynucleotide or polypeptide are detected. In another embodiment, the presence of an indel is detected.
[0152] The term "detectable label" means a composition that enables a molecule of interest to be detected when linked thereto, by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (such as those commonly used in ELISA), biotin, di goxygenin, or haptens.
[0153] "Fragment" means a part of a polypeptide or nucleic acid molecule. This part is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the full length of the reference nucleic acid molecule or polypeptide, or 90%. Fragments can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 00, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0154] "Hybridization" means hydrogen bonding between complementary nucleic acid bases and can be Watson- Crick, Hoogsteen or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleic acid bases that form hydrogen bonds to form a pair.
[0155] The term "inhibitor of base repair" or "IBR" refers to a protein that can inhibit the activity of nucleic acid repair enzymes, such as base excision repair enzymes. In one embodiment, IBR is an inhibitor of inosine base excision repair. Examples of inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL and hAAG. In certain embodiments, I BR is an inhibitor of Endo V or hAAG. In certain embodiments, IBR is catalytically inactive It is either EndoV or catalytically inactive hAAG.
[0156] The terms "isolated", "purified", or "biologically pure" refer to a substance from which the components that are normally associated with it in its natural state have been removed to varying degrees. "Isolation" indicates the degree of separation from the original source or the surrounding environment. "Purification" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein has other substances sufficiently removed so that impurities do not substantially affect the biological properties of the protein or cause other adverse consequences. That is, the nucleic acids or peptides of the present invention are purified when they are substantially free of cellular material, viral material, or medium when produced by recombinant DNA technology, or are substantially free of chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to essentially one band on an electrophoretic gel. For proteins that can undergo modifications such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be purified separately.
[0157]
[0157] "Isolated polynucleotide" means a nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene of interest in the natural genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, this term includes, for example, nucleic acids incorporated into a vector; plasmids or viruses that replicate autonomously; Integrated; integrated into the genomic DNA of prokaryotes or eukaryotes; or separate from other sequences as a distinct molecule (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). Further, this term includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences.
[0158] The term "isolated polypeptide" means a polypeptide of the present invention that is separated from the components that accompany it in its natural state. Typically, a polypeptide is isolated when it is at least 60% free by weight from the proteins and natural organic molecules with which it associates in its natural state. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99% by weight the polypeptide of the present invention. The isolated polypeptide of the present invention can be obtained, for example, by extraction from a natural source, expression of a recombinant nucleic acid encoding such a polypeptide; or by chemically synthesizing the protein. Purity can be measured by any suitable method, e.g., column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0159] The term "linker" as used herein refers to a bond (e.g., a covalent bond), chemical group, or molecule that links two molecules or moieties (e.g., two domains of a fusion protein). In certain embodiments, the linker is between the gRNA-binding domain of an RNA-programmable nuclease that includes a Cas9 nuclease domain and a nucleic acid editing protein (e.g., cytidine Connect the catalytic domain of deaminase or adenosine deaminase). In certain embodiments the linker connects dCas9 and the nucleic acid editing protein. Typically, the linker is placed between or adjacent to two groups, molecules, or other moieties and is covalently linked to each via a covalent bond, thus linking the two. In certain embodiments the linker is an amino acid or a plurality of amino acids (e.g., a peptide or a protein). In certain embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In certain embodiments, the linker is 5-200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length. Longer or shorter linkers are also contemplated. In certain embodiments, the linker can also be called an XTEN linker and comprises the amino acid sequence SGSETPGTSESATPES. In certain embodiments, the linker comprises the amino acid sequence SGGS. In certain embodiments, the linker is (SGGS) n, (GGGS) n , (GGGGS) n , (G) n n(EA n、 AAK) n, (GGS) n n, SGSETPGTSESATPES, or the (XP) n motif, or any combination thereof, where n is independently an integer from 1 to 30 and X is any n amino acid. In certain embodiments In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0160] In some embodiments, the domain of the nucleobase editor is fused via a linker comprising the amino acid sequence SGGSSGSETPGTSESATP ESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPG TSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domain of the nucleobase editor is fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some aspects, the linker is 24 amino acids in length. In some aspects, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some aspects, the linker is 40 amino acids in length. In some aspects, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some aspects, the linker is 64 amino acids in length. In some aspects, the linker comprises the amino acid sequence SG GSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some aspects, the linker is 92 amino acids in length. In some aspects, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GTSTEPSEGS It contains APGTSESATPESGPGSEPATS.
[0161] As used herein, the term "mutation" refers to the substitution of a residue within a sequence, e.g., within a nucleic acid or amino acid sequence, by another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are typically described herein by identifying the original residue, then identifying the position of the residue within the sequence, and then identifying the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)).
[0162] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds that contain nucleobases and an acidic moiety, such as nucleosides, nucleotides, or polynucleotides of nucleotides. Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides, is a linear molecule in which adjacent nucleotides are linked to each other via phosphodiester bonds. In certain embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In certain embodiments, "nucleic acid" refers to an oligonucleotide chain containing three or more individual nucleotide residues. As used herein, the terms "oligo nucleotide" and "polynucleotide" refer to polymers of nucleotides (e.g., at least three). nucleotides). can be used interchangeably to refer to a chain of at least three nucleotides. In certain embodiments the term "nucleic acid" includes RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be present in nature, for example in association with a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome chromatid, or other naturally occurring nucleic acid molecule. Alternatively a nucleic acid molecule can be a non-naturally occurring molecule, such as a recombinant DNA or RNA, artificial chromosome, engineered genome, or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or other non-naturally occurring molecule that includes non-naturally occurring nucleotides or nucleosides. Further, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources produced using recombinant expression systems and optionally purified, chemically synthesized etc. In the case of chemically synthesized molecules, nucleic acids can include, where appropriate for example, nucleoside analogs such as chemically modified bases or sugars, and analogs having backbone modifications Nucleic acid sequences are represented in the 5' to 3' direction unless otherwise indicated In certain embodiments, nucleic acids include natural nucleosides (e.g., adenosine, thymidine, guanosine cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2 -aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5 -methylcytidine); etc. -Propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadeno sine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguan ine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribo se, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) or contain them.
[0163] The terms "nuclear localization sequence", "nuclear localization signal" or "NLS" mean an amino acid sequence that promotes the translocation of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, by Plank et al. in the international PCT application PCT / EP2000 / 011690, filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is, for example, the optimized NLS described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In certain embodiments, the NLS contains the amino acid sequences KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0164] The present disclosure provides nucleic acid programmable nucleic acid (e.g., DNA or RNA) binding proteins. Nucleic acid programmable nucleic acid binding proteins are, for example, "nucleic acid programmable DNA binding proteins" or "napDNAbp". The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to a protein that associates with a nucleic acid (e.g., DNA or RNA) such as a guide nucleic acid that directs the napDNAbp to a specific nucleic acid sequence. For example, the Cas9 protein can bind to a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, for example, nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpfl, Cas12b / C2c1, and Cas12c / C2c3. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure even if not specifically listed herein. As used herein, "obtaining" as in "obtaining an agent" includes synthesizing, purchasing, or otherwise acquiring the agent. The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to cleavage ... ... ...
[0165] ... ...
[0166] ... Used in conjunction with (e.g., binds to or associates with) one or more non-target RNAs In some embodiments, the RNA programmable nuclease, when complexed with RNA, Typically, the bound RNA is called a guide RNA (gRNA). gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule is called a single guide RNA (sgRNA). Although sometimes referred to as a single molecule, "gRNA" may be used as a complex of two or more molecules. are used interchangeably to refer to guide RNAs that exist as a single RNA species. The gRNA present in the target nucleic acid has (1) a domain that shares homology with the target nucleic acid (e.g., a domain that facilitates Cas9 amplification of the target nucleic acid). (2) a domain that binds to the Cas9 protein In one embodiment, domain (2) is a polypeptide that is directed against a sequence known as tracrRNA. For example, in some embodiments, domain (2) comprises: Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. The gRNA (e.g., domain) is identical or homologous to the tracrRNA as provided in Another example of a method for detecting 2-dependent cas9 nucleases is the Switchable Cas9 Nucleases and Uses Thereof. Provisional patent application USSN 61 / 874,682, filed September 6, 2013, and "Delivery System F or Functional Nucleases” filed on September 6, 2013. It can be found in 1 / 874,746, and the entire content of each of them is incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2) and can be referred to as an " extended gRNA". By way of example, an extended gRNA can, as described herein, for example, bind to two or more Cas9 proteins and bind to target nucleic acids in two or more different regions. The gRNA comprises a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides the sequence specificity of the nuclease:RNA complex. In one aspect, the RNA programmable nuclease is a (CRISPR associated system) Cas9 endonuclease, such as Cas9 (Csnl) from Streptococcus pyogenes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes s." Ferretti J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Prime aux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G. , Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A. , McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658 - 4663(2001); "CRISPR RNA , Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A. , McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658 - 4663(2001); "CRISPR RNA "maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Voge l J., Charpentier E., Nature 471:602-607(2011) (see reference).
[0167] As used herein, the term "recombinant" in relation to a protein or nucleic acid refers to a protein or nucleic acid that does not exist in nature but is a product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 mutations as compared to any naturally occurring sequence.
[0168] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.
[0169] "Reference" means a standard or control condition.
[0170] "Reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset or the entirety of a particular sequence; for example, a segment of a full-length cDNA or gene sequence or a full-length cDNA or gene sequence. For a polypeptide, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least At least about 25 amino acids, more preferably about 35 amino acids, about 50 amino acids, or about 10 0 amino acids. For nucleic acids, the length of the reference nucleic acid sequence is generally at least about 50 nu cleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nu cleotides or about 300 nucleotides or any integer around or between them is.
[0171] "Specifically binds" means recognizing and binding to the polypeptide and / or nucleic acid molecule of the present invention, but substantially not recognizing and binding to other molecules in a sample (e.g., a biological sample) and binding to nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DN A binding domain and guide nucleic acid), compounds, or molecules.
[0172] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridize" means species At least one strand of the double-stranded nucleic acid molecule can hybridize. "Hybridize" means species Under various stringency conditions, it means forming pairs of double-stranded molecules between complementary polynucleotide sequences (e.g., the genes described herein), or between parts thereof. (See, for example, Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507).
[0173] For example, stringent salt concentrations are usually less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of an organic solvent such as formamide, whereas high stringency hybridization can be obtained in the presence of at least about 35% form amide, more preferably in the presence of at least about 50% formamide. Stringent temperature conditions will usually include a temperature of at least about 30°C, more preferably at least about 37 °C, and most preferably at least about 42°C. During hybridization, the concentration of a surfactant (e.g., sodium dodecyl sulfate (SDS)), and various additional parameters such as the inclusion or exclusion of carrier DNA are well known to those skilled in the art. By combining these various conditions as needed, various levels of stringency are achieved. In one embodiment, hybridization occurs at 30°C in 750 mM NaCl, 75 mM trisodium citrate and 1% SDS. In another embodiment, hybridization Hybridization occurs in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37°C. In another embodiment, hybridization occurs in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA at 42°C. Useful variations of these conditions will be readily apparent to those skilled in the art.
[0174] In most applications, the washing step following hybridization is also of varying stringency. Washing stringency conditions can be defined by salt concentration and temperature. As noted above, washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, a stringent salt concentration for the wash step is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash step generally include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the wash step is performed in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS at 25°C. In a more preferred embodiment, the wash step is performed in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 42°C. In a more preferred embodiment, , The washing step is carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York. "Substantially identical" means a polypeptide or nucleic acid molecule that shows at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). In one embodiment, such sequences have at least 60%, 80%, 85%, 90%, 95%, or 99% identity at the amino acid level or in nucleic acids to the sequences used for comparison.
[0175]
[0176] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Grou p, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison , Wis. 53705's Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PIL EUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glyc ine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; phenylalan ine, tyrosine. In an exemplary approach for determining the degree of identity, the BLAST pro gram can be used, and sequences with probability scores between e -3 and e -100 that are closely related are shown.
[0177] "Subject" means a mammal including, but not limited to, a human or non-human mammal such as a cow, horse, dog, sheep or cat. A subject includes a breeding animal that is bred to supply commodities such as labor and food for consumption, including, but not limited to, cows, goats, chickens, horses, pigs, rabbits, and sheep.
[0178] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleic refers to the array to be generated. In one embodiment, the target site is deaminated by a deaminase or a fusion protein containing the same (e.g., cytidine or adenine deaminase).
[0179] RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, so in principle these proteins can target any sequence specified by the guide RNA. Methods of using RNA-programmable nucleases such as Cas9 for site-specific cleavage (e.g., for genome modification) are known in the art (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823 -826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a C RISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA -programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, J.E. . Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacteri al genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013) for reference. The entire contents of each of these are hereby incorporated by reference). ).
[0180] The ranges provided herein are to be understood as being shorthand for all of the values within the range. For example, a range of 1 to 50 is to be understood as including any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50). ). ).
[0181] Unless otherwise specified or clear from the context, the term " or" as used herein is to be understood as being inclusive. Unless otherwise specified or clear from the context, the terms "a", "an", and "the" as used herein are to be understood as being singular or plural. ). ).
[0182] Unless otherwise specified or clear from the context, the term "about" as used herein is to be understood as being within the normal tolerance in the art, e.g., within 2 standard deviations of the mean value. About can be 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5 It can be understood to be within %, 0.1%, 0.05%, or 0.01%. Unless otherwise apparent from the context, all numerical values provided in this specification are modified by the term "about".
[0183] The description of the list of chemical groups in the definition of a variable factor in this specification includes the definition of that variable factor as any single group or combination of the listed groups. The description of an embodiment regarding a variable factor or aspect in this specification includes that embodiment as any single embodiment, or in combination with any other embodiment or part thereof.
[0184] The compositions or methods provided in this specification can be combined with one or more of any other compositions and methods provided in this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0185]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
DETAILED DESCRIPTION OF THE INVENTION
[0186] As described below, the present invention relates to a base editor with reduced off-target deamination, a method of using such base editor, and an assay for characterizing a base editor with reduced off-target deamination (e.g., as compared to programmed on-target deamination). is characterized.
[0187] [Adenosine deaminase] In certain embodiments, the nucleic acid base editor of the present invention comprises an adenosine deaminase domain. In some embodiments, the adenosine deaminase provided herein can deaminate adenine. In some embodiments, the adenosine deaminase provided herein can deaminate adenine in the deoxyadenosine residues of DNA. The adenosine deaminase can be derived from any suitable organism (e.g., Escherichia coli). In some embodiments, the adenine deaminase is a native adenosine deaminase that contains one or more mutations corresponding to any of the mutations provided herein (e.g., mutations in ecTadA). One of ordinary skill in the art can identify corresponding residues in any homologous protein, for example, by sequence alignment and determination of homologous residues. Thus, one of ordinary skill in the art can generate mutations corresponding to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) in any native adenosine deaminase (e.g., one having homology to ecTadA). In certain embodiments, the adenosine deaminase is of prokaryotic origin. In certain embodiments, Moreover, adenosine deaminase is of bacterial origin. In certain embodiments, the adenosine de aminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewane lla putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In certain embodiments, the adenosine deaminase is from E. coli .
[0188] In one embodiment, the fusion protein of the invention comprises wild-type TadA linked to TadA7.10, which is linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single TadA7.10 domain (e.g., provided as a monomer). In other embodiments , the ABE7.10 editor comprises TadA7.10 and TadA(wt) that can form a heterodimer. The relevant sequences are as follows.
[0189] TadA(wt): SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLY VTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD
[0190] TadA7.10: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD
[0191] In one aspect, adenosine deaminase is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% with respect to any of the amino acid sequences described in any of the adenosine deaminases provided herein, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. The adenosine deaminases provided herein are understood to be capable of including one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides deaminase domains having a particular percent identity that additionally include any one or combination of the mutations described herein. In some embodiments, the adenosine deaminase has an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 compared to a reference sequence or any of the adenosine deaminases provided herein 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations. comprises. In some embodiments, adenosine deaminase is known in the art or compared to any of the amino acid sequences described herein, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues. The amino acid sequence is included.
[0192] In one aspect, adenosine deaminase comprises the D108X mutation in the TadA reference sequence or the corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In one aspect, adenosine deaminase comprises the D108G, D108N, D108V, D108A or D108Y mutation in the TadA reference sequence or the corresponding mutation in another adenosine deaminase. However, it should be understood that further deaminases can be aligned in the same way to identify homologous amino acid residues that can be mutated as provided herein.
[0193] In one aspect, adenosine deaminase comprises the A106X mutation in the TadA reference sequence or the corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine Indicates any amino acid other than the corresponding amino acid in nosin deaminase. In certain embodiments the adenosine deaminase comprises the A106V mutation in the TadA reference sequence, or another corresponding mutation in an adenosine deaminase.
[0194] In certain embodiments, the adenosine deaminase comprises the E155X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises the E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase .
[0195] In certain embodiments, the adenosine deaminase comprises the D147X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises the D147Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0196] It should be understood that any of the mutations provided herein (e.g., based on the ecTadA amino acid sequence of the TadA reference sequence) can be introduced into other adenosine deaminases such as Staphylococcus aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). How homologous the mutated residues in ecTadA are should be understood. This will be apparent to those skilled in the art. Thus, any of the mutations identified in ecTadA can be made in other adenosine deaminases having the same amino acid residues. Also, any of the mutations provided in this specification can be made, either individually or in any combination, in ecTadA or another adenosine deaminase. For example, adenosine deaminase can include the D108N, A106V, E155V, and / or D147Y mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase. In certain embodiments, adenosine deaminase includes the following groups of mutations (groups of mutations are separated by ";") in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E55V; D108N, A106V, and D147Y; D108N, E55V, and D147Y; A106V, E55V, and D147Y; and D108N, A106V, E55V, and D147Y; provided that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (e.g., ecTadA). It should be understood that. In some embodiments, adenosine deaminase has H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X in the TadA reference sequence. It should be understood that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (e.g., ecTadA). It should be understood that.
[0197] In some embodiments, adenosine deaminase has H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X in the TadA reference sequence. T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X , A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I one or more of 156X and / or K157X mutations, or one or more corresponding mutations in another adenosine deaminase wherein the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments the adenosine deaminase has H8Y, T17S, L18E, W23L associated with the TadA reference sequence, L34S, W45L, R51H, A56E, or A56S, E59G, E85K, or E85G, M94L, 1951, V102A , F104L, A106V, R107C, or R107H, or R107P, D108G, or D108N, or D108A , D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or one or more of K157R mutations, or one or more corresponding mutations in another adenosine deaminase . In one aspect, the adenosine deaminase includes one or more of H8X, D108X, and / or N127X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments
[0198] the adenosine deaminase has one or more of H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase . In some embodiments, the adenosine deaminase includes one or more of H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase . In some embodiments, the adenosine deaminase includes one or more of H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase .
[0199] In some embodiments, adenosine deaminase has an H8X in the TadA reference sequence , R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X , K161X, Q163X, and / or one or more of the T166X mutations, or one or more corresponding mutations in another adenosine deaminase , where X indicates the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments , adenosine deaminase has one or more of the H8Y, R26W, M61I, L68Q mutations in the TadA reference sequence, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H, and / or one or more of the T166P mutations, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, adenosine deaminase has one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in the TadA reference sequence, or corresponding
[0200] mutations (s) in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adeno sine deaminase has one, two, three, four, five, six, seven, or selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E1 55X, and Q163X in the TadA reference sequence, or corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adeno sine deaminase has one, two, three, four, five, six, seven, or more mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E1 comprises eight mutations, or corresponding mutations (plural possible) in another adenosine deaminase, where X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase indicating the presence of. In certain embodiments, the adenosine deaminase is TadA selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the reference sequence, one , two, three, four, or five mutations, or corresponding mutations (plural possible) in another adenosine deaminase, where X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase indicating the presence of. In certain embodiments, the adenosine deaminase comprises 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of H8X, A106X, D108X, mutations (plural possible) in another adenosine deaminase, where X is the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase indicating. In certain embodiments, the adenosine deaminase is TadA and comprises 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of H8X, R 126X, L68X, D108X, N127X, D147X, and E155X in the reference sequence, one, two, three four, five, six, seven, or eight mutations, or corresponding mutations (plural possible) in another adenosine deaminase, where X is the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase indicating. In certain embodiments, the adenosine deaminase is TadA and comprises 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the reference sequence, one two, three, four, or five mutations, or corresponding mutations (plural possible) in another adenosine deaminase where X is the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase indicating. In certain embodiments, the adenosine deaminase is TadA and comprises 1, 2, 3, 4, or 5 mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the reference sequence, one two, three, four, or five mutations, or corresponding mutations (plural possible) in another adenosine Comprising the corresponding mutation(s) in adenosine deaminase, where X is wild-type adenosine de Indicates the presence of any amino acid other than the corresponding amino acid in deaminase.
[0201] In some embodiments, the adenosine deaminase has H8Y in the TadA reference sequence , D108N, N127S, D147Y, R152C, and Q154H, and is selected from the group consisting of 1, 2, 3, 4, 5, or 6 mutations, or the corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase has H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the TadA reference sequence and is selected from the group consisting of 1, 2, 3, 4, 5, 6, 7, or 8 mutations, or the corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase has H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence and is selected from the group consisting of 1, 2, 3, 4, or 5 mutations, or the corresponding mutation(s) in another adenosine deaminase. In some embodiments, the adenosine deaminase has H8Y, A106T, D108N, N127S, E1 55D, and K161Q in the TadA reference sequence and is selected from the group consisting of 1, 2, 3, 4, 5, or 6 mutations, or the corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase has H8Y, R126W, L68Q in the TadA reference sequence and is selected from the group consisting of 1, 2, 3, 4, 5, or 6 mutations, or the corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase has H8Y, R126W, L68Q in the TadA reference sequence , one, two, three, four, five selected from the group consisting of D108N, N127S, D147Y, and E155V two, six, seven, or eight mutations, or corresponding mutations in other adenosine deaminases. In some embodiments, the adenosine deaminase is TadA reference One selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the alignment Two, three, four, or five mutations, or corresponding Mutations (plural possible) in other adenosine deaminases.
[0202] In one aspect, the adenosine deaminase comprises one or more corresponding mutations in another adenosine deaminase. In one aspect, the adenosine deaminase is Ta dA reference sequence contains the D108N, D108G, or D108V mutation, or corresponding Mutations in other adenosine deaminases. In one aspect, the adenosine deaminase Contains the A106V and D108N mutations in the TadA reference sequence, or corresponding Mutations in other adenosine deaminases. In one aspect, the adenosine deaminase is Ta dA reference sequence contains the R107C and D108N mutations, or corresponding Mutations in other adenosine deaminases. In one aspect, the adenosine deaminase is the TadA reference Contains the H8Y, D108N, N127S, D147Y, and Q154H mutations in the sequence, or corresponding Mutations in other adenosine deaminases. In one aspect, the adenosine deaminase Contains the H8Y, R24W, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence , or a corresponding mutation in another adenosine deaminase. In certain embodiments the adenosine deaminase comprises the D108N, D147Y, and E155V mutations in the TadA reference sequence, , or a corresponding mutation in another adenosine deaminase. In certain embodiments the adenosine deaminase comprises the H8Y, D108N, and S127S mutations in the TadA reference sequence, , or a corresponding mutation in another adenosine deaminase. In certain embodiments the adenosine deaminase comprises the A106V, D108N, D147Y, and E155V mutations in the TadA reference sequence, , or a corresponding mutation in another adenosine deaminase. .
[0203] In some embodiments, the adenosine deaminase comprises one or more of the S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X mutations in the tadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S mutations in the TadA reference sequence, , or one or more corresponding mutations in another adenosine deaminase.
[0204] In certain embodiments, the adenosine deaminase comprises the L84X mutant adenosine deaminase, where X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. represents an amino acid. In certain embodiments, adenosine deaminase comprises the L8 4F mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0205] In certain embodiments, adenosine deaminase comprises the H123X mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type ad enosine deaminase. In certain embodiments, adenosine deaminase comprises the H123Y mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0206] In certain embodiments, adenosine deaminase comprises the I157X mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type ad enosine deaminase. In certain embodiments, adenosine deaminase comprises the I157F mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0207] In certain embodiments, adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X , D108X, H123X, D147X, E155X, and I156X in the TadA reference sequence, or the corresponding mutation(s) in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments, adenosine The deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the tadA reference sequence, or corresponding mutation(s) in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase. In certain embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in the TadA reference sequence.
[0208] In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S in the TadA reference sequence.
[0209] or five mutations, or corresponding mutations in another adenosine deaminase, or comprise a mutation.
[0210] In some embodiments, the adenosine deaminase has one or more of the E25 X, R26X, R107X, A142X, and / or A143X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In one aspect, the adenosine deaminase has E25M, E25D, E25A , E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A , R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A 143W, A143M, A143S, A143Q and / or A143R mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase has one or more of the mutations described herein corresponding to the TadA reference sequence or one or more of the corresponding mutations in another adenosine deaminase.
[0211] In one aspect, the adenosine deaminase has the E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X is the wild-type ade Indicates any amino acid other than the corresponding amino acid in adenosine deaminase. In certain embodiments adenosine deaminase comprises an E25M, E25D, E25A, E25R, E2 5V, E25S, or E25Y mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0212] In certain embodiments, adenosine deaminase comprises an R26X mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase, where X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adenosine deaminase comprises an R26G, R26N, R26Q , R26C, R26L, or R26K mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0213] In certain embodiments, adenosine deaminase comprises an R107X mutation in the TadA reference sequence , or the corresponding mutation in another adenosine deaminase, where X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments adenosine deaminase comprises an R107P, R07K, R107A, R107 N, R107W, R107H, or R107S mutation in the TadA reference sequence, or the corresponding mutation in another adenosine deaminase.
[0214] In certain embodiments, adenosine deaminase comprises an A142X mutation in the TadA reference sequence , or the corresponding mutation in another adenosine deaminase, where X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments represents any amino acid other than the corresponding amino acid in deoxyinosine deaminase. In one embodiment in, deoxyinosine deaminase contains the A142N, A142D, A142G mutations in the TadA reference sequence, or the corresponding mutations in another deoxyinosine deaminase.
[0215] In one embodiment, deoxyinosine deaminase contains the A143X mutation in the TadA reference sequence, or the corresponding mutation in another deoxyinosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In one embodiment in, deoxyinosine deaminase contains the A143D, A143G, A143E, A14 3L, A143W, A143M, A143S, A143Q and / or A143R mutations, or the corresponding mutations in another adenosine deaminase. In some embodiments, deoxyinosine deaminase contains one or more of the H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or K161X mutations in the TadA reference sequence, or one or more corresponding mutations in another deoxyinosine deaminase, where the presence of X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adenosine
[0216] deaminase contains one or more of the H36L, N37T, N37S, P48T, P48L, I49V, R51H in the TadA reference sequence, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T mutations, or one or more corresponding mutations in another adenosine deaminase. Here, the presence of X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adenosine deaminase contains one or more of the H36L, N37T, N37S, P48T, P48L, I49V, R51H in the TadA reference sequence, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T mutations, or one or more corresponding mutations in another adenosine deaminase. One or more mutations in the adenosine deaminase, or one or more corresponding mutations in another adenosine deaminase comprises.
[0217] In some embodiments, the adenosine deaminase comprises an H36X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase In some embodiments, the adenosine deaminase comprises an H36L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0218] In one aspect, the adenosine deaminase comprises an N37X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adeno sine deaminase. In one aspect, the adenosine deaminase comprises an N37T or N37S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0219] In one aspect, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild type adenosine deaminase. In one aspect, the adenosine deaminase comprises a P48T or P48L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0220] In certain embodiments, the adenosine deaminase comprises an R51X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises an R51H or R51L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0221] In certain embodiments, the adenosine deaminase comprises an S146X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises an S146R or S146C mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0222] In certain embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In certain embodiments, the adenosine deaminase comprises a K157N mutation in the TadA reference sequence, or another adenosine deaminase with a corresponding mutation.
[0223] In certain embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X is a wild Indicates any amino acid other than the corresponding amino acid in type adenosine deaminase. In one embodiment, the adenosine deaminase comprises a P48S, P48T, or P4 8A mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0224] In one embodiment, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence , or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In one embodiment , the adenosine deaminase comprises an A142N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0225] In one embodiment, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In one embodiment , the adenosine deaminase comprises a W23R or W23L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0226] In one embodiment, the adenosine deaminase comprises an R152X mutation in the TadA reference sequence , or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In one embodiment, the adenosine deaminase comprises an R152P or R52H mutation in the TadA reference sequence comprises a spontaneous mutation, or a corresponding mutation in another adenosine deaminase.
[0227] In one embodiment, the adenosine deaminase can comprise the mutations H36L, R51L, L84F, A106V, D10 8N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments the adenosine deaminase comprises the following combinations of mutations related to the tadA reference sequence where each mutation in the combination is separated by "_" and each combination of mutations is within parentheses: (A106V_D108N), (R107C_D108N), (H8Y_D108N_S 127S_D 147Y_Q154H), (H8Y_R24W_D108N_N127S_D147Y_E155V), (D108N_D147 Y_E155V), (H8Y_D108N_S 127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V D108N D147Y E155V) (D108Q D147Y E155V) (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147 Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V), (E59A cat dead_A106V_D108N_D147Y_E155V), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D103A_D014N), (G22P_D 103 A_D 104N), (G22P_D 103 A_D 104N_S 138 A), (D 103 A_D 104N_S 138A), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I15 6F), (R 26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (R26C_L 84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_A1 42N_A143L_D147Y_E155V_I156F), (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I15 6F), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (A106V_D108N_A142N_D147Y_E155V), (R26G_A106V_D108N_A142N_D147Y_E155V), (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V), (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V), (E25D_R26G_A106V_D108N_A142N_D147Y_E155V), (A106V_R107K_D108N_A142N_D147Y_E155V), (A106V_D108N_A142N_A143G_D147Y_E155V), (A106V_D108N_A142N_A143L_D147Y_E155V), (H36L_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F), (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T), (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F), (N72S_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F), (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F), 57N), (H36L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F_K161T), (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N), (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E), (H36L_G67V_L84F_A106V_D108N_H123Y_S 146T_D147Y_E155V_I156F), (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F), (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F), (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F), (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F), (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E), (R74Q L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F), (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (P48S_A142N), (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N), (P48T_I49V_A142N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F (H36L_P48T _I49V_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_ I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152H_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_E155 V_I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_R152P _E155V_I156 F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_R152P_E155 V_I156F _K1 57N).
[0228] [Cytidine deaminase] In one embodiment, the fusion protein of the present invention comprises cytidine deaminase. In certain aspects, the cytidine deaminase provided herein can deaminate cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytosine deaminase provided herein can deaminate cytosine in DNA. The cytidine deaminase can be derived from any suitable organism. It can be. In some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase that contains one or more mutations corresponding to any of the mutations provided herein. One of ordinary skill in the art can identify the corresponding residues in any homologous protein, for example, by sequence alignment and determination of homologous residues. Thus, one of ordinary skill in the art can introduce a mutation corresponding to any of the mutations described herein into any naturally occurring cytidine deaminase. In one embodiment, the cytidine deaminase is of prokaryotic origin. In one embodiment, the cytidine deaminase is of bacterial origin. In one embodiment, the cytidine deaminase is of mammalian (e.g., human) origin. In some embodiments, the cytidine deaminase has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the cytidine deaminase amino acid sequences described herein. It should be understood that the cytidine deaminases provided herein may include one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain having a specific percent identity to which any of the mutations or combinations thereof described herein are added. In some embodiments, the cytidine deaminase is a reference sequence
[0229] In some embodiments, the cytidine deaminase is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the cytidine deaminase amino acid sequences described herein. It should be understood that the cytidine deaminases provided herein may include one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain having a specific percent identity to which any of the mutations or combinations thereof described herein are added. In some embodiments, the cytidine deaminase is a reference sequence including an amino acid sequence that is. The cytidine deaminases provided herein may include one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain having a specific percent identity to which any of the mutations or combinations thereof described herein are added. In some embodiments, the cytidine deaminase is a reference sequence It should be understood that the cytidine deaminase may include one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain having a specific percent identity to which any of the mutations or combinations thereof described herein are added. In some embodiments, the cytidine deaminase is a reference sequence Please understand that it may include one or more mutations (for example, any of the mutations provided herein). The present disclosure provides any deaminase domain having a specific percent identity to which any of the mutations or combinations thereof described herein are added. In some embodiments, the cytidine deaminase is a reference sequence It should be understood that any of the mutations described herein or combinations thereof are added to any deaminase domain having a specific percent identity. In some embodiments, the cytidine deaminase is a reference sequence In some embodiments, the cytidine deaminase is a reference sequence or an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the cytidine deaminases provided herein. In some embodiments the cytidine deaminase, compared to any of the amino acid sequences known in the art or described herein has at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues .
[0230] The fusion protein of the present invention comprises a nucleic acid editing domain. In one aspect, the nucleic acid editing do main can catalyze a base change from C to U. In one aspect, the nucleic acid editing do main is a deaminase domain. In one aspect, the deaminase is cytidine deaminase or adenosine deaminase. In one aspect, the deaminase is , an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In one In one aspect, the deaminase is an APOBEC deaminase. In certain aspects, the deaminase is APOBEC 2 deaminase. In certain aspects, the deaminase is APOBEC 3 deaminase. In certain aspects, the deaminase is APOBEC3A deaminase . In certain aspects, the deaminase is APOBEC3B deaminase. In certain aspects the deaminase is APOBEC3C deaminase. In certain aspects, the deaminase is APOBEC3D deaminase. In certain aspects, the deaminase is APOBEC3E deaminase . In certain aspects, the deaminase is APOBEC3F deaminase. In certain aspects the deaminase is APOBEC3G deaminase. In certain aspects, the deaminase is APOBEC3H deaminase. In certain aspects, the deaminase is APOBEC4 deaminase. In certain aspects, the deaminase is activation-induced deaminase (AID) . In certain aspects, the deaminase is a vertebrate deaminase. In certain aspects the deaminase is an invertebrate deaminase. In certain aspects, the deaminase is a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse deaminase . In certain aspects, the deaminase is a human deaminase. In certain aspects the deaminase is a rat deaminase, such as rAPOBEC1. In certain aspects the deaminase is Petromyzon marinus cytidine deaminase 1 (pmCDA1) . In certain aspects, the deaminase is human APOBEC3G. In certain aspects, the de The deaminase is a fragment of human APOBEC3G. In certain embodiments, the deaminase is a human APOBEC3G variant comprising the D316R and D317R mutations. In certain embodiments, the deaminase is a fragment of human APOBEC3G and comprises mutations corresponding to the D316R and D317R mutations. In certain embodiments, the nucleic acid editing domain has at least 80%, at least 85%, at least 90%, at least 92% at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the deaminase domain of any of the deaminases described herein. In certain embodiments, the fusion proteins provided herein include one or more features that improve the base
[0231] editing activity of the fusion protein. For example, the fusion proteins provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein include a Cas9 domain that has no nuclease activity (dCas9), or a Cas9 nickase (nCas9), which cleaves one strand of a double-stranded DNA molecule. (dCas9), or a Cas9 nickase (nCas9), which cleaves one strand of a double-stranded DNA molecule.
[0232] [Cas9 domain of nucleic acid base editor] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is the Cas9 domain. Exemplary non-limiting Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In certain embodiments, the Cas9 domain is a nuclease-active domain It is. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain contains any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 8 5%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 amino acid sequence having mutations of 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 compared to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80 compared to any one of the amino acid sequences described herein, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least one of the amino acid sequences described herein. In some embodiments, the Cas9 domain is at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80 30, at least 40, at least 50, at least 60, at least 70, at least 80 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues.
[0233] In certain embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10X and H840X mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10A and H840A mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences described herein. As an example, the nuclease-inactive Cas9 domain comprises the following amino acid sequence described in the cloning vector pPlatTET-gRNA2 (accession number BAV54124): MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, see Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-s pecific control of gene expression.” Cell. 2013; 152(5):1173-83. The entire content is incorporated herein by reference.).
[0234] Further suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and the knowledge in the art and are within the scope of the present disclosure. Such further exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H84 0A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains ([[]] For example, Prashant et al., CAS9 transcriptional activators for target specificity Screening and paired nickases for cooperative genome engineering. Nature Biotech See 2013; 31(9): 833-838 of Nature Biotechnology (the entire content of which is incorporated herein by reference). In some embodiments, the dCas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 100 mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000 mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000 mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000 mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000 mutations compared to any of the amino acid sequences described herein. 0, at least 1100, or at least 1200 identical consecutive amino acid residues. It contains the amino acid sequence.
[0235] In one embodiment, the Cas9 domain is a Cas9 nickase. Cas9 protein, which can cut only one strand of a nucleic acid molecule (e.g., a double-stranded DNA molecule) In some embodiments, the Cas9 nickase can be used to target the double-stranded nucleic acid molecule. This is because the Cas9 nickase cleaves the gRNA (e.g., sgRNA) bound to the Cas9. It means to cleave paired (complementary) strands. The Cas9 nickase contains a D10A mutation and has a histidine at position 840. In embodiments, the Cas9 nickase cleaves the non-targeted, non-base edited strand of the double-stranded nucleic acid molecule; This is because Cas9 nickase base pairs with the gRNA (e.g., sgRNA) bound to Cas9. In one embodiment, the Cas9 nickase cleaves the strand that is not bound to the H840A end. or a corresponding mutation, containing an aspartic acid residue at position 10. In some embodiments, the Cas9 nickase is any of the Cas9 nickases provided herein. or at least 60%, at least 65%, at least 70%, at least 75%, at least 80% , at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, Contains an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to Additional suitable Cas9 nickases may be identified based on this disclosure and knowledge in the art. It is obvious to one skilled in the art and within the scope of this disclosure.
[0236] [Cas9 domain with reduced PAM exclusivity] Typically, a Cas9 protein such as Cas9 from S. pyogenes (spCas9) requires a standard NGG PAM sequence to bind to a specific nucleic acid region, where "N" in "NGG" is adenine (A), thymidine (T), or cytosine (C), and G is guanosine. This can limit the ability to edit a desired base within the genome. In some embodiments, the base editing fusion proteins provided in this specification may need to be placed in a region containing the target base at an exact position, e.g., upstream of the PAM. See, for example, Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016) (the entire contents of which are incorporated herein by reference). Thus, in some embodiments, any of the fusion proteins provided herein may include a Cas9 domain that can bind to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences are described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). base at an exact position, e.g., upstream of the PAM. See, for example, Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016) (the entire contents of which are incorporated herein by reference). Thus, in some embodiments, any of the fusion proteins provided herein may include a Cas9 domain that can bind to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences are described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). ammable editing of a target base in genomic DNA without double-stranded DNA cleavage vage” Nature 533, 420-424 (2016) (the entire contents of which are incorporated herein by reference). Thus, in some embodiments, any of the fusion proteins provided herein may include a Cas9 domain that can bind to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences are described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). protein may include a Cas9 domain that can bind to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences are described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). ine, B. P., et al., “Engineered CRISPR-Cas9 nucleases with broad PAM compatibility and high specificity” Nature Biotechnology 32, 279-284 (2014). leases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins Tiver, B. P., et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2 015), the entire contents of which are incorporated herein by reference. Table 1 below lists several PAM variants.
[0237] Table 1: Cas9 Protein and Corresponding PAM Sequences
Table 1
[0238] In some embodiments, the Cas9 domain is the Cas9 domain from Staphylococcus aureus (SaCas9). In one aspect, the SaCas9 domain is nuclease-active SaCas9 , nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, SaCas9 comprises the N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0239] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a non-standard PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having the NNGRRT PAM sequence. In some embodiments, the SaCas9 domain , one or more of the E781X, N967X, and R1014X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the E781K, N96 7K, or R1014H mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0240] Exemplary SaCas9 sequence KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE N SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG The above residue N579, underlined and in bold, is mutated (e.g., to A579) to generate a SaCas9 nickase. Zero can be generated.
[0241] Exemplary SaCas9n sequence KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG The above residue A579 can be mutated from N579 to generate SaCas9 nickase, and is underlined and in bold characters as shown.
[0242] Exemplary SaKKH Cas9 KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNR K LINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFY K NDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPP H IIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. The above residue A579 can be mutated from N579 to generate SaCas9 nickase, and is indicated by underline and bold letters. The above residues K781, K967, and H1014 can be mutated from E781, N967, and R1014 to generate SaKKH Cas9, and are indicated by underline and italics.
[0243] In certain embodiments, the Cas9 domain is the Cas9 domain from Streptococcus pyogenes (SpCas9). In certain embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactive SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 has a D9X mutation, or an amino acid sequence provided herein comprising a corresponding mutation in any of the columns, where X is any amino acid other than D . In some embodiments, SpCas9 comprises the D9A mutation, or a corresponding mutation in any of the amino acid sequences provided herein . In one aspect, the SpCas9 domain , the SpCas9d domain or the SpCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM . In one aspect, the SpCas9 domain, the SpCas9d domain or the SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA or NGCG PAM sequence. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations , or a corresponding mutation in any of the amino acid sequences provided herein , where X is any amino acid. In some embodiments, the S pCas9 domain comprises one or more of the D1134E, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein . In some embodiments, the SpCas9 domain comprises the D1134E, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein . In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations, or a corresponding mutation in any of the amino acid sequences provided herein , where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein . In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations, or a corresponding mutation in any of the amino acid sequences provided herein , where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein It contains corresponding mutations in any of the resulting amino acid sequences. In some embodiments the SpCas9 domain contains the D1134V, R1334Q, and T1336R mutations, or corresponding mutations in any of the amino acid sequences provided herein It contains corresponding mutations in any of the resulting amino acid sequences. In some embodiments the SpCas9 domain contains one or more of the D1134X, G1217X, R1334X, and T1336X mutations or corresponding mutations in any of the amino acid sequences provided herein where X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D 1134V, G1217R, R1334Q, and T1336R mutations, or corresponding mutations in any of the amino acid sequences provided herein It contains corresponding mutations in any of the resulting amino acid sequences. In some embodiments the SpCas9 domain contains the D1134V, G1217R, R1334Q, and T1336R mutations, or corresponding mutations in any of the amino acid sequences provided herein It contains corresponding mutations in any of the amino acid sequences provided herein.
[0244] In some embodiments, any Cas9 domain of the fusion proteins provided herein contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 polypeptide described herein. In some embodiments any Cas9 domain of the fusion proteins provided herein contains the amino acid sequence of any Cas9 polypeptide described herein In some embodiments, the present specification Any Cas9 domain of the fusion proteins provided in the detailed description consists of the amino acid sequence of any Cas9 polypeptide described herein.
[0245] Exemplary SpCas9 DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0246] Exemplary SpCas9n DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0247] Exemplary SpEQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF E SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues E1134, Q1334, and R1336 are mutated from D1134, R1334, and T1336 and can give rise to SpEQR Cas9, which is underlined and in bold.
[0248] Exemplary SpVQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQ Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, Q1334, and R1336 are mutated from D1134, R1334, and T1336 to generate SpVQR Cas9, which is underlined and in bold.
[0249] Exemplary SpVRER Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, R1217, Q1334, and R1336 are D1134, G1217, R1334, and T1336 It can be mutated to produce SpVRER Cas9, which is indicated by underlining and boldface.
[0250] The Cas9 nuclease has two functional endonuclease domains, RuvC and HNH. When Cas9 binds to the target DNA, it undergoes a conformational change that positions the nuclease domains and cleaves the strand opposite the target DNA. The final result of DNA cleavage via Cas9 is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is repaired by one of two general repair pathways: (1) the error-prone but efficient non-homologous end joining (NHEJ) pathway; or (2) the homology-directed repair (HDR) pathway, which is less efficient but highly faithful.
[0251] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be calculated by any convenient method. For example, in some cases, efficiency can be expressed as the percentage of successful HDR. For example, nuclease assays for testing can be used to generate cleavage products, and percentages can be calculated using the ratio of products to substrate. For example, a nuclease enzyme for measurement that directly cleaves DNA containing a newly incorporated restriction sequence as a result of successful HDR can be used. The higher the proportion of substrate being cleaved, the higher the percentage of HDR (the more efficient HDR is). As an illustrative example, the percentage of HDR can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (e.g., (b + c) / (a + b + c), where "a" is the band intensity of the DNA substrate, "b" and and "c" are cleavage products.).
[0252] In some cases, efficiency can be represented by the success rate of NHEJ. For example, cleavage products are generated using a T7 endonuclease I assay, and the percentage of NHEJ can be calculated using the ratio of the product to the substrate. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJ produces small random insertions or deletions (indels) at the initial cleavage site). More cleavage indicates a higher proportion of NHEJ (higher efficiency of NHEJ). As an illustrative example, the proportion (percentage) of NHEJ can be calculated using the formula (1-(1-(b + c) / (a + b + c)) ×100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are cleavage products ( Ran et. al., 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308). The NHEJ repair pathway is the most active repair mechanism and frequently causes small nucleotide insertions or deletions (indels) at the DSB site. The randomness of NHEJ-mediated DSB repair has important practical implications because a cell population expressing Cas9 and gRNA or guide polynucleotide results in a diverse array of mutations. In most cases, NHEJ causes small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that introduce premature stop codons within the open reading frame (ORF) of the target gene. (1-(1-(b + c) / (a + b + c)) 1 / 2 ) ×100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are cleavage products ( Ran et. al., 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308). Ran et. al., 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308). ).
[0253] The NHEJ repair pathway is the most active repair mechanism and frequently causes small nucleotide insertions or deletions (indels) at the DSB site. The randomness of NHEJ-mediated DSB repair has important practical implications because a cell population expressing Cas9 and gRNA or guide polynucleotide results in a diverse array of mutations. In most cases, NHEJ causes small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that introduce premature stop codons within the open reading frame (ORF) of the target gene. ). The ideal final outcome is a loss-of-function mutation within the target gene.
[0254] NHEJ-mediated DSB repair often disrupts the gene's open reading frame, while homology-directed repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions such as the addition of a fluorophore or tag.
[0255] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to the target cell type along with the gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences immediately upstream and downstream of its target (referred to as left and right homology arms). The length of each homology arm can depend on the size of the change to be introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, double-stranded oligonucleotide, or double-stranded DNA plasmid and can be obtained. The efficiency of HDR is generally low (less than 10% modified alleles) even in cells expressing Cas9, gRNA, and the exogenous repair template. Since HDR occurs during the S and G2 phases of the cell cycle, the efficiency of HDR can be increased by synchronizing the cells. Chemically or genetically inhibiting genes involved in NHEJ can also increase the HDR frequency.
[0256] In some embodiments, Cas9 is a modified Cas9. A given gRNA target sequence can have additional sites of partial homology across the genome. These sites are called off-targets and need to be considered when designing the gRNA, although optimizing the gRNA design In addition, the specificity of CRISPR can be increased by modifying Cas9. Cas9 induces double-strand breaks (DS cleavage) through the combined activity of two nuclease domains, RuvC and HNH. B) The D10A mutant of SpCas9, Cas9 nickase, contains one nuclease domain. and generate DNA nicks rather than DSBs. HDR-mediated gene editing for specific gene editing Editing can also be combined with nickase.
[0257] In some cases, the Cas9 is a variant Cas9 protein. Variant Cas9 Polypeptides differs by a single amino acid compared to the amino acid sequence of the wild-type Cas9 protein (e.g. In some instances, the amino acid sequence may include a sequence having deletions, insertions, substitutions, and fusions. The modified Cas9 polypeptide contains an amino acid sequence that reduces the nuclease activity of the Cas9 polypeptide. For example, in some instances, The antCas9 polypeptides exhibited less than 50% of the nuclease activity of the corresponding wild-type Cas9 protein. less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. In this case, the variant Cas9 protein has no substantial nuclease activity. If the protein is a variant Cas9 protein that does not have substantial nuclease activity, This may be referred to as "dCas9."
[0258] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, A variant Cas9 protein may be a variant of a wild-type Cas9 protein (e.g., a wild-type Cas9 protein). Less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% nuclease activity is shown.
[0259] In some cases, the variant Cas9 protein can cleave the complementary strand of the guide-target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide-target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10) and thus can cleave the complementary strand of the double-stranded guide-target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence (thus, when this variant Cas9 protein cleaves a double-stranded target nucleic acid, a single-strand break (SSB) occurs instead of a double-strand break (DSB)) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).
[0260] In some cases, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded guide-target sequence, but has a reduced ability to cleave the complementary strand of the guide-target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine to alanine at amino acid position 840) mutation and thus the non-complementary strand of the guide-target sequence It can cut the DNA, but the ability to cut the complementary strand of the guide target sequence is reduced (therefore, when this variant Cas9 protein cuts the double-stranded guide target sequence, a single-strand break (SSB) occurs instead of a double-strand break (DSB)). Such a Cas9 protein has a reduced ability to cut the guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to the guide target sequence (e.g., a single-stranded guide target sequence).
[0261] In some cases, the variant Cas9 protein has a reduced ability to cut both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some cases, the variant Cas9 protein has both D10A and H840A mutations, and as a result, the polypeptide has a reduced ability to cut both the complementary and non-complementary strands of double-stranded target DNA. Such a Cas9 protein has a reduced ability to cut the target DNA (e.g., single-stranded target DNA), but retains the ability to bind to the target DNA (e.g., single-stranded target DNA).
[0262] As another non-limiting example, in some cases, the variant Cas9 protein has W4 76A and W1126A mutations, and as a result, the polypeptide has a reduced ability to cut the target DNA (e.g., single-stranded target DNA), but retains the ability to bind to the target DNA (e.g., single-stranded target DNA).
[0263] As another non-limiting example, in some cases, the variant Cas9 protein has P4 75A, W476A, N477A, D1125A, W1126A, and D1127A mutations, and as a result, the polypeptide The peptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target D NA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0264] As another non-limiting example, in some cases, the variant Cas9 protein has H8 40A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 has the catalytic His residue at position 840 of the Cas9 HNH domain restored (A840H). .
[0265] As another non-limiting example, in some cases, the variant Cas9 protein has H8 40A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide reduces the ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some In such cases, the variant Cas9 protein has D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). When the variant Cas9 protein has W476A and W1126A mutations, or when the variant Cas9 protein has P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to the PAM sequence. Thus, in such cases, using such a variant Cas9 protein in a binding method, this method does not require a PAM sequence. In other words, in certain cases, when using such a variant Cas9 protein in a binding method, this method may include guide RNA, but this method can be performed in the absence of a PAM sequence (thus, the binding specificity is provided by the target segment of the guide RNA). To achieve the above effects, other residues can be mutated (i.e., one or the other nuclease moiety is inactivated). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted).
[0266] In one aspect, a variant Cas9 protein having reduced catalytic activity (e.g., Ca The s9 protein has D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H98 2A, H983A, A984A, and / or D986A, it can bind to target DNA in a site-specific manner as long as it retains the ability to interact with the guide RNA (because it is guided to the target DNA sequence by the guide RNA).
[0267] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9 -VRER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas 9-LRVSQL.
[0268] As an alternative to S. pyogenes Cas9, RNA-guided endonucleases from the Cpf1 family that show cleavage activity in mammalian cells can be mentioned. CRISPR (CRISPR / Cpf1) derived from Prevotella and Francisella is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and Overcome some. Unlike Cas9 nuclease, the result of DNA cleavage via Cpf1 is a short Double-strand break with a 3' overhang. The alternating cleavage pattern of Cpf1 opens up the possibility of directional gene insertion similar to traditional restriction enzyme Cloning, which can enhance the efficiency of gene editing. Similar to the Cas9 variants and orthologs described above, Cpf1 Can also expand the number of sites that CRISPR can target to regions rich in AT lacking the NGG PAM site preferred by SpCas9 or to AT-rich genomes. The Cpf1 locus contains an α / β mixed domain , a following helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9 . Furthermore, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helix recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain composition indicates that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system . The Cpf1 locus encoded Cas1, Cas2, and Cas4 proteins more similar to type I and type III than to the type II system. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); thus, only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing . In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex has a motif Indicates that it is functionally unique and is classified as a class 2, type V CRISPR system . The Cpf1 locus encoded Cas1, Cas2, and Cas4 proteins more similar to type I and type III than to the type II system. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); thus, only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing . Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); thus, only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing . Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing . Since it has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), this is beneficial for genome editing . In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex has a motif Target DNA or RNA is cleaved by identifying a protospacer adjacent to the -ph5'-YTN-3' . After identification of the PAM, Cpf1 introduces a sticky end-like DNA double-strand break with a 4 or 5 nucleotide overhang. NA double-strand break.
[0269] [Fusion protein containing two napDNAbps and a deaminase domain] Some embodiments of the present disclosure provide a fusion protein comprising a nickase-active napDNAbp domain (e.g., nCas domain main), and a catalytically inactive napDNAbp (e.g., dCas domain), and a nucleic acid base e ditor (e.g., adenosine deaminase domain, cytidine deaminase domain), wherein at least the napDNAbp domain is linked by a linker . It should be understood that the Cas domain may be any of the Cas domains or Cas proteins provided herein (e.g., dCas9 and nCas9). In some embodiments, any of the Cas domain, DNA-binding protein domain, or Cas protein may be Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, including but not limited to these. An example of a programmable polynucleotide binding protein having a PAM specificity different from Cas9 is Clustered Regularly from Prevotella and Francisella 1 Interspaced Short Palindromic Repeats (Cpf1). Similar to Cas9, Cpf1 is also class 2 Interspaced Short Palindromic Repeats (Cpf1). Similar to Cas9, Cpf1 is also class 2 Interspaced Short Palindromic Repeats (Cpf1). Similar to Cas9, Cpf1 is also class 2 is a CRISPR effector. For example, but not limited to, in some embodiments the fusion protein comprises the following structure where the deaminase is adenosine deaminase or cytidine de aminase: NH2 - [Deaminase] - [nCas domain] - [dCas domain] - COOH; NH2 - [Deaminase] - [dCas domain] - [nCas domain] - COOH; NH2 - [nCas domain] - [dCas domain] - [Deaminase] - COOH; NH2 - [dCas domain] - [nCas domain] - [Deaminase] - COOH; NH2 - [nCas domain] - [Deaminase] - [dCas domain] - COOH; NH2 - [dCas domain] - [Deaminase] - [nCas domain] - COOH;
[0270] In some embodiments, the "-" used in the above general configuration indicates the presence of an optional linker. In some embodiments, the deaminase and napDNAbp ( e.g., Cas domain) are directly fused rather than being linked by a linker sequence. In certain embodiments, the linker is present between the deaminase domain and the napDNAbp. In certain embodiments, the deaminase or other nucleic acid base editor is directly fused to dCas, and the linker connects dCas and nCas9. In some embodiments, the deamin ase and napDNAbps are fused via any of the linkers provided herein. For example, in some embodiments, the deaminase and napDNAbp are fused via any of the linkers provided in the section entitled "Linker". In some In some embodiments, the dCas domain and the deaminase are immediately adjacent, and the nC as domain is linked to one of these domains (either 5' or 3') via a linker .
[0271] [Fusion protein with internal insertion] The present disclosure provides a fusion protein comprising a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding protein (e.g., napDNAbp). The heterologous polypeptide can be a polypeptide not found in the native or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to napDNAbp at the C-terminus of napDNAbp, or at the N-terminus of napDNAbp , or inserted at an internal position of napDNAbp. In certain embodiments, the heterologous polypeptide is inserted at an internal position of napDNAbp . In some embodiments, the heterologous polypeptide is a deaminase or a functional fragment thereof
[0272] . For example, the fusion protein can comprise a deaminase flanked by an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide . The deaminase in the fusion protein can be a cytidine de aminase. The deaminase in the fusion protein can be an adenosine deaminase . The deaminase can be a circular permutant deaminase. For example
[0273] , the deaminase can be a circular permutant adenosine deaminase or a circular permutant cytidine deamin ase. In some embodiments, the deaminase is in the TadA reference sequence . It is a circular permutated TadA that is circularly permutated at amino acid residue 116 with the numbering in the TadA reference sequence. In some embodiments, the deaminase is a circular permutated TadA that is circularly permutated at amino acid residue 136 with the numbering in the TadA reference sequence. In some embodiments the deaminase is a circular permutated TadA that is circularly permutated at amino acid residue 65 with the numbering in the TadA reference sequence.
[0274] The fusion protein can comprise multiple deaminases. The fusion protein can comprise, for example, 1, 2, 3, 4, 5 or more deaminases. In certain embodiments the fusion protein comprises one deaminase. In certain embodiments, the fusion protein comprises two deaminases. Two or more deaminases in the fusion protein can be adenosine deaminase, cytidine deaminase, or a combination thereof. Two or more deaminases can be homodimers. Two or more deaminases can be heterodimers. Two or more deaminases can be inserted tandemly into napDNAbp. In some embodiments, two or more deaminases may not be tandem in napDNAbp.
[0275] In certain embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In certain embodiments the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In certain embodiments the Cas9 polypeptide is a nuclease dead (dCas9) polypeptide or a fragment thereof. In certain embodiments, the Cas9 polypeptide is a nuclease inactive (dead) Cas9 It is a (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein can be the full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be the full-length Cas9 polypeptide. The Cas9 polypeptide can be, for example, truncated at the N-terminus or C-terminus relative to the native Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, portion, or domain of the Cas9 polypeptide that can still bind to the target polynucleotide and the guide nucleic acid sequence. In certain embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9),
[0276] Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant thereof. The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the native Cas9 polypeptide.
[0277] The Cas9 polypeptide of the fusion protein can be, for example, truncated at the N-terminus or C-terminus relative to the native Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the native Cas9 polypeptide.
[0278] The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 amino acid sequence shown in SEQ ID NO: 1. The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the native Cas9 polypeptide. The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the native Cas9 polypeptide. The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the native Cas9 polypeptide.
[0279] The Cas9 polypeptide of the fusion protein can include an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 amino acid sequence shown in SEQ ID NO: 1. At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least The amino acid sequence may be at least 99%, or at least 99.5%, identical to the amino acid sequence of the first amino acid.
[0280] The heterologous polypeptide (e.g., a deaminase) can be used to bind, for example, napDNAbp to a target polynucleotide. and napDNAbp (e.g., Cas9) at an appropriate position to maintain the ability to bind to the guide nucleic acid. ) can be inserted into the deaminase function (e.g., base editing activity) or napDNA The nucleotide sequence can be modified without compromising the function of the bp (e.g., its ability to bind to the target nucleic acid and the guide nucleic acid). Deaminases can be inserted into the napDNAbp. Deaminases can be used for example for crystallographic studies. The disordered regions or high temperature factors indicated by Or it may be inserted into napDNAbp in the region containing the B factor. The structure or structure of the unstructured, disordered or unstructured regions, such as solvent exposed regions and loops, may be determined by the Deaminases can be used to insert flexible loops without impairing function. In one embodiment, the napDNAbp may be inserted in a loop region or in a solvent-exposed region. The deaminase is inserted into a flexible loop of the Cas9 polypeptide.
[0281] In one embodiment, the insertion position of the deaminase is B of the crystal structure of the Cas9 polypeptide. In some embodiments, the deaminase is more efficient than the average Higher B-factors (e.g., compared to the total protein or to protein domains containing disordered regions) is inserted into a region of the Cas9 polypeptide that includes a higher B factor). The B factor or temperature factor can indicate the fluctuation from the average position of atoms (e.g., due to temperature-dependent atomic vibrations in a crystal lattice or as a result of static disorder). A high B factor (e.g., higher than the average B factor) for backbone atoms indicates a region with relatively high local mobility and can do so. Such regions can be used to insert deaminase without impairing the structure or function. The deaminase is 50%, 60%, 7 0%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more higher than the average B factor for the total protein and can be inserted at the position of a residue having a Cα atom with a B factor. The deamin ase can be inserted at the position of a residue having a Cα atom with a B factor that is 50%, 60%, 7 0%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more higher than the average B factor for the Cas9 protein domain containing the residue. A Cas9 polypeptide position containing a B factor higher than the average, for example, residues 76 8, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248 can be included as numbered in SEQ ID NO: 1. A Cas9 polypeptide region containing a B factor higher than the average, for example, residues 792-872, 792-906, and 2-791 can be included as numbered in SEQ ID NO: 1. A heterologous polypeptide (e.g., deaminase) is 768, 791 as numbered in SEQ ID NO: 1 in the numbering of SEQ ID NO: 1 in the numbering of SEQ ID NO: 1 can include residues 792-872, 792-906, and 2-791.
[0282] The heterologous polypeptide (e.g., deaminase) can be 768, 791 as numbered in SEQ ID NO: 1 , 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1 An amino acid residue selected from the group consisting of 246, 1247, and 1248, or the corresponding amino acid residue in another Cas9 polypeptide Can be inserted into napDNAbp at the corresponding amino acid residue in the peptide. In some embodiments In the form, the heterologous polypeptide is at amino acid positions 768 - 769 numbered in SEQ ID NO: 1 , 791 - 792, 792 - 793, 1015 - 1016, 1022 - 1023, 1026 - 1027, 1029 - 1030, 1040 - 1041, 1052 - 1053, 1054 - 1055, 1067 - 1068, 1068 - 1069, 1247 - 1248, or between 1248 - 1249, or between the corresponding amino acid positions Thereof. In some embodiments, the heterologous polypeptide The peptide is at amino acid positions 769 - 770, 792 - 793, 793 - 794, 1016 - 1017, 1023 - 1024, 1027 - 1028, 1030 - 1031, 1041 - 1042, 1053 - 1054, 1055 - 1056, 1068 - 10 69, 1069 - 1070, 1248 - 1249, or between 1249 - 1250, or between the corresponding amino acid positions Thereof. In some embodiments, the heterologous polypeptide is at SEQ ID NO: 1 The numbered 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 105 4, 1067, 1068, 1069, 1246, 1247, and an amino acid residue selected from the group consisting of 1248 , or replace the corresponding amino acid residue in another Cas9 polypeptide. It should be understood that referring to SEQ ID NO: 1 with respect to the insertion position is for illustrative purposes only. Discussed here It should be understood that referring to SEQ ID NO: 1 with respect to the insertion position is for illustrative purposes only. Discussed here The insertion is not limited to the Cas9 polypeptide sequence of SEQ ID NO: 1, and for example, Cas9 nickase (n Cas9), nuclease-inactivated Cas9 (dCas9), Cas9 variants lacking a nuclease domain , truncated Cas9, or Cas9 domains lacking a partial or complete HNH domain, etc. also includes insertions at corresponding positions in variant Cas9 polypeptides.
[0283] The heterologous polypeptide (e.g., deaminase) is numbered in SEQ ID NO: 1, and amino acid residues selected from the group consisting of 768, 79 2, 1022, 1026, 1040, 1068, and 1247, or at the corresponding amino acid residues in another Cas9 polypeptide, can be inserted into napDNAbp. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1068-1069 in SEQ ID NO: 1, or between the corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 793-794, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1069-1070 in SEQ ID NO: 1, or between the corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of 768, 792, 1022, 1 026, 1040, 1068, and 1247 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 793-794, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1069-1070, or between 1248-1249, or between the corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of 768, 792, 1022, 1 026, 1040, 1068, and 1247 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of 768, 792, 1022, 1026, 1040, 1068, and 1247 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9
[0284] Heterologous polypeptides (e.g., deaminases) can be inserted at the amino acid residues shown in Figures 4, 5, 6, or 7, or at the corresponding amino acid residues in another Cas9 polypeptide at napDNA bp. Heterologous polypeptides (e.g., deaminases) can be inserted at the amino acid residues selected from the group consisting of numbered 1002, 1003, 1025, 1052 - 1056, 1242 - 1247, 1061 - 1077, 943 - 947, 686 - 691, 569 - 578, 530 - 539, and 1060 - 1077 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide at napDNA bp. The deaminase can be inserted at the N-terminus or C-terminus of the residue or can replace the residue. In some embodiments, the deaminase is inserted at the C-terminus of the residue. In the amino acid residues, or in the corresponding amino acid residues in another Cas9 polypeptide, napDNA bp can be inserted. Heterologous polypeptides (e.g., deaminases) can be inserted at the amino acid residues numbered 1002, 1003, 1025, 1052 - 1056, 1242 - 1247, 1061 - 1077, 943 - 947, 686 - 691, 569 - 578, 530 - 539, and selected from the group consisting of 1060 - 1077 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Ca s9 polypeptide at napDNA bp. The deaminase can be inserted at the N-terminus or C-terminus of the residue or can replace the residue. In some embodiments, the deaminase is inserted at the C-terminus of the residue. In some embodiments, ABE (e.g., TadA) is inserted at the amino acid residues selected from the group consisting of numbered 10 15, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 in SEQ ID NO: 1, or at the corresponding .
[0285] In some embodiments, ABE (e.g., TadA) is inserted at the amino acid residues selected from the group consisting of numbered 10 15, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE (e.g., TadA) is replaced at the numbered residues 792 - 872, 792 - 906, or 2 - 791 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide and inserted. In some embodiments, ABE is selected from the group consisting of numbered 1015, 1022, 1029, 10 40, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 in SEQ ID NO: 1 and inserted. is inserted at the N-terminus of the amino acid to be introduced, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at amino acids numbered 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 in SEQ ID NO: 1, or at the C-terminus of the corresponding amino acid residue of another Cas9 polypeptide, selected from the group consisting of In some embodiments, ABE is inserted by replacing the amino acid numbered 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, CBE (e.g., APOBEC1) is inserted at amino acid residues selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid selected from the group consisting of
[0286] 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted with the amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted with the amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 in SEQ ID NO: 1, or at the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, ABE is inserted with the amino acid selected from the group consisting of It is inserted at the C-terminus of the base. In some embodiments, ABE replaces or inserts an amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247 at the numbering in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide.
[0287] In some embodiments, the deaminase is inserted at amino acid residue 768 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 768 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 768 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE replaces or inserts the corresponding amino acid residue of amino acid 768 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide.
[0288] In some embodiments, the deaminase is inserted at amino acid residue 791 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 791 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 791 at the numbering in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, In this case, ABE is inserted by replacing the amino acid 791 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. is inserted by replacing the corresponding amino acid residue.
[0289] In some embodiments, the deaminase is inserted at amino acid residue 792 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. is inserted at the corresponding amino acid residue. In some embodiments, ABE is inserted at the N-terminus of amino acid 792 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 792 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 792 numbered in SEQ ID NO: 1, or of another Cas 9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1016 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. is inserted by replacing the corresponding amino acid residue.
[0290] In some embodiments, the deaminase is inserted at amino acid residue 1016 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. is inserted at the corresponding amino acid residue. In some embodiments, ABE is inserted at the N-terminus of amino acid 1016 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1016 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1016 numbered in SEQ ID NO: 1, or of another Ca s9 polypeptide. In some embodiments, ABE is inserted at amino acid 1016 numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. is inserted by replacing the corresponding amino acid residue.
[0291] In some embodiments, the deaminase is inserted at amino acid residue 1022 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1022 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some In some embodiments, ABE is inserted at the C-terminus of amino acid 1022 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1022 in the numbering of SEQ ID NO: 1, or of another Cas9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some In some embodiments, ABE is inserted at the C-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1023 in the numbering of SEQ ID NO: 1, or of another Cas9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some
[0292] In some embodiments, ABE is inserted at the C-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1023 in the numbering of SEQ ID NO: 1, or of another Cas9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some In some embodiments, ABE is inserted at the C-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1023 in the numbering of SEQ ID NO: 1, or of another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1023 in the numbering of SEQ ID NO: 1, or of another Cas9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some
[0293] In some embodiments, the deaminase is inserted at amino acid residue 1023 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. It is inserted at the acid residue 1026 or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1026 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1026 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1026 numbered in SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, the deaminase is inserted at the acid residue 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide.
[0294] In some embodiments, the deaminase is inserted at the acid residue 1040 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. It is inserted at the acid residue 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1029 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1029 numbered in SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, the deaminase is inserted at the acid residue 1040 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1040 numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1040 numbered in SEQ ID NO: 1, or another Cas9 polypeptide.
[0295] In some embodiments, the deaminase is inserted at the acid residue 1040 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. It is inserted at the acid residue 1040 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1040 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1040 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1040 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, the deaminase is inserted at the amino acid residue 1052 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1052 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9
[0296] polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1052 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1052 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, the deaminase is inserted at the amino acid residue 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of the amino acid 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of the amino acid 1054 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide.
[0297] In some embodiments, the deaminase is inserted at the amino acid residue 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the amino acid 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of the amino acid 1054 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue in another Cas9 is inserted at the N-terminus of the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1054 as numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments ABE is inserted by replacing the corresponding amino acid residue of amino acid 1054 as numbered in SEQ ID NO: 1, or another Cas9 polypeptide.
[0298] In some embodiments, the deaminase is inserted at amino acid residue 1067 as numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1067 as numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1067 as numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments ABE is inserted by replacing the corresponding amino acid residue of amino acid 1067 as numbered in SEQ ID NO: 1, or another Cas9 polypeptide.
[0299] In some embodiments, the deaminase is inserted at amino acid residue 1068 as numbered in SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1068 as numbered in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some In some embodiments, ABE is inserted at the C-terminus of amino acid 1068 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1068 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, the deaminase is inserted at amino acid residue 1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1069 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide.
[0300] In some embodiments, the deaminase is inserted at amino acid residue 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1069 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted by replacing the corresponding amino acid residue of amino acid 1069 in the numbering of SEQ ID NO: 1, or another Cas9 polypeptide.
[0301] In some embodiments, the deaminase is inserted at amino acid residue 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, ABE is inserted at the N-terminus of amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, ABE is inserted at the C-terminus of amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. It is inserted at the C-terminus of the corresponding amino acid residue in the s9 polypeptide. In some embodiments the ABE is inserted by replacing the amino acid 1246 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide.
[0302] In some embodiments, the deaminase is inserted at the amino acid residue 1247 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the N-terminus of the amino acid 1247 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the C-terminus of the amino acid 1247 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide by replacing it. In some embodiments, the ABE is inserted at the C-terminus of the amino acid 1247 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. It is inserted at the C-terminus of the corresponding amino acid residue in the s9 polypeptide. In some embodiments the ABE is inserted by replacing the amino acid 1247 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide.
[0303] In some embodiments, the deaminase is inserted at the amino acid residue 1248 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the N-terminus of the amino acid 1248 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the C-terminus of the amino acid 1248 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the C-terminus of the amino acid 1248 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the ABE is inserted at the C-terminus of the amino acid 1248 in the numbering of SEQ ID NO: 1, or at the corresponding amino acid residue of another Cas9 polypeptide. It is inserted at the C-terminus of the corresponding amino acid residue in the s9 polypeptide. In some embodiments In one state, ABE replaces and inserts the amino acid 1248 numbered in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. It is inserted by replacing the corresponding amino acid residue of the peptide.
[0304] In certain embodiments, a heterologous polypeptide (e.g., a deaminase) is inserted into a flexible loop of the Cas9 polypeptide. The flexible loop portion is numbered 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. It can be selected from the group consisting of. The flexible loop portion can be selected from the group consisting of amino acid residues numbered 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297 in SEQ ID NO: 1, or the corresponding amino acid residue in another Cas9 polypeptide. -537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1 300, or the corresponding amino acid residue in another Cas9 polypeptide. The flexible loop portion can be selected from the group consisting of amino acid residues numbered 1-529, 538-568, 580-6 85, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297, or another Cas9 It can be selected from the group consisting of the corresponding amino acid residues in the polypeptide.
[0305] A heterologous polypeptide (e.g., a deaminase) corresponds to the amino acid residues numbered 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077 in SEQ ID NO: 1, or the corresponding amino acid residue of another Cas9 polypeptide. It can be inserted into the Cas9 polypeptide region. -579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077, or It can be inserted into the Cas9 polypeptide region corresponding to the corresponding amino acid residue of another Cas9 polypeptide. It can be inserted.
[0306] A heterologous polypeptide (e.g., a deaminase) is inserted instead of the deletion region of the Cas9 polypeptide. can be introduced. The deletion region may correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. In some embodiments, the deletion region corresponds to amino acid residues 792-872 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deletion region corresponds to amino acid residues 792-906 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deletion region corresponds to amino acid residues 2-791 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deletion region corresponds to amino acid residues 1017-1069 in the numbering of SEQ ID NO: 1, or the corresponding amino acid residues.
[0307] A heterologous polypeptide (e.g., deaminase) can be inserted within the structural or functional domain of the Cas9 polypeptide. A heterologous polypeptide (e.g., deaminase) can be inserted between two structural or functional domains of the Cas9 polypeptide. A heterologous polypeptide (e.g., deaminase) can be inserted instead of a structural or functional domain of the Cas9 polypeptide, for example, after deleting the domain from the Cas9 polypeptide. The structural or functional domains of the Cas9 polypeptide can include, for example, RuvC I, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH.
[0308] In some embodiments, the Cas9 polypeptide lacks one or more domains s...
Claims
1. 1. A fusion protein comprising a TadA variant inserted within a flexible loop of a Cas9 polypeptide, wherein said Cas9 polypeptide is a nickase or is nuclease inactive, and said flexible loop comprises a region selected from the group consisting of amino acid residues 768-793, 1002-1040, 1052-1077, and 1232-1248, numbered in SEQ ID NO: 1, or a region corresponding thereto; the TadA variant of the fusion protein deaminates a target nucleobase in a target DNA molecule, resulting in reduced deamination at non-target sites compared to an end-terminal fusion protein comprising the TadA variant fused to the N-terminus or C-terminus of a Cas9 polypeptide. Fusion proteins.
2. The fusion protein comprises the following structure: NH 2 -[N-terminal fragment of Cas9]-[TadA variant]-[C-terminal fragment of Cas9]-COOH The fusion protein of claim 1, wherein each symbol "]-[" is an optional linker.
3. 2. The fusion protein of Claim 1, wherein the C-terminus of said N-terminal fragment or the N-terminus of said C-terminal fragment comprises a portion of a flexible loop of said Cas9 polypeptide.
4. 4. The fusion protein of claim 1, wherein the flexible loop comprises an amino acid that is adjacent to the target nucleobase when the fusion protein deaminates the target nucleobase.
5. 5. The fusion protein of Claim 4, wherein said flexible loop comprises a portion of an alpha helix structure of a Cas9 polypeptide.
6. (b) the TadA variant is inserted at amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1052-1053, 1054-1055, 1067-1068, 1068-1069, or 1247-1248 in the numbering of SEQ ID NO: 1, or at the amino acid positions corresponding thereto; or (c) the TadA variant is inserted at amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1068-1069, or 1247-1248 in the numbering of SEQ ID NO: 1, or at the amino acid positions corresponding thereto; or (d) the TadA variant is inserted at amino acid positions 1016-1017, 1023-1024, 1029-1030, 1069-1070 or 1247-1248 in the numbering of SEQ ID NO: 1, or at amino acid positions corresponding thereto; The fusion protein according to any one of claims 1 to 5.
7. the N-terminal fragment comprises amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, and / or 1078-1231 of the Cas9 polypeptide, as numbered in SEQ ID NO: 1, or residues corresponding thereto; and / or the C-terminal fragment comprises amino acid residues 1301-1368, 1248-1297, 1078-1231, 1026-1051, and / or 948-1001 of the Cas9 polypeptide, as numbered in SEQ ID NO: 1; and / or an N-terminal or C-terminal fragment of a Cas9 polypeptide comprising a DNA-binding domain; and / or an N-terminal or C-terminal fragment of a Cas9 polypeptide comprising a RuvC domain; and / or an N-terminal or C-terminal fragment of a Cas9 polypeptide comprising an HNH domain; and / or Neither the N-terminal nor the C-terminal fragment of the Cas9 polypeptide contains an HNH domain; and / or Neither the N-terminal nor the C-terminal fragment of the Cas9 polypeptide contains the RuvC domain. The fusion protein according to any one of claims 1 to 6.
8. the Cas9 polypeptide comprises a partial or complete deletion in one or more structural domains; and / or the TadA variant is inserted into the Cas9 polypeptide at the position of the partial or complete deletion; or the deletion is within the RuvC domain, or the deletion is within the HNH domain, or The deletion bridges the RuvC domain and the C-terminal domain, the LI domain and the HNH domain, or the RuvC domain and the LI domain; A fusion protein according to any one of claims 1 to 7.
9. the Cas9 polypeptide comprises a deletion of amino acids 1017 to 1069, or the amino acids corresponding thereto, as numbered in SEQ ID NO: 1; or the Cas9 polypeptide comprises a deletion of amino acids 792 to 872, or the amino acids corresponding thereto, as numbered in SEQ ID NO: 1; or the Cas9 polypeptide comprises a deletion of amino acids 792 to 906, or the amino acids corresponding thereto, as numbered in SEQ ID NO: 1; The fusion protein of claim 8.
10. 1. A fusion protein comprising a TadA variant inserted within a Cas9 polypeptide, comprising the following structure: NH 2 -[Cas9 N-terminus fragment]-[TadA variant]-[C-terminal fragment of Cas9]-COOH Here, each of the symbols "]-[" is an optional linker, the Cas9 polypeptide comprises a partial or complete deletion of the HNH domain; A fusion protein in which the TadA variant is inserted in place of the partial or complete deletion.
11. The fusion protein of claim 10, wherein the C-terminal amino acid of the N-terminal fragment is amino acid 791, 907, or 873 according to the numbering in SEQ ID NO:
1.
12. 1. A fusion protein comprising a TadA variant inserted within a Cas9 polypeptide, comprising the following structure: NH 2 -[Cas9 N-terminus fragment]-[TadA variant]-[C-terminal fragment of Cas9]-COOH Here, each symbol "]-[" is an optional linker, Cas9 contains a partial or complete deletion of the RuvC domain, and the TadA variant is inserted in the position of the partial or complete deletion.
13. The fusion protein of claim 12 further comprising a UGI domain.
14. The optional linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n , (EAAAK) n , (GGS) n , SGSETPGTSESATPES, or (XP) n motif or combinations thereof, where n is independently an integer from 1 to 30; or an N-terminal fragment of a Cas9 polypeptide fused to said TadA variant without a linker; or a C-terminal fragment of Cas9 fused to the TadA variant without a linker; A fusion protein according to claim 12 or 13.
15. The fusion protein of any one of claims 1 to 14, further comprising a nuclear localization signal.
16. The fusion protein of any one of claims 1 to 15, wherein the fusion protein forms a complex with a guide nucleic acid sequence to effect deamination of a target nucleobase.
17. A polynucleotide encoding the fusion protein of any one of claims 1 to 16.
18. An expression vector comprising the polynucleotide of claim 17.
19. the expression vector is a mammalian expression vector; and / or the vector is a viral vector selected from the group consisting of an adeno-associated virus (AAV), a retroviral vector, an adenoviral vector, a lentiviral vector, a Sendai virus vector, and a herpes virus vector; and / or the vector comprises a promoter, 19. The expression vector of claim 18.
20. A cell comprising the fusion protein of any one of claims 1 to 16, the polynucleotide of claim 17, or the vector of claim 18 or 19.
21. A kit comprising the fusion protein of any one of claims 1 to 16, the polynucleotide of claim 17, or the vector of claim 18 or 19.
22. 17. An in vitro or ex vivo base editing method comprising contacting a target DNA molecule with a fusion protein of any one of claims 1 to 16, wherein the TadA variant of the fusion protein deaminates a nucleobase in the target DNA molecule, thereby editing the target DNA molecule.
23. 23. The method of claim 22, wherein the target DNA molecule is contacted with a guide nucleic acid sequence to effect deamination of the target nucleobase.
24. 1. An in vitro or ex vivo method for editing a target nucleobase in a target DNA molecule, comprising contacting the target DNA molecule with a fusion protein comprising a TadA variant flanked by N-terminal and C-terminal fragments of a Cas9 polypeptide, wherein the Cas9 polypeptide is a nickase or is nuclease-inactive, wherein the TadA variant of the fusion protein deaminates the target nucleobase in the target DNA molecule, and wherein the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a portion of a flexible loop of a Cas9 polypeptide, wherein the flexible loop comprises a region selected from the group consisting of amino acid residues 768-793, 1002-1040, 1052-1077, and 1232-1248, as numbered in SEQ ID NO: 1, or a region corresponding thereto, and wherein the fusion protein results in reduced deamination at non-target sites compared to an end-terminal fusion protein comprising the TadA variant fused to the N-terminus or C-terminus of a Cas9 polypeptide.
25. 1. An in vitro or ex vivo method for editing a target nucleobase in a target DNA molecule, comprising contacting the target DNA molecule with a fusion protein comprising a TadA variant inserted within a flexible loop of a Cas9 polypeptide, wherein the Cas9 polypeptide is a nickase or is nuclease inactive, and wherein the fusion protein comprises the following structure: NH 2 -[Cas9 N-terminus fragment]-[TadA variant]-[C-terminal fragment of Cas9]-COOH wherein each "]-[" represents an optional linker, the TadA variant of the fusion protein deaminates a target nucleobase in a target DNA molecule, the flexible loop comprises a region selected from the group consisting of amino acid residues 768-793, 1002-1040, 1052-1077, and 1232-1248, as numbered in SEQ ID NO: 1, or a region corresponding thereto, and the fusion protein results in reduced deamination at non-target sites compared to an end-terminal fusion protein comprising the TadA variant fused to the N-terminus or C-terminus of a Cas9 polypeptide.
26. The method of any one of claims 22 to 25, wherein the contacting is performed intracellularly.
27. 1. A pharmaceutical composition for treating a genetic condition in a subject, the pharmaceutical composition comprising: a fusion protein comprising a TadA variant flanked by an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide, or a polynucleotide encoding said fusion protein; and a guide nucleic acid sequence or a polynucleotide encoding said guide nucleic acid sequence, wherein said guide nucleic acid sequence directs said fusion protein to deaminate a target nucleobase in a target DNA molecule of said subject, thereby treating said genetic condition; wherein said Cas9 polypeptide comprises a partial or complete deletion of an HNH domain or a RuvC domain, and wherein said TadA variant is inserted in place of the partial or complete deletion.
28. 1. A pharmaceutical composition for treating a genetic condition in a subject, the composition comprising a fusion protein comprising a TadA variant inserted within a flexible loop of a Cas9 polypeptide, wherein the Cas9 polypeptide is a nickase or is nuclease inactive, wherein the fusion protein comprises the following structure: NH 2 -[Cas9 N-terminus fragment]-[TadA variant]-[C-terminal fragment of Cas9]-COOH Here, each of the symbols "]-[" is an optional linker, 10. The pharmaceutical composition of claim 1, wherein the TadA variant of the fusion protein deaminates a target nucleobase in a target DNA molecule of a subject, thereby treating the genetic condition, and the flexible loop comprises a region selected from the group consisting of amino acid residues 768-793, 1002-1040, 1052-1077, and 1232-1248, as numbered in SEQ ID NO: 1, or a region corresponding thereto, and wherein the fusion protein results in reduced deamination at non-target sites compared to an end-terminal fusion protein comprising the TadA variant fused to the N-terminus or C-terminus of a Cas9 polypeptide.
29. the target nucleobase comprises a mutation associated with the genetic condition; deamination of said target nucleobase replaces said target nucleobase with a wild-type nucleobase; or deamination of the target nucleobase replaces the target nucleobase with a non-wild-type nucleobase; that deamination of said target nucleobase ameliorates the symptoms of said genetic condition.
29. The pharmaceutical composition of claim 27 or 28.
30. the target DNA molecule comprises a mutation associated with the genetic condition at a nucleobase other than the target nucleobase; and / or deamination of said target nucleobase ameliorates the symptoms of said genetic condition; 29. The pharmaceutical composition of claim 27 or 28.
31. A protein library for optimized base editing comprising a plurality of fusion proteins, each of the plurality of fusion proteins comprising a TadA variant flanked by N-terminal and C-terminal fragments of a Cas9 polypeptide, wherein the N-terminal fragment of each of the fusion proteins is different from the N-terminal fragment of the remaining fusion proteins of the plurality, or the C-terminal fragment of each of the fusion proteins is different from the C-terminal fragment of the remaining fusion proteins of the plurality, wherein each TadA variant of the fusion proteins deaminates a target nucleobase adjacent to a protospacer adjacent motif (PAM) sequence in a target DNA molecule, and the N-terminal fragment or the C-terminal fragment binds to the target DNA molecule; (a) each of the TadA variants is inserted within a flexible loop of the Cas9 polypeptide, the flexible loop comprising a region selected from the group consisting of amino acid residues 768-793, 1002-1040, 1052-1077, and 1232-1248, as numbered in SEQ ID NO: 1, or a region corresponding thereto; and / or (b) the Cas9 polypeptide comprises a deletion in the HNH domain or the RuvC domain, wherein the HNH domain comprises the amino acid positions shown in bold plain text in the amino acid sequence below, and the RuvC domain comprises the amino acid positions shown in bold italic text in the amino acid sequence below: wherein the TadA variant is inserted at the position of the deletion, at least one of the plurality of fusion proteins deaminates the target nucleobase with less off-target deamination compared to an end-terminal fusion protein comprising the TadA variant fused to the N-terminus or C-terminus of SEQ ID NO: 1; Protein library.
32. 32. The protein library of claim 31 , wherein for each nucleobase that is 1 to 20 nucleobases away from the PAM sequence, at least one of the plurality of fusion proteins deaminates that nucleobase.
33. the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a variant thereof; or the Cas9 polypeptide is a modified Cas9 and has specificity for a modified protospacer adjacent motif (PAM); 33. A protein library according to claim 31 or 32.
34. 34. The protein library of any one of claims 31 to 33, wherein the Cas9 polypeptide is a nickase or an inactive nuclease.