Composition and method for delivering nucleobase editing system

JP2025028857A5Active Publication Date: 2025-10-09BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024192680
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-13
Filing Date
2024-11-01
Publication Date
2025-10-09
Estimated Expiration
2039-09-07

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a composition and a method for helping achievement of delivery of CRISPR / Cas9 component to a cell.SOLUTION: Provided are a composition and a method for delivering first and second polynucleotide severally encoding fragments of A→G base editor fusion protein containing one or more deaminase (e.g., adenosine deaminase) and nCas9, where the first polynucleotide encodes an N-terminal fragment of nCas9 fused to intein -N of a divided intein pair, and the second polynucleotide encodes a C-terminal fragment of nCas9 fused to intein -C of a divided intein pair. A method for delivering (e.g., AAV delivery) these fragments together with sgRNA cell is also provided, where these fragments are connected together by a divided intein system, by which a functional base editing system is reconstituted in a cell.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Patent Application No. 62 / 728,703, filed September 7, 2018, and U.S. Provisional Patent Application No. 62 / 728,703, filed December 13, 2018. This application claims the benefit of U.S. Provisional Patent Application No. 62 / 779,404, filed on 2006-01-16, all of which are hereby incorporated by reference. No. 6,239,999, the entire contents of which are incorporated herein by reference.

[0002] The discovery of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) It has revolutionized the field of molecular biology. Much of this enthusiasm stems from the fact that CRISPR / Cas9 is helping to reverse human disease. It focuses on the clinical potential of treating and editing the human genome to identify disease-causing mutations. These genes could potentially be repaired using CRISPR or CRISPR-based systems. One challenge to achieve is delivering the elements necessary for genome editing. For example, with respect to CRISPR / Cas9, SpCas9 and sgRNA are encoded in a DNA plasmid vector. It can be delivered via adeno-associated virus (AAV). However, AAV is packaged their small binding capacity helps achieve delivery of CRISPR / Cas9 components into cells and / or or other elements (e.g., polypeptide domains) to fulfill the desired gene editing purpose. DNA templates for promoters, reporters, fluorescent tags, multiple sgRNAs, or HDR It is difficult to get the government to include other entities (such as the government) in the system. Summary of the Invention

[0003] In some embodiments, (a) a fusion protein comprising a deaminase and an N-terminal fragment of Cas9. a first polynucleotide encoding a polypeptide, wherein the N-terminal fragment of Cas9 is a polypeptide comprising the N-terminal fragment of Cas9 A continuous sequence beginning at the end of and ending at positions A292 to G364 of Cas9 in the numbering in SEQ ID NO: 2 and the N-terminal fragment of Cas9 is fused to a split intein-N. and (b) a second polynucleotide encoding a C-terminal fragment of Cas9, The C-terminal fragment of Cas9 begins at positions A292 to G364 of Cas9 according to the numbering in SEQ ID NO: 2. It is a continuous sequence that terminates at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused to the split intein-C. Provided herein are compositions comprising a second polynucleotide comprising:

[0004] In some embodiments, (a) a first polynucleotide encoding an N-terminal fragment of Cas9 wherein the N-terminal fragment of Cas9 begins at the N-terminus of Cas9 and is a sequence corresponding to the numbered fragments in SEQ ID NO:2. The N-terminal fragment of Cas9 is a continuous sequence that terminates at positions A292 to G364 of Cas9. (b) a first polynucleotide fused to the in-N and the C-terminal fragment of Cas9 and deamina a second polynucleotide encoding a fusion protein comprising a Cas9 enzyme; and The C-terminal fragment of Cas9 begins at positions A292 to G364 of Cas9 in the numbering of SEQ ID NO: 2 and ends at the C The second sequence is a continuous sequence terminating at the end, where the C-terminal fragment of Cas9 is fused to the split intein-C. Provided herein are compositions comprising a polynucleotide of

[0005] In some embodiments, (a) a fusion protein comprising a deaminase and an N-terminal fragment of Cas9. a first polynucleotide encoding a polypeptide, wherein the N-terminal fragment of Cas9 is a polypeptide comprising the N-terminal fragment of Cas9 A292 to G364, F445 to K483, or E565 of Cas9 in the numbering of SEQ ID NO: 2, starting from the end The N-terminal fragment of Cas9 is fused to the split intein-N. (b) a first polynucleotide encoding a C-terminal fragment of Cas9; and (b) a second polynucleotide encoding a C-terminal fragment of Cas9. wherein the C-terminal fragment of Cas9 is A292 of Cas9 according to the numbering in SEQ ID NO: 2. Consecutive sequences beginning at positions ~G364, F445~K483, or E565~T637 and terminating at the C-terminus of Cas9 and the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted for Ala, Ser, or Thr; A composition comprising a second polynucleotide, wherein the C-terminal fragment of Cas9 is fused to a split intein-C.

[0006] An article is provided herein.

[0006] In some embodiments, (a) a first polynucleotide encoding an N-terminal fragment of Cas9 wherein the N-terminal fragment of Cas9 begins at the N-terminus of Cas9 and is a sequence corresponding to the numbered fragments in SEQ ID NO:2. A continuous sequence that terminates at positions A292 to G364, F445 to K483, or E565 to T637 of Cas9. , a first polynucleotide in which an N-terminal fragment of Cas9 is fused to a split intein-N; and b) A second polynucleotide encoding a fusion protein containing the C-terminal fragment of Cas9 and the deaminase. wherein the C-terminal fragment of Cas9 is A292 of Cas9 according to the numbering in SEQ ID NO: 2. Consecutive sequences starting at positions ~G364, F445~K483, or E565~T637 and terminating at the C-terminus of Cas9 and the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted for Ala, Ser, or Thr; A composition comprising a second polynucleotide, wherein the C-terminal fragment of Cas9 is fused to a split intein-C.

[0006] An article is provided herein.

[0007] In some embodiments, the N-terminal fragment of Cas9 is an amino acid sequence according to the numbering in SEQ ID NO:2. Acids 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 576, 588, or 589 In some embodiments, the C-terminal fragment of Cas9 or the N-terminal fragment of Cas9 comprises SEQ ID NO: Amino acids S303, T310, T313, S355, A456, S460, A463, T466, and S46 in the numbering in No. 2 9, Ala / Cys, Ser / Cy at residues corresponding to T472, T474, C574, S577, A589, or S590 In some embodiments, the composition comprises a single guide nucleotide sequence, a ... In some embodiments, the sgRNA further comprises a polynucleotide encoding the sgRNA. In some embodiments, the first and second polynucleotides are linked. and the second polynucleotide are separately expressed. In some embodiments, the deaminase is a wild-type TadA or Ta In some embodiments, the deaminase is a TadA dimer. In some embodiments, the TadA dimer comprises wild-type TadA and TadA 7.10. The fusion protein comprises a nuclear localization signal (NLS). In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is linked to an NLS. In some embodiments, the NLS is a bipartite NLS. The fusion protein is linked to a base editor containing a deaminase and SpCas9. In some embodiments, the C-terminal fragment of Cas9 and the fusion protein are linked together. These are then combined to form deaminases and base editor proteins, including SpCas9. In some embodiments, SpCas9 has nickase activity or is catalytically inactive.

[0008] In some embodiments, the fusion protein disclosed herein and an N-terminal fragment of Cas9 Provided herein are compositions comprising: Compositions comprising a fusion protein and a C-terminal fragment of Cas9 are provided herein. In embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 and the deaminase are linked by a linker. In some embodiments, the linker is a peptide linker.

[0009] In some embodiments, the nucleic acid sequence comprises the first and second polynucleotides disclosed herein. In one embodiment, the vector comprises a promoter. In some embodiments, the promoter is a constitutive promoter. The constitutive promoter is a CMV or CAG promoter. , retroviral vectors, adenoviral vectors, lentiviral vectors, herpes The vector is selected from the group consisting of an adenovirus vector and an adeno-associated virus vector. In some embodiments, the vector is an adeno-associated virus vector.

[0010] In some embodiments, the compositions disclosed herein or the vectors disclosed herein are Provided herein is a cell comprising a vector. In some embodiments, the cell is a mammalian cell. be.

[0011] In some embodiments, Ala / Cys, Ser / Cys, or Thr / Cys mutations are used herein.

[0023] A reconstituted A to G base editor protein is provided, comprising a Cas9 domain comprising: In some embodiments, the mutations are at amino acids S303, T310, T313, S355, Compatible with A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, or S590 in the residues.

[0012] In some embodiments, (a) a sequence beginning at the N-terminus of Cas9 and continuing in the sequence numbering of SEQ ID NO:2 to C A continuous sequence ending at positions A292 to G364 of as9, fused to the split intein-N, Cas (b) an N-terminal fragment of Cas9 beginning at positions A292 to G364 of Cas9 according to the numbering in SEQ ID NO: 2; A C-terminal fragment of Cas9, which is a contiguous sequence terminating at the C-terminus of Cas9 and fused to split intein-C. and (iii) a composition comprising one or more polynucleotides encoding the polypeptides of the present invention.

[0013] In some embodiments, (a) a sequence beginning at the N-terminus of Cas9 and continuing in the sequence numbering of SEQ ID NO:2 to C amino acids 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 576, 588, or 5 of as9 (b) an N-terminal fragment of Cas9, which is a contiguous sequence ending at 89 and fused to the split intein-N; Cas9: 303, 310, 313, 355, 456, 460, 463, 466, 472 ... A contiguous sequence beginning at positions 74, 577, 589, or 590 and ending at the C-terminus of Cas9, and one or more polynucleotides encoding a C-terminal fragment of Cas9 fused to Cas9-C. Compositions comprising:

[0014] In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 contains a nuclear localization signal. In some embodiments, the N-terminal fragment of Cas9 and the N-terminal fragment of Cas9 are linked to a null nucleotide sequence (NLS). Both C-terminal fragments are linked to an NLS. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the N-terminal fragment of Cas9 and the C-terminal fragment of Cas9 are linked In some embodiments, SpCa9 has nickase activity. or catalytically inactive.

[0015] In some embodiments, the N-terminal fragment of Cas9 in (a) disclosed herein is ) is provided herein a composition comprising a C-terminal fragment of Cas9.

[0016] In some embodiments, a vector comprising one or more polynucleotides disclosed herein is In some embodiments, the vector comprises a promoter. In embodiments, the promoter is a constitutive promoter. The promoter is a CMV or CAG promoter. Rovirus vector, adenovirus vector, lentivirus vector, herpesvirus a vector selected from the group consisting of a viral vector, an adeno-associated virus vector, and an adeno-associated virus vector. In the method, the vector is an adeno-associated virus vector.

[0017] In some embodiments, the compositions disclosed herein or the vectors disclosed herein are Provided herein is a cell comprising a vector. In some embodiments, the cell is a mammalian cell. be.

[0018] In some embodiments, Cas9 variants containing Ala / Cys, Ser / Cys, or Thr / Cys mutations are used. Ant polypeptides are provided herein. In some embodiments, Cys residues at amino acids 303, 310, 313, 355, 456, 460, 463, 466, 472, or 474

[0023] Provided are Cas9 variant polypeptides comprising:

[0019] In some embodiments, provided herein are methods for delivering a base editor system to a cell. The present invention provides a method for transfecting a cell with a deaminase and an N-terminal fragment of Cas9. and a first polynucleotide encoding a fusion protein comprising: The terminal fragment begins at the N-terminus of Cas9 and extends from positions A292 to G364 of Cas9 in the numbering of SEQ ID NO: 2. The first is a continuous sequence ending at the N-terminal end of the Cas9 fragment, and the N-terminal fragment of Cas9 is fused to the split intein-N. (b) a second polynucleotide encoding a C-terminal fragment of Cas9, Here, the C-terminal fragment of Cas9 begins at positions A292 to G364 of Cas9 according to the numbering in SEQ ID NO: 2. It is a continuous sequence that terminates at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused to the split intein-C. a second polynucleotide, and (c) a single guide RNA (sgRNA) or a coding sequence thereof. The method includes contacting the target gene with a polynucleotide that encodes the target gene.

[0020] In some embodiments, provided herein are methods for delivering a base editor system to a cell. The present invention provides a method for transfecting a cell with a fusion protein comprising an N-terminal fragment of Cas9. a first polynucleotide encoding an N-terminal fragment of Cas9, A continuous sequence beginning at the end of and ending at positions A292 to G364 of Cas9 in the numbering in SEQ ID NO: 2 wherein the N-terminal fragment of Cas9 is fused to the split intein-N; b) a second polynucleotide encoding a C-terminal fragment of Cas9 and a deaminase, Here, the C-terminal fragment of Cas9 begins at positions A292 to G364 of Cas9 according to the numbering in SEQ ID NO: 2. It is a continuous sequence that terminates at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused to the split intein-C. a second polynucleotide, and (c) a single guide RNA (sgRNA) or a sequence encoding the same. The method includes contacting the target gene with a polynucleotide that encodes the target gene.

[0021] In some embodiments, methods for delivering a base editor system to a cell are disclosed herein. The method includes: (a) injecting cells with a fusion protein containing a deaminase and an N-terminal fragment of Cas9; a first polynucleotide encoding a fusion protein, wherein the N-terminal fragment of Cas9 is , starting from the N-terminus of Cas9 and including A292 to G364, F445 to K483 of Cas9 in the numbering of SEQ ID NO: 2. or a continuous sequence ending at positions E565 to T637, and the N-terminal fragment of Cas9 is the split intein-N (b) a first polynucleotide encoding a C-terminal fragment of Cas9, fused to a nucleotide, wherein the C-terminal fragment of Cas9 is A2 of Cas9 according to the numbering in SEQ ID NO: 2; A continuous sequence beginning at positions 92 to G364, F445 to K483, or E565 to T637 and ending at the C-terminus of Cas9 the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted with Ala, Ser, or Thr, and (c) a second polynucleotide in which the C-terminal fragment of s9 is fused to split intein-C; and and contacting the sample with a single guide RNA (sgRNA) or a polynucleotide encoding the same. .

[0022] In some embodiments, methods for delivering a base editor system to a cell are disclosed herein. The method includes: (a) injecting a cell with a fusion protein comprising an N-terminal fragment of Cas9; a first polynucleotide that encodes a Cas9 N-terminal fragment, wherein the Cas9 N-terminal fragment begins at the N-terminus of Cas9; That is, A292 to G364, F445 to K483, or E565 to T637 of Cas9 in the numbering of SEQ ID NO: 2 and the N-terminal fragment of Cas9 is fused to the split intein-N. (b) a second polynucleotide encoding a C-terminal fragment of Cas9 and a deaminase; nucleotides, wherein the C-terminal fragment of Cas9 is Consecutive sequences beginning at positions A292–G364, F445–K483, or E565–T637 and ending at the C-terminus of Cas9 wherein the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted for Ala, Ser, or Thr; (c) a second polynucleotide in which the C-terminal fragment of Cas9 is fused to the split intein-C; and contacting the cells with a single guide RNA (sgRNA) or a polynucleotide encoding the same. nothing.

[0023] In some embodiments, the sgRNA is complementary to the target polynucleotide. Thus, the target polynucleotide is present in the genome of the organism. In some embodiments, the first polynucleotide is an animal, a plant, or a bacterium. The second polynucleotide, and / or the polynucleotide encoding them may be a vector. In some embodiments, the vector is a retroviral vector. -, adenovirus vector, lentivirus vector, herpesvirus vector, and and adeno-associated virus vectors. In one embodiment, the vector In some embodiments, the vector is an adeno-associated virus vector. The fragment or N-terminal fragment of Cas9 is a fragment of amino acids S303, T310, T313, T314, T315, T316, T317, T318, T319, T320, T321, T322, T323, T324, T325, T326, T327, T328, T329, T330, T331, T332, T , S355, A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, or S590 It includes Ala / Cys, Ser / Cys, or Thr / Cys mutations at the corresponding residues. In some embodiments, the deaminase is adenosine deaminase. In some embodiments, the deaminase is TadA or a variant thereof. wild-type TadA or Tad7.10. In some embodiments, the deaminase is a TadA dimer. In some embodiments, the TadA dimer comprises wild-type TadA and TadA7.10. In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 comprises an NLS. In some embodiments, both the N-terminal fragment of Cas9 and the C-terminal fragment of Cas9 comprise an NLS. In some embodiments, the NLS is a bipartite NLS. The N-terminal fragment of Cas9 and the C-terminal fragment of Cas9 are linked to form SpCa9. In various forms, SpCas9 either has nickase activity or is catalytically inactive.

[0024] In some embodiments, a polynucleotide encoding a fusion protein as described herein wherein the fusion protein comprises a deaminase and an N-terminal fragment of Cas9, The N-terminal fragment of Cas9 begins at the N-terminus of Cas9 and is numbered from A292 to G364 of Cas9 in SEQ ID NO: 2. The N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is wherein the fusion protein comprises a deaminase and an N-terminal fragment of Cas9, The N-terminal fragment of the present invention begins at the N-terminus of Cas9 and is numbered from F445 to K483 of Cas9 in SEQ ID NO: 2. The N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is wherein the fusion protein comprises a deaminase and an N-terminal fragment of Cas9, The N-terminal fragment of the present invention begins at the N-terminus of Cas9 and is numbered from E565 to T637 of Cas9 in SEQ ID NO: 2. The N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is wherein the fusion protein comprises a deaminase and a C-terminal fragment of Cas9, The C-terminal fragment of Cas9 begins at positions A292 to G364 of Cas9 in the numbering of SEQ ID NO: 2 and ends at the C It is a continuous sequence that terminates at the end, and the C-terminal fragment of Cas9 is fused to the split intein-C. In some embodiments, the polynucleotide encoding the fusion protein is wherein the fusion protein comprises a deaminase and a C-terminal fragment of Cas9, The C-terminal fragment of Cas9 begins at positions F445 to K483 of Cas9 in the numbering of SEQ ID NO: 2 and ends at the C It is a continuous sequence that terminates at the end, and the C-terminal fragment of Cas9 is fused to the split intein-C. In some embodiments, provided herein is a polynucleotide encoding a fusion protein, Here, the fusion protein comprises a deaminase and a C-terminal fragment of Cas9, and the C-terminal fragment of Cas9 begins at positions E565 to T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9. In one embodiment, the C-terminal fragment of Cas9 is fused to the split intein-C. Thus, provided herein are polynucleotides encoding fusion proteins, wherein the fusion The protein comprises a deaminase and a C-terminal fragment of Cas9, the C-terminal fragment of Cas9 being represented by SEQ ID NO: Starting at positions A292 to G364, F445 to K483, or E565 to T637 of Cas9 in the numbering in 2 , a continuous sequence terminating at the C-terminus of Cas9, and the N-terminal residues of the C-terminal fragment of Cas9 are Ala, Ser, or Cys substituted for Thr, and the C-terminal fragment of Cas9 is fused to the split intein C.

[0025] In some embodiments, the C-terminal fragment of Cas9 or the N-terminal fragment of Cas9 is Ala / Cys, S In some embodiments, the mutation comprises an er / Cys, or Thr / Cys mutation. Amino acids S303, T310, T313, S355, A456, S460, A463, and T466 as numbered in column 2 , S469, T472, T474, C574, S577, A589, or S590. In some embodiments, the deaminase is adenosine deaminase. In some embodiments, the deaminase is TadA or a variant thereof. is wild-type TadA or Tad7.10. In some embodiments, the fusion proteins are linked together In some embodiments, the fusion protein comprises two deaminases, wild-type TadA and In some embodiments, the fusion protein comprises both TadA7.10 and TadA7.10. In some embodiments, the fusion protein comprises an NLS. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the C-terminal fragment of Cas9 comprises the amino acid sequence of SpCas9. The end fragment or C-terminal fragment of Cas9 contains one or more amino acids associated with reduced nuclease activity. Contains substitutions.

[0026] In some embodiments, amino acids 302, 309, 312, 354, 455, 459, 46 N-terminal fragments of the Cas9 protein containing 2, 465, 471, or 473 were fused to the split intein-N. In some embodiments, a combination of Cas9 proteins is provided herein. C-terminal protein fragments are provided, wherein the N-terminal amino acid of the C-terminal fragment is amino acid 303, Cys substitutions at 310, 313, 355, 456, 460, 463, 466, 472, or 474, resulting in split-in In some embodiments, the A to G base editor fusion protein is fused to a TEIN-C. A polynucleotide encoding a fragment of a fusion protein is provided, and the fusion protein comprises one The deaminase and the N-terminal fragment of Cas9 are contained therein, and the N-terminal fragment is fused to split intein-N. In some embodiments, the A to G base editor fusion protein is described herein. A polynucleotide encoding a fragment of the fusion protein is provided, and the fusion protein comprises one or more deoxyribonucleotides. The C-terminal fragment of Cas9 is fused to a split intein-C. In some embodiments, the protein of an A to G base editor fusion protein is described herein. A fusion protein is provided that combines one or more deaminases with an N-terminal fragment of Cas9. In some embodiments, the N-terminal fragment is fused to a split intein-N. The specification provides protein fragments of A→G base editor fusion proteins, and the fusion proteins The protein comprises one or more deaminases and a C-terminal fragment of Cas9, the C-terminal fragment being a split-in. It is fused to tein-C.

[0027] In some embodiments, the A→G bases each comprise one or more deaminases and Cas9. A combination comprising first and second polynucleotides encoding fragments of an editor fusion protein. A composition is provided herein, wherein the first polynucleotide is fused to a split intein-N. The second polynucleotide encodes an N-terminal fragment of Cas9 fused to the split intein-C. In some embodiments, the C-terminal fragment of Cas9 encodes one or more deaminators. Compositions Comprising N- and C-Terminal Fragments of an A→G Base Editor Fusion Protein Comprising Cas9 and Caspase provided herein, wherein the N-terminal fragment is a fragment of SpCas9 fused to a split intein-N. The C-terminal fragment contains the remainder of SpCas9 fused to the split intein-C.

[0028] In some embodiments, methods for delivering a base editor system to a cell are disclosed herein. The method includes treating a cell with one or more deaminases and a Cas9 to transform an A→G base pair. First and second polynucleotides, each encoding a fragment of an editor fusion protein. wherein the first polynucleotide is fused to a split intein-N. a second polynucleotide encoding an N-terminal fragment of Cas9 fused to split intein-C; and either the first or second polynucleotide encodes a C-terminal fragment of Cas9. In some embodiments, the base editor system encodes a single guide RNA. Provided herein are methods for delivering to cells, the methods comprising delivering to cells one or more derivatives of N- and C-terminal fragments of the A→G base editor fusion protein containing amine and SpCas9, and contacting the N-terminal fragment with a guide RNA, wherein the N-terminal fragment is a split intein-N. The C-terminal fragment contains the fragment of SpCas9 fused to the split intein-C. In some embodiments, the method for editing a target polynucleotide in a cell includes the steps of: Methods are provided herein, which include transforming cells with one or more deaminases and Cas9. The first and second polynucleotides each encode a fragment of the A→G Base Editor fusion protein. contacting the first polynucleotide with a split intein-N; The second polynucleotide encodes an N-terminal fragment of Cas9 fused to the split intein-C. and wherein either the first or second polynucleotide encodes a C-terminal fragment of Cas9 fused thereto. and the encoded protein and single guide RNA are expressed in cells. This includes expressing the gene in cells.

[0029] Other features and advantages of the invention will become apparent from the detailed description and claims. There will be.

[0030] [Definition] Unless otherwise defined, all technical and scientific terms used herein are and have the meanings commonly understood by those skilled in the art to which this invention pertains. The literature provides those skilled in the art with general definitions of many of the terms used in this invention: leton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); T he Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossa ry of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). used in this specification. When used, the following terms are referred to below unless otherwise specified: It has the meaning.

[0031] "Adenosine deaminase" refers to the enzyme that hydrolyzes the deamination of adenine or adenosine. In one embodiment, the term "de" refers to a polypeptide or fragment thereof that is capable of catalyzing the deactivation of a protein. The aminase or deaminase domain hydrolyzes adenosine to inosine. Catalyzes the hydrolytic deamination of deoxyadenosine or deoxyadenosine to deoxyinosine In some embodiments, the adenosine deaminase is a deoxyadenosine deaminase. It catalyzes the hydrolytic deamination of adenine or adenosine in DNA. The adenosine deaminases provided herein (e.g., engineered adenosine deaminases) aminase, evolved adenosine deaminase) can be derived from any organism, e.g., bacteria. In some embodiments, the deaminase or deaminase domain may be In some embodiments, the deaminase is a variant of a naturally occurring deaminase derived from a plant. Alternatively, the deaminase domain may be non-naturally occurring. In a similar manner, the deaminase or deaminase domain is a naturally occurring deaminase At least 50%, at least 55%, at least 60%, at least 65%, at least 70% %, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the adenosine deaminase is derived from a gene encoding adenosine deaminase from a plant, e.g., E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA (ecTadA) deaminase or This is a fragment of it.

[0032] For example, truncated ecTadA may have one or more N-terminal amino acids compared to full-length ecTadA. In some embodiments, the truncated ecTadA is a truncated ecTadA fragment of the full-length ecTadA. Compared to TadA, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, Alternatively, the 20 N-terminal amino acid residues may be deleted. The packed ecTadA has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 compared to the full-length ecTadA. The C-terminal amino acid residues of 4, 15, 6, 17, 18, 19, or 20 may be deleted. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. In embodiments, the TadA deaminase is an N-terminal truncated TadA. In PCT / US2017 / 045381, the entire contents of which are incorporated herein by reference, TadA is There is one of the listed TadAs.

[0033] In certain embodiments, the adenosine deaminase comprises the following amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPT AHAEIMALRQGGLVMQNYRLIDAT LYVTLEPCVMCAGAMIHSRIGRVVFGARDAKT GAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKA QKKAQSSTD This is called the "TadA reference sequence."

[0034] In some embodiments, the TadA deaminase is a full-length E. coli TadA deaminase. For example, In certain embodiments, the adenosine deaminase comprises the following amino acid sequence: MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VMQNYRLIDATLYVTLEPCVMCAGAMIHSRIG RVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSD FFRMRRQEI KAQKKAQSSTD

[0035] However, additional adenosine deaminases useful herein will be apparent to those skilled in the art. It should be understood that these are obvious and within the scope of the present disclosure. The aminase may be a homolog of adenosine deaminase (AD AT) that acts on tRNA Exemplary AD AT homologs include, but are not limited to, the following:

[0036] Staphylococcus aureus TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCS GS LMNLLQQS NFNHRAIVDKG VLKE AC S TLLTTFFKNL RANKKS TN

[0037] Bacillus subtilis TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGC S GTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAAR KNLSE

[0038] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV

[0039] Shewanella putrefaciens (S. putrefaciens) TadA: MDE YWMQVAMQM AEKAEAAGE VPVGA VLVKDGQQIATGYNLS IS QHDPT AHAEI LCLRSAGKKLENYRLLDA TLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKK ALKLAQRAQQGIE

[0040] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQS DPT ΑΗ AEIIALRNG AKNI QN YRLLNS TLY VTLEPCTMC AG AILHS RIKRLVFG AS D YK TGAIGSRFHFFDDYKMNHTLEITSGVLAEE CSQKLSTFFQKRREEKKIEKALLKSLSD K

[0041] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAAKLGNYRL TDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRK AKI

[0042] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP

[0043] TadA7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0044] An "agent" is any small molecule compound, antibody, nucleic acid molecule, or polypeptide. , or fragments thereof.

[0045] "Alteration" refers to any alteration that can be detected by standard art known methods such as those described herein. This refers to a change in the structure, expression level, or activity of a gene or polypeptide. As used herein, an alteration (e.g., an increase or decrease) is defined as a 10% change in expression level, a 2% change in expression level, or a 10% change in expression level. These include a 5% change, a 40% change, and a 50% or greater change in expression level.

[0046] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analog may possess the biological activity of the corresponding naturally occurring polypeptide. Increased functionality of the analog compared to the native polypeptide while retaining at least some of the Such modifications may enhance, for example, polynucleotide binding activity. Increased protease resistance, membrane permeability, or half-life of analogs without altering In another example, the polynucleotide analog can be a natural polynucleotide. The analogs may have certain modifications that enhance the function of the analogs compared to the corresponding native polynucleotides. Such modifications retain the biological activity of the nucleotides. The analogs may increase the affinity, half-life, and / or nuclease resistance of the non-naturally occurring The amino acid sequence may contain nucleotides or amino acids.

[0047] A "base editor (BE)" or "nucleobase editor (NBE)" is a In one embodiment, the drug is a drug that binds to a nucleotide and has nucleobase-modifying activity. The agent has a domain with base editing activity, i.e., a domain that modifies bases (e.g., For example, a fusion protein containing a domain capable of modifying an amino acid sequence (e.g., A, T, C, G, U). In some embodiments, the domain having base editing activity modifies a base in a nucleic acid molecule. In some embodiments, the base editor can aminolate a base within a DNA molecule. In some embodiments, the base editor aminated a cytosine (C ) or adenosine can be deaminated. can deaminate cytosine (C) and adenosine (A) in DNA. In embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In form, the base editor is an adenosine base editor (ABE) and a cytidine salt. In some embodiments, the base editor is adenosine deaminase (ADM)-containing base editor (CBE). In some embodiments, the target gene is a nuclease-inactivated Cas9 (dCas9) fused to a nuclease. Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9). Circularly permuted Cas9 is known in the art, see, for example, Oakes et al., Cell 176, 254- 267, 2019. In some embodiments, the base editor is a base removal repair In some embodiments, the fusion protein is fused to a second inhibitor, e.g., a UGI domain. Cas9 nicks fused to a deaminase and an inhibitor of base excision repair, such as a UGI domain. In other embodiments, the base editor is an abasic base editor.

[0048] Nucleobase components of base editor systems and polynucleotide programmable nucleases The nucleotide binding moieties can be covalently or non-covalently bound to one another. For example, In some embodiments, the deaminase domain is a polynucleotide programming Targeted to a target nucleotide sequence by a nucleotide-binding domain that can be In some embodiments, the polynucleotide has a programmable nucleotide binding domain. In some embodiments, the deaminase domain may be fused or linked to the deaminase domain. The polynucleotide programmable nucleotide binding domain is a deaminase The deaminase domain binds to the ATPase domain by non-covalently interacting with or binding to the ATPase domain. It can be targeted to a target nucleotide sequence. For example, in some embodiments In the method, a nucleobase editing component, e.g., a deaminase component, is used to program polynucleotides. and a further heterologous moiety or domain that is part of the nucleotide-binding domain that can be bound. Additional heterologous moieties or domains that can interact, associate, or form complexes. In some embodiments, the additional heterologous moiety can include a polypeptide. Some examples are capable of binding, interacting, associating, or forming complexes with In embodiments, the additional heterologous moiety binds to, interacts with, or associates with the polynucleotide, or can form a complex. In some embodiments, the additional heterologous moiety In some embodiments, additional The heterologous moiety can be attached to a polypeptide linker. In this case, the additional heterologous moiety can be attached to the polynucleotide linker. The species portion may be a protein domain. In some embodiments, additional heterologous sequences may be present. The seed portion contains the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein domain. Main, SfMu Com coat protein domain, steryl α motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, can be an RNA recognition motif.

[0049] The base editor system can further comprise a guide polynucleotide component. The components of the base editor system may be linked by covalent bonds, non-covalent interactions, or both. It should be understood that the compounds may be coupled to one another via any combination of bonds and interactions. In some embodiments, the deaminase domain is selected from the group consisting of: It can be targeted to a target nucleotide sequence. For example, in some embodiments The nucleobase editing component of the base editor system, e.g., the deaminase component, interacts with and binds to a portion or segment of a nucleotide (e.g., a polynucleotide motif) or further heterologous moieties or domains (e.g. RNA or DNA binding) that can form complexes. In some embodiments, the polypeptide may comprise a polynucleotide binding domain, such as a ligation protein. In the present invention, an additional heterologous moiety or domain (e.g., a polypeptide such as an RNA or DNA binding protein) may be used. The nucleotide-binding domain can be fused or linked to the deaminase domain. In some embodiments, the additional heterologous moiety binds to, interacts with, or associates with the polypeptide. In some embodiments, the polypeptide may be bound to or complexed with the polypeptide. Thus, the additional heterologous moiety binds to, interacts with, associates with, or binds to the polynucleotide. In some embodiments, the additional heterologous nucleotide can be conjugated to a heterologous nucleotide. The moiety is capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous moiety can be attached to the polypeptide linker. In some embodiments, the additional heterologous moiety can be attached to the polynucleotide linker. The additional heterologous moiety may be a protein domain. The additional heterologous portion contains the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein domain. Protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein The amino acid sequence may be a protein or an RNA recognition motif.

[0050] In some embodiments, the base editor system comprises one or more proteins, fusion proteins, The present invention may include a polypeptide, or a polynucleotide encoding the polypeptide. In an embodiment, the base editor system comprises a deaminase and an N-terminal fragment of napDNAbp. a first polynucleotide encoding a fusion protein comprising: For example, in certain embodiments, napDNAb The N- and C-terminal fragments of p are reconstituted to form the base editor protein. , the N-terminal fragment of napDNAbp is fused to intein-N, and the C-terminal fragment of napDNAbp is fused to intein-C. can be fused to

[0051] In one embodiment, the base editor system includes a base excision repair (BER) component. In some embodiments, the base editor system may further comprise an inhibitor of the base editor system. The system can further comprise an inhibitor of a base excision repair (BER) component. The components of the group editor system can be linked by covalent bonds, non-covalent interactions, or their It should be understood that the molecules may be coupled to one another via any combination of bonds and interactions. Inhibitors of BER components can include base excision repair inhibitors. ,The inhibitor of base excision repair may be uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of base excision repair can be an inosine base excision repair inhibitor. In some embodiments, the inhibitor of base excision repair is a polynucleotide programmable The nucleotide sequence can be targeted to a target nucleotide sequence by a nucleotide-binding domain capable of binding to the target nucleotide sequence. In one embodiment, the polynucleotide programmable nucleotide binding domain , may be fused or linked to an inhibitor of base excision repair. The programmable nucleotide binding domain is a deaminase domain and a salt In some embodiments, the polynucleotide may be fused or linked to an inhibitor of base excision repair. Nucleotide-programmable nucleotide-binding domains are inhibitors of base excision repair by noncovalently interacting with or associating with inhibitors of base excision repair. This allows for targeting of inhibitors of base excision repair to target nucleotide sequences. For example, in some embodiments, an inhibitor of a base excision repair component is a further heterologous moiety or moieties that are part of the nucleotide-programmable nucleotide-binding domain or a further heterologous moiety or domain that can interact, associate, or complex with the In some embodiments, the inhibitor of base excision repair may comprise a guide polynucleotide. For example, in some embodiments, In the present invention, the inhibitor of base excision repair is a part or segment of the guide polynucleotide ( Further molecules that can interact with, associate with, or complex with the target molecule (e.g., a polynucleotide motif). a heterologous moiety or domain (e.g., a polynucleotide such as an RNA or DNA binding protein) that In some embodiments, the guide polynucleotide may comprise a nucleotide-binding domain. Further heterologous moieties or domains (e.g., polynucleotides such as RNA or DNA binding proteins) The nucleotide-binding domain can be fused or linked to an inhibitor of base excision repair. In some embodiments, the additional heterologous moiety binds, interacts, associates with the polynucleotide. In some embodiments, the additional heterologous moiety can bind or complex with the In some embodiments, the additional Additional heterologous moieties can be attached to the polypeptide linker. In the polynucleotide linker, the additional heterologous moiety can be attached to the polynucleotide linker. The heterologous moiety may be a protein domain. In some embodiments, additional The heterologous portion contains the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein. domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, Or it may be an RNA recognition motif.

[0052] "Base editing activity" refers to the ability to chemically change bases within a polynucleotide. In one embodiment, the first base is converted to the second base. In one embodiment, the base editing activity is , a cytidine deaminase activity, e.g., converting the target C·G to T·A. In embodiments, the base editing activity is adenosine deaminase activity, e.g., A It is the activity that converts ·T to G·C.

[0053] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Cas9 Active, inactive, or partially active DNA cleavage domains of Cas9 and / or gRNA binding of Cas9 Cas9 nuclease refers to an RNA-guided nuclease containing a c domain. asn1 nuclease or CRISPR (clustered regularly interspaced short palindromic CRISPR is also known as a repeat-associated nuclease. The adaptive immune system provides defense against transposable elements (transposable elements, conjugative plasmids). CRIS comprises a spacer, a sequence complementary to the preceding mobile element, and a target invader nucleic acid. The PR cluster is transcribed and processed into CRISPR RNA (crRNA). In the endothelial cell line, the correct processing of pre-crRNA is mediated by a small trans-coding RNA (tracrRNA), an endogenous RNA. tracrRNA requires ribonuclease 3 (rnc) and Cas9 protein. This guides the processing of pre-crRNA by Cas9 / crRNA / tracrRNA. A endonuclease cleaves a linear or circular dsDNA target complementary to the spacer The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically cleaved. Both the crRNA and tracrRNA are nucleolytically trimmed to 3'-5' ends. Engineer a single guide RNA ("sgRNA," or simply "gRNA") to integrate into species A For example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, See Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated by reference. Cas9 targets short motifs (PAM or PAM) in the CRISPR repeats. Ca recognizes the protospacer adjacent motif and helps distinguish self from non-self. The sequence and structure of s9 nuclease are well known to those skilled in the art (see, e.g., "Complete Genome"). ome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by tr ans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., S harma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentie r E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endo nuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., See Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entirety of which is incorporated herein by reference. Cas9 orthologs include, but are not limited to: Although not widely known, it has been described in various species, including S. pyogenes and S. thermophilus. Suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, comprising: Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA B and Cas9 sequences from the organisms and loci disclosed in Biology 10:5, 726-737. The entire contents of which are incorporated herein by reference.

[0054] The nuclease-inactivated Cas9 protein is interchangeably referred to as the “dCas9” protein (nuclease-“ It may also be referred to as "dead" Cas9 or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having the following structure are known (see, e.g., Jinek et al. al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression”(2013) Cell. 28; 152 (5): 1173-83, the contents of each of which are incorporated herein by reference. For example, the DN of Cas9 The A cleavage domain is composed of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA and binds to RuvC1. The subdomains cleave the non-complementary strand. Mutations within these subdomains allow Cas9 to For example, mutations D10A and H840A inhibit the nuclease activity of S. pyogenes Cas9. completely inactivates cleavage activity (Jinek et al., Science. 337:816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, the Cas9 nuclease is and having an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is referred to as "nCas9." In one embodiment, the protein is a nickase called Cas9 ("nickase"). Proteins comprising fragments of as9 are provided. For example, in some embodiments, The protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; 2) The DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is These are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof. For example, a Cas9 variant may have at least about 70% identity, at least about 8% identity, or at least about 9% identity to wild-type Cas9. 0% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least or at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or or at least about 99.9% identical. In some embodiments, the Cas9 variant is a wild-type Compared to type 1 Cas9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes In some embodiments, the Cas9 variant may be a fragment of Cas9 (e.g., a gRNA-binding fragment). domain or DNA cleavage domain), the fragments of which have at least one identical sequence to the corresponding fragment of wild-type Cas9. at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, At least about 98% identical, at least about 99% identical, at least about 99.5% identical In some embodiments, the fragment is identical to or at least about 99.9% identical to the corresponding wild-type C At least 30%, at least 35%, at least 40%, at least 45%, or at least at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% do.

[0055] In some embodiments, the fragment is at least 100 amino acids in length. In the above, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 60 0, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or In some embodiments, the wild-type Cas9 is at least 1300 amino acids in length. Cas9 from Escherichia coli pyogenes (NCBI reference sequence: NC_17053.1, nucleotide sequence and The amino acid sequence is as follows: ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA (SEQ ID NO: 1) JPEG2025028857000002.jpg161169 (single underline: HNH domain; double underline: RuvC domain)

[0056] In one embodiment, the wild-type Cas9 has the following nucleotide and / or amino acid sequence: Corresponding to or containing the following sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025028857000003.jpg157169 (single underline: HNH domain; double underline: RuvC domain)

[0057] In some embodiments, wild-type Cas9 is directed against Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: String: Q99ZW2 (amino acid sequence below). ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAAGAGCATCCT GTTGAAAATACTCAAATTGCAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025028857000004.jpg156168(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0058] In some embodiments, Cas9 is derived from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Strept ococcus iniae(NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB I Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jej uni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_00 2342100.1), or Cas9 from any other organism.

[0059] In some embodiments, dCas9 contains one or more nucleotides that inactivate Cas9 nuclease activity. It corresponds to or comprises part or all of a mutated Cas9 amino acid sequence. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or contains a corresponding mutation in another Cas9. In some embodiments, dCas9 is The amino acid sequence of dCas9 (D10A and H840A) is as follows: JPEG2025028857000005.jpg157168 (single underline: HNH domain; double underline: RuvC domain)

[0060] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.

[0061] In other embodiments, D10A, e.g., resulting in nuclease-inactivated Cas9 (dCas9), and dCas9 variants having mutations other than H840A. For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9 is , at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, those having at least about 99.9% identity are provided. About 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more, or longer. Variants of dCas9 having amino acid sequences are provided.

[0062] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Does not contain the string, only one or more fragments of it.

[0063] Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and Cas9 Further suitable sequences of domains and fragments will be apparent to those skilled in the art.

[0064] In some embodiments, Cas9 is isolated from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense(NCBI Ref: NC_021846.1); Strepto coccus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB I Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jej uni (NCBI Ref: YP_002344900.1); or Neisseria. meningitidis (NCBI Ref: YP_0023 This refers to Cas9 derived from 42100.1).

[0065] Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nCas9, or nuclease-active Cas9, its variants and homologs It should be understood that within the scope of this disclosure are any and all Cas9 proteins, including: In some embodiments, Cas9 includes, but is not limited to, those provided below. The protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). The quality is nuclease-active Cas9.

[0066] Exemplary catalytically inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0067] Exemplary catalyst of Cas9ニッカーゼ(nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0068] Exemplary catalyst activityCas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0069] In some embodiments, Cas9 is expressed in the Archaea (Archaea), which comprise the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the Cas protein refers to, for example, Cas9 from Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 refers to CasX or CasY, and all of its The contents of this document are incorporated herein by reference. Many CRISPR-Cas systems have been identified, including the first reported Cas9 in the archaeal domain of life This branched Cas9 protein was shown to function in little-studied nanoarchaea. It was discovered as part of an active CRISPR-Cas system. Two systems, CRISPR-CasX and CRISPR-CasY, have been discovered that are among the most In some embodiments, Cas9 is a CasX or a variant of CasX. In some embodiments, Cas9 represents a variant of CasY. Nucleic acid programmable DNA binding protein (napDNAbp) Other RNA-guided DNA binding proteins may also be used as the RNA-guided DNA binding protein of the present disclosure. It should be understood that the range is

[0070] In some embodiments, the napDNAbp is a Cas9 domain, e.g., a nuclease active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCa s9 and nCas9), type II Cas effector proteins, type V Cas effector proteins proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or Other nucleic acid programmable nucleic acids include modified or engineered versions thereof. DNA binding proteins are also within the scope of this disclosure, even if not specifically listed in this disclosure. For example, Makarova et al. "Classification and Nomenclature of CRISPR-Cas" Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.20 18.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. See doi: 10.1126 / science.aav7271 (full contents of each article) (The entirety of which is incorporated herein by reference).

[0071] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. is at least 85%, at least 90%, or At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the napDNAbp comprises a naturally occurring amino acid sequence having identity to the napDNAbp. In some embodiments, the napDNAbp is a CasX or CasY protein present. At least 85%, at least at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least CasX and CasY from other bacterial species share at least 99.5% amino acid sequence identity. It should also be understood that they may be used in accordance with the present disclosure.

[0072] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus islandicu s (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG

[0073] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandic us (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0074] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacte rium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0075] The term "CRISPR-Cas domain" or "CRISPR-Cas DNA binding domain" refers to a CRISPR-associated ( RNA-guided proteins containing Cas proteins or fragments thereof (e.g., Active, inactive, or partially active DNA cleavage domains and / or Cas proteins CRISPR clusters are transcribed into CRISPR RNA (a protein containing an RNA-binding domain). The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In some CRISPR systems, the correct processing of pre-crRNA requires transcription. The small RNA encoding tracrRNA (tracrRNA), endogenous ribonuclease 3 (rnc) and Cas proteins This tracrRNA is required for the processing of pre-crRNA, which is assisted by ribonuclease 3. Cas9 / crRNA and / or tracrRNA then bind to the spacer endonucleolytic cleavage of linear or circular dsDNA targets complementary to the target. may require both RNAs for DNA binding and cleavage. However, the role of crRNA and tracrRNA A single guide RNA ("sgRNA" or simply "gRNA") is designed to incorporate both flanks into a single RNA species. A") can be genetically engineered. See, e.g., Jinek M., Chylinski K., Fonfara I., Hau See er M., Doudna JA, Charpentier E. Science 337:816-821(2012) (the entire contents of which are incorporated herein by reference) (The text is incorporated herein by reference.) Cas proteins bind short motifs within the CRISPR repeats. It recognizes the PAM or protospacer adjacent motif and distinguishes self from non-self. CRISPR-Cas proteins include Cas9, CasX, CasY, Cpf1, C2c1, and C2 Further suitable CRIS include, but are not limited to, C3 or an active fragment thereof. PR-Cas proteins and sequences will be apparent to those of skill in the art based on the present disclosure.

[0076] Nuclease-inactivated CRISPR-Cas proteins are interchangeably referred to as "dCas" proteins ( It can be called a nuclease (dead Cas) or catalytically inactive Cas. Inactive DNA Methods for generating Cas proteins (or fragments thereof) with cleavage domains are known ( For example, Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (20 13) Cell. 28;152(5):1173-83, the entire contents of which are incorporated herein by reference. For example, the DNA cleavage domain of Cas9 consists of the HNH nuclease subdomain and the RuvC1 subdomain. It is known that the HNH subdomain is a gRNA subdomain. The RuvC1 subdomain cleaves the strand complementary to the target strand, while the RuvC2 subdomain cleaves the non-complementary strand. Mutations within the subdomains suppress the nuclease activity of Cas9. For example, mutations D10A and and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., S Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)). In this form, Cas nucleases contain an inactive (e.g., inactivated) DNA cleavage domain. Cas is a nickase called "nCas" protein ("nickase" Cas Cas variants share homology with CRISPR-Cas proteins or fragments thereof. For example, a Cas variant may have at least about 70% identity to a wild-type CRISPR-Cas protein. sex, at least about 80% identity, at least about 90% identity, at least about 95% identity, At least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity In some embodiments, the Cas variant has a wild-type CRISPR-Cas protein Compared to quality, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 , 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 , 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes In some embodiments, the Cas variant is a fragment of a wild-type CRISPR-Cas tag. at least about 70% identity, at least about 80% identity to the corresponding fragment of the protein; At least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least CRISPR-Cas proteins may be cloned to have at least about 99.5% identity, or at least about 99.9% identity. a fragment of a protein (e.g., a gRNA programmable DNA binding domain or DNA cleavage domain) In some embodiments, the fragment comprises the amino acid length of the corresponding wild-type CRISPR-Cas protein. At least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80% identical, at least 85%, at least 90%, at least 95% identical, at least 96%, At least 97%, at least 98%, at least 99%, or at least 99.5%. In some embodiments, the fragment is at least 100 amino acids in length. , length is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 70 0, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least It has 1,300 amino acids.

[0077] In this disclosure, the terms "comprises," "comprising," and "contain" are used interchangeably. "container," "having," and the like shall have the meanings ascribed to them in U.S. patent law. It can mean "includes," "including," etc., and "essentially from "consisting essentially of" or "consists essentially of" "y")" also has the meaning given to it in U.S. patent law, and such term is open-ended. The basic or novel properties of the described thing are not present in other things than the described thing. Unless otherwise varied, other than those described are permitted, except that the prior art implementation Exclude aspects.

[0078] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "catalyzed polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing , cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine PmCDA1 (Petromyzon marinus cytosine deaminase) from Petromyzon marinus 1, "PmCDA1"), AID (active AID) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.) Activation-induced cytidine deaminase (AICDA), and APOBEC are exemplary cytidine deaminases. It is Ze.

[0079] The nucleotide and amino acid sequences of PmCDA1 and the nucleotide and amino acid sequences of the CDS of human AID The acid sequence is shown below.

[0080] >tr|A5H718|A5H718_PETMA Cytosine deaminase OS=Petromyzon marinus OX=7757 PE=2 SV =1 MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSG TERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLK IWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKT LKRAEKRRSELSIMIQVKILHTTKSPAV

[0081] >EF094822.1 Petromyzon marinus isolate PmCDA.21 cytosine deaminase mRNA, complet e cds TGACACGACACAGCCGTGTATATGAGGAAGGGTAGCTGGATGGGGGGGGGGGGAATACGTTCAGAGAGGA CATTAGCGAGCGTCTTGTTGGTGGCCTTGAGTCTAGACACCTGCAGACATGACCGACGCTGAGTACGTGA GAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACAACAAAAAATCCGTGTCGCA TAGATGCTACGTTCTCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTTGGGGCTATGCTGTG AATAAACCACAGAGCGGGACAGAACGTGGAATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAAT ACCTGCGCGACAACCCCGGACAATTCACGATAAATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGC TGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGC AAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGT TGAATGTAATGGTAAGTGAACACTACCAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATT GAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATG ATTCAGGTAAAAATACTCCACACCACTAAGAGTCCTGCTGTTTAAGAGGCTATGCGGATGGTTTTC

[0082] >tr|Q6QJ80|Q6QJ80_HUMAN Activation-induced cytidine deaminase OS=Homo sapiens OX =9606 GN=AICDA PE=2 SV=1 MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELL FLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRK AEPEGLRRLHRAGVQIAIMTFKAPV

[0083] >NG_011588.1:5001-15681 Homo sapiens activation induced cytidine deaminase (AICD A), RefSeqGene (LRG_17) on chromosome 12 AGAGAACCATCATTAATTGAAGTGAGATTTTTCTGGCCTGAGACTTGCAGGGAGGCAAGAAGACACTCTG GACACCACTATGGACAGGTAAAGAGGCAGTCTTCTCGTGGGTGATTGCACTGGCCTTCCTCTCAGAGCAA ATCTGAGTAATGAGACTGGTAGCTATCCCTTTCTCTCATGTAACTGTCTGACTGATAAGATCAGCTTGAT CAATATGCATATATTTTTTGATCTGTCTCCTTTTCTTCTATTCAGATCTTATACGCTGTCAGCCCAAT TCTTTCTGTTTCAGACTTCTCTTGATTTCCCTCTTTTCATGTGGCAAAAGAAGTAGTGCGTACAATGTA CTGATTCGTCCTGAGATTTGTACCATGGTTGAAACTAATTTATGGTAATAATATTAACATAGCAAATCTT TAGAGACTCAAATCATGAAAAGGTAATAGCAGTACTGTACTAAAAACGGTAGTGCTAATTTTCGTAATAA TTTTGTAAATATTCAAACAGTAAAACACTTGAAGACACACTTTCCTAGGGAGGCGTTTACTGAAATAATTT AGCTATAGTAAGAAAATTTGTAATTTTAGAAATGCCAAGCATTCTAAATTAATTGCTTGAAAGTCACTAT GATTGTGTCCATTATAAGGAGACAAATTCATTCAAGCAAGTTATTTAATGTTAAAGGCCCAATTGTTAGG CAGTTAATGGCACTTTTACTATTAACTAATCTTTCCATTTGTTCAGACGTAGCTTAACTTACCTCTTAGG TGTGAATTTGGTTAAGGTCCTCATAATGTCTTTATGTGCAGTTTTTGATAGGTTATTGTCATAGAACTTA TTCTATTCCTACATTTTATGATTACTATGGATGTATGAGAATAACACCTAATCCTTTACTTTACCTCAAT TTAACTCCTTTATAAAGAACTTACATTACAGAATAAAGATTTTTTAAAAATATATTTTTTTGTAGAGACA GGGTCTTAGCCCAGCCGAGGCTGGTCTCTAAGTCCTGGCCCAAGCGATCCTCCTGCCTGGGCCTCCTAAA GTGCTGGAATTATAGACATGAGCCATCACATCCAATATACAGAATAAAGATTTTTAATGGAGGATTTAAT GTTCTTCAGAAAATTTTCTTGAGGTCAGACAATGTCAAATGTCTCCTCAGTTTACACTGAGATTTTGAAA ACAAGTCTGAGCTATAGGTCCTTGTGAAGGGTCCATTGGAAATACTTGTTCAAAGTAAAATGGAAAGCAA AGGTAAAATCAGCAGTTGAAATTCAGAGAAAGACAGAAAAGGAGAAAAGATGAAATTCAACAGGACAGAA GGGAAATATATTATCATTAAGGAGGACAGTATCTGTAGAGCTCATTAGTGATGGCAAAATGACTTGGTCA GGATTATTTTTAACCCGCTTGTTTCTGGTTTGCACGGCTGGGGATGCAGCTAGGGTTCTGCCTCAGGGAG CACAGCTGTCCAGAGCAGCTGTCAGCCTGCAAGCCTGAAACACTCCCTCGGTAAAGTCCTTCCTACTCAG GACAGAAATGACGAGAACAGGGAGCTGGAAACAGGCCCCTAACCAGAGAAGGGAAGTAATGGATCAACAA AGTTAACTAGCAGGTCAGGATCACGCAATTCATTTCACTCTGACTGGTAACATGTGACAGAAACAGTGTA GGCTTATTGTATTTTCATGTAGAGTAGGACCCAAAAATCCACCCAAAGTCCTTTATCTATGCCACATCCT TCTTATCTATACTTCCAGGACACTTTTTCTTCCTTATGATAAGGCTCTCTCTCTCTCCACACACACACAC ACACACACACACACACACACACACACACACACAAACACACACCCCGCCAACCAAGGTGCATGTAAAAAGA TGTAGATTCCTCTGCCTTTCTCATCTACACAGCCCAGGAGGGTAAGTTAATATAAGAGGGATTTATTGGT AAGAGATGATGCTTAATCTGTTTAACACTGGGCCTCAAAGAGAGAATTTCTTTTCTTCTGTACTTATTAA GCACCTATTATGTGTTGAGCTTATATATACAAAGGGTTATTATATGCTAATATAGTAATAGTAATGGTGG TTGGTACTATGGTAATTACCATAAAAATTATTATCCTTTTAAAATAAAGCTAATTATTATTGGATCTTTT TTAGTATTCATTTTATGTTTTTTATGTTTTTGATTTTTTAAAAGACAATCTCACCCTGTTACCCAGGCTG GAGTGCAGTGGTGCAATCATAGCTTTCTGCAGTCTTGAACTCCTGGGCTCAAGCAATCCTCCTGCCTTGG CCTCCCAAAGTGTTGGGATACAGTCATGAGCCACTGCATCTGGCCTAGGATCCATTTAGATTAAAATATG CATTTTAAATTTTAAAATAATATGGCTAATTTTTACCTTATGTAATGTGTATACTGGCAATAAATCTAGT TTGCTGCCTAAAGTTTAAAGTGCTTTCCAGTAAGCTTCATGTACGTGAGGGGAGACATTTAAAGTGAAAC AGACAGCCAGGTGTGGTGGCTCACGCCTGTAATCCCAGCACTCTGGGAGGCTGAGGTGGGTGGATCGCTT GAGCCCTGGAGTTCAAGACCAGCCTGAGCAACATGGCAAAACGCTGTTTCTATAACAAAAATTAGCCGGG CATGGTGGCATGTGCCTGTGGTCCCAGCTACTAGGGGGCTGAGGCAGGAGAATCGTTGGAGCCCAGGAGG TCAAGGCTGCACTGAGCAGTGCTTGCGCCACTGCACTCCAGCCTGGGTGACAGGACCAGACCTTGCCTCA AAAAAATAAGAAGAAAAATTAAAAATAAATGGAAACAACTACAAAGAGCTGTTGTCCTAGATGAGCTACT TAGTTAGGCTGATATTTTGGTATTTAACTTTTAAAGTCAGGGTCTGTCACCTGCACTACATTATTAAAAT ATCAATTCTCAATGTATATCCACACAAAGACTGGTACGTGAATGTTCATAGTACCTTTATTCACAAAACC CCAAAGTAGAGACTATCCAAATATCCATCAACAAGTGAACAAATAAACAAAATGTGCTATATCCATGCAA TGGAATACCACCCTGCAGTACAAAGAAGCTACTTGGGGATGAATCCCAAAGTCATGACGCTAAATGAAAG AGTCAGACATGAAGGAGGAGATAATGTATGCCATACGAAATTCTAGAAAATGAAAGTAACTTATAGTTAC AGAAAGCAAATCAGGGCAGGCATAGAGGCTCACACCTGTAATCCCAGCACTTTGAGAGGCCACGTGGGAA GATTGCTAGAACTCAGGAGTTCAAGACCAGCCTGGGCAACACAGTGAAACTCCATTCTCCACAAAAATGG GAAAAAAAGAAAGCAAATCAGTGGTTGTCCTGTGGGGAGGGGAAGGACTGCAAAGAGGGAAGAAGCTCTG GTGGGGTGAGGGTGGTGATTCAGGTTCTGTATCCTGACTGTGGTAGCAGTTTGGGGTGTTTACATCCAAA AATATTCGTAGAATTATGCATCTTAAATGGGTGGAGTTTACTGTATGTAAATTATACCTCAATGTAAGAA AAAATAATGTGTAAGAAAACTTTCAATTCTCTTGCCAGCAAACGTTATTCAAATTCCTGAGCCCTTTACT TCGCAAATTCTCTGCACTTCTGCCCCGTACCATTAGGTGACAGCACTAGCTCCACAAATTGGATAAATGC ATTTCTGGAAAAGACTAGGGACAAAATCCAGGCATCACTTGTGCTTTCATATCAACCATGCTGTACAGCT TGTGTTGCTGTCTGCAGCTGCAATGGGGACTCTTGATTTCTTTAAGGAAACTTGGGTTACCAGAGTATTT CCACAAATGCTATTCAAATTAGTGCTTATGATATGCAAGACACTGTGCTAGGAGCCAGAAAACAAAGAGG AGGAGAAATCAGTCATTATGTGGGAACAACATAGCAAGATATTTAGATCATTTTGACTAGTTAAAAAAGC AGCAGAGTACAAAATCACACATGCAATCAGTATAATCCAAATCATGTAAATATGTGCCTGTAGAAAGACT AGAGGAATAAACACAAGAATCTTAACAGTCATTGTCATTAGACACTAAGTCTAATTATTATTATTAGACA CTATGATATTTGAGATTTAAAAAATCTTTAATATTTTAAAATTTAGAGCTCTTCTATTTTTCCATAGTAT TCAAGTTTGACAATGATCAAGTATTACTCTTTCTTTTTTTTTTTTTTTTTTTTTTTTTGAGATGGAGTTT TGGTCTTGTTGCCCATGCTGGAGTGGAATGGCATGACCATAGCTCACTGCAACCTCCACCTCCTGGGTTC AAGCAAAGCTGTCGCCTCAGCCTCCCGGGTAGATGGGATTACAGGCGCCCACCACCACACTCGGCTAATG TTTGTATTTTTAGTAGAGATGGGGTTTCACCATGTTGGCCAGGCTGGTCTCAAACTCCTGACCTCAGAGG ATCCACCTGCCTCAGCCTCCCAAAGTGCTGGGATTACAGATGTAGGCCACTGCGCCCGGCCAAGTATTGC TCTTATACATTAAAAAACAGGTGTGAGCCACTGCGCCCAGCCAGGTATTGCTCTTATACATTAAAAAATA GGCCGGTGCAGTGGCTCACGCCTGTAATCCCAGCACTTTGGGAAGCCAAGGCGGGCAGAACACCCGAGGT CAGGAGTCCAAGGCCAGCCTGGCCAAGATGGTGAAACCCCGTCTCTATTAAAAATACAAACATTACCTGG GCATGATGGTGGGCGCCTGTAATCCCAGCTACTCAGGAGGCTGAGGCAGGAGGATCCGCGGAGCCTGGCA GATCTGCCTGAGCCTGGGAGGTTGAGGCTACAGTAAGCCAAGATCATGCCAGTATACTTCAGCCTGGGCG ACAAAGTGAGACCGTAACAAAAAAAAAAAAATTTAAAAAAAGAAATTTAGATCAAGATCCAACTGTAAAA AGTGGCCTAAACACCACATTAAAGAGTTTGGAGTTTATTCTGCAGGCAGAAGAGAACCATCAGGGGGTCT TCAGCATGGGAATGGCATGGTGCACCTGGTTTTTGTGAGATCATGGTGGTGACAGTGTGGGGAATGTTAT TTTGGAGGGACTGGAGGCAGACAGACCGGTTAAAAGGCCAAGCACAACAGATAAGGAGGAAGAAGATGAGG GCTTGGACCGAAGCAGAGAAGAGCAAACAGGGAAGGTACAAATTCAAGAAATATTGGGGGGTTGAATCA ACACATTTAGATGATTAATTAAATATGAGGACTGAGGAATAAGAAATGAGTCAAGGATGGTTCCAGGCTG CTAGGCTGCTTACCTGAGGTGGCAAAGTCGGGAGGAGTGGCAGTTTAGGACAGGGGGCAGTTGAGGAATA TTGTTTTGATCATTTTGAGTTTGAGGTACAAGTTGGACACTTAGGTAAAGACTGGAGGGGAAATCTGAAT ATACAATTATGGGGACTGAGGAACAAGTTTTTTATTTTTTGTTTCGTTTTCTTGTTGAAGAACAAATTT AATTGTAATCCCAAGTCATCAGCATCTAGAAGACAGTGGCAGGAGGTGACTGTCTTGTGGGTAAGGGTTT GGGGTCCTTGATGAGTATCTCTCAATTGGCCTTAAATATAAGCAGGAAAAGGAGTTTATGATGGATTCCA GGCTCAGCAGGGCTCAGGGAGGGCTCAGGCAGCCAGCAGAGGAAGTCAGAGCATCTTCTTTGGTTTAGCCCC AAGTAATGACTTCCTTAAAAAAGCTGAAGGAAAATCCAGAGTGACCAGATTATAAACTGTACTCTTGCATT TTCTCTCCCTCCTTCACCCACAGCCTCTTGATCGAACCGGAGGAAGTTTCTTTACCAATTCAAAAATGTC CGCTGGGCTAAGGGTCGGCGTGAGACCTACCTGTGCTACGTAGTGAAGAGGCGTGACAGTGCTACATCCT TTTCACTGGACTTTGGTTATCTTCGCAATAAGGTATCAATTAAAGTCGGCTTTGCAAGCAGTTTAATGGT CAACTGTGAGTGCTTTTAGAGCCACCTGCTGATGGTATTACTTCCATCCTTTTTTGGCATTTGTGTCTCT ATCACATTCCTCAAATCCTTTTTTTTATTTCTTTTTCCATGTCCATGCACCCATATTAGACATGGCCCAA AATATGTGATTTAATTCCTCCCCAGTAATGCTGGGCACCCTAATACCACTCCTTCCTTCAGTGCCAAGAA CAACTGCTCCCAAACTGTTTACCAGCTTTCCTCAGCATCTGAATTGCCTTTGAGATTAATTAAGCTAAAA GCATTTTTATATGGGAGAATATTATCAGCTTGTCCAAGCAAAAATTTTAAATGTGAAAAACAAATTGTGT CTTAAGCATTTTTGAAAATTAAGGAAGAAGAATTTGGGAAAAAATTAACGGTGGCTCAATTCTGTCTTCC AAATGATTTCTTTTCCCTCCTACTCACATGGGTCGTAGGCCAGTGAATACATTCAACATGGTGATCCCCA GAAAACTCAGAGAAGCCTCGGCTGATGATTAATTAAATTGATCTTTCGGCTACCCGAGAGAATTACATTT CCAAGAGACTTCTTCACCAAAATCCAGATGGGTTTACATAAACTTCTGCCCACGGGTATCTCCTCTCTCC TAACACGCTGTGACGTCTGGGCTTGGTGGAATCTCAGGGAAGCATCCGTGGGGTGGAAGGTCATCGTCTG GCTCGTTGTTTGATGGTTATATTACCATGCAATTTTCTTTGCCTACATTTGTATTGAATACATCCCAATC TCCTTCCTATTCGGTGACATGACACATTCTATTTCAGAAGGCTTTGATTTTCAAGCACTTTTCATTTAC TTCTCATGGCAGTGCCTATTACTTCTCTTACAATACCCATCTGTCTGCTTTACCAAAATCTATTTCCCCT TTTCAGATCCTCCCAAATGGTCCTCATAAACTGTCCTGCCTCCACCTAGTGGTCCAGGTATATTTCCACA ATGTACATCAACAGGCACTTCTAGCCATTTTCCTTCCAAAAGGTGCAAAAAGCAACTTCATAAACACA AATTAAATCTTCGGTGAGGTAGTGTGATGCTGCTTCCTCCCAACTCAGCGCACTTCGTCTTCCTCATTCC ACAAAAACCCATAGCCTTCCTTCACTCTGCAGGACTAGTGCTGCCAAGGGTTCAGCTCTACCTACTGGTG TGCTCTTTTGAGCAAGTTGCTTAGCCTCTCTGTAACAAGGACAATAGCTGCAAGCATCCCCAAAGATC ATTGCAGGAGACAATGACTAAGGCTACCAGAGCCGCAATAAAAGTCAGTGAATTTTAGCGTGGTCCTCTC TGTCTCTCCAGAACGGCTGCCACGTGGAATTGCTCTTCCTCCGCTACATCTCGGACTGGGACCTAGACCC TGGCCGCTGCTACCGCGTCACCTGGTTCACCTCCTGGAGCCCCTGCTACGACTGTGCCCGACATGTGGCC GACTTTCTGCGAGGGAACCCCAACCTCAGTCTGAGGATCTTCACCGCGCGCCTCTACTTCTGTGAGGACC GCAAGGCTGAGCCCGAGGGGCTGCGGCGGCTGCACCGCGCCGGGTGCAAATAGCCATCATGACCTTCAA AGGTGCGAAAGGGCCTTCCGCGCAGGCGCAGTGCAGCAGCCCGCATTCGGGATTGCGATGCGGAATGAAT GAGTTAGTGGGGAAGCTCGAGGGGAAGAAGTGGGCGGGGATTCTGGTTCACCTCTGGAGCCGAAATTAAA GATTAGAAGCAGAGAAAAGAGTGAATGGCTCAGAGACAAGGCCCCGAGGAAATGAGAAAATGGGGCCAGG GTTGCTTCTTTCCCCTCGATTTGGAACCTGAACTGTCTTCTACCCCCATATCCCCGCCTTTTTTTCCTTT TTTTTTTTTTGAAGATTATTTTTACTGCTGGAATACTTTTGTAGAAAACCACGAAAGAACTTTCAAAGCC TGGGAAGGGCTGCATGAAAATTCAGTTCGTCTCTCCAGACAGCTTCGGCGCATCCTTTTGGTAAGGGGCT TCCTCGCTTTTTAAATTTTCTTTCTTTCTCTACAGTCTTTTTTGGAGTTTCGTATATTTCTTATATTTTC TTATTGTTCAATCACTCTCAGTTTTCATCTGATGAAAACTTTATTTCTCCTCCACATCAGCTTTTTCTTC TGCTGTTTCACCATTCAGAGCCCTCTGCTAAGGTTCCTTTTCCCTCCCTTTTCTTTCTTTTGTTGTTTCA CATCTTTAAATTTCTGTCTCTCCCCAGGGTTGCGTTTCCTTCCTGGTCAGAATTCTTTTCTCCTTTTTTT TTTTTTTTTTTTTTTTTTTTAAACAAACAAACAAAAAACCCAAAAAAACTCTTTCCCAATTTACTTTCTT CCAACATGTTACAAAGCCATCCACTCAGTTTAGAAGACTCTCCGGCCCCACCGACCCCCAACCTCGTTTT GAAGCCATTCACTCAATTTGCTTCTCTCTTTCTCTACAGCCCCTGTATGAGGTTGATGACTTACGAGACG CATTTCGTACTTTGGGACTTTGATAGCAACTTCCAGGAATGTCACACACGATGAAATATCTCTGCTGAAG ACAGTGGATAAAAAAACAGTCCTTCAAGTCTTCTCTGTTTTTATTCTTCAACTCTCACTTTCTTAGAGTTT ACAGAAAAATATTTATATACGACTCTTTAAAAAGATCTATGTCTTGAAAATAGAGAAGGAACACAGGTC TGGCCAGGGACGTGCTGCAATTGGTGCAGTTTTGAATGCAACATTGTCCCCTACTGGGAATAACAGAACT GCAGGACCTGGGAGCATCCTAAAGTGTCAACGTTTTTCTATGACTTTTAGGTAGGATGAGAGCAGAAGGT AGATCCTAAAGCATGGTGAGAGGATCAAATGTTTTTATATCAAACATCCTTTATTATTTGATTCATTTG AGTTAACAGTGGTGTTAGTGATAGATTTTTCTATTCTTTTCCCTTGACGTTTACTTTCAAGTAACACAAA CTCTTCCATCAGGCCATGATCTATAGGACCTCCTAATGAGAGTATCTGGGTGATTGTGACCCCAAACCAT CTCTCCAAAGCATTAATATCCAATCATGGCCTGTAGTTTTAATCAGCAGAAGCATGTTTTTATGTTTGT ACAAAAGAAGATTGTTATGGGTGGGGATGGAGGTATAGACCATGCATGGTCACCTTCAAGCTACTTTAAT AAAGGATCTTAAAATGGGCAGGAGGACTGTGAACAAGACACCCTAATAATGGGTTGATGTCTGAAGTAGC AAATCTTCTGGAAACGCAAACTCTTTTAAGGAAGTCCCTAATTTGAAACACCCACAAACTTCACATATC ATAATTAGCAAACAATTGGAAGGAAGTTGCTTGAATGTTGGGGAGAGGAAAATCTATTGGCTCTCGTGGG TCTCTTCATCTCAGAAATGCCAATCAGGTCAAGGTTTGCTACATTTTGTGTGTGTGATGCTTCTCCCA AAGGTATATTAACTATATAAGAGTTGTGACAAAACAGAATGATAAAGCTGCGAACCGTGGCACACGCT CATAGTTCTAGCTGCTTGGGAGGTTGAGGAGGGAGGATGGCTTGAACACAGGTGTTCAAGGCCAGCCTGG GCAACATAACAAGATCCTGTCTCTCAAAAAAAAAAAAAAAAAAGAAAGAGAGGGCCGGGCGTGGTG GCTCACGCCTGTAATCCCAGCACTTTGGGAGGCCGAGCCGGGCGGATCACCTGTGGTCAGGAGTTTGAGA CCAGCCTGGCCAACATGGCAAAACCCCGTCTGTACTCAAAATGCAAAAATTAGCCAGGCGTGGTAGCAGG CACCTGTAATCCCAGCTACTTGGGAGGCTGAGGCAGGAGAATCGCTTTGAACCCAGGAGGTGGAGGTTGCA GTAAGCTGAGATCGTGCCGTTGCACTCCAGCCTGGGCGACAAGAGCAAGACTCTGTCTCAGAAAAAAAAA AAAAAAAGAGAGAGAGAGAAAGAGAACAATTTGGGAGAGAAGGATGGGAAGCATTGCAAGGAAAT TGTGCTTTATCCAACAAAATGTAAGGAGCCAATAAGGGATCCCTATTTGTCTCTTTGGTGTCTATTTGT CCCTAACAACTGTCTTTGACAGTGAGAAAAATATTCAGAATAACCATATCCCTGTGCCGTTATTACCTAG CAACCCTTGCAATGAAGATGAGCAGATCCACAGGAAAACTTGAATGCACAACTGTCTTATTTTAATCTTA TTGTACATAAGTTTGTAAAAGAGTTAAAAATTGTTACTTCATGTATTCATTTATATTTTATATTATTTTG CGTCTAATGATTTTTTATTAACATGATTTCCTTTTCTGATATATTGAAATGGAGTCTCAAAGCTTCATAA ATTTATAACTTTAGAAATGATTCTAATAACAACGTATGTAATTGTAACATTGCAGTAATGGTGCTACGAA GCCATTTCTCTTGATTTTTAGTAAACTTTTATGACAGCAAATTTGCTTCTGGCTCACTTTCAATCAGTTA AATAAATGATAAATAATTTTGGAAGCTGTGAAGATAAAATACCAAATAAAATAATATAAAAGTGATTTAT ATGAAGTTAAAATAAAAAATCAGTATGATGGAATAAACTTG

[0084] Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like (APOBEC) mRNA editing enzyme, catalytic polypeptide-like, The deaminase family. Members of this family are C to U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain. APOBEC family members are important for cytidine deamination. OBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now known as "APOBEC3E" (refers to APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) Deaminases include, but are not limited to, SaBE3, SaKKH-BE3, VQR-BE3, and EQR-B Numerous modifications including E3, VRER-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3 Cytidine deaminases are commercially available and are available from Addgene (Pla Sumido 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177).

[0085] Other exemplary deaminases that can be fused to Cas9 according to embodiments of the present disclosure are provided below. In some embodiments, the active domain of each sequence, e.g., the localization signal, Domains without nulls (nuclear localization sequence, nuclear export signal, cytoplasmic localization signal) It should be understood that it can be used.

[0086] Human AID: JPEG2025028857000006.jpg30169 (Underline: nuclear localization sequence; Double underline: nuclear export signal)

[0087] Mouse AID: JPEG2025028857000007.jpg29169 (Underline: nuclear localization sequence; Double underline: nuclear export signal)

[0088] Canine AID: JPEG2025028857000008.jpg28169 (Underline: nuclear localization sequence; Double underline: nuclear export signal)

[0089] Bovine AID: JPEG2025028857000009.jpg29169 (Underline: nuclear localization sequence; Double underline: nuclear export signal)

[0090] Rat AID JPEG2025028857000010.jpg29169 (Underline: nuclear localization sequence; Double underline: nuclear export signal)

[0091] Mouse APOBEC-3 JPEG2025028857000011.jpg45167 (italics: nucleic acid editing domain)

[0092] Rat APOBEC-3: JPEG2025028857000012.jpg45167 (italics: nucleic acid editing domain)

[0093] Rhesus macaque APOBEC-3 G: JPEG2025028857000013.jpg41167 (italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0094] Chimpanzee APOBEC-3 G: JPEG2025028857000014.jpg38167 (italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0095] Green monkey APOBEC-3G: JPEG2025028857000015.jpg40167 (italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0096] Human APOBEC-3G: JPEG2025028857000016.jpg38168 (italics: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0097] Human APOBEC-3F: JPEG2025028857000017.jpg38169 (italics: nucleic acid editing domain)

[0098] Human APOBEC-3B: JPEG2025028857000018.jpg38169 (italics: nucleic acid editing domain)

[0099] Rat APOBEC-3B: MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRK QGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCR LIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRY YRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILR IYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL

[0100] Bovine APOBEC-3B: DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQL KQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHW IKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLRQRILTAPI

[0101] Chimpanzee APOBEC-3B: MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQL SAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENF VYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCN EAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYK EALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPQSPGP CLPLCSEPPLGSLLPTGRPAPSLPFLLTASSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWA SWEDISH

[0102] Human APOBEC-3C: JPEG2025028857000019.jpg 20169

[0103] Gorilla APOBEC3C JPEG2025028857000020.jpg20169 (italics: nucleic acid editing domain)

[0104] Human APOBEC-3A: JPEG2025028857000021.jpg20169 (italics: nucleic acid editing domain)

[0105] Rhesus macaque APOBEC-3A: JPEG2025028857000022.jpg20169 (italics: nucleic acid editing domain)

[0106] Bovine APOBEC-3A: JPEG2025028857000023.jpg20169 (italics: nucleic acid editing domain)

[0107] Human APOBEC-3H: JPEG2025028857000024.jpg20169 (italics: nucleic acid editing domain)

[0108] Rhesus macaque APOBEC-3H: MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYL TWSPCPSCAGELVDFIKAHRHLNLRIFASRLYYHWRPNYQEGLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPS EKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR

[0109] Human APOBEC-3D: JPEG2025028857000025.jpg37169 (italics: nucleic acid editing domain)

[0110] Human APOBEC-1: MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSM SCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR

[0111] Mouse APOBEC-1 : MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNT RCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYP PSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK

[0112] Rat APOBEC-1 : MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK

[0113] Human APOBEC-2: MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKG GQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEP EIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK

[0114] Mouse APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKG GQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0115] Rat APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0116] Bovine APOBEC-2: MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEP EIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0117] Petromyzon marinus CDA1 (pmCDAl) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLR DNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCR KIFIQSSHNQLNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV

[0118] Human APOBEC3G D316R D317R MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKL HRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKFNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGF LCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTC FTSWSPCFSCAQEMAKFISKKHVSLCIFTARIYRRQGRC QEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN

[0119] Human APOBEC3G chain A MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDY RVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISF TYSEFKHCWDTFVDHQGCP FQPWDGLD EHSQDLSGRLRAILQ

[0120] Human APOBEC3G chain A D120R D121R MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQD YRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGC PFQPWDGLDEHSQDLSGRLRAILQ

[0121] The term "deaminase" or "deaminase domain" refers to a protein that catalyzes a deamination reaction. It refers to a protein or a fragment thereof.

[0122] "Detect" refers to identifying the presence, absence, or amount of the analyte to be detected. In one embodiment, sequence variations in a polynucleotide or polypeptide are detected. In another embodiment, the presence of indels is detected.

[0123] A "detectable label" means a label that, when attached to a molecule of interest, is detectable by spectroscopic, photochemical, or biochemical means. By "detectable" is meant a composition that renders it detectable through biological, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent Dyes, electron-dense reagents, enzymes (e.g., those commonly used in ELISA), biotin, zygotes xygenin, or hapten.

[0124] "Disease" means any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ. In certain embodiments, a disease or disorder suitable for treatment with the compositions of the present invention is may be due to point mutations, splicing events, premature stop codons, or misfolding events. Related to elephants.

[0125] "DNA-binding protein domain" means a polypeptide or fragment thereof that binds to DNA. In some embodiments, the DNA binding protein domain exhibits sequence-specific DNA binding. In another embodiment, the DNA is a zinc finger or TALE domain having binding activity. The A-binding protein domain is the domain of the CRISPR-Cas protein that binds to DNA (e.g. Cas9), including those that bind to protospacer adjacent motifs (PAMs). In some embodiments, the DNA binding protein domain binds to a polynucleotide (e.g., a single guide The complex is bound to the gRNA and the protospacer adjacent motif. In some embodiments, the DNA binding protein domain binds to a DNA sequence identified by the They contain enzyme activity (e.g., nCas9) or are catalytically inactive (e.g., dCas9, diCas9). In yet another embodiment, the DNA binding protein domain The main one is a catalytically inactive variant of the homing endonuclease I-SceI, or The DNA-binding domain of the TALE protein AvrBs4. See, e.g., Gabsalilow et al., Nucleic Acids See ids Research, Volume 41, Issue 7, 1 April 2013, Pages e83. The DNA binding protein domain is fused to a catalytic domain (e.g., FokI, MutH). In certain embodiments, the zinc finger domain is an endonuclease In another embodiment, the TALE is fused to the catalytic domain of FokI. It is fused to MutH, which contains the KING activity.

[0126] As used herein, the term "effective amount" refers to an amount sufficient to elicit a desired biological response. In certain embodiments, an effective amount refers to the amount of a biologically active agent transfected with a plasmid. a base editor sufficient to express an active base editing system in the transfected cell. - the amount of two or more plasmids that comprise parts of the system. As will be understood by those skilled in the art, An effective amount of an agent (e.g., a fusion protein) can be, for example, a dose that induces a desired biological response, e.g., editing. The specific allele, genome, or target site to be targeted, the cells or tissues to be targeted, and The time to delivery may vary depending on various factors, such as the type of medication and the type of drug used.

[0127] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, which portion is identical to that of a reference nucleic acid At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the entire length of the molecule or polypeptide , or 90%. The fragments may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 nucleotides or amino acids. do.

[0128] "Hybridization" means hydrogen bonding between complementary nucleobases, as defined by Watson- It can be a Crick, Hoogsteen or reversed Hoogsteen hydrogen bond. For example, Adenine and thymine are complementary nucleobases that form hydrogen bonds to form pairs.

[0129] The term "inhibitor of base repair" or "IBR" refers to an inhibitor of base repair by a nucleic acid repair enzyme, such as base excision repair. IBR refers to a protein that can inhibit the activity of an enzyme. It is an inhibitor of base excision repair. Examples of base repair inhibitors include APE1 and Endo III. , Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL and inhibitors of hAAG. In certain embodiments, the IBR is an inhibitor of Endo V or hAAG. In one embodiment, the IBR is a catalytically inactive EndoV or a catalytically inactive It is hAAG.

[0130] An "intein" excises itself and releases the remaining fragment (the extein) (extein)) to form peptide bonds in a process known as protein splicing Inteins are fragments of proteins that can be linked together using a "protein intron." The intein excises itself and links the rest of the protein. The process is referred to herein as "protein splicing" or "intein-mediated In some embodiments, precursor proteins (integrins) are spliced. Inteins (intein-containing proteins before intein-mediated protein splicing) Such inteins are referred to herein as split inteins. These are called split inteins (e.g., split intein-N and split intein-C). In Bacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is expressed in two separate genes. Encoded by the genes dnaE-n and dnaE-c. Encoded by the dnaE-n gene The intein may be referred to herein as "intein N." The loaded intein may be referred to herein as "intein C."

[0131] Other intein systems can also be used. For example, the dnaE intein, i.e., Cfa-N (e.g., Based on the intein pair Cfa-C (e.g., split intein-N) and Cfa-C (e.g., split intein-C) Synthetic inteins have been described (see, e.g., Stevens, J. Med. Chem. Soc. 1999, 144:111-112, which is incorporated herein by reference). s et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Use in accordance with this disclosure Non-limiting examples of intein pairs that can be used include Cfa DnaE intein, Ssp GyrB intein, and Intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein Cne intein, and Cne Prp8 intein (see, e.g., U.S. Pat. No. 6,223,629, incorporated herein by reference). Examples include those described in Patent No. 8,394,604.

[0132] Exemplary nucleotide and amino acid sequences of inteins are provided.

[0133] DnaE intein-N DNA:TGCCTGTCATACGAACCGAGATACTGACAGTAGAATATGGCCTTCTGCC AATCGGG AAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCG ATAACAATGGTAACATTTATACTCAGCCAGTTGCCC AGTGGCACGACCGG GGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAG GGCCACTAAGGACC ACAAATTTATGACAGTCGATGGCCAGATGCTGCCTA TAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGA CAACCTT CCTAAT

[0134] DnaE Intein-N Protein:CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN

[0135] DnaE intein-C DNA:ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGA GTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT

[0136] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0137] Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCC TATTGGAAAGATTGTCGAAGAGAGAATTG AATGCACAGTATATACTGTAG ACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGC GGCGAAC AAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACG AGCAACTAAAGATCATAAATTCATGACCACTGACGG GCAGATGTTGCCAA TAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTG CCA

[0138] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNR GEQEVFEYCLEDGSIIRATKDHKFMTTDG QMLPIDEIFERGLDLKQVDGL P

[0139] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAG GAAAGTAAAGATAATATCTCGAAAAAGTC TTGGTACCCAAAATGTCTATG ATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTA GCCAGCA AC

[0140] Cfa-C protein:MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLV ASN

[0141] To join the N-terminal part of the split Cas9 with the C-terminal part of the split Cas9, intein N and intein B were used. The tein C can be fused to the N-terminal end of the split-Cas9 and the C-terminal end of the split-Cas9, respectively. For example, In some embodiments, intein-N is C-terminal to the N-terminal portion of the split Cas9. The split-Cas9 is fused to the intein-N domain, forming the structure N--[N-terminal portion of split-Cas9]-[intein-N]--C. In some embodiments, the intein-C is located at the N-terminus of the C-terminal portion of the split Cas9. The intein is fused to the C-terminal portion of the split Cas9, forming the structure N-[intein-C]-[C-terminal portion of the split Cas9]-C. Intein for linking the protein (e.g., split-Cas9) to which the intein is fused The mechanism of tein-mediated protein splicing is described, for example, in the publications As described in Shah et al., Chem Sci. 2014; 5(1):446-461, the present invention Methods for designing and using inteins are known in the art. and are disclosed, for example, in WO2014004336, WO2017132580, US20150344549 and US20180127780. and JP 2003-1026634, each of which is incorporated herein by reference in its entirety.

[0142] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its native state. Substances that have been removed to varying degrees from components normally associated with them when found in the natural state. "Isolated" refers to the degree of separation from the original source or surrounding environment. "Purified" refers to the degree of separation from the original source or surrounding environment. A "purified" or "biologically pure" protein is one that is free from impurities. that the substance does not materially affect the biological properties of the protein or cause other adverse consequences. In other words, the nucleic acid or peptide of the present invention is or, if produced by recombinant DNA technology, cellular material, viral material, or culture medium. If the substance is not naturally contained in the substance, or if it is chemically synthesized, chemical precursors or other chemical substances Purity and homogeneity are typically determined by analytical methods. Chemical techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography The term "purified" refers to the degree to which a nucleic acid or protein is purified by electrophoresis. This can mean that the resulting protein essentially produces one band. For proteins that can undergo glycosylation, the different modifications are purified separately. This can result in different isolated proteins that can be isolated.

[0143] An "isolated polynucleotide" is a nucleic acid molecule that does not occur in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. It means nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene. The term refers to, for example, those incorporated into vectors; autonomously replicating plasmids or viruses. integrated into the genomic DNA of prokaryotes or eukaryotes; or independently of other sequences another molecule (e.g., cDNA generated by PCR or restriction endonuclease digestion) or genomic or cDNA fragments). RNA molecules transcribed from DNA molecules, as well as hybrids encoding additional polypeptide sequences. It contains recombinant DNA that is part of the hybrid gene.

[0144] An "isolated polypeptide" is a polypeptide of the invention separated from components that naturally accompany it. Typically, a polypeptide is a polypeptide that is a protein with which it is naturally associated. A substance is isolated if it is at least 60% by weight free from substances and naturally occurring organic molecules. Preferably, the preparation comprises at least 75% by weight, more preferably at least 90%, and most preferably Preferably, at least 99% of the isolated polypeptides of the present invention are polypeptides of the present invention. A peptide can be prepared by, for example, extraction from a natural source, recombinant nucleic acid encoding such a polypeptide, or the like. or by chemically synthesizing the protein. Any suitable method, e.g., column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.

[0145] As used herein, the term "linker" refers to a molecule that connects two molecules or moieties (e.g., a fusion protein). refers to a bond (e.g., a covalent bond), chemical group, or molecule that connects two domains of a protein (e.g., two domains of a protein). In one embodiment, the linker is an RNA program containing a Cas9 nuclease domain. the gRNA-binding domain of a ramming-capable nuclease and the catalytic domain of a nucleic acid-editing protein In some embodiments, a linker connects the dCas9 and the nucleic acid editing protein. Typically, a linker is positioned between or connects two groups, molecules, or other moieties. are adjacent to each other and linked to each other via covalent bonds, thus forming the two In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., In some embodiments, the linker is an organic molecule, group, polymer, or peptide. In some embodiments, the linker is a polymer or chemical moiety. Acids, e.g., lengths of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102 , 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids Longer or shorter linkers are also contemplated. In some embodiments, the linker is In one embodiment, the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker, is included. In some embodiments, the linker comprises the amino acid sequence SGGS. n , (GG GS) n , (GGGGS) n , (G) n、 (EAAAK) n , (GGS) n , SGSETPGTSESATPES, or (XP) n motif or any combination thereof, wherein n is independently an integer from 1 to 30; and X In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. , 10, 11, 12, 13, 14, or 15.

[0146] In some embodiments, the nucleobase editor domain is SGGSSGSETPGTSESATP ESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPG TSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS Ami In some embodiments, the nucleobase is fused via a linker comprising a nucleobase sequence. The editor domain contains the amino acid sequence SGSETPGTSESATPES, which may also be called the XTEN linker. In one embodiment, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. , comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. The linker is 64 amino acids in length. In some embodiments, the linker has the amino acid sequence SG GSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In one embodiment, the linker is 92 amino acids in length. Acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GTSTEPSEGS Contains APGTSESATPESGPGSEPATS.

[0147] A "marker" is any molecule that has an altered expression level or activity that is associated with a disease or disorder. The term "protein" refers to any protein or polynucleotide.

[0148] As used herein, the term "mutation" refers to a change in a sequence, e.g., a nucleic acid or amino acid sequence. The substitution of a residue in a sequence of amino acids by another residue, or the deletion of one or more residues in the sequence. Mutations, as used herein, are typically made by identifying the original residue and then and identifying the newly substituted residue, Various methods for making amino acid substitutions (mutations) are provided herein. are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).

[0149] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a nucleic acid molecule comprising a nucleobase and Compounds containing an acidic moiety, such as nucleosides, nucleotides, or polynucleotides Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides is a linear sequence in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or or nucleosides). In some embodiments, a "nucleic acid" refers to three or more individual nucleosides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing nucleotide residues. "Nucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., A chain of at least three nucleotides may be used interchangeably to refer to a nucleotide sequence. "Nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids are, for example, , genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, It may naturally occur in the context of a chromatid or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules include, for example, non-naturally occurring molecules, recombinant DNA or RNA, artificial chromosomes, engineered engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or is a non-naturally occurring molecule containing a non-naturally occurring nucleotide or nucleoside Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms refer to nucleic acids. Nucleic acids include those derived from natural sources, such as those derived from phosphodiesters. purified from, produced using a recombinant expression system and optionally purified, chemically synthesized In the case of chemically synthesized molecules, nucleic acids may, where appropriate, be Nucleotides, such as analogs with chemically modified bases or sugars, and backbone modifications, are also suitable. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. In some embodiments, nucleic acids contain natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine cytidine, and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine, 2-thiocytidine, Othymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2 -Aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5 -propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadeno sine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine guanine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, 2'-deoxyribose, 2'-deoxyribose, arabinose, and hexose); and / or modified Is there a decorative phosphate group (e.g., phosphorothioate and 5'-N-phosphoramidite linkage)? , or including them.

[0150] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to the localization of a protein. The term "nuclear localization sequence" refers to an amino acid sequence that promotes import into the cell nucleus. For example, WO / 2001 / 038547 filed on November 23, 2000 and published on May 31, 2001 is incorporated herein by reference. and Plank et al. in published international PCT application PCT / EP 2000 / 011690, in which No. 6,239,493, which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In embodiments, the NLS may be any of the NLSs described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4 172. In some embodiments, the NLS is Sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR , RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0151] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to the Associating with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid, that directs the Abp to a specific nucleic acid sequence For example, the Cas9 protein refers to a protein that binds the Cas9 protein to a guide RNA. It may bind to a guide RNA that directs it to a complementary specific DNA sequence. , napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 nickase ( nCas9), or nuclease-inactivated Cas9 (dCas9). Examples of binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, and C2. Other nucleic acid programmable DNA fragments include, but are not limited to, c1, and C2c3. A binding proteins are also within the scope of this disclosure even if not specifically listed in this disclosure. .

[0152] As used herein, "obtaining," as in "obtaining a drug," means This includes synthesizing, purchasing, or otherwise obtaining the drug.

[0153] As used herein, a "patient" or "subject" refers to a person who has been diagnosed with a disease or disorder. or suspected of having or developing them, such as a mammalian subject. In certain embodiments, the term "patient" refers to a person who has a higher than average likelihood of developing a disease or disorder. Refers to a mammalian subject. Exemplary patients include humans, non-human primates, cats, dogs, pigs, and cats. Dogs, cats, horses, goats, sheep, rodents (e.g., mice, rabbits, rats, guinea pigs) and other mammals that may benefit from the treatments disclosed herein. Exemplary human patients can be male and / or female. A "subject in need thereof," as used herein, is a person who has been diagnosed with a disease or disorder. or referred to as a patient suspected of having such a condition.

[0154] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to Used in conjunction with (e.g., bound to or associated with) one or more non-target RNAs In some embodiments, the RNA-programmable nuclease, when complexed with RNA, Typically, the bound RNA is called a guide RNA (gRNA). gRNAs can exist as a complex of two or more RNAs, or as a single RNA molecule. A gRNA that exists as a single RNA molecule is called a single guide RNA (sgRNA). Although sometimes referred to as "gRNA," "gRNA" can occur as a single molecule or as a complex of two or more molecules. are used interchangeably to refer to guide RNAs present as a single RNA species. The gRNA present in the target nucleic acid has (1) a domain that shares homology with the target nucleic acid (e.g., a domain that facilitates Cas9 duplication to the target). and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) is directed against a sequence known as tracrRNA. For example, in some embodiments, domain (2) comprises: Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. The gRNA (e.g., domain) is identical to or homologous to the tracrRNA provided in the Another example of a method for detecting nucleotides containing Cas9 (including those containing nucleotides containing Cas9) is the "Switchable Cas9 Nucleases and Uses Thereof" U.S. Provisional Patent Application USSN 61 / 874,682 filed September 6, 2013 and "Delivery System Fo U.S. Provisional Patent Application USSN 61 filed September 6, 2013, entitled "Functional Nucleases" / 874,746, the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2), and is referred to as an "extension" domain. By way of example, an extended gRNA may be referred to as a "longered gRNA" as described herein. , e.g., by binding two or more Cas9 proteins to target nucleic acids in two or more different regions. The gRNA contains a nucleotide sequence complementary to the target site, which binds to the target site. Mediates the binding of nuclease / RNA complexes to the target and determines the sequence specificity of the nuclease:RNA complex. In some embodiments, the RNA programmable nuclease is a CRISPR-associated system ) Cas9 endonuclease, e.g., Cas9 (Csnl) from Streptococcus pyogenes (For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes.") Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Na jar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, Mc Laughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA mat uration by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chy linski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J. , Charpentier E., Nature 471:602-607(2011)).

[0155] The term "recombinant" as used herein with respect to a protein or nucleic acid means that the protein or nucleic acid is It refers to proteins or nucleic acids that do not occur in nature but are the product of human engineering. For example, In some embodiments, the recombinant protein or nucleic acid molecule is any naturally occurring At least one, at least two, at least three, at least four, or at least at least five, at least six, or at least seven amino acid or nucleotide mutations Contains an octide sequence.

[0156] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100% .

[0157] "Reference" refers to a standard or control condition. In one embodiment, a reference Nucleobase editors expressed in intein-containing fragments for intein-dependent reassembly Full-length nucleobase edema expressed on a single plasmid in the same cells under the same conditions as The activity of the filter.

[0158] A "reference sequence" is a defined sequence used as a basis for sequence comparison. It can be a subset or the entirety of a particular sequence; for example, a full-length cDNA or gene sequence. A segment of a gene, or the complete cDNA or gene sequence. For polypeptides, see the reference polypeptide. The length of the peptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, At least about 25 amino acids, even more preferably about 35 amino acids, about 50 amino acids, or about 10 For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 amino acids. nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides nucleotides or about 300 nucleotides or any integer therebetween be.

[0159] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific position in the genome. where each mutation is present in the population to a noticeable extent (e.g., >1%). For example, At certain base positions in the human genome, C nucleotides can occur in most individuals, In a small number of individuals, the position is occupied by A. This means that there is a SNP at this particular position, and C A variation of two nucleotides, SN or A, is said to be an allele at this position. P underlies differences in susceptibility to disease. The severity of disease and the body's response to treatment also affect genetic SNPs are a representation of genetic variation. SNPs occur in the coding region of genes, non-coding regions of genes, and In some embodiments, the codon may be located in a nucleotide sequence, or in an intergenic region (the region between genes). SNPs in the gene sequence can affect the amino acid sequence of the protein produced due to the degeneracy of the genetic code. SNPs in coding regions are of two types: synonymous and non-synonymous. Synonymous SNPs do not affect the protein sequence, whereas non-synonymous SNPs change the amino acid sequence of a protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not located in the region affect gene splicing, transcription factor binding, and messenger RNA degradation. These SNPs can affect the sequence of genes involved in the transcription ... The gene expression that is expressed is called an eSNP (expressed SNP) and can be upstream or downstream of the gene. A single nucleotide variant (SNV) is a single nucleotide variation with unlimited frequency that can occur somatically. Somatic single base variations may also be referred to as single base modifications.

[0160] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. and bind to other molecules in the sample (e.g., biological sample) but do not substantially recognize or bind to other molecules in the sample. Nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA) A binding domain and guide nucleic acid), compound, or molecule.

[0161] Nucleic acid molecules useful in the methods of the present invention may encode a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. Typically, but not necessarily, substantial identity is shown. A polynucleotide having a nucleotide sequence typically hybridizes with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention can be modified to The present invention also includes any nucleic acid molecule encoding an endogenous nucleotide sequence or a fragment thereof. It need not be 100% identical to the nucleic acid sequence, but typically will show substantial identity. Polynucleotides having "substantial identity" to a double-stranded nucleic acid molecule typically have a small number of double-stranded nucleic acid molecules. "Hybridize" means to hybridize with at least one of the strands of a given molecule. Complementary polynucleotide sequences (e.g., those described herein) are synthesized under various stringency conditions. This means that a double-stranded molecule is formed between the two genes (or a part of them). For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (1987) Methods Enzymol. 152:507).

[0162] For example, a stringent salt concentration is typically less than about 750 mM NaCl and 75 mM citrate. trisodium, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably Preferably, the concentration is less than about 250 mM NaCl and 25 mM trisodium citrate. See hybridization can be obtained in the absence of organic solvents, e.g., formamide. whereas high stringency hybridization requires at least about 35% formaldehyde. It can be obtained in the presence of at least about 50% formamide. Stringent temperature conditions are typically at least about 30°C, more preferably at least about 37°C. 0 C, and most preferably at least about 42°C. The concentration of detergent (e.g., sodium dodecyl sulfate (SDS)) and carrier DNA content Various additional parameters, such as inclusion or exclusion, are well known to those skilled in the art. By combining these various conditions, various levels of stringency can be achieved. In one embodiment, hybridization is achieved at 30° C. in 750 mM NaCl, 75 In another embodiment, hybridization occurs in 10 mM trisodium citrate and 1% SDS. The solution was incubated at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide. In another embodiment, the hybridization occurs in 100 μg / ml denatured salmon sperm DNA (ssDNA). Hybridization was performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% The reaction takes place in SDS, 50% formamide, and 200 μg / ml ssDNA. The implementation will be readily apparent to one skilled in the art.

[0163] For most applications, the washing steps that follow hybridization are also stringent. Wash stringency conditions are defined by salt concentration and temperature. As mentioned above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. This can be increased by increasing the thickness of the strip for the cleaning process. The appropriate salt concentration is preferably less than about 30 mM NaCl and 3 mM trisodium citrate. and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the process are typically at least about 25°C, more preferably In one embodiment, the temperature is at least about 42°C, and even more preferably at least about 68°C. Washing steps were performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash steps are carried out in 15 mM NaCl, 1.5 mM NaCl, at 42°C. 10 mM trisodium citrate, and 0.1% SDS. Washing steps were performed at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton et al. nd Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wil ey Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular C loning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York It has been done.

[0164] "Split" means divided into two or more pieces.

[0165] A "split Cas9 protein" or "split-Cas9" is a protein that is composed of two separate nucleotides. Cas9 protein provided as N-terminal and C-terminal fragments encoded by the sequences The polypeptides corresponding to the N-terminal and C-terminal parts of the Cas9 protein are spliced ​​together. In certain embodiments, the Cas9 protein can be reconstituted to form a "reconstituted" Cas9 protein. The Cas9 protein is described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-9 49, 2014, or Jiang et al. (2016) Science 351: 867-871. PDB As described in file: 5F9R (each of which is incorporated herein by reference), The protein is divided into two fragments within a disordered region of the protein. The disordered region has been characterized by X-ray crystallography, NMR spectroscopy, and Spectroscopy, electron microscopy (e.g., cryo-EM), and / or in silico protein modeling one or more protein structure determination techniques known in the art, including but not limited to, In some embodiments, the protein can be determined by the method of approximately Any C, T, A, or or any other Cas9, Cas9 variant (e.g., nCas9, dCas9), if or at a corresponding position in another nap DNA bp. The protein split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting a protein into two fragments is , referred to as "splitting" the protein.

[0166] "Substantially identical" means that the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein) is substantially identical to the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein). any one of the sequences) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) It means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to In this embodiment, such sequences may be identical at the amino acid level or to the sequences used for comparison. have at least 60%, 80%, or 85%, 90%, 95%, or even 99% identity in the nucleic acid .

[0167] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE UP / PRETTYBOX program). Such software is available with various permutations. By assigning degrees of homology to the sequences, deletions, and / or other modifications, it is possible to determine whether the sequences are identical or Similar sequences are matched. Conservative substitutions typically include substitutions within the following groups: leucine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, Paragine, glutamine; serine, threonine; lysine, arginine; phenylalanine In an exemplary approach to determining the degree of identity, the BLAST program You can use the RAM, -3 and e -100 Probability scores between indicate closely related sequences .

[0168] A "subject" includes, but is not limited to, a cow, horse, dog, sheep, or cat. Subjects include livestock, animals that produce labor and food, and mammals that are either human or non-human. Includes domestic animals kept to provide food or other commodities, such as cattle, goats, chickens, and horses. , pigs, rabbits, and sheep.

[0169] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase or a fusion enzyme containing the same. Deaminated by a ligase protein (e.g., cytidine or adenine deaminase) do.

[0170] RNA-programmable nucleases (e.g., Cas9) use RNA:D to target DNA cleavage sites. Since NA hybridization is used, these proteins can, in principle, be used as guide RNAs. Any sequence specified by A can be targeted. For site-specific cleavage, C Methods using RNA-programmable nucleases such as as9 (e.g., to modify genomes) For this purpose, methods for generating multiplexed genomic DNA (e.g., Cong, L. et al., Multiplex gene expression) are known in the art (see, e.g., Cong, L. et al., Multiplex gene expression). ome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013 ); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programme d genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., G enome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. See Nature biotechnology 31, 233-239 (2013). the entire contents of each of which are incorporated herein by reference).

[0171] Ranges provided herein are understood to be shorthand for all values ​​within the range. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 , 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood to include any number, combination of numbers, or subrange.

[0172] As used herein, the terms "treat," "treating," and "treatment" refer to "Treatment" refers to the treatment or other treatment that reduces or improves the disorder and / or its associated symptoms. Treating a disorder or condition means completely alleviating the associated disorder, condition or symptoms. It will be understood that it is not necessary to completely remove the But not.)

[0173] The term "uracil glycosylase inhibitor" or "UGI" as used herein refers to a In this case, a protein capable of inhibiting uracil-DNA glycosylase base excision repair enzyme is In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In certain embodiments, the UGI proteins provided herein include fragments of UGI, and UGI or For example, in some embodiments, the UGI domain may be a UGI fragment. In some embodiments, the polypeptide comprises a fragment of the amino acid sequence provided below. UGI fragments may contain at least 60%, at least 65%, at least 80%, or at least 90% of the exemplary UGI sequences provided herein. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% In some embodiments, the UGI comprises an amino acid sequence comprising the amino acid sequence provided herein below. Amino acid sequences homologous to the amino acid sequences or amino acid sequences provided herein below In some embodiments, the amino acid sequence is homologous to a fragment of the UGI or U Proteins containing fragments of GI or homologs of UGI or fragments of UGI are referred to as "UGI variants." UGI variants are referred to as "UGI variants." UGI variants share homology with UGI or a fragment thereof. For example, UGI variants A variant has at least 70% identity, at least 70% identity to a wild-type UGI or a UGI described herein. at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, At least 98% identity, at least 99% identity, at least 99.5% identity, or at least In some embodiments, the UGI variant has at least 99.9% identity to the UGI and a fragment of UGI, which fragment is a fragment of wild-type UGI or a fragment of UGI provided below. At least 70% identity, at least 80% identity, at least 90% identity, at least at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity identity, at least 99% identity, at least 99.5% identity, or at least 99.9% identity In some embodiments, the UGI comprises the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SD APE YKPW ALVIQDS NGENKIKML

[0174] Unless otherwise stated or clear from the context, as used herein, " The term "or" is understood to be inclusive. Unless otherwise clear, as used herein, the terms "a," "an," and "the" refer to It is understood to be singular or plural.

[0175] As used herein, unless otherwise stated or clear from the context, The term "about" means within normal tolerance in the art, e.g., within 2 standard deviations of the mean. Approximately means 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5% of the stated value. %, 0.1%, 0.05%, or 0.01%. Unless otherwise clear from the context, All numerical values ​​provided herein are modified by the term about.

[0176] The recitation of a list of chemical groups in a variable definition herein is intended to encompass any single This includes the definition of the variables as a group or combination of the listed groups. The description of an embodiment with respect to a variable or aspect does not necessarily mean that it should be taken as any single embodiment or as a single or multiple embodiment. or any other embodiment or portion thereof.

[0177] The compositions or methods provided herein may be used in combination with any other compositions and methods provided herein. The method may be combined with one or more of the following methods: [Brief explanation of the drawings]

[0178] [Figure 1] Figure 1 shows a schematic diagram of the A→G base editor (ABE) fusion protein nucleobase editor, which contains two adenosine deaminase domains, wild-type (wt) TadA and evolved (evo) TadA, fused to the S. pyogenes (Sp) Cas9 nickase (nCas9) with a C-terminal bipartite nuclear localization signal (NLS). The diagram also identifies three regions of the unstructured Cas9 protein that can be split into N- and C-terminal fragments that can be reconstituted using a split-intein system (i.e., fusing intein-N and intein-C to the N- and C-terminal fragments, respectively).

[0179] [Figure 2]Figure 2 provides three graphs quantifying the base editing activity of a base editing system containing a spliced ​​nCas9 nucleobase editor fusion protein (i.e., ABE), reproducing the fusion protein described in Figure 1. The ABE was split at the indicated amino acid positions (e.g., T310, T313, A456, S469, and C574 relative to the SpCas9 amino acid sequence). The N- and C-terminal fragments of the ABE were fused to intein-N and intein-C, respectively. These fragments and the indicated guide RNAs were expressed on separate plasmids in cultured HEK293 cells expressing a protein containing the ABCA4 gene with the 5882G>A mutation. The base editing activity of the reconstituted ABE against the ABCA4 5882A>G target was compared to that of a control ABE. Base editing activity was dependent on the presence of both the N- and C-terminal fragments of the ABE. No base editing activity was observed when only the N- or C-terminal fragments of the ABE were expressed.

[0180] [Figure 3] Figure 3 is a graph confirming base editing activity with the 21-nt guide described in Figure 2. Note that the experiment in Figure 3 was performed in a different format than the experiment in Figure 2. All reconstituted ABEs exhibited good base editing activity. This activity was dependent on the presence of both the N- and C-terminal fragments of nCas9. No activity was observed when only the N- or C-terminal fragments of nCas9 were expressed. The activity of the reconstituted ABEs was compared to that of the control ABE7.09 and ABE7.10 fusion proteins.

[0181] [Figure 4] Figure 4 is a graph quantifying the base editing activity described in Figure 2. In Figure 4, a 20-nucleotide (nt) guide RNA containing a hammerhead ribozyme (HRz) was used. The activity of the various reconstituted ABEs was compared to that of the control ABE7.09 and ABE7.10 fusion proteins.

[0182] [Figure 5]Figures 5A-5D show the determination of the multiplicity of infection (MOI) for AAV2 coinfection of ARPE-19 cells. Figure 5A is a graph showing 1:1 coinfection of AAV2 / CMV-mCherry and AAV2 / CMV-EmGFP at various viral loads (vg / cell). Figure 5B shows fluorescent images detecting EmGFP (left) and mCherry (center), as well as an overlay image (right) showing the colocalization of EmGFP and mCherry. Figure 5C is a graph showing the percentage of cells expressing mCherry at various viral loads (vg / cell) when infected with AAV2 / CMV-mCherry. Figure 5D is a graph showing the percentage of cells expressing EmGFP at various viral loads (vg / cell) when infected with AAV2 / EmGFP.

[0183] [Figure 6] Figure 6 is a series of graphs showing that delivery of split editors into ARPE-19 cells via dual AAV2 infection results in high A>G conversion rates in ABCA4 5882A. Multiplicities of infection (MOI) for dual infection are shown: 20,000 vg / cell (top left); 30,000 vg / cell (top right); 40,000 vg / cell (bottom left); and 60,000 vg / cell (bottom right). DETAILED DESCRIPTION OF THE INVENTION

[0184] As described below, the present invention provides compositions and methods for delivering base editing systems. The present invention provides a method for translating a split intein into a nucleic acid sequence using an A-to-G nucleobase editor (ABE). This is based, at least in part, on the discovery that a segmented image can be "split" and reconstructed using a split image. N- and C-terminal fragments of ABE fused to the in-pair intein-N and intein-C, respectively The polynucleotides encoding the nucleotides are delivered to the cell on separate vectors along with the single guide RNA. The encoded ABE fragments are spliced ​​together to form the targeted ABE fragment of the nucleic acid sequence. and reconstituted functional nucleobase editor fusion proteins that are particularly useful for targeted editing. Ta.

[0185] [Intein] Intein ( in tervening pro tein ) is a self-processing enzyme found in a wide variety of organisms. domain, which is responsible for the process known as protein splicing Protein splicing is a multi-step biochemical process that involves both the cleavage and formation of peptide bonds. Endogenous substrates for protein splicing are found in intein-containing organisms. Although inteins are proteins that are expressed in a variety of ways, they can also be used to create virtually any polypeptide backbone. It can also be used to chemically manipulate.

[0186] In protein splicing, inteins cleave two peptide bonds. Thus, it excises itself from one precursor polypeptide, thereby forming the adjacent extein. (external protein) sequences are linked through the formation of new peptide bonds. Occurs post-translationally (but may also occur co-translationally). Intein-mediated protein synthesis Issuing occurs spontaneously and requires only the folding of the intein domain.

[0187] Approximately 5% of inteins are split inteins, which are called N-inteins and C-inteins. are transcribed and translated as two separate polypeptides, each of which contains one extein. During translation, the intein fragment spontaneously and noncovalently binds to the standard intein. They assemble into a trans-protein structure and perform protein splicing in trans. The mechanism of splicing involves a series of acyl transfer reactions, which result in the formation of intein-encoding nucleotides. Cleavage of the two peptide bonds at the exteme junction and the N-exteme and C-exteme This process results in the formation of a new peptide bond between the N-extein and It is initiated by the activation of a peptide bond that connects the N-terminus of the intein. All inteins have an N-terminal cysteine ​​or serine, which is connected to the C-terminus of the N-extein. This N-to-O / S acyl group transfer is a conserved transcription factor. leonine and histidine (called the TXXH motif), and the commonly found asparagus This is promoted by carboxylic acid, leading to the formation of a linear (thio)ester intermediate. The intermediates are those where the first residue of the C-extein is a cysteine, serine, or threonine ( +1) undergoes trans-(thio)esterification. The resulting branched (thio)ester The intermediate is a unique cyclization of the highly conserved C-terminal asparagine of inteins. This process is reversed by the histidine (in the highly conserved HNF motif). It is promoted by the penultimate histidine (as found in the ribosomal ribosomal nucleotides) and the penultimate histidine, and aspartic acid is also involved. This succinimide-forming reaction excises the intein from the reaction complex, allowing the non-peptide to be released. This structure leaves the exteins linked via a tide bond. rapidly rearranges to form a stable peptide bond.

[0188] [Adenosine deaminase] In some embodiments, the fusion proteins of the invention comprise an adenosine deaminase domain. In some embodiments, the adenosine deaminase provided herein comprises can deaminate adenine. The provided adenosine deaminase removes adenine from deoxyadenosine residues in DNA. Adenosine deaminase can be derived from any suitable organism, such as the large intestine. In some embodiments, the adenine deaminase can be derived from any of the enzymes described herein. One or more mutations corresponding to any of the provided mutations (e.g., mutations in ecTadA) A person skilled in the art can easily identify adenosine deaminase by, for example, the sequence By alignment and determination of homologous residues, the corresponding Thus, one skilled in the art can identify the residues of any naturally occurring adenosine dehydrogenase. The present invention relates to a method for the detection of tyrosine kinases (e.g., those with homology to ecTadA) and a method for the detection of tyrosine kinases (e.g., those with homology to ecTadA). corresponding to any of the mutations identified in ecTadA Mutations can be generated. In some embodiments, the adenosine deaminase is prokaryotic. In some embodiments, the adenosine deaminase is of bacterial origin. In bacteria, adenosine deaminase is expressed in Escherichia coli, Staphylococcus aureus, S. almonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter cr In some embodiments, the adenosine deaminase is derived from Bacillus escentus or Bacillus subtilis. The enzyme is derived from Escherichia coli.

[0189] In one embodiment, the fusion protein of the invention comprises wild-type TadA linked to TadA7.10. In certain embodiments, the fusion protein comprises a single The TadA7.10 domain comprises one TadA7.10 domain (e.g., provided as a monomer). In this study, the ABE7.10 editor was found to be capable of forming heterodimers with TadA7.10 and Tad A (wt). The relevant sequences are:

[0190] TadA(wt): SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATLY VTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD

[0191] TadA7.10: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD

[0192] In some embodiments, TadA (e.g., with double-stranded substrate activity) or Ta dA7.10 is provided as a homodimer or a monomer.

[0193] In some embodiments, the adenosine deaminase is an adenosine deaminase provided herein. at least 60%, at least 10%, or at least 15% of any of the amino acid sequences described in any of the At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or or an amino acid sequence that is at least 99.5% identical to the amino acid sequence provided herein. Adenosine deaminase may contain one or more mutations (e.g., the mutations provided herein). It should be understood that the present disclosure may include specific percentages of the same or different. The deaminase domain may additionally have any of the mutations described herein. In some embodiments, the present invention provides a method for treating adenosine dehydrogenase (ADD)-induced leukemia, ... or a combination thereof. The enzyme is compared to either the reference sequence or the adenosine deaminase provided herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid sequences In some embodiments, the adenosine deaminase is a nucleotide sequence known in the art. at least as compared to any of the amino acid sequences known or described herein at least 5, at least 10, at least 15, at least 20, at least 25, at least At least 30, at least 35, at least 40, at least 45, at least 50, at least At least 60, at least 70, at least 80, at least 90, at least 100, At least 110, at least 120, at least 130, at least 140, at least 150, Amino acids having at least 160 or at least 170 identical consecutive amino acid residues Contains the acid sequence.

[0194] In some embodiments, the adenosine deaminase comprises a D108X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In this case, the adenosine deaminase is expressed by the following amino acids: D108G, D108N, D108V, D108A in the TadA reference sequence. , or the D108Y mutation, or a corresponding mutation in another adenosine deaminase However, additional deaminases may be similarly aligned and are included herein. It is understood that homologous amino acid residues can be identified that can be mutated to provide It should be.

[0195] In some embodiments, the adenosine deaminase comprises an A106X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In this study, adenosine deaminase was detected by the A106V mutation in the TadA reference sequence or by another Includes the corresponding mutation in adenosine deaminase.

[0196] In some embodiments, the adenosine deaminase comprises an E155X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and the presence of X indicates Any amino acid other than the corresponding amino acid in the native adenosine deaminase is indicated. In some embodiments, the adenosine deaminase is selected from the group consisting of E155D, E155G, or contains the E155V mutation or a corresponding mutation in another adenosine deaminase .

[0197] In some embodiments, the adenosine deaminase comprises a D147X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and the presence of X indicates Any amino acid other than the corresponding amino acid in the native adenosine deaminase is indicated. In some embodiments, the adenosine deaminase comprises a D147Y mutation in the TadA reference sequence, or or a corresponding mutation in another adenosine deaminase.

[0198] The mutations provided herein (e.g., based on the ecTadA amino acid sequence of the TadA reference sequence) Neither of these enzymes (i.e., the ATPases) are related to other adenosine deaminases, such as Staphylococcus aureus TadA (saTadA), or other adenosine deaminase (e.g., bacterial adenosine deaminase) It should be understood that the mutated residues in ecTadA may be homologous to those in ecTadA. It will be clear to one skilled in the art whether any of the mutations identified in ecTadA are related The present invention can be made in other adenosine deaminases having similar amino acid residues. Any of the mutations provided herein may be used in ecTadA or another adenosine deaminase. It should also be understood that the following may be produced individually or in any combination. For example, Adenosine deaminase is a variant of TadA that contains D108N, A106V, E155V, and / or It may contain a D147Y mutation, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase is derived from the following mutation in the TadA reference sequence: Different groups (groups of mutations separated by a ";"), or mutations in different adenosine deaminases, The corresponding mutations are: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V;A106V and D147Y;E155V and D147Y;D108N, A106V, and E55V; D108N, A106V, and D147Y;D108N, E55V, and D147Y;A106V, E55V, and D147Y; and D108N, A106V, E55V, and D147Y; provided, however, that the corresponding variants provided herein It is believed that any combination of variations can be made in adenosine deaminase (e.g., ecTadA). I want to be understood.

[0199] In some embodiments, the adenosine deaminase is H8X in the TadA reference sequence. , T17X, ​​L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X , A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I 156X, and / or K157X mutations, or one or more of other adenosine deaminases wherein the presence of X is indicative of a wild-type adenosine dehydrogenase (ADD) mutation in Any amino acid other than the corresponding amino acid in the enzyme is indicated. In TadA, the adenosine deaminase is expressed as H8Y, T17S, L18E, and W23L, which are related to the TadA reference sequence. , L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, 1951, V102A , F104L, A106V, R107C, or R107H, or R107P, D108G, or D108N, or D108A , D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or or one or more of the K157R mutations, or one in other adenosine deaminases Contains the corresponding mutations above.

[0200] In some embodiments, the adenosine deaminase is selected from the group consisting of H8X, D108X, and and / or one or more of the N127X mutations, or adenosine deaminase The amino acid sequence includes one or more corresponding mutations, where X indicates the presence of any amino acid. In this form, adenosine deaminase is expressed by the residues H8Y, D108N, and / in the TadA reference sequence. or one or more N127S mutations, or one or more counterparts in another adenosine deaminase This includes mutations that

[0201] In some embodiments, the adenosine deaminase is H8X in the TadA reference sequence. , R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X , one or more of the K161X, Q163X, and / or T166X mutations, or another adenosine deaminase and one or more corresponding mutations in adenosine deaminase, where X is a wild-type adenosine deaminase. It indicates the presence of any amino acid other than the corresponding amino acid in the In this state, adenosine deaminase is expressed by the following amino acids: H8Y, R26W, M61I, L68Q in the TadA reference sequence. , M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or one or more of the E155V or E155D, K161Q, Q163H, and / or T166P mutations, or or one or more corresponding mutations in other adenosine deaminases.

[0202] In some embodiments, the adenosine deaminase is H8X in the TadA reference sequence. , D108X, N127X, D147X, R152X, and Q154X; four, five, or six mutations, or the corresponding mutations in another adenosine deaminase mutation(s), where X is the corresponding mutation(s) in wild-type adenosine deaminase. In some embodiments, the presence of any amino acid other than the amino acid adeno The syndeaminase is a TadA reference sequence with H8X, M61X, M70X, D108X, N127X, Q154X, and E1 1, 2, 3, 4, 5, 6, 7, or more selected from the group consisting of 55X, and Q163X 8 mutations or corresponding mutations in other adenosine deaminases (multiple where X is an amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase is TadA. one selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the reference sequence , two, three, four, or five mutations, or a counterpart in another adenosine deaminase and X is a corresponding mutation(s) in wild-type adenosine deaminase. In some embodiments, adenosine is present. The deaminase is a mutant of H8X, A106X, D108X, another adenosine deaminase mutation ( and wherein the mutations are 1, 2, 3, 4, 5, or 6 selected from the group consisting of: X is the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase is present at the H8X, R in the TadA reference sequence. one, two, three or more selected from the group consisting of 126X, L68X, D108X, N127X, D147X, and E155X; One, four, five, six, seven, or eight mutations, or a different adenosine deaminase and X is the corresponding mutation(s) in wild-type adenosine deaminase. It indicates the presence of any amino acid other than the corresponding amino acid. The deaminase is derived from H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence. one, two, three, four, or five mutations selected from the group consisting of: and the corresponding mutation(s) in adenosine deaminase, where X is the wild-type adenosine deaminase. The presence of any amino acid other than the corresponding amino acid in the amino acid sequence of ...

[0203] In some embodiments, the adenosine deaminase is H8Y in the TadA reference sequence. , D108N, N127S, D147Y, R152C, and Q154H; four, five, or six mutations, or the corresponding mutations in another adenosine deaminase In some embodiments, the adenosine deaminase comprises a TadA reference sequence. Selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G and Q163H in the sequence One, two, three, four, five, six, seven or eight mutations selected, or other adenosines In some embodiments, the adenovirus comprises a corresponding mutation in adenovirus deaminase. The syndeaminase is derived from H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence. or another adenovirus. In some embodiments, the corresponding mutation(s) in the syndeaminase. The adenosine deaminase is expressed at the residues H8Y, A106T, D108N, N127S, and E1 in the TadA reference sequence. 1, 2, 3, 4, 5, or 6 mutations selected from the group consisting of K161Q, K161D, K161E, K161F, K161G, K161H ... mutations, or corresponding mutations in other adenosine deaminases. In embodiments, the adenosine deaminase is H8Y, R126W, L68Q in the TadA reference sequence. , D108N, N127S, D147Y, and E155V. One, six, seven, or eight mutations, or their counterparts in other adenosine deaminases In some embodiments, the adenosine deaminase comprises a mutation that inhibits TadA. one selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the reference sequence; Two, three, four, or five mutations, or their counterparts in another adenosine deaminase Contains a mutation(s) that

[0204] In some embodiments, the adenosine deaminase is In some embodiments, the adenosine deaminase comprises one or more corresponding mutations. D108N, D108G, or D108V mutations in the dA reference sequence, or another adenosine deoxynucleotide In one embodiment, the enzyme comprises a corresponding mutation in adenosine deaminase. The A106V and D108N mutations in the TadA reference sequence, or another adenosine deaminase In some embodiments, the adenosine deaminase comprises a corresponding mutation in Ta The R107C and D108N mutations in the dA reference sequence, or in another adenosine deaminase In some embodiments, the adenosine deaminase comprises a corresponding mutation in TadA. H8Y, D108N, N127S, D147Y, and Q154H mutations in the sequence, or another adenosine In some embodiments, the adenosine deaminase includes a corresponding mutation in the deaminase. The enzyme harbors the H8Y, R24W, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase. Therefore, adenosine deaminase is expressed at the D108N, D147Y, and E155V mutations in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase. In this study, adenosine deaminase was identified at the H8Y, D108N, and S127S residues in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase. In TadA, adenosine deaminase is expressed by the A106V, D108N, D147Y and and E155V mutations, or corresponding mutations in another adenosine deaminase .

[0205] In some embodiments, the adenosine deaminase is S2X in the tadA reference sequence. , one or more of the H8X, I49X, L84X, H123X, N127X, I156X and / or K160X mutations, and contains one or more corresponding mutations in another adenosine deaminase, and the presence of X Any amino acid other than the corresponding amino acid in wild-type adenosine deaminase In some embodiments, the adenosine deaminase is S2 in the TadA reference sequence. one or more of the following mutations: A, H8Y, I49F, L84F, H123Y, N127S, I156F and / or K160S , or one or more corresponding mutations in another adenosine deaminase.

[0206] In some embodiments, the adenosine deaminase is an L84X mutant adenosine deaminase. X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase In some embodiments, the adenosine deaminase is L8 in the TadA reference sequence. 4F mutation, or a corresponding mutation in another adenosine deaminase.

[0207] In some embodiments, the adenosine deaminase comprises a H123X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase Any amino acid other than the corresponding amino acid in adenosine deaminase is indicated. In TadA, adenosine deaminase is expressed by the H123Y mutation in the TadA reference sequence or by another containing the corresponding mutation in the adenosine deaminase of

[0208] In some embodiments, the adenosine deaminase comprises a I157X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase Any amino acid other than the corresponding amino acid in adenosine deaminase is indicated. In TadA, adenosine deaminase is expressed by the I157F mutation in the TadA reference sequence, or by another containing the corresponding mutation in the adenosine deaminase of

[0209] In some embodiments, the adenosine deaminase comprises the amino acid sequence L84X, A106X in the TadA reference sequence. , D108X, H123X, D147X, E155X, and I156X; Four, five, six, or seven mutations, or their counterparts in another adenosine deaminase wherein X is a mutation(s) corresponding to a corresponding amino acid sequence in wild-type adenosine deaminase. In some embodiments, adenosine is present. The deaminase is S2X, I49X, A106X, D108X, D147X, and E155X in the tadA reference sequence. or another mutation selected from the group consisting of: and corresponding mutation(s) in adenosine deaminase, where X is the wild type indicates the presence of any amino acid other than the corresponding amino acid in adenosine deaminase In some embodiments, the adenosine deaminase comprises H8X, A106X, and H8X residues in the TadA reference sequence. 1, 2, 3, 4, or 5 mutations selected from the group consisting of D108X, N127X, and K160X or a corresponding mutation(s) in another adenosine deaminase, X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase Indicates the presence of acid.

[0210] In some embodiments, the adenosine deaminase comprises the following amino acids in the TadA reference sequence: L84F, A106V , D108N, H123Y, D147Y, E155V, and I156F; Four, five, six, or seven mutations, or their counterparts in another adenosine deaminase In some embodiments, the adenosine deaminase comprises a mutation that and one or two selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V. Contains one, three, four, five, or six mutations.

[0211] In some embodiments, the adenosine deaminase is H8Y in the TadA reference sequence. , A106T, D108N, N127S, and K160S. or five mutations, or corresponding mutations in other adenosine deaminases. This includes mutations.

[0212] In some embodiments, the adenosine deaminase is E25 in the TadA reference sequence. one or more of the following mutations: X, R 26 X, R 107 X, A 142 X, and / or A 143 X; or and one or more corresponding mutations in another adenosine deaminase, and the presence of X indicates a wild-type Any amino acid other than the corresponding amino acid in type adenosine deaminase is indicated. In some embodiments, the adenosine deaminase is selected from the group consisting of E25M, E25D, and E25A in the TadA reference sequence. , E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A , R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A one or more of the following mutations: A143W, A143M, A143S, A143Q, and / or A143R, or other mutations In some embodiments, the enzyme comprises one or more corresponding mutations in adenosine deaminase. wherein the adenosine deaminase is a mutation described herein corresponding to the TadA reference sequence. or a corresponding mutation in another adenosine deaminase Contains one or more of the following:

[0213] In some embodiments, the adenosine deaminase comprises an E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In this case, adenosine deaminase is expressed as E25M, E25D, E25A, E25R ... E25V, E25S, or E25Y mutations, or the corresponding mutations in another adenosine deaminase. Includes natural mutations.

[0214] In some embodiments, the adenosine deaminase comprises a R26X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X is the wild-type adenosine deaminase It indicates any amino acid other than the corresponding amino acid in nosine deaminase. In embodiments, the adenosine deaminase is selected from the group consisting of R26G, R26N, R26Q in the TadA reference sequence. , R26C, R26L, or R26K mutations, or the corresponding mutations in another adenosine deaminase. This includes mutations that

[0215] In some embodiments, the adenosine deaminase comprises a R107X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase Any amino acid other than the corresponding amino acid in adenosine deaminase is indicated. In TadA, adenosine deaminase is expressed by R107P, R07K, R107A, and R107 in the TadA reference sequence. N, R107W, R107H, or R107S mutation, or a counterpart in another adenosine deaminase It contains the corresponding mutations.

[0216] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase Any amino acid other than the corresponding amino acid in adenosine deaminase is indicated. In TadA, adenosine deaminase is expressed at the A142N, A142D, and A142G aberrations in the TadA reference sequence. mutation, or a corresponding mutation in another adenosine deaminase.

[0217] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase Any amino acid other than the corresponding amino acid in adenosine deaminase is indicated. In TadA, adenosine deaminase is expressed at A143D, A143G, A143E, and A14 3L, A143W, A143M, A143S, A143Q and / or A143R mutations, or another adenosine with the corresponding mutation in the deaminase.

[0218] In some embodiments, the adenosine deaminase is H36X in the TadA reference sequence. , N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or or one or more K161X mutations, or one or more in another adenosine deaminase a corresponding mutation, where the presence of X is the corresponding mutation in wild-type adenosine deaminase. In some embodiments, the amino acid sequence of the adenovirus is The syndeaminase contains the following amino acids in the TadA reference sequence: H36L, N37T, N37S, P48T, P48L, I49V, and R51H , R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T one or more of the mutations, or one or more corresponding mutations in another adenosine deaminase Includes.

[0219] In some embodiments, the adenosine deaminase is H36X in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase is H36L mutation, or a corresponding mutation in another adenosine deaminase.

[0220] In some embodiments, the adenosine deaminase comprises a N37X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X is a wild-type adenosine deaminase. It refers to any amino acid other than the corresponding amino acid in syndeaminase. In this case, adenosine deaminase is expressed by the N37T or N37S mutation in the TadA reference sequence, or contains the corresponding mutation in another adenosine deaminase.

[0221] In some embodiments, the adenosine deaminase is a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X is a wild-type It indicates any amino acid other than the corresponding amino acid in the type adenosine deaminase. In embodiments, the adenosine deaminase is a P48T or P48L mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase.

[0222] In some embodiments, the adenosine deaminase comprises a R51X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X is a wild-type adenosine deaminase. It refers to any amino acid other than the corresponding amino acid in syndeaminase. In this case, adenosine deaminase is expressed by the R51H or R51L mutation in the TadA reference sequence, or contains the corresponding mutation in another adenosine deaminase.

[0223] In some embodiments, the adenosine deaminase comprises a S146X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In the present invention, the adenosine deaminase is a S146R or S146C mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.

[0224] In some embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In this study, adenosine deaminase was detected by the K157N mutation in the TadA reference sequence or by another Includes the corresponding mutation in adenosine deaminase.

[0225] In some embodiments, the adenosine deaminase is a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X is a wild-type It indicates any amino acid other than the corresponding amino acid in the type adenosine deaminase. In embodiments, the adenosine deaminase is P48S, P48T, or P48S in the TadA reference sequence. 8A mutation, or a corresponding mutation in another adenosine deaminase.

[0226] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, and X is the wild-type adenosine deaminase It represents any amino acid other than the corresponding amino acid in nosine deaminase. In this study, adenosine deaminase was detected by the A142N mutation in the TadA reference sequence or by another Includes the corresponding mutation in adenosine deaminase.

[0227] In some embodiments, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X is a wild-type It indicates any amino acid other than the corresponding amino acid in the type adenosine deaminase. In embodiments, the adenosine deaminase is a variant of the TadA gene, comprising a W23R or W23L mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase.

[0228] In some embodiments, the adenosine deaminase comprises a R152X mutation in the TadA reference sequence. or a corresponding mutation in another adenosine deaminase, where X is Any amino acid other than the corresponding amino acid in the native adenosine deaminase is indicated. In some embodiments, the adenosine deaminase is a mutant of the TadA reference sequence, e.g., R152P or R52H. The present invention includes a natural mutation in the adenosine deaminase gene, or a corresponding mutation in another adenosine deaminase.

[0229] In one embodiment, the adenosine deaminase has the mutations H36L, R51L, L84F, A106V, D10 8N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In this form, adenosine deaminase is expressed by the following combination of mutations relative to the tadA reference sequence: where each mutation in the combination is separated by a "_" and each combination of mutations is In brackets: (A106V_D108N), (R107C_D108N), (H8Y_D108N_S 127S_D 147Y_Q154H), (H8Y_R24W_D108N_N127S_D147Y_E155V), (D108N_D147 Y_E155V), (H8Y_D108N_S 127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V D108N D147Y E155V) (D108Q D147Y E155V) (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147 Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V), (E59A cat dead_A106V_D108N_D147Y_E155V), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D103A_D014N), (G22P_D 103 A_D 104N), (G22P_D 103 A_D 104N_S 138 A) , (D 103 A_D 104N_S 138A), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I15 6F), (R 26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (R26C_L 84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_A1 42N_A143L_D147Y_E155V_I156F), (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I15 6F), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (A106V_D108N_A142N_D147Y_E155V), (R26G_A106V_D108N_A142N_D147Y_E155V), (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V), (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V), (E25D_R26G_A106V_D108N_A142N_D147Y_E155V), (A106V_R107K_D108N_A142N_D147Y_E155V), (A106V_D108N_A142N_A143G_D147Y_E155V), (A106V_D108N_A142N_A143L_D147Y_E155V), (H36L_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F), (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T), (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F), (N72S_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F), (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F), 57N), (H36L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F_K161T), (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N), (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E), (H36L_G67V_L84F_A106V_D108N_H123Y_S 146T_D147Y_E155V_I156F), (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F), (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F), (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F), (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F), (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E), (R74Q L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F), (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (P48S_A142N), (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N), (P48T_I49V_A142N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F (H36L_P48T _I49V_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_ I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152H_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_E155 V_I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_R152P _E155V_I156 F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_R152P_E155 V_I156F _K1 57N).

[0230] [Cytidine deaminase] In one embodiment, a fusion protein of the invention comprises a cytidine deaminase. In some embodiments, the cytidine deaminase provided herein deactivates cytosine or can deaminate 5-methylcytosine to uracil or thymine. In some embodiments, the cytosine deaminase provided herein deactivates cytosine in DNA. Cytidine deaminase can be derived from any suitable organism. In some embodiments, the cytidine deaminase is a naturally occurring one or more cytidine deaminases corresponding to any of the mutations provided herein. Those skilled in the art will be able to identify the mutations by, for example, sequence alignment and comparison. Determination of homologous residues allows the identification of corresponding residues in any homologous protein. Thus, one skilled in the art can identify a mutation corresponding to any of the mutations described herein. can be generated in any naturally occurring cytidine deaminase. In embodiments, the cytidine deaminase is of prokaryotic origin. In some embodiments, the cytidine deaminase is derived from a mammalian source. It is of mammalian (e.g., human) origin.

[0231] In some embodiments, the cytidine deaminase is a cytidine deaminase described herein. At least 60%, at least 65%, or at least At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. The enzyme may contain one or more mutations (e.g., any of the mutations provided herein). It is understood that the present disclosure does not limit the scope of any deaminat- ing sequences having a particular percent identity. The enzyme domain is modified with any of the mutations described herein or combinations thereof. In some embodiments, the cytidine deaminase is a nucleotide sequence that is a reference sequence or a or any cytidine deaminase provided herein, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, It includes amino acid sequences with 47, 48, 49, 50 or more mutations. In some embodiments, the cytidine deaminase may be any of those known in the art or described herein. at least 5, at least 10, or fewer amino acids compared to any of the amino acid sequences described in the literature At least 15, at least 20, at least 25, at least 30, at least 35, at least 40, At least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues The amino acid sequence includes:

[0232] The fusion proteins of the present invention comprise a nucleic acid editing domain. The domain can catalyze a C to U base change. In some embodiments, the nucleic acid editing domain The main domain is a deaminase domain. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase is adenosine deaminase. , apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In embodiments, the deaminase is APOBECl deaminase. In some embodiments, the deaminase is an APOBEC 2 deaminase. In some embodiments, the deaminase is an APOBEC3A deaminase. In some embodiments, the deaminase is an APOBEC3B deaminase. In one embodiment, the deaminase is an APOBEC3C deaminase. In some embodiments, the deaminase is an APOBEC3E deaminase. In some embodiments, the deaminase is an APOBEC3F deaminase. In some embodiments, the deaminase is an APOBEC3G deaminase. In some embodiments, the deaminase is an APOBEC3H deaminase. In some embodiments, the deaminase is an activation-induced deaminase (AID). D). In some embodiments, the deaminase is a vertebrate deaminase. In some embodiments, the deaminase is an invertebrate deaminase. The minase was synthesized in humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, or mice. In some embodiments, the deaminase is a human deaminase. In embodiments, the deaminase is a rat deaminase, e.g., rAPOBEC1. In this case, the deaminase is Petromyzon marinus cytidine deaminase 1 (pmCDAl). In some embodiments, the deaminase is human APOBEC3G. The aminase is a fragment of human APOBEC3G. In some embodiments, the deaminase is D316R In one embodiment, the deaminase is a human APOBEC3G variant comprising a D317R mutation. is a fragment of human APOBEC3G and contains mutations corresponding to the D316R D317R mutations. In embodiments, the nucleic acid editing domain is a deaminase of any of the deaminases described herein. At least 80%, at least 85%, at least 90%, at least 92% for the enzyme domain , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%), or at least 99.5% identical.

[0233] In certain embodiments, the fusion proteins provided herein comprise a base of the fusion protein. The fusion proteins provided herein contain one or more features that improve editing activity. The material may comprise a Cas9 domain with reduced nuclease activity. Therefore, the fusion proteins provided herein contain a Cas9 domain that does not have nuclease activity. (dCas9), or Cas9 nickase (nCas9), which cuts one strand of a double-stranded DNA molecule. It may have a Cas9 domain.

[0234] [Other Nucleobase Editors] The present invention relates to a method for producing a cytidine deaminase or adenosine deaminase in a fusion protein according to the present invention. The deaminase domain can be modified with virtually any nucleobase editor known in the art. The present invention provides a nucleobase editor fusion protein in which the amino acid sequence of the present invention is replaced with a sequence of the present invention.

[0235] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) , a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The main ones are nuclease-active Cas9 domains, nuclease-inactive Cas9 domains, or Ca In some embodiments, the Cas9 domain can be a nuclease active domain. For example, the Cas9 domain binds both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain may be a Cas9 domain that cleaves one strand of the target gene (the other strand). The amino acid sequence comprises any one of the amino acid sequences described herein. Thus, the Cas9 domain may be at least one of the amino acid sequences described herein. At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 8 5%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, Some contain amino acid sequences that are at least 99%, or at least 99.5%, identical to one another. In embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 , 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. at least 10, at least 15, at least 20, or at least At least 30, at least 40, at least 50, at least 60, at least 70, at least 80 , at least 90, at least 100, at least 150, at least 200, at least 250, at least at least 300, at least 350, at least 400, at least 500, at least 600, at least at least 700, at least 800, at least 900, at least 1000, at least 1100, or less It comprises an amino acid sequence having an identical stretch of at least 1200 amino acid residues.

[0236] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can cleave a double-stranded nucleic acid molecule without cleaving either strand. In some embodiments, the nucleic acid molecule can be linked to the nucleic acid molecule (e.g., via a gRNA molecule). The enzyme-inactive dCas9 domain contains the D10X mutation of the amino acid sequence described herein. and H840X mutation, or in any of the amino acid sequences provided herein and the corresponding mutation, and X is any amino acid change. The nuclease-inactive dCas9 domain contains the D10A mutation of the amino acid sequence described herein. and H840A mutations, or the corresponding mutations in any of the amino acid sequences described herein. As an example, a nuclease-inactive Cas9 domain can be used in a cloning vector. pPlatTET-gRNA2 (accession number BAV54124) contains the following amino acid sequence: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific c control of gene expression.” Cell. 2013; 152(5):1173-83, the entire contents of which are incorporated herein by reference. , which is incorporated herein by reference.

[0237] Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. Such further examples will be apparent to those skilled in the art based on their knowledge and are within the scope of the present disclosure. Suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H84 0A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains ( For example, Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotech 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference. In some embodiments, the dCas9 domain is a dCas9 domain provided herein. At least 60%, at least 65%, at least 70%, or at least At least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 9 6%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity In some embodiments, the Cas9 domain comprises an amino acid sequence having the amino acid sequence described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 compared to any of the amino acid sequences listed , 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32 , 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or In some embodiments, the Cas9 domain comprises an amino acid sequence having one or more mutations. The amino acid sequence has at least 10,000 or more amino acids compared to any of the amino acid sequences described herein. At least 15, at least 20, at least 30, at least 40, at least 50, at least 60, At least 70, at least 80, at least 90, at least 100, at least 150, at least At least 200, at least 250, at least 300, at least 350, at least 400, at least 500 , at least 600, at least 700, at least 800, at least 900, at least 1000, Amino acid sequences having at least 1100, or at least 1200 identical consecutive amino acid residues Contains columns.

[0238] In some embodiments, the Cas9 domain is a Cas9 nickase. Cas9 protein, which can cut only one strand of a nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase can target the double-stranded nucleic acid molecule. This allows the Cas9 nickase to separate the gRNA (e.g., sgRNA) bound to the Cas9 from the nucleotide sequence. It means cleaving paired (complementary) strands. The Cas9 nickase contains a D10A mutation and has a histidine at position 840. In embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule; This is because the Cas9 nickase base pairs with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase cleaves the H840A end strand. containing a natural mutation with an aspartic acid residue at position 10, or the corresponding mutation. In some embodiments, the Cas9 nickase is any of the Cas9 nickases provided herein. or at least 60%, at least 65%, at least 70%, at least 75%, at least 80% , at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, containing an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to the Additional suitable Cas9 nickases may be identified based on this disclosure and knowledge in the art. It is obvious to one skilled in the art and is within the scope of this disclosure.

[0239] [Cas9 domain with reduced PAM exclusivity] In one particular embodiment, the invention comprises a Cas9 domain split into two fragments. Each of the two fragments has a terminal intein, i.e., N The C-terminal fragment is fused at its C-terminus to one member of the intein system, and the C-terminal fragment is It has a member of the intein system at its N-terminus.

[0240] Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), bind to specific nucleic acids. It requires the standard NGG PAM sequence to bind to the region, where the "N" in "NGG" is adenovirus. A is cytosine (A), thymidine (T) or cytosine (C), and G is guanosine. This may limit the ability to edit desired bases in the genome. The base editing fusion proteins provided herein can be used to target a target at a precise location, e.g., upstream of the PAM. It may be necessary to place the base in a region containing the base. ammable editing of a target base in genomic DNA without double-stranded DNA clear See, Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Thus, in some embodiments, the fusion proteins provided herein Either protein binds to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. The Cas9 domain that binds to the non-canonical PAM sequence may comprise a Cas9 domain that can bind to the non-canonical PAM sequence. These have been described in the art and will be apparent to those skilled in the art. For example, The Cas9 domain to be bound is described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 nuc leases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleins Tiber, BP, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2 015), the entire contents of each of which are incorporated herein by reference. Table 1 below lists the Several PAM variants are described.

[0241] Table 1. Cas9 proteins and corresponding PAM sequences [Table 1]

[0242] In some embodiments, the PAM is NGC. In some embodiments, the NGC PA M is recognized by the Cas9 variant. In some embodiments, the NGC PAM variant The models are D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E and T1337R (combined) and (hereafter referred to as "MQKFRAER").

[0243] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCa In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, nuclease The enzyme-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In embodiments, SaCas9 contains an N579A mutation, or an amino acid sequence provided herein. containing the corresponding mutation in either of the sequences.

[0244] In some embodiments, a SaCas9 domain, a SaCas9d domain, or a SaCas9n domain The nucleotides can bind to nucleic acid sequences with non-canonical PAMs, and in some embodiments Therefore, the SaCas9 domain, SaCas9d domain, or SaCas9n domain has an NNGRRT PAM sequence. In some embodiments, the SaCas9 domain can bind to a nucleic acid sequence that encodes the , E781X, N967X, and R1014X mutations, or one or more of the amino acid sequences provided herein and corresponding mutations in any of the sequences, where X is any amino acid. In some embodiments, the SaCas9 domain contains the E781K, N967K, and R1014H mutations. or one or more of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain contains the following mutations: E781K, N96 7K, or R1014H mutation, or any of the amino acid sequences provided herein The corresponding mutations are included.

[0245] Exemplary SaCas9 Sequences KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE N SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENNYEVNSKCYEEAKKLKKISNQAE FIASFYNDDLIKINGLYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG The underlined and bolded residue above, N579, is mutated (e.g., to A579) to form the SaCas9 nick. This can produce ze.

[0246] Exemplary SaCas9n Sequences KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG Residue A579 above can be mutated from N579 to generate SaCas9 nickase, underlined and bolded. is indicated by letters.

[0247] Exemplary SaKKH Cas9 KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNR K LINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFY K NDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPP H IIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. Residue A579 above can be mutated from N579 to generate SaCas9 nickase, underlined and bolded. The residues K781, K967, and H1014 are shown in bold. The residues K781, K967, and H1014 are shown in bold. The residues K781, K967, and H1014 are shown in bold. It can be mutated to give SaKKH Cas9 and is shown underlined and italicized.

[0248] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus pyogenes. In some embodiments, the SpCas9 domain comprises a nuclease-active SpCas9, a nuclease-binding domain (SpCas9). SpCas9 is a nucleotide-binding protein that binds to the nucleotides in the nucleotide sequence of the target molecule, either ase-inactive SpCas9 (SpCas9d) or SpCas9 nickase (SpCas9n). In this embodiment, SpCas9 has a D9X mutation, or an amino acid sequence provided herein. and a corresponding mutation in any of the following sequences, where X is any amino acid other than D. In some embodiments, SpCas9 is a D9A mutation or a variant of any of the variants provided herein. In some embodiments, the SpCas9 domain contains a corresponding mutation in either of the amino acid sequences. Main, SpCas9d domain or SpCas9n domain binds to nucleic acid sequences with non-canonical PAM In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain The domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain contains D1134X, R1334X, and T1336X mutations. or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, S The pCas9 domain may contain one or more of the D1134E, R1334Q, and T1336R mutations, or any of the mutations described herein. It includes corresponding mutations in any of the amino acid sequences provided. In this embodiment, the SpCas9 domain contains the D1134E, R1334Q, and T1336R mutations, or the mutations described herein. The amino acid sequences of the present invention may be modified in various ways, including corresponding mutations in any of the amino acid sequences provided in the publication. In embodiments, the SpCas9 domain contains one of a D1134X, R1334X, and T1336X mutation. or one or more corresponding hits in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 The domain may contain one or more of the D1134V, R1334Q, and T1336R mutations, or any of the mutations provided herein. In some embodiments, the amino acid sequence of the nucleotide ... In this case, the SpCas9 domain contains the D1134V, R1334Q, and T1336R mutations, or the mutations described herein. and corresponding mutations in any of the amino acid sequences provided. wherein the SpCas9 domain contains one or more of D1134X, G1217X, R1334X, and T1336X mutations. or a corresponding mutation in any of the amino acid sequences provided herein. wherein X is any amino acid. In some embodiments, the SpCas9 domain is D one or more of the 1134V, G1217R, R1334Q, and T1336R mutations, or any of the mutations provided herein In some embodiments, the amino acid sequence of the nucleotide ... , the SpCas9 domain contains the D1134V, G1217R, R1334Q, and T1336R mutations, or the mutations described herein. The amino acid sequences provided herein include corresponding mutations in any of the amino acid sequences provided herein.

[0249] In some embodiments, the Cas9 of any of the fusion proteins provided herein The domain is at least 60%, at least At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or or an amino acid sequence that is at least 99.5% identical to the amino acid sequence of the target gene. The Cas9 domain of any of the fusion proteins provided herein may be any of the fusion proteins described herein. In some embodiments, the amino acid sequence of any Cas9 polypeptide of the present invention is The Cas9 domain of any of the fusion proteins provided herein may be any of the fusion proteins described herein. The Cas9 polypeptide of any one of the invention may be a Cas9 polypeptide.

[0250] Exemplary SpCas9 DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0251] Exemplary SpCas9n DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0252] Exemplary SpEQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF E SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues E1134, Q1334, and R1336 were mutated from D1134, R1334, and T1336. SpEQR-Cas9 can be generated and is shown in bold and underlined.

[0253] Exemplary SpVQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, Q1334, and R1336 were mutated from D1134, R1334, and T1336. SpVQR Cas9 can be generated and is shown in bold and underlined.

[0254] Exemplary SpVRER Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF VSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, R1217, Q1334, and R1336 are D1134, G1217, R1334, and T1336 can be mutated to yield SpVRER Cas9, and is shown underlined and in bold.

[0255] In certain embodiments, the fusion proteins of the invention comprise a dCas9 domain that binds to a canonical PAM sequence. nCas9 domains that bind to non-canonical and non-canonical PAM sequences (e.g., non-canonical PAMs identified in Table 1) In another embodiment, the fusion protein of the invention comprises an nCas9 domain that binds to a canonical PAM sequence. dCas9 domains that bind to main and non-canonical PAM sequences (e.g., non-canonical PAMs identified in Table 1) Including Inn.

[0256] [High-fidelity Cas9 domain] Some aspects of the present disclosure provide high-fidelity Cas9 domains. In this study, high-fidelity Cas9 domains were characterized by a Cas9 domain with a higher fidelity than the corresponding wild-type Cas9 domain. those containing one or more mutations that reduce the electrostatic interactions between the phosphate and the sugar-phosphate backbone of DNA Without wishing to be bound by any particular theory, it is believed that the sugar-phosphate backbone of DNA High-fidelity Cas9 domains with reduced electrostatic interactions have fewer off-target effects In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) may have a It contains one or more mutations that reduce the bond between the phosphate and the sugar-phosphate backbone of DNA. In this manner, the Cas9 domain reduces the bond between the Cas9 domain and the sugar-phosphate backbone of DNA. At least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, At least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least The gene may also contain one or more mutations that reduce the gene expression level by 65%, or at least 70%.

[0257] In some embodiments, any of the Cas9 fusion proteins provided herein , N497X, R661X, Q695X, and / or Q926X mutations, or any of the mutations provided herein. and corresponding mutations in any of the amino acid sequences described above, where X is any amino acid. In some embodiments, any of the Cas9 fusion proteins provided herein is an acid. Any of the mutations may contain one or more of the N497A, R661A, Q695A, and / or Q926A mutations, or any of the mutations described herein. In some embodiments, the amino acid sequence of the present invention may be a mutation in the corresponding amino acid sequence of the present invention. In some embodiments, the Cas9 domain may comprise a D10A mutation, or the amino acid sequence provided herein. The Cas9 domains with high fidelity are well known in the art and contain corresponding mutations in either of the sequences. These techniques are well known in the art and will be apparent to those skilled in the art. For example, Cas9 domains with high fidelity can be used. The article is based on Kleinstiver, B.P., et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016); and S laymaker, IM, et al. “Rationally engineered Cas9 nucleases with improved spec ificity.” Science 351, 84-88 (2015), the full contents of which are hereby incorporated by reference. and is hereby incorporated by reference.

[0258] High-fidelity Cas9 domain mutations to Cas9 are shown in bold and underlined. DKKYSIGL A IGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMT A FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFM A LIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR A ITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0259] Cas9 nuclease has two functional endonuclease domains, RuvC and HNH. Upon binding to the target DNA, Cas9 undergoes a conformational change that induces nuclease degradation. The final result of Cas9-mediated DNA cleavage is that the target DNA strand is cut by a domain. is a double-strand break (DSB) in the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). SBs are repaired by one of two general repair paths: (1) an efficient but error-prone (2) the less efficient but more fidelity homology-guided repair pathway; Return (HDR) path.

[0260] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) is not known by any simple method. It can be calculated in a convenient way. For example, in some cases, the efficiency is For example, a test nuclease assay can be used to determine the percentage of cleavage product. The ratio of product to substrate can be used to calculate the percentage. For example, direct cleavage of DNA containing the newly integrated restriction sequence as a result of successful HDR. A nuclease enzyme can be used to measure the amount of substrate cleaved, which increases the rate of HDR. As an illustrative example, the HDR percentage (percentage) The percentage of cleavage can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (e.g., (b+c) / (a+b+c) where "a" is the band intensity of the DNA substrate, and "b" and and "c" are cleavage products.

[0261] In some cases, efficiency can be expressed as the success rate of NHEJ. For example, T7 endonuclease Cleavage products were generated using a cleavage enzyme I assay, and the ratio of product to substrate was used to determine the percentage of NHEJ. The conversion can be calculated using wild-type and mutant T7 endonuclease I. High levels of NHEJ (NHEJ generates small random insertions or deletions (indels) at the initial cleavage site) It cleaves the mismatched heteroduplex DNA resulting from hybridization. This indicates a high rate of NHEJ (high efficiency of NHEJ). (Percentage) is calculated using the formula (1-(1-(b+c) / (a+b+c)) 1 / 2 ) × 100 , where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products ( Ran et. al., 2013 Sep. 12; 154(6):1380-9; and Ran et. al., Nat Protoc. 2013 No v.; 8(11): 2281-2308).

[0262] The NHEJ repair pathway is the most active repair mechanism, resulting in small nucleotide insertions or deletions at DSB sites. The random nature of NHEJ-mediated DSB repair is due to the lack of Cas9 and gRNA. Alternatively, a population of cells expressing the guide polynucleotide may result in a diverse array of mutations. In most cases, NHEJ results in small indentations in the target DNA, which has important practical implications. This results in premature deletion of the target gene's open reading frame (ORF). Amino acid deletions, insertions, or frameshift mutations resulting in undesired stop codons occur. The ideal end result is a loss-of-function mutation in the target gene.

[0263] NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, but Homologous-directed repair (HDR) is a method for repairing single nucleotide changes, such as the addition of a fluorophore or tag. These can be used to generate specific nucleotide changes ranging from large insertions such as

[0264] To utilize HDR for gene editing, a DNA repair template containing the desired sequence is prepared using gRNA and It can be delivered to the cell type of interest along with Cas9 or Cas9 nickase. The rate is determined by the desired edit and the regions immediately upstream and downstream of the target (called left and right homology arms). The length of each homologous arm determines the magnitude of the change introduced. The repair template may depend on the size of the insert, with larger inserts requiring longer homology arms. a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid The efficiency of HDR was measured in cells expressing Cas9, gRNA, and exogenous repair template. HDR occurs between the S and G2 phases of the cell cycle, so the risk of developing a mutation is generally low (less than 10% correcting alleles). The efficiency of HDR can be increased by synchronizing cells. Chemically or genetically inhibiting offspring can also increase HDR frequency.

[0265] In some embodiments, the Cas9 is a modified Cas9. There may be additional sites of partial homology throughout the genome. These targets must be considered when designing gRNAs, but they are often used to optimize gRNA design. In addition to this, the specificity of CRISPR can also be increased by modifying Cas9. Cas9 induces double-strand breaks (DS) through the combined activity of two nuclease domains, RuvC and HNH. B) Cas9 nickase, the D10A mutant of SpCas9, contains one nuclease domain This allows for the generation of DNA nicks rather than DSBs. Editing can also be combined with nickase.

[0266] In some cases, the Cas9 is a variant Cas9 protein. Variant Cas9 polypeptide differs by a single amino acid compared to the amino acid sequence of the wild-type Cas9 protein (e.g., In some instances, the amino acid sequence may be a sequence of a variant of the amino acid sequence, such as a nucleotide sequence of a nucleotide, a nucleotide sequence of ... The modified Cas9 polypeptide contains an amino acid sequence that reduces the nuclease activity of the Cas9 polypeptide. have alterations (e.g., deletions, insertions, or substitutions). For example, in some instances, The ant-Cas9 polypeptides exhibited less than 50% of the nuclease activity of the corresponding wild-type Cas9 protein. less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. In this case, the variant Cas9 protein does not have substantial nuclease activity. If the protein is a variant Cas9 protein that does not have substantial nuclease activity, This may be referred to as "dCas9."

[0267] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, A variant Cas9 protein may be a variant of a wild-type Cas9 protein (e.g., a wild-type Cas9 protein). Less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or or less than about 0.1%.

[0268] In some cases, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. but have a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. Variant Cas9 proteins contain mutations (amino acid substitutions) that reduce the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant The Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10). , thus, can cleave the complementary strand of the double-stranded guide target sequence, but This variant of Cas9 has a reduced ability to cleave the non-complementary strand of the target sequence (hence, When a protein cleaves a double-stranded target nucleic acid, it creates a single-strand break (S) instead of a double-strand break (DSB). (See, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) .

[0269] In some cases, the variant Cas9 protein cleaves the non-complementary strand of the double-stranded guide target sequence. can cleave the complementary strand of the guide target sequence, but with a reduced ability to cleave the complementary strand of the guide target sequence. Variant Cas9 proteins have mutations (amino acid substitutions) that reduce the function of the HNH domain. (RuvC / HNH / RuvC domain motif). In embodiments, the variant Cas9 protein is H840A (His at amino acid position 840). alanine to cysteine) mutation, thus inhibiting the non-complementary strand of the guide target sequence. can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (Thus, this variant Cas9 protein cleaves the double-stranded guide target sequence.) (When the Cas9 protein is inserted into the guide target sequence, it generates an SSB instead of a DSB.) The ability to cleave the guide target sequence (e.g., single-stranded guide target sequence) is reduced, but the ability to cleave the guide target sequence (e.g., single-stranded guide target sequence) is reduced. The strand retains the ability to bind to the guide target sequence.

[0270] In some cases, the variant Cas9 protein is capable of cleaving complementary and non-complementary strands of a double-stranded target DNA. As a non-limiting example, in some cases, the variant The Cas9 protein has both the D10A and H840A mutations, resulting in a polypeptide These have a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). However, they retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0271] As another non-limiting example, in some cases, the variant Cas9 protein is 76A and W1126A mutations, such that the polypeptide binds to the target DNA (e.g., single-stranded target have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but have reduced ability to bind to target DNA (e.g., single-stranded target DNA). The power is retained.

[0272] As another non-limiting example, in some cases, the variant Cas9 protein may be P4 75A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, resulting in Such Cas9 proteins have a reduced ability to cleave target DNA. Although the target DNA (e.g., single-stranded target DNA) has a reduced ability to cleave the target DNA (e.g., single-stranded target DNA), It retains the ability to bind to DNA.

[0273] As another non-limiting example, in some cases, the variant Cas9 protein is H8 40A, W476A, and W1126A mutations, so that the polypeptide binds to target DNA (e.g., The target DNA (e.g., single-stranded target DNA) has a reduced ability to cleave, but As another non-limiting example, in some cases, The antCas9 protein has H840A, D10A, W476A, and W1126A mutations, resulting in Such Cas9 proteins have a reduced ability to cleave target DNA. has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but In some embodiments, the variant retains the ability to bind to single-stranded target DNA. This Cas9 has a restored catalytic His residue at position 840 of the Cas9 HNH domain (A840H). .

[0274] As another non-limiting example, in some cases, the variant Cas9 protein is H8 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, resulting in The polypeptide has reduced ability to cleave target DNA (e.g., single-stranded target DNA), but not target DNA. As another non-limiting example, some In some cases, the variant Cas9 protein is D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide binds to the target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded targets). have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but have reduced ability to bind to target DNA (e.g., single-stranded target DNA). If the variant Cas9 protein has the W476A and W1126A mutations, or variant Cas9 proteins P475A, W476A, N477A, ​​D1125A, W1126A, and D1127 A mutation, the variant Cas9 protein does not bind efficiently to the PAM sequence. Therefore, in such cases, such variant Cas9 proteins are used in methods of conjugation. In other words, in some cases, such a barrier is not required. When a recombinant Cas9 protein is used in the method of attachment, the method may include a guide RNA. The method can be performed in the absence of a PAM sequence (thus, the specificity of binding depends on the guide RNA). (This is brought about by the target segment of A). To achieve the above effect, other residues were varied. Non-limiting examples include: Residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and and / or A987 can be altered (i.e., substituted). is also suitable.

[0275] In some embodiments, a variant Cas9 protein (e.g., Ca) with reduced catalytic activity is provided. The s9 protein was found to be D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H98 2A, H983A, A984A, and / or D986A) that interacts with the guide RNA can bind to target DNA in a site-specific manner as long as it retains the ability to (This is because the target DNA sequence is guided by the

[0276] In some embodiments, the variant Cas protein is spCas9, spCas9-VRQR, spCas9 -VRER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas It can be 9-LRVSQL.

[0277] As an alternative to S. pyogenes Cas9, the Cpf1 family exhibits cleavage activity in mammalian cells. RNA-guided endonucleases derived from Prevotella and Francisella 1 may be mentioned. The current CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. 1 is a class II CRISPR / Cas RNA-guided endonuclease. This adaptive immune mechanism is involved in the Pre The Cpf1 gene is associated with the CRISPR locus and is found in the bacteria Vogelia and Francisella. It encodes an endonuclease that uses a guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and plays a role in the restriction of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in short The staggered cleavage pattern of Cpf1 is similar to that of traditional restriction enzymes. This opens up the possibility of directional gene transfer, similar to cloning, which is Similar to the Cas9 variants and orthologues described above, Cpf1 The number of sites that CRISPR can target is limited to ATs lacking the NGG PAM site preferred by SpCas9. It can also extend into AT-rich regions or into the AT-rich genome. The Cpf1 locus is an α / β mixed domain. In this figure, RuvC-I and the subsequent helical region, RuvC-II, and zinc finger-like domain are The Cpf1 protein contains a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 is Cas9. The Cpf1 CRISPR-Cas domain organization is important for Cpf1 to function. The results show that the CRISPR system is unique to CRISPR and is classified as a Class 2, Type V CRISPR system. The pf1 locus encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and III than to type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA). therefore, only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9, but also Because it has a small sgRNA molecule (about half the number of nucleotides of Cas9), it is suitable for genome editing. In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex targets the G-rich PAM motif Cleavage of target DNA or RNA occurs through identification of the protospacer adjacent to 5'-YTN-3'. After the identification of Cpf1, sticky-end-like DNA doublets with overhangs of 4 or 5 nucleotides were identified. Introduces a double-strand break.

[0278] [protospacer adjacent motif] The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a CRISPR-bacterial 2-6 base pairs of DNA immediately following the DNA sequence targeted by Cas9 nuclease in the immune system In some embodiments, the PAM refers to a 5' PAM (i.e., the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located upstream of the protostaglandin A). The nucleotide sequence may be located downstream of the 5' end of the sequence.

[0279] The PAM sequence is essential for target binding, but its exact sequence varies depending on the type of Cas protein. It exists.

[0280] The base editors provided herein can be used to generate bases adjacent to canonical or non-canonical protospacers. CRISPR proteins that can bind to nucleotide sequences containing the PAM motif (PAM) sequence The PAM site may comprise a nucleotide sequence adjacent to the target polynucleotide sequence. Some aspects of the present disclosure provide CRISPR targets with different PAM specificities. For example, a base editor comprising all or part of a protein derived from S. pyogenes is provided. Cas9 proteins, such as Cas9 (spCas9), are typically engineered to bind to specific nucleic acid regions. Requires the standard NGG PAM sequence, where the "N" in "NGG" stands for adenine (A) or thymine (T). , guanine (G), or cytosine (C), where G is guanine. PAM is a CRISPR protein. Different base editors can be protein-specific and contain domains from different CRISPR proteins. The PAM can be located 5' or 3' of the target sequence. The PAM can be located upstream or The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. PAMs can be any length. Often, PAMs are 2-6 nucleotides long. Some PAMs The variants are listed in Table 1.

[0281] In some embodiments, SpCas9 targets the PAM nucleic acid sequence 5'-NGC-3' or In various embodiments of the above aspects, the SpCas9 has specificity for Cas9 or any of the proteins listed in Table 1. In various embodiments of the above aspects, the modified SpCas9 is In some embodiments, the variant Cas protein is spCas9-MQKFRAER. Cas9, spCas9-VRQR, spCas9-VRER, xCas9 (sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER , spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL. In this state, amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R (SpCas9-MQKFRAER) with specificity for the engineered PAM 5'-NGC-3' A mutant SpCas9 is used.

[0282] In some embodiments, the PAM is an NGT. In some embodiments, the NGT PAM is a variant. In some embodiments, the NGT PAM variant has one or more residues 1335-1339. , 1337, 1135, 1136, 1218, and / or 1219 through targeted mutations In some embodiments, the NGT PAM variant comprises one or more of residues 1219, 1335, 133 7, 1218. In some embodiments, the NGT PAM The variant contains targeted mutations at one or more of residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variants are those listed in Table 2 and Table 3 below. and 3. The targeted mutations are selected from the set of targeted mutations provided in

[0283] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218 [Table 2]

[0284] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219, and 1335 [Table 3]

[0285] In some embodiments, the NGT PAM variant is variants 5, 7, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 8, 31, or 36. In some embodiments, the variant has improved NGT Has PAM awareness.

[0286] In some embodiments, the NGT PAM variant is at residues 1219, 1335, 1337, and / or or 1218. In some embodiments, the NGT PAM variant has a mutation at Selected with mutations to improve recognition from the variants provided in Table 4 .

[0287] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218 [Table 4]

[0288] In some embodiments, the NGT PAM is selected from the variants provided in Table 5 below. .

[0289] Table 5: NGT PAM variants [Table 5]

[0290] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus pyogenes. In some embodiments, the SpCas9 domain comprises a nuclease-active SpCas9, a nuclease-binding domain (SpCas9). SpCas9 is a nucleotide-binding protein that binds to the nucleotides in the nucleotide sequence of the target molecule, either ase-inactive SpCas9 (SpCas9d) or SpCas9 nickase (SpCas9n). In this embodiment, SpCas9 has a D9X mutation, or an amino acid sequence provided herein. and a corresponding mutation in any of the following sequences, where X is any amino acid other than D. In some embodiments, SpCas9 is a D9A mutation or a variant of any of the variants provided herein. In some embodiments, the SpCas9 domain contains a corresponding mutation in either of the amino acid sequences. Main, SpCas9d domain or SpCas9n domain binds to nucleic acid sequences with non-canonical PAM In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain The domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence.

[0291] In some embodiments, the SpCas9 domain contains the D1135X, R1335X, and T1336X aberrant sequences. One or more of the mutations or corresponding amino acid sequences in any of the amino acid sequences provided herein In some embodiments, the SpCas comprises a mutation, where X is any amino acid. The 9 domain may contain one or more of the D1135E, R1335Q, and T1336R mutations, or any of the mutations provided herein. In some embodiments, the amino acid sequence of the target gene is a nucleotide sequence. In the present invention, the SpCas9 domain contains the D1135E, R1335Q, and T1336R mutations, or the mutations described herein. It includes corresponding mutations in any of the amino acid sequences provided. In this embodiment, the SpCas9 domain contains one or more of the following mutations: D1135X, R1335X, and T1336X. or corresponding mutations in any of the amino acid sequences provided herein. wherein X is any amino acid. The inclusion may comprise one or more of the D1135V, R1335Q, and T1336R mutations, or any of the mutations provided herein. In some embodiments, the amino acid sequence of the nucleotide ... In one embodiment, the SpCas9 domain contains the D1135V, R1335Q, and T1336R mutations, or the mutations described herein. In some embodiments, the amino acid sequence of the nucleotide ... In this case, the SpCas9 domain contains one or more of the following mutations: D1135X, G1217X, R1335X, and T1336X. or corresponding mutations in any of the amino acid sequences provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain comprises D1135 V, G1217R, R1335Q, and T1336R mutations, or one or more of the amino acids provided herein. In some embodiments, the Sp The Cas9 domain may contain D1135V, G1217R, R1335Q, and T1336R mutations, or the mutations described herein. The amino acid sequences provided herein include corresponding mutations in any of the amino acid sequences provided herein.

[0292] In some embodiments, the Cas9 of any of the fusion proteins provided herein The domain is at least 60%, at least At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or or an amino acid sequence that is at least 99.5% identical to the amino acid sequence of the target gene. The Cas9 domain of any of the fusion proteins provided herein may be any of the fusion proteins described herein. In some embodiments, the amino acid sequence of any Cas9 polypeptide of the present invention is The Cas9 domain of any of the fusion proteins provided herein may be any of the fusion proteins described herein. The Cas9 polypeptide of any one of the invention may be a Cas9 polypeptide.

[0293] In some examples, the base editors disclosed herein are derived from CRISPR proteins. The PAM recognized by the domain is inserted into an insert (e.g., AAV The nucleic acid sequence may be provided to the cell on a separate oligonucleotide from the insert. In this embodiment, providing the PAM on a separate oligonucleotide allows for a sequence that is otherwise not aligned with the target sequence. A target that cannot be cleaved due to the absence of an adjacent PAM on the same polynucleotide Allows for cleavage of the sequence.

[0294] In one embodiment, S. pyogenes Cas9 (SpCas9) is used as a CRISPR enzyme for genome engineering. However, other enzymes may also be used. In some embodiments, different endonucleases are used to target specific genomic targets. In some embodiments, synthetic SpCas9-derived variants with non-NGG PAM sequences can be used. Additionally, other Cas9 orthologs from various species have been identified. These "non-SpCas9" molecules can bind to a variety of PAM sequences that may also be useful in this disclosure. For example, the relatively large size of SpCas9 (approximately 4 kilobases (kb) of coding sequence) This may result in an SpCas9 cDNA plasmid that cannot be efficiently expressed in vitro. Conversely, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kilobase weaker than SpCas9. Because it is short, it can be efficiently expressed in cells. The enzyme demonstrated the ability to modify target genes in mammalian cells in vitro and in mice in vivo. In some embodiments, the Cas proteins can target different PAM sequences. In some embodiments, the target gene can be flanked by a Cas9 PAM, e.g., 5'-NGG. In other embodiments, other Cas9 orthologs may have different PAM requirements. For example, S. ther PAMs like those in M. mophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) ) and other PAMs such as that of Neisseria meningiditis (5'-NNNNGATT) also target genes can be found adjacent to

[0295] In some embodiments, for S. pyogenes systems, the target gene sequence is 5'-NGG P The AM may be preceded by (i.e., 5' to) a 20 nt guide RNA sequence that base pairs with the opposite strand. In one embodiment, the adjacent cleavage site is formed to mediate Cas9 cleavage adjacent to the PAM. The cut can be (approximately) 3 base pairs upstream of the PAM. In one embodiment, the adjacent cut is (approximately) 3 base pairs upstream of the PAM. In one embodiment, the adjacent cleavage may be approximately 0-20 base pairs upstream of the PAM. For example, adjacent cleavages can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 upstream of the PAM. , 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 salt The adjacent cleavage can be 1 to 30 base pairs downstream of the PAM. An exemplary SpCas9 protein sequence that can bind:

[0296] The amino acid sequence of an exemplary PAM-bound SpCas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0297] The amino acid sequence of an exemplary PAM-bound SpCas9n is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0298] The amino acid sequence of an exemplary PAM-bound SpEQR Cas9 is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF ESPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD In this sequence, D1135, R1335, and T1337 were mutated to generate SpEQR Cas9. The possible residues E1135, Q1335, and R1337 are underlined and in bold.

[0299] The amino acid sequence of an exemplary PAM-linked SpVQR Cas9 is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRI DLSQLGGD In this sequence, D1135, R1335, and T1336 were mutated to generate SpVQR Cas9. The possible residues V1135, Q1335, and R1336 are underlined and in bold.

[0300] The amino acid sequence of an exemplary PAM-linked SpVRER Cas9 is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELINGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRI DLSQLGGD

[0301] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. The recombinant Cas9 domain is a SpyMacCas9 domain. The acCas9 domain is used to generate nuclease-active SpyMacCas9 and nuclease-inactive SpyMacCas9 (SpyM acCas9d), or SpyMacCas9 nickase (SpyMacCas9n). In this case, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain has a non-standard PAM. In some embodiments, the SpyMacCas9 domain can bind to a nucleic acid sequence that encodes the SpyMacCas9 domain. The SpCas9d domain or the SpCas9n domain binds to a nucleic acid sequence having a NAA PAM sequence. It is possible.

[0302] Exemplary SpyMacCas9 MDKKYSIGLDIGTNSVGWAVITDDYKVPSKKFKVLGNTDRHSIKKNLIGALLFGSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLADSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQIYNQLFEENPINASRVDAKAILSARLSKSRRLENLIAQLPGEKRNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNSEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGAYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRGMIERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGHSL HEQIANLAGSPAIKKGILQTVKIVDELVKVMGHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFIKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEIQTVGQNGGLFDDNPKSPLEVT PSKLVPLKKELNPKKYGGYQKPTTAYPVLLITDTKQLIPISVMNKKQFEQNPVKFLRDRGYQQVGKNDFIKLPKYTLVDI GDGIKRLWASSKEIHKGNQLVVSKKSQILLYHAHHLDSDLSNDYLQNHNQQFDVLFNEIISFSKKCKLGKEHIQKIENVY SNKKNSASIEELAESFIKLLGFTQLGATSPFNFLGVKLNQKQYKGKKDYILPCTEGTLIRQSITGLYETRVDLSKIGED.

[0303] In some cases, the variant Cas9 protein is H840A, P475A, W476A, N477A, ​​D1125A, W1 126A, D1218A mutations, resulting in a reduced ability to cleave target DNA or RNA. Such Cas9 proteins have the ability to cleave target DNA (e.g., single-stranded target DNA). Although the binding capacity is reduced, it still retains the ability to bind to target DNA (e.g., single-stranded target DNA). By way of non-limiting example, in some cases, the variant Cas9 protein may be D10A, H8 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1218A mutations, resulting in The polypeptide has a reduced ability to cleave target DNA (e.g., single-stranded target DNA). Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). The variant Cas9 protein retains the ability to bind to target DNA (e.g., single-stranded target DNA). If the protein has the W476A and W1126A mutations, or if the variant Cas9 protein has P475A, Variant Cas9 proteins with W476A, N477A, ​​D1125A, W1126A, and D1218A mutations Proteins do not bind efficiently to PAM sequences. When a variant Cas9 protein is used in the conjugation method, the method does not require a PAM sequence. In other words, in some cases, such variant Cas9 proteins are used in methods of conjugation. In some cases, the method may include a guide RNA, but the method may be performed in the absence of a PAM sequence. (Thus, binding specificity is provided by the target segment of the guide RNA) Other residues may be mutated (i.e., one or the other nucleoside) to achieve the above effect. (This inactivates the cleavage portion.) Non-limiting examples include residues D10, G12, G17, E762, H840 , modify (i.e., replace) N854, N863, H982, H983, A984, D986, and / or A987 Mutations other than alanine substitution are also suitable.

[0304] In some embodiments, the CRISPR protein-derived domain of the base editor is a canonical PAM sequence. In other embodiments, the Cas9 protein may comprise all or a portion of the Cas9 protein having the base sequence (NGG). The Cas9-derived domain of the editor can use non-canonical PAM sequences. Sequences are described in the art and will be apparent to those skilled in the art. For example, non-standard PAM sequences may be used. The Cas9 domains that bind to the array are described in Kleinstiver, B.P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature, 523, 481-485 (2015); and Kleinstiver, BP, et al., “Broadening the targeting range of Staphylococcus a ureus CRISPR-Cas9 by modifying PAM recognition”Nature Biotechnology, 33, 1293-1 298 (2015), the entire contents of each of which are incorporated herein by reference.

[0305] [Fusion protein containing a nuclear localization sequence (NLS)] In some embodiments, the fusion proteins provided herein comprise one or more (e.g., two, 3, 4, 5) nuclear targeting sequence, e.g., nuclear localization sequence (NLS). In some embodiments, a bipartite NLS is used. an amino acid sequence that promotes the import of a protein into the cell nucleus (e.g., by nuclear transport), including In some embodiments, any of the fusion proteins provided herein In some embodiments, the NLS is a fusion In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 domain. In some embodiments, the NLS is located at the C-terminus of the nCas9 domain or dCas9 domain. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. wherein the NLS is any one of the amino acids in the NLS sequences provided or referenced herein. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, in which: The disclosure of which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In the same manner, the NLS has the amino acid sequence PKKKRKVEGADKRTADGSEFES PKKKRKV, KRTADGSEFESPKKKRK V, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKK KRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, NL S is present in a linker, or the NLS is present in a linker, e.g., a linker described herein. In some embodiments, the N- or C-terminal NLS is flanked by a bipartite NLS. A bipartite NLS consists of two basic amino acids separated by a relatively short spacer sequence. It contains an amino acid cluster (hence the name bipartite), and one part (monopartite) The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is a ubiquitous bipartite The prototype of the signal is a sequence of approximately 10 amino acids consisting of two clusters of basic amino acids. An exemplary bipartite NLS sequence is PKKKRKVEGADKRTA DGSEFES PKKKRKV.

[0306] In some embodiments, the fusion proteins of the present invention do not include a linker sequence. In embodiments, there is a linker sequence between one or more domains or proteins.

[0307] It should be understood that the fusion proteins of the present disclosure may include one or more additional features. For example, in some embodiments, the fusion protein is an inhibitor, cytoplasmic localization sequences, export sequences such as nuclear export sequences, or other localization sequences, as well as the possibility of fusion proteins. The nucleic acid sequences provided herein may include sequence tags useful for solubilization, purification, or detection. Suitable protein tags include, but are not limited to, biotin carboxylase. Carrier tag (BCCP) tag, myc-tag, calmodulin tag, FLAG-tag, hemagglutinin Nin (HA)-tag, polyhistidine tag (also called histidine tag or His-tag) , maltose-binding protein (MBP)-tag, nus-tag, glutathione-S-transferase GST-tag, green fluorescent protein (GFP)-tag, thioredoxin tag, S-tag, So Ftag (e.g. Softag1, Softag3), Streptag, Biotin ligase tag, FLASH tag , V5-tag, and SBP-tag. Further suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein comprises one or more His tags.

[0308] [Linker] In one embodiment, the peptides or peptide domains of the present invention are linked A linker can be used to connect the two. The linker can be as simple as a covalent bond, or It can be a polymer linker that is multiple atoms long. The linker can be a peptide linker. or a non-peptide linker. In certain embodiments, the linker is a UV-cleavable linker. In some embodiments, the linker can be a polynucleotide linker, e.g., For example, it can be an RNA linker. In one embodiment, the linker is a polypeptide. or is amino acid-based. In other embodiments, the linker is a peptide-like In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, In one embodiment, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is cyclic or non-cyclic. It is a cyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In embodiments, the linker is a polymer (e.g., polyethylene, polyethylene glycol). In certain embodiments, the linker is an alkylene, a methacrylate, a methacrylate copolymer ... In some embodiments, the phosphorus phosphate group is a phosphate group, and the phosphorus phosphate group is a phosphate group. Carbohydrates are aminoalkanoic acids (e.g., glycine, ethanoic acid, alanine, beta-alanine) , 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In the present invention, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In some embodiments, the linker is a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker is based on a polyethylene glycol moiety (PEG) In other embodiments, the linker comprises an amino acid. In some embodiments, the linker comprises an aryl or heteroaryl group. In certain embodiments, the linker is based on a phenyl ring. Facilitates binding of nucleophiles (e.g., thiols, aminos) from the peptide to the linker Any electrophile can be used as part of the linker. Exemplary electrophiles include activated esters, activated amides, Michael acceptors, Alkyl halides, aryl halides, acyl halides, and isothiocyanates These include, but are not limited to:

[0309] In certain embodiments, the linker may be one or more amino acids (e.g., a peptide). In some embodiments, the linker is a bond (e.g., a bond or a protein). covalent bond), organic molecule, group, polymer, or chemical moiety. The linker may be about 3 to about 104 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16) in length. , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 , 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 61, 62, 63, 64 , 65, 70, 75, 80, 85, 90, 95, or 100) amino acids.

[0310] [Cas9 complex with guide RNA] Some aspects of the disclosure include multiple fusion proteins comprising any of the fusion proteins provided herein. Provides guidance to achieve optimal length for nucleobase editor activity. RNA can be utilized (e.g., (GGGS) n , (GGGGS) n , and (G) n Highly flexible phosphorus in the form of From Carr, (EAAAK) n , (SGGS) n , to a more rigid linker in the form of SGSETPGTSESATPES (See e.g. Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechn ol. 2014; 32(6): 577-82, the entire contents of which are incorporated herein by reference), and (XP) n In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 , 13, 14 or 15. In some embodiments, the linker is (GGS) n Contains motifs, wherein n is 1, 3, or 7. In some embodiments, The Cas9 domain of the fusion protein is attached via a linker containing the amino acid sequence SGSETPGTSESATPES. and are fused together.

[0311] In some embodiments, the guide nucleic acid (e.g., guide RNA) is 15 to 100 nucleotides in length, It contains a sequence of at least 10 contiguous nucleotides that is complementary to the target sequence. In embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 , 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 In some embodiments, the length of the guide is 47, 48, 49, or 50 nucleotides. The RNA is complementary to the target sequence. 8, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides In some embodiments, the target sequence is a DNA sequence. The sequence may be a sequence in a bacterial, yeast, fungal, insect, plant or animal genome. In some embodiments, the target sequence is a sequence in the human genome. In some embodiments, the 3' end of the target sequence is immediately adjacent to a canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence may contain a non-canonical PAM sequence (e.g., a sequence listed in Table 1). In some embodiments, the guide nucleic acid (e.g., guide RNA) is is complementary to the sequence associated with the disorder.

[0312] Some aspects of the present disclosure involve the use of the fusion proteins or complexes provided herein. For example, some aspects of the disclosure provide methods for preparing DNA molecules according to the methods provided herein. and at least one guide RNA. wherein the guide RNA is about 15 to 100 nucleotides in length and A sequence of at least 10 contiguous nucleotides complementary to the sequence. In this configuration, the 3' end of the target sequence is immediately adjacent to an AGC, GAG, TTT, GTG, or CAA sequence. In some embodiments, the 3' end of the target sequence is NGA, NGCG, NGN, NNGRRT , NNNRRT, NGCG, NGCN, NGTN, NGTN, NGTN, or immediately adjacent to the 5' (TTTV) sequence .

[0313] In some embodiments, the fusion proteins of the invention are capable of mutagenizing a target of interest. These mutations can affect the function of the target. For example, Targeting regulatory regions with a protein editor alters their function and regulates downstream proteins. The expression of quality is reduced.

[0314] The numbering of specific positions or residues in each sequence is based on the specific protein used. It will be understood that the numbering scheme will depend on the protein quality and the numbering scheme. The numbering of protein precursors and the mature proteins themselves may differ, and the numbering may vary by species. Sequence differences may affect numbering. Those skilled in the art will recognize and use the numbers using methods well known to those skilled in the art, e.g. By sequence alignment and determination of homologous residues, any homologous proteins and their Each residue in each encoding nucleic acid can be identified.

[0315] Any of the fusion proteins disclosed herein can be used to target a target site, e.g., a target gene to be edited. To target the mutation-containing site, the fusion protein is co-expressed with a guide RNA. It will be apparent to those skilled in the art that it is typically necessary to realize As described in more detail elsewhere herein, guide RNAs typically comprise a transcription factor that allows for Cas9 binding. Conferring sequence specificity to the crRNA framework and Cas9:nucleic acid editing enzyme / domain fusion protein Alternatively, the guide RNA and tracrRNA may be combined as two nucleic acid molecules. In some embodiments, the guide RNA may be provided separately from the target sequence. The guide sequence is typically 20 nucleotides in length. Cas9: A nucleic acid editing enzyme / domain fusion protein targeted to a specific genomic target site. The sequence of a suitable guide RNA for use in translating the gene will be apparent to one of skill in the art based on the present disclosure. Suitable guide RNA sequences such as these are typically within 50 nucleotides of the target nucleotide to be edited. The fusion sequence includes a guide sequence complementary to a nucleic acid sequence within the upstream or downstream region of the nucleotide sequence. Some exemplary guides suitable for targeting any of the proteins to a specific target sequence are: The RNA sequences are provided herein.

[0316] [Fusion protein containing cytidine deaminase, adenosine deaminase, and Cas9 domains How to use the Some aspects of the present disclosure involve the use of the fusion proteins or complexes provided herein. For example, some aspects of the disclosure provide methods for preparing DNA molecules according to the methods provided herein. and at least one guide RNA. wherein the guide RNA is about 15 to 100 nucleotides in length and A sequence of at least 10 contiguous nucleotides complementary to the sequence. In some configurations, the 3' end of the target sequence is immediately adjacent to the canonical PAM sequence (NGG). In embodiments, the 3' end of the target sequence is not immediately adjacent to the canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is an AGC, GAG, TTT, GTG, or CAA sequence. In some embodiments, the 3' end of the target sequence is immediately adjacent to NGA, NGCG, Immediately adjacent to the NGN, NNGRRT, NNNRRT, NGCG, NGCN, NGTN, NGTN, or 5' (TTTV) sequence It is connected.

[0317] In some embodiments, the fusion proteins of the invention are capable of mutagenizing a target of interest. In particular, the multi-effector nucleobase editors described herein are used for can generate multiple mutations in the target sequence. These mutations affect the function of the target. For example, multi-effector nucleobase editors can be used to modify regulatory regions. Targeting alters the function of the regulatory region, resulting in decreased expression of downstream proteins.

[0318] The numbering of specific positions or residues in each sequence is based on the specific protein used. It will be understood that the numbering scheme will depend on the protein quality and the numbering scheme. The numbering of protein precursors and the mature proteins themselves may differ, and the numbering may vary by species. Sequence differences may affect numbering. Those skilled in the art will recognize and use the numbers using methods well known to those skilled in the art, e.g. By sequence alignment and determination of homologous residues, any homologous proteins and their Each residue in each encoding nucleic acid can be identified.

[0319] The Cas9 domain and cytidine deaminase or adenosine deaminase as disclosed herein The fusion protein containing the enzyme is then inserted into the target site, e.g., the target site containing the mutation to be edited. To target the site, the fusion protein is co-transfected with a guide RNA, e.g., an sgRNA. It will be apparent to one skilled in the art that expression is typically required. As described in more detail in section 1, the guide RNA typically comprises a sequence that allows for Cas9 binding. Conferring sequence specificity to the racrRNA framework and Cas9:nucleic acid editing enzyme / domain fusion protein Alternatively, the guide RNA and tracrRNA may be two nucleic acid molecules. In some embodiments, the guide RNA may be provided separately from the guide sequence. The guide sequence is typically 20 nucleotides long. Cas9: A nucleic acid editing enzyme / domain fusion protein targeted to a specific genomic target site. The sequence of a suitable guide RNA for targeting will be apparent to one of skill in the art based on the present disclosure. Such a suitable guide RNA sequence typically contains 50 nucleotides of the target nucleotide to be edited. The nucleic acid sequence provided herein includes a guide sequence complementary to a nucleic acid sequence within the upstream or downstream region of the nucleic acid sequence. Some exemplary guides suitable for targeting any of the fusion proteins to a specific target sequence are: Id RNA sequences are provided herein.

[0320] [Base editor efficiency] The fusion proteins of the present invention can be used to cleave specific nuclei without generating significant rates of indels. By modifying the bases, the base editor efficiency is improved. As used herein, "indel" refers to the insertion or deletion of a nucleotide base within a nucleic acid. Such insertions or deletions result in frameshift mutations within the coding region of a gene. In some embodiments, multiple insertions or deletions in the nucleic acid may occur. Efficiently amplify specific nucleotides within a nucleic acid without introducing deletions (i.e., indels) In certain embodiments, it is desirable to generate base editors that modify (e.g., mutate) is intended modification relative to an indel. A greater proportion of variations (e.g., mutations) can be generated. In some embodiments, the base editors provided herein provide for greater than 1:1 targeted mutations. In some embodiments, a ratio of indels to indels can be generated. The base editors used are at least 1.5:1, at least 2:1, at least 2.5:1, or less at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least At least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600: 1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1; or higher. The number of mutations and indels can be determined using any suitable method.

[0321] In some embodiments, the base editors provided herein modify regions of a nucleic acid. In some embodiments, the region can limit the formation of indels in the base at the nucleotide targeted by the editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by In some embodiments, any of the base editors provided herein The percentage of indels in the nucleic acid region is less than 1%, less than 1.5%, less than 2%, less than 2.5%, less than 3%. Full, Less than 3.5%, Less than 4%, Less than 4.5%, Less than 5%, Less than 6%, Less than 7%, Less than 8%, Less than 9%, Less than 10% The amount of the nucleic acid region formed can be limited to less than 12%, less than 15%, or less than 20%. The number of bases represents the time that a nucleic acid (e.g., a nucleic acid in a cell's genome) is exposed to a base editor. In some embodiments, the number or percentage of indels may depend on the amount of the nucleic acid (e.g., For example, nucleic acids in the genome of a cell are exposed to the base editor for at least 1 hour and at least 2 hours, at least 6 hours, at least 12 hours, at least 24 hours, at least 36 hours, At least 48 hours, at least 3 days, at least 4 days, at least 5 days, at least 7 days, at least A decision will be made in at least 10 days, or at least 14 days.

[0322] Some aspects of the disclosure include, but are not limited to, any of the base editors provided herein. in a nucleic acid (e.g., a nucleic acid within a subject's genome) without generating an inordinate number of unintended mutations. This is based on the recognition that it is possible to efficiently generate intended mutations using several In this embodiment, the intended mutation is specifically designed to produce the intended mutation. These mutations are generated by specific base editors bound to gRNAs designed for this purpose. In some embodiments, the intended mutation is a stop codon, e.g., a stop codon in the In some embodiments, the mutation is a mutation that creates a premature stop codon within the gene region. The intended mutation is one that removes a stop codon. The mutations that are targeted are mutations that alter the splicing of the gene. In this case, the intended mutation may be in a regulatory sequence of the gene (e.g., a gene promoter or In some embodiments, the mutations described herein are mutations that alter the function of the gene encoding the gene (the gene repressor). Any of the base editors offered in Mutation ratios (e.g., intended mutations:unintended mutations) greater than 1:1 In some embodiments, any of the base editors provided herein can The ratio of intended to unintended mutations should be at least 1.5:1, but not less than 1.5:1. at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least at least 50:1, at least 100:1, at least 150:1, at least 200:1, at least 250:1, It can be at least 500:1, or at least 1000:1. The features of base editors described in the "Base Editor Efficiency" section are useful in the fusion It is understood that the present invention can be applied to methods using either proteins or fusion proteins. It should be.

[0323] [Method for editing nucleic acids] Some aspects of the present disclosure provide methods for editing nucleic acids. In one embodiment, the method comprises: In some embodiments, the method comprises: a) detecting a nucleic acid (e.g., a double-stranded DNA sequence) contacting a target region of the target sequence with a complex comprising a base editor and a guide nucleic acid (e.g., gRNA). b) inducing strand separation of said target region; and c) isolating a single strand of the target region. d) converting a first nucleobase of said target nucleobase pair to a second nucleobase; and and cleaving no more than one strand of the target region using A third nucleobase complementary to the first nucleobase is linked to a fourth nucleobase complementary to the second nucleobase. In one embodiment, the method comprises the step of: It should be understood that in some embodiments, step b is omitted. In some embodiments, the method provides a 100% or less ... Less than 12%, less than 10%, less than 8%, less than 6%, less than 4%, less than 2%, less than 1%, less than 0.5%, less than 0.2%, or In some embodiments, the method results in less than 0.1% or less indel formation. a fifth nucleobase that is complementary to the fourth nucleobase, thereby achieving the intended In some embodiments, the method further comprises generating an edited base pair (e.g., G·C to A·T). In some embodiments, at least 5% of the intended base pairs are edited. At least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the depicted base pairs are fused. To be gathered.

[0324] In some embodiments, the intended versus unintended product at the target nucleotide is The ratio of ingredients is at least 2:1, at least 5:1, at least 10:1, at least 20:1, at least 30:1, at least 40:1, at least 50:1, at least 60:1, at least 70:1, at least 80:1, at least 90:1, at least 100:1, or at least 200:1, or more In some embodiments, the ratio of intended mutation to indel formation is 1: greater than 1, greater than 10:1, greater than 50:1, greater than 100:1, greater than 500:1, or greater than 1000:1 or more. In some embodiments, the cleaved single strand (nicked strand) hybridizes to the guide nucleic acid. In some embodiments, the cleaved single strand is cleaved to a region containing the first nucleobase. In some embodiments, the base editor comprises a dCas9 domain. In some embodiments, the base editor protects or bonds the unedited strand. In some embodiments, the intended editing base pair is upstream of the PAM site. In one embodiment, the intended editing base pair is 1, 2, 3, 4, 5, 6, 7, 8 of the PAM site. , 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream. In some embodiments, the intended editing base pair is downstream of the PAM site. In this embodiment, the intended editing base pairs are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 3 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream. In some embodiments, the method does not require a canonical (e.g., NGG) PAM site. In embodiments, the nucleobase editor comprises a linker. In some embodiments, the linker is a long In some embodiments, the linker is 5 to 20 amino acids in length. In some embodiments, the linker is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or or 20 amino acids. In one embodiment, the linker is 32 amino acids in length In another embodiment, a "long linker" is at least about 60 amino acids in length. In another embodiment, the linker is about 3 to 100 amino acids in length. In this embodiment, the target region comprises a target window, and the target window comprises a target nucleic acid base pair. In some embodiments, the target window comprises 1 to 10 nucleotides. In this embodiment, the target window has a length of 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, In some embodiments, the target window is 1 to 3, 1 to 2, or 1 nucleotide. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 In some embodiments, the intended editing base pair is a target wink. In some embodiments, the target window is within the intended editing base pair. In some embodiments, the method comprises: It is carried out using either

[0325] In some embodiments, the present disclosure provides a method for editing nucleotides. In some embodiments, the present disclosure provides a method for editing nucleic acid base pairs in a double-stranded DNA sequence. In some embodiments, the method comprises: a) detecting a target of double-stranded DNA sequence; contacting the region with a complex comprising a base editor and a guide nucleic acid (e.g., gRNA), wherein the target region comprises a target nucleic acid base pair; and b) inducing strand separation of the target region. and c) replacing a first nucleobase of the target nucleobase pair in the single strand of the target region with a second nucleobase. and d) cleaving no more than one strand of the target region. wherein a third nucleobase complementary to the first nucleobase is complementary to a second nucleobase. a fourth nucleobase, and the second nucleobase is replaced by a fourth nucleobase that is complementary to the fourth nucleobase. the fifth nucleobase is replaced by the fifth nucleobase, thereby generating the intended edited base pair, In some embodiments, the efficiency of generating the intended base pair is at least 5%. It should be understood that b is omitted. In some embodiments, the intended base In some embodiments, at least 5% of the intended base pairs are edited. In some embodiments, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the image is edited. In the present invention, the present invention provides a method for reducing the , less than 6%, less than 4%, less than 2%, less than 1%, less than 0.5%, less than 0.2%, or less than 0.1% indel forms In some embodiments, the intended product pair at the target nucleotide is The ratio of unintended products is at least 2:1, at least 5:1, at least 10:1, at least at least 20:1, at least 30:1, at least 40:1, at least 50:1, at least 60:1, at least at least 70:1, at least 80:1, at least 90:1, at least 100:1, or at least 200:1 In some embodiments, the intended mutation versus indel is The ratio of formation is greater than 1:1, greater than 10:1, greater than 50:1, greater than 100:1, greater than 500:1, or greater than 1000:1, or In some embodiments, the cleaved single strand h...

Claims

1. 1. A composition comprising an adeno-associated virus (AAV) vector, (a) a first AAV vector comprising a polynucleotide encoding a fusion protein comprising, in N-terminal to C-terminal order, a deaminase and an N-terminal fragment of Cas9, wherein the N-terminal fragment of Cas9 is a contiguous sequence beginning at the N-terminus of Cas9 and ending at position 302, 309, 312, or 354 of Cas9 as numbered in SEQ ID NO:2, wherein the N-terminal fragment of Cas9 is fused at its C-terminus to a split intein-N; and (b) a second AAV vector comprising a polynucleotide encoding a C-terminal fragment of Cas9, the C-terminal fragment of Cas9 being a contiguous sequence beginning at position S303 of Cas9 if the N-terminal fragment terminates at position 302, at position T310 of Cas9 if the N-terminal fragment terminates at position 309, at position T313 of Cas9 if the N-terminal fragment terminates at position 312, or at position S355 of Cas9 if the N-terminal fragment terminates at position 354, as numbered in SEQ ID NO:2, and ending at the C-terminus of Cas9, wherein the C-terminal fragment of Cas9 is fused at its N-terminus to a split intein-C. A composition comprising:

2. 1. A composition comprising an adeno-associated virus (AAV) vector, (a) a first AAV vector comprising a polynucleotide encoding a fusion protein comprising, in N-terminal to C-terminal order, a deaminase and an N-terminal fragment of Cas9, wherein the N-terminal fragment of Cas9 is a contiguous sequence beginning at the N-terminus of Cas9 and ending at position 302, 309, 312, 354, 465, 471, or 576 of Cas9 as numbered in SEQ ID NO: 2, wherein the N-terminal fragment of Cas9 is fused at its C-terminus to a split intein-N; and (b) a second AAV vector comprising a polynucleotide encoding a C-terminal fragment of Cas9, wherein the C-terminal fragment of Cas9 is at position S303 of Cas9 if the N-terminal fragment terminates at position 302, at position T310 of Cas9 if the N-terminal fragment terminates at position 309, at position T313 of Cas9 if the N-terminal fragment terminates at position 312, or at position S355 of Cas9 if the N-terminal fragment terminates at position 354, according to the numbering in SEQ ID NO:2; a contiguous sequence beginning at position T466 of Cas9 if the N-terminal fragment terminates at position 465, at position T472 of Cas9 if the N-terminal fragment terminates at position 471, or at position S577 of Cas9 if the N-terminal fragment terminates at position 576, and terminating at the C-terminus of Cas9, wherein the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted for Ala, Ser, or Thr, and wherein the C-terminal fragment of Cas9 is fused at its N-terminus to a split intein-C. A composition comprising:

3. a) the composition further comprises a single guide RNA (sgRNA) or a polynucleotide encoding same; b) the composition further comprises an AAV vector comprising an sgRNA or a polynucleotide encoding same; or c) the first AAV vector or the second AAV vector comprises an sgRNA or a polynucleotide encoding the sgRNA; The composition according to claim 1 or 2.

4. 3. The composition of claim 1 or 2, wherein the deaminase is adenosine deaminase.

5. 3. The composition of claim 1 or 2, wherein the deaminase is TadA or a variant thereof.

6. 3. The composition of claim 1 or 2, wherein the deaminase is wild-type TadA or TadA7.

10.

7. the deaminase is a TadA dimer; or the deaminase is a TadA dimer comprising wild-type TadA and TadA7.10; The composition according to claim 1 or 2.

8. The composition of claim 1 or 2, wherein the fusion protein comprises a nuclear localization signal (NLS).

9. 3. The composition of claim 1 or 2, wherein the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is linked to a nuclear localization signal (NLS).

10. 3. The composition of claim 1 or 2, wherein the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is linked to a bipartite NLS.

11. 3. The composition of claim 1 or 2, wherein the Cas9 has nickase activity or is catalytically inactive.

12. The first AAV vector and / or the second AAV vector are a) comprises a promoter; or b) comprises a constitutive promoter; or c) containing a constitutive promoter that is a CMV or CAG promoter; The composition according to claim 1 or 2.

13. 3. The composition of claim 1 or 2, wherein the C-terminal fragment of Cas9 comprises an Ala / Cys, Ser / Cys, or Thr / Cys mutation at a residue corresponding to amino acid S303, T310, T313, or S355 in the numbering of SEQ ID NO:

2.

14. 3. The composition of claim 1 or 2 for use in a method for delivering a base editor system to a cell, the method comprising contacting the cell with the first and second AAV vectors and a single guide RNA (sgRNA) or a polynucleotide encoding same.

15. 10. An in vitro or ex vivo method for delivering a base editor system to a cell, the method comprising contacting the cell with a first and a second AAV vector of claim 1 or 2 and a single guide RNA (sgRNA) or a polynucleotide encoding same.