Novel nucleic acid base editor and method of use thereof

Novel adenosine deaminase domains and programmable nucleobase editors, like TadA*9, enhance the specificity and efficiency of nucleotide modifications in genomic DNA, addressing the limitations of current base editors.

JP2025170240APending Publication Date: 2025-11-18BEAM THERAPEUTICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025122950
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2025-07-23
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Current base editors lack high specificity and efficiency in modifying target nucleic acid sequences, particularly in inducing desired modifications in genomic DNA.

Method used

Development of adenosine deaminase domains, such as TadA*9, and novel programmable nucleobase editors, including fusion proteins with DNA-binding domains and base editor domains, to enhance specificity and efficiency in targeted nucleotide modifications.

Benefits of technology

The novel base editors achieve precise and efficient editing of genomic sequences, reducing off-target effects and improving the accuracy of nucleotide modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170240000115
    Figure 2025170240000115
  • Figure 2025170240000116
    Figure 2025170240000116
  • Figure 2025170240000117
    Figure 2025170240000117
Patent Text Reader

Abstract

To provide an improved base editor capable of inducing modifications within a target sequence with higher specificity and efficiency.SOLUTION: In one embodiment, provided is an adenosine deaminase comprising a modification at a specific amino acid position of a specific sequence, or a corresponding modification in another adenosine deaminase.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 897,777, filed September 9, 2019. International PCT application No. PCT / US2020 / 018195, filed February 13, 2020. No. 60 / 699,493, filed on May 1, 2003, the entire contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Targeted editing of nucleic acid sequences, e.g., targeted cleavage or specific modification (alteration) of genomic DNA Targeted delivery of nucleotides is a very promising approach for studying gene function and has been shown to be effective in the study of human genetic Currently available base editors target C Cytidine base editors (e.g., BE4) that convert G base pairs to T·A and A·T to G·C Examples of suitable adenine base editors include ABE7.10 and other adenine base editors that are well known in the art. There is a need for improved base editors that can induce modifications in target sequences with high specificity and efficiency. It has been done. Summary of the Invention

[0003] As described below, the present invention provides adenosine deaminase domains (e.g., TadA*9 or ABE9), and a novel programmable nucleobase editor, In some embodiments, the ABE9 of the present invention is used for editing a gene encoding a protein. Polynucleotides, e.g., polynucleotides containing pathogenic mutations associated with genetic disorders. Gather.

[0004] In one embodiment, 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139 of SEQ ID NO: 1 , 146, and 158, or another adenosine Adenosine deaminase with the corresponding modifications in the deaminase: In one embodiment, the adenosine deaminase is selected from the group consisting of R21N, R23H, E25 of SEQ ID NO: 1. F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C 146R, and A158K, or another adenosine deaminase In one embodiment, the adenosine deaminase comprises the corresponding modification in V of SEQ ID NO: 1. 82T modification, or a corresponding modification in another adenosine deaminase. In embodiments, the adenosine deaminase is selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71 of SEQ ID NO: 1. , 72, 73, 94, 124, 133, 139, 146, and 158 or a corresponding modification in another adenosine deaminase. The adenosine deaminase of this aspect and its embodiments comprises two or more modifications. In embodiments, the adenosine deaminase of this aspect and its embodiments comprises three or more of the In one embodiment, the adenosine deaminase of this aspect and embodiments thereof comprises the modification Further comprising one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In embodiments, the adenosine deaminase of this aspect and embodiments thereof comprises one of the following groups of modifications: Contains one of the following: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K_V82S+Y123H+Y147R+Q154R; Q71M_V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In one embodiment, the adenosine deaminase variant is any of those listed in Tables 14 or 18. In one embodiment, the adenosine derivative of this aspect and its embodiments comprises a modification or group of modifications. The aminase is selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the C-terminal deletion of this aspect and its embodiments includes a C-terminal deletion beginning at the residue Adenosine deaminase is derived from Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R In one embodiment, the present invention further comprises a modification selected from the group consisting of: The adenosine deaminase may be any of the adenosine deaminases listed in Table 14, Table 18, or Figures 3A-3C. It is an aminase variant.

[0005] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide protease. a configurable DNA binding domain and the following sequences: 21, 23, 25, 38, 51, 54, 70, 71, 72 of SEQ ID NO: 1 , 73, 94, 124, 133, 139, 146, and 158. adenosine deaminase containing a corresponding modification in another adenosine deaminase and at least one base editor domain that is a basease variant: JPEG2025170240000002.jpg41165

[0006] In one embodiment, the adenosine deaminase variant is R21N, R23H, E25 of SEQ ID NO: 1. F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C 146R, and A158K, or another adenosine deaminase This includes the corresponding modifications in

[0007] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide protease. a mutable DNA-binding domain and R21N, R23H, E25F, N38G, L51W, P54C, M70V of SEQ ID NO: 1 , Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K. or a corresponding modification in another adenosine deaminase. at least one base editor domain that is a denosine deaminase variant; include.

[0008] In any embodiment of the fusion protein of any of the above aspects and embodiments thereof The adenosine deaminase variant may be the V82T modification of SEQ ID NO: 1, or another adenosine deaminase variant. It further includes corresponding modifications in the deaminase.

[0009] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide protease. a mutable DNA binding domain and the modifications V82T and R21N, R23H, E25F, N38G, L51W of SEQ ID NO: 1 , P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A15 8K, or another adenosine deaminase adenosine deaminase variants containing at least one base residue corresponding to the corresponding modification and a debate domain.

[0010] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, adeno Syndeaminase variants are 21, 23, 25, 38, 51, 54, 70, 71, 72, 73 of SEQ ID NO: 1 , 94, 124, 133, 139, 146, and 158 or a corresponding modification in another adenosine deaminase. In one embodiment, the adenosine deaminase variant comprises two or more modifications. The adenosine deaminase variant comprises three or more modifications. The deaminase variant further comprises one or more of the following modifications: Y147T, Y147R , Q154S, Y123H, and Q154R. In one embodiment, the adenosine deaminase variant is 49, 150, 151, 152, 153, 154, 155, 156, and 157. It contains a C-terminal deletion.

[0011] In embodiments of the above fusion proteins and embodiments thereof, the base editor domain comprises an adenosine deaminase variant monomer, and the adenosine deaminase monomer is , R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M of SEQ ID NO: 1 One or more selected from the group consisting of 94V, P124W, T133K, D139L, D139M, C146R, and A158K In one embodiment, the base editor domain comprises a modification of the wild-type adenosine deaminase Adenosine deaminase heterodimers containing phosphodiesterase domains and adenosine deaminase variants In one embodiment, the adenosine deaminase variant comprises a dimer of Y147T, Y147R, Further comprising a modification selected from the group consisting of Q154S, Y123H, V82S, T166R, and Q154R. In embodiments, the base editor domain comprises a TadA*7.10 domain and an adenosine deaminase domain. In one embodiment, the present invention includes an adenosine deaminase heterodimer comprising an adenosine deaminase variant domain. In the present specification, the adenosine deaminase variant contains two or more modifications.

[0012] In another embodiment of the fusion protein of any of the above aspects and embodiments thereof, the adenovirus The syndeaminase variants are ABE9 (TadA*9 deaminase) variants listed in Table 14, Table 18, or Figures 3A-3C. minase variant).

[0013] In another embodiment of the fusion protein of any of the above aspects and embodiments thereof, the adenovirus Syndeaminase variants 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 compared to full-length ABE9 , truncations with deletion of 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues It is type ABE8 or ABE9.

[0014] In another embodiment of the fusion protein of any of the above aspects and embodiments thereof, the polynucleotide Nucleotide-programmable DNA-binding domains include Cas9, Cas12a / Cpf1, Cas12b / C2c1, and Cas 12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ It's the main thing.

[0015] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide comprising the sequence: Nucleotide-programmable DNA-binding domains include: JPEG2025170240000003.jpg166160 In the sequence, the bolded sequence represents a sequence derived from Cas9, and the italicized sequence represents a linker sequence. The underlined sequences indicate bipartite nuclear localization sequences, and at least one The base editor domains are 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 9 of SEQ ID NO: 1. The alterations are at amino acid positions selected from the group consisting of 4, 124, 133, 138, 139, 146, and 158. In one embodiment, the adenosine deaminase variant comprises The variants are R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, and N72K of SEQ ID NO: 1. , Y73S, M94V, P124W, T133K, D138M, D139L, D139M, C146R, and A158K. In another embodiment, the adenosine deaminase variant comprises a selected modification of the sequence In one embodiment, the adenosine deaminase variant comprises two or more modifications V82T, numbered 1. In one embodiment, the adenosine deaminase variant comprises three or more of the above modifications. In one embodiment, the adenosine deaminase variant comprises the modifications Y147T, Y14 7R, Q154S, Y123H, V82S, T166R, and Q154R. In one embodiment, the adenosine deaminase variant has two or more of the following modifications: Including: Y147T, Y147R, Q154S, Y123H, and Q154R.

[0016] In any of the above fusion proteins and embodiments thereof, adenosine deaminase is The enzyme variant comprises any one of the following groups of modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R In one embodiment, the adenosine deaminase variant is selected from the group consisting of those listed in Table 14 or 18, or those listed in FIG. 3A. 3C to 3C.

[0017] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the polynucleotide The programmable DNA-binding domain is Staphylococcus aureus Cas9 (SaCas9), St reptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof.

[0018] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the polynucleotide The otide programmable DNA-binding domain contains an engineered protospacer adjacent motif In one embodiment, the modified SaCas9 comprises a modified SaCas9 with PAM specificity. The amino acid substitutions include E782K, N968K, and R1015H, or their corresponding amino acid substitutions.

[0019] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the polynucleotide The otide programmable DNA-binding domain contains an engineered protospacer adjacent motif SpCas9 variants with PAM specificity include: the nucleic acid sequence 5'-NGA-3', 5'-NGC-3', 5'-NGG-3', 5'-NGT-3', or 5''-NGN-3' In one embodiment, the variant SpCas9 has specificity for the following: D1135M, S1136Q an amino acid substitution selected from G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R; or the corresponding amino acid substitutions: I322V, S409I, E427G, R654L, R753G (MQKFRAE R) or their corresponding amino acid substitutions: I322V, S409I, E427G, R654L, R753G, R111 4G, or a corresponding amino acid substitution thereof; or an amino acid substitution as set forth in Figures 3A-3C nothing.

[0020] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the polynucleotide The oxidase-programmable DNA-binding domain is nuclease-inactive or nickase-inactive In one embodiment, the nickase variant has the amino acid substitution D10A or Contains the amino acid substitutions corresponding to:

[0021] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, adenosine The deaminase domain deaminates adenine in deoxyribonucleic acid (DNA). can be done.

[0022] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, adenosine The deaminase is a modified adenosine deaminase that does not occur in nature.

[0023] In an adenosine deaminase embodiment of the above aspects and embodiments thereof, adenosine deaminase The aminase is TadA deaminase. In a protein embodiment, the adenosine deaminase is TadA deaminase. In an embodiment, the TadA deaminase is a TadA*7.10 variant.

[0024] In a fusion protein embodiment of any of the above aspects and embodiments thereof, the fusion protein The protein contains a polynucleotide-programmable DNA-binding domain and an adenosine deaminase A linker is included between the domains. In one embodiment, the linker has the amino acid sequence: SGGSSG Includes GSSGSETPGTSESATPES.

[0025] In a fusion protein embodiment of any of the above aspects and embodiments thereof, the fusion protein The protein comprises one or more nuclear localization signals. In one embodiment, the nuclear localization signal is a Hypertite nuclear localization signal.

[0026] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the Cas9 is It's Cas9.

[0027] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, the Cas9 is Cas9 or SpCas9.

[0028] In an embodiment of the fusion protein of any of the above aspects and embodiments thereof, Cas9 is In one embodiment, the modified SaCas9 has the amino acid substitution E782K , N968K, and R1015H, or the amino acid substitutions corresponding thereto. The modified SaCas9 comprises the following amino acid sequence: KRNYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG.

[0029] In another aspect, a method for producing a fusion protein comprising the steps of: A polynucleotide that

[0030] In another aspect, there is provided a cell, the cell being a cell of any one of the above aspects and embodiments thereof. A polynucleotide encoding one fusion protein and a targeting base editor , one or more guide polynucleotides that result in an A·T to G·C change at a SNP associated with a genetic disease The antibody is produced by introducing a nucleotide into a cell, or a precursor cell thereof. In one embodiment, the cells are human cells. In one embodiment, the cells are in vitro or in vivo. In one embodiment, the genetic disease is alpha-1 antitrypsin deficiency (A1AD). In embodiments, the fusion protein and one or more guide polynucleotides form a complex within the cell. Form.

[0031] In another aspect, an isolated cell grown or expanded from the cells of the above aspects and embodiments thereof. A cell or population of cells is provided.

[0032] In one aspect, a method for treating a genetic disease in a subject in need thereof is provided. and a method for producing a cell, an isolated cell, or a cell of any one of the above aspects and embodiments thereof. In one embodiment of the method, the cell, isolated cell, or population is administered to the subject. The population of cells may be autologous, allogeneic, or xenogeneic to the subject.

[0033] In one aspect, a base editor system is provided, the base editor system comprising: and the following sequences: 21, 23, 25, 38, 51, 54 of SEQ ID NO: 1; , 70, 71, 72, 73, 82, 94, 124, 133, 139, 146, and 158 adenosine deaminase containing a modification at the amino acid position, adenosine deaminase containing a corresponding modification in another adenosine deaminase and at least one base editor domain that is a deaminase variant: In an embodiment of the JPEG2025170240000004.jpg49160 base editor system, the adenosine deaminase variant is R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, and M94 in SEQ ID NO: 1 a modification selected from the group consisting of V, P124W, T133K, D139L, D139M, C146R, and A158K, or or a corresponding modification in another adenosine deaminase. The editor system targets base editor domains to identify genes associated with genetic disorders. Further comprising one or more guide polynucleotides that result in the SNP A·T to G·C change In one embodiment of the base editor system, the adenosine deaminase variant is It can deaminate adenine in hydroxyribonucleic acid (DNA). In one embodiment of the stem, the guide polynucleotide is ribonucleic acid (RNA), or deoxyribonucleic acid (DEA). In one embodiment of the base editor system, the guide polynucleotide The nucleotides are CRISPR RNA (crRNA) sequences, trans-activating CRISPR RNA (tracrRNA) sequences, or a combination thereof. In one embodiment, the base editor system comprises a second In one embodiment, the second guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). The guide polynucleotides include CRISPR RNA (crRNA) sequences, trans-activating CRISPR RNA (traRNA) sequences, and crRNA) sequences, or a combination thereof. In an embodiment of the present invention, the polynucleotide programmable DNA binding domain is 9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas1 In one embodiment, the polynucleotide protease comprises a Cas12h, Cas12i, or Cas12j / CasΦ domain. The programmable DNA binding domain is nuclease inactive. The nucleotide-programmable DNA-binding domain is a nickase. The polynucleotide programmable DNA binding domain comprises a Cas9 domain. In this state, the Cas9 domain is expressed as a nuclease-inactive (dead) Cas9 (dCas9), a Cas9 nickase (nCas9), or nuclease-active Cas9. In one embodiment, the Cas9 domain comprises a Cas In one embodiment, the polynucleotide comprises a programmable DNA binding domain. The gene is an engineered or modified polynucleotide programmable DNA binding domain. In an embodiment of the above base editor system and its embodiments, a genetic disorder is alpha-1 antitrypsin deficiency (A1AD).

[0034] In another aspect, a method for correcting a single nucleotide polymorphism (SNP) in a polynucleotide is provided, the method comprising: a target nucleotide sequence (at least a portion of which is a polynucleotide or its reverse complement) (located in the fusion protein of any one of the above aspects and embodiments thereof, or contacting with a base editor system of any one of the above aspects and embodiments thereof; When targeting a base editor to a target nucleotide sequence, the SNP or its complementary nucleic acid and editing the SNP by deaminating the base, wherein the SNP or its complementary nucleic acid salt is In one embodiment, the SNP is corrected by deaminating the alpha-1 antigen. In one embodiment, the SNP is located in the SERPINA1 gene and is associated with trypsin deficiency (A1AD). The modification includes the alteration of E342K (PiZ allele).

[0035] In one aspect, a method for editing a polynucleotide is provided, the method comprising: , a fusion protein according to any one of the above aspects and embodiments thereof, or a fusion protein according to any one of the above aspects and embodiments thereof. and contacting the polynucleotide with a base editor system according to any one of the embodiments. In one embodiment of the method, the editing comprises editing less than 20% indels (i indel) formation, less than 15% indel formation, less than 10% indel formation; less than 5% indel formation; <4% indel formation; <3% indel formation; <2% indel formation; <1% indel formation less than 0.5% indel formation; or less than 0.1% indel formation. In embodiments, the editing does not result in a translocation.

[0036] In another embodiment, a TadA*7.10 adenosine deaminase variant selected from: ABE9 (TadA*9 deaminase variant) and Cas9 endonuclease domain containing hmm: monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+A109S of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+T111R of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+D119N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+H122N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147d+Q154S of SEQ ID NO: 1, and the mutations I322V, S409 I, spCas9 (MQKFRAER) with E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+F149Y of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+T166I of SEQ ID NO: 1, and the mutation I322V spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; and monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+D167N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G. monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+L36H+N157K of SEQ ID NO: 1, and the mutations spCas9 (MQKFRAER) with I322V, S409I, E427G, R654L, R753G, and R1114G; monoTadA*7 with the mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K of SEQ ID NO:1. 10, and SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G; monoT with the mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W of SEQ ID NO: 1 adA*7.10 and SpCas9 (MQKFR AER); monoTadA with the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N of SEQ ID NO: 1 *7.10, and SpCas9, MQKFRAER, with mutations I322V, S409I, E427G, R654L, R753G, and R1114G ; and A moiety having the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W of SEQ ID NO: 1 noTadA*7.10, and SpCas9(MQ KFRAER); and adenosine deaminase variant domain targeting to One or more guide polynucleotides that result in an A·T to G·C change at a SNP associated with a genetic disease. Ochido In one embodiment of the base editor, the SNP is It is associated with antitrypsin deficiency (A1AD).

[0037] In another embodiment, a TadA adenosine deaminase domain and an SpCas9 enzyme selected from: one or more polynucleotides encoding ABE9 base editors containing a endonuclease domain; A vector containing the nucleotide is provided. monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+A109S and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+T111R and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+D119N and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+H122N and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147d+Q154S and mutations I322V, S409I, E427G, R65 4L, spCas9 with R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+F149Y and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+T166I and mutations I322V, S409I, E427 spCas9 (MQKFRAER) with G, R654L, and R753G; and monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+D167N and mutations I322V, S409I, E427 G, spCas9 (MQKFRAER) with R654L and R753G. monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+L36H+N157K, mutations I322V, S409I, E spCas9 (MQKFRAER) with 427G, R654L, R753G, and R1114G; monoTadA*7.10 with mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K and mutation I SpCas9 (MQKFRAER) with 322V, S409I, E427G, R654L, R753G, and R1114G; monoTadA*7.10 and monoTadA*7.10 with mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W and SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G; monoTadA*7.10 with mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N and SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G; and monoTadA*7.10 with mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W , SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G. In embodiments, the vector is a plasmid, viral, or mRNA vector.

[0038] In another aspect, the fusion protein of any one of the above aspects and embodiments thereof; or A composition comprising a base editor system according to any one of the above aspects and embodiments thereof is provided. In one embodiment, the composition comprises a pharmaceutically acceptable excipient, diluent, or carrier. Further includes:

[0039] In another aspect, a fusion of any one of the above aspects and embodiments thereof linked to a guide RNA. The present invention provides a composition comprising a protein, wherein the guide RNA is a gene encoding a gene that inhibits alpha-1 antitrypsin deficiency (A1AD). The present invention comprises a nucleic acid sequence complementary to the SERPINA1 gene associated with

[0040] In another aspect, a base of any one of the above aspects and embodiments thereof attached to a guide RNA The present invention provides a composition comprising an editor system, wherein the guide RNA is a gene encoding a gene for an alpha-1 antitrypsin deficiency. It comprises a nucleic acid sequence complementary to the SERPINA1 gene associated with (A1AD).

[0041] In a composition embodiment of any one of the above aspects and embodiments thereof, adenosine The adenine deamination variant can deaminate adenine in deoxyribonucleic acid (DNA). can.

[0042] In a composition embodiment of any one of the above aspects and embodiments thereof, the fusion protein or a base editor system, (i) containing Cas9 nickase; (ii) containing a nuclease-inactive Cas9; (iii) SpCas9 variants containing a combination of amino acid substitutions shown in Figures 3A-3C; and (iv) I322V, S409I, E427G, R654L, R753G (MQKFRAER); or I322V, S409I, E427G R654L, R753G, and R1114G (MQKFRAER). SpCas9 variants including

[0043] In a composition embodiment of any one of the above aspects and embodiments thereof, the composition comprises a pharmaceutical The composition further comprises a physiologically acceptable excipient, diluent, or carrier, ie, a pharmaceutical composition.

[0044] In one embodiment, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier. In one embodiment of the pharmaceutical composition, the pharmaceutical composition comprises: The disease or disorder is alpha-1 antitrypsin deficiency (A1AD). In this form, the fusion protein or base editor system is attached to a guide RNA and The nuclear RNA is complementary to the SERPINA1 gene, which is associated with alpha-1 antitrypsin deficiency (A1AD). In one embodiment of the pharmaceutical composition, the gRNA and the base editor are In the above pharmaceutical compositions and embodiments thereof, the gRNA is formulated separately from the 5' From 3', 5'-ACCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCA AC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-CCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAA C UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-CAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-UCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC U UGAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; or 5'-CGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UU GAAAAAAGU GGCACCGAGU CGGUGCUUUU-3', or one thereof; The pharmaceutical composition and its implementation are also described. In an embodiment, the pharmaceutical composition further comprises a vector suitable for expression in a mammalian cell. wherein the vector comprises a polynucleotide encoding a base editor. In one embodiment, the polynucleotide encoding the base editor is mRNA. In one embodiment of the composition, the vector is a viral vector. Viral vectors include retroviral vectors, adenoviral vectors, and lentiviral vectors. virus vector, herpesvirus vector, or adeno-associated virus vector (AAV In an embodiment of the pharmaceutical composition of any one of the above aspects and embodiments thereof, The pharmaceutical composition further comprises a ribonucleoparticle suitable for expression in a mammalian cell. In a pharmaceutical composition embodiment of any one of the above and its embodiments, the pharmaceutical composition comprises a lipid It further includes quality.

[0045] In another aspect, a method for treating alpha-1 antitrypsin deficiency (A1AD) is provided, the method comprising: administering the pharmaceutical composition of any one of the above aspects and embodiments thereof to a subject in need thereof. This includes:

[0046] In another aspect, the method of any of the above in treating alpha-1 antitrypsin deficiency (A1AD) in a subject. The present invention provides a use of the pharmaceutical composition of any one of the aspects and embodiments thereof.

[0047] In an embodiment of the above method or use, the subject is a human.

[0048] The fusion protein or base editor system according to any one of the above aspects and embodiments thereof. In some embodiments, the adenosine deaminase variant may be any of the following groups of modifications: or one of: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In one embodiment, an adenosine deaminase variant, e.g., a TadA*9 deaminase variant, The variant includes any modification or group of modifications set forth in Tables 14 or 18.

[0049] It will be understood by those skilled in the art in relation to the adenosine deaminase of the above aspects and embodiments thereof. As shown, other adenosine deaminase inhibitors corresponding to the amino acid modifications set forth in SEQ ID NO: 1 may be used. The amino acid modifications in the nucleotide sequence are determined by performing a standard sequence alignment and aligning the amino acid sequence of SEQ ID NO: 1. sequence and that of other adenosine deaminases, such as TadA deaminase, as described above. By assessing the relatedness and / or identity of the sequence or its related protein portion In one embodiment, the amino acid sequence of another adenosine deaminase can be readily determined. The sequence has at least 85% sequence identity to SEQ ID NO: 1. In one embodiment, another The amino acid sequence of adenosine deaminase has at least 90% sequence identity to SEQ ID NO: 1. In one embodiment, the amino acid sequence of the alternative adenosine deaminase is SEQ ID NO: 1 In one embodiment, the amino acid sequence of the present invention is a sequence having at least 95% sequence identity to another adenosine deaminase (ADD) amino acid. The amino acid sequence of the enzyme has at least 98% sequence identity to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the alternative adenosine deaminase is at least as similar to SEQ ID NO:1 as SEQ ID NO:1. Both share 99% sequence identity.

[0050] In another aspect, an adenosine deaminase or an amino acid modification or group of modifications as follows: adenosine deaminase variants, including TadA*7.10 variants containing one of adenosine deaminase, fusion protein, base editor, or base editor described above. The present invention provides computer systems and their embodiments: V82T; I76Y+V82T; or I76Y+V82T+Y 147T+Q154S.

[0051] In another embodiment, a TadA*7.10 variant comprising any one of the following amino acid modifications or groups of modifications: Provides adenosine deaminase variants: V82T; I76Y+V82T; or I 76Y+V82T+Y147T+Q154S.

[0052] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide protease. a mutatable DNA-binding domain and any one of the following amino acid modifications or groups of modifications: TadA*7.10 adenosine deaminase variants with at least one base editor domain In one embodiment of the fusion protein, the fusion protein contains: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S. In this configuration, the polynucleotide programmable DNA binding domain binds to the Cas9 endonuclease. In one embodiment of the fusion protein, the Cas9 endonuclease domain The gene contains spCas9 (MQKFRAER) with the mutations I322V, S409I, E427G, R654L, and R753G.

[0053] The above adenosine deaminase variants and embodiments thereof, or the above fusion proteins In an embodiment of the protein and its embodiments, TadA7*10 is a monomer.

[0054] In another aspect, a nucleobase editor is provided, the nucleobase editor being selected from: TadA*7.10 adenosine deaminase variant domain and Cas9 endonuclease Domains include: monoTadA*7.10 with mutation V82T and mutations I322V, S409I, E427G, R654L, and R753G spCas9(MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T and mutations I322V, S409I, E427G, R654L, and R753G spCas9 (MQKFRAER); or monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S and mutations I322V, S409I, E427G, R65 4L, spCas9 (MQKFRAER) with R753G.

[0055] definition The following definitions supplement those in the art and are intended for the present application, examples of which are provided herein. Attribution to related or unrelated cases, such as commonly owned patents or applications It is not intended that any methods or materials similar or equivalent to those described herein be used in any way without the express written permission of the present disclosure. Although any of the materials and methods described herein may be used in the practice or testing of the present invention, preferred materials and methods are Therefore, the terms used herein describe specific embodiments. For illustrative purposes only and are not intended to be limiting.

[0056] Unless otherwise defined, all technical and scientific terms used herein are defined by the principles of the present invention. The meanings of the terms have the meanings commonly understood by those skilled in the art. The reference provides one of the most important documents in the art, including general definitions of many of the terms used in this invention: ingleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glos sary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein Where used, the following terms have the meanings ascribed to them below unless otherwise specified.

[0057] In this application, the use of the singular includes the plural unless expressly stated otherwise. When used herein, the singular forms "a," "an," and "the" shall be used unless the context clearly indicates otherwise. It should be noted that the plural referents are included unless otherwise specified. The use of "or" means "and / or" unless stated otherwise. The words "including", as well as "include", "includes", and the use of other forms such as "included" is not limiting.

[0058] As used in this specification and claim(s), the word "comprising" (and any form of compr such as "comprise" and "comprises" "having" (and "have" and "has"), etc. Any form of "having," "including" (and "includes") " and "include" and any form of including or "contai ning)" (and any form of including, such as "contains" and "contain" (Containing) is inclusive or open-ended and includes additional elements or methods not cited. Any embodiment referred to herein does not exclude any method or step of the present disclosure. It is contemplated that the method may be performed on a composition or a method for preparing a drug, and vice versa. The compositions of the present disclosure can be used to achieve the methods of the present disclosure.

[0059] The term "about" or "approximately" refers to an acceptable error for a particular value as determined by one of ordinary skill in the art. It means that the difference is within the range and how the value is measured or determined, i.e., measurement For example, "about" may mean, in accordance with the convention in the art, 1 to 100% of the total amount of a substance, and may be used in part depending on the constraints of the system. Alternatively, "about" can mean within 20%, 10%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, %, or a range up to 1%. Alternatively, particularly in biological systems or processes With respect to a value, the term may mean within the same order of magnitude, e.g., within 5 times, within 2 times, etc. Unless otherwise indicated, when specific values ​​are recited in this application and claims, the term "About" is understood to mean within an acceptable range of error for that particular value. It should be.

[0060] References herein to "some embodiments," "an embodiment," "an embodiment," or "an "Other embodiments" refers to embodiments in which a particular feature, structure, or characteristic described in connection with an embodiment is different from, or is incompatible with, another embodiment. It is included in at least some embodiments, but not necessarily in all embodiments of the present disclosure. means there is no

[0061] "Adenosine deaminase" refers to the hydrolytic deamination of adenine or adenosine. In some embodiments, the term "polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing deaminase or deaminase domain converts adenosine to inosine or Adenosine catalyzes the hydrolytic deamination of deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase is a deoxyadenosine deaminase. It catalyzes the hydrolytic deamination of adenine or adenosine in ribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., modified adenosine deaminases) The adenosine deaminase (evolved adenosine deaminase) can be derived from any organism, such as a bacterium.

[0062] In some embodiments, the deaminase or deaminase domain is selected from the group consisting of human, human thymocyte, Naturally occurring materials derived from organisms such as monkeys, gorillas, monkeys, cows, dogs, rats, or mice. In some embodiments, the deaminase or deaminase For example, in some embodiments, the deaminase domain is not naturally occurring. The enzyme or deaminase domain has at least 50% or less of the amino acid sequence of the native deaminase. At least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 8 0%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, At least 98%, at least 99%, or at least 99.5% identity. In embodiments, the adenosine deaminase is selected from the group consisting of E. coli, S. aureus, S. typhi, S. putrefaciens, In some embodiments, the virus is derived from bacteria such as H. ens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. The TadA deaminase is E. coli TadA (ecTadA) deaminase or a fragment thereof.

[0063] For example, the deaminase domain may be any of those described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078 ) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is , the entire contents of which are incorporated herein by reference. Also, Komor, AC, et al., “Programming mmable editing of a target base in genomic DNA without double-stranded DNA cleavage age” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base e diting of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision repair inhibition and ba cteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017)) and Rees, HA, et a l., “Base editing: precision chemistry on the genome and transcriptome of livin g cells.” Nat Rev Genet. 2018 Dec; 19 (12): 770-788. doi: 10.1038 / s41576-018-00 See also, 59-1, the entire contents of which are incorporated herein by reference.

[0064] The wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence): ) has: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTD.

[0065] In some embodiments, the adenosine deaminase comprises the following sequence modification: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD (also known as TadA*7.10).

[0066] The present invention provides novel nucleobase editors that are modified relative to the TadA*7.10 reference sequence. It is characterized by:

[0067] In some embodiments, TadA*7.10 comprises at least one modification. In certain embodiments, TadA*7.10 comprises modifications at amino acids 82 and / or 166. , variants of the above reference sequence include one or more of the following modifications: Y147T, Y147R, Q 154S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H is the modification H123Y of TadA*7.10. In another embodiment, the barrier of the TadA*7.10 sequence is The compound has the following modifications of SEQ ID NO: 1: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72 Contains one or more of K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K In some embodiments, the variant of the TadA*7.10 sequence is selected from the group consisting of: The following combinations of alterations are included: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V8 2S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H +Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.

[0068] In other embodiments, the present invention provides a method for identifying a target gene in TadA*7.10, a TadA reference sequence, or another TadA. Starting at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157 compared to the corresponding mutations adenosine deaminase variants containing deletions, e.g., TadA*8, including C-terminal deletions provide.

[0069] In yet other embodiments, the adenosine deaminase variant is TadA*7.10, TadA reference. Each contains the following modifications compared to the reference sequence or the corresponding mutation in another TadA: Y147T, Y Two ampoules with one or more of 147R, Q154S, Y123H, V82S, T166R, and / or Q154R In another embodiment, the adenosine deaminase domain is a homodimer containing the adenosine deaminase domain. Deaminase variants are identified as TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. Compared to: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+ Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V8 2S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y1 From the group 47R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R Two adenosine deaminase domains (e.g., , TadA*8).

[0070] In other embodiments, the adenosine deaminase variant is a wild-type TadA adenosine deaminase. The amino acid sequence of TadA*7.10 was compared with the corresponding mutation in the TadA*7.10, the TadA reference sequence, or another TadA. In comparison, the following modifications were present: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or adenosine deaminase variant domains containing one or more of Q154R (e.g., Ta dA*8). In another embodiment, adenosine deaminase variants The wild-type TadA adenosine deaminase domain and TadA*7.10, the TadA reference sequence, or or compared with the corresponding mutations in other TadA: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82 S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y 123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T 166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+ adenosine deaminase variants comprising a combination of alterations selected from the group consisting of Q154R, Q154R, and Q154R; It is a heterodimer containing a domain (e.g., TadA*8).

[0071] In another embodiment, the adenosine deaminase variant comprises a TadA*7.10 domain and a T The following modifications were made compared to adA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA: i.e., one or more of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R adenosine deaminase variant domain (e.g., TadA*8), In another embodiment, the adenosine deaminase variant is a TadA*7.10 domain. The following mutations were observed compared to TadA*7.10, the TadA reference sequence, or the corresponding mutations in another TadA: Modification combinations: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y12 3H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q1 adenovirus containing 54R+I76Y; V82S+Y123H+Y147R+Q154R; or I76Y+V82S+Y123H+Y147R+Q154R It is a heterodimer containing a syndeaminase variant domain (e.g., TadA*8). In one embodiment, the adenosine deaminase is a compound having adenosine deaminase activity, such as or a fragment thereof. MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0072] In some embodiments, TadA*8 is truncated. The truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and 13 amino acids compared to the full-length TadA*8. They lack 3, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues. In this embodiment, the truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 6 It lacks 0, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues. In some embodiments, the adenosine deaminase variant is full-length TadA*8. .

[0073] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain, and comprising an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKS TN Bacillus subtilis (B.subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKN LSE Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEI LCLRSAGKKLENYRLLDATLYITLE PCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFKRRRDEKKALKLAQR AQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗ AEIIALRNGAKNIQNY RLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYK TGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKR REEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: JPEG2025170240000005.jpg29168Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD.

[0074] "Adenosine deaminase base editor 8 (ABE8) polynucleotide" means a polynucleotide encoding ABE8. It refers to a polynucleotide that encodes a

[0075] "Adenosine deaminase base editor 9 (ABE9) polypeptide" or "ABE9" , an adenosine deaminase barrier comprising one or more modifications at positions sssssss of the sequence shown below: In one embodiment, the term "base editor" refers to a base editor as defined herein, including the base (TadA*9). In the adenosine deaminase variant (TadA*9), the following modifications are included: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V in the sequence , P124W, T133K, D139L, D139M, C146R, and A158K: The relevant bases to be modified in the JPEG2025170240000006.jpg45167 reference sequence are underlined and in bold. , ABE9 contains further modifications as described herein compared to the reference sequence.

[0076] "Adenosine deaminase base editor 9 (ABE9) polynucleotide" means a polynucleotide that encodes ABE9. It means a polynucleotide that encodes

[0077] "Alpha-1 antitrypsin (A1AT) protein" refers to UniProt accession number P01009 It refers to a polypeptide or fragment thereof having at least about 95% amino acid sequence identity. In certain embodiments, the A1AT protein has one or more modifications compared to the following reference sequences: In one particular embodiment, the A1AT protein associated with A1AD comprises an E342K mutation. A typical A1AT amino acid sequence is >sp|P01009|A1AT_HUMAN α-1-antitrypsin OS=Homo sapi ens OX=9606 GN=SERPINA1 PE=1 SV=3 and has the following amino acid sequence: JPEG2025170240000007.jpg55168In this A1AT protein sequence, the first 24 amino acids constitute the signal peptide (underlined Position 342 of the sequence mutated in A1AD (i.e., E342K) is located in the signal sequence The amino acid residue "E" following the formula is set as amino acid "1".

[0078] "Administering" refers to providing one or more compositions described herein to a patient or subject. As examples, but not limited to, administration of compositions, e.g. Injections include intravenous (iv), subcutaneous (sc), intradermal (id), and intraperitoneal (ip) injections. The administration can be by intramuscular (i.m.) injection or intramuscular (i.m.) injection. Parenteral administration can be by, for example, bolus injection or gradual administration over time. Alternatively, or simultaneously, administration may be by oral route. It is possible to do so.

[0079] "Drug" means any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or or fragments thereof.

[0080] "Modification" refers to the process of modifying a protein by standard methods known in the art, such as those described herein. Changes (increases) in the sequence, expression level, or activity of a gene or polypeptide detected by As used herein, modification means a 10% change in expression levels. These include a 25% change, a 40% change, and a 50% or greater change in expression level.

[0081] "Ameliorate" means, for example, to reduce, suppress, attenuate, decrease, or inhibit the onset or progression of a disease. It means to stop or stabilize.

[0082] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide. On the one hand, they have certain biochemical modifications that enhance the function of the analogue relative to the native polypeptide. Such biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life, e.g. For example, analogs can increase ligand binding without altering ligand binding. It's okay to do that.

[0083] "Base editor (BE)" or "nucleobase editor (NBE)" refers to a polynucleobase In various embodiments, a salt refers to an agent that binds to a nucleotide and has nucleobase-modifying activity. Base editors can be used to generate nucleobase-modifying polypeptides (e.g., deaminases) and guide polypeptides. Polynucleotide programmable nucleic acids in combination with nucleotides (e.g., guide RNAs) In various embodiments, the agent comprises a nucleotide-binding domain. a protein domain that binds to bases (e.g., A, T, C) in a nucleic acid molecule (e.g., DNA). It is a biomolecular complex containing a domain that can modify a target molecule (such as a nucleotide, ... In an embodiment, the polynucleotide programmable DNA binding domain is a deaminase In one embodiment, the agent is fused or linked to a domain having base editing activity. In another embodiment, the fusion protein has base editing activity. The protein domain is linked to a guide RNA (e.g., an RNA-binding motif on the guide RNA). In some embodiments, the RNA-binding domain is fused to a ribozyme and a deaminase. , a domain with base editing activity can deaminate bases within a nucleic acid molecule. In some embodiments, a base editor deaminates one or more bases in a DNA molecule. In some embodiments, the base editor can modify a cytosine (C) or or adenosine (A). In some embodiments, the base enzyme Deter can deaminate cytosine (C) and adenosine (A) in DNA In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) or and a cytidine base editor (CBE). In some embodiments, the base editor is , a nuclease-inactive Cas9 (dCas9) fused to adenosine deaminase. In some embodiments, the Cas9 is a circular permutant Cas9 (e.g., Circularly permuted Cas9s are known in the art, for example, Oakes et al., Cell 176, 254-267, 2019. In some embodiments, The base editor can be an inhibitor of base excision repair, e.g., a UGI domain, or a dISN domain. In some embodiments, the fusion protein is fused to a deaminase. These include Cas9 nickase and inhibitors of base excision repair, such as UGI or dISN domains. In embodiments, the base editor is a base deletion base editor.

[0084] In some embodiments, the adenosine deaminase is evolved from TadA. In embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated ( In some embodiments, the base editor is a deoxyribonuclease (e.g., a Cas or Cpf1) enzyme. It is a catalytically inactive Cas9 (dCas9) fused to an aminase domain. In embodiments, the base editor comprises a Cas9 nickase (nC) fused to a deaminase domain. In some embodiments, the base editor is a polypeptide that inhibits base excision repair (BER). In some embodiments, the inhibitor of base excision repair is fused to a uracil DNA glycoprotein. In some embodiments, the inhibitor of base excision repair is a guanosine glutamate inhibitor (UGI). Base editors are inhibitors of inosine base excision repair. For more information on base editors, see International PCT Application No. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632) , each of which is incorporated herein by reference in its entirety. Also, Komor, A. C., et al., “Programmable editing of a target base in genomic DNA without double e-stranded DNA cleavage” Nature 533,420-424 (2016); Gaudelli, NM, et al., “ Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision repair r inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors wit h higher efficiency and product purity” Science Advances 3: eaao 4774 (2017), and and Rees, HA, et al., “Base editing: precision chemistry on the genome and tra nscriptome of living cells.” Nat Rev Genet. 2018 Dec; 19 (12): 770-788. doi: 1 See also 0.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference. (This is the case.)

[0085] In some embodiments, the base editor is an adenosine deaminase variant (e.g., For example, TadA*8) is combined with a circularly permuted Cas9 (e.g., spCAS9) and a bipartite nuclear localization sequence (BNS). These are generated by cloning into a backbone containing the circulating ABE8 or ABE9 vector (e.g., ABE8 or ABE9). Ring-substituted Cas9s are known in the art and are described, for example, in Oakes et al., Cell 176, 254-267, Exemplary circular permutation sequences are listed below, with the bolded sequences representing Cas9 The sequences in italics represent linker sequences, and the underlined sequences represent bypass sequences. The β-tite nuclear localization sequence is shown. CP5 (MSP) NGC = Pam variant with mutation; normal Cas9 prefers NGG PID = protein phase interacting domain and "D10A" nickase): JPEG2025170240000008.jpg184169

[0086] In some embodiments, ABE8 is selected from the base editors in Tables 10, 11, or 13 below. In some embodiments, ABE8 is an adenosine deaminase evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is a TadA*8 variant as described in Tables 8, 10, 11, or 13 below. In an embodiment, the adenosine deaminase variant is Y147T, Y147R, Q154S, Y123H, TadA*7.1 containing one or more modifications selected from the group consisting of V82S, T166R, and / or Q154R 0 variant (e.g., TadA*8). In various embodiments, ABE8 is Y147T+Q154R ; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y +V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q1 54R; and TadA* comprising a combination of modifications selected from the group of I76Y+V82S+Y123H+Y147R+Q154R 7.10 variants (e.g., TadA*8).

[0087] In some embodiments, ABE8 encodes one copy of TadA deaminase, e.g., one Ta In some embodiments, ABE8 is a monomeric construct comprising the same or a dA*8 variant. are multiple, e.g., two, different TadA deaminases, e.g., wild-type TadA and a TadA*8 variant. The construct is a dimeric or heterodimeric construct containing two copies of the ribozyme.

[0088] In some embodiments, ABE9 is selected from the base editors in Table 14 below. In some embodiments, ABE9 is an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE9 includes any of the variants listed in Table 14. In some embodiments, the adenosine deaminase variant is a TadA*7.10 variant described above. The variant is selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In various embodiments, ABE9 is TadA*7.10 containing one or more modifications as set forth in Table 14. In addition to those described above, TadA*7.10 having a modification selected from the following: Y147R+Q154 R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q15 4S; V82T+Q154S and Y123H+Y147R+Q154R+I76Y. In some embodiments, ABE9 is a TadA derivative. A monomeric construct containing one copy of the aminase, for example, the TadA*9 variant. In some embodiments, ABE9 is a mutant of the same or a different TadA deaminase, e.g., wild-type TadA. and dimeric or heterodimeric constructs containing multiple, e.g., two, copies of the TadA*9 variant. It is a thing.

[0089] In some embodiments, the ABE9 base editor comprises the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0090] Exemplary adenine salts for use in the base editing compositions, systems, and methods described herein include: The base editor ABE has the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23; 551 (7681): 464-471. doi: 10.10 38 / nature 24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct; 36 (9): 843-846. d oi: 10.1038 / nbt.4172.) having at least 95% identity to the ABE nucleic acid sequence Polynucleotide sequences are also encompassed. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGAGCGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCGCTGGCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGGAGAAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGGCCGACCTGTTTT CTGGCCGCCAAGAACCTGCTCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0091] "Base editing activity" refers to the ability to chemically change bases within a polynucleotide. In one embodiment, the first base is converted to a second base. Base editing activity refers to cytidine deaminase activity, e.g., the conversion of target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity. In another embodiment, the base editing activity is a cytogenetic or nucleotide sequence, e.g., A·T to G·C conversion. Deaminase activity, e.g., conversion of target C·G to T·A and adenosine or adenosine deaminase activity, Nine deaminase activity, for example, conversion of A·T to G·C.

[0092] The term "base editor system" refers to a system that uses nucleic acid bases to edit a target nucleotide sequence. In various embodiments, the base editing (BE) system comprises: (1) a base editing system for editing a target nucleotide; Polynucleotide programmable nucleoside for deaminating nucleic acid bases in a sequence a cytidine binding domain, a deaminase domain (e.g., a cytidine deaminase or an adenosine deaminase) (2) polynucleotide programmable nucleotide binding The nucleic acid sequence includes one or more guide polynucleotides (e.g., guide RNAs) in combination with a binding domain. In various embodiments, the base editing (BE) system is adenosine deaminase or cytochrome P450 (C1)-dependent cleavage. Nucleic acid base editor domains selected from gin deaminases, and nucleic acid sequence-specific binding In some embodiments, the base editor system comprises a domain having an activity. 1) a polynucleoside analog for deaminating one or more nucleobases in a target nucleotide sequence; A base editor containing a programmable DNA binding domain and a deaminase domain (BE); and (2) in combination with a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide programmable The programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor. In some embodiments, the base editor is an adenine or are adenosine base editors (ABEs) or cytidine base editors (CBEs).

[0093] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Active, inactive, or partially active DNA cleavage by Cas9 and / or the gRNA-binding domain of Cas9 Cas9 nuclease refers to an RNA-guided nuclease containing a nucleotide sequence (a protein containing a nucleotide sequence). , casnl nuclease or CRISPR (clustered regularly interspaced short palindrome) An exemplary Cas9 is a nuclease that encodes a nucleotide sequence encoding a nucleotide sequence encoding a repeat-associated nuclease (RCR). The gene Cas9 (spCas9) has the following amino acid sequence: JPEG2025170240000009.jpg182169

[0094] The term "Cas12b" or "Cas12b domain" refers to the Cas12b / C2c1 protein or a fragment thereof. RNA-guided nucleases containing fragments (e.g., active, inactive, or partially active Cas12b) a protein containing a DNA cleavage domain and / or a gRNA binding domain of Cas12b), The contents of each are incorporated herein by reference. Cas12b orthologs are also known as Alicyclo bacillus acidoterrestris, Alicyclobacillus acidophilus (Teng et al., Cell Discov 2018 Nov 27; 4: 63), Bacillus hisashi, and Bacillus species V3-13, Additional suitable Cas12b nucleases and sequences are available from , will be apparent to those skilled in the art based on this disclosure.

[0095] In some embodiments, a protein comprising Cas12b or a fragment thereof is a "Cas12b barrier protein." Cas12b variants are called "Cas12b variants." Cas12b variants share homology with Cas12b or fragments thereof. For example, For example, a Cas12b variant may be at least about 70% identical, or at least about 80% identical to wild-type Cas12b. 1. At least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 9 7% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or In some embodiments, the Cas12b variant is at least about 99.9% identical to the wild-type C Compared to as12b, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes In some embodiments, the Cas12b variant is a fragment of Cas12b (e.g., a gRNA-binding fragment). domain or DNA cleavage domain) whose fragments are similar to the corresponding fragments of wild-type Cas12b. At least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical , at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, the fragments are at least 30%, at least 10%, or at least 30% of the amino acid length of the corresponding wild-type Cas12b. at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 8 5%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98 %, at least 99%, or at least 99.5%. Exemplary Cas12b polypeptides are listed below. Please write in. Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|C2C1_ALIAG CRISPR - related endonuclease C2c1 from Alicyclobacillus acido - terrestris (strain ATCC49025 / DSM3922 / CIP106132 / NCIMB13137 / GD3B) from Alicyclobacillus acido - terrestris (strain ATCC49025 / DSM3922 / CIP1061 32 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRRLCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMVNQRIEGYLVKQIRSRVPLQD 2019 AacCas12b (Alicyclobacillus acidiphilus)-WP_067623834 MAVKSMKVKLRLDNMPEIRAGLWKLHTEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECYKTAEECKAELLERLRARQ VENGHCGPAGSDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGGLGIAKAGNKPRWVRMREAGEPGWE EEKAKAEARKSTDTRDADVLRALADFGLKPLMRVYTDSDMSSVQWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGE AYAKLVEQKSRFEQKNFVGQEHLVQLVNQLQQDMKEASHGLESKEQTAHYLTGRALRGSDKVFEKWEKLDPDAPFDLYDT EIKNVQRRNTRRFGSHDLFAKLAEPKYQALWREDASFLTRYAVYNSIVRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGEGRHAIRFQKLLTVEDGVAKEVDDVTVPISMSAQLDDLLPRDPHELVALYFQDYGAEQHLAGEFGGAK IQYRRDQLNHLHARRGARDVYLNLSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSEGRVPFCFPIEGNENLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPMDANQMTPDWREAFEDELQKLKSLYGICGDREWTEAVYESVRR VWRHMGKQVRDWRKDVRSGERPKIRGYQKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREHI DHAKEDRLKKLADRIIMEALGYVYALDDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELL NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCAREQNPEPFPWWLNKFVAEHKLDGCPLRADDLIPTGEGEFF VSPFSAEEGDFHQIHADLNAAQNLQRRLWSDFDISQIRLRCDWGEVDGEPVLIPRTTGKRTADSYGNKVFYTKTGVTYYE RERGKKRRKVFAQEELSEEEAELLVEADEAREKSVVLMRDPSGIINRGDWTRQKEFWSMVNQRIEGYLVKQIRSRVRLQE SACENTGDI BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515 JPEG2025170240000010.jpg148170BvCas12b V4, a variant of Cas12b, contains S893R, K846R, and and E837G changes. BvCas12b (Bacillus sp. V3-13) NCBI reference sequence: WP_101661451.1 MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEH GSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEER KAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTE SYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESAPEELWKVVAEQQNKMS EGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEK QKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIK NHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLLEG MRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQ IRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQI VSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIM TALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEY SSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVI HADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPK SQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTD TTFDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL

[0096] The term "conservative amino acid substitution" or "conservative mutation" refers to a substitution of an amino acid with a common characteristic. It refers to the substitution of an amino acid with another amino acid that has a specific characteristic. A functional method is to calculate the normalized frequency of amino acid changes between corresponding proteins of homologous organisms. (Schulz, GE and Schirmer, RH, Principles of Protein S structure, Springer-Verlag, New York (1979). According to such an analysis, the group Amino acids within a group are preferentially exchanged with each other, thus affecting the overall protein structure. groups of amino acids can be defined where the amino acids are most similar to each other (Schulz , GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include those at the amino acid substitution substitutions, for example, arginine to lysine, which allows the positive charge to be maintained, and and vice versa; aspartic acid to glutamic acid so that the negative charge can be maintained. and vice versa; threonine to serine so as to maintain a free -OH; and free -NH2 Examples include substitution of asparagine with glutamine, which allows the protein to be maintained at a constant level.

[0097] The terms "coding sequence" or "protein-coding sequence" are used interchangeably herein. A coding sequence refers to a segment of a polynucleotide that encodes a protein. The region or sequence may also be called the open reading frame. codons at the 3' end and a stop codon near the 3' end. Stop codons useful in base editors include: JPEG2025170240000011.jpg38170

[0098] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "cytididinyl" refers to a polypeptide or fragment thereof that is capable of catalyzing a cytididinyl polypeptide. Deaminase converts cytosine to uracil or 5-methylcytosine to thymine Petromyzon marinus (Petromyzon marinus cytosine deaminase 1, "PmCDA1") PmCDA1 derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and AID ( Activation-induced cytidine deaminase (AICDA), and APOBEC are exemplary cytidine deaminases. It is Ze.

[0099] As used herein, the term "deaminase" or "deaminase domain" refers to a deaminase domain that It refers to a protein or enzyme that catalyzes an amination reaction. The enzyme or deaminase domain converts cytidine or deoxycytidine, respectively. cytidine deamination, which catalyzes the hydrolytic deamination to uridine or deoxyuridine In some embodiments, the deaminase or deaminase domain is Cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil In some embodiments, the deaminase hydrolyzes adenine to hypoxanthine. In some embodiments, the enzyme is an adenosine deaminase that catalyzes the selective deamination of adenosine. Aminosine deamination hydrolytically deaminates adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or The deaminase domain converts adenosine or deoxyadenosine to iodophosphate, respectively. Adenosine deamina catalyzes the hydrolytic deamination to inosine or deoxyinosine In some embodiments, the adenosine deaminase is a deoxyribonucleic acid ( The adenosine provided herein catalyzes the hydrolytic deamination of adenosine in DNA. adenosine deaminase (e.g., engineered adenosine deaminase, evolved adenosine deaminase) The enzyme (deaminase) can be derived from any organism, such as a bacterium. Adenosine deaminase is a ubiquitous enzyme found in E. coli, S. aureus, S. typhi, S. putrefaciens, and H. influenzae. In some embodiments, the adenovirus is derived from a bacterium such as C. nzae, or C. crescentus. In some embodiments, the deaminase is a TadA deaminase. The deaminase domain is expressed in human, chimpanzee, gorilla, monkey, cow, dog, and rat. or variants of naturally occurring deaminases from organisms such as mice. In embodiments, the deaminase or deaminase domain is not naturally occurring. For example, in some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase. minase and at least 50%, at least 55%, at least 60%, at least 65%, at least At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91% , at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7 %, at least 99.8%, or at least 99.9% identical.

[0100] "Detection" refers to identifying the presence, absence, or amount of an analyte being detected. In an embodiment, alterations in the sequence of a polynucleotide or polypeptide are detected. In this case, the presence of indels is detected.

[0101] A "detectable label" means a label that, when attached to a molecule of interest, is capable of detecting a given molecule spectroscopically, photochemically, or biochemically. "detectable" refers to a composition that allows a molecule of interest to be detected through immunological, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloids, Particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., in enzyme-linked immunosorbent assays (ELISAs)) Commonly used enzymes), biotin, digoxigenin, or haptens.

[0102] "Disease" means a pathological condition or condition that damages or interferes with the normal function of a cell, tissue, or organ. It means disability.

[0103] An "effective amount" is an amount that is comparable to an untreated patient or an individual without the disease, i.e., a healthy individual. In comparison, drugs or active compounds required to ameliorate the symptoms of a disease, e.g., an amount of a base editor described herein or agent sufficient to elicit a desired biological response. or the amount of active compound. The effective amount of active compound(s) used will depend on the mode of administration, the age, weight, and general health of the subject. Ultimately, your doctor or veterinarian will determine the appropriate dosage and dosing regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is sufficient to introduce a change into a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is an amount of a base editor of the present invention sufficient to achieve a therapeutic effect. Such therapeutic effects may be observed in a subject, tissue, or It is not necessary that the amount of the gene required to alter the pathogenic gene in every cell of a target, Approximately 1%, 5%, 10%, 25%, 50%, 75% or more of the intracellular pathogens present in a tissue or organ In one embodiment, the effective amount is sufficient to alter the gene. is sufficient to improve the symptoms of

[0104] In some embodiments, the fusion proteins provided herein, e.g., nCas9 domains, adenosine deaminase and deaminase domains (e.g., adenosine deaminase, cytidine deaminase An effective amount of a nucleobase editor, including a nucleobase editor, is a compound that is particularly useful for treating nucleobase disorders, such as rheumatoid arthritis, ... and rheumatoid arthritis. It refers to an amount of a target site that is heterologously bound and edited that is sufficient to induce editing. As will be appreciated, an effective amount of an agent, e.g., a fusion protein, can be determined, for example, by the amount of the agent that is effective to produce a desired product. biological response, e.g., the specific allele, genome, or target site to be edited, target may vary depending on various factors, such as the cells or tissues involved and / or the drugs used. .

[0105] In some embodiments, the fusion proteins provided herein, e.g., nCas9 domains, An effective amount of a fusion protein comprising a deaminase and deaminase domain is an amount of fusion protein sufficient to induce editing of the target site to which it specifically binds and edits; As will be appreciated by those skilled in the art, agents such as fusion proteins, nucleic acids, enzymes, hybrid proteins, protein dimers, proteins (or protein dimers) An effective amount of a complex of a dimer and a polynucleotide, or a polynucleotide, can be, for example, the biological response of, for example, the particular allele, genome or target site to be edited; It varies depending on various factors such as the target cell or tissue and / or the drug used. obtain.

[0106] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. or at least about 10%, 20%, 30%, 40%, 50% of the entire length of the reference nucleic acid molecule or polypeptide. , 60%, 70%, 80%, or 90%. Fragments may contain 10, 20, 30, 40, 50, 60, 70, 80, 90 , or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides may contain amino acids.

[0107] "Guide RNA" or "gRNA" refers to a polynucleotide promoter specific for a target sequence. Complex with a gram-competent nucleotide-binding domain protein (e.g., Cas9 or Cpf1) In one embodiment, the guide polynucleotide The primer is a guide RNA (gRNA). gRNAs can be present as a complex of two or more RNAs or as a single A gRNA that exists as a single RNA molecule is called a single guide RNA (s Although sometimes referred to as gRNA, "gRNA" can be used as a single molecule or as a set of two or more molecules. Used interchangeably to refer to guide RNAs that exist as a complex; typically a single RNA species gRNAs, which exist as a nucleotide sequence, contain two domains: (1) a domain that shares homology with the target nucleic acid (e.g., (1) a domain that directs the binding of the Cas9 complex to a target; and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) comprises a domain that binds to the tracrRNA gene. For example, in some embodiments, the domain In (2), as described in Jinek et al., Science 337: 816-821 (2012), It is identical to or homologous to acrRNA, the entire contents of which are incorporated herein by reference. Other examples (e.g., those containing domain 2) are described in "Switchable Cas9 Nucleases and Uses" Thereof,” and “Delivery System For Functional Nucleases ", the entire contents of each of which are incorporated by reference in their entirety. In some embodiments, the gRNA comprises two or more domains (1) and (2) and may be referred to as an "extended gRNA." Extended gRNAs are those described herein. The Cas9 protein binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions, so that The gRNA contains a nucleotide sequence complementary to the target site, and this sequence is used to target the target site. mediates binding of the nuclease / RNA complex and provides sequence specificity for the nuclease:RNA complex .

[0108] A "heterodimer" refers to a heterodimer of a wild-type TadA domain and a variant of the TadA domain (e.g., TadA*8 or TadA*9) or two variant TadA domains (e.g., TadA*7.10 and TadA *8 or two TadA*8 domains; or TadA*7.10 and TadA*9 or two TadA*9 domains The term "fusion protein" refers to a fusion protein containing two domains, such as a fusion protein (e.g., a fusion protein with a nucleotide sequence) and a fusion protein with a nucleotide sequence (e.g., a fusion protein with a nucleotide sequence).

[0109] "Hybridization" means hydrogen bonding, which means that complementary nuclei Acid-base Watson-Crick, Hoogsteen, or reverse Hoogsteen For example, adenine and thymine pair through the formation of hydrogen bonds. are complementary nucleobases.

[0110] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0111] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or their grammatical equivalents The equivalents are proteins capable of inhibiting the activity of nucleic acid repair enzymes, e.g., base excision repair enzymes. In some embodiments, IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, and h These include inhibitors of OGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In embodiments, the IBR is an inhibitor of Endo V or hAAG. , catalytically inactive EndoV, or catalytically inactive hAAG. The base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI is a protein that can inhibit the base excision repair enzyme uracil DNA glycosylase. In some embodiments, the UGI domain is a wild-type UGI or a wild-type In some embodiments, the UGI proteins provided herein include fragments of: Fragments of UGI and proteins homologous to UGI or UGI fragments are included. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. base repair inhibitors are known to inhibit "catalytically inactive inosine-specific nucleases" or "inactive The term "inosine-specific nuclease" is used to refer to a specific inosine-specific nuclease. However, catalytically inactive inosine glycosylases (e.g., alkyladenine glycosylases) ABA (AAG) can bind to inosine, but does not create an abasic site or convert inosine to ATP. As a result, the newly formed inosine moiety cannot be removed or eliminated from DNA damage / repair. In some embodiments, catalytically inactive inosine is sterically blocked from the cleavage mechanism. A specific nuclease can bind to inosine in nucleic acid but does not cleave the nucleic acid. Non-limiting exemplary catalytically inactive inosine-specific nucleases include, for example, human a catalytically inactive alkyl adenosine glycosylase (AAG nuclease) derived from For example, catalytically inactive endonuclease V (EndoV nuclease) from E. coli In some embodiments, the catalytically inactive AAG nuclease has an E125Q mutation. or a corresponding mutation in another AAG nuclease.

[0112] "Inteins" are proteins that bind to each other in a process called protein splicing. Proteins that can be excised and the remaining fragments (exteins) joined by peptide bonds Inteins are also called "protein introns." The process of excising itself and attaching the remainder of the protein is referred to herein as "tapping." This process is called "protein splicing" or "intein-mediated protein splicing." In some embodiments, the intein of the precursor protein (intein-mediated protein The intein-containing protein (pre-spliced ​​intein-containing protein) is derived from two genes. Such inteins are referred to herein as split inteins (e.g., split inteins). Intein-N and split intein-C), for example, in cyanobacteria, DnaE, ​​the catalytic subunit of DNA polymerase III, is expressed by two separate genes, dnaE-n and The intein encoded by the dnaE-c gene is the intein encoded by the dnaE-n gene. Intein N is an intein encoded by the dnaE-c gene. may be referred to herein as "intein-C."

[0113] Other intein systems can also be used, for example, dnaE intein-based syntheses. Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C). Intein pairs (intein-C) have been described (see, e.g., (See Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138 (7): 2162-5). Non-limiting examples of intein pairs that can be used include: Cfa DnaE intein, Spp GyrB Intein, Spp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma Dna B intein and Cne Prp8 intein (see, e.g., U.S. Pat. No. 6,223,629, incorporated herein by reference). (described in U.S. Pat. No. 8,394,604).

[0114] 1 shows exemplary nucleotide and amino acid sequences of inteins.

[0115] DnaE intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGA ATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGG AAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAG ATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDG QMLPIDEIFERELDLMRVDNL PN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTG CTCTGAAGAACGGATTCATAG CTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGA ATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAG AAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAG ATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCT TGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0116] To join the N-terminal part of split Cas9 and the C-terminal part of split Cas9, Intein-N and Intein-C are the N-terminal part of split-Cas9 and split-Cas, respectively. 9. For example, in some embodiments, intein-N may be fused to the C-terminal portion of The N-terminal portion of split Cas9 was fused to the C-terminus, i.e., N--[N-terminus of split Cas9]. In some embodiments, the intein-C is formed by the following structure: was fused to the N-terminus of the C-terminal portion of split Cas9, i.e., N-[intein-C]-[ssp The C-terminal portion of split Cas9 is then fused to the protein ( For example, intein-mediated protein splicing to bind split Cas9 The binding mechanism is described, for example, in Shah et al., Chem Sci. 2014; 5, incorporated herein by reference. Inteins are well known in the art, as described in (1): 446-461. Methods for using such compounds are known in the art, for example, WO2014004336, WO2017132580 , US20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety. is incorporated herein by reference.

[0117] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its natural state. Substances that are removed to varying degrees from the constituents that normally accompany them as recognized "Isolate" refers to the degree of separation from the original source or surroundings. "Purified" or "biologically pure" protein refers to a degree of separation that is greater than isolation. The protein is purified to ensure that impurities do not materially affect the biological properties of the protein or cause other adverse effects. Other substances are sufficiently removed that they do not cause any adverse reactions. Nucleic acids or peptides may be produced by recombinant DNA technology, or may be derived from cellular material, viral material, or other substantially free of substances or culture medium, or, if chemically synthesized, chemical precursors A substance is said to be purified if it is substantially free of trace elements or other chemicals. Purity and homogeneity are usually , analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" means that a nucleic acid or protein is purified by electrophoresis. It can be shown that the phospholipase C12 gives rise to essentially one band in a running gel. In the case of proteins that can be subject to modifications such as cosylation, the different modifications can be purified separately. This can result in different isolated proteins that can be isolated.

[0118] An "isolated polynucleotide" is a polynucleotide that is isolated from the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. It refers to nucleic acid (e.g., DNA) that does not contain genes adjacent to the gene. Thus, the term includes, for example, recombinant DNA that is incorporated into a vector; recombinant DNA incorporated into a plasmid or virus; or into a prokaryotic or eukaryotic organism Recombinant DNA that is integrated into genomic DNA; or a separate molecule independent of other sequences (e.g. cDNA or genomic fragments generated by PCR or restriction endonuclease digestion In addition, the term includes recombinant DNA molecules that exist as fragments of DNA or cDNA fragments. and a hybrid gene encoding an additional polypeptide sequence. This includes recombinant DNA, which is a part of the

[0119] An "isolated polypeptide" is a polypeptide of the present invention that has been separated from components that naturally accompany it. Generally, a polypeptide refers to a polypeptide that is naturally associated with A substance is considered isolated when it is at least 60% by weight free of proteins and naturally occurring organic molecules. Preferably, the formulation comprises at least 75% by weight, more preferably at least 90% by weight, and most preferably Preferably, the isolated polypeptide of the present invention is at least 99% by weight of the polypeptide of the present invention. The polypeptides may be prepared by, for example, extracting them from natural sources. by expression of a recombinant nucleic acid encoding the protein; or by chemically synthesizing the protein. The purity can be determined by any suitable method, e.g., column chromatography, polyacrylamide gel electrophoresis, or the like. This can be measured by amide gel electrophoresis or HPLC analysis.

[0120] As used herein, the term "linker" refers to a molecule or moiety that connects two molecules or moieties, e.g., a protein. Two components of a protein or ribonucleoconjugate, or two dopants of a fusion protein a main, e.g., a polynucleotide programmable DNA binding domain (e.g., dCas9) and deaminase domains (e.g., adenosine deaminase, cytidine deaminase) or adenosine deaminase and cytidine deaminase) It can refer to a linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule. The linkers are used to connect various components or different parts of the components of the base editor system. For example, in some embodiments, the linker can be a polynucleotide program Guide to possible nucleotide binding domains Polynucleotide binding domains and deaminases In some embodiments, the linker can link the catalytic domains of CRISPR The polypeptide and the deaminase can be linked. In some embodiments, the linker In some embodiments, the linker can link Cas9 and the deaminase. can link dCas9 and the deaminase. In some embodiments, the linker can link nCas9 and the deaminase. In some embodiments, the linker can bind the guide polynucleotide and the deaminase. In this embodiment, the linker connects the deamination component of the base editor system to the polynucleotide protease. In some embodiments, the nucleotide-binding moiety can be grammable. The linker binds the RNA-binding portion of the deaminating component of the base editor system to the polynucleotide. In some embodiments, the programmable nucleotide binding component can bind to the target molecule. In this study, the linker acts as a bridge between the RNA-binding moiety of the deaminating component of the base editor system and the polynucleotide Able to bind the RNA-binding moiety of a nucleotide-programmable binding component A linker is a molecule that is located between or adjacent to two groups, molecules, or other moieties and connect to each other through covalent or non-covalent interactions, thus connecting the two In some embodiments, the linker can be an organic molecule, group, polymer, or chemical In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker may be a DNA linker. In some embodiments, the linker may be an RNA linker. In some embodiments, the ligand may comprise an aptamer that can bind to the ligand. The peptide may be a carbohydrate, a peptide, a protein, or a nucleic acid. Alternatively, the linker can include an aptamer, which can be derived from a riboswitch. The riboswitches that are used are theophylline riboswitch and thiamine pyrophosphate (TPP) riboswitch. Chi, adenosine cobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetranucleotide Hydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch riboswitch, GlmS riboswitch, or prequosin 1 (PreQ1) riboswitch. In some embodiments, the linker is a polypeptide or a polypeptide ligand. In some embodiments, the polynucleotide may comprise an aptamer bound to a protein domain. Peptide ligands include the K homology (KH) domain, the MS2 coat protein domain, and the PP7 Coat protein domain, SfMu Com coat protein domain, Sterile α motif , telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 In some embodiments, the polypeptide may be a protein or RNA recognition motif. The ligand can be part of a base editor system component. For example, a nucleic acid base editing component The molecule may comprise a deaminase domain and an RNA recognition motif.

[0121] In some embodiments, the linker may be one amino acid or multiple amino acids (e.g., In some embodiments, the linker may be of length about 5 to 100 amino acids, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 , 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 In some embodiments, the linker is about 100 to 150 amino acids in length. 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids Longer or shorter linkers may also be used. Shorter linkers are also contemplated. In some embodiments, the linker is It comprises the sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. SGGS)n, (GGGS)n, (GGGGS)n, (G)n, (EAAAK)n, (GGS)n, SGSETPGTSESATPES, or (XP)n motifs, or any combination thereof, wherein n is are independently an integer from 1 to 30, and X is any amino acid. are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In embodiments, the linker comprises multiple proline residues and is 5 to 21, 5 to 14, 5 to 9, 5 to 7, or Amino acids, such as PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, P (AP) 10 Such proline-rich linkers are also called "rigid" linkers. can be.

[0122] In some embodiments, the linker comprises an RNA promoter comprising a Cas9 nuclease domain. The gRNA-binding domain of a tunable nuclease and a nucleic acid editing protein (e.g., cytidine The catalytic domain of the enzyme (deaminase or adenosine deaminase) binds to the In embodiments, a linker connects the dCas9 and the nucleic acid editing protein. For example, the linker is Located between or adjacent to two groups, molecules, or other moieties and connected via a covalent bond In some embodiments, the linker , a single amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is between 5 and 200 amino acids in length, e.g., 5, 6, 7, or ,8,9,10,11,12,13,14,15,16,17,18,19,20,25,35,45,50,55,60,60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130 , 140, 150, 160, 175, 180, 190, or 200 amino acids.

[0123] In some embodiments, the base editor domain is selected from the group consisting of SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEG Amino acid sequence of SAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS In some embodiments, the domain of the base editor is fused via a linker comprising The XTEN linker is a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. The anchor contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In embodiments, the linker is 64 amino acids in length. The amino acid sequence is SGGSSGGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS S In some embodiments, the linker is 92 amino acids in length. In an embodiment, the linker has the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSP Contains AGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0124] A "marker" is any molecule whose expression level or activity is altered in association with a disease or disorder. It means a protein or a polynucleotide.

[0125] As used herein, the term "mutation" refers to a change in a sequence (e.g., a nucleic acid or amino acid sequence). Refers to the substitution of a residue by another residue, or the deletion or insertion of one or more residues within a sequence Mutations usually identify the original residue, followed by the position of the residue within the sequence, and then The newly substituted residues are described herein by identifying the residues. Various methods for making the amino acid substitutions (mutations) provided in , for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., C Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) In some embodiments, the base editors of the present disclosure reduce a significant number of unintended mutations. , e.g., to modify a nucleic acid (e.g., a nucleic acid in a subject's genome) without generating unintended point mutations. ) can be efficiently used to generate "intended mutations" such as point mutations. In embodiments, the intended mutation is a mutation specifically designed to produce the intended mutation. A specific base editor (e.g., a sequence encoding a specific base editor) bound to a guide polynucleotide (e.g., a gRNA) Mutations generated by nucleotide-binding proteins (nucleotide editors or adenosine base editors) .

[0126] Generally, created or identified within a sequence (e.g., an amino acid sequence described herein) Mutations are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. Those skilled in the art will be able to determine the location of amino acid and nucleic acid sequence variations relative to a reference sequence. You will easily understand how to do this.

[0127] The term "non-conservative mutation" includes, for example, a tryptophan to lysine, a serine to phenyl This includes amino acid substitutions between different groups, such as alanine. Preferably, the acid substitutions do not interfere with or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions may enhance the biological activity of functional variants, resulting in The biological activity of the functional variant is increased compared to the wild-type protein.

[0128] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to a sequence that directs a protein to the cell nucleus. Nuclear localization sequences refer to amino acid sequences that promote protein import. Nuclear localization sequences are known in the art and include: For example, P.I., filed on November 23, 2000, and published on May 31, 2001 as WO / 2001 / 038547, Lank et al., International PCT Application No. PCT / EP2000 / 011690, the contents of which are incorporated by reference in their entirety. The present invention is incorporated by reference herein for disclosure of nuclear localization sequences. S is described, for example, by Koblan et al., Nature Biotech. 2018 doi: 10.1038 / nbt.4172 In some embodiments, the NLS is an optimized NLS having the amino acid sequence KRTADG SEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAA Contains IVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0129] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. refers to nitrogen-containing biological compounds that form nucleosides, the building blocks of nucleotides. The ability of acids and bases to form base pairs and stack with each other is what makes ribonucleic acid (RNA) and deoxyribonucleic acid (DENA) so unique. Directly generates long chain helical structures such as DNA. Five nucleic acid bases (adenine (A) , cytosine (C), guanine (G), thymine (T), and uracil (U) are primary nucleobases or standard nucleobases. Adenine and guanine are derived from purines, while cytosine, uranium Cysteine ​​and thymine are derived from pyrimidines. DNA and RNA contain modified other (non-primary) Non-limiting exemplary modified nucleobases include hypoxanthine, chymotrypsin, hydroxybenzoates, and hydroxybenzoates. Santhin, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5 Hypoxanthine and xanthine may be present in the presence of mutagens. Both are produced by deamination (replacement of an amine group with a carbonyl group). Hypoxanthine can be produced from adenine. Xanthine can be produced from It can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" is a nucleic acid consisting of a nucleic acid base and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine Nucleosides with modified nucleobases include xylidine and deoxycytidine. Examples of methylguanosine include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrogen These include uridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" is a nucleic acid consisting of a nucleic acid base, a five-carbon sugar (either ribose or deoxyribose), or ), and at least one phosphate group.

[0130] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to nucleic acids containing nucleic acid bases and acidic moieties. The term "nucleoside" refers to a compound containing a nucleotide, such as a nucleoside, a nucleotide, or a polymer of nucleotides. Generally, large nucleic acids, e.g., nucleic acid molecules containing three or more nucleotides, are linear molecules. In this case, adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleic acids). In some embodiments, a "nucleic acid" refers to three or more individual nucleosides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing tetroxide residues. "Nucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., at least and nucleotide sequences (strings of three nucleotides) can be used interchangeably. In this embodiment, "nucleic acid" encompasses RNA as well as single- and / or double-stranded DNA. Acids include, for example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cos It may occur naturally in the context of a sequence, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules include non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, modified It may be a whole genome, or a fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid. or non-natural nucleotides or nucleosides. "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., phosphodiesterase Nucleic acids can be purified from natural sources and can be synthesized using a variety of methods. They can be produced using recombinant expression systems, for example, and optionally purified or chemically synthesized. If necessary, for example, in the case of chemically synthesized molecules, the nucleic acid may be chemically modified. Nucleoside analogs, such as analogs with modified bases or sugars, and backbone modifications may also be included. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise specified. In general, nucleic acids are made up of natural nucleosides (e.g., adenosine, thymidine, guanosine, cysteine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxythiazol-1-yl oxycytidine); nucleoside analogues (e.g., 2-aminoadenosine, 2-thiothymidine) , inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-amino Denosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propionyl leu-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7- Deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O( 6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified modified bases (e.g., methylated bases); intercalated bases; modified Sugars (2'-e.g., fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5 '-N-phosphoramidite linkages).

[0131] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to a It is sometimes used interchangeably with "nucleotide-programmable nucleotide-binding domain." and a guide nucleic acid or guide polynucleotide (e.g., a nucleotide sequence) that guides the napDNAbp to a specific nucleic acid sequence. It refers to a protein that binds to a nucleic acid (e.g., DNA or RNA) such as a gRNA. In some embodiments, the polynucleotide-programmable nucleotide binding domain is In some embodiments, the polynucleotide is a programmable DNA binding domain. A programmable nucleotide binding domain is a polynucleotide programmable In some embodiments, the polynucleotide programmable RNA binding domain A possible nucleotide binding domain is the Cas9 protein. The Cas9 protein binds the guide RN A guide RNA can bind to a specific DNA sequence complementary to A, which guides the Cas9 protein. In some embodiments, the napDNAbp can contain a Cas9 domain, e.g., a nuclease activity Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of programmable DNA binding proteins include Cas9 (e.g., dCas9 and nC as9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Non-limiting examples of Cas enzymes include Cas1, Cas12h, Cas12i, and Cas12j / CasΦ. , Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, C as8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas1 2a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas 12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cm r5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S , Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, V Type I Cas effector proteins, CARF, DinG, homologs thereof, or modifications or alterations thereof Other nucleic acid programmable DNA binding proteins may also be used in the present disclosure. , which may not be specifically described in this disclosure. , Makarova et al. “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J.2018 Oct; 1:325-336. doi: 10.1089 / crispr.2018.0033; Yan e t al., “Functionally diverse type V CRISPR-Cas systems” Science.2019 Jan 4; 36 3 (6422): 88-91. See doi:10.1126 / science.aav7271 (the full contents of each are available here). (Incorporated herein by reference).

[0132] As used herein, the terms "nucleobase-editing domain" or "nucleobase-editing protein" refer to " refers to a nucleobase modification in RNA or DNA, e.g., a cytosine (or cytidine) to to uracil (or uridine) or thymine (or thymidine), and to adenine (or adenosine) to hypoxanthine (or inosine), and template-free deamination refers to a protein or enzyme that can catalyze the addition and insertion of nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenosine deaminase). adenosine deaminase or adenosine deaminase; or cytidine deaminase or cytidine In some embodiments, the nucleobase editing domain is a nucleobase-editing domain comprising multiple a deaminase domain (e.g., adenine deaminase or adenosine deaminase) In some embodiments, the enzyme is a cytosine deaminase or a cytidine deaminase. In some embodiments, the nucleobase-editing domain may be a naturally occurring nucleobase-editing domain. In this embodiment, the nucleobase-editing domain is modified from a naturally occurring nucleobase-editing domain. The nucleobase editing domain may be derived from a bacterial, Any living organism, such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse It can come from things.

[0133] As used herein, the term "obtaining" in "obtaining a drug" includes synthesizing, purchasing, etc. or obtained by other means.

[0134] As used herein, a "patient" or "subject" refers to a person who has a disease or disorder or mammals at risk of developing the disease or diagnosed as being affected or suspected of developing the disease In some embodiments, the term "patient" refers to a subject or individual suffering from a disease or disorder. An exemplary patient is a mammalian subject who has a higher than average chance of developing the disease disclosed herein. Humans, non-human primates, cats, dogs, pigs, cattle, and rats that can benefit from this therapy Dogs, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs) and other mammals. An exemplary human patient is a male and / or It can be a woman.

[0135] A "patient in need thereof" or a "subject in need thereof" as used herein means a person who is suffering from a disease. have, are at risk for, or have been diagnosed with or have a disease or disorder This refers to patients who have been predetermined to have or are suspected of having the condition.

[0136] The terms "pathogenic mutation", "pathogenic variant", "pathogenic mutation", "pathogenic variant", " A "deleterious mutation" or "predisposing mutation" is a mutation that increases an individual's susceptibility to a particular disease or disorder. In some embodiments, a pathogenic variant refers to a genetic change or mutation that increases a person's risk of developing a disease or predisposition to a disease. A sexual mutation results in at least one wild-type amino acid sequence in the protein encoded by the gene. In some cases, the nucleotide sequence of ...

[0137] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and are linked together by a peptide (amide) bond. These terms refer to polymers of amino acid residues of any size, structure, or function. Refers to a protein, peptide, or polypeptide. A polypeptide is at least three amino acids in length. A protein, peptide, or A polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a peptide, or polypeptide, may be linked to a chemical entity, e.g., a carbohydrate group. , hydroxyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, bond, It can be modified by adding a linker or the like for functionalization or other modifications. Proteins, peptides, or polypeptides may also be single molecules or multi-molecular complexes. A protein, peptide, or polypeptide may be a naturally occurring protein or A protein, peptide, or polypeptide may be a simple fragment of a peptide. It may be natural, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a protein derived from at least two different proteins. A hybrid polypeptide containing a protein domain. Located at the amino-terminal (N-terminal) portion of a protein or at the carboxy-terminal (C-terminal) protein and therefore amino-terminal or carboxy-terminal fusion proteins, respectively. Proteins can be made up of different domains, e.g., nucleic acid binding domains. a domain (e.g., the gRNA binding domain of Cas9 that directs the protein to bind to the target site) and a nucleus. The nucleic acid editing protein may comprise an acid cleavage domain or a catalytic domain of a nucleic acid editing protein. In this form, the protein comprises a proteinaceous portion, e.g., an amino acid that constitutes a nucleic acid binding domain. and organic compounds, e.g., compounds that can act as nucleic acid cleaving agents. In some embodiments, the protein forms a complex with a nucleic acid, e.g., RNA or DNA. Any of the proteins provided herein are can also be produced by any method known in the art. The proteins provided in can be produced via recombinant protein expression and purification. This is particularly suitable for fusion proteins containing peptide linkers. Methods for expression and purification of .beta.-LysRNA are well known and are described in Green and Sambrook, Molecular Cloning: A Labora atory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference. will be done.

[0138] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) The amino acids (including the amino acids) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Synthetic amino acids such as acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylaminomethyl- Stain, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine , 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthyl Cyclohexylalanine, cyclohexylglycine, indoline-2-carbo carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid Phosphonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxybenzoic acid monoamide, hydroxylidine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexyl Xanthanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane )-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine Polypeptides and proteins include polypeptides, This may involve post-translational modification of one or more amino acids of the peptide construct. Typical examples include phosphorylation, acylation including acetylation and formylation, glycosylation ( including N-linked and O-linked), amidation, hydroxylation, methylation and ethylation alkylation, ubiquitination, pyrrolidone carboxylic acid addition, disulfide bridge formation, sulfation ation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, These include glypiation, lipoylation and iodination.

[0139] The term "recombinant" as used herein with respect to a protein or nucleic acid means a protein or nucleic acid that is not naturally occurring It refers to proteins or nucleic acids that do not exist but are the product of human engineering. For example, some In embodiments, the recombinant protein or nucleic acid molecule has at least one amino acid sequence that is at least as long as any naturally occurring sequence. at least one, at least two, at least three, at least four, at least five, at least The amino acid or nucleotide sequence may contain at least six or at least seven mutations.

[0140] By "reduction" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0141] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type or In other embodiments, the reference is, but is not limited to, a cell that has been subjected to a test condition. No or placebo or normal saline, medium, buffer, and / or Untreated cells that are subjected to a control vector that does not carry the polynucleotide.

[0142] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A subset or the entirety of a specified sequence; e.g., a segment of a full-length cDNA or gene sequence In the case of a polypeptide, the reference polypeptide may be a sequence of a polypeptide or a complete cDNA or gene sequence. The length of the polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids. , at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, and at least Each of the sequences is about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides. nucleotide, or any integer thereabout or therebetween. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. The columns are polynucleotide sequences encoding the wild-type proteins.

[0143] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to a nuclease that Used with (e.g., bind to or associate with) one or more non-target RNA(s) In some embodiments, the RNA-programmable nuclease is in a complex with RNA. When the nuclease is bound to the RNA, it can be called a nuclease:RNA complex. Usually, the bound RNA is a guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is referred to as (CRISPR Cas9 endonuclease, e.g., Cas9 (Csn1) from Streptococcus pyogenes (For example, "Complete genome sequence of an M1 strain of Streptococcus pyog Ferretti JJ et al., Proc. Natl. Acad. Sci. USA 98: 4658-4663 (2001 ); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. ” See Deltcheva E. et al., Nature 471: 602-607 (2011).

[0144] RNA-programmable nucleases (e.g., Cas9) are capable of catalyzing RNA:DNA hybridization. These proteins, in principle, utilize guided translation to target DNA cleavage sites. Any sequence specified in the RNA can be targeted. Use RNA-programmable nucleases such as Cas9 to modify the genome Methods for this purpose are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering ineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al, RNA-guided human genome engineering via Cas9. Science 339,823-826(2013); Hwang, WY et al, Efficient genome editing in zebrafish using a CRISPR-Cas system. Nat ure biotechnology 31, 227-229 (2013); Jinek, M. et al, RNA-programmed genome edi ting in human cells. eLife 2, e00471(2013); Dicarlo, JE et al, Genome enginee ring in Saccharomyces cerevisiae using CRISPR-Cas systems. h (2013); Jiang, W. et al RNA-guided editing of bacterial genomes using CRISPR-C as systems. Nature biotechnology 31, 233-239 (2013) (see the full text of each). The contents of which are incorporated herein by reference).

[0145] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific location in the genome. association, and each variation is present to a significant degree (e.g., >1%) in the population. For example, at certain base positions in the human genome, C nucleotides occur in many individuals. However, in a small number of individuals, the position is occupied by A. This is because there is a SNP at this particular position. This means that there are two possible nucleotide variations, C or A. The SNPs underlie differences in susceptibility to disease. The severity of the disease and how the body responds to treatment are also manifestations of genetic variation. can be found in the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes In some embodiments, SNPs within a coding sequence may be present due to the degeneracy of the genetic code. Therefore, it does not necessarily change the amino acid sequence of the protein produced. There are two types of SNPs: synonymous and non-synonymous. Synonymous SNPs do not affect the protein sequence, but Nonsynonymous SNPs alter the amino acid sequence of a protein. Nonsynonymous SNPs include missense and nonsense SNPs. There are two types of SNPs: SNPs that are not in protein-coding regions and SNPs that are gene splices. affect isolating, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNAs. Gene expression affected by this type of SNP is called an eSNP (expressed SNP), and It can be located upstream or downstream of the gene. Single nucleotide variants (SNVs) are frequency-limited. There is an unlimited number of single nucleotide variations that can occur in somatic cells. Nucleotide variations are also called single nucleotide modifications.

[0146] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. a nucleic acid that binds to a target molecule but does not substantially recognize or bind to other molecules in a sample (e.g., a biological sample). Acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding domains) "Main and guide nucleic acids," as used herein, refer to a compound or molecule.

[0147] Nucleic acid molecules useful in the methods of the invention include those encoding a polypeptide of the invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to an endogenous nucleic acid sequence. Although it is not necessary for the endogenous sequence to be identical, it generally indicates substantial identity. Polynucleotides that have "identity" generally have at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include those capable of hybridizing to the Any nucleic acid molecule encoding a specific polypeptide or fragment thereof is included. The nucleic acid molecule need not be 100% identical to the endogenous nucleic acid sequence, but generally will have substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence is generally , capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule.

[0148] "Hybridizing" refers to the formation of complementary polynucleotides under various stringency conditions. duplex with a nucleotide sequence (e.g., a gene described herein), or a portion thereof (See, e.g., Wahl, GM and SL Be rger (1987) Methods Enzymol. 152: 399; Kimmel, AR (1987) Methods Enzymol. 152 : 507).

[0149] For example, a stringent salt concentration is typically about 750 mM NaCl and 75 mM trisodium citrate. less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 500 mM NaCl and 50 mM trisodium citrate, Less than about 250 mM NaCl and 25 mM trisodium citrate. In the absence of a nucleotide, low stringency hybridization can be obtained. While at least about 35% formamide, more preferably at least about 50% formamide In the presence of mide, high stringency hybridization can be obtained. Stringent temperature conditions typically include at least about 30°C, and more preferably at least about 50°C. This includes temperatures of at least about 37°C, and most preferably at least about 42°C. The reaction time, concentration of detergent (e.g., sodium dodecyl sulfate (SDS)), and calibration Various additional parameters, such as the inclusion or exclusion of DNA, are well known to those skilled in the art. Depending on the situation, various combinations of these conditions can be used to achieve different levels of stringency. In a preferred embodiment, hybridization is carried out in 750 mM NaCl. 1, 75 mM trisodium citrate, and 1% SDS at 30°C. In this condition, hybridization was performed in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS , 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37°C. In the most preferred embodiment, hybridization is carried out in 250 mM NaCl, 25 mM trisodium citrate, The assay is carried out at 42°C in 5% thorium, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art.

[0150] For most applications, the wash steps following hybridization are Wash stringency conditions vary depending on salt concentration and temperature. As mentioned above, the stringency of the wash can be defined as the decreasing salt concentration. This can be increased by increasing the stringency of the wash step or by increasing the temperature. The preferred salt concentration is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, most preferably Preferably, the concentration is less than about 15 mM NaCl and 1.5 mM trisodium citrate. The appropriate temperature conditions typically include at least about 25°C, more preferably at least about 42°C. and even more preferably at least about 68° C. In one embodiment, the cleaning step The assay is performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In one embodiment, the wash steps are performed in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SD In a more preferred embodiment, the wash steps are carried out in 15 mM NaCl, 1.5 mM S at 42°C. The procedure is carried out in 0.1% trisodium citrate and 0.1% SDS at 68°C. Additional variations of these conditions are Hybridization techniques will be readily apparent to those skilled in the art. The method is well known in the art, for example, by Benton and Davis (Science 196: 180, 1977); Grunstein and H ogness (Proc. Natl. Acad. Sci., USA 72: 3961, 1975); Ausubel et al. (Current Pro tocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kim mel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Labo ratory Press, New York.

[0151] "Split" means divided into two or more pieces.

[0152] A "split Cas9 protein" or "split Cas9" is a protein that contains two separate nucleotides. The Cas9 protein is provided as an N-terminal fragment and a C-terminal fragment encoded by the The polypeptides corresponding to the N-terminal and C-terminal parts of the Cas9 protein are spliced ​​together. In certain embodiments, the Cas9 tag may be fused to form a "reconstituted" Cas9 protein. The protein is, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 20 14 or as described in Jiang et al. (2016) Science 351: 867-871. As shown, the protein is split into two fragments within a disordered region. PDB file: 5F9R, each of which is incorporated herein by reference. So, the protein is roughly between amino acids A292 and G364, F445 and K483, or E565 and T637. any C, T, A, or S within the SpCas9 region, or any other Cas9, Cas9 variant It splits into two fragments at the corresponding position in the target gene (e.g., nCas9, dCas9) or other napDNAbp. In some embodiments, the protein is SpCas9 T310, T313, A456, S469, or C574 into two fragments. In some embodiments, the protein is split into two fragments. The process of dividing the protein into smaller fragments is called "cleavage."

[0153] In other embodiments, the N-terminal portion of the Cas9 protein is S. pyogenes Cas9 wild type (SpCas9 ) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2) The C-terminal portion of the Cas9 protein contains amino acids 574 to 1368 or 6 in the wild-type SpCas9. Includes parts 38 to 1368.

[0154] The C-terminal part of split Cas9 was combined with the N-terminal part of split Cas9 to form a complete Cas9 tag. In some embodiments, the C-terminal portion of the Cas9 protein can form a protein. The sequence begins at the end of the N-terminal portion of the Cas9 protein. In some embodiments, the C-terminal portion of the split Cas9 is from amino acids (551-651) to 1368 of spCas9. "(551-651)-1368" refers to the amino acids between 551 and 651 (inclusive). This means that the sequence begins with amino acid 1368 and ends with amino acid 1368. For example, the C The terminal parts are amino acids 551 to 1368, 552 to 1368, 553 to 1368, 554 to 1368, and 555 to 1368 of spCas9. 8, 556~1368, 557~1368, 558~1368, 559~1368, 560~1368, 561~1368, 562~1368, 563~1368, 564~1368, 565~1368, 566~1368, 567~1368, 568~1368, 569~1368, 570 ~1368, 571~1368, 572~1368, 573~1368, 574~1368, 575~1368, 576~1368, 577~1 368, 578~1368, 579~1368, 580~1368, 581~1368, 582~1368, 583~1368, 584~1368 , 585~1368, 586~1368, 587~1368, 588~1368, 589~1368, 590~1368, 591~1368, 5 92~1368, 593~1368, 594~1368, 595~1368, 596~1368, 597~1368, 598~1368, 599 ~1368, 600~1368, 601~1368, 602~1368, 603~1368, 604~1368, 605~1368, 606~1 368, 607~1368, 608~1368, 609~1368, 610~1368, 611~1368, 612~1368, 613~1368 , 614~1368, 615~1368, 616~1368, 617~1368, 618~1368, 619~1368, 620~1368, 6 21~1368, 622~1368, 623~1368, 624~1368, 625~1368, 626~1368, 627~1368, 628 ~1368, 629~1368, 630~1368, 631~1368, 632~1368, 633~1368, 634~1368, 635~1 368, 636~1368, 637~1368, 638~1368, 639~1368, 640~1368, 641~1368, 642~1368 , 643~1368, 644~1368, 645~1368, 646~1368, 647~1368, 648~1368, 649~1368, 6 50-1368, or 651-1368. The C-terminal portion of the split Cas9 protein is located between amino acids 574 and 1368 or 638 of SpCas9. Includes the part ~1368.

[0155] A "subject" is a human or non-human mammal, e.g., a non-human primate (monkey), cow, horse, "Animal" means a mammal, including but not limited to a dog, sheep, or cat. In some embodiments, the subject described herein comprises a pathogenic mutation in a polynucleotide sequence. nothing.

[0156] "Substantially identical" means that a polypeptide or nucleic acid molecule has a similar sequence to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequences (e.g., It means that the nucleic acid sequence has at least 50% identity to the nucleic acid sequence of any one of the nucleic acid sequences listed above. In embodiments, such sequences have, at the amino acid level, the following sequence relative to the sequence used for comparison: or at least 60%, 80%, or 85%, 90%, 95%, or even 99% identical at the nucleic acid level It is sex.

[0157] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group, Univ. versity of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 sequence analysis software packages, BLAST, BESTFIT, COBALT, EMBOSS Needle, GA P, or PILEUP / PRETTYBOX program). by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions are usually made in the following groups: Lysine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, Asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine An exemplary method for measuring the degree of identity is to use a substitution between e-3 and e-1. The BLAST program, in which a probability score of 0.00 indicates closely related sequences, can be used. , for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1 , b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on, and c) Query clustering parameters: Use query clusters on; Word Size 4; Max cluster ter distance 0.8; Alphabet Regular. EMBOSS Needle can be used, for example, with the following parameters: a) Matrix: BLOSUM62, b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair, e) END GAP PENALTY: false, f) END GAP OPEN: 10, and g) END GAP EXTEND:0.5.

[0158] The term "target site" refers to a site that is targeted by a deaminase (e.g., cytidine deaminase or adenine deaminase). adenosine deaminase) or a fusion protein containing a deaminase (e.g., dCas9-adenosine deaminase fusion protein or base editor disclosed herein It refers to a sequence within a nucleic acid molecule that is amino-conjugated.

[0159] As used herein, the terms "treat," "treating," and "treatment" include reduces or ameliorates the disorder and / or its associated symptoms, or produces the desired pharmacological effect. Treating a disorder or condition refers to obtaining a specific biological and / or physiological effect. Does not preclude the complete elimination of a disorder, condition, or its associated symptoms It will be understood that the effect does not necessarily have to be curative. The therapeutic effect is, without limitation, the reduction of the disease and / or adverse symptoms resulting from the disease. To partially or completely reduce, reduce, prevent, mitigate, alleviate, or lessen the intensity of a condition In some embodiments, the effect is preventative, i.e., the effect is curative or curative of the disease. To this end, the present invention provides a method for preventing or preventing the onset or recurrence of a disease or a condition. The method comprises administering a therapeutically effective amount of a composition as described herein.

[0160] "Uracil glycosylase inhibitors" or "UGIs" are compounds that inhibit the uracil excision repair system. In one embodiment, the agent binds to the host's uracil DNA glycosylase. and a protein or fragment thereof that prevents the removal of uracil residues from DNA. In its morphological form, UGI is capable of inhibiting the base excision repair enzyme uracil DNA glycosylase. In some embodiments, UG is a protein, fragment, or domain thereof that can The I domain comprises wild-type UGI or a modified version thereof. I-domains include fragments of the exemplary amino acid sequences set forth below. In the present specification, a UGI fragment is at least 60%, at least 65%, or at least 100% of the exemplary UGI sequences shown below. At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 Contains 5%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% In some embodiments, the UGI comprises an exemplary amino acid sequence, as described below. Some embodiments include amino acid sequences that are homologous to the UGI amino acid sequence or a fragment thereof. In certain embodiments, UGI or a portion thereof may be a wild-type UGI or UGI sequence, or a portion thereof, as described below. or part thereof, at least 70%, at least 75%, at least 80%, at least 85% , at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identity. A typical UGI contains the following amino acid sequence: >splP14739IUNGI_BPPB2 uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGE NKIKML.

[0161] Ranges provided herein are intended to be shorthand notations for all values ​​within the range. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 , 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 , 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood that the range includes any number, combination of numbers, or subrange from

[0162] The recitation of a list of chemical groups in any definition of a variable herein does not include any single group or sequence. The combination of any of the listed groups includes the definition of that variable. The recitation of an embodiment of any aspect may be used as any single embodiment or as any other embodiment. It includes the embodiment in combination with the embodiment or any portion thereof.

[0163] Any composition or method provided herein may be used in combination with any other composition or method provided herein. The present invention can be combined with one or more of the compositions and methods.

[0164] The description and examples herein set forth in detail embodiments of the present disclosure. It is understood that the present invention is not limited to the particular embodiments described, as such may vary. Those skilled in the art will recognize that there are many variations and modifications of this disclosure that fall within the scope of this disclosure. You will realize that it is included.

[0165] All terms are intended to be understood as would be understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein are It has the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0166] Practice of the several embodiments disclosed herein is by those skilled in the art unless otherwise indicated. The skills covered include immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, and genomics. and conventional recombinant DNA techniques are used. See, e.g., Sambrook and Green, Molecular Cloning. ning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in M olecular Biology (FM Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Ham es and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Lab oratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Spe. See Generalized Applications, 6th Edition (RI Freshney, ed. (2010)).

[0167] Various features of the disclosure may be described in the context of a single embodiment, but those features may also be combined as separate components. The present disclosure is not to be construed as limiting the scope of the present invention. Although described herein in the context of separate embodiments for purposes of illustration, the present disclosure also provides for a single implementation. The section headings used herein are for organizational purposes only. and should not be construed as limiting the subject matter described.

[0168] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be apparent from the following detailed description which describes exemplary embodiments, in which the principles of the present disclosure are utilized. The following detailed description can be obtained by reference to the accompanying drawings, which are set forth below. [Brief explanation of the drawings]

[0169] [Figure 1]

[0023] Figure 1 shows a series of graphs showing the percent A>G editing activity of the indicated adenosine base editors. Each editor is referenced by a number; for example, 433 indicates pNMG-B433, which is ABE8.32. Each editor referenced in the graphs was tested with gRNAs HRB03, HRB04, HRB08, HRB12, and ng-424, respectively. gRNA sequences are provided in Example 3. [Figure 2] A heatmap showing the percent A>G editing activity in gray shading for the designated adenosine base editors (ABE8 and ABE9) listed in Table 14. Each editor listed in the figure was tested with a different gRNA: HRB03, HRB04, HRB08, HRB12, and ng-424. [Figure 3] A table showing the TadA deaminase variant (e.g., TadA*9; ABE9) and Cas9 (e.g., SpCas9) variant components of the adenosine base editors described herein is provided. These ABE9 base editors have A>G editing activity and are useful for correcting SNP mutations associated with alpha-1 antitrypsin disease (A1AD), such as PiZ mutations in the SERPINA1 gene. In some cases, SpCas9 variants have specificity for 5'-NGC-3' PAM. Figure 3A lists the adenosine base editors by their plasmid number. Figures 3B and 3C show various TadA deaminase variants and the amino acid mutations contained in the Tad*7.10 amino acid sequence, as well as PAM variants and the amino acid mutations contained therein. [Figure 4]Nucleic acid sequences, tables, and graphs related to the enhanced rate of nucleobase modification generation via base editor modifications are shown. Figures 4A and 4B show nucleic acid sequences and tables related to the enhanced rate of nucleobase modification generation in primary PiZZ fibroblasts via base editor modifications, as described in Figures 4C and 4D, and the increased serum alpha-1 antitrypsin (A1AT) levels produced by lipid nanoparticle (LNP)-mediated delivery and base editing in NSG-PiZ transgenic mice, as described in Figures 5A and 5B below. In particular, Figure 4A shows a target DNA sequence containing the target site (A at position 7 of the target DNA sequence) encoding the PiZZ mutation associated with A1AD. This sequence contains a 20-nucleotide protospacer and a non-canonical spCas9 NGC PAM. Beneficial editing at position A7 = wild type (WT), and editing at positions A5 and A7 = WT + D341G are also shown. Figure 4B shows a table describing the components of the various base editor TadA deaminase variants and Cas9 PAM variants used to correct PiZ mutations. The table lists the variants (e.g., variants (Var) 1-9) used to obtain the results shown in Figures 4C, 4D, 5A, and 5B. In this table, the amino acid mutations of SpCas9 (SpCas9 variants) are listed in the right-most column (PAM variant) of the table. The "RVRFRAR" SpCas9 variant contains the following mutations: L1111R + D1135V + G1218R + E1219F + A1322R + R1335A + T1337R. Figures 4C and 4D show bar graphs depicting the editing rates observed in patient-derived PiZZ fibroblasts (GM11423, Corriel Biorepository) transfected with base editing reagents using the Neon electroporation system. Each treatment consisted of 10 μl of electroporation buffer containing 70,000 fibroblasts, 100 ng of mRNA encoding the base editor, and 50 ng of α-1-modified gRNA. After 48 hours of recovery, cells were lysed and the locus of interest was interrogated by targeted amplicon sequencing. Data were obtained from two independent experiments.These data and results indicate that both optimizing NGC PAM recognition (variants 1–3, Figures 4B and 4C) and optimizing TadA deaminase (e.g., ABE9) by incorporating mutations into the TadA deaminase (variants 4–9, Figures 4B–4D) improve the efficiency of target base editing. [Figure 5] Figures 5A and 5B show graphs related to the increase in serum A1AT produced by lipid nanoparticle (LNP)-mediated delivery and base editing in NSG-PiZ transgenic mice. Tables of the target site DNA sequences of the various editors used to correct PiZ mutations, as well as the TadA deaminase variant and Cas9 PAM variant components, are provided in Figures 4A and 4B above. Figure 5A shows a graph depicting the editing rate observed in whole liver gDNA from the NSG-PiZ transgenic mouse model 7 days after treatment with 1.5 mg / kg LNPs containing a 1:1 weight ratio of gRNA and mRNA encoding the base editor. Commercially available NSG-PiZ mice (The Jackson Laboratory, Mount Desert Island, ME) express mutant human SERPINA1 (Glu342Lys mutation) in an immunodeficient NOD-SCIDγ (NSG) background, which provides a stable background for human hepatocytes after partial hepatectomy. The results show that ngcABEvar9 (Figure 4B) resulted in higher editing rates than the previous version, variant 8. Figure 5B shows a graph demonstrating that increases in serum alpha-1 antitrypsin (A1AT) (post-blood draw) compared to pre-treatment samples (pre-blood draw), as measured by MSD sandwich immunoassay, correlated with editing rates. Based on these results, base editing using the TadA deaminase variants described herein can address alpha-1 antitrypsin deficiency and its potential pulmonary sequelae. DETAILED DESCRIPTION OF THE INVENTION

[0170] The present invention provides novel adenine base editors (e.g., ABE9) and and methods of using these adenosine deaminase variants.

[0171] Nucleobase Editor As used herein, editing, modifying, or altering a target nucleotide sequence of a polynucleotide Novel base editors (e.g., ABE8 and ABE9) or nucleobase editors for modifying In particular, the novel ABE9 base editor and its component adenosine deaminase are disclosed. , as set forth in Tables 14 and 18 below. As used herein, polynucleotide programmable Nucleotide binding domains and nucleobase editing domains (e.g., adenosine deaminase ) and a base editor. The nucleotide-binding domain capable of binding to the bound guide polynucleotide (e.g., gR NA), specifically binds to the target polynucleotide sequence (i.e., Complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence (via a nucleic acid sequence encoding a base editor), thereby localizing the base editor to the target nucleic acid sequence desired to be edited. In some embodiments, the target polynucleotide sequence can be single-stranded. In some embodiments, the target polynucleotide sequence comprises RNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. Includes:

[0172] Polynucleotide-programmable nucleotide-binding domains Polynucleotide-programmable nucleotide-binding domains also bind to RNA It should be understood that nucleic acids may include programmable proteins. A programmable nucleotide-binding domain guides it to RNA Other nucleic acid programmable DNA binding proteins are also within the scope of this disclosure. Although within the scope of the present disclosure, they are not specifically described.

[0173] The polynucleotide-programmable nucleotide-binding domain of the base editor It may itself contain one or more domains. For example, a polynucleotide programmable nucleic acid The nucleotide binding domain may comprise one or more nuclease domains. In embodiments, the nuclease of the polynucleotide programmable nucleotide binding domain The enzyme domain may comprise an endonuclease or an exonuclease. In this context, the term "exonuclease" refers to an enzyme that removes nucleic acids (e.g., RNA or DNA) from free ends. The term "endonuclease" refers to a protein or polypeptide that can be digested " refers to catalyzing (e.g., cleaving) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease refers to a protein or polypeptide that is capable of The enzyme is capable of cleaving a single strand of a double-stranded nucleic acid. Endonucleases are capable of cleaving both strands of a double-stranded nucleic acid molecule. In embodiments, the polynucleotide programmable nucleotide binding domain is a deoxyribonucleotide. In some embodiments, the polynucleotide programmable The capable nucleotide binding domain may be a ribonuclease.

[0174] In some embodiments, the polynucleotide-programmable nucleotide binding domain The nuclease domain of the ribosomal protein cleaves zero, one, or two strands of the target polynucleotide. In some cases, polynucleotides can be programmable to bind nucleotides. The binding domain may comprise a nickase domain. As used herein, the term "nickase" refers to the cutting of only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). Polynucleotides containing nuclease domains capable of programmable nucleotide binding In some embodiments, a nickase is an active polynucleotide protease. By introducing one or more mutations into the programmable nucleotide-binding domain, A fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide For example, a polynucleotide programmable nucleic acid sequence may be derived from a nucleic acid sequence that binds to a nucleic acid. When the peptide-binding domain contains a nickase domain derived from Cas9, the nickase domain derived from Cas9 is The case domain may contain a D10A mutation and a histidine at position 840. In such a case, residue H 840 retains catalytic activity, thereby enabling single-strand cleavage of nucleic acid duplexes. For example, the nickase domain from Cas9 can contain a H840A mutation, while the 10th position The amino acid residue at position 1 remains D. In some embodiments, the nickase is fully catalytic. Polynucleotide-programmable nucleotide binding in functionally active (e.g., native) form domain, all or part of the nuclease domain that is not required for nickase activity For example, the polynucleotide programmable nucleic acid sequence can be induced by removing the When the peptide-binding domain contains a nickase domain derived from Cas9, the nickase domain derived from Cas9 is The case domain may include a deletion of all or part of the RuvC domain or the HNH domain. .

[0175] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0176] catalytically inactive (i.e., incapable of cleaving a target polynucleotide sequence) Polynucleotide base editors containing programmable nucleotide binding domains are also available. As used herein, the terms "catalytically inactive" and "nucleic acid" are used interchangeably. "Nuclease inactivity" refers to the inability of a nucleic acid strand to be cleaved by one or more enzymes. Polynucleotides with mutations and / or deletions on programmable nucleotide binding In some embodiments, catalytically inactive domains are used interchangeably to refer to catalytically inactive domains. A polynucleotide programmable nucleotide-binding domain base editor is As a result of specific point mutations in the above nuclease domains, nuclease activity is lost. For example, in the case of a base editor containing a Cas9 domain, Cas9 can convert the D10A mutation and the H840A Such mutations may inactivate both nuclease domains, This results in a loss of nuclease activity. The polynucleotide programmable nucleotide binding domain can be a catalytic domain (e.g. , RuvC1 and / or HNH domains). In some embodiments, catalytically inactive polynucleotides can be used to programmably bind nucleotides. The domains contain point mutations (e.g., D10A or H840A), and all or part of the nuclease domain. or a partial deletion.

[0177] Also provided herein are polynucleotide-programmable nucleotide binding domains. From the previous functional version, catalytically inactive polynucleotide programmable nucleic acid Mutations that can generate peptide-binding domains are contemplated. For example, catalytically inactive In the case of deoxyribonucleotide Cas9 ("dCas9"), mutants with mutations other than D10A and H840A are provided, and This results in a nuclease-inactivated Cas9. Such mutations include, for example, D10 and H84 0, or other amino acid substitutions within the nuclease domain of Cas9 (e.g., HNH nuclease subdomain and / or RuvC1 subdomain). Suitable nuclease-inactive dCas9 domains can be prepared by one of skill in the art based on this disclosure and knowledge in the art. Such additional exemplary suitable nuclei may be apparent to those skilled in the art and are within the scope of this disclosure. The cleavage-inactive Cas9 domains include D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H These include, but are not limited to, the 840A / N863A mutation domain (see, e.g., Prashant et al. ., CAS9 transcriptional activators for target specificity screening and paired n ickases for cooperative genome engineering. Nature Biotechnology. 2013; 31 (9): 833-838, the entire contents of each of which are incorporated herein by reference.

[0178] Polynucleotide programmable nucleotide sequences that can be incorporated into base editors Non-limiting examples of domain-binding domains include domains derived from CRISPR proteins, restriction nuclease nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases In some cases, base editors include natural or modified proteins. Polynucleotide-programmable nucleotide-binding domains containing proteins or portions thereof CRISPR (i.e., Clustered Regularly Interleaved) Interspaced Short Palindromic Repeats (IRRs)-mediated nucleic acid modification is a process that binds to nucleic acid sequences. Such proteins are referred to herein as "CRISPR proteins." Thus, as used herein, a polynucleotide comprising all or a portion of a CRISPR protein is Base editors containing programmable nucleotide-binding domains (i.e., The entire CRISPR protein, also known as the "CRISPR protein-derived domain" of the The present invention also discloses a base editor that includes a domain of the base editor. The domains derived from CRISPR proteins are similar to those in the wild-type or naturally occurring versions of CRISPR proteins. For example, as described below, CRISPR protein-derived domains can be modified relative to the In is a gene that contains one or more mutations compared to the wild-type or native version of a CRISPR protein, It may include insertions, deletions, rearrangements and / or recombinations.

[0179] CRISPR targets mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains a spacer, the aforementioned mobile The CRISPR cluster contains a sequence complementary to the transcription element and a target invader nucleic acid. In type II CRISPR systems, pre-crRNA is generated and processed into CRISPR RNA (crRNA). For proper processing, trans-encoded small RNAs (tracrRNAs), endogenous RNAs, and tracrRNA requires the transcription factor 3 (rnc) and Cas9 proteins. It acts as a guide for ribonuclease 3-assisted processing. NA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then split into 3 strands. In nature, DNA is bound and cleaved by exonucleolytic trimming. Usually, both the protein and the RNA are required. However, both the crRNA and tracrRNA are required. A single guide RNA ("sgRNA," or simply "gRNA") is used to incorporate aspects of the present invention into a single RNA species. ) can be prepared. For example, Jinek M., Chylinski K., Fonfara I., Hauer M., See Doudna JA, Charpentier E. Science 337: 816-821 (2012) (these The entire contents of which are incorporated herein by reference. Cas9 binds to short motifs of CRISPR repeat sequences. recognizes the PAM or protospacer adjacent motif to distinguish self from non-self It is useful.

[0180] In some embodiments, the methods described herein utilize engineered Cas proteins. Guide RNA (gRNA) contains a scaffold sequence required for Cas binding and a target sequence to be modified. User-defined sequences that define genomic (or polynucleotide, e.g., DNA or RNA) targets It is a short synthetic RNA consisting of a defined spacer of about 20 nucleotides. By modifying the target sequence present in the gRNA, the Cas protein can be targeted to the genome or polynucleotides. The specificity of the Cas protein can be altered by modifying the gRNA targeting sequence. How well a sequence is aligned with a polynucleotide target sequence in the genome compared to other parts of the genome In one embodiment, the Cas protein is SpCas9. is.

[0181] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGCUAGAAAUAGCAA GUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU.

[0182] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGCUAGAAAUAGCAA GUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGGACCGAGUCGGUGCUUUU.

[0183] In one embodiment, the terminal uracils (U) of the gRNA scaffold described above are optionally "mU*mU*mU*U" which exhibit 2'OMe and have phosphorothioate linkages.

[0184] In one embodiment, the RNA scaffold comprises a stem loop. Sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACA Contains AAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.

[0185] In one embodiment, the S. pyrogenes sgRNA scaffold polynucleotide sequence is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.

[0186] In one embodiment, the S. aureus sgRNA scaffold polynucleotide sequence is as follows: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGC GAGA.

[0187] In one embodiment, the sgRNA scaffold of BhCas12b has the following polynucleotide sequence: GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUC UUACGAGGCAUUAGCAC.

[0188] In one embodiment, the BvCas12b sgRNA scaffold has the following polynucleotide sequence: GACCUAUAGGGUCAAUGAAUCUGGGCGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUG CUUGGCAC.

[0189] In some embodiments, the CRISPR protein-derived domain that is incorporated into the base editor The nucleic acid is capable of binding to a target polynucleotide when combined with a bound guide nucleic acid. Endonucleases (e.g., deoxyribonucleases or ribonucleases) that can In some embodiments, the CRISPR protein that is incorporated into the base editor The protein-derived domain, when combined with the bound guide nucleic acid, binds to the target polynucleotide. In some embodiments, the base editor is a nickase that can bind to The domain from the CRISPR protein to be integrated is combined with the bound guide nucleic acid. It is a catalytically inactive domain that can bind to a target polynucleotide when In some embodiments, the target to which the CRISPR protein-derived domain of the base editor binds is In some embodiments, the target polynucleotide is DNA. The target polynucleotide to which the protein-derived domain binds is RNA.

[0190] Cas proteins that may be used herein include class 1 and class 2. Non-limiting examples of proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, as5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, C sy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Cs n2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2 , Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Cs f4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf 1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, their homologs, or modified or altered versions thereof. Cas9, which has two functional endonuclease domains, RuvC and HNH. The unmodified CRISPR enzyme may have DNA cleavage activity. The CRISPR enzyme may cleave DNA at a target sequence, e.g. Induce cleavage of one or both strands within the target sequence and / or within the complementary strand of the target sequence. For example, the CRISPR enzyme can target the first or last nucleotide of the target sequence. Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more The cleavage of one or both strands can be induced within more base pairs of the target site.

[0191] The mutant CRISPR enzyme cleaves one or both strands of the target polynucleotide containing the target sequence. Encode a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme so that it lacks the ability to cleave the target gene. Cas9 can be expressed as a wild-type exemplary Cas9 polypeptide (e.g., For example, Cas9 from S. pyogenes), at least (or at least about) 50%, 60% , 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity Cas9 may refer to a polypeptide having sequence identity and / or sequence homology. For a suitable Cas9 polypeptide (e.g., from S. pyogenes), at most (or (At least approximately) 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, Alternatively, it may refer to a polypeptide having 100% sequence identity and / or sequence homology. may be a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof. The term "Cas9" may refer to wild-type or modified forms of the Cas9 protein, which may contain amino acid changes such as combinations of amino acids.

[0192] In some embodiments, the CRISPR protein-derived domain of the base editor is bacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphth eria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwane nse (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Bellie lla baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721. 1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI R ef: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria men ingitidis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococc The Cas9 vector may comprise all or part of Cas9 from U. aureus.

[0193] Nucleobase editor Cas9 domain The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, e.g., "Complete genome sequencing"). sequence of an M1 strain of Streptococcus pyogenes.”.Ferretti et al., Proc. Nat l. Acad. Sci. USA 98: 4658-4663 (2001); “CRISPR RNA maturation by trans-enco ded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471: 602- 607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive ba See Jinek M. et al., Science 337: 816-821 (2012) The entire contents of each of these are incorporated herein by reference. Cas9 orthologs include It has been described in various species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpenti, er, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2 013) Cas9 sequences derived from the organisms and loci disclosed in RNA Biology 10: 5, 726-737. columns are listed; the entire contents of which are incorporated herein by reference.

[0194] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 driver. Non-limiting exemplary Cas9 domains are provided herein. The main ones are nuclease-active Cas9 domains, nuclease-inactive Cas9 domains, or Ca In some embodiments, the Cas9 domain may be a nuclease. For example, the Cas9 domain binds both strands of a double-stranded nucleic acid (e.g., double-stranded DNA). In some embodiments, the Cas9 domain may be a Cas9 domain that cleaves both strands of the molecule. The domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain may be a Cas9 domain having at least one amino acid sequence similar to any one of the amino acid sequences described herein. At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 8 5%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, Some contain amino acid sequences that are at least 99%, or at least 99.5%, identical to one another. In embodiments, the Cas9 domain is compared to any one of the amino acid sequences described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 2 2, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 4 Contains an amino acid sequence with 2, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. Compared to one, at least 10, at least 15, at least 20, at least 30, at least At least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least at least 100, at least 150, at least 200, at least 250, at least 300, at least At least 350, at least 400, at least 500, at least 600, at least 700, at least 800 , at least 900, at least 1000, at least 1100, or at least 1200 identical sequences The amino acid sequence includes the following amino acid residues:

[0195] In some embodiments, proteins comprising fragments of Cas9 are provided. For example, some In embodiments, the protein is (1) the gRNA binding domain of Cas9; or (2) the DNA binding domain of Cas9. In some embodiments, the Cas9 domain comprises one of two Cas9 domains, the cleavage domain. Proteins containing s9 or fragments thereof are referred to as "Cas9 variants." share homology with Cas9 or a fragment thereof. For example, a Cas9 variant may be similar to wild-type Cas9. , at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% Identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In embodiments, the Cas9 variant has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, , 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 , 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 In some embodiments, the Cas9 variant may have 50 or more amino acid changes. The construct contains a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) and As a result, the fragments are at least about 70% identical, at least about 80% identical to the corresponding fragment of wild-type Cas9. , at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, the fragment is at least about 99.9% identical to the corresponding wild-type Cas9. , at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least Both are 97%, at least 98%, at least 99%, or at least 99.5% amino acids in length. In some embodiments, the fragment is at least 100 amino acids in length. In this embodiment, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600 , 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids in length.

[0196] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The Cas9 sequence may comprise a full-length amino acid sequence of interest, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. First, it includes only one or more fragments thereof. As used herein, suitable Cas9 domains and Cas9 Exemplary amino acid sequences of the fragments are provided, further suitable sequences of Cas9 domains and fragments are: It will be apparent to one skilled in the art.

[0197] The Cas9 protein binds to specific DNA sequences that are complementary to its guide RNA. In some embodiments, the polynucleotide primer may be linked to a guide RNA. The programmable nucleotide binding domain can be a Cas9 domain, e.g., a domain with nuclease activity. Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas 12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Ca Non-limiting examples of Cas enzymes include, but are not limited to, s12i, s12j / CasΦ, and Cas12j / CasΦ. Examples include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, and Cas6. , Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas1 0, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas1 2g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4 , Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 , Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Cs x3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh 2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effectors -protein, type VI Cas effector protein, CARF, DinG, or their homologs, or Modified or altered versions thereof are included.

[0198] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NCBI Reference Reference sequence: NC_017053.1, corresponds to the following nucleotide and amino acid sequences: JPEG2025170240000012.jpg42169JPEG2025170240000013.jpg228157JPEG2025170240000014.jpg239157JPEG2025170240000015.jpg154167

[0199] In some embodiments, the wild-type Cas9 has the following nucleotide sequence and / or amino acid sequence: Corresponding to or containing sequences: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAAACATGAACGGCACCCCATCTTTGGAAACATAGGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAAGATGGATGGGAGCGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGAACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCAAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025170240000016.jpg188167

[0200] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NCBI reference Sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: Q99ZW2 (nucleotide sequence below) corresponding to the amino acid sequence: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAATTAACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAGGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AAATTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAGGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATTTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025170240000017.jpg195168

[0201] These genes are Cas9 and are Corynebacterium ulcerans (NCBI Refs: NC_0156). 83.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_01678). 6.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (N CBI Ref: NC_017861.1); Spiroplasm taiwan (NCBI Ref: NC_021846.1); Streptococcus ccus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); P.S sychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI). Ref: YP_820832.1); Listeria harmless (NCBI Ref: NP_472073.1); Campylobacter jejun i (NCBI Ref: YP_002344900.1); or Neisseria meningitidis (NCBI Ref: YP_00234 2100.1) or Cas9 from any other organism.

[0202] Additional Cas9 proteins, including variants and homologs (e.g., nuclease-inactive Cas9 proteins, e.g., Cas9 proteins ... s9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9) are within the scope of this disclosure. It is understood that the present invention is within the scope of the present invention. Exemplary Cas9 proteins include the proteins shown below. In some embodiments, the Cas9 protein includes, but is not limited to, In some embodiments, the Cas9 protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is a nickase (nCas9). It is a lyase-active Cas9.

[0203] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9 For example, the dCas9 domain can cleave either strand of a double-stranded nucleic acid molecule. In some embodiments, the nucleic acid sequence may be linked to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule). The nuclease-inactive dCas9 domain may be a D10X mutation in the amino acid sequence described herein. and H840X mutations, or the corresponding mutations in any of the amino acid sequences provided herein. and X is any amino acid change. The s9 domain may have the D10A and H840A mutations of the amino acid sequence described herein, or The present invention also includes corresponding mutations in any of the amino acid sequences provided herein. The enzyme-inactive Cas9 domain was inserted into the cloning vector pPlatTET-gRNA2 (accession no. BAV54124) The amino acid sequence described in

[0204] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence -specific control of gene expression.” Cell.2013;152(5):1173-83; the entire contents of which are incorporated herein by reference).

[0205] Additional suitable nuclease-inactive dCas9 domains may be selected based on this disclosure and the knowledge in the art. Such additional exemplary suitable methods will be apparent to those skilled in the art based on the present disclosure. Suitable nuclease-inactive Cas9 domains include D10A / H840A, D10A / D839A / H840A, and D10A / D 839A / H840A / N863A mutant domain (e.g., Prasha nt et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the entire contents of which are incorporated herein by reference.

[0206] In some embodiments, the Cas9 nuclease induces inactive (e.g., inactivating) DNA cleavage. domain, i.e., Cas9 is referred to as the "nCas9" protein ("nickase" Cas9). The nuclease-inactivated Cas9 protein is called the "dCas9" protein. This is referred to interchangeably as protein (nuclease "inactive" Cas9) or catalytically inactive Cas9. Generation of Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains Methods are known (e.g., Jinek et al., Science. 337: 816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152 (5): 1173-83 (all of the respective The contents of which are incorporated herein by reference. For example, the DNA cleavage domain of Cas9 includes HNH It is known to contain two subdomains: a nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain Mutations within these subdomains abolish the nuclease activity of Cas9. For example, mutations D10A and H840A can silence the nuclease of S. pyogenes Cas9. completely inactivates enzyme activity (Jinek et al., Science. 337: 816-821 (2012); Qi et al. ., Cell. 28; 152 (5): 1173-83 (2013)).

[0207] In some embodiments, the dCas9 domain is any of the dCas9 domains provided herein. For any one of them, at least 60%, at least 65%, at least 70%, at least 75%, At least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least Amino acids that are at least 97%, at least 98%, at least 99%, or at least 99.5% identical In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. Compared to any one of the sequences, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 domain comprises an amino acid sequence having the amino acid sequence of any one of the amino acids described herein. At least 10, at least 15, at least At least 20, at least 30, at least 40, at least 50, at least 60, at least 70, At least 80, at least 90, at least 100, at least 150, at least 200, at least 2 50, at least 300, at least 350, at least 400, at least 500, at least 600, At least 700, at least 800, at least 900, at least 1000, at least 1100, or or an amino acid sequence having at least 1200 identical contiguous amino acid residues.

[0208] In some embodiments, dCas9 contains one or more mutations that inactivate Cas9 nuclease activity. The amino acid sequence of the Cas9 gene may correspond to, or may partially or entirely contain, a variant of the Cas9 amino acid sequence. In some embodiments, the nuclease-inactive dCas9 domain comprises an amino acid sequence described herein. D10X and H840X mutations in the amino acid sequence, or any of the amino acid sequences provided herein. and X is any amino acid change. The nuclease-inactive dCas9 domain contains the D10A mutation and H 840A mutation, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the nuclease-inactive Cas9 domain is It contains the amino acid sequence described in pPlatTET-gRNA2 (accession number BAV54124).

[0209] In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG2025170240000018.jpg189167

[0210] In some embodiments, the amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is , as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence -specific control of gene expression." Cell. 2013; 152 (5): 1173-83 (the entire contents of which are incorporated herein by reference).

[0211] In some embodiments, the amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is , as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD

[0212] In some embodiments, the Cas9 domain comprises a D10A mutation, while the residue at position 840 is , the histidine in the amino acid sequence provided above or the histidine in the amino acid sequence provided herein. The amino acid sequence is located at a corresponding position in either of the amino acid sequences.

[0213] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, which For example, this mutation results in nuclease-inactivated Cas9 (dCas9). and other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9. Substitutions (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) In some embodiments, the amino acid sequence is at least about 70% identical, at least about 80% identical, or at least about 100% identical. At least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, A variant or relative of dCas9 that is at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, the isomers are selected from the group consisting of about 5 amino acids, about 10 amino acids, and about 15 amino acids. , about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids dCas9 variants with short or long amino acid sequences, about 100 amino acids or more Provide a riant.

[0214] Additional suitable nuclease-inactive dCas9 domains may be selected based on this disclosure and the knowledge in the art. Such additional exemplary suitable methods will be apparent to those skilled in the art based on the present disclosure. Suitable nuclease-inactive Cas9 domains include D10A / H840A, D10A / D839A / H840A, and D10A / D 839A / H840A / N863A mutant domain (e.g., Prasha nt et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the entire contents of which are incorporated herein by reference.

[0215] In some embodiments, the Cas9 domain is a Cas9 nickase. , Cas9, which can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule) In some embodiments, the Cas9 nickase can be a protein. The Cas9 nickase cleaves the target strand, which is the gRNA (e.g., sgRNA) that is bound to the Cas9. ) means to cleave the strand that is base-paired (complementary) to the In this example, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule. However, this is because the Cas9 nickase base-pairs with the gRNA (e.g., sgRNA) that binds to Cas9. In some embodiments, the Cas9 nickase is H840A mutation and has an aspartic acid residue at position 10, or has a corresponding mutation. In some embodiments, the Cas9 nickase is any of the Cas9 nickases provided herein. For one, at least 60%, at least 65%, at least 70%, at least 75%, At least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 9 7%, at least 98%, at least 99%, or at least 99.5% identity Additional suitable Cas9 nickases include those described herein based on this disclosure and knowledge in the art. It is clear to one skilled in the art and is within the scope of this disclosure.

[0216] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD

[0217] In some embodiments, Cas9 is directed against archaeal organisms that comprise the domain and kingdom of unicellular prokaryotic microorganisms. refers to Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, programmable Suitable nucleotide-binding proteins are described, for example, in Burstein et al., "New CRISPR-Cas system s from uncultivated microbes.” Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 The CasX or CasY protein may be any of the CasX or CasY proteins described herein, the entire contents of which are incorporated herein by reference. The first to use genome-resolved metagenomics to characterize the archaeal domain of life. Several CRISPR-Cas systems have been identified, including the reported Cas9. The various Cas9 proteins It was discovered in little-studied nanoarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered. These are among the most compact systems ever discovered. In some embodiments, in the base editor systems described herein, Cas9 may be fused to CasX, or In some embodiments, the salts described herein are replaced by variants of sX. In the base editor system, Cas9 is replaced by CasY or a variant of CasY. Other RNA-guided DNA-binding proteins can be classified as nucleic acid programmable DNA-binding proteins. It should be understood that such materials may be used as a substrate (napDNAbp) and are within the scope of the present disclosure. It is.

[0218] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The gram-competent DNA-binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX protein. or at least 85%, at least 90%, at least 91%, or at least at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least or at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. In some embodiments, the programmable nucleotide binding protein comprises a sequence , naturally occurring CasX or CasY proteins. In some embodiments, programmable The nucleotide binding protein can be any of the CasX or CasY proteins described herein. In contrast, at least 85%, at least 90%, at least 91%, at least 92%, at least 93% %, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, Contains amino acid sequences that are at least 99%, or at least 99.5%, identical to those from other bacterial species. It should be understood that the original CasX and CasY may also be used in accordance with the present disclosure.

[0219] Exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53)tr|F0N N87|F0NN87_SULIH CRISPR-associated Casx protein OS=Sulfolobus islandicus (strain HVE10 / 4) GN= SiH_0402 PE=4 SV=1) The amino acid sequence is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.

[0220] Exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS = Sulfolobus C. islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1) amino acid sequence is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.

[0221] デルタブローオバクテリア(Deltaproteobacteria)CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSY KSGKQPFVGAWQAFYKRRLKEVWKPNA

[0222] An exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein The amino acid sequence of the protein CasY (an uncultured Parcubacteria group bacterium) is as follows: As follows: MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI.

[0223] Cas9 nuclease contains two functional endonuclease domains, RuvC and HNH. Upon target binding, Cas9 undergoes a conformational change, which allows it to bind to the opposite side of the target DNA. The nuclease domain is positioned to cleave the opposite strand. The end result is a double-strand break (DSB) in the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). The resulting DSBs are repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) ) the less efficient but high-fidelity homology-directed repair (HDR) pathway.

[0224] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be determined by any convenient For example, in some cases, the efficiency can be calculated as the For example, the Surveyor Nuclease Assay can be used to measure the cleavage yield. The ratio of product to substrate can be used to calculate the rate. For example, direct insertion of DNA containing newly integrated restriction enzyme sequences as a result of successful HDR. Surveyor nuclease enzymes can be used to cleave substrates that are highly cleavable. The lower the value, the higher the HDR ratio (the higher the HDR efficiency). The percentage of HDR can be calculated using the following formula: [(cleavage product) / (base)] and cleavage products)] (e.g., (b+c) / (a+b+c), where "a" is the band intensity of the DNA substrate and "b" is the band intensity of the DNA substrate). and "c" is the cleavage product).

[0225] In some cases, efficiency can be expressed as the percentage of successful NHEJ. For example, The endonuclease I assay was used to generate cleavage products, and the ratio of product to substrate was used to determine the The rate of NHEJ can be calculated using the following formula: It cleaves mismatched heteroduplex DNA resulting from hybridization of two DNA strands ( NHEJ involves the creation of small random insertions or deletions at the original cleavage site. The more breaks there are, the higher the rate of NHEJ (the more efficient the NHEJ). As an illustrative example, the NHEJ rate (percentage) can be calculated using the following formula: It can be calculated using: (1-(1-(b+c) / (a+b+c)) 1 / 2 ) × 100, where "a" is the DNA substrate The band intensities are shown in Table 1, and "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154 (6): 1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8 (11): 2281-2308) .

[0226] The NHEJ repair pathway is the most active repair mechanism, resulting in the insertion or deletion of small nucleotides at DSB sites. Cas9 and gRNA or guide polynucleotide are expressed, which frequently leads to deletions (indels). The random nature of NHEJ-mediated DSB repair is important because a population of cells expressing the gene can develop diverse mutations. In most cases, NHEJ introduces small indels into the target DNA. resulting in amino acid deletions, insertions, or frameshift mutations, resulting in the target gene Generates a premature stop codon within the gene's open reading frame (ORF). Ideal The end result is typically a loss-of-function mutation in the target gene.

[0227] NHEJ-mediated DSB repair often disrupts gene open reading frames. However, homology-directed repair (HDR) can be used to detect single nucleotide changes and fluorophores. It is possible to generate specific nucleotide changes up to large insertions such as addition of nucleotides or tags. To utilize HDR for gene editing, a DNA repair template containing the sequence of interest is introduced into the gene editing system using a gRNA ( The target cell type can be delivered along with the target gene (or genes) and Cas9 or Cas9 nickase. The replica template contains the desired edit, as well as additional homologous sequences immediately upstream and downstream of the target ( The length of each homology arm can be determined by the This may depend on the size of the change to be made; the larger the insertion, the longer the homology arms will be required. The repair template may be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded oligonucleotide. It can be a double-stranded DNA plasmid and can express Cas9, gRNA, and an exogenous repair template. Even in cells that do, the efficiency of HDR is generally low (less than 10% of alleles are modified). Because HDR occurs during the S and G2 phases of the cell cycle, synchronizing cells can increase the efficiency of HDR. Chemical or genetic inhibition of genes involved in NHEJ also increases HDR frequency. It is possible.

[0228] In some embodiments, the Cas9 is a modified Cas9. A given gRNA targeting sequence is There may be additional sites where there is partial homology across the genome. These sites are called off-targets and must be taken into consideration when designing gRNAs. In addition to optimization, the specificity of CRISPR can also be increased through modifications of Cas9. It generates double-strand breaks (DSBs) through the combined activity of the two nuclease domains of S and HNH. The D10A variant of pCas9, Cas9 nickase, retains one nuclease domain and inhibits DSBs. The nickase system generates DNA nicks rather than cleavage, allowing for HDR-mediated targeting for specific gene editing. It can also be combined with gene editing.

[0229] In some cases, the Cas9 is a variant Cas9 protein. The peptide differs in one amino acid from the amino acid sequence of the wild-type Cas9 protein. The amino acid sequence may be any of the following: In some cases, the variant Cas9 polypeptide may inhibit the nuclease activity of the Cas9 polypeptide. have amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the In some cases, the variant Cas9 polypeptides may be nucleotide equivalents of the corresponding wild-type Cas9 protein. Less than 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the enzyme activity In some cases, the variant Cas9 protein has substantial nuclease activity. A variant of the subject Cas9 protein that does not have substantial nuclease activity. When it is a Cas9 protein, it may be referred to as "dCas9."

[0230] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, a variant Cas9 protein can be a variant of a wild-type Cas9 protein (e.g., a wild-type Cas9 protein Less than about 20%, less than about 15%, less than about 10%, less than about 5%, or less than about 1% of the endonuclease activity of the protein less than, or less than about 0.1%.

[0231] In some cases, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. can cleave the non-complementary strand of the double-stranded guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, variant Cas9 proteins contain mutations (amino acid substitutions) that reduce the function of the RuvC domain. As a non-limiting example, in some embodiments, a variant Cas9 protein may have Proteins have D10A (aspartic acid to alanine at amino acid position 10), thus It is possible to cleave the complementary strand of the single-stranded guide target sequence but the non-complementary strand of the double-stranded guide target sequence. The variant Cas9 protein has a reduced ability to cleave double-stranded targets. When cleaving nucleic acids, it generates single-strand breaks (SSBs) instead of double-strand breaks (DSBs). See, e.g., Jinek et al., Science. 2012 Aug. 17; 337 (6096): 816-21).

[0232] In some cases, the variant Cas9 protein binds to the non-complementary strand of the double-stranded guide target sequence. can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, variant Cas9 proteins have the HNH domain (RuvC / HNH / RuvC domain motif) It may have mutations (amino acid substitutions) that reduce its function. In embodiments, the variant Cas9 protein is H840A (histidine to acetyltransferase at amino acid position 840). lanin) mutation and therefore capable of cleaving the non-complementary strand of the guide target sequence but have a reduced ability to cleave the complementary strand of the guide target sequence (hence, variant Ca When the s9 protein cleaves the double-stranded guide target sequence, it generates an SSB rather than a DSB. Such a Cas9 protein cleaves a guide target sequence (e.g., a single-stranded guide target sequence). The guide target sequence (e.g., a single-stranded guide target sequence) is bound to the target sequence, but has a reduced ability to cleave the target sequence. It retains the ability to do so.

[0233] In some cases, the variant Cas9 protein is non-complementary to the complementary strand of the double-stranded target DNA. As a non-limiting example, in some cases, The variant Cas9 protein contains both the D10A and H840A mutations, thereby allowing the polypeptide to The peptides have a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). while retaining the ability to bind to target DNA (e.g., single-stranded target DNA).

[0234] As another non-limiting example, in some cases, the variant Cas9 protein is and W1126A mutations, which reduces the ability of the polypeptide to cleave target DNA. Such Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). have a reduced ability to bind to target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA There are.

[0235] As another non-limiting example, in some cases, the variant Cas9 protein comprises P475A , W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, thereby Such Cas9 proteins have a reduced ability to cleave target DNA. a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but They retain the ability to bind to target DNA.

[0236] As another non-limiting example, in some cases, the variant Cas9 protein is H840A , W476A, and W1126A mutations, which allows the polypeptide to cleave the target DNA. Such Cas9 proteins have reduced ability to target DNA (e.g., single-stranded target DNA). have a reduced ability to cleave target DNA but retain their ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, variant Cas9 proteins The polypeptide has H840A, D10A, W476A, and W1126A mutations, thereby Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., , single-stranded target DNA), but have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 retains the ability to bind to Cas9 The catalytic His residue at position 840 of the HNH domain (A840H) is restored.

[0237] As another non-limiting example, in some cases, the variant Cas9 protein is H840A , P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, thereby Such Cas9 proteins have a reduced ability to cleave target DNA. have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but As another non-limiting example, some In this case, the variant Cas9 proteins are D10A, H840A, P475A, W476A, N477A, ​​D1125A , W1126A, and D1127A mutations, which allows the polypeptide to cleave target DNA. Such Cas9 proteins have a reduced ability to target DNA (e.g., single-stranded target DNA). ), but has a reduced ability to cleave target DNA (e.g., single-stranded target DNA) In some cases, the variant Cas9 protein retains the W476A and W1126A mutations. If the variant Cas9 protein contains mutations P475A, W476A, N477A, ​​D1125A, W1126 When the variant Cas9 protein contained the A and D1127A mutations, it did not bind efficiently to the PAM sequence. Therefore, in some such cases, it is necessary to bind such variant Cas9 proteins. When used in a combined method, the method does not require a PAM sequence. In the case of using such variant Cas9 proteins in a conjugation method, the method involves the use of a guide RNA. A, but the method can be performed without a PAM sequence (and Specificity is thus provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve this (i.e., one or more nucleotides can be mutated). Non-limiting examples include residues D10, G12, G17, E762, H840, N Modify (i.e., replace) 854, N863, H982, H983, A984, D986, and / or A987 ) Mutations other than alanine substitutions are also suitable.

[0238] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., Ca The s9 protein was found to be D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A , H983A, A984A, and / or D986A), the variant Cas9 protein is As long as they retain the ability to interact with the target RNA, they can still target DNA in a site-specific manner. It is able to bind to the target DNA sequence (because it is guided to it by the guide RNA).

[0239] In some embodiments, the variant Cas protein is spCas9, spCas9-VRQR, spCas9 -VRER, xCas9(sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER, spCas9-MQKSER, spCas9-LRK IQK, or spCas9-LRVSQL.

[0240] In some embodiments, the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1 332A, R1335E, and T1337R (SpCas9-MQKFRAER), which contains the modified PAM5'-NGC-3' A modified SpCas9 with isomerism is used.

[0241] Alternatives to S. pyogenes Cas9 include RNA guides of the Cpf1 family that exhibit cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1 is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is a class II CRISPR / Cas This adaptive immune mechanism is a key player in the development of Prevotella and Franc The Cpf1 gene is associated with the CRISPR locus and is found in the fungus Gastrointestinal tract. Cp encodes an endonuclease that uses the id RNA to find and cut viral DNA. f1 is a smaller and simpler endonuclease than Cas9, and is less susceptible to the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in: The staggered cleavage pattern of Cpf1 is Similar to traditional restriction enzyme cloning, this allows for directional gene transfer. This allows for increased gene editing efficiency. Similarly, Cpf1 also influences the number of sites that can be targeted by CRISPR, which SpCas9 prefers. The C gene can be extended to AT-rich regions or AT-rich genomes lacking NGG PAM sites. The pf1 locus contains a mixed α / β domain, RuvC-I followed by a helical region, RuvC-II, and The Cpf1 protein contains a RuvC domain and a zinc finger-like domain. Furthermore, Cpf1 possesses a RuvC-like endonuclease domain similar to that of HNH endonuclease. It does not have a cleavage domain, and the N-terminus of Cpf1 lacks the α-helical recognition lobe of Cas9. 1. The CRISPR-Cas domain architecture of Cpf1 is functionally unique and is a class 2, type V The Cpf1 locus is more likely to be classified as a CRISPR system than a type I or III system. It encodes Cas1, Cas2, and Cas4 proteins similar to those of the Cpf1 gene. Functional Cpf1 is trans-activated. It does not require activating CRISPR RNA (tracrRNA), and therefore only CRISPR (crRNA) is required. This is because Cpf1 is not only smaller than Cas9, but also contains a smaller sgRNA molecule (approximately half the size of Cas9). The Cpf1-crRNA complex contains 10 nucleotides, making it effective for genome editing. In contrast to the G-rich PAM targeted by s9, the protospacer adjacent motif 5'-YTN-3' After PAM identification, Cpf1 cleaves the target DNA or RNA. The overhang of the nucleotide introduces sticky end-like DNA double-strand breaks.

[0242] In some embodiments, the Cas9 is a Cas9 vector with specificity for the engineered PAM sequence. In some embodiments, the additional Cas9 variant and PAM sequence are r, SM et al. Continuous evolution of SpCas9 variants compatible with non-G PA Ms, Nat. Biotechnol. (2020), which is incorporated herein by reference in its entirety. In some embodiments, the Cas9 variant does not have specific PAM requirements. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, comprises an NRNH PAM (in the sequence R is A or G, and H is A, C, or T. In some embodiments, the SpCas9 variant comprises a PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, or C. In some embodiments, the SpCas9 variant has specificity for AC. 1134, 1135, 1137, 1339, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 12 64, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339 In some embodiments, the SpCas9 variant has an amino acid substitution at the corresponding position. Ant is at positions 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, Contains an amino acid substitution at positions 1335, 1337, or corresponding thereto. In embodiments, the SpCas9 variants are located at positions 1114, 1134, 1135, 1137, 1339, 1151, 1180, 11 88, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333 or In some embodiments, the SpCas9 variant has an amino acid substitution at the corresponding position. Position 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253 , 1286, 1293, 1320, 1321, 1332, 1335, 1339, or amino acid substitutions at the corresponding positions In some embodiments, the SpCas9 variant has positions 1114, 1227, 1135, 1180, , 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349 or their corresponding Exemplary amino acid substitutions and PAM specificities of SpCas9 variants include amino acid substitutions at positions Shown in Tables 1A to 1D. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]

[0243] In some embodiments, the Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or its derivatives. In some embodiments, NmeCas9 is a variant of the NNNNGAYW PAM. In some embodiments, Y is C or T and W is A or T. In this example, NmeCas9 has specificity for the NNNNGYTT PAM, where Y is C or T. In some embodiments, NmeCas9 has specificity for the NNNNGTCT PAM. In embodiments, NmeCas9 is Nme1 Cas9. In some embodiments, NmeCas9 is NNN NGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM , NNNNCCGGPAM, NNNCCCA PAM, NNNNCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG In some embodiments, the PAM has specificity for a PAM, a NNNNCCAT PAM, or a NNNGATT PAM. Nme1Cas9 has NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, or NN In some embodiments, NmeCas9 has specificity for the CAA PAM, CA In some embodiments, NmeCas9 has specificity for Nm e2 Cas9. In some embodiments, NmeCas9 is specific for NNNNCC (N4CC) PAM. In some embodiments, N is any one of A, G, C, or T. NmeCas9 has NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, and NNNNCCCC PAM. Specificity for AM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM, or NNNGATT PAM In some embodiments, the NmeCas9 is Nme3Cas9. NmeCas9 has specificity for NNNNCAAA, NNNNCC, or NNNNCNNN PAMs Additional NmeCas9 sequences are described in Edraki et al. Mol. Cell. (2019) 73 (4): 714-726. The properties and PAM sequences are incorporated herein by reference in their entirety. An exemplary amino acid sequence of Nme1Cas9 is shown below: Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002235162.1 JPEG2025170240000023.jpg131167

[0244] An exemplary amino acid sequence of Nme2Cas9 is shown below: Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002230835.1 JPEG2025170240000024.jpg130170

[0245] Nucleobase editor Cas12 domain Microbial CRISPR-Cas systems are generally divided into class 1 and class 2 systems. Class 1 systems are Class 2 systems have a multi-subunit effector complex, whereas class 3 systems have a single protein effector For example, Cas9 and Cpf1 have different types (type II and type V, respectively). In addition to Cpf1, class 2 V-type CRISPR-Cas systems also include Cas12a / Cpfl , Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and For example, Shmakov et al., “Discovery and Functional Cha acterization of Diverse Class 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 6 0 (3): 385-397; Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1 (5): 325-336; and Yan et al. , “Functionally Diverse Type V CRISPR-Cas Systems, “Science, 2019 Jan. 4; 363: See, e.g., pp. 88-91, the entire contents of each of which are incorporated herein by reference. The protein contains a RuvC (or RuvC-like) endonuclease domain. Mature CRISPR Although the production of RNA (crRNA) is generally tracrRNA-independent, for example, Cas12b / C2c1 tracrRNA is required for crRNA production. Cas12b / C2c1 mediates DNA cleavage by both crRNA and tracrRNA. Depends on.

[0246] The nucleic acid programmable DNA binding proteins contemplated by the present invention include those of class 2 type V. Cas proteins (Cas12 proteins) classified as Cas class 2 and V proteins Non-limiting examples of nucleotides include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, and Cas1 2e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ homologs, or modified versions thereof As used herein, Cas12 proteins include Cas12 nucleases, In some embodiments, the Cas12 domain may be referred to as a Cas12 protein domain. The Cas12 protein of the present invention can be used to bind to internally fused proteins such as deaminase domains. It contains an amino acid sequence interrupted by a protein domain.

[0247] In some embodiments, the Cas12 domain is a nuclease-inactive Cas12 domain or is a Cas12 nickase. In some embodiments, the Cas12 domain has nuclease activity. For example, the Cas12 domain is a functional domain that binds double-stranded nucleic acids (e.g., double-stranded DNA molecules). In some embodiments, the Cas12 domain may be a Cas12 domain that nicks a single strand. In some embodiments, the amino acid comprises any one of the amino acid sequences described herein. , the Cas12 domain may comprise at least one of the amino acid sequences described herein. 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, At least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the amino acid sequence is at least 99%, or at least 99.5% identical. In one embodiment, the Cas12 domain has, compared to any one of the amino acid sequences described herein: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 2 3, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 4 The amino acid sequence may contain 3, 44, 45, 46, 47, 48, 49, 50 or more mutations. In some embodiments, the Cas12 domain comprises any one of the amino acid sequences described herein. Compared to at least 10, at least 15, at least 20, at least 30, at least 40 , at least 50, at least 60, at least 70, at least 80, at least 90, at least At least 100, at least 150, at least 200, at least 250, at least 300, at least 350 , at least 400, at least 500, at least 600, at least 700, at least 800, At least 900, at least 1000, at least 1100, or at least 1200 identical consecutive It comprises an amino acid sequence having amino acid residues.

[0248] In some embodiments, proteins comprising fragments of Cas12 are provided. For example, several In some embodiments, the protein comprises: (1) the gRNA binding domain of Cas12; or (2) the gRNA binding domain of Cas12. In some embodiments, the Cas12 domain comprises one of two domains: a DNA cleavage domain. Proteins containing Cas12 or fragments thereof are referred to as "Cas12 variants." Variants share homology with Cas12 or fragments thereof. For example, Cas12 variants At least about 70% identical, at least about 80% identical, at least about 90% identical, or at least about 100% identical to the wild-type Cas12. at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, At least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, the Cas12 variant has 1, 2, 3, 4, 5 , 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 , 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 In some embodiments, the amino acid sequence may have 47, 48, 49, 50 or more amino acid changes. Cas12 variants are fragments of Cas12 (e.g., the gRNA binding domain or DNA cleavage domain). so that the fragment is at least about 70% identical to the corresponding fragment of wild-type Cas12. , at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% Identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least In some embodiments, the fragment is about 99.5% identical, or at least about 99.9% identical. At least 30%, at least 35%, at least 40% of the amino acid length of the corresponding wild-type Cas12; At least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least or at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400 , 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150 , 1200, 1250, or at least 1300 amino acids in length.

[0249] In some embodiments, Cas12 contains one or more mutations that alter Cas12 nuclease activity. corresponding to, or containing a portion or the entirety of, a Cas12 amino acid sequence having a mutation. Such mutations include, for example, amino acid substitutions within the RuvC nuclease domain of Cas12. In some embodiments, the sequence is at least about 70% identical to wild-type Cas12, at least about 70% identical to wild-type Cas12, About 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, Cas12 variants that are at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to each other. In some embodiments, variants or homologs of the nucleotide sequence of the present invention are provided. about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids amino acids, about 75 amino acids, about 100 amino acids or more, or longer amino acid sequences The present invention provides a variant of Cas12 having the following structure:

[0250] In some embodiments, the Cas12 fusion proteins provided herein comprise a Cas12 protein. The full-length amino acid sequence of the protein, for example, one of the Cas12 sequences provided herein, but However, in other embodiments, the fusion proteins provided herein contain a full-length Cas12 sequence. Examples of suitable Cas12 domains herein include, but are not limited to, one or more fragments thereof. Exemplary amino acid sequences are provided, and additional suitable sequences for Cas12 domains and fragments are within the skill of the art. It will be clear.

[0251] Generally, class 2 V-type Cas proteins encode a single functional RuvC endonuclease enzyme. (See, e.g., Chen et al., "CRISPR-Cas12a target binding unleashes discriminate single-stranded DNase activity,” Science 360: 436-439 (2018) In some cases, the Cas12 protein is a variant Cas12b protein. (See Strecker et al., Nature Communications, 2019, 10 (1): Art. No.: 212 In one embodiment, the variant Cas12 polypeptide is a variant of the wild-type Cas12 protein. When compared with the amino acid sequence, 1, 2, 3, 4, 5 or more amino acids differ (e.g. In some instances, the amino acid sequence may be a sequence having deletions, insertions, substitutions, or fusions. Altered Cas12 polypeptides may contain amino acid changes (e.g., For example, a variant Cas1 2 indicates less than 50%, less than 40%, less than 30%, or less than the nickase activity of the corresponding wild-type Cas12b protein. Cas12b polypeptides having less than 20%, less than 10%, less than 5%, or less than 1%. In this case, the variant Cas12b protein does not have substantial nickase activity.

[0252] In some cases, the variant Cas12b protein has reduced nickase activity. For example, variant Cas12b proteins may suppress the nickase activity of wild-type Cas12b proteins. This represents less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the total.

[0253] In some embodiments, the Cas12 protein is Cas12a / , which is active in mammalian cells. Contains RNA-guided endonucleases from the Cpf1 family. Prevotella and Francisella CRISPR derived from 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is , a class II CRISPR / Cas RNA-guided endonuclease. This adaptive immune mechanism is It is found in the bacteria Prevotella and Francisella 1. The Cpf1 gene is associated with the CRISPR locus. It is an endonuclease that uses a guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, and encodes CRISPR. Unlike Cas9 nuclease, Cpf1-mediated D The result of NA cleavage is a double-strand break with a short 3' overhang. The fragmentation pattern, similar to conventional restriction enzyme cloning, allows for the possibility of directional gene transfer. This can broaden the scope of gene editing and improve the efficiency of gene editing. Similarly, Cpf1 limits the number of sites that can be targeted by CRISPR to the number of NGGs preferred by SpCas9. It can also be extended to AT-rich regions or AT-rich genomes lacking PAM sites. The locus contains a mixed α / β domain, RuvC-I followed by a helical region, RuvC-II, and a zinc finger. The Cpf1 protein contains a RuvC domain similar to the RuvC domain of Cas9. Furthermore, unlike Cas9, Cpf1 possesses a HNH endonuclease domain. It does not have a cleavage domain, and the N-terminus of Cpf1 lacks the α-helical recognition lobe of Cas9. 1. The CRISPR-Cas domain architecture of Cpf1 is functionally unique and is a class 2, type V The Cpf1 locus contains Cas1, Cas2, and Cas4 proteins. It encodes a protein that is more similar to types I and III than to type II systems. It does not require activated CRISPR RNA (tracrRNA), and therefore only CRISPR (crRNA) is required. This is because not only is Cpf1 smaller than Cas9, but the sgRNA molecule is also small (approximately half the size of Cas9). The Cpf1-crRNA complex is a target of Cas9. In contrast to G-rich PAMs, the protospacer adjacent motifs 5'-YTN-3' or 5'-T Cpf1 cleaves the target DNA or RNA by identifying TTN-3'. After identifying the PAM, Cpf1 cleaves the target DNA or RNA. or introduce sticky end-like DNA double-strand breaks with 5-nucleotide overhangs.

[0254] In some embodiments of the invention, CRISPR enzymes that are mutated relative to the corresponding wild-type enzyme (The result is that the mutated CRISPR enzyme encodes a target polynucleotide containing the target sequence.) Vectors (which lack the ability to cleave one or both strands of the nucleotide) can be used. Cas12 is an exemplary wild-type Cas12 polypeptide (e.g., Cas12 from Bacillus hisashii). ) at least (or at least approximately) 50%, 60%, 70%, 80%, 90%, 91%, 92%, 9 3%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology Cas12 may refer to a polypeptide having a wild-type exemplary Cas12 polypeptide (e.g., For example, Bacillus hisashii (BhCas12b), Bacillus sp. V3-13 (BvCas12b), and Alicyclobacill us acidiphilus (AaCas12b)), up to or up to approximately 50%, 60%, 70%, 80%, and 90% , 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or can refer to polypeptides with sequence homology. Cas12 can be used to identify deletions, insertions, substitutions, variants, The amino acid sequence may include amino acid alterations such as substitutions, mutations, fusions, chimeras, or any combination thereof. It may refer to wild-type or modified forms of the Cas12 protein.

[0255] In some embodiments, the BhCas12b guide polynucleotide has the following sequence: BhCas12b sgRNA scaffold (underlined) + 20nt–23nt guide sequence (N n (denoted by JPEG2025170240000025.jpg20170

[0256] In some embodiments, the BvCas12b and AaCas12b guide polynucleotides have the following sequences: Having columns: BvCas12b sgRNA scaffold (underlined) + 20nt–23nt guide sequence (N n (denoted by JPEG2025170240000026.jpg19170AaCas12b sgRNA scaffold (underlined) + 20nt–23nt guide sequence (N n (denoted by JPEG2025170240000027.jpg22169

[0257] Nucleic acid programmable DNA binding proteins Some embodiments of the present disclosure act as nucleic acid programmable DNA binding proteins. Fusion proteins containing domains that specifically target proteins such as base editors are provided. The method can be used to guide a specific nucleic acid (e.g., DNA or RNA) sequence. In this embodiment, the fusion protein comprises a nucleic acid programmable DNA binding protein domain and a DNA Non-limiting examples of nucleic acid programmable DNA binding proteins include: Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12 Cas enzymes include Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ. Non-limiting examples include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12) Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12 e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2 , Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5 , Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx 16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst 2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V C Cas effector proteins, type VI Cas effector proteins, CARF, DinG, and their interactions Other nucleic acid programs are possible. DNA binding proteins capable of binding to the nucleotides of the present invention are also within the scope of this disclosure, although they are not specifically mentioned in this disclosure. For example, Makarova et al. "Classification and Nomencl Warning 10ature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR- 2019 Jan 4; 363 (6422): 88-91. doi: 10.1126 / science.aav72 71, the entire contents of each of which are incorporated herein by reference.

[0258] An example of a nucleic acid programmable DNA binding protein with a PAM specificity different from Cas9 is Clustered Regularly Interspaced Short P from Prevotella and Francisella 1 (Cpf1) Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with functions distinct from those of Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and targets T-rich proteins. It utilizes a rotospacer adjacent motif (TTN, TTTN, or YTN). It cleaves DNA through differential DNA double-strand breaks. It is composed of 16 Cpf1 family proteins. Two enzymes from Acidaminococcus and Lachnospiraceae are efficient genomic enzymes in human cells. Cpf1 protein has been shown to have genome editing activity. Previously, for example, Yamano et al., "Crystal structure of Cpf1 in complex with g guide RNA and target DNA.” Cell (165) 2016, pp. 949-962, and the entire contents are The contents of which are incorporated herein by reference.

[0259] used as a programmable DNA-binding protein domain with a guide nucleotide sequence Nuclease-inactive Cpf1 (dCpf1) variants that can be used in the present compositions and methods are also useful. The Cpf1 protein is a RuvC-like endonuclease enzyme similar to the RuvC domain of Cas9. Cpf1 has a nucleotide sequence but no HNH endonuclease domain, and the N-terminus of Cpf1 contains the α-terminal domain of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (referenced in Honmei) In the cleavage of both DNA strands, the RuvC-like domain of Cpf1 is involved in the cleavage of both DNA strands, and Ruv This indicates that inactivation of the C-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 In some embodiments, the dCpfl of the present disclosure comprises D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1 Any mutation that inactivates the RuvC domain of Cpf1, e.g., mutation corresponding to 255A. It is understood that substitution mutations, deletions, or insertions may be used in accordance with the present disclosure.

[0260] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The gram-competent DNA-binding protein (napDNAbp) may be the Cpf1 protein. In some embodiments, the Cpf1 protein is Cpf1 nickase (nCpf1). In some forms, the Cpf1 protein is nuclease-inactive Cpf1 (dCpf1). In embodiments, Cpf1, nCpf1, or dCpf1 has at least one amino acid sequence similar to that of the Cpf1 sequences disclosed herein. at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least At least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% In some embodiments, the amino acid sequence is at least 99.5% identical to the amino acid sequence of dCpf1 is at least 85%, at least 90%, or at least 100% identical to the Cpf1 sequences disclosed herein. at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% agreement and the amino acid sequences D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A are identical. , E1006A / D1255A, or D917A / E1006A / D1255A mutations. It should be understood that Cpf1 may also be used in accordance with the present disclosure.

[0261] Wild-type Francisella novicida Cpf1 (D917, E1006, and D1255) are underlined in bold. (I'm JPEG2025170240000028.jpg162169

[0262] Francisella novicida Cpf1 D917A (A917, E1006, and D1255 are underlined in bold) (I'm JPEG2025170240000029.jpg164170

[0263] Francisella novicida Cpf1 E1006A (D917, A1006, and D1255 are underlined in bold) (I'm JPEG2025170240000030.jpg164169

[0264] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are underlined in bold) (I'm JPEG2025170240000031.jpg167170

[0265] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are in bold and underlined) being pulled) JPEG2025170240000032.jpg167167

[0266] Francisella novicida Cpf1 D917A / D1255A (A917, E1006, and A1255 are in bold and underlined) being pulled) JPEG2025170240000033.jpg162169

[0267] Francisella novicida Cpf1 E1006A / D1225A (D917, A1006, and A1255 are in bold and underlined) is drawn) JPEG2025170240000034.jpg167170

[0268] Francisella novicida Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are in bold) (underlined in JPEG2025170240000035.jpg164169

[0269] In some embodiments, one of the Cas9 domains present in the fusion protein is fused to a PAM sequence. DNA-binding protein domains that can be programmed with guide nucleotide sequences without the need for may be replaced.

[0270] In some embodiments, the Cas9 domain is CaCas9 from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCa s9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, SaCas9 contains an N579A mutation, or an amino acid sequence as provided herein. The present invention includes corresponding mutations in either the amino acid sequence.

[0271] In some embodiments, a SaCas9 domain, a SaCas9d domain, or a SaCas9n domain can bind to nucleic acid sequences with non-canonical PAMs. The as9 domain, SaCas9d domain, or SaCas9n domain is not included in the NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain can bind to a nucleic acid sequence having a sequence. The amino acid sequence may contain one or more of the E781X, N967X, and R1014X mutations, or the amino acid sequence provided herein. and X is any amino acid. In some embodiments, the SaCas9 domain contains one or more of the following mutations: E781K, N967K, and R1014H; or one or more corresponding mutations in any of the amino acid sequences provided herein In some embodiments, the SaCas9 domain comprises an E781K, N967K, or R1014H mutation. or the corresponding mutation in any of the amino acid sequences provided herein.

[0272] Exemplary SaCas9 Sequences JPEG2025170240000036.jpg133169The residue N579 shown above, underlined and in bold, was mutated (e.g., to A579) to obtain Sa Cas9 nickase may be obtained.

[0273] Exemplary SaCas9n Sequences JPEG2025170240000037.jpg131169Residue N579 can be mutated to obtain SaCas9 nickase. Lines are drawn and shown in bold.

[0274] Exemplary SaKKH Cas9 JPEG2025170240000038.jpg121159Residue N579 can be mutated to obtain SaCas9 nickase. The lines are underlined and shown in bold. E781, N967, and R1014 were mutated to create SaKKH Cas9. The residues K781, K967, and H1014 obtained above are underlined and italicized. are.

[0275] In some embodiments, the napDNAbp is a circular permutant. In the sequence, the plain text indicates the adenosine deaminase sequence, and the bolded sequence is derived from Cas9. The sequences in italics represent linker sequences, and the underlined sequences represent bipartite sequences. Nuclear localization sequences are shown, and double underlined sequences indicate mutations. CP5 (with MSP "NGC" PID and "D10A" nickase): JPEG2025170240000039.jpg180169

[0276] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) comprises: It is the single effector of the microbial CRISPR-Cas system. Examples of such proteins include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Microbial CRISPR-Cas systems are generally divided into class 1 and class 2 systems. Class 1 systems are multisubsystems. Class 1 systems have a single protein effector complex, whereas Class 2 systems have a single protein effector For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three Different class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) are described in Shmakov et al., “Di discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60 (3): 385-397 (the entire contents of which are incorporated by reference). The two effectors of the system, Cas12b / C2c1 and Cas12c / C2c3, are Cpf The third system contains a RuvC-like endonuclease domain related to 1. EPN contains an effector with an RNase domain. Mature CRISPR RNA is generated by Cas12b / C2c1. Unlike CRISPR RNA production by tracrRNA, Cas12b / C2c1 does not depend on tracrRNA. This depends on both CRISPR RNA and tracrRNA.

[0277] Alicyclobaccillus acidot complexed with chimeric single-molecule guide RNA (sgRNA) The crystal structure of A. errastris Cas12b / C2c1 (AacC2c1) has been reported. For example, Liu et al. “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. See Cell, 2017 Jan. 19; 65 (2): 310-322, the entire contents of which are incorporated herein by reference. The crystal structure shows Alicyclobacillus bound to target DNA as a ternary complex. For example, Yang et al., "PAM-dependent Tar get DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 D ec. 15; 167(7): 1814-1828, the entire contents of which are incorporated herein by reference. Both the target and non-target DNA strands are independently located within a single RuvC catalytic pocket. The catalytically competent conformation of AacC2c1 was captured, in which Cas12b / C2c1-mediated cleavage results in a staggered cut of the target DNA at seven nucleotides. Cas12b / C2c1 Structural comparison of the ternary complex with previously identified Cas9 and Cpf1 counterparts in the CRISPR-Cas9 system This shows the diversity of mechanisms used.

[0278] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The grammable DNA-binding protein (napDNAbp) binds to the Cas12b / C2c1 or Cas12c / C2c3 proteins. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In this state, napDNAbp binds at least At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 9 4%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. The apDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp has at least one sequence similar to any one of the napDNAbp sequences provided herein. at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least At least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or an amino acid sequence that is at least 99.5% identical to Cas12b / from other bacterial species. It should be understood that C2c1 or Cas12c / C2c3 may also be used in accordance with the present disclosure. be.

[0279] Cas12b / C2c1 ((uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|C2C1_ALIAG CRISPR-related enzyme Denuclease C2c1 OS = Alicyclobacillus acido-terrestris (strain ATCC49025 / DSM3922 / CIP 106132 / NCIMB13137 / GD3B)GN=c2c1 PE=1 SV=1)The amino acid sequence is as follows: MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMVNQRIEGYLVKQIRSRVPLQD 2019 AacCas12b (Alicyclobacillus acidiphilus)-WP_067623834 MAVKSMKVKLRLDNMPEIRAGLWKLHTEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECYKTAEECKAELLERLRARQ VENGHCGPAGSDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGGLGIAKAGNKPRWVRMREAGEPGWE EEKAKAEARKSTDTRDADVLRALADFGLKPLMRVYTDSDMSSVQWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGE AYAKLVEQKSRFEQKNFVGQEHLVQLVNQLQQDMKEASHGLESKEQTAHYLTGRALRGSDKVFEKWEKLDPDAPFDLYDT EIKNVQRRNTRRFGSHDLFAKLAEPKYQALWREDASFLTRYAVYNSIVRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGEGRHAIRFQKLLTVEDGVAKEVDDVTVPISMSAQLDDLLPRDPHELVALYFQDYGAEQHLAGEFGGAK IQYRRDQLNHLHARRGARDVYLNLSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSEGRVPFCPFIEGNENLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPMDANQMTPDWREAFEDELQKLKSLYGICGDREWTEAVYESVRR VWRHMGKQVRDWRKDVRSGERKPIRGYQKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREHI DHAKEDRLKKLADRIIMEALGYVYALDDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELL NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCAREQNPEPFPWWLNKFVAEHKLDGCPLRADDLIPTGEGEF VSPFSAEEGDFHQIHADLNAAQNLQRRLWSDFDISQIRRLCDWGEVDGEPVLIPRTTGKRTADSYGNKVFYTKTGVTYYE RERGKKRRKVFAQEEELSEEEAELLVEADEAREKSVVLMRDPSGIINRGDWTRQKEFWSMVNQRIEGYLVKQIRSRVRLQE SACENTGDI BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515 MAPKKKRKVGIHGVPAAATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKVSKAE IQAELWDFVLKMQKCNSFTHEVDKDEVFNILRELYEELVPSSVEKKGEANQLSNKFLYPLVDPNSQSGKGTASSGRKPRW YNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAEYGLIPLFIPYTDSNEPIVKEIKKWMEKSRNQSVRRLDKDMFIQAL ERFLSWESWNLKVKEEYEKVEKEYKTLEERIKEDIQALKALEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREIIQKWL KMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYATFCEIDKKKKDAKQQATFTLADPIN HPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTVQLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKH AFTYKDESIKFPLKGTLGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKVVNFKPKE LTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLFFPIKGTELYAVHRASFNIKLPGET LVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFEDITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWV AFLKQLHKRLEVEIGKEVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQLNHL NALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERSRFENSKLMKWSRREIPRQVALQGE IYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKLQDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDR KCVTTHADINAAQNLQKRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWVNAGKLKIK KGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFGKLERILISKLTNQYSISTIEDDSS KQSMKRPAATKKAGQAKKKK

[0280] In some embodiments, Cas12b is a variant of BhCas12b, and compared to BhCas12b BvCas12b (V4) contains the following changes: S893R, K846R, and E837G. BhCas12b (V4) , represented as follows: 5'mRNA Cap---5'UTR---bhCas12b---STOP sequence---3'UTR---12 0 polyA tail. 5'UTR: GGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGAGCCACC 3'UTR (TriLink standard UTR) GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCG TGGTCTTTGAATAAAGTCTGA Nucleic acid sequence of bhCas12b (V4) ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCCGCCACCAGATCCTTCATCCTGAAGATCGA GCCCAACGAGGAAGTGAAGAAAGGCCTCTGGAAAACCCACGAGGTGCTGAACCACGGAATCGCCTACTACATGAATATCC TGAAGCTGATCCGGCAAGAGGCCATCTACGAGCACCACGAGCAGGACCCCAAGAATCCCAAGAAGGTGTCCAAGGCCGAG ATCCAGGCCGAGCTGTGGGATTTCGTGCTGAAGATGCAGAAGTGCAACAGCTTCACACACGAGGTGGACAAGGACGAGGT GTTCAACATCCTGAGAGAGCTGTACGAGGAACTGGTGCCCAGCAGCGTGGAAAAGAAGGGCGAAGCCAACCAGCTGAGCA ACAAGTTTCTGTACCCTCTGGTGGACCCCAACAGCCAGTCTGGAAAGGGAACAGCCAGCAGCGGCAGAAAGCCCAGATGG TACAACCTGAAGATTGCCGGCGATCCCTCCTGGGAAGAAGAGAAGAAGAAGTGGGGAAGAAGAATAAGAAAAAGGACCCGCT GGCCAAGATCCTGGGCAAGCTGGCTGAGTACGGACTGATCCCTCTGTTCATCCCCTACACCGACAGCAACGAGCCCATCG TGAAAGAAATCAAGTGGATGGAAAGTCCCGGAACCAGAGCGTGCGGCGGCTGGATAAGGACATGTTCATTCAGGCCCTG GAACGGTTCCTGAGCTGGGAGAGCTGGAACCTGAAAGTGAAAGAGGAATACGAGAAGGTCGAGAAAGAGTACAAGACCCT GGAAGAGAGGATCAAAGAGGACATCCAGGCTCTGAAGGCTCGGAACAGTATGAGAAGAGCGGCAAGAACAGCTGCTGC GGGACACCCTGAACACCAACGAGTACCGGCTGAGCAAGAGAGGCCTTAGAGGCTGGCGGGAAATCATCCAGAAATGGCTG AAAATGGACGAGAACGAGCCCTCCGAGAAGTACCTGGAAGTTGTTCAAGGACTACCAGCGGAAGCACCCTAGAGAGGCCCGG CGATTACAGCGTGTACGAGTTCCTGTCCAAGAAAGAGAACCACTTCATCTGGCGGAATCACCCTGAGTACCCCTACCTGT ACGCCACCTTCTGCGAGATCGACAAGAAAAAGAAGGACGCCAAGCAGCAGGCCACCTTCACACTGGCCGATCCTATCAAT CACCCTCTGTGGGTCCGATTCGAGGAAAGAAGCGGCAGCAACCTGAACAAGTACAGAATCCTGACCGAGCAGCTGCACAC CGAGAAGCTGAAGAAAAAGCTGACAGTGCAGCTGGACCGGCTGATCTACCCTACAGAATCTGGCGGCTGGGAAGAGAAGG GCAAAGTGGACATTGTGCTGCTGCCCAGCCGGCAGTTCTACAACCAGATCTTCCTGGACATCGAGGAAAAGGGCAAGCAC GCCTTCACCTACAAGGATGAGAGCATCAAGTTCCCTCTGAAGGGCACACTCGGCGGAGCCAGAGTGCAGTTCGACAGAGA TCACCTGAGAAGATACCCTCACAAGGTGGAAAGCGGCAACGTGGGCAGAATCTACTTCAACATGACCGTGAACATCGAGC CTACAGAGTCCCCAGTGTCCAAGTCTCTGAAGATCCACCGGGACGACTTCCCCAAGGTGGTCAACTTCAAGCCCAAAGAA CTGACCGAGTGGATCAAGGACAGCAAGGGCAAGAAACTGAAGTCCGGCATCGAGTCCCTGGAAATCGGCCTGAGAGTGAT GAGCATCGACCTGGGACAGAGACAGGCCGCTGCCGCCTCTATTTTCGAGGTGGTGGATCAGAAGCCCGACATCGAAGGCA AGCTGTTTTTCCCAATCAAGGGCACCGAGCTGTATGCCGTGCACAGAGCCAGCTTCAACATCAAGCTGCCCGGCGAGACA CTGGTCAAGAGCAGAGAAGTGCTGCGGAAGGCCAGAGAGGACAATCTGAAACTGATGAAACCAGAAGCTCAACTTCCTGCG GAACGTGCTGCACTTCCAGCAGTTCGAGGACATCACCGAGAGAGAGAAGCGGGTCACCAAGTGGATCAGCAGCAAGAGA ACAGCGACGTGCCCCTGGTGTACCAGGATGAGCTGATCCAGATCCGCGAGCTGATGTACAAGCCTTACAAGGACTGGGTC GCCTTCCTGAAGCAGCTCCACAAGAGACTGGAAGTCGAGATCGGCAAAGAAGTGAAGCACTGGCGGAAGTCCCTGAGCGA CGGAAGAAAGGGCCTGTACGGCATCTCCCTGAAGAACATCGACGAGATCGATCGGACCCGGAAGTTCCTGCTGAGATGGT CCCTGAGGCCTACCGAACCTGGCGAAGTGCGTAGACTGGAACCCGGCCAGAGATTCGCCATCGACCAGCTGAATCACCTG AACGCCCTGAAAGAAGATCGGCTGAAGAAGATGGCCAACACCATCATCATGCACGCCCTGGGCTACTGCTACGACGTGCG GAAGAAGAAATGGCAGGCTAAGACCCCGCCTGCCAGATCATCCTGTTCGAGGATCTGAGCAACTACAACCCCTACGAGG AAAGGTCCCGCTTCGAGAACAGCAAGCTCATGAAGTGGTCCAGACGCGAGATCCCCAGACAGGTTGCACTGCAGGGCGAG ATCTATGGCCTGCAAGTGGGAGAAGTGGGCGCTCAGTTCAGCAGCAGATTCCACGCCAAGACAGGCAGCCCTGGCATCAG ATGTAGCGTCGTGACCAAAGAGAGCTGCAGGACAATCGGTTCTTCAAGAATCTGCAGAGAGAGGGCAGACTGACCCTGG ACAAAATCGCCGTGCTGAAAGAGGGCGATCTGTACCCAGACAAAAGGCGGCGAGAAGTTCATCAGCCTGAGCAAGGATCGG AAGTGCGTGACCACACACGCCGACATCAACGCCGCTCGAACCTGCAGAAGCGGTTCTGGACAAGAACCCACGGCTTCTA CAAGGTGTACTGCAAGGCCTACCAGGTGGACGGCCAGACCGTGTACATCCCTGAGAGCAAGGACCAGAAGCAGAAGATCA TCGAAGAGTTCGGCGAGGGCTACTTCATTCTGAAGGACGGGGTGTACGAATGGGTCAACGCCGGCAAGCTGAAAATCAAG AAGGGCAGCTCCAAGCAGAGCAGCAGCGAGCTGGTGGATAGCGACATCCTGAAAGACAGCTTCGACCTGGCCCTCCGAGCT GAAAGGCGAAAAGCTGATGCTGTACAGGGACCCCAGCGGCAATGTGTTCCCCAGCGACAAATGGATGGCCGCTGGCGTGT TCTTCGGAAAGCTGGAACGCATCCTGATCAGCAAGCTGACCAACCAGTACTCCATCAGCACCATCGAGGACGACAGCAGC AAGCAGTCTATGAAAAGGCCGGCGGCCACGAAAAAGGCCGGCCAGGCAAAAAAAGAAAAAG

[0281] In some embodiments, Cas12b is BvCas12B. In some embodiments, Cas12b represents the amino acid substitutions S893R, K846R, and S893R, as numbered in the exemplary sequence of BvCas12b shown below: and E837G. BvCas12b (Bacillus sp. V3-13) NCBI reference sequence: WP_101661451.1 MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEH GSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEER KAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTE SYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESAPEELWKVVAEQQNKMS EGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEK QKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIK NHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLLEG MRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQ IRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQI VSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIM TALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEY SSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVI HADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPKSQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTDTT FDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL

[0282] In some embodiments, the Cas12b is BTCas12b. lovorans)NCBI reference sequence: WP_041902512 MATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKV SKAEIQAELWDFVLKMQKCNSFTHEVDKDVVFNILRELYEELVPSSVEKKGEANQLSNKF LYPLVDPNSQSGKGTASSGRKPRWYNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAE YGLIPLFIPFTDSNEPIVKEIKWMEKSRNQSVRRLDKDMFIQALERFLSWESWNLKVKEE YEKVEKEHKTLEERIKEDIQAFKSLEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREII QKWLKMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYAT FCEIDKKKKDAKQQATFTLADPINHPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTV QLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKHAFTYKDESIKFPLKGT LGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKFVNF KPKELTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLF FPIKGTELYAVHRASFNIKLPGETLVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFE DITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWVAFLKQLHKRLEVEIGK EVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQ LNHLNALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERS RFENSKLMKWSRREIPRQVALQGEIYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKL QDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDRKLVTTHADINAAQNLQ KRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWGNAGK LKIKKGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFG KLERILISKLTNQYSISTIEDDSSKQSM

[0283] In some embodiments, napDNAbp refers to Cas12c. The 2c protein is Cas12c1 or a variant of Cas12c1. The Cas12 protein is Cas12c2 or a variant of Cas12c2. The Cas12 protein is the Cas12c protein from Oleiphilus species HI0009 (i.e., OspCa s12c) or variants of OspCas12c. These Cas12c molecules are described in Yan et al., “Fun ctionally Diverse Type V CRISPR-Cas Systems,” Science, 2019 Jan. 4; 363: 88-91 the entire contents of which are incorporated herein by reference. In this state, napDNAbp binds to the native Cas12c1, Cas12c2, or OspCas12c protein. At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the amino acid sequence of the present invention is at least 99%, or at least 99.5%, identical to the amino acid sequence of the present invention. In the present specification, napDNAbp is a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp is any Cas12c1, Cas12c2, or Os DNAbp described herein. For pCas12c protein, at least 85%, at least 90%, at least 91%, at least At least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97% , an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to Cas12c1, Cas12c2, or OspCas12c from other bacterial species may also be used in accordance with the present disclosure. It should be understood that it can be used. Cas12c1 MQTKKTHLHLISAKASRKYRRTIACLSDTAKKDLERRKQSGAADPAQELSCLKTIKFKLEVPEGSKLPSFDRISQIYNAL ETIEKGSLSYLLFALILSGFRIFPNSSAAKTFASSSCYKNDQFASQIKEIFGEMVKNFIPSELESILKKGRRKNNKDWTE ENIKRVLNSEFGRKNSEGSSALFDSFLSKFSQELFRKFDSWNEVNKKYLEAAELLDSMLASYGPFDSVCKMIGDSDSRNS LPDKSTIAFTNNAEITVDIESSVMPYMAIAALLREYRQSKSKAAPVAYVQSHLTTTNGNGLSWFFKFGLDLIRKAPVSSK QSTSDGSKSLQELFSVPDDKLDGLKFIKEACEALPEASLLCGEKGELLGYQDFRTSFAGHIDSWVANYVNRLFELIELVN QLPESIKLPSILTQKNHNLVASLGLQEAEVSHSLELFEGLVKNVRQTLKKLAGIDISSSPNEQDIKEFYAFSDVLNRLGS IRNQIENAVQTAKKDKIDLESAIEWKEWKLKKLPKLNGLGGGVPKQQELLDKALESVKQIRHYQRIDFERVIQWAVNEH CLETVPKFLVDAEKKKINKESSTDFAAKENAVRFLLEGIGAAARGKTDSVVSKAAYNWFVVNNFLAKKDLNRYFINCQGCI YKPPYSKRRSLAFALRSDNKDTIEVVWEKFETFYKEISKEIEKFNIFSQEFQTFLHLENLRMKLLLRRIQKPIPAEIAFF SLPQEYYDSLPPNVAFLALNQEITPSEYITQFNLYSSFLNGNLILLRRSRSYLRAKFSWVGNSKLIYAAKEARLWKIPNA YWKSDEWKMILDSNVLVFDKAGNVLPAPTLKKVCEREGDLLRFPLLRQLPHDWCYRNPFVKSVGREKNVIEVNKEGEPK VASALPGSLFRLIGPAPFKSLLLDDCFFNPLDKDLRECMLIVDQEISQKVEAQKVEASLESCTYSIAVPIRYHLEEPKVSN QFENVLAIDQGEAGLAYAVFSLKSIGEAETKPIAVGTIRIPSIRRLIHSVSTYRKKKQRLQNFKQNYDSTAFIMRENVTG DVCAKIVGLMKEFNAFPVLEYDVKNLESGSRQLSAVYKAVNSHFLYFKEPGRDALRKQLWYGGDSWTIDGIEIVTRERKE DGKEGVEKIVPLKVFPGRRSVSARFTSKTCSCCCGRNVFDWLFTEKKAKTNKKFNVNSKGELTTADGVIQLFEADRSKGPKF YARRKERTPLTKPIAKGSYSLEEIERRVRTNLRRAPKSKQSRDTSQSQYFCVYKDCALHFSGMQADENAAINIGRRFLTA LRKNRRSDFPSNVKISDRLDN Cas12c2 MTKHSIPLHAFRNSGADARKWKGRIALLAKRGKETMRTLQFPLEMSEPEAAAINTTPFAVAYNAIEGTGKGTLFDYWAKL HLAGFRFFPSGGAATIFRQQAVFEDASWNAAFCQQSGKDWPWLVPSKLYERFTKAPREVAKKDGSKKSIEFTQENVANES HVSLVGASITDKTPEDQKEFFLKMAGALAEKFDSWKSANEDRIVAMKVIDEFLKSEGLHLPSLENIAVKCSVETKPDNAT VAWHDAPMSGVQNLAIGVFATCASRIDNIYDLNGGKLSKLIQESATTPNVTALSWLFGKGLEYFRTTDIDTIMQDFNIPA SAKESIKPLVESAQAIPTMTVLGKKNYAPFRPNFGKIDSWIANYASRLMLLNDILEQIEPGFELPQALLDNETLMSGID MTGDELKELIEAVYAWVDAAKQGLATLLGRGGNVDDAVQTFEQFSAMMDTLNGTLNTISARYVRAVEMAGKDEARLEKLI ECKFDIPKWCKSVPKLVGISGGLPKVEEEIKVMNAAFKDVRARMFVRFEEIAAYVASKGAGMDVYDALEKRELEQIKKLK SAVPERAHIQAYRAVLHRIGRAVQNCSEKTKQLFSSKVIEMGVFKNPSHLNNFIFNQKGAIYRSPFDRSRHAPYQLHADK LLKNDWLELLAEISATLMASESTEQMEDALRLERTRLQLQLSGLPDWEYPASLAKPDIEVEIQTALKMQLAKDTVTSDVL QRAFNLYSSVLSGLTFKLLRRSFSLKMRFSVADTTQLIYVPKVCDWAIPKQYLQAEGEIGIAARVVTESSPAKMVTEVEM KEPKALGHFMQQAPHDWYFDASLGGTQVAGRIVEKGKEVGKERKLVGYRMRGNSAYKTVLDKSLVGNTELSQCSMIIEIP YTQTVDADFRAQVQAGLPKVSINLPVKETITASNKDEQMLFDRFVAIDLGERGLGYAVVFDAKTLELQESGHRPIKAITNL LNRTHHYEQRPNQRQKFQAKFNVNLSELRENTVGDVCHQINRICAYYNAFPVLEYMVPDRLDKQLKSVYESVTNRYIWSS TDAHKSARVQFWLGGETWEHPYLKSAKDKKPLVLSPGRGASGKGTSQTCSCCGRNPFDLIKDMKPRAKIAVVDGKAKLEN SELKLFERNLESKDDMLARRHRNERAGMEQPLTPGNYTVDEIKALLRANLRRAPKNRRTKDTTVSEYHCVFSDCGKTMHA THESE ARE THE WAYS TO GO OspCas12c MTKLRHRQKKLTHDWAGSKKREVLGSNGKLQNPLLMPVKKGQVTEFRKAFSAYARATKGEMTDGRKNMFTHSFEPFKTKP SLHQCELADKAYQSLHSYLPGSLAHFLLSAHALGFRIFSKSGEATAFQASSKIEAYESKLASELACVDLSIQNLTISTLF NALTTSVRGKGEETSADPLIARFYTLLTGKPLSRDTQGPERDLAEVISRKIASSFGTWKEMTANPLQSLQFFEEELHALD ANVSLSPAFDVLIKMNDLQGDLKNRTIVFDPAPVFEYNAEDPADIIIKLTARYAKEAVIKNQNVGNYVKNAITTTNANG LGWLLNKGLSLLPVSTDDELLEFIGVERSHPSCHALIELIAQLEAPELFEKNVFSDTRESEVQGMIDSAVSNHIARLSSSR NSLSMDSEELERLIKSFQIHTPHCSLFIGAQSLSQQLESLPEALQSGVNSADILLGSTQYMLTNSLVEESIATYQRTLNR INYLSGVAGQINGAIKRKAIDGEKIHLPAAWSELISLPFIGQPVIDVESDLAHLKNQYQTLSNEFDTLISALQKNFDLNF NKALLNRTQHFEAMCRSTKKNALSKPEIVSYRDLLARLTSCLYRGSLVLRRAGIEVLKKHKIFESNSELREHVHERKHFV FVSPLDRKAKKLLRLTDSRPDLLHVIDEILQHDNLENKDRESLWLVRSGYLLAGLPDQLSSSFINLPIITQKGDRRLIDL IQYDQINRDAFVMLVTSAFKSNLSGLQYRANKQSFVVTRTLSPYLGSKLVYVPKDKDWLVPSQMFEGRFADILQSDYMVW KDAGRLCVIDTAKHLSNIKKSVFSSEEVLAFLRELPHRTFIQTEVRGLGVNVDGIAFNNGDIPSLKTFSNCVQVKVSRTN TSLVQTLNRWFEGGKVSPPSIQFERAYYKKDDQIHEDAAKRKIRFQMPATELVHASDDAGWTPSYLLGIDPGEYGMGLSL VSINNGEVLDSGFIHINSLINFASKKSNHQTKVVPRQQYKSPYANYLEQSKDSAAGDIAHILDRLIYKLNALPVFEALSG NSQSAADQVWTKVLSFYTWGDNDAQNSIRKQHWFGASHWDIKGMLRQPPTEKKPKPYIAFPGSQVSSYGNSQRCSCCGRN PIEQLREMAKDTSIKELKIRNSEIQLFDGTIKLFNPDPSTVIERRRHNLGPSRIPVADRTFKNISPSSLEFKELITIVSR SIRHSPEFIAKKRGIGSEYFCAYSDCNSSLNSEANAAANVAQKFQKQLFFEL

[0284] In some embodiments, napDNAbp refers to Cas12g, Cas12h, or Cas12i, For example, see Yan et al., “Functionally Diverse Type V CRISPR-Cas Systems,” Science nce, 2019 Jan. 4; 363: 88-91; the entire contents of each of which are incorporated herein by reference. By aggregating over 10 terabytes of sequence data, Cas12g, Ca It shows weak similarity to previously characterized class V proteins, including s12h and Cas12i. A novel class of type V Cas proteins has been identified. In some embodiments, the Cas12 protein The protein is Cas12g or a variant of Cas12g. In some embodiments, the Cas12 protein The protein is Cas12h or a variant of Cas12h. In some embodiments, the Cas12 protein The protein is Cas12i or a variant of Cas12i. It is understood that any of these may be used as a napDNAbp and are within the scope of the present disclosure. In some embodiments, the napDNAbp is directed against a naturally occurring Cas12g, Cas12h, or Cas12i protein. and at least 85%, at least 90%, at least 91%, at least 92%, at least 93% , at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, Some examples include sequences that are at least 99%, or at least 99.5%, identical to one another. In embodiments, the napDNAbp is a naturally occurring Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp is any Cas12g, Cas12h, or Cas1 gene described herein. For 2i protein, at least 85%, at least 90%, at least 91%, at least 92% , at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, containing an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to the Cas12g, Cas12h, or Cas12i from other bacterial species may also be used in accordance with the present disclosure. It should be understood that in some embodiments, Cas12i is Cas12i1 or Cas It is 12i2. Cas12g1 MAQASSTPAVSPRPRPRYREERTLVRKLLPRPGQSKQEFRENVKKLRKAFLQFNADVSGVCQWAIQFRPRYGKPAEPTET FWKFFLEPETSLPPNDSRSPEFRRLQAFEAAAGINGAAALDDPAFTNELRDSILAVASRPKTKEAQRLFSRLKDYQPAHR MILAKVAAEWIESRYRRAHQNWERNYEEWKKEKQEWEQNHPELTPEIREAFNQIFQQLEVKEKRVRICPAARLLQNKDNC QYAGKNKHSVLCNQFNEFKKNHLQGKAIKFFYKDAEKYLRCGLQSLKPNVQGPFREDWNKYLRYMNLKEETLRGKNGGRL PHCKNLGQECEFNPHTALCKQYQQQLSSRPDLVQHDELYRKWRREYWREPRKPVFRYPSVKRHSIAKIFGENYFQADFKN SVVGLRLDSMPAGQYLEFAFAPWPRNYRPQPGETEISSVHLHFVGTRPRIGFRFRVPHKRSRFDCTQEELDELRSRTFPR KAQDQKFLEAARKRLLETFPGNAEQELRLLAVDLGTDSARAAFFIGKTFQQAFPLKIVKIEKLYEQWPNQKQAGDRRDAS SKQPRPGLSRDHVGRHLQKMRAQASEIAQKRQELTGTPAPETTTDQAAKKATLQPFDLRGLTVHTARMIRDWARLNARQI IQLAEENQVDLIVLESLRGFRPPGYENLDQEKKRRVAFFAHGRIRRKVTEKAVERGMRVVTVPYLASSKVCAECRKKQKD NKQWENKRKRGLFKCEGCGSQAQVDENAARVLGRVFWGEIELPTAIP Cas12h1 MKVHEIPRSQLLKIKQYEGSFVEWYRDLQEDRKKFASLLFRWAAFGYAAREDDGATYISPSQALLERRLLLGDAEDVAIK FLDVLFKGGAPSSSCYSLFYEDFALRDKAKYSGAKREFIEGLATMPLDKIIERIRQDEQLSKIPAEEWLILGAEYSPEEI WEQVAPRIVNVDRSLGKQLRERLGIKCRRPHDAGYCKILMEVVARQLRSHNETYHEYLNQTHEMKTKVANNLTNEFDLVC EFAEVLEEKNYGLGWYVLWQGVKQALKEQKKPTKIQIAVDQLRQPKFAGLLTAKWRALKGAYDTWKLKKRLEKRKKAFPYM PNWDNDYQIPVGLTGLGVFTLEVKRTEVVVDLKEHGKLFCSHSHYFGDLTAEKHPSRYHLKFRHKLKLRKRDSRVEPTIG PWIEAALREITIQKKPNGVFYLGLPYALSHGIDNFQIAKRFFSAAKPDKEVINGLPSEMVVGAADLNLSNIVAPVKARIG KGLEGPLHALDYGYGELIDGPKILTPDGPRCGELISLKRDIVEIKSAIKEFKACQREGLTMSEETTTWLSEVESPSDSPR CMIQSRIADTSRRLNSFKYQMNKEGYQDLAEALRLLDAMDSYNSLLESYQRMHLSPGEQSPKEAKFDTKRASFRDLLRRR VAHTIVEYFDDCDIVFFEDLDGPSDSDSRNNALVKLLSPRTLLLYIRQALEKRGIGMVEVAKDGTSQNNPISGHVGWRNK QNKSEIYFYEDKELLVMDADEVGAMNILCRGLNHSVCPYSFVTKAPEKKNDEKKEGDYGKRVKRFLKDRYGSSNVRFLVA SMGFVTVTTKRPKDALVGKRLYYHGGELVTHDLHNRMKDEIKYLVEKEVLARRVSLSDSTIKSYKSFAHV Cas12i1 MSNKEKNASETRKAYTTKMIPRSHDRMKLLGNFMDYLMDGTPIFFELWNQFGGGIDRDIISGTANKDKISDDLLLAVNWF KVMPINSKPQGVSPSNLANLFQQYSGSEPDIQAQEYFASNFDTEKHQWKDMRVEYERLLAELQLSRSDMHHDLKLMYKEK CIGLSLSTAHYITSVMFGTGAKNNRQTKHQFYSKVIQLLEESTQINSVEQLASIILKAGDCSDSYRKLRIRCSRKGATPSI LKIVQDYELGTNHDDEVNVPSLIANLKEKLGRFEYECEWKCMEKIKAFLASKVGPYYLGSYSAMLENALSPIKGMTKNC KFVLKQIDAKNDIKYENEPFGKIVEGFFDSPYFESDTNVKWVLHPHHIGESNIKTLWEDLNAIHSKYEEDIASLSEDDKKE KRIKVYQGDVCQTINTYCEEVGKEAKTPLVQLLRYLYSRKDDIAVDKIIDGITFLSKKHKVEKQKINPVIQKYPSFNFGN NSKLLKGIISPKDKLKHNLKCNRNQVDNYIWIEIKVLNTKTMRWEKHHYALSSTRFLEEVYYPATSENPPDALAARFRTK TNGYEGKPALSAEQIEQIRSAPVGLRKVKKRQMRLEAARQQNLLPRYTWGKDFNINICKRGNNFEVTLATKVKKKKEKNY KVVLGYDANIVRKNTYAAIEAHANGDGVIDYNDLPVKPIESGFVTVESQVRDKSYDQLSYNGVKLLYCKPHVESRRSFLE KYRNGTMKDNRGNNIQIDFMKDFEAIADDETSLYYFNMKYCKLLQSSIRNHSSQAKEYREEIFELLRDGKLSVLKLSSLS NLSFVMFKVAKSLIGTYFGHLLKKPKNSKSDVKAPPITDEDKQKADPEMFALRLALEEKRLNKVKSKKEVIANKIVAKAL ELRDKYGPVLIKGENISDTTKKKGKKSSTNSFLMDWLARGVANKVKEMVMMHQGLEFVEVNPNFTSHQDPFVHKNPENTFR ARYSRCTPSELTEKNRKEILSFLSDKPSKRPTNAYYNEGAMAFLATYGLKKNDVLGVSLEKFKQIMANILHQRSEDQLLF PSRGGMFYLATYKLDADATSVNWNGKQFWVCNADLVAAYNVGLVDIQKDFKKK Cas12i2 MSSAIKSYKSVLRPNERKNQLLKSTIQCLEDGSAFFFKMLQGLFGGITPEIVRFSTEQEKQQQDIALWCAVNWFRPVSQD SLTHTIASDNLVEKFEEYYGGTASDAIKQYFSASIGESYYWNDCRQQYYDLCRELGVEVSDLTHDLEILCREKCLAVATE SNQNNSIISVLFGTGEKEDRSVKLRITKKILEAISNLKEIPKNVAPIQEIILNVAKATKETFRQVYAGNLGAPSTLEKFI AKDGQKEFDLKKLQTDLKKVIRGKSKERDWCCQEELRSYVEQNTIQYDLWAWGEMFNKAHTALKIKSTRNYNFAKQRLEQ FKEIQSLNNLLVVKKLNDFFDSEFFSGEETYTICVHHLGGKDLSKLYKAWEDDPADPENAIVVLCDDLKNNFKKEPIRNI LRYIFTIRQECSAQDILAAAKYNQQLDRYKSQKANPSVLGNQGFTWTNAVILPEKAQRNDRPNSLDLRIWLYLKLRHPDG RWKKHHIPFYDTRFFQEIYAAGNSPVDTCQFRTPRFGYHLPKLTDQTAIRVNKKHVKAAKTEARIRLAIQQGTLPVSNLK ITEISATINSKGQVRIPVKFDVGRQKGTLQIGDRFCGYDQNQTASHAYSLWEVVKEGQYHKELGCFVRFISSGDIVSITE NRGNQFDQLSYEGLAYPQYADWRKKASKFVSLWQITKKNKKKEIVTVEAKEKFDAICKYQPRLYKFNKEYAYLLRDIVRG KSLVELQQIRQEIFRFIEQDCGVTRLGSLSLSTLETVKAVKGIIYSYFSTALNASKNNPISDEQRKEFDPELFALLEKLE LIRTRKKKQKVERIANSLIQTCLENNIKFIRGEGDLSTTNNATKKKANSRSMDWLARGVFNKIRQLAPMHNITLFGCGSL YTSHQDPLVHRNPDKAMKCRWAAIPVKDIGDWVLRKLSQNLRAKNIGTGEYYHQGVKEFLSHYELQDLEEELLKWRSDRK SNIPCWVLQNRLAEKLGNKEAVVYIPVRGGRIYFATHKVATGAVSIVFDQKQVWVCNADHVAAANIALTVKGIGEQSSDE ENPDGSRIKLQLTS

[0285] Representative nucleic acid and protein sequences of base editors are as follows: BhCas12b GGSGGS-ABE8-Xten20 (P153) JPEG2025170240000040.jpg215169JPEG2025170240000041.jpg244164JPEG2025170240000042.jpg212169BhCas12b GGSGGS-ABE8-Xten20 (K255) JPEG2025170240000043.jpg30169JPEG2025170240000044.jpg243168JPEG2025170240000045.jpg234157JPEG2025170240000046.jpg133163BhCas12b GGSGGS-ABE8-Xten20 (D306) JPEG2025170240000047.jpg108169JPEG2025170240000048.jpg221157JPEG2025170240000049.jpg227153JPEG2025170240000050.jpg57165BhCas12b GGSGGS-ABE8-Xten20 (D980) JPEG2025170240000051.jpg193169JPEG2025170240000052.jpg229160JPEG2025170240000053.jpg214163BhCas12b GGSGGS-ABE8-Xten20 (K1019) JPEG2025170240000054.jpg232157JPEG2025170240000055.jpg229153JPEG2025170240000056.jpg147168

[0286] In the above sequences, the Kozak sequences are underlined in bold; the dashed line indicates the Kozak sequences. indicates an N-terminal nuclear localization signal (NLS) followed by a line; lowercase letters indicate a GGGSGGS linker; The line indicates the sequence encoding ABE8; the unmodified sequence encodes BhCas12b. ; double underline indicates Xten20 linker; single underline indicates C-terminal NLS; dotted underline GGAT CC indicates the GS linker; and the italic letters are the code for the 3× hemagglutinin (HA) tag. Represents an array.

[0287] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The grammable DNA binding protein (napDNAbp) may be a Cas12j / CasΦ protein. stomach. Cas12j / CasΦ is known as Pausch et al., “CRISPR-CasΦ from huge phages is a hypercom pact genome editor,” Science, 17 July 2020, Vol. 369, Issue 6501, pp. 333-337 No. 6,239,999, which is incorporated herein by reference in its entirety. In this state, napDNAbp is at least 85% more efficient than the native Cas12j / CasΦ protein. At least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least In some embodiments, the napDNAbp comprises an amino acid sequence that is 99.5% identical to a naturally occurring Cas In some embodiments, the napDNAbp is a nuclease-inactive 12j / CasΦ protein. Cas12j / CasΦ from other species may also be used in this study. It should be understood that the invention may be used in accordance with the teachings. An exemplary Cas12j / CasΦ amino acid sequence is as follows: >CasΦ-1 MADTPTLFTQFLRHHLPGQRFRKDILKQAGRILANKGEDATIAFLRGKSEESPPDFQPPVKCPIIACSRPLTEWPIYQAS VAIQGYVYGQSLAEFEASDPGCSKDGLLGWFDKTGVCTDYFSVQGLNLIFQNARKRYIGVQTKVTNRNEKRHKKLKRINA KRIAEGLPELTSDEPESALDETGHLIDPPGLNTNIYCYQQVSPKPLALSEVNQLPTAYAGYSTSGDDPIQPMVTKDRLSI SKGQPGYIPEHQRALLSQKKHRRMRGYGLKARALLVIVRIQDDWAVIDLRSLLRNAYWRRIVQTKEPSTITKLLKLVTGD PVLDATRMVATFTYKPGIVQVRSAKCLKNKQGSKLFSERYLNETVSVTSIDLGSNNLVAVATYRLVNGNTPELLQRFTLP SHLVKDFERYKQAHDTLEDSIQKTAVASLPQGQQTEIRMWSMYGFREAQERVCQELGLADGSIPWNVMTATSTILTDLFL ARGGDPKKCMFTSEPKKKKNSKQVLYKIRDRAWAKMYRTLLSKETREAWNKALWGLKRGSPDYARLSKRKEELARRCVNY TISTAEKRAQCGRTIVALEDLNIGFFHGRGKQEPGWVGLFTRKKENRWLMQALHKAFLELAHHRGYHVIEVNPAYTSQTC PVCRHCDPDNRDQHNREAFHCIGCGFRGNADLDVATHNIAMVAITGESLKRARGSVASKTPQPLAAE* >CasΦ-2 MPKPAVESEFSKVLKKHFPGERFRSSYMKRGGKILAAQGEEAVVAYLQGKSEEEPPNFQPPAKCHVVTKSRDFAEWPIMK ASEAIQRYIYALSTTERAACKPGKSSESHAAWFAATGVSNHGYSHVQGLNLIFDHTLGRYDGVLKKVQLRNEKARARLES INASRADEGLPEIKAEEEEAVATNETGHLLQPPGINPSFYVYQTISPQAYRPRDEIVLPPEYAGYVRDPNAPIPLGVVRNR CDIQKGCPGYIPEWQREAGTAISPKTGKAVTVPGLSPKKNRMRRYWRSEKEKAQDALLVTVRIGTDWVVIDVRGLLRNA RWRTIAPKDISNLALLDLFTGDPVIDVRRNIVTFTYLTDACGTYARKWTLKGKQTKATLDKLTATQTVALVAIDLGQTNP ISAGISRVTQENGALQCEPLDRFTLPDDLLKDISAYRIAWDRNEEELRARSVEALPEAQQAEVRALDGVSKETARTQLCA DFGLDPKRLPWDKMSSNTTFISEALLSNSVSRDQVFFTPAPKKGAKKAPVEVMRKDRTWARAYKPRLSVEAQKLKNEAL WALKRTSPEYLKLSRRKEELCRRSINYVIEKTRRRTQCQIVIPVIEDLNVRFFHGSGKRLPGWDNFFTAKKENRWFIQGL HKAFSDLRTHRSFYVFEVRPERTSITCPKCGHCEVGNRDGEAFQCLSCGKTCNADLDVATHNLTQVALTGKTMPKREEPR DAQGTAPARTKKASKCAPPAEREDQTPAQEPSQTS >CasΦ-3 MEKEITELTKIRREFPNKKFSSTDMKKAGKLLKAEGPDAVRDFLNSCQEIIGDFKPPVKTNIVSISRPFEEWPVSMVGRA IQEYFSLTKEELESVHPGTSSEDHKSFFNITGLSNYNYTSVQGLNLIFKNAKAIYDGTLVKANNKNKKLEKKFNEINHK RSLEGLPIITPDFEEPFDENGHLNNPPGINRNIYGYQGCAAKVFVPSKHKMVSLPKEYEGYNRDPNLSLAGFRNRLEIPE GEPGHVPWFQRMDIPEGQIGHVNKIQRFNFVHGKNSGKVKFSDKTGRVKRYHHSKYKDATKPYFLEESKKVSALDSILA IITIGDDWVVFDIRGLYRNVFYRELAQKGLTAVQLLDLFTGDPVIDPKKGVVTFSYKEGVVPVFSQKIVPRFKSRDTLEK LTSQGPVALLSVDLGQNEPVAARVCSLKNINDKITLDNSCRISFLDDYKKQIKDYRDSLDELEIKIRLEAINSLETNQQV EIRDLDVFSADRAKANTVDMFDIDPNLISWDSMSDARVSTQISDLYLKNGGDESRVYFEINNKRIKRSDYNISQLVRPKL SDSTRKNLNDSIWKLKRTSEEYLKLSKRKLELSRAVVNYTIRQSKLLSGINDIVIIDELDVKKKFNGRGIRDIGWDNFF SSRKENRWFIPAFHKAFSELSSNRGLCVIEVNPAWTSATCPDCGFCSKENRDGINFTCRKCGVSYHADIDVATLNIARVA VLGKPMSGPADRERLGDTKKPRVARSRKTMKRKDISNSTVEAMVTA* >CasΦ-4 MYSLEMADLKSEPSLAKLLRDRFPGKYWLPKYWLKLAEKKRLTGGEEAACEYMADKQLDSPPPNFRPPARCVILAKSRPF EDWPVHRVASKAQSFVIGLSEQGFAALRAAPPSTADARRDWLRSHGASEDDLMALEAQLLETIMGNAISLHGGVLKKIDN ANVKAAKRLSGRNEARLNKGLQELPPEQEGSAYGADGLLVNPPGLNLNIYCRKSCPKPVPKNTARFVGHYPGYLRDSDI LISGTMDRLTIIEGMPGHIPAWQREQGLVKPGGRRRRLSGSESNMRQKVDPSTGPRRSTRSGTVNRSNQRTGRNGDPLLV EIRMKEDWVLLDARGLLRNLRWRESKRGLSCDHEDLSLSGLLALFSGDPVIDPVRNEVVFLYGEGIIPVRSTKPVGTRQS KKLLERQASMGPLTLISCDLGQTNLIAGRASAISLTHGSLGVRSSVRIELDPEIIKSFERLRKDADRLETEILTAAKETL SDEQRGEVNSHEKDSPQTAKASLCRELGLHPPSLPWGQMGPSTTFIADMLISHGRDDDAFLSHGEFPTLEKRKKFDKRFC LESRPLLSSETRKALNESLWEVKRTSSEYARLSQRKKEMARRAVNFVVEISRRKTGLSNVIVNIEDLNVRIFHGGGKQAP GWDGFFRPKSENRWFIQAIHKAFSDLAAHHGIPVIESDPQRTSMTCPECGHCDSKNRNGVRFLCKGCGASMDADFDAACR NLERVALTGKPMPKPSTSCERLLSATTGKVCSDHSLSHDAIEKAS* >CasΦ-5 MSSLPTPLELLKQKHADLFKGLQFSSKDNKMAGKVLKKDGEEAALAFLSERGVSRGELPNFRPPAKTLVVAQSRPFEEFP IYRVSEAIQLYVYSLSVKELETVPSGSSTKKEHQRFFQDSSVPDFGYTSVQGLNKIFGLARGIYLGVITRGENQLQKAKS KHEALNKKRRASGEAETEFDPTPYEYMTPERKLAKPPGVNHSIMCYVDISVDEFDFRNPDGIVLPSEYAGYCREINTAIE KGTVDRLGHLKGGPGYIPGHQRKESTTEGPKINFRKGRIRRSYTALYAKRDSRRVRQGKLALPSYRHHMMRLNSNAESAI LAVIFFGKDWVVFDLRGLLRNVRRWRNLFVDGSTPSTLLGMFGDPVIDPKRGVVAFCYKEQIVPVVSKSITKMVKAPELLN KLYLKSEDPLVLVAIDLGQTNPVGVGVYRVMNASLDYEVVTRFALESELLREIESYRQRTNAFEAQIRAETFDAMTSEEQ EEITRVAFSASKAKENVCHRFGMPVDAVDWATMGSNTIHIAKWVMRHGDPSLVEVLEYRKDNEIKLDKNGVPKKVKLTD KRIANLTSIRLRFSQETSKHYNDTMWELRRKHPVYQKLSKSKADFSRRVVNSIIRRNHLVPRARIVFIIEDLKNLGKVF HGSGKRELGWDSYFEPKSENRWFIQVLHKAFSETGKHKGYYIIECWPNWTSCTCPKCSCCDSENRHGEVFRCLACGYTCN TDFGTAPDNLVKIATTGKGLPPGKPKRCKGSSKGKNPKIARSSETGVSVTESGAPKVKKSSPTQTSQSSSQSAP* >CasΦ-6 MNKIEKEKTPLAKLMNENFAGLRFPFAIIKQAGKKLLKEGELKTIEYMTGKGSIEPLPNFKPPVKCLIVAKRRDLKYFPI CKASCEIQSYVYSLNYKDFMDYFSTPMTSQKQHEEFFKKSGLNIEYQNVAGLNLIFNNVKNTYNGVILKVKNRNEKLKK AIKNNYEFEEIKTFNDDGCLINKPGINNVIYCFQSISPKILKNITHLPKEYNDYDCSVDRNIIQKYVSRLDIPESQPGHV PEWQRKLPEFNNTNNPRRRRKWYSNGRNISKGYSVDQVNQAKIEDSLLAQIKIGEDWIILDIRGLLRDNLNRRELISYKNK LTIKDVLGFFSDYPIIDIKNLVTFCYKEGVIQVVSQKSIGNKKSKQLLEKLIENKPIALVSIDLGQTNPVSVKISKLNK INKISIESFTYRFLNEEILKEIEKYRKDYDKLELKLINEA >CasΦ-7 MSNTAVSTREHMSNKTTPPSPLSLLLRAHFPGLKFESQDYKIAGKKLRDGGGPEAVISYLTGKGQ...

Claims

1. SEQ ID NO:1: 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 or a modification at an amino acid position selected from the group consisting of: Adenosine deaminase with corresponding modifications.

2. R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P in SEQ ID NO: 1 an alteration selected from the group consisting of 124W, T133K, D139L, D139M, C146R, and A158K; or The adenosine deaminase of claim 1, including the corresponding modification in another adenosine deaminase. Deaminase.

3. The V82T modification of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase, may also be used. The adenosine deaminase according to claim 1 or 2, further comprising:

4. 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and and 158, or another adenosine The adenosine deaminase according to any one of claims 1 to 3, comprising the corresponding modification in the deaminase. Deaminase.

5. The adenosine derivative according to any one of claims 1 to 4, comprising two or more of the above modifications. Minase.

6. The adenosine derivative of any one of claims 1 to 5, comprising three or more of the above modifications. Minase.

7. further comprising one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R; The adenosine deaminase according to any one of claims 1 to 6.

8. The adenosine deaminase may be selected from the group consisting of the following modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K_V82S+Y123H+Y147R+Q154R; Q71M_V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R The adenosine deaminase according to any one of claims 1 to 7, comprising any one of:

9. a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157; The adenosine deaminase according to any one of claims 1 to 8, comprising a C-terminal deletion starting from 。

10. A modification selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R The adenosine deaminase according to any one of claims 1 to 6, further comprising:

11. adenosine deaminase variants listed in Table 14, Table 18, or Figures 3A-3C. Item 7. The adenosine deaminase according to any one of Items 1 to 6.

12. A polynucleotide programmable DNA binding domain and SEQ ID NO:1: 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 or a modification at an amino acid position selected from the group consisting of: adenosine deaminase variants containing the corresponding modifications, and a fusion protein comprising a nucleotide sequence of ...

13. The adenosine deaminase variant is selected from the group consisting of R21N, R23H, E25F, N38G, L51 of SEQ ID NO: 1 W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A1 58K, or a corresponding modification in another adenosine deaminase.

13. The fusion protein of claim 12, comprising the following modification:

14. Polynucleotide-programmable DNA binding domain and R21N, R23H, E25F of SEQ ID NO: 1 , N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C1 46R, and A158K, or another adenosine deaminase at least one base that is an adenosine deaminase variant containing a corresponding modification in A fusion protein comprising an editor domain.

15. The adenosine deaminase variant may be the V82T modification of SEQ ID NO: 1, or another adenosine deaminase variant.

15. The method according to any one of claims 12 to 14, further comprising a corresponding modification in syndeaminase. The fusion protein described above.

16. A polynucleotide programmable DNA binding domain and the modified V82T of SEQ ID NO: 1, and R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, one or more modifications selected from the group consisting of D139L, D139M, C146R, and A158K, or another Adenosine deaminase variants containing corresponding modifications in adenosine deaminase and at least one base editor domain which is

17. The adenosine deaminase variant is selected from the group consisting of 21, 23, 25, 38, 51, 54, and 70 of SEQ ID NO:

1. , 71, 72, 73, 94, 124, 133, 139, 146, and 158. adenosine deaminase, or a corresponding modification in another adenosine deaminase, A fusion protein according to any one of claims 12 to 16.

18. 10. The adenosine deaminase variant of claim 1, wherein the adenosine deaminase variant comprises two or more of the modifications.

18. A fusion protein according to any one of 2 to 17.

19. 10. The adenosine deaminase variant of claim 1, wherein the adenosine deaminase variant comprises three or more of the modifications.

18. A fusion protein according to any one of 2 to 17.

20. The adenosine deaminase variant has the following modifications: Y147T, Y147R, Q154S, Y123H 20. The fusion protein of claim 12, further comprising one or more of: Q154R; and Q154R. Protein.

21. The adenosine deaminase variant may be selected from the group consisting of the following modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R The fusion protein according to any one of claims 12 to 20, comprising any one of:

22. The adenosine deaminase variant is selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, and 156.

21. Any of claims 12 to 20, comprising a C-terminal deletion starting at a residue selected from the group consisting of 157 and 158. A fusion protein according to any one of claims 1 to 4.

23. the base editor domain comprises a monomer of an adenosine deaminase variant; The adenosine deaminase monomer comprises R21N, R23H, E25F, N38G, L51W, and P of SEQ ID NO:

1. 54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and 21. The method according to any one of claims 12 to 20, comprising one or more modifications selected from the group consisting of: A158K; The fusion protein described above.

24. the base editor domain comprises a wild-type adenosine deaminase domain and an adenosine and a deaminase variant.

18. A fusion protein according to any one of claims 1 to 17.

25. The adenosine deaminase variant is selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166 25. The fusion protein of claim 24, further comprising a modification selected from the group consisting of Q154R, Q155R, and Q154R. Quality.

26. The base editor domain is a TadA*7.10 domain and an adenosine deaminase barrier domain.

18. The method of claim 12, comprising an adenosine deaminase heterodimer comprising a nucleotide sequence and a nucleotide sequence. The fusion protein according to any one of claims 1 to 4.

27. 27. The adenosine deaminase variant of claim 26, wherein the adenosine deaminase variant comprises two or more modifications. Fusion proteins.

28. the base editor comprising a TadA*7.10 domain and the following group of modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R and a heterodimer comprising an adenosine deaminase variant comprising any one of A fusion protein according to any one of claims 12 to 17.

29. The adenosine deaminase variant is selected from the group consisting of ABE9, ABE9B, ABE9C, ABE9D, ABE9E ... or a TadA*9 deaminase variant. protein.

30. The adenosine deaminase variant has 1, 2, 3, 4, 5, 6, or 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues 30. The fusion protein according to any one of claims 12 to 29, which is a truncated ABE8 or ABE9 lacking Plagiarism.

31. The polynucleotide-programmable DNA binding domain is Cas9, Cas12a / Cpf1, Cas1 2b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Ca The fusion protein of any one of claims 12 to 30, which is an s12j / CasΦ domain.

32. The following array: A fusion protein containing a polynucleotide containing a programmable DNA binding domain. In the sequences, the bolded sequences represent the Cas9-derived sequences, and the italicized sequences represent the linker sequences. The underlined sequence indicates the bipartite nuclear localization sequence, and at least one base The interdomains are 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 13 of SEQ ID NO: 1 Adenoviruses containing modifications at amino acid positions selected from the group consisting of 3, 138, 139, 146, and 158. The fusion protein comprises a syndeaminase variant.

33. The adenosine deaminase variant is selected from the group consisting of R21N, R23H, E25F, N38G, L51 of SEQ ID NO: 1 W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D138M, D139L, D139M, C146R 33. The fusion protein of claim 32, comprising a modification selected from the group consisting of A158K and A158K.

34. 34. The method of claim 33, wherein the adenosine deaminase variant comprises the modification V82T of SEQ ID NO:

1. The fusion protein described.

35. 3. The adenosine deaminase variant of claim 3, wherein the adenosine deaminase variant comprises two or more of the modifications.

35. The fusion protein according to claim 3 or 34.

36. 3. The adenosine deaminase variant of claim 3, wherein the adenosine deaminase variant comprises three or more of the modifications.

35. The fusion protein according to claim 3 or 34.

37. The adenosine deaminase variant is selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166 35. The fusion protein of claim 33 or 34, further comprising a modification selected from the group consisting of Q154R, Q155R, and Q154R. Synthetic protein.

38. The adenosine deaminase variant has the following modifications: Y147T, Y147R, Q154S, Y123H 35. The fusion protein of claim 33 or 34, comprising two or more of Q154R, and Q154R.

39. The adenosine deaminase variant may be selected from the group consisting of the following modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; or any other modification or group of modifications in Table 14.

3. The fusion protein according to claim 2.

40. The polynucleotide programmable DNA binding domain is selected from the group consisting of Staphylococcus aureus C as9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogene 40. The method according to any one of claims 12 to 39, wherein the nucleic acid sequence is a Cas9 (SpCas9), or a variant thereof. The fusion protein described above.

41. The polynucleotide programmable DNA binding domain is an engineered protospace Any of claims 12 to 40, comprising a modified SaCas9 with PAM specificity. The fusion protein according to any one of claims 1 to 4.

42. The modified SaCas9 contains the amino acid substitutions E782K, N968K, and R1015H, or 42. The fusion protein of claim 41, comprising a corresponding amino acid substitution.

43. The polynucleotide programmable DNA binding domain is an engineered protospace Any of claims 12 to 40, comprising a variant of SpCas9 with para-adjacent motif (PAM) specificity. A fusion protein according to any one of claims 1 to 4.

44. The modified PAM is selected from the group consisting of the nucleic acid sequences 5'-NGA-3', 5'-NGC-3', 5'-NGG-3', 5'-NGT-3' 44. The fusion protein of claim 43, having specificity for 5''-NGN-3', or 5''-NGN-3'.

45. wherein the variant SpCas9: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or any of these corresponding amino acid substitutions; I322V, S409I, E427G, R654L, R753G (MQKFRAER) or their corresponding amino acid substitutions; I322V, S409I, E427G, R654L, R753G, R1114G or their corresponding amino acid substitutions; or the amino acid substitutions set forth in Figures 3A-3C 45. The fusion protein of claim 43 or 44, comprising an amino acid substitution selected from:

46. The polynucleotide programmable DNA binding domain is nuclease inactive or The fusion protein of any one of claims 12 to 45, wherein is a nickase variant.

47. The nickase variant has the amino acid substitution D10A or a corresponding amino acid substitution.

47. The fusion protein of claim 46, comprising

48. The adenosine deaminase domain removes adenine from deoxyribonucleic acid (DNA). The fusion protein of any one of claims 12 to 47, which can be aminated.

49. the adenosine deaminase is a non-naturally occurring modified adenosine deaminase; A fusion protein according to any one of claims 12 to 47.

50. 50. Any one of claims 12 to 49, wherein the adenosine deaminase is TadA deaminase. The fusion protein according to claim 1.

51. 51. The fusion protein of claim 50, wherein the TadA deaminase is a TadA*7.10 variant. Quality.

52. the polynucleotide programmable DNA binding domain and the adenosine deaminase The fusion protein according to any one of claims 12 to 51, comprising a linker between the nucleotide domain and the nucleotide domain. quality.

53. 53. The method of claim 52, wherein the linker comprises the amino acid sequence: SGGSSGGSSGSETPGTSESATPES Fusion proteins.

54. The fusion protein of any one of claims 12 to 53, comprising one or more nuclear localization signals. Quality.

55. 55. The fusion protein of claim 54, wherein the nuclear localization signal is a bipartite nuclear localization signal. Synthetic protein.

56. 56. The fusion protein of any one of claims 12 to 55, wherein the Cas9 is StCas9.

57. The fusion transcription factor of any one of claims 12 to 55, wherein the Cas9 is SaCas9 or SpCas9. Protein.

58. 56. The method of any one of claims 12 to 55, wherein the Cas9 is a modified SaCas9 or a modified SpCas9. Fusion protein.

59. The modified SaCas9 has the amino acid substitutions E782K, N968K, and R1015H, or the amino acid substitutions corresponding thereto.

59. The fusion protein of claim 58, comprising an amino acid substitution.

60. The modified SaCas9 has the amino acid sequence: KRNYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG 60. The fusion protein of claim 59, comprising:

61. A polynucleotide encoding the fusion protein of any one of claims 12 to 60. 。

62. A polynucleotide encoding the fusion protein of any one of claims 12 to 60. 、 The base editor can be targeted to convert SNPs associated with genetic diseases from A and T to G and C. one or more guide polynucleotides that effect the modification; into a cell, or a precursor cell thereof.

63. 63. The cell of claim 62, wherein the cell is a human cell.

64. 64. The cell of claim 62 or 63, wherein the cell is in vitro or in vivo.

65. Any of claims 62 to 64, wherein the genetic disease is alpha-1 antitrypsin deficiency (A1AD). The cell described in any one of claims 1 to 4.

66. The fusion protein and the one or more guide polynucleotides are complexed within the cell.

66. The cell of any one of claims 62 to 65, which forms a body.

67. 67. An isolated cell grown or expanded from a cell according to any one of claims 62 to 66. or a population of cells.

68. 62. A method for treating a genetic disease in a subject in need thereof, comprising: The method comprises administering to the subject a cell described in any one of claims 1 to 67.

69. 69. The method of claim 68, wherein the cells are autologous, allogeneic, or xenogeneic to the subject. How to do it.

70. A polynucleotide programmable DNA binding domain and SEQ ID NO:1: 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 82, 94, 124, 133, 139, 146, and 158 or in another adenosine deaminase, at least one base that is an adenosine deaminase variant containing a corresponding modification in and an editor domain.

71. The adenosine deaminase variant is selected from the group consisting of R21N, R23H, E25F, N38G, L51 of SEQ ID NO: 1 W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K, or another adenosine deaminase 71. The base editor system of Claim 70, comprising the corresponding modification.

72. The base editor domain is targeted to convert the SNP A-T to G associated with a genetic disease.

70. The method of claim 70, further comprising one or more guide polynucleotides that result in a modification to C.

71. A base editor system according to claim 71.

73. The adenosine deaminase variant decomposes adenine in deoxyribonucleic acid (DNA).

73. The base editor system of any one of claims 70 to 72, which is capable of deaminating Tem.

74. The guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA).

74. The base editor system of claim 73.

75. The guide polynucleotide may comprise a CRISPR RNA (crRNA) sequence, a transactivating CRISPR RNA (crRNA) sequence, or a A (tracrRNA) sequence, or a combination thereof. system.

76. 73. The base editor system of Claim 72, further comprising a second guide polynucleotide. Hmm.

77. The second guide polynucleotide may be ribonucleic acid (RNA) or deoxyribonucleic acid (DNA 77. The base editor system of claim 76, comprising:

78. The second guide polynucleotide may comprise a CRISPR RNA (crRNA) sequence, a transactivating CRISPR 77. The base sequence of claim 76, comprising a nucleotide sequence of tracrRNA, a nucleotide sequence of tracrRNA, or a combination thereof. Computer system.

79. The polynucleotide-programmable DNA binding domain is Cas9, Cas12a / Cpf1, Cas1 2b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Ca 79. The base editor system of any one of claims 70 to 78, comprising an s12j / CasΦ domain. Hmm.

80. The polynucleotide programmable DNA binding domain is nuclease inactive.

80. The base editor system of claim 79.

81. 8. The method of claim 7, wherein the polynucleotide-programmable DNA-binding domain is a nickase.

10. The base editor system according to claim 9.

82. The polynucleotide programmable DNA binding domain comprises a Cas9 domain.

80. The base editor system of claim 79.

83. The Cas9 domain may be a nuclease-inactive Cas9 (dCas9), a Cas9 nickase (nCas9), or or a nuclease-active Cas9.

84. 84. The base editor system of Claim 83, wherein said Cas9 domain comprises a Cas9 nickase. Hmm.

85. The polynucleotide-programmable DNA binding domain is modified or Any of claims 70 to 84, wherein the polynucleotide is a programmable DNA binding domain. The base editor system of any one of claims 1 to 4.

86. 73. The base of claim 72, wherein the genetic disease is alpha-1 antitrypsin deficiency (A1AD). Editor system.

87. 1. A method for correcting a single nucleotide polymorphism (SNP) in a polynucleotide, comprising: a target nucleoside, at least a portion of which is located in said polynucleotide or its reverse complement; The peptide sequence is a fusion protein according to any one of claims 12 to 60 or a fusion protein according to any one of claims 70 to 85. contacting the base editor with the base editor system of any one of claims 1 to 4; deamination of the SNP or its complementary nucleobase upon targeting to the target nucleotide sequence. and deamination of the SNP or its complementary nucleobase is performed by The method, wherein the SNP is corrected.

88. 88. The method of claim 87, wherein the SNP is associated with alpha-1 antitrypsin deficiency (A1AD). The method described.

89. The SNP is in the SERPINA1 gene, and the modification comprises the E342K (PiZ allele) modification.

89. The method of claim 87 or 88.

90. 13. A method for editing a polynucleotide, said method comprising: fusion protein according to any one of claims 1 to 60 or according to any one of claims 70 to 85 contacting said polynucleotide with a base editor system, thereby editing said polynucleotide; The method comprising:

91. The editing may be less than 20% indel formation, less than 15% indel formation, less than 10% indel formation, <5% indel formation; <4% indel formation; <3% indel formation; <2% indel formation indel formation; less than 1% indel formation; less than 0.5% indel formation; or less than 0.1% indel formation 91. The method of claim 90, resulting in the formation of a nucleus.

92. 92. The method of Claim 91, wherein said editing does not result in a translocation.

93. an ABE comprising a TadA*7.10 adenosine deaminase variant domain selected from: 9 and Cas9 endonuclease domains: monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+A109S of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+T111R of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+D119N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+H122N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147d+Q154S of SEQ ID NO: 1, and the mutations I322V, S409 I, spCas9 (MQKFRAER) with E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+F149Y of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+T166I of SEQ ID NO: 1, and the mutation I322V spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; and monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+D167N of SEQ ID NO: 1, and the mutation I322V , spCas9 (MQKFRAER) with S409I, E427G, R654L, and R753G; monoTadA*7.10 with the mutations I76Y+V82T+Y147T+Q154S+L36H+N157K of SEQ ID NO: 1, and the mutations spCas9 (MQKFRAER) with I322V, S409I, E427G, R654L, R753G, and R1114G; monoTadA*7 with the mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K of SEQ ID NO:

1. 10, and SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G; monoT with the mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W of SEQ ID NO: 1 adA*7.10 and SpCas9 (MQKFR AER); monoTadA with the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N of SEQ ID NO: 1 *7.10, and SpCas9, MQKFRAER, with mutations I322V, S409I, E427G, R654L, R753G, and R1114G ; and A moiety having the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W of SEQ ID NO: 1 noTadA*7.10 and SpCas9 (MQ KFRAER); and The adenosine deaminase variant domain is targeted to detect a gene associated with a genetic disease. one or more guide polynucleotides that result in an A / T to G / C change at the SNP Base editors, including:

94. 94. The method of claim 93, wherein the SNP is associated with alpha-1 antitrypsin deficiency (A1AD). The base editor described.

95. monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+A109S and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+T111R and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+D119N and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+H122N and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147d+Q154S and mutations I322V, S409I, E427G, R65 4L, spCas9 with R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+F149Y and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+T166I and mutations I322V, S409I, E427 spCas9 (MQKFRAER) with G, R654L, and R753G; and monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+D167N and mutations I322V, S409I, E427 G, spCas9 with R654L, R753G (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S+L36H+N157K, mutations I322V, S409I, E spCas9 (MQKFRAER) with 427G, R654L, R753G, and R1114G; monoTadA*7.10 with mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K and mutation I SpCas9 (MQKFRAER) with 322V, S409I, E427G, R654L, R753G, and R1114G; monoTadA*7.10 and monoTadA*7.10 with mutations I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G monoTadA*7.10 with mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N and SpCas9 (MQKFRAER) with the mutations I322V, S409I, E427G, R654L, R753G, and R1114G; and monoTadA*7.10 with mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W , and SpCas9 (MQKFRAER) with mutations I322V, S409I, E427G, R654L, R753G, and R1114G. a TadA adenosine deaminase domain and a SpCas9 endonuclease domain selected from a vector comprising one or more polynucleotides encoding an ABE9 base editor, -.

96. 96. The vector of claim 95, which is a plasmid, viral, or mRNA vector.

97. A fusion protein according to any one of claims 12 to 60 or any one of claims 70 to 85. A composition comprising the base editor system of claim 1.

98. 98. The composition of claim 97, further comprising a pharmaceutically acceptable excipient, diluent, or carrier. thing.

99. A composition comprising the fusion protein of any one of claims 12 to 60 bound to a guide RNA. The composition, wherein the guide RNA is a SERPI associated with alpha-1 antitrypsin deficiency (A1AD). The composition, comprising a nucleic acid sequence complementary to the NA1 gene.

100. The base editor system of any one of claims 70 to 85, attached to a guide RNA. wherein the guide RNA is a gene encoding a gene encoding a gene for a gene encoding a gene for a gene associated with alpha-1 antitrypsin deficiency (A1AD). The composition comprises a nucleic acid sequence complementary to the SERPINA1 gene.

101. The adenosine deaminase variant decomposes adenine in deoxyribonucleic acid (DNA).

101. The composition of any one of claims 97 to 100, which can be deaminated.

102. the fusion protein or base editor system comprising: (i) whether it contains Cas9 nickase; (ii) contain nuclease-inactive Cas9; (iii) contain SpCas9 variants containing combinations of amino acid substitutions shown in Figures 3A-3C; or (iv) I322V, S409I, E427G, R654L, R753G (MQKFRAER); or I322V, S409I, E427 G, R654L, R753G, R1114G (MQKFRAER) SpCas9 variants, including 102. The composition of any one of claims 97 to 101.

103. Any of claims 99-102, further comprising a pharmaceutically acceptable excipient, diluent, or carrier. The composition according to any one of claims 1 to 4.

104. 99. A pharmaceutical composition for the treatment of a disease or disorder, comprising the composition of claim 98.

105. 105. The method of claim 104, wherein the disease or disorder is alpha-1 antitrypsin deficiency (A1AD). A pharmaceutical composition comprising:

106. the fusion protein or the base editor system is bound to a guide RNA; The guide RNA is targeted to the SERPINA1 gene, which is associated with alpha-1 antitrypsin deficiency (A1AD).

106. The pharmaceutical composition of claim 105, comprising a nucleic acid sequence that is complementary to said nucleic acid sequence.

107. Claim 106, wherein the gRNA and the base editor are formulated together or separately. The pharmaceutical composition described in

108. The gRNA, from 5' to 3', 5'-ACCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAU CAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-CCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUC AAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-CAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCA AC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAA C UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; 5'-UCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3'; or 5'-CGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' or 1, 2, 3, 4, or 5 nucleotides thereof selected from one or more of:

108. The pharmaceutical composition of any one of claims 98 or 103 to 107, comprising a 5' truncated fragment of said

109. Any of claims 98 or 103-108, further comprising a vector suitable for expression in mammalian cells.

10. The pharmaceutical composition of claim 1, wherein the vector encodes the base editor. The pharmaceutical composition comprising a polynucleotide.

110. 109. The method of Claim 109, wherein the polynucleotide encoding the base editor is an mRNA. Pharmaceutical compositions.

111. 110. The pharmaceutical composition of claim 109, wherein the vector is a viral vector.

112. The viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector, or viral vectors, herpes virus vectors, or adeno-associated virus vectors (AA 112. The pharmaceutical composition of claim 111, wherein

113. Claims 98 or 103-108 further comprising a ribonucleoparticle suitable for expression in a mammalian cell. The pharmaceutical composition according to any one of the preceding claims.

114. 109. The pharmaceutical composition of any one of claims 98 or 103 to 108, further comprising a lipid.

115. 1. A method for treating alpha-1 antitrypsin deficiency (A1AD), said method comprising administering to a subject in need thereof administering to a subject a pharmaceutical composition according to any one of claims 98 or 103 to 114. The method of treatment comprises:

116. Claim 98 or 103 in treating alpha-1 antitrypsin deficiency (A1AD) in a subject Use of a pharmaceutical composition according to any one of claims 1 to 114.

117. 117. The method of claim 115 or the use of claim 116, wherein the subject is a human.

118. The adenosine deaminase variant may be selected from the group consisting of the following modifications: E25F+V82S+Y123H; T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+V82S+Y123H+T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; V82S+Y123H+P124W+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R; P54C+V82S+Y123H+Y147R+Q154R; Y73S+V82S+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; R23H+V82S+Y123H+Y147R+Q154R; R21N+V82S+Y123H+Y147R+Q154R; V82S+Y123H+Y147R+Q154R+A158K; N72K+V82S+Y123H+D139L+Y147R+Q154R; E25F+V82S+Y123H+D139M+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R 87. The base editor system of any one of claims 70 to 86, comprising any one of Tem.

119. The adenosine deaminase or adenosine deaminase variant is selected from the group consisting of: Acid modification or group of modifications: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S 10. A TadA*7.10 variant comprising any one of the following: adenosine deaminase, fusion protein, base editor, or base editor described herein system.

120. The following amino acid modification or group of modifications: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S adenosine deaminase barrier, which is a TadA*7.10 variant containing any one of nt.

121. Polynucleotide-programmable DNA binding domains with the following amino acid alterations or modifications: Strange group: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S At least one TadA*7.10 adenosine deaminase variant comprising any one of and one base editor domain.

122. the polynucleotide-programmable DNA-binding domain is a Cas9 endonuclease 122. The fusion protein of claim 121, comprising a domain.

123. The Cas9 endonuclease domain contains the mutations I322V, S409I, E427G, R654L, and R753G.

123. The fusion protein of Claim 122, comprising spCas9 (MQKFRAER) having the following structure:

124. 122. The adenosine deaminase variant of claim 121, wherein the TadA7*10 is a monomer. Or a fusion protein described in any one of claims 121 to 123.

125. A TadA*7.10 adenosine deaminase variant domain and a Cas9 enzyme selected from the following: Nucleobase editors containing endonuclease domains: monoTadA*7.10 with mutation V82T and mutations I322V, S409I, E427G, R654L, and R753G spCas9 (MQKFRAER); monoTadA*7.10 with mutations I76Y+V82T and mutations I322V, S409I, E427G, R654L, and R753G spCas9 (MQKFRAER); or monoTadA*7.10 with mutations I76Y+V82T+Y147T+Q154S and mutations I322V, S409I, E427G, R65 4L, spCas9 (MQKFRAER) with R753G.

Citation Information

Patent Citations

  • Novel Nucleic Acid Base Editor and Method of Using the Same

    JP7717684B2

  • Adenosine nucleobase editors and uses thereof

    WO2018027078A1