Composition and method for complement activation modification

KR1020260124082APending Publication Date: 2026-08-14BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020267018385
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-20
Publication Date
2026-08-14

Smart Images

  • Figure PCT00210_ABST
    Figure PCT00210_ABST
Patent Text Reader

Abstract

Composition and method for reducing complement activation by introducing one or more modifications to an intracellular complement factor B [CFB] polynucleotide. In certain embodiments, the invention of the present disclosure features a base editor system for modifying a CFB polynucleotide (e.g., a fusion protein or complex comprising a programmable DNA binding protein, a nucleobase editor, and gRNA), wherein the modification is associated with a reduction in the expression and / or activity of the CFB polypeptide encoded by the polynucleotide.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] 관련 출원의 상호 참조

[0002] This application claims priority to U.S. Provisional Application No. 63 / 601,145 filed on November 20, 2023, the entire contents of said Provisional Application are incorporated herein by reference in their entirety.

[0003] 서열목록

[0004] The present application includes a sequence list electronically submitted in XML format, the entire sequence list incorporated herein by reference. The name of the sequence list XML file created on November 20, 2024 is 180802-055802PCT_SL.xml, and its size is 3,905,996 bytes. Background Technology

[0005] The complement system is an important part of the innate immune system and is involved in the removal of microorganisms and cellular debris, as well as inflammation and the activation of various immune pathways. While overactivation of the complement system or inappropriate targeting of self-cells can lead to disease, it has been successfully and safely proven that inhibiting complement system activity provides therapeutic benefits to patients suffering from an overactive complement system. Therefore, there is interest in improved methods to reduce complement system activation in these patients.

[0006] As described below, the present disclosure features a composition and a method for reducing complement activation by introducing one or more modifications to a complement factor B [CFB] polynucleotide within a cell. In certain embodiments, the invention of the present disclosure features a base editor system for modifying a CFB polynucleotide (e.g., a fusion protein or complex comprising a programmable DNA binding protein, a nucleobase editor, and gRNA), wherein the modification is associated with a reduction in the expression and / or activity of the factor B polypeptide encoded by the polynucleotide. Non-limiting examples of modifications include base editing.

[0007] In one embodiment, the present disclosure provides a method for modifying the nucleobase of a complement factor B [CFB] polynucleotide. The method comprises the step of contacting the CFB polynucleotide with a base editor system containing one or more guide polynucleotides or one or more polynucleotides encoding one or more guide polynucleotides, and a base editor containing one or more polynucleotides encoding a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain or the base editor. The method comprises (a), (b), and / or (c), wherein in (a), one or more guide polynucleotides target the base editor to modify the nucleobase of the CFB polynucleotide by: i. disrupting the splice site of the CFB polynucleotide and / or; ii. changing the start codon of the CFB polynucleotide and / or; iii. changing the TATA box of the CFB polynucleotide and / or; or iv. introducing a new stop codon into the CFB polynucleotide. In (b), the deaminase domain

[0008] A TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ No. 1) or a fragment thereof lacking only N-terminal methionine, wherein TadA* comprises, compared with the TadA*7.10 amino acid sequence, i. Y123H, Y147R, and Q154R, ii. I76Y, Y133H, Y147R, and Q154R, iii. V82S and Q164R, iv. It further contains combinations of amino acid changes selected from one or more of I76Y, V82S, Y123H, Y147R, and Q154R, v. I76Y, V82T, Y123H, Y147R, and Q154R, and vi. I76Y, V82T, Y123H, Y147T, and Q154S. In (c), one or more guide polynucleotides contain a nucleic acid sequence selected from CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443), and UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467), or contain at least 10 to 23 consecutive nucleotides of a spacer nucleic acid sequence listed in any one of Tables 2A to 2F. The method changes the nucleobase of the CFB polynucleotide.

[0009] In another aspect, the present disclosure provides a method for modifying the nucleobase of a complement factor B [CFB] polynucleotide. The method comprises the step of contacting the CFB polynucleotide with one or more guide polynucleotides or one or more polynucleotides encoding one or more guide polynucleotides or one or more guide polynucleotides and a base editor or one or more polynucleotides encoding a base editor containing a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain. The method comprises (a) and (b), wherein in (a), the deaminase domain is cytidine deaminase or

[0010] A TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ No. 1) or a fragment thereof lacking only N-terminal methionine, wherein TadA* comprises, compared with the TadA*7.10 amino acid sequence, i. Y123H, Y147R, and Q154R, ii. I76Y, Y133H, Y147R, and Q154R, iii. V82S and Q164R, iv. It further contains combinations of amino acid changes selected from one or more of I76Y, V82S, Y123H, Y147R, and Q154R, v. I76Y, V82T, Y123H, Y147R, and Q154R, and vi. I76Y, V82T, Y123H, Y147T, and Q154S. In (b), one or more guide polynucleotides contain a spacer containing a nucleotide sequence selected from one or more of UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835). The method involves modifying the nucleobase of the CFB polynucleotide.

[0011] In another aspect, the present disclosure provides a cell produced by a method of any aspect or embodiment of the present disclosure.

[0012] In another aspect, the present disclosure provides a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or a base editor system comprising one or more polynucleotides encoding the base editor and one or more guide polynucleotides, or one or more polynucleotides encoding one or more guide polynucleotides. One or more guide polynucleotides contain a nucleotide sequence selected from CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443) and UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467), and / or contain at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleobases of a spacer containing a nucleotide sequence listed in Tables 2A to 2F.

[0013] In another aspect, the present disclosure provides a polynucleotide encoding a base editor or a component thereof of any aspect or embodiment of the present disclosure.

[0014] In another aspect, the present disclosure provides a base editor system or a vector containing a polynucleotide of any aspect or embodiment of the present disclosure.

[0015] In another aspect, the present disclosure provides a base editor system or lipid nanoparticles containing polynucleotides of any aspect or embodiment of the present disclosure.

[0016] In another aspect, the present disclosure provides lipid nanoparticles containing A) a polynucleotide encoding a base editor and B) a guide RNA or a polynucleotide encoding a guide RNA.

[0017] In another aspect, the present disclosure provides a pharmaceutical composition comprising an effective amount of a base editor system of any aspect or embodiment of the present disclosure, a polynucleotide, a vector, or lipid nanoparticles, and a pharmaceutically acceptable excipient.

[0018] In another aspect, the present disclosure provides a kit containing a container in which a base editor system, vector, lipid nanoparticle, or pharmaceutical composition of any aspect or embodiment of the present disclosure is contained.

[0019] In another aspect, the present disclosure provides a guide polynucleotide containing a sequence listed in any one of Tables 1A to 2F.

[0020] In another aspect, the present disclosure provides a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or a base editor system comprising one or more polynucleotides encoding the base editor and one or more guide polynucleotides, or one or more polynucleotides encoding one or more guide polynucleotides. The method comprises (a), (b), (c) and (d), wherein in (a), the deaminase domain is a cytidine deaminase, or

[0021] A TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ No. 1) or a fragment thereof lacking only N-terminal methionine, wherein TadA* comprises, compared with the TadA*7.10 amino acid sequence, i. Y123H, Y147R, and Q154R, ii. I76Y, Y133H, Y147R, and Q154R, iii. V82S and Q164R, iv. It further contains combinations of amino acid changes selected from one or more of I76Y, V82S, Y123H, Y147R, and Q154R, v. I76Y, V82T, Y123H, Y147R, and Q154R, and vi. I76Y, V82T, Y123H, Y147T, and Q154S. In (b), one or more guide polynucleotides contain a spacer containing a nucleotide sequence selected from one or more of UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835). In (c), one or more guiding polynucleotides are,

[0022] Terminal modification SpCas9 guide polynucleotide

[0023] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ ID NO: 440),

[0024] HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440), and

[0025] NLS(bpsv40): containing a sequence selected from one or more of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV (SEQ Nos. 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of nucleotide 'N', and 'Ns' represents that nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), where the number of N nucleotides is 15 to 25. (d) napDNAbp is a spCas9 clefting enzyme polypeptide that binds to a protospacer adjacent motif (PAM) selected from one or more of NGA, NGC and NGG, where 'N' is any nucleotide.

[0026] In another aspect, the present disclosure provides lipid nanoparticles containing a base editor system or a component thereof of any aspect or embodiment of the present disclosure.

[0027] In another aspect, the present disclosure provides a composition comprising a polynucleotide and a guide RNA. The polynucleotide encodes a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain. The deaminase domain is a cytidine deaminase or contains a TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1. Compared to the TadA*7.10 amino acid sequence, TadA* is i. It further contains a combination of amino acid changes selected from one or more of I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147R and Q154R.The guide RNA is a nucleotide sequence selected from one or more of CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835), and at least one of the spacer nucleic acid sequences listed in any one of Tables 2A to 2F. It contains a spacer containing a nucleotide sequence containing 10 to 23 consecutive nucleotides.

[0028] In another aspect, the present disclosure provides a lipid nanoparticle (LNP) composition containing mRNA and guide RNA. The mRNA encodes a base editor containing a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain. The deaminase domain is a cytidine deaminase or contains a TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1. Compared to the TadA*7.10 amino acid sequence, TadA* is i. It further contains a combination of amino acid changes selected from one or more of I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147R and Q154R.The guide RNA is a nucleotide sequence selected from one or more of CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835), and at least one of the spacer nucleic acid sequences listed in any one of Tables 2A to 2F. It contains a spacer containing a nucleotide sequence containing 10 to 23 consecutive nucleotides.

[0029] In another aspect, the present disclosure provides a therapeutic method comprising the step of administering a lipid nanoparticle (LNP) composition containing mRNA and guide RNA to a subject. The mRNA encodes a base editor containing a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain. The deaminase domain is cytidine deaminase or contains a TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1. Compared to the TadA*7.10 amino acid sequence, TadA* is i. It further contains a combination of amino acid changes selected from one or more of I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147R and Q154R.The guide RNA is a nucleotide sequence selected from one or more of CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835), and at least one of the spacer nucleic acid sequences listed in any one of Tables 2A to 2F. It contains a spacer containing a nucleotide sequence containing 10 to 23 consecutive nucleotides.

[0030] In any aspect or embodiment of the present disclosure, the splice site is located near the 3' end of exon 1, exon 10, exon 11, exon 12, exon 14, exon 15, or exon 16 of the CFB polynucleotide. In any aspect or embodiment of the present disclosure, the splice site is located near the 5' end of exon 5, exon 8, exon 9, exon 10, exon 11, exon 14, or exon 18 of the CFB polynucleotide.

[0031] In any aspect or embodiment of the present disclosure, one or more guide polynucleotides contain a spacer that is complementary to both the human CFB polynucleotide and the non-human primate CFB polynucleotide. In any aspect or embodiment of the present disclosure, one or more guide polynucleotides contain a spacer that is complementary to the human CFB polynucleotide but not complementary to the non-human primate CFB polynucleotide. In any aspect or embodiment of the present disclosure, one or more guide polynucleotides contain a spacer that contains only 20 or 21 nucleotides. In any aspect or embodiment of the present disclosure, one or more guide polynucleotides contain a spacer containing a nucleotide sequence selected from one or more of UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835).

[0032] In any aspect or embodiment of the present disclosure, the deaminase domain is an adenosine deaminase containing a TadA*7.10 amino acid sequence further containing a combination of amino acid modifications selected from one or more of i. Y123H, Y147R and Q154R, ii. I76Y, Y133H, Y147R and Q154R, iii. V82S and Q164R, iv. I76Y, V82S, Y123H, Y147R and Q154R, v. I76Y, V82T, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147T and Q154S.

[0033] In any aspect or embodiment of the present disclosure, napDNAbp is a clefting enzyme. In any aspect or embodiment of the present disclosure, napDNAbp binds to a protospacer adjacent motif (PAM) selected from one or more of NGA, NGC, NGG, and NNNRRT, wherein 'N' is any nucleotide and 'R' is A or G. In any aspect or embodiment of the present disclosure, napDNAbp is a Cas9 polypeptide.

[0034] In any aspect or embodiment of the present disclosure, one or more guide polynucleotides contain modified nucleotides. In any aspect or embodiment of the present invention, one or more guide polynucleotides are,

[0035] Terminal modification SpCas9 guide polynucleotide

[0036] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ ID NO: 440),

[0037] Terminal modification SaCas9 guide polynucleotide

[0038] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUsmUsmUsmU (SEQ ID NO: 3128);

[0039] HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),

[0040] HM07: mNsmNsmNsmNmNmNmNmNmNmNNNNNNNNNNNmGUUUUAGmAmGmCmUmAmGmAmAmAmUmAmGmCmAmAGUUmAAmAAmUAmAmGmGm CmUmAGUmCmCGUUAmUmCAAmCmUmUmGmAmAmAmAmAmGmUmGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),

[0041] NLS(bpsv40): mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV(sequence numbers 440 and 446),

[0042] LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGUGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmUsmU(Sequence No. 445);

[0043] NLS+LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmAmGmUmGmGmGmAmGmGmAmGmGmGmGmUmGmGmUsmUsmUsmU-NHC5-CrossL- CKRTADGSEFESPKKKRKV(SEQ ID NOs 445 and 446); and

[0044] LONGEST+GOLD:

[0045] The sequence contains a sequence selected from one or more of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmGmGmAmAmAmCmGmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAmUmCAAmCmUmUGGACUUCGGUCCmAmAmGUGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmUsmU (Sequence No. 447), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of nucleotide 'N', and 'Ns' represents that nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), where the number of N nucleotides is 15 to 25.

[0046] In any aspect or embodiment of the present disclosure, the CFB polynucleotide is present within a cell. In any aspect or embodiment of the present disclosure, the cell is a mammalian cell. In any aspect or embodiment of the present disclosure, the cell is a retinal cell or other cell of the eye, a neuron, or a hepatocyte.

[0047] In any aspect or embodiment of the present disclosure, one or more guide polynucleotides target a base editor and modify the nucleobase of the CFB polynucleotide to destroy the splice site of the CFB polynucleotide.

[0048] In any aspect or embodiment of the present disclosure, napDNAbp is a splitting enzyme.

[0049] In any aspect or embodiment of the present disclosure, napDNAbp binds to a protospacer adjacent motif (PAM) selected from one or more of NGA, NGC, NGG and NNNRRT, wherein 'N' is any nucleotide and 'R' is A or G.

[0050] In any aspect or embodiment of the present disclosure, napDNAbp is a Cas9 polypeptide.

[0051] In any aspect or embodiment of the present disclosure, CFB activity, protein concentration and / or mRNA concentration is reduced by at least about 15% compared to control cells with no change.

[0052] In any aspect or embodiment of the present disclosure, the CFB polynucleotide is in contact with two or more guide polynucleotides, each guide polynucleotide being bonded to a different position within the CFB polynucleotide.

[0053] In any aspect or embodiment of the present disclosure, the deaminase domain comprises a TadA variant (TadA*) containing an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of cytidine deaminase or MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEVENT. 1), wherein TadA* comprises, compared with the TadA*7.10 amino acid sequence, i. Y123H, Y147R, and Q154R, ii. I76Y, Y133H, Y147R, and Q154R, iii. iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, v. I76Y, V82T, Y123H, Y147R and Q154R, and vi. further contain a combination of amino acid changes selected from one or more of I76Y, V82T, Y123H, Y147T and Q154S.

[0054] In any aspect or embodiment of the present disclosure, the method is not a process for altering the genetic identity of human germ cells.

[0055] In any aspect or embodiment of the present disclosure, the adenosine deaminase domain 표 5G It contains a combination of mutations selected from those listed.

[0056] In any aspect or embodiment of the present disclosure, the guide RNA is,

[0057] Terminal modification SpCas9 guide polynucleotide

[0058] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ ID NO: 440),

[0059] HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440), and

[0060] NLS(bpsv40): containing a sequence selected from one or more of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV (SEQ Nos. 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of nucleotide 'N', and 'Ns' represents that nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), where the number of N nucleotides is 15 to 25.

[0061] In any embodiment of the present disclosure, the polynucleotide is mRNA.

[0062] In any aspect or embodiment of the present disclosure, the composition is formulated with lipid nanoparticles (LNP).

[0063] 정의

[0064] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by a person skilled in the art to which this disclosure pertains. The following references provide general definitions of many terms used in this disclosure to a person skilled in the art. Reference [Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994)]; Reference [The Cambridge Dictionary of Science and Technology (Walker ed., 1988)]; Reference [The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991)]; and Reference [Hale & Marham, The Harper Collins Dictionary of Biology (1991)]. As used herein, unless otherwise specified, the following terms have the meanings assigned below.

[0065] 'Adenine' or '9 H '-purine-6-amine' is It refers to a purine nucleobase with the molecular formula C5H5N5, having the structure and corresponding to CAS number 73-24-5.

[0066] 'Adenosine' or '4-amino-1-[(2 R ,3 R ,4 S ,5 R )-3,4-dihyroxy-5-(hydroxymethyl)oxolane-2-yl]pyrimidine-2(1 H )-on' is It refers to an adenine molecule attached to a ribose sugar via a glycosidic bond, having the structure and corresponding to CAS number 65-46-3. Its molecular formula is C 10 H 13 It is N5O4.

[0067] "Adenosine deaminase" or "adenine deaminase" means a polypeptide or a fragment thereof capable of catalyzing the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination from adenosine to inosine or from deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically modified adenosine deaminases, evolved adenosine deaminases) may be derived from any organism (e.g., eukaryotes, prokaryotes) including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase variant with one or more modifications, capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA), and may be referred to as a 'double deaminase'. Non-limiting examples of double deaminases include those described in PCT / US22 / 22050. In some embodiments, the target polynucleotide is single or double-stranded. In some embodiments, the adenosine deaminase variant can deaminate both adenine and cytosine of DNA. In some embodiments, the adenosine deaminase variant can deaminate both adenine and cytosine of single-stranded DNA. In some embodiments, the adenosine deaminase variant can deaminate both adenine and cytosine of RNA.In the embodiments, the adenosine deaminase variants are selected for all purposes from those described in PCT / US2020 / 018192, PCT / US2020 / 049975, PCT / US2017 / 045381, PCT / US2021 / 016827, PCT / US2022 / 073781, PCT / US24 / 34189, or PCT / US2020 / 028568, the entire contents of which are incorporated herein by reference. Additional non-limiting examples of adenosine deaminases include those designed using artificial intelligence and disclosed or referenced in the literature [Rufflow, et al., "Design of highly functional genome editors by modeling of the universe of CRISPR-Cas Sequences," bioRxiv, posted April 22, 2024, doi: 10.1101 / 2024.04.22.590591], the full disclosure thereof is incorporated herein by reference for all purposes. Additional exemplary adenosine deaminase amino acid sequences include TadA-8e (Sequence No. 3575), Tad1 (Sequence No. 3576), Tad2 (Sequence No. 3577), Tad3 (Sequence No. 3578), Tad4 (Sequence No. 3579), Tad6 (Sequence No. 3580), Tad6-SR (Sequence No. 3581), TadA9 (Sequence No. 3582), TadA20 (Sequence No. 3583), Staphylococcus aureus (. Staphylococcus aureus ) TadA (sequence number 3584), Bacillus subtilis ( Bacillus subtilis ) TadA (sequence number 3585), Salmonella typhimurium ( Salmonella typhimurium ) TadA (sequence number 3586), Shewanella putrepathiens ( Shewanella putrefaciens )(Sequence No. 3587), Haemophilus influenzae( Haemophilus influenzae ) F3031 TadA (sequence number 3588), Caullobacter crescentus ( Caulobacter crescentus ) TadA (sequence number 3589), Geobacter sulfur reducens ( Geobacter sulfurreducens ) TadA (sequence number 3590), Streptococcus phyogenes ( Streptococcus pyogenes ) TadA (sequence number 3591), Aquifex aeolicus ( Aquifex aeolicus ) TadA (Sequence No. 3592) and Escherichia coli[ E. coli ] Includes TadA deaminase (ecTadA) (Sequence No. 3593).

[0068] 'Adenosine deaminase activity' refers to the catalyzing the deamination of adenine from polynucleotides or from adenosine to guanine.

[0069] An 'adenosine base editor (ABE)' refers to a base editor containing adenosine deaminase.

[0070] 'Adenosine base editor (ABE) polynucleotide' refers to a polynucleotide that encodes ABE.

[0071] 'Adenosine base editor 8 (ABE8) polypeptide' or 'ABE8' is 표 5B One or more changes listed in, 표 5B One of the combinations of changes listed in, or 표 5B Meaning a base editor as defined herein comprising an adenosine deaminase or an adenosine deaminase variant comprising a change at one or more amino acid positions listed in, wherein such change is for the reference sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ No. 1) or for a corresponding position of another adenosine deaminase. In embodiments, ABE8 comprises a change at amino acid 82 and / or 166 of SEQ No. 1. In some embodiments, ABE8 comprises additional changes as described herein relative to the reference sequence.

[0072] 'Adenosine base editor 8 (ABE8) polynucleotide' refers to a polynucleotide that encodes the ABE8 polypeptide.

[0073] "Administering" refers to providing one or more of the compositions described herein to a patient or subject. By example and without limitation, administration of the composition (e.g., injection) may be performed by intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. One or more such routes may be utilized. Parenteral administration may be performed over time, for example, by bolus injection or gradual perfusion. In some embodiments, parenteral administration includes infusion or injection into the blood vessel, intravenous, intramuscular, intra-arterial, intradural, intratumoral, intradermal, intraperitoneal, percutaneous, subcutaneously, subcuticularly, intra-articular, subcapsular, subcutaneous, and intrasternal. Alternatively, or concurrently, administration may be performed by an oral route.

[0074] 'Agent' means any small molecule chemical compound, antibody, nucleic acid molecule, polypeptide, or functional fragment thereof.

[0075] “Change” means a change in the numerical, structural, or activity of an analyte, gene, or polypeptide as detected by known standard technical methods such as those described herein. As used herein, a change includes a change in expression levels (e.g., increase or decrease). In the embodiments, the increase or decrease in expression levels is 10%, 25%, 40%, 50%, or more. In some embodiments, a change includes the insertion, deletion, or substitution of a nucleobase or amino acid (e.g., by genetic manipulation).

[0076] 'To improve' means to reduce, inhibit, weaken, shrink, stop, or stabilize the development or progression of a disease.

[0077] "Analogous" refers to a molecule that is not identical but possesses similar functional or structural characteristics. For example, polypeptide analogs retain the biological activity of the corresponding naturally occurring polypeptide while possessing specific biochemical modifications that enhance the function of the analog relative to the naturally occurring polypeptide. Such biochemical modifications can, for example, increase the protease resistance, membrane permeability, or half-life of the analog without altering ligand binding. Analogs may contain non-natural amino acids.

[0078] 'Base editor [BE]' or 'nucleobase editor polypeptide (NBE)' refers to an agent that binds to a polynucleotide and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase modification polypeptide (e.g., a deaminase) and a polynucleotide programmable nucleotide binding domain (e.g., Cas9 or Cpf1). Representative nucleic acid and protein sequences of the base editor include sequences corresponding to SEQ ID NOs 2 through 11, which have about or at least about 85% sequence identity with any base editor sequence provided in the sequence listing.

[0079] 'BE4 cytidine deaminase (BE4) polypeptide' refers to a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain, a cytidine deaminase domain, and two uracil glycosylase inhibitor (UGI) domains. In an embodiment, napDNAbp is a Cas9n(D10A) polypeptide. Non-limiting examples of the cytidine deaminase domain include rAPOBEC, ppAPOBEC, RrA3F, AmAPOBEC1, and SsAPOBEC3B.

[0080] 'BE4 cytidine deaminase (BE4) polynucleotide' means a polynucleotide encoding a BE4 polypeptide.

[0081] 'Base editing activity' refers to the action of chemically altering a base within a polynucleotide. In one embodiment, the first base is converted to the second base. In one embodiment, the base editing activity is cytidine deaminase activity, e.g., target C G to T It is to convert to A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, e.g., A From T to G It is to convert it to C.

[0082] The term 'base editor system' refers to an intermolecular complex for editing nucleosides of a target nucleotide sequence. In various embodiments, the base editor [BE] system comprises (1) a polynucleotide programmable nucleotide binding domain, a deaminase domain (e.g., cytidine deaminase or adenosine deaminase) for deaminating nucleosides in a target nucleotide sequence, and (2) one or more guide polynucleotides (e.g., guide RNA) together with the polynucleotide programmable nucleotide binding domain. In various embodiments, the base editor [BE] system comprises a nucleoside editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises (1) a base editor [BE] comprising a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence, and (2) one or more guide RNAs together with the polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor [ABE] or a cytidine or cytosine base editor (CBE).In some embodiments, a base editor system (e.g., a base editor system comprising cytidine deaminase) comprises a uracil glycosylationase inhibitor or other agent or peptide that inhibits an inosine base excision repair system (e.g., a uracil stabilizing protein provided in WO2022015969, the entire disclosure thereof incorporated herein by reference for all purposes).

[0083] The terms 'Cas9' or 'Cas9 domain' refer to RNA-guided nucleases comprising the Cas9 protein or a fragment thereof (e.g., a protein containing the active, inactive, or partially active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Additionally, Cas9 nucleases sometimes refer to casnl nucleases or clustered regularly interspaced short palindromic repeat (CRISPR) associated nucleases.

[0084] The terms 'conservative amino acid substitution' or 'conservative mutation' refer to the replacement of a single amino acid with another amino acid that shares a common characteristic. A functional method for defining common characteristics between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (reference [Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)]). According to this analysis, amino acid groups can be defined as those where amino acids within the group are preferentially exchanged, and thus have the most similar effects on the overall protein structure (ibid. [Schulz, GE and Schirmer, RH]). Non-limiting examples of conservative mutations include amino acid substitutions, for example, substituting arginine with lysine and vice versa so that a positive charge is maintained, aspartic acid with glutamic acid and vice versa so that a negative charge is maintained, threonine with serine so that a free -OH is maintained, and asparagine with glutamine so that a free -NH2 is maintained.

[0085] Amino acids can generally be classified into classes based on the following common side chain characteristics.

[0086] (1) Hydrophobic: Norleucine, Met, Ala, Val, Leu, He;

[0087] (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gin;

[0088] (3) Acids: Asp, Glu;

[0089] (4) Basics: His, Lys, Arg;

[0090] (5) Residues affecting chain orientation: Gly, Pro;

[0091] (6) Aromatic: Trp, Tyr, Phe.

[0092] In some embodiments, a conservative substitution may involve exchanging a member of one of these classes for another member of the same class. In some embodiments, a non-conservative substitution involves exchanging a member of one of these classes for another class.

[0093] The terms 'coding sequence' or 'protein coding sequence,' as used interchangeably herein, refer to a segment of polynucleotide that codes for a protein. A coding sequence may also be referred to as an open decoding frame. A region or sequence is bounded by a start codon closer to the 5' end and by a stop codon closer to the 3' end. Stop codons useful for the base editors described herein include TAG, TAA, and TGA.

[0094] 'Complement factor B [CFB] polypeptide' or 'factor B [factor B, FB] polypeptide' means a factor B protein or a fragment thereof capable of mediating the activation of the complement system having at least about 85% amino acid sequence identity with GenBank accession number AAA16820.1 provided below. In the embodiments, CFB can produce C3a and C3b by cleaving an Arg-Ser bond in complement component C3, and / or can produce C5a and C5b by cleaving an Arg-Ser bond in complement component C5.

[0095] >AAA16820.1 Complement Factor B[Homo sapiens]

[0096] MGSNLSPQLCLMPFILGLLSGGVTTTPWSLAQPQGSCSLEGVEIKGGSFRLLQEGQALEYVCPSGFYPYPVQTRTCRSTGSWSTLKTQDQKTVRKAECRAIHCPRPHDFENGEYWPRSPYYNVSDEISFHCYDGYTLRGSANRTCQVNGRWSGQTAICDNGAGYCSNPGIPIGTRKVGSQYRLEDSVTYHCSRGLTLRGSQRRTCQEGGSWSGTEPSCQDSFMYDTPQEVAEAFLSSLTETIEGVDAEDGHGPGEQQKRKIVLDPSGSMNIYLVLDGSDSIGASNFTGAKKCLVNLIEKVASYGVKPRYGLVTYATYPKIWVKVSEADSSNADWVTKQLNEINYEDHKLKSGTNTKKALQAVYSMMSWPDDVPPEGWNRTRHVIILMTDGLHNMGGDPITVIDEIRDLLYIGKDRKNPREDYLDVYVFGVGPLVNQVNINALASKKDNEQHVFKVKDMENLEDVFYQMIDESQSLSLCGMVWEHRKGTDYHKQPWQAKISVIRPSKGHESCMGAVVSEYFVLTAAHCFTVDDKEHSIKVSVGGEKRDLEIEVVLFHPNYNINGKKEAGIPEFYDYDVALIKLKNKLKYGQTIRPICLPCTEGTTRALRLPPTTTCQQQKEELLPAQDIKALFVSEEEKKLTRKEVYIKNGDKKGSCERDAQYAPGYDKVKDISEVVTPRFLCTGGVSPYADPNTCRGDSGGPLIVHKRSRFIQVGVISWGVVDVCKNQKRQKQVPAHARDFHINLFQVLPWLKEKLQDEDLGFL(서열번호 426)

[0097] 'Complement factor B [CFB] polynucleotide' or 'factor B [FB] polynucleotide' refers not only to a nucleic acid molecule encoding a CFB polypeptide, but also to introns, exons, 3' untranslating regions, 5' untranslating regions, and regulatory sequences or fragments thereof associated with their expression. In the embodiments, the CFB polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for CFB expression. Homo sapiens ( Homo Sapiens An exemplary CFB nucleotide sequence (GenBank: L15702.1:41-2335; Ensembl: ENST00000425368.7) in ) is provided below.

[0098] L15702.1:41-2335 Human complement factor B mRNA, intact cds

[0099]

[0100] chromosome:GRCh38:6:31945050:31952686:1(ENST00000425368.7), where the exon is 굵은 글씨 It is marked as , and the non-translated area is 밑줄 It is represented as, and the intron is 굵은 글씨 It corresponds to the area marked with regular text between the areas marked with , and the TATA box is indicated by a double underline. The sequence contains 18 exons, where exon 1 corresponds to the first exon at the 5' end of the sequence, exon 2 corresponds to the second exon at the 5' end of the sequence, and this continues to exon 18.

[0101] GGGAAGGGAATGTGACCAGGTCTAGGTCTGGAGTTTCAGCTTGGACACTGAGCCAAGCAGACAAGCAAAGCAAGCCAGGACACACCATCCTGCCCCAGGCCCAGCTTCTCTCCTGCCTTCCAACGCC ATGGGGAGCAATCTCAGCCCCCAACTCTGCCTGATGCCCTTTATCTTGGGCCTCTTGTCTGGAG GTAAGCGAGGGTAACCTTCCCTTCCTGCTGTCTCCAGCATCCCTCCTTGGCCTTTTGGGGCCAGGCTTCATCAGCCTTTCTCTTCAG GTGTGACCACCACTCCATGGTCTTTGGCCCGGCCCCAGGGATCCTGCTCTCTGGAGGGGGTAGAGATCAAAGGCGGCTCCTTCCGACTTCTCCAAGAGGGCCAGGCACTGGAGTACGTGTGTCCTTCTGGCTTCTACCCGTACCCTGTGCAGACACGTACCTGCAGATCTACGGGGTCCTGGAGCACCCTGAAGACTCAAGACCAAAAGACTGTCAGGAAGGCAGAGTGCAGAG GTTTGAGGGCAATGAGTGTGGGCAGTGGCCTAAGGCAGAAACAGGGCAGGCGGCAGCAAGGTCAGGACTAGGATGAGACTAGGCAGGGTGACAAGGTGGGCTGACCGGGAGTAGGAGCAGTTTTAGGGTGGCAGGCGGAAAGGGGGCAAGAAAAAGCGGAGTTAACCCTTACTAAGCATTTACCCTGGGCTTCCAGGCAGCCCTGGAAGTCAAGAGAACACTCAGAAATGGGGAGGGAGAAGCAGTGGAAATCCATATGGGTTGAGGAGTAGGTAAGATGCTGCTTCTGCGGGACTGGGAATGCGCTGTTTCTCAGTGACATGGTCTCCGAGACCAGGAGGGATACACCTAAGGCAGCCTTTCCCTCTTGATGACTTCTACTTGTCCCCCCTTCTCAAAG CAATCCACTGTCCAAGACCACACGACTTCGAGAACGGGGAATACTGGCCCCGGTCTCCCTACTACAATGTGAGTGATGAGATCTCTTTCCACTGCTATGACGGTTACACTCTCCGGGGCTCTGCCAATCGCACCTGCCAAGTGAATGGCCGATGGAGTGGGCAGACAGCGATCTGTGACAACGGAG GTGAGAAGCATCCCCTCCCCCTACATTGCTGTCTCCCTGACGGCGCCCAGCCCGAGGAGTGGGCACTCGGCTCCGGACACTGTAACTCTTGCTCTCTACCTTGCTCACGGGGCCTCAGGCTTCAGTGCTTACCTCGATGTCTCATACCTCTGCAG CGGGGTACTGCTCCAACCCGGGCATCCCCATTGGCACAAGGAAGGTGGGCAGCCAGTACCGCCTTGAAGACAGCGTCACCTACCACTGCAGCCGGGGGCTTACCCTGCGTGGCTCCCAGCGGCGAACGTGTCAGGAAGGTGGCTCTTGGAGCGGGACGGAGCCTTCCTGCCAAG GTGACCTTTGACCTGTACCCCCAGGTCAGATCCTGGTCTTCCATCCTACTGTCTTCTCTCCCCACCTCAACCCTGCTCTTTCCTCACTTTGTTTAAACCTCCCTGTACAACTATCTCACTTCTGAGCCTTTTATACCCTGGAAACCCATGATCCCCCGTCTCTTTGGTCACTGTATCCCTGACACTCCCAGACATTTGACCTCATTTCTGACTCTCCCAG ACTCCTTCATGTACGACACCCCTCAAGAGGTGGCCGAAGCTTTCCTGTCTTCCCTGACAGAGACCATAGAAGGAGTCGATGCTGAGGATGGGCACGGCCCAG GTTTGAAGACAGAGAAGGGAGGCAGGGCAGGGAACTGGGGGAAAATGGAGAAGGGACAGAACTGTTAATGCTGGAGCCTGAGCCACTCTCCTGGCACCCAG GGGAACAACAGAAGCGGAAGATCGTCCTGGACCCTTCAGGCTCCATGAACATCTACCTGGTGCTAGATGGATCAGACAGCATTGGGGCCAGCAACTTCACAGGAGCCAAAAAGTGTCTAGTCAACTTAATTGAGAAG GTGGAATCCTCCTATCCCTGAACTCGGGGGAATGGAATCTCGCTGATCTTCCAGGACTAGCTCCCTGATCATTCCAGCCCCTCTGAACAACAGGGCCCCAGGAAAATCTCCAGGTCCTATTCTGTCCTCCTTCCCTTTTACTTGAAGCAGTTTCTTGACTGGTAATTCCTCCATGAACCTCAGCCCTTGAGCCTCTTACTGAGAGCCTCCCTGTCCCAGCAAAGTCGCTGAAATCTCCCAATCACAGTATTCTATTTTCAATGCCATGGCGCCTTGTTCTCCTCACCCACAG GTGGCAAGTTATGGTGTGAAGCCAAGATATGGTCTAGTGACATATGCCACATACCCCAAAATTTGGGTCAAAGTGTCTGAAGCAGACAGCAGTAATGCAGACTGGGTCACGAAGCAGCTCAATGAAATCAATTATGAAG GTCAGAGGTTAGGGAATGGTGGGAGGTTCACTTTGGGGTCAGGAGGTTCAGGGTGGAGGGGGTCATGAGACTACCTTGAGGGCGACAGGGAGGACCACTTTGTAGTCAAAAGTTGAACAGCAGGATCGTTGGGCAATGGAGGTTAGTGGGAACCTGTTGGGGGCTGGAAGGGCCACTTTGTGGTCAAAGGGAAGTCCGTGTAATGATGATTAACTTAAAAAGTTGAAAGATGTGGGATTTCAGTTGCAGATTGGTCTCTGGGGTTAAAAGATGGCTTGGAAGACCAGGTGAGGTGATGGTCTCTTCCCTCTCCACAG ACCACAAGTTGAAGTCAGGGACTAACACCAAGAAGGCCCTCCAGGCAGTGTACAGCATGATGAGCTGGCCAGATGACGTCCCTCCTGAAGGCTGGAACCGCACCCGCCATGTCATCATCCTCATGACTGATG GTCAGAAGGGACCTCTCTCCTGTCCCAGCCTCCCCACCTTCTCAGACCAGCATGTGGCCCTTAAGTCCACTTGTAACACTATACCCATGGTTGGGGCCCTGAATGTGACTCATAGCTGGCTGTTCATCTCTCCTGTGACCCTTCATAAGGAATTCTTCCTAAGCCCTGTGATCAACTATCTCTAACCCTTCCTCAACTTGCTCACCCTGCCATGTGTATCCCTGCCTTTAGCCAGTTTATCTTCCTTATCTCCTACCCTCATGGTCCTGTCTCTTCTGCAG GATTGCACAACATGGGCGGGGACCCAATTACTGTCATTGATGAGATCCGGGACTTGCTATACATTGGCAAGGATCGCAAAAACCCAAGGGAGGATTATCTGG GTGAGTAACCTGCCTAGGACCCAGCACCCCACTTCCTCAGGGCTTGGACCCTCATCCTTCCTTTTTATCCCTCAG ATGTCTATGTGTTTGGGGTCGGGCCTTTGGTGAACCAAGTGAACATCAATGCTTTGGCTTCCAAGAAAGACAATGAGCAACATGTGTTCAAAGTCAAGGATATGGAAAACCTGGAAGATGTTTTCTACCAAATGATCG GTAGGGAGATACAAGGGAATAAAGAACACAACTCTCCTCAGGTTCCCCTGAAGTAATTCATTCTTCCTCTACACCTGAAGCTCTAGTTGCCTGGAAAGCCTTCTTCATTCCTCCTTCTCTACCTCAGTGTCACTATTCTTGTTTCCTGGCACTGTTCACTTAACCTTAGAATCACAGAGCTCTGAGCACTTCAGAGATCTTTCTATAGTCCTACATTTGACACGTGGAAACAGAAGCCAAAGGAGGTCAAGGGACAGCAAGTTAGCAACAAGGGTGGGCTTGAAAACAGCCAGGCCTCTGACAGCTTGATCCCAAGTTCTTTCCCTTTTCAGTCCACCATAGCAGTTTTCTCCTAACACGAGGAAACAAATACCCGTGGTCTTTCCCTTTCTCCTTTTGGGCCTTTGCTCCCCATAGACTCCTACCCAAAAGGCTGCTGCCATTTGGGAATGAAGTGTTCCGAGTTTTCAGCACATTCTCCTTCTCTGCCAG ATGAAAGCCAGTCTCTGAGTCTCTGTGGCATGGTTTGGGAACACAGGAAGGGTACCGATTACCACAAGCAACCATGGCAGGCCAAGATCTCAGTCATT GTAAGCACAGAATCCCAGTAGTGGGGACTTGGGGGAGGTGAGGTCAAGGTGAAATGGGAGTAGGGGAAGGAAAAAATGGCCATAAGAGATGGTGGTTTGTGAAAGTTGAGCTTTCCCTCTCTACTGTTGTGTCCCCAG CGCCCTTCAAAGGGACACGAGAGCTGTATGGGGGCTGTGGTGTCTGAGTACTTTGTGCTGACAGCAGCACATTGTTTCACTGTGGATGACAAGGAACACTCAATCAAGGTCAGCGTAG GTAAGGATGCAACTGAAGGTCCTGGGCTGCACCTATGCTCTCCAGGCAACACCTCCCACTTTCTACAGATCCTACACTCCACCCATCCTCAATGCAGCCCCATTCCTTGCACCCCAGACCAGTCAGGGATGGGGGAAGACGTGAAGTTAGGAATGACACGGGGCCAGAGGCAGGAAGCTGCCCACAAAGAGGTGGTACCTACTCTCCTACTTCAG GAGGGGAGAAGCGGGACCTGGAGATAGAAGTAGTCCTATTTCACCCCAACTACAACATTAATGGGAAAAAAGAAGCAGGAATTCCTGAATTTTATGACTATGACGTTGCCCTGATCAAGCTCAAGAATAAGCTGAAATATGGCCAGACTATCAG GTGAGAGCGTCCAGATCCCTGAGGAAAGGCTGGGAAAGGCTGGAGGACTGGGGTGAGGAGCAGGCCTGGTTTGCTGTTCTCCTTGTCCTTTATAG GCCCATTTGTCTCCCCTGCACCGAGGGAACAACTCGAGCTTTGAGGCTTCCTCCAACTACCACTTGCCAGCAACAAA GTAAGACATACTTGGCAAGAGGATAAGGATGAGATCCCAAGAGACAAGTGGGGCATGAGAGGGAGGTGCAATAGGAAGAGATGATGCCTGGCCCAGAACCTAGCTCTAGAAGGGCTTAGGGGACATCTACTGAGTGACAAAGGCAATGGGGAGATGACAGTGGTGGGAGCAGCTGAAGTGACGCAGTCTATTCGTCCAG AGGAAGAGCTGCTCCCTGCACAGGATATCAAAGCTCTGTTTGTGTCTGAGGAGGAGAAAAAGCTGACTCGGAAGGAGGTCTACATCAAGAATGGGGATAAG GTGAGAAACGGGCATCCTAAGGAGGCACTCTAGGCCCCAATCCTTCCTAAGCCACTTCTGTTCATTACTTCTCCATGCTTCCCACCTCCCCTACAG AAAGGCAGCTGTGAGAGAGATGCTCAATATGCCCCAGGCTATGACAAAGTCAAGGACATCTCAGAGGTGGTCACCCCTCGGTTCCTTTGTACTGGAGGAGTGAGTCCCTATGCTGACCCCAATACTTGCAGAG GTGAGAGAATGCTCTTTGGTTGTGCTACAAGTGCCCAAGGCCCAACAGTCCTTTTCTCTACAGCTTCTCCTCTCCTTGCAG GTGATTCTGGCGGCCCCTTGATAGTTCACAAGAGAAGTCGTTTCATTCAA GTGAGTCCTCCCTTTCCTATCTGGGGAGATGCCAAGTGGTCAGCATGGGCCCCAAAGCAGGAAAGCTCAATGCATGTGGCTAGTAATTCGAGGTAGGCAGAGCCTGCCTCACCTTAGGACCGCATGTCTTGCCTGCGTGTGTCAAGAACGAGGCTGAGCTGGGTCCCTAGTCTGATTCCTTTAGGTCAGCTAAGACACAAGCAGGAACAGCCATGCTTCCAGGATTAGGAATTCTACTGAATGATCCATGGCACCCCACTGCCTCTGCAG GTTGGTGTAATCAGCTGGGGAGTAGTGGATGTCTGCAAAAACCAGAAGCGGCAAAAGCAGGTACCTGCTCACGCCCGAGACTTTCACATCAACCTCTTTCAAGTGCTGCCCTGGCTGAAGGAGAAACTCCAAGATGAGGATTTGGGTTTTCTATAA GGGGTTTCCTGCTGGACAGGGGCGTGGGATTGAATTAAAACAGCTGCGACAACA CCTGTGTTCCAGATCCTTTTGGGGCAAGGGAGTGGGGAACAGGCACTGGCCATGTTGTTACACTGAGATCAAACCTGACAGCCGTTTTTAAAGGTTTAACCCCAATCCCAAGTGCTGAAAAACCAGAGGCTGAGGGAGATGTGTAAGCTTCCACCTCAGTGTTTTACTGAGACCAGCATTGGGGCATATGAGGCACAAGGAATCCAGCTCTGTTCCCTAGAAGCCATCCACAAGGTTTTCCTTGTAGACGTCATCACTGTAGACAATCTGGGTCCTCTTGTCCCGGTGGCAACCCTTAGGGCTGTTCTGGACAGCTAGGGAGGGAGGAGAGGAACAGTTAAGGTCTAAAGGAGATCATAGAACAGACCCTGAGGCTGACTCCTGACCACCTCACTCCTGGCCACTGGCCCCTGGAAGCCCAGTTTCCACGCTGCCCTCTGGTGGCCAGGATGGCCTGTCTTCCTTAGCTCCTTTGTGCCAACCCATGGCCAAGAAAAGTATAAGTGGACATTTTGATGAATGTTTTGTTCTTAGAAAAATCCCAAATGTCATTGTTGAGACACGTGAATGATATTAACCCACTACTTACAGTCAGTATGTCA(서열번호 428)

[0102] "Complex" refers to a combination of two or more molecules whose interactions depend on intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and π-effects. In one embodiment, the complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, the complex comprises one or more polypeptides that associate to form a base editor (e.g., a base editor including a nucleic acid programmable DNA binding protein such as Cas9 and a deaminase) and a polynucleotide (e.g., guide RNA). In one embodiment, the complex is held together by hydrogen bonds. It should be understood that one or more components of a base editor (e.g., a deaminase or a nucleic acid programmable DNA binding protein) may be associated covalently or non-covalently. As an example, the base editor may include a deaminase covalently linked to a nucleic acid programmable DNA binding protein (e.g. by a peptide bond). Alternatively, the base editor may include a nucleic acid programmable DNA binding protein that is associated non-covalently with the deaminase (e.g., wherein one or more components of the base editor are supplied trans- and associated directly or through another molecule such as a protein or nucleic acid). In one embodiment, one or more components of the complex are held together by hydrogen bonds.

[0103] 'Cytosine' or '4-aminopyrimidine-2(1 H )-on' is It refers to a purine nucleobase having the structure and the molecular formula C4H5N3O corresponding to CAS number 71-30-7.

[0104] 'Cytidine' is It refers to a cytosine molecule attached to a ribose sugar via a glycosidic bond, having the structure and corresponding to CAS number 65-46-3. Its molecular formula is C9H 13 It is N3O5.

[0105] 'Cytidine base editor (CBE)' refers to a base editor containing cytidine deaminase. Non-limiting examples of amino acid sequences in the cytidine deaminase base editor are BE4max (SEQ No. 3658), YE1-BE4 (SEQ No. 3659), YE2-BE4 (SEQ No. 3660), YEE-BE4 (SEQ No. 3661), EE-BE4 (SEQ No. 3662), R33A-BE4 (SEQ No. 3663), R33A+K34A-BE4 (SEQ No. 3664), APOBEC3A(A3A)-BE4 (SEQ No. 3665), APOBEC3B(A3B)-BE4 (SEQ No. 3666), APOBEC3G(A3G)-BE4 (SEQ No. 3667), AID-BE4 (SEQ No. 3668), CDA-BE4 (SEQ No. 3669), FERNY-BE4 (SEQ No. 3670), evolved APOBEC3A (eA3A)-BE4 (SEQN 3671), AALN-BE4 (SEQN 3672), BE4max modified to SpCas9-NG (SEQN 3673), YE1-SpCas9-NG (YE1-NG) (SEQN 3674), YE2-SpCas9-NG (SEQN 3675), YEE-SpCas9-NG (SEQN 3676), EE-SpCas9-NG (SEQN 3677), R33A+K34A-SpCas9-NG (SEQN 3678), YE1-CP1028 (YE1-BE4-CP1028 or YE1-CP) (SEQN 3679), YE2-CP1028 Includes the amino acid sequences of (YE2-BE4-CP1028)(SEQ No. 3680), YEE-CP1028 (YEE-BE4-CP1028)(SEQ No. 3681), EE-CP1028 (EE-BE4-CP1028)(SEQ No. 3682), R33A+K34A-CP1028 (R33A+K34A-BE4-CP1028)(SEQ No. 3683), BE4max(containing cleft enzyme)(SEQ No. 3702), BE4(SEQ No. 3703), BE4 containing His tag (SEQ No. 3704), BE4max(SEQ No. 3705), AncBE4max 689(SEQ No. 3706), and AncBE4max 687(SEQ No. 3707).

[0106] 'Cytidine base editor (CBE) polynucleotide' means a polynucleotide encoding CBE. Non-limiting examples of polynucleotide sequences encoding cytidine deaminase base editors include sequences encoding BE4max (SEQ No. 3721), AncBE4max689 (SEQ No. 3722), and AncBE4max687 (SEQ No. 3723).

[0107] "Cytidine deaminase" or "cytosine deaminase" means a polypeptide or a fragment thereof capable of deaminating cytidine or cytosine. In embodiments, cytidine or cytosine is present within the polynucleotide. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout this application. Petromizone Marinus Cytosine Deaminase 1 ( petromyzon marinus Cytidine deaminase 1, PmCDA1) (SEQ NOs 13–14), activation-induced cytidine deaminase (AICDA) (SEQ NOs 15–21), and APOBEC (SEQ NOs 12–61) are exemplary cytidine deaminases. Additional exemplary cytidine deaminase (CDA) sequences are provided in the sequence list as SEQ NOs 62–66 and SEQ NOs 67–189. Non-limiting examples of cytidine deaminases include those described in these disclosures, PCT / US20 / 16288, PCT / US2018 / 021878, 180802-021804 / PCT, PCT / US2018 / 048969, PCT / US2016 / 058344, PCT / US2020 / 062428 and PCT / US2019 / 033848, the whole of which is incorporated herein by reference for all purposes. Non-limiting examples of cytidine deaminase amino acid sequences include rat APOBEC1 (SEQ No. 3684), human APOBEC1 (SEQ No. 3685), human APOBEC3 (SEQ No. 3686), human APOBEC3B (SEQ No. 3687), human APOBEC3G (SEQ No. 3688), evoAPOBEC3A(eA3A)(SEQ No. 3689), evoCDA(SEQ No. 3690), evoAPOBECl(SEQ No. 3691), YE1(SEQ No. 3692), YE2(SEQ No. 3693), YEE(SEQ No. 3694), EE(SEQ No. 3695), R33A(SEQ No. 3696), R33A+K34A(SEQ No. 3697), AALN(SEQ No. 3698), FERNY (sequence number 3699), evoFERNY (sequence number 3700), APOBEC (sequence number 3724), Anc686 APOBEC (sequence number 3725), human APOBEC-3G D316R_D317R (sequence number 5726), human APOBEC-3G chain A (sequence number 3727), human APOBEC3-G chain A D120R_D121R (sequence number 3728),Mouse APOBEC3 (Sequence No. 3729), Rat APOBEC3 (Sequence No. 3730), Rhesus APOBEC-3G (Sequence No. 3731), Chimpanzee APOBEC-3G (Sequence No. 3732), Green monkey APOBEC-3G (Sequence No. 3733), Human APOBEC-3G (Sequence No. 3734), Human APOBEC-3F (Sequence No. 3735), APOBEC-3B (Sequence No. 3736), Rat APOBEC-3B (Sequence No. 3737), Cattle APOBEC-3B (Sequence No. 3738), Chimpanzee APOBEC-3B (Sequence No. 3739), Gorilla APOBEC-3C (Sequence No. 3740), Human APOBEC-3A (Sequence No. 3741), Rhesus APOBEC-3A (Sequence No. Includes the amino acid sequences of 3742), bovine APOBEC-3A (SEQ No. 3743), human APOBEC-3H (SEQ No. 3744), human APOBEC-3D (SEQ No. 3745), rat ABOPEC1 (SEQ No. 3746), Anc689 APOBEC (SEQ No. 3747), Anc687 APOBEC (SEQ No. 3748), Anc686 APOBEC (SEQ No. 3749), Anc655 APOBEC (SEQ No. 3750), and Anc733 APOBEC (SEQ No. 3751).

[0108] 'Cytidine deaminase polynucleotide' refers to a polynucleotide encoding cytidine deaminase. Non-limiting examples of polynucleotide sequences encoding a cytidine deaminase domain include sequences encoding rat APOBEC1 (SEQ No. 3709), Anc689 APOBEC (SEQ No. 3710), Anc687 APOBEC (SEQ No. 3711), Anc686 APOBEC (SEQ No. 3712), Anc655 APOBEC (SEQ No. 3713), Anc733 APOBEC (SEQ No. 3714), rat APOBEC1 (SEQ No. 3715), Anc689 APOBEC (SEQ No. 3716), Anc687 APOBEC (SEQ No. 3717), Anc686 APOBEC (SEQ No. 3718), Anc655 APOBEC (SEQ No. 3719), and Anc733 APOBEC (SEQ No. 3720).

[0109] 'Cytosine deaminase activity' refers to a step that catalyzes the deamination of cytosine or cytidine. In one embodiment, a polypeptide having cytosine deaminase activity converts an amino group to a carbonyl group. In one embodiment, the cytosine deaminase converts cytosine to uracil (i.e., C to U) or 5-methylcytosine to thymine (i.e., 5mC to T). In some embodiments, the cytosine deaminase as provided herein has increased cytosine deaminase activity compared to a reference cytosine deaminase (e.g., at least 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or more).

[0110] As used herein, the terms 'deaminase' or 'deaminase domain' refer to a protein or a fragment thereof that catalyzes a deamination reaction.

[0111] "Detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.

[0112] "Detectable label" means a composition that, when linked to a target molecule, makes the latter detectable through spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metallic beads, colloidal particles, fluorescent dyes, electron density reagents, enzymes (e.g., commonly used in enzyme-linked immunosorbent assays [ELISA]), biotin, digoxigenin, or haptens.

[0113] "Disease" means any pathological condition or disorder that impairs or interferes with the normal function of a cell, tissue, or organ. Exemplary diseases include therapeutic diseases in which the activity and / or expression of intracellular complement factor B [CFB] polypeptides are reduced, including the introduction of alterations to intracellular complement factor B [CFB] polynucleotides. In some cases, a disease is a condition associated with the inappropriate activation of the subject's complement system. Non-limiting examples of diseases associated with the inappropriate activation of the complement system include blood disorders, transplant or graft rejection, inflammatory diseases or disorders, eye diseases or disorders, kidney diseases or disorders, heart diseases, respiratory diseases or disorders, autoimmune disorders, inflammatory bowel diseases or disorders, arthritis, neurodegenerative diseases or disorders, musculoskeletal diseases or disorders associated with inflammation, disorders affecting the integumentary system, diseases or disorders affecting the central nervous system, diseases or disorders affecting the circulatory system, diseases or disorders affecting the gastrointestinal system, diseases or disorders affecting the thyroid, chronic pain, allergies, and lung diseases. Additional non-limiting examples of diseases associated with inappropriate activation of the complement system include paroxysmal nocturnal hemoglobinuria [PNH], atypical hemolytic syndrome [aHUS], Hellep syndrome, autoimmune hemolytic anemia, graft rejection, ischemia / reperfusion injury, graft injury, hyperacute rejection, graft rejection or failure, acute antibody-mediated rejection, chronic inflammation, chronic allograft angiopathy, chronic rejection of a graft or graft, age-related macular degeneration (e.g., wet or dry age-related macular degeneration), diabetic retinopathy, glaucoma, uveitis, autoimmune diseases, myasthenia gravis, neuromyelitis optica (NMO), renal disease, membranoproliferative glomerulonephritis [MPGN] (e.g., MPGN type I, II, or III), IgA nephropathy [IgAN], primary membranous nephropathy, C3 glomerulopathy, proteinuria, Neurodegenerative disease, neuropathic pain, sinusitis, nasal polyposis, cancer, sepsis, respiratory distress syndrome, anaphylaxis,Infusion reactions, respiratory diseases or disorders (e.g., asthma or chronic obstructive pulmonary disease (COPD), idiopathic pulmonary fibrosis or asthma), Th2-related disorders (e.g., disorders associated with high levels or high activation of Th2 subtype CD4+ helper T cells), disorders associated with high levels or inappropriate activation of Th17 subtype CD4+ helper T cells, inflammatory bowel diseases (e.g., Crohn's disease or ulcerative colitis), inflammatory skin diseases, chronic inflammatory diseases, psoriasis, atopic dermatitis, systemic scleroderma, sclerosis, Behcet's disease, dermatomyositis, polymyositis, multiple sclerosis (MS), dermatitis, meningitis, encephalitis, uveitis, osteoarthritis, lupus nephritis, rheumatoid arthritis (RA), Shoren syndrome, vasculitis, central nervous system (CNS) inflammatory disorders, chronic hepatitis, chronic pancreatitis, glomerulonephritis, sarcoidosis, thyroiditis, Pathological immune response to tissue / organ transplantation, bronchiolitis, hypersensitivity pneumonitis, idiopathic pulmonary fibrosis [IPF], periodontitis, gingivitis, disorders associated with excessive or inappropriate activity of IgE-producing cells, neuromyelitis optica, pemphigus pseudopemphigus, pulmonary fibrosis (e.g., idiopathic pulmonary fibrosis), radiation-induced lung injury, allergic bronchopulmonary aspergillosis, hypersensitivity pneumonitis, eosinophilic pneumonia, interstitial pneumonia, sarcoidosis, Wegener's granulomatosis, bronchiolitis obliterans, allergic rhinitis, inflammatory joint conditions (e.g., arthritis such as rheumatoid arthritis or psoriatic arthritis, juvenile chronic arthritis, spondyloarthritis, Reiter's syndrome, or gout), dermatomyositis, polymyositis, chronic muscle inflammation, pemphigus, systemic lupus erythematosus, dermatomyositis, scleroderma, scleroderma, Sjögren's syndrome, chronic urticaria, demyelinating disease, amyotrophic lateral sclerosis, chronic pain, stroke, Allergic neuritis, Huntington's disease, Alzheimer's disease, Parkinson's disease, circulatory system diseases, polyarteritis nodosum, Wegener's granulomatosis, giant cell arteritis, Churg-Strauss syndrome, microscopic polyangiitis, Henoch-Schönlein purpura,Takayasu's arteritis, Kawasaki disease, Behcet's disease, ulcerative colitis, thyroiditis (e.g., Hashimoto's thyroiditis, Graves' disease, postpartum thyroiditis), myocarditis, hepatitis (e.g., hepatitis C), pancreatitis, glomerulonephritis (e.g., membranoproliferative glomerulonephritis or membranous glomerulonephritis), pancreaticitis, eye disorders, choroidal neovascularization (CNV), retinal neovascularization (RNV), ocular inflammation, retinopathy of prematurity, proliferative vitreoretinopathy, uveitis, keratitis, conjunctivitis and scleritis, geographic atrophy, conjunctivitis, keratitis, scleritis, iritis, iridociliitis, ciliitis, flatulitis, choroiditis, persistent asthma and allergic asthma. In some cases, the disease is selected from neurological disorders such as glaucoma, diabetic retinopathy, age-related macular degeneration, amyotrophic lateral sclerosis [ALS], multiple sclerosis [MS], Alzheimer's disease, and various tauopathy.

[0114] "Dual editing activity" or "dual deaminase activity" means having adenosine deaminase and cytidine deaminase activities. In one embodiment, a base editor having dual editing activity has both A→G and C→T activities, wherein the two activities are approximately equal to each other or differ by about 10% or 20%. In another embodiment, the dual editor has A→G activity that is about 10% or less or greater than 20% of C→T activity. In yet another embodiment, the dual editor has A→G activity that is about 10% or less or less than 20% of C→T activity. In some embodiments, the adenosine deaminase variant has mainly cytosine deaminase activity and has little to no adenosine deaminase activity. In some embodiments, the adenosine deaminase variant has cytosine deaminase activity and has no significant or detectable adenosine deaminase activity. Non-limiting examples of proteins having dual deaminase activity include those described in these disclosures, International Patent Applications WO 2024 / 040083 and WO 2022 / 204574, the entirety of which is incorporated herein by reference for all purposes.

[0115] "Effective dose" refers to an amount of agent (e.g., base editor, cell) as described herein, which is an amount necessary to improve the symptoms of the disease compared to an untreated patient or a healthy individual, i.e., a healthy individual, or an amount of agent sufficient to induce a desired biological response. The effective dose of the active compound(s) used to carry out the embodiments of the present disclosure for the therapeutic treatment of the disease depends on the method of administration, the age, weight, and overall health condition of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. Such an amount is referred to as the "effective" dose. In one embodiment, the effective dose is an amount of the base editor of the present disclosure sufficient to introduce a modification to the target gene within a cell (e.g., in vitro or in vivo cell). In one embodiment, the effective dose is an amount of the base editor required to achieve a therapeutic effect. These therapeutic effects do not need to be sufficient to alter pathogenic genes in all cells of the target, tissue, or organ, but are sufficient to alter pathogenic genes in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the target, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of the disease.

[0116] "Fragment" means a part of a polypeptide or nucleic acid molecule. This part contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the total length of the reference nucleic acid molecule or polypeptide. The fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. In some embodiments, the fragment is a functional fragment.

[0117] "Guide polynucleotide" means a polynucleotide or a polynucleotide complex that is specific to a target sequence and can form a complex with a polynucleotide programmable nucleotide binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA [gRNA]. The gRNA may exist as a complex of two or more RNAs or as a single RNA molecule.

[0118] 'Hybridization' refers to hydrogen bonding, which can be Watson-Crick, Hoogstein, or reverse Hoogstein hydrogen bonds between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that form a pair through the formation of hydrogen bonds.

[0119] In the context of Factor B, ‘inappropriate activation’ means any increase in complement activation associated with a disease or disorder. In one embodiment, inappropriate activation is an increase or elevation in activation locally (e.g., in organs or tissues such as the central nervous system or eyes) or systemically compared to a healthy standard (e.g., a healthy subject). In some cases, ‘inappropriate activation’ is activation associated with chronic inflammation in the subject (e.g., lasting longer than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 weeks). In some cases, inappropriate activation is activation that occurs directly to the subject’s tissues, cells, or organs and / or activation that causes unwanted damage to the subject’s tissues, cells, or organs. In the embodiments, a disease or disorder associated with inappropriate activation of the complement system may be treated by any method or composition provided herein for reducing or eliminating the expression and / or activity of Factor B polypeptide. In one embodiment, complement activation is detected by measuring levels of factor B polypeptide and / or cleaved factor B polypeptide (e.g., Ba fragment or Bb fragment), wherein inappropriate activation can be determined by high levels of factor B polypeptide and / or cleaved factor B polypeptide compared to a healthy reference subject.

[0120] 'Increase' means a change in amount of at least 10%, 25%, 50%, 75%, or 100%, or about 1.5 times, about 2 times, about 3 times, about 4 times, about 5 times, about 6 times, about 7 times, about 8 times, about 9 times, about 10 times, about 15 times, about 20 times, about 25 times, about 30 times, about 35 times, about 40 times, about 45 times, about 50 times, or about 100 times.

[0121] The terms 'inhibitor of base repair', 'inhibitor of base repair', 'IBR', or their grammatical equivalents refer to proteins capable of inhibiting the activity of nucleic acid repair enzymes, e.g., base excision repair enzymes.

[0122] "Intein" is a fragment of protein that can self-excise and connect the remaining fragment [extein] by peptide bonds in a process known as protein splicing. The process in which the intein self-excises and connects the remaining parts of the protein is referred herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the intein is a trans-splicing intein (also referred to as "split intein"). In the case of a trans-splicing intein, the full-length polypeptide is separated into two separate fragments, the C-terminus of the separated intein (N-intein) is fused to the N-terminus of the separated intein, and the N-terminus of the remaining C-terminus is fused to the C-terminus of the separated intein (C-intein). Without being bound by theory or mechanism of action, contacting two polypeptide sequences together results in the excision of entain and the junctioning of the two polypeptide sequences to form a full-length polypeptide sequence. In the embodiments, contacting two polypeptide fragments fused to an entain fragment or a peptide derived from an entain fragment, respectively, is associated with greater intracellular catalytic activity (e.g., deamination of nucleobases in the polynucleotide sequence) than observed when two polypeptide fragments are contacted together in a cell and no entain fragment is present. Non-limiting examples of N-entain and C-entain sequences are 표 A 또는 표 B It includes sequences that share at least 85% sequence identity with respect to the amino acid sequences or functional fragments listed in ).

[0123]

[0124]

[0125] The terms “isolated,” “purified,” or “biologically pure” refer to a substance that is free to varying degrees from the components typically associated with it as found in nature. “Isolated” indicates the degree of separation from the original source or surroundings. “Purified” indicates a higher degree of separation than isolation. A “purified” or “biologically pure” protein is a protein that is substantially free of other substances so that any impurities do not substantially affect the biological properties of the protein or cause other adverse effects. That is, the nucleic acid or peptide of the present disclosure is purified when produced by recombinant DNA techniques in which cellular material, viral material, or culture medium is substantially absent, or when chemically synthesized in which chemical precursors or other chemicals are substantially absent. Purity and homogeneity are generally determined using analytical chemistry techniques, for example, polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term “purified” may be indicated by the nucleic acid or protein producing essentially a single band on an electrophoretic gel. In the case of proteins that can be modified, for example, by phosphorylation or glycosylation, different modifications can produce different isolated proteins that can be purified individually.

[0126] "Isolated polynucleotide" means a nucleic acid molecule that is devoid of a gene located adjacent to a gene in the naturally occurring genome of the organism from which the nucleic acid molecule of the present disclosure originated. Accordingly, the term includes recombinant DNA that is incorporated, for example, into a vector; into a self-replicating plasmid or virus; or into the genomic DNA of a prokaryote or eukaryote; or exists as a separate molecule (e.g., cDNA or a genomic or cDNA fragment generated by PCR or restriction nuclease digestion) independent of other sequences. Additionally, the term includes recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence as well as an RNA molecule transcribed from a DNA molecule.

[0127] "Isolated polypeptide" means a polypeptide of the present disclosure isolated from a naturally associated component. Generally, the polypeptide is isolated when at least 60% by weight is absent from a protein and a naturally occurring organic molecule that naturally associates with it. In the embodiments, the formulation is at least 75%, at least 90%, or at least 99% by weight of the polypeptide of the present disclosure. The isolated polypeptide of the present disclosure may be obtained, for example, by extraction from a natural source, by the expression of a recombinant nucleic acid encoding such polypeptide, or by chemical synthesis of the protein. Purity may be measured by any suitable method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0128] As used herein, the term “linker” refers to a molecule that connects two moieties. In one embodiment, the term “linker” refers to a covalent linker (e.g., a covalent bond) or a non-covalent linker.

[0129] "Marker" means any protein or polynucleotide having an alteration in expression, numerical value, structure, or activity associated with a disease or disorder. In the embodiments, the disease or disorder is associated with the inappropriate activation of the complement system. In some cases, the marker is a factor B polynucleotide or polypeptide.

[0130] The term 'mutation' as used herein refers to the substitution of one residue with another within a sequence, e.g., a nucleic acid or amino acid sequence, or the deletion or insertion of one or more residues within a sequence. A mutation is generally described herein by identifying the position of a residue within the sequence following the original residue and identifying the identity of the newly substituted residue. Various methods for creating the amino acid substitutions (mutations) provided herein are well known in the art, for example, [Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th Provided by Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).

[0131] The terms 'nucleic acid' and 'nucleic acid molecule' as used herein refer to compounds comprising nucleobases and acidic moieties, e.g., nucleosides, nucleotides, or polymers of nucleotides. Generally, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules, wherein adjacent nucleotides are connected to each other via phosphodiester links. In some embodiments, 'nucleic acid' refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, 'nucleic acid' refers to an oligonucleotide chain comprising three or more individual nucleotide residues. The terms 'oligonucleotide' and 'polynucleotide' as used herein may be used interchangeably to refer to polymers of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, 'nucleic acid' encompasses RNA as well as single and / or double-stranded DNA. Nucleic acids may occur naturally in the context of, for example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. Meanwhile, nucleic acid molecules may be non-naturally occurring molecules, for example, recombinant DNA or RNA, artificial chromosomes, genetically engineered genomes, or fragments thereof, or synthetic DNA, RNA, or DNA / RNA hybrids, or may include non-naturally occurring nucleotides or nucleosides. Additionally, the terms 'nucleic acid', 'DNA', 'RNA', and / or similar terms include nucleic acid analogs, for example, analogs having a backbone other than a phosphodiester backbone. Nucleic acids may be purified from natural sources, generated using recombinant expression systems and optionally purified, chemically synthesized, etc.Where appropriate, for example, in the case of chemically synthesized molecules, nucleic acids may include chemically modified bases or sugars, and nucleoside analogs such as analogs having backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated. In some embodiments, the nucleic acid is a natural nucleoside (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothimidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) and / or modified phosphate groups (e.g., phosphorothioates and 5'-. N - Phosphoamidite linkage) or includes it.

[0132] The terms ‘nuclear localization sequence’, ‘nuclear localization signal’, or ‘NLS’ refer to an amino acid sequence that promotes the influx of a protein into the cell nucleus. Nuclear localization sequences are known in the art and, for example, are described in the international PCT application No. PCT / EP2000 / 011690 filed on November 23, 2000, published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference to the disclosure of an exemplary nuclear localization sequence. In other embodiments, the NLS is, for example, an optimized NLS described in the literature [Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172]. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (Sequence No. 190), KRPAATKKAGQAKKKK (Sequence No. 191), KKTELQTTNAENKTKKL (Sequence No. 192), KRGINDRNFWRGENGRKTR (Sequence No. 193), RKSGKIAAIVVKRPRK (Sequence No. 194), PKKKRKV (Sequence No. 195), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (Sequence No. 196), PKKKRKVEGADKRTADGSEFESPKKKRKV (Sequence No. 328) or RKSGKIAAIVVKRPRKPKKKRKV (Sequence No. 329).

[0133] The terms 'nucleobase,' 'nitrogenous base,' or 'base,' used interchangeably herein, refer to nitrogen-containing biological compounds that ultimately form nucleosides, which are the components of nucleotides. The ability of nucleobases to form base pairs and stack one on top of another directly leads to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases—adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)—are called primary or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other modified (non-primary) bases. Non-limiting exemplary modified nucleobases may include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine may be generated through the presence of mutagens, and both may be generated through deamination (replacing an amine group with a carbonyl group). Hypoxanthine may be modified from adenine. Xanthine may be modified from guanine. Uracil may result from the deamination of cytosine. A 'nucleoside' consists of a nucleobase and a pentose sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (5-methyluridine, m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine.Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A 'nucleotide' consists of a nucleobase, a pentose sugar (ribose or deoxyribose), and at least one phosphate group. Non-limiting examples of modified nucleobases and / or chemical modifications with modified nucleobases include pseudo-uridine, 5-methylcytosine, 2'-. O- Methyl-3'-phosphonoacetate, 2'- O- MethylthioPACE(2'- O -methyl thioPACE, MSP), 2'- O- Methyl-PACE(2'- O -methyl-PACE, MP), 2'-fluoroRNA (2'-fluoroRNA, 2'-F-RNA), constrained ethyl [constrained ethyl, S-cEt], 2'- O- methyl(2'- O -methyl, 'M'), 2'- O -methyl-3'-phosphorothioate(2'- O -methyl-3'-phosphorothioate, 'MS'), 2'- O- Methyl-3'-thiophosphonoacetate(2'- O It may include -methyl-3'-thiophosphonoacetate, 'MSP'), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine.

[0134] The term 'nucleic acid programmable DNA binding protein' or 'napDNAbp' may be used interchangeably with 'polynucleotide programmable nucleotide binding domain' to refer to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), which guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein may associate with a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 cleftase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi).Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Includes Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Type II Cas effector protein, Type V Cas effector protein, Type VI Cas effector protein, CARF, DinG, homologs thereof, or modified or genetically engineered versions thereof. Other nucleic acid programmable DNA-binding proteins are also within the scope of this disclosure, but may not be specifically listed in this disclosure. For example, the full contents of each of these are incorporated herein by reference [Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033]; and the reference [Yan et al.See , "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271]. Exemplary nucleic acid programmable DNA-binding proteins and nucleic acid sequences encoding nucleic acid programmable DNA-binding proteins are provided in the sequence list as SEQ NOs 197–231, 232–245, 254–257, 260, and 378. In some embodiments, napDNAbp is (CRISPR-associated system) Cas9 nuclease, e.g., Streptococcus pyogenes (. Streptococcus pyogenes Cas9(Csnl) of ) (e.g., sequence number 197), Neisseria meningitidis ( Neisseria meningitidis )'s Cas9 (NmeCas9, sequence number 208), Nme2Cas9 (sequence number 209), Streptococcus constellatus ( Streptococcus constellatusIt is )(ScoCas9) or a derivative thereof (e.g., a sequence having at least about 85% sequence identity with Cas9, e.g., Nme2Cas9 or spCas9). Additional non-limiting examples of nucleic acid programmable DNA-binding proteins are those designed using artificial intelligence and disclosed or referenced in the literature [Rufflow, et al., "Design of highly functional genome editors by modeling of the universe of CRISPR-Cas Sequences," bioRxiv, posted April 22, 2024, doi: 10.1101 / 2024.04.22.590591], the entire disclosure of which is incorporated herein by reference for all purposes. In some embodiments, napDNAbp is OpenCRISPR-1 or a variant thereof (e.g., a variant containing a D10A amino acid change and / or lacking N-terminal methionine). Additional non-limiting examples of nucleic acid programmable DNA-binding proteins include those disclosed in International Patent Application PCT / US2019 / 047996.

[0135] The terms ‘nuclear base editing domain’ or ‘nuclear base editing protein’ as used herein refer to a protein or enzyme capable of catalyzing nuclear base modification in RNA or DNA, such as deamination from cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine) and from adenine (or adenosine) to hypoxanthine (or inosine), as well as non-templated nucleotide addition and insertion. In some embodiments, the nuclear base editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase or cytidine deaminase or cytosine deaminase).

[0136] As used herein, "obtaining" in "obtaining an agent" includes synthesizing, purchasing, or acquiring an agent by other means.

[0137] 'OpenCRISPR-1 polypeptide' means a protein or a fragment thereof having an amino acid sequence having at least about 85% amino acid sequence identity with respect to SEQ ID NO. 3568, which associates with a nucleic acid such as a guide nucleic acid or guide polynucleotide that guides napDNA bp to a specific nucleic acid sequence. Further details regarding the OpenCRISPR-1 polypeptide are disclosed in the literature [Rufflow, et al., "Design of highly functional genome editors by modeling of the universe of CRISPR-Cas Sequences," bioRxiv posted April 22, 2024, doi: 10.1101 / 2024.04.22.590591], the full disclosure of which is incorporated herein by reference for all purposes.

[0138] 'OpenCRISPR-1 polynucleotide' means a nucleic acid molecule encoding an OpenCRISPR-1 polypeptide, as well as introns, exons, 3' untranslated regions, 5' untranslated regions, and regulatory sequences or fragments thereof associated with their expression. In embodiments, the OpenCRISPR-1 polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for OpenCRISPR-1 expression. An exemplary OpenCRISPR-1 nucleotide sequence is provided in SEQ ID NO. 3569.

[0139] A guide RNA suitable for use in combination with an OpenCRISPR-1 polypeptide in various embodiments contains a scaffold or a fragment thereof having at least 85% sequence identity with respect to a nucleotide sequence selected from the following sequences capable of binding to the OpenCRISPR-1 polypeptide:

[0140] GUUUUAGAGCUGUGUUGAAAAACACAGCAAGUUAAAAUAAGGCUUUGUCCGUAUCCAACUUGAAAAAGUGAGCACCGAUUCGGUGC(Sequence No. 3570);

[0141] GUUUUAGAGCUGGAAACAGCAAGUUAAAAUAAGGCUUUGUCCGUAUCCAACUUGAAAAAGUGAGCACCGAUUCGGUGC(SEQ No. 3571); and GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(SEQ No. 3572).

[0142] “Subject” or “Patient” means a mammal, including but not limited to human or non-human mammals. In embodiments, mammals are cattle, horses, dogs, sheep, rabbits, rodents, non-human primates, or cats. In one embodiment, “Patient” refers to a mammalian subject with a higher-than-average likelihood of developing a disease or disorder. Exemplary patients may be humans, non-human primates, cats, dogs, pigs, cattle, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that may benefit from the therapy disclosed herein. Exemplary human patients may be males and / or females.

[0143] "Patients requiring this" or "objects requiring this" are referred to in this institution as patients who have been diagnosed with a disease or disability, who suffer from or are at risk of suffering from a disease or disability, who have been determined in advance to suffer from a disease or disability, or who are suspected of suffering from a disease or disability.

[0144] The terms 'pathogenic mutation', 'pathogenic variant', 'disease-causing mutation', 'disease-causing variant', 'harmful mutation', or 'predisposing mutation' refer to genetic alterations or mutations that are associated with a disease or disorder, or increase an individual's vulnerability or predisposition to a specific disease or disorder. In some embodiments, a pathogenic mutation comprises at least one wild-type amino acid substituted by at least one pathogenic amino acid within a protein encoded by a gene.

[0145] The terms 'protein', 'peptide', 'polypeptide', and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide(amide) bonds. Proteins, peptides, or polypeptides may be naturally occurring, recombinant, or synthetic, or any combination thereof.

[0146] The term 'fusion protein' as used herein refers to a hybrid polypeptide comprising protein domains from at least two different proteins.

[0147] The term 'recombinant' as used herein in the context of proteins or nucleic acids refers to proteins or nucleic acids that do not occur in nature but are products of human genetic engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence containing at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0148] 'Decrease' means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0149] "Reference" refers to standard or control conditions. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments and without limitation, the reference is an untreated cell applied to a placebo or a control vector that does not contain the test conditions, or a standard saline, medium, buffer, and / or the target polynucleotide. In an embodiment, the reference is a healthy subject or cell without inappropriate activation of the complement system. In some cases, the reference is an unedited or untreated cell (e.g., hepatocyte), tissue (e.g., components of the central nervous system or organs such as the liver or eyes), and / or subject. In an embodiment, the reference is a subject to which the composition of the present disclosure or its components have not been administered. In some cases, the reference is a subject prior to a change in treatment.

[0150] A 'reference sequence' is a defined sequence used as a standard for sequence comparison. The reference sequence may be a subset or the whole of the specified sequence, for example, a full-length cDNA or a segment of a gene sequence, or an intact cDNA or gene sequence. In the case of polypeptides, the length of the reference polypeptide sequence will generally be at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. In the case of nucleic acids, the length of the reference nucleic acid sequence will generally be at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer near or in between. In some embodiments, the reference sequence is the wild-type sequence of the target protein. In other embodiments, the reference sequence is a polynucleotide sequence encoding the wild-type protein.

[0151] The terms 'RNA-programmable nuclease' and 'RNA-guide nuclease' refer to a nuclease that forms a complex (e.g., binds to or associates with) one or more RNA(s) that are not targets for cleavage. In some embodiments, when the RNA-programmable nuclease forms a complex with RNA, it may be referred to as a nuclease-RNA complex. Generally, the bound RNA(s) are referred to as guide RNA [gRNA]. In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 nuclease, e.g., Streptococcus pyogenes ( Streptococcus pyogenes Cas9(Csnl) of ) (e.g., sequence number 197), Neisseria meningitidis ( Neisseria meningitidis )'s Cas9 (NmeCas9, sequence number 208), Nme2Cas9 (sequence number 209), Streptococcus constellatus ( Streptococcus constellatus It is )(ScoCas9) or a derivative thereof (e.g., a sequence having at least about 85% sequence identity with Cas9, e.g., Nme2Cas9 or spCas9).

[0152] "Specifically binds" means a nucleic acid molecule, polypeptide, polypeptide / polynucleotide complex, compound, or molecule that recognizes and binds to the polypeptide and / or nucleic acid molecule of the present disclosure, but does not substantially recognize and bind to other molecules in a sample, e.g., a biological sample.

[0153] "Substantially identical" means a polypeptide or nucleic acid molecule that exhibits at least 50% identity with respect to a reference amino acid sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any one of the amino acid or nucleic acid sequences described herein. In one embodiment, such sequences are at least about 60%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or even 99.99% identical to the sequence used for comparison at the amino acid level or nucleic acid level.

[0154] Sequence identity is generally measured using sequence analysis software [e.g., the Genetics Computer Group's Sequence Analysis Software Package, located at the University of Wisconsin Center for Biotechnology, 1710 University Avenue, Madison, Wisconsin, 53705; BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs]. This software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions generally include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0155] Nucleic acid molecules useful for the method of the present disclosure include any nucleic acid molecule encoding the polypeptide of the present disclosure or a functional fragment thereof. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. Polynucleotides having 'substantial identity' with respect to the endogenous sequence can generally hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful for the method of the present disclosure include any nucleic acid molecule encoding the polypeptide of the present disclosure or a functional fragment thereof. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. Polynucleotides having 'substantial identity' with respect to the endogenous sequence can generally hybridize with at least one strand of a double-stranded nucleic acid molecule. 'Hybridize' means forming a pair that forms a double-stranded molecule between complementary polynucleotide sequences (e.g., the gene described herein) or parts thereof under various degrees of strictness. (For example, see the literature [Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507]).

[0156] 'Division' means being divided into two or more fragments.

[0157] "Cut-off polypeptide" or "cut-off protein" refers to a protein provided with an N-terminal fragment and a C-terminal fragment that translate into two distinct polypeptides in nucleotide sequence(s). In some embodiments, the polypeptides corresponding to the N-terminal and C-terminal portions of the cut-off protein may be spliced ​​to form a "reconstituted" protein. In embodiments, the cut-off polypeptide is a nucleic acid programmable DNA binding protein (e.g., Cas9) or a base editor.

[0158] The term 'target site' refers to a target nucleobase within a nucleotide sequence or a modified nucleic acid molecule. In the embodiments, the modification is the deamination of the base. The deaminase may be a cytidine or adenine deaminase. A fusion protein or base editing complex containing a deaminase may include a dCas9-adenosine deaminase fusion protein, a Cas12b-adenosine deaminase fusion protein, or a base editor disclosed herein.

[0159] Terms such as 'treat,' 'treating,' and 'treatment' as used herein refer to reducing or improving a disorder and / or associated symptoms, or obtaining a desired pharmacological and / or physiological effect. It will be understood that, without exclusion, treating a disorder or condition does not require the complete elimination of the associated disorder, condition, or symptom. In some embodiments, the effect is therapeutic, that is, without limitation, the effect partially or completely reduces, alleviates, eliminates, weakens, improves, reduces, or cures the intensity of the disease and / or adverse symptoms caused by the disease. In some embodiments, the effect is preventive, that is, the effect prevents or prevents the occurrence or recurrence of the disease or condition. To this end, the method disclosed herein comprises the step of administering a therapeutically effective amount of a composition as described herein.

[0160] 'Uracil glycosylation enzyme inhibitor' or 'UGI' refers to an agent that inhibits the uracil-excision repair system. Base editors containing cytidine deaminase convert cytosine to uracil and then to thymine through DNA replication or repair. In various embodiments, uracil DNA glycosylation enzymes (UGIs) prevent base excision repair that converts U back to C. In some cases, contacting cells and / or polynucleotides with the UGI and the base editor prevents base excision repair that converts U back to C. An exemplary UGI includes the following amino acid sequence.

[0161] >splP14739IUNGI_BPPB2 Uracil-DNA glycosylation enzyme inhibitor

[0162] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 231).

[0163] In some embodiments, the agent that inhibits the uracil resection repair system is a uracil stabilizing protein [USP]. For example, refer to WO 2022015969 A1 incorporated herein by reference.

[0164] The term 'vector' as used herein refers to a means of producing a transformed cell by introducing a nucleic acid molecule into a cell. Vectors include plasmids, transposons, phages, viruses, liposomes, lipid nanoparticles, and episomes.

[0165] The ranges provided herein are understood as abbreviations for all values ​​within the range. For example, the range from 1 to 50 is understood to include any number, combination of numbers, or sub-range in the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0166] In any definition of a variable of the present invention, a reference to a list of chemical functional groups includes the definition of any single functional group or a combination of the listed functional groups. A reference to an embodiment regarding a variable or aspect of the present invention includes an embodiment as any single embodiment or an embodiment combined with any other embodiment or part thereof.

[0167] All terms are intended to be understood as they would be understood by a person skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by a person skilled in the art to which this disclosure pertains.

[0168] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, unless the context clearly indicates otherwise, the singular forms 'a,' 'an,' and 'the' should be understood to include the plural. In this application, the use of 'or' means 'and / or' unless otherwise stated. Furthermore, the use of the term 'including' as well as other forms, such as 'includes,' 'includes,' and 'included,' is not limited.

[0169] As used in this specification and claims(s), the words 'comprising' (and any form of including, such as 'comprising' and 'to include'), 'having' (and any form of having, such as 'to have' and 'to have'), 'comprising [including]' (and any form of including, such as 'comprising' and 'to include'), or 'containing' (and any form of containing, such as 'to contain' and 'to contain') are inclusive or open. These expressions indicate the presence of specific elements, features, components, and / or method steps, but do not exclude the presence of other elements, features, components, and / or method steps. Any embodiment referred to as 'comprising' a specific component(s) or element(s) is considered to be 'consisting of' or 'essentially composed of' that specific component(s) or element(s) in some embodiments. Any embodiment discussed in this specification may be implemented in connection with any method or composition of this disclosure, and vice versa. Additionally, the composition of the present disclosure may be used to achieve the method of the present disclosure.

[0170] The terms 'approximately' or 'roughly' mean within an acceptable margin of error for a specific value as determined by a person skilled in the art, and this will vary in part depending on how the value is measured or determined, that is, the limitations of the measurement system.

[0171] In this specification, references to "some embodiments," "an embodiment," "one embodiment," or "other embodiments" mean that certain features, structures, or characteristics described in connection with the embodiments are included in at least some embodiments, but are not necessarily included in all embodiments of this disclosure. Brief explanation of the drawing

[0172] Fig. 1It provides a schematic diagram showing alternative pathways for complement amplification. Fig. 2 This provides a bar graph showing the percentage of base editing from maximum A to G of factor B polynucleotides measured in HEK293T cells transfected with a base editor system containing adenosine deaminase and a guide indicated on the x-axis. A base editor system containing guide sg23 and adenosine deaminase was used as a positive control for base editing. Fig. 3 This provides a bar graph showing the percentage of maximum C to T base editing of factor B polynucleotides measured in HEK293T cells transfected with a base editor system containing cytidine deaminase and a guide indicated on the x-axis. A base editor system containing guide sg23 and cytidine deaminase was used as a positive control for base editing. Fig. 4 It provides a bar graph showing the percentage of base editing from maximum A to G of factor B polynucleotides measured in HEK293T cells transfected with a base editor system containing the indicated adenosine deaminase and guide polynucleotides. Fig. 4 In this context, the term 'NHP X-reactivity' refers to 'non-human primate cross-reactivity'. A base editor system possessing NHP X-reactivity will edit both human and non-human primate factor B polynucleotides. Each rod group, from left to right, corresponds to a base editor system containing the guide polynucleotides gRNA1193, gRNA1120, gRNA1230, gRNA1217, gRNA1204, gRNA1218, gRNA1203, gRNA1202, gRNA1190, gRNA1213, gRNA1210, and sg23, respectively. The guide polynucleotide sg23 was used as a positive control. Fig. 5It provides a bar graph showing human complement factor B (hCFB) protein levels (left axis and left bar of each bar pair) in primary human hepatocytes (PHH) at day 11 (D11) after transfection (P-TF) with the indicated base editor system, and the percentage of base edits from maximum A to G of factor B polynucleotides measured in PHH at day 13 (D13) after P-TF with the indicated base editor system (right axis and right bar of each bar pair). Fig. 5 The editors listed are base editors containing the indicated TadA* adenosine deaminase domain, and the term 'NHP X-reactivity' refers to 'non-human primate cross-reactivity'. Base editor systems possessing NHP X-reactivity will edit both human and non-human primate Factor B polynucleotides. The guide sg23 was used as a positive control. The guide polynucleotide gRNA1204 targeted the human Factor B polynucleotide sequence GCTTACAATGACTGAGATCTTGG (Sequence No. 429), which is GCTTACA G There is a difference in the bolded G of the non-human primate (cyno) factor B polynucleotide sequence of TGACTGAGATCTTGG (SEQ No. 430). Protein levels were measured using an Abcam ELISA kit [Human Factor B ELISA kit (ab137973)] (range: 4.375 ng / ml–140 ng / ml, lower limit of quantitation (LLOQ): 0.8 ng / mL). Fig. 5 In the editor 'spCas9', 'spCas9' refers to the spCas9 nuclease capable of inducing double-strand breaks in DNA. Fig. 6This provides a bar graph showing the effect of guide polynucleotide spacer length on the percentage of base editing from A to G of factor B polynucleotides in HEK293T cells. Cells were base edited using a base editor system containing guide RNA containing the indicated adenosine deaminase base editor and spacers with indicated nucleotide (nt) lengths ranging from 19 to 23 nucleotides. The first five bars from the left correspond to base editor ABE8.8, which is specific for NGG PAM sequences; the second five bars from the left correspond to base editor ABE 8.13, which is specific for NGG PAM sequences; the third five bars from the left correspond to base editor ABE 8.8, which is specific for NGG PAM sequences; and the rightmost bar corresponds to ABE8.8, which is specific for NGG PAM sequences. Figs. 7a and 7b It provides bar graphs and schematic diagrams related to the optimization of guide spacer lengths. Fig. 7a It provides a bar graph showing the levels of human complement factor B [hCFB] protein in human hepatocytes (PXB cells) isolated from PXB-mice at day 11 [P-TF] after transfection with the indicated base editor system [D11] (left axis and left bar of each bar pair) and the percentage of base editing from maximum A to G of factor B polynucleotides measured in PXB cells at day 13 [D13] after P-TF with the indicated base editor system (right axis and right bar of each bar pair). Fig. 7aThe editors listed herein are base editors containing the indicated TadA* adenosine deaminase domain, the term 'NHP X-reactivity' refers to 'non-human primate cross-reactivity', and the term 'protospacer length (nt)' represents the spacer length (19–23 nucleotides) corresponding to the indicated guide polynucleotide in nucleotides (nt). A base editor system having NHP X-reactivity will edit both human and non-human primate Factor B polynucleotides. Guide sg23 was used as a positive control. The guide polynucleotide gRNA1204 targeted the human Factor B polynucleotide sequence GCTTACAATGACTGAGATCTTGG (Sequence No. 429), which is GCTTACA targeted by the guide polynucleotide gRNA1999 (a non-human primate surrogate of gRNA1204). G There is a difference in the bolded G of the non-human primate (cyno) factor B polynucleotide sequence of TGACTGAGATCTTGG (Sequence No. 430). Protein levels were measured using the Abcam ELISA kit [Human Factor B ELISA kit (ab137973)] (Range: 4.375 ng / ml–140 ng / ml, Lower Limit of Quantification [LLOQ]: 0.8 ng / mL). Protein levels were measured using the Abcam ELISA kit [Human Factor B ELISA kit (ab137973)] (Range: 4.375 ng / ml–140 ng / ml, Lower Limit of Quantification [LLOQ]: 0.8 ng / mL). Fig. 7b Is Fig. 7a Provides a schematic diagram describing the experiment used to collect the data shown in. Fig. 7b In this, 'NGS' stands for Next-Generation Sequencing. Figures 8a and 8bThe provides bar graphs and Western blot images showing complement factor B polynucleotide base editing efficiency measured in primary cyno hepatocytes (PCH) transfected with a base editor system containing one of the indicated guide polynucleotides that have cross-reactivity to adenosine deaminase and non-human primates and human factor B, or are non-human primate surrogate guide polynucleotides. Fig. 8a It provides a bar graph showing the percentage of base edits from maximum A to G of factor B polynucleotides measured in PCH transfected with a base editor system containing adenosine deaminase and a labeled guide polynucleotide. The guide polynucleotide gRNA2072 targeted the human factor B polynucleotide sequence GCTTACAATGACTGAGATCTTGG (SEQ No. 429), which is GCTTACA G There is a difference in the bolded G of the non-human primate (cyno) factor B polynucleotide sequence of TGACTGAGATCTTGG (sequence number 430). Fig. 8b The top panel provides a Western blot showing Factor B levels measured in monkey serum, PCH supernatant, humanized mouse serum, and Abcam’s human complement factor B [CFB] ELISA standard using an anti-complement factor B monoclonal antibody (Ab-CFB). Fig. 8b The bottom panel of is Fig. 8a Provides a bar graph showing the cyano-CFB protein level normalized to the pretreatment level for the corresponding cell. Figures 9a–9dIt provides bar graphs and plots representing the maximum percentage of base editing from A to G of factor B polynucleotides measured in primary human hepatocytes [PHH] or human liver cancer cells (HepG2 cells) transfected with a base editor system containing labeled mRNA encoding adenosine deaminase and labeled guide polynucleotides. The base editor encoded by the mRNA molecule (e.g., m3534 / MRNA3534) referenced in the figure is Table 9 It is listed in. Fig. 9a This provides a bar graph showing the percentage of base edits from maximum A to G of factor B polynucleotides measured in PHH transfected with the indicated base editor system. The base editor system sg23 / m3534 was used as a positive control. Figs. 9b~9d It provides a plot showing the percentage of base editing from maximum A to G of factor B polynucleotides measured in HepG2 cells transfected with a base editing system containing different doses of the respective guide polynucleotides TSBTx3826, TSBTx3837, and TSBTx3935 and a constant dose of the indicated mRNA molecule encoding the base editor. The base editor encoded by the mRNA molecule referenced in the figure (e.g., MRNA3534) Table 9 It is listed in. Figs. 10a and 10b FRG administered a base editor system containing the guide polynucleotide gRNA1193 and the ABE8.8 adenosine deaminase base editor. TM Provides a bar graph showing the percentage of base edits from maximum A to G, insertion / deletion (indel) mutation rates, and protein levels of human complement factor B [hCFB] in liver-humanized mice. Fig. 10a FRGs transfected with a base editor system containing an adenosine deaminase base editor and the terminal-modification guide polynucleotide gRNA1193 at 2 mg / kg (mpk) or 0.3 mpk.TM We provide a bar graph showing the percentage of maximum A to G base edits and indel mutation rates of hCFB measured in liver-humanized mice. Tris buffered saline (TBS) was administered to mice as a negative control. Fig. 10b FRGs transfected with a base editor system containing an adenosine deaminase base editor and 2 mg / kg (mpk) of the terminal-modification guide polynucleotide gRNA1193. TM Provides a bar graph showing the concentration (Conc.) of hCFB protein (hCFB Pr.) measured in liver-humanized mice. Fig. 10b Each group of the three bars corresponds to the values ​​measured from left to right at day 0 (D0) before transfection (i.e., 'before administration'), day 7 after transfection, and the end of the experiment (i.e., 'final') on day 14 after transfection. Fig. 10b The top panel of shows unnormalized protein concentration, and Fig. 10b The bottom panel shows protein concentrations normalized to day 0 (D0) concentrations. Fig. 11 FRGs transfected with a base editor system containing the ABE8.8 adenosine deaminase base editor and 2 mg / kg (mpk) or 0.3 mpk of the end-modification guide polynucleotide gRNA1193. TM Provides a series of plots showing a negative correlation between serum hC3 and hCFB protein levels measured in liver-humanized mice. The x-axis represents the days elapsed since transfection, indicating the time point of measurement. Fig. 11 The arrows extending from each curve indicate the axis corresponding to each curve. Mice were administered Tris-buffered saline (TBS) as a negative control. Figures 12a and 12bFRG administered a base editor system containing the guide polynucleotide TSBTx3826 having an NLS nucleotide modification scheme and one of the indicated adenosine deaminase base editors (i.e., ABE8.8, ABE8.20, or ABE9.52). TM Percentage of base editing from A to G of human complement factor B [hCFB] in liver-humanized mice ( Fig. 12a ) and protein levels( Fig. 12b Provides a bar graph representing ). A base editor system containing the guide polynucleotide sg23 was used as a positive control. Figures 12a and 12b In this, the term 'Mod Scheme selection' refers to the nucleotide modification scheme of the guide polynucleotide, the term 'BE selection (NLS)' refers to the selection of a base editor using a guide polynucleotide having an NLS nucleotide modification scheme, the term 'pre-administration' refers to a value measured before administering the base editor system to the mouse, and the terms '0.5 mpk' and '0.3 mpk' refer to the doses of the guide polynucleotide administered to the mouse. The TSBTx3826 guide polynucleotide exhibited cross-reactivity (i.e., targeting of base editing) for both human and cyno CFB polynucleotides, and the target base editing site was a splice site at the 5' end of exon 3 of the factor B polynucleotide. Figures 12a and 12b In this, the term 'ABE9.52' refers to a base editor containing the adenosine deaminase domain of TadA*8.20 with amino acid changes of V82T, Y147T, and Q154S. Fig. 13 FRG administered a base editor system containing a guide polynucleotide TSBTx3826 targeting a splice site at the 5' end of exon 10 and one of the indicated adenosine deaminase base editors. TMProvides a bar graph showing the levels of human complement factor B [hCFB] exons indicated in mRNA collected from liver-humanized mouse tissues. Measurements were performed on day 14 after administration of the base editor system. Fig. 13 The mRNA level was normalized to the mRNA level measured for the actin beta (ACTB) gene. Fig. 13 In this context, the term 'ALAS1(sg23)' refers to the level of 5'-aminolevulinate synthase 1 (5'-Aminolevulinate Synthase 1, ALAS1) transcripts in mice administered a base editor system containing the guide polynucleotide sg23. Fig. 13 In this, the term 'ABE9.52' refers to a base editor containing the adenosine deaminase domain of TadA*8.20 with amino acid changes of V82T, Y147T, and Q154S. Figures 14a and 14b FRG administered a base editor system containing the guide polynucleotide TSBTx3837 having the HM01 nucleotide modification scheme and one of the indicated adenosine deaminase base editors (i.e., ABE8.8 or ABE8.20). TM Percentage of base editing from A to G of human complement factor B [hCFB] in liver-humanized mice ( Fig. 14a ) and protein levels( Fig. 14b Provides a bar graph representing ). A base editor system containing the guide polynucleotide sg23 or TSBTx3826 having an NLS nucleotide modification method was used as a control. Figures 14a and 14bIn this, the term 'selection of modification method' refers to the nucleotide modification method of the guide polynucleotide, the term 'BE selection (NLS)' refers to the selection of a base editor using a guide polynucleotide having an NLS nucleotide modification method, the term 'before administration' refers to a value measured before administering the base editor system to the mouse, and the terms '0.5 mpk' and '0.3 mpk' refer to the doses of the guide polynucleotide administered to the mouse. The TSBTx3837 guide polynucleotide targeted hCFB, and the target site of the TSBTx3837 guide polynucleotide differed from the corresponding cyno-CFB target site by one (1) nucleotide, and since the target base editing site was located at the splice site of the 3' end of exon 11 of the factor B polynucleotide, no cross-reactivity (i.e., targeting of base editing) to the cyno-CFB polynucleotide was observed. Figures 14a and 14b In this, the term 'ABE9.52' refers to a base editor containing the adenosine deaminase domain of TadA*8.20 with amino acid changes of V82T, Y147T, and Q154S. Fig. 15 FRG administered a base editor system containing a guide polynucleotide TSBTx3837 targeting the splice site at the 3' end of exon 11 and one of the indicated adenosine deaminase base editors. TM Provides a bar graph showing the levels of human complement factor B [hCFB] exons indicated in mRNA collected from liver-humanized mouse tissues. Measurements were performed on day 14 after administration of the base editor system. Fig. 15 The mRNA level was normalized to the mRNA level measured for the actin beta (ACTB) gene. Fig. 15In this context, the term 'ALAS1(sg23)' refers to the level of 5'-aminolevulinate synthase 1 (ALAS1) transcripts in mice administered a base editor system containing the guide polynucleotide sg23. Figures 16a and 16b FRG administered a base editor system containing the guide polynucleotide TSBTx3835 having a terminal-modification nucleotide modification scheme and one of the indicated adenosine deaminase base editors (i.e., ABE8.13 or ABE9.52). TM Percentage of base editing from A to G of human complement factor B [hCFB] in liver-humanized mice ( Fig. 16a ) and protein levels( Fig. 16b Provides a bar graph representing ). A base editor system containing the guide polynucleotide sg23 or TSBTx3826 having an NLS nucleotide modification method was used as a control. Figures 16a and 16b In this, the term 'selection of modification method' refers to the nucleotide modification method of the guide polynucleotide, the term 'BE selection (NLS)' refers to the selection of a base editor using a guide polynucleotide having an NLS nucleotide modification method, the term 'pre-administration' refers to a value measured before administering the base editor system to the mouse, and the terms '0.5 mpk' and '0.3 mpk' refer to the doses of the guide polynucleotide administered to the mouse. The TSBTx3835 guide polynucleotide targeted hCFB and exhibited cross-reactivity (i.e., targeting of base editing) to human and cyno CFB polynucleotides, and the target base editing site was a splice site at the 3' end of exon 16 of the factor B polynucleotide. Figures 16a and 16b In this, the term 'ABE9.52' refers to a base editor containing the adenosine deaminase domain of TadA*8.20 with amino acid changes of V82T, Y147T, and Q154S. Fig. 17FRG administered a base editor system containing a guide polynucleotide TSBTx3835 targeting a splice site at the 3' end of exon 16 and one of the indicated adenosine deaminase base editors. TM Provides a bar graph showing the levels of human complement factor B [hCFB] exons indicated in mRNA collected from liver-humanized mouse tissues. Measurements were performed on day 14 after administration of the base editor system. Fig. 17 The mRNA level was normalized to the mRNA level measured for the actin beta (ACTB) gene. Fig. 17 In this context, the term 'ALAS1(sg23)' refers to the level of 5'-aminolevulinate synthase 1 (ALAS1) transcripts in mice administered a base editor system containing the guide polynucleotide sg23. Fig. 17 In this, the term 'ABE9.52' refers to a base editor containing the adenosine deaminase domain of TadA*8.20 with amino acid changes of V82T, Y147T, and Q154S. Fig. 18This provides a plot showing the percentage of editing from maximum A to G of complement factor B polynucleotides in primary human hepatocytes [PHH] or primary cyno hepatocytes [PCH] as indicated, transfected with a base editor system containing the indicated guide polynucleotide and adenosine deaminase base editor. A base editor system containing adenosine deaminase and the guide polynucleotide sg23 was used as a positive control. Cells were transfected with mRNA encoding the guide polynucleotide and the base editor at a mass ratio of 1 to 3 (1:3). The TSBTx3837 guide polynucleotide was used in combination with the base editor ABE8.20, and the guide polynucleotide contained the HM01 nucleotide modification scheme. The TSBTx3826 guide polynucleotide was used in combination with a base editor containing the TadA*8.20 adenosine deaminase domain with amino acid changes V82T, Y147T, and Q154S, and the guide polynucleotide included an NLS nucleotide modification scheme. Figs. 19a–19d It provides schematic diagrams and plots. Fig. 19a This provides a schematic diagram showing the sequences of polynucleotide constructs used to compare the efficacy of guide polynucleotides targeting human complement factor B [CFB] and / or non-human primate CFB for base editing. Binding sites of guide polynucleotides targeting human CFB polynucleotides (i.e., 'CFB guide-human') and non-human primate CFB polynucleotides (i.e., 'CFB guide-NHP') are indicated. Fig. 19a In this, the term '10 bp' refers to a 10-nucleotide spacer, and the term '30 bp random spacer' refers to a randomized sequence of 30 nucleotides. Fig. 19a The two nucleotide sequences shown in are inversely complementary to each other. Fig. 19a The upper nucleotide sequence in is CATGGCAGGCCAAGATCTCAGTCATTGTAAGCACAGAATCCCATATGGAAGGTCATTAGCTCCGGCAAGCAATCATGGCAGGCCAAGATCTCAGTCACTGTAAGCACAGAATCCCA (Sequence No. 431), and the amino acid sequences are HGRPRSQSL (Sequence No. 432) and AQNPIWKVISSGKQSWQAKISVTVSTES (Sequence No. 433). Fig. 19a In the amino acid sequence, the term '*' indicates a stop codon, and 'CFB insert' indicates that a polynucleotide construct has been inserted into the genome of HEK293T cells. Figs. 19b–19d represents the percentage of base editing in three distinct experiments (i.e., Batch 1, Batch 2, and Batch 3, respectively) at the 'CFB-guided-human' and 'CFB-guided-NHP' sites in HEK293T cells transfected with a base editing system containing adenosine deaminase and indicated doses of guide gRNA2067 (TSBTx3837; targeting the CFB-guided-human site) or guide gRNA2072 (TSBTx2072; targeting the CFB-guided-HNP site). Fig. 20 is a labeled base editor system containing a guide polynucleotide and adenosine deaminase (i.e., Table 12.1A We provide a bar graph showing the editing of the complement factor B [CFB] TATA box from A to G in human liver cancer cells (HepG2 cells) transfected with samples 1 to 16) described in [the document]. Cells were transfected with mRNA encoding a guide polynucleotide and a base editor at a total saturation dose of 800 ng in a mass ratio of 1:3. The CFB TATA box was located at positions -157 to -151 relative to the CFB start codon. A base editor system containing an adenosine deaminase base editor and guide sgRNA_088 (sg23) was used as a positive control. Fig. 20 Each bar is in order from left to right Table 12.1AIt corresponds to a base editor system containing the base editors listed in ). Figures 21a and 21b is a base editor system containing a guide polynucleotide and adenosine deaminase (i.e., Table 12.1B Listed in Fig. 21a Samples 1 to 8 and Fig. 21b Human liver cancer cells (HepG2 cells) transfected with samples 1 to 3) Fig. 21a We provide a bar graph showing the editing of the complement factor B [CFB] start codon from A to G in primary human hepatocytes [PHH] monolayer cells. An adenosine deaminase base editor and a base editor system containing guide sgRNA_088 (sg23) were used as positive controls. Figures 21a and 21b Below the x-axis of the bar graph, CFB amino acid changes corresponding to the base edits corresponding to each bar (e.g., M1T, G2E, G2R, L5P, or S3P) are listed. Cells were transfected with a total saturation dose of 800 ng by mixing the guide polynucleotide and mRNA encoding the base editor in a mass ratio of 1:3. Figs. 22a and 22b It provides a bar graph and schematic diagram related to the TATA box and initiation codon disruption of complement factor B [CFB] in primary human hepatocytes [PHH] for protein knockdown. Fig. 22a is a base editor system containing adenosine deaminase and guide polynucleotides (i.e., Table 12.1C Provides a bar graph showing the human complement factor B [hCFB] protein level in PHH on day 12 (D12) after transfection with samples 1 to 16 described in [P-TF] (left axis and left bar of each bar pair) and the percentage of base editing from maximum A to G of factor B polynucleotides measured in PXB cells on day 13 (D13) of P-TF using a base editing system (right axis and right bar of each bar pair). Fig. 22aA base editor system containing adenosine deaminase and the guide polynucleotide sgRNA_088 (sg23) was used as a positive control for base editing. Fig. 22b Is Fig. 22a Provides a schematic diagram explaining the experiment used to collect the data presented in. Fig. 22b In this context, the term 'MC' refers to medium replacement, and the term 'NGS' refers to next-generation sequencing. Fig. 22a The data was not normalized, but a similar pattern was observed even when the data was normalized to the protein level before transfection (i.e., day 0). Fig. 23 This provides a set of bar graphs showing high editing and excellent reduction of complement factor B [CFB] protein levels in co-cultures of primary human hepatocytes [PHH] transfected with a base editor system containing one of the indicated PAM specificities (e.g., NGC, NGG, or NGA) and one of the indicated guide polynucleotides targeting the CFB initiation codon for base editing. A base editor system containing adenosine deaminase and the guide polynucleotide sg23 or gRNA1193 (TSBTx3826) was used as a positive control. Fig. 23 In this context, the term 'dABE (-) control' refers to a defective or 'dead' adenosine base editor. Fig. 23 The top panel displays data collected using cells from a donor designated as 'JGC', and Fig. 23The bottom panel of the document shows data collected using donor cells designated as 'MRW'. The base editing system administered mRNA encoding guide polynucleotides and adenosine deaminase to the cells at a total saturation dose of 800 ng. None of the guide polynucleotides cross-reacted with non-human primate target sites (i.e., guide polynucleotides target human CFB polynucleotides for base editing but not cyano CFB polynucleotides). The target site of gRNA 3657 is TGCTCCC CAT GGC G TT G GAAGGC (Sequence No. 434), wherein the corresponding non-human primate [NHP] target site is TGCTCCC CAT GGC A TT A It is GAAGGC (Sequence No. 435), where the nucleotides in bold represent the parts of the human gRNA 3657 target site that differ from the corresponding NHP target site, and the nucleotide corresponding to the CFB start codon is underlined. The target site of gRNA 3658 is T TGCTCCC CAT GGC G TT G GAAGG (Sequence No. 436), but the corresponding non-human primate [NHP] target site is C TGCTCCC CAT GGC A TT A It is GAAGG (Sequence No. 437), where the nucleotides in bold indicate the part of the human gRNA 3658 target site that differs from the corresponding NHP target site, and the nucleotide corresponding to the CFB start codon is underlined. The target site of gRNA 3660 is CCC CAT GGC G TT G GAAGGCAGGA (Sequence No. 438), wherein the corresponding non-human primate [NHP] target site is CCC CAT GGC A TT A GAAGGCAGGA (sequence number 439), where the nucleotides in bold represent parts of the human gRNA 3660 target site that differ from the corresponding NHP target site, and the nucleotides corresponding to the CFB start codon are underlined. Fig. 23 In all groups consisting of three bars, the first two bars from the left correspond to the hCFB protein levels measured on day 7 (D7) and day 13 (D13) after transfection, normalized to the levels measured before transfection (i.e., day 0), and the right bar corresponds to the A-to-G edit of the CFB polynucleotide measured on day 13. Fig. 24 It provides a schematic diagram illustrating the guidance-dependent and guidance-independent deamination of nucleotides of polynucleotides and lists representative methods that can predict or measure them. Fig. 24 For all purposes, the entire disclosure is taken from the literature [Kempton and Lei, Science, 364:234-236 (2019)], the whole of which is incorporated herein by reference. Fig. 25 is related to the disruption of the CFB start codon in primary human hepatocyte co-cultures Fig. 23 Provides a bar graph that represents the data in an alternative way. Fig. 25 Each pair of bars represents the hCFB protein level and the edit from A to G from left to right. Fig. 25 In this, 'dABE(-) control' refers to a negative control base editor system containing a catalytically inactive base editor. Fig. 25 The base editor system of (i.e., samples 1 through 9) Table 12.1D It is listed in. FIGS. 26a and FIGS. 26b It provides bar graphs and schematic diagrams related to the functional evaluation of guide polynucleotides targeting initiation codons in long-term HepG2 culture systems. Fig. 6 silver Table 12.1E A bar graph is provided showing the base editing rate for the indicated target site achieved using an active or inactive base editing system corresponding to samples 1 to 8 described in [the document]. Fig. 26b Is Fig. 26a Provides a schematic diagram describing the experiment used to collect the data presented in. Fig. 26b In this, 'MC' indicates a medium change, 'TF' indicates transfection using a base editor system, and 'NGS' indicates next-generation sequencing. Fig. 27 silver Table 12.1E Provides a plot showing human complement factor B [hCFB] protein levels in a long-term HepG2 culture system containing cells transfected with an active (left panel) or inactive (right panel) base editing system corresponding to samples 1 to 8 described in [the document], and targets the sites indicated for editing on the indicated dates post-transfection [post-TF]. Fig. 27 In the left panel, the lines at day 10 after transfection correspond from top to bottom to Sample 1, Sample 7, Sample 2, Sample 6 / Sample 5, Sample 3, and Sample 4, and the third line from the bottom at day 22 corresponds to Sample 5. Fig. 27 In the right panel, the lines at day 10 after transfection correspond from top to bottom to Sample 1, Sample 6, Sample 2, Sample 3, Sample 7, Sample 4, and Sample 5. The 'inactive editor' contains a catalytically inactive base editor. Specific details for implementing the invention

[0173] Base editors, nucleases, and guide RNA [gRNA] used to edit, modify, or alter target polynucleotides are provided herein. In certain embodiments, the base editor or nuclease of the present disclosure modifies a complement factor B [CFB] polynucleotide. In certain embodiments, the base editor of the present invention introduces a stop codon modification to the CFB polynucleotide or disrupts the TATA box, initiation site, or splice site of the CFB polynucleotide. The modification is associated with a decrease in the activity or levels of the CFB polypeptide and / or polynucleotide within the cell.

[0174] The invention of the present disclosure is based at least in part on the discovery that alternative pathways of the complement system require protein factor B for complement pathway amplification and function. The present invention is also based at least in part on the discovery that base editing (e.g., disruption of splice receptors or splice donors or introduction of stop codons) can be used to reduce the expression of factor B polypeptides in cells associated with dysregulated complement systems (e.g., inappropriate activation). In particular, reducing the activity and / or expression of factor B polypeptides in subjects diagnosed with diseases or disorders associated with the overactivation of the complement system can be an effective therapeutic strategy. This reduction in activity and / or expression can be achieved using any base editing system and / or nuclease and method provided herein. Accordingly, the present disclosure features compositions and methods for editing factor B polynucleotides. Editing for factor B polynucleotides is associated with a reduction in the expression and / or activity of factor B polypeptides in the target's cells, tissues, and / or body fluids, as well as a reduction in symptoms associated with the overactivation of the target's complement system or other pathogenic activations.

[0175] Accordingly, as described in the examples provided herein, a base editor system that disrupts complement system activity by destroying the function of Factor B at the gene level has been successfully developed. The destruction of Factor B was performed through the silencing / knockout of the Factor B gene.

[0176] In an embodiment, the method of the present disclosure involves interfering with the splicing of a factor B polynucleotide transcript. For example, a base editor or base editor system provided herein may be used to edit a nucleus of a splice receptor located at the 5' of an exon of a factor B polynucleotide. In some embodiments, the target sequence is a splice receptor in an intron portion adjacent to an exon of a factor B polynucleotide, and editing the nucleus of the splice receptor is associated with a change in the splice receptor compared to a wild-type splice receptor site. In some embodiments, deamination of an A or C nucleus at the splice receptor interferes with the splicing of mRNA transcripts during or after transcription. In some embodiments, the target suffers from dysregulation and / or hyperactivation of the complement system and any associated disease or disorder, or is likely to develop such a condition.

[0177] In some cases, the method of the present disclosure comprises the step of modifying a factor B polynucleotide to introduce a stop codon, initiation site destruction, or TATA box destruction associated with a decrease in the number or activity of the complement factor B polynucleotide and / or polypeptide. The modification may be made by a base editor system such as those described herein.

[0178] In some embodiments, the present disclosure provides a base editor that efficiently generates an intended mutation, such as a point mutation within a nucleic acid molecule (e.g., a nucleic acid within the target genome), without generating a significant number of unintended mutations, such as unintended point mutations. In some embodiments, the intended mutation is a mutation generated by a base editor containing a specific base editor (e.g., an adenosine base editor or a cytidine base editor), wherein the base editor system is specifically designed to generate the intended mutation. In some embodiments, the intended mutation is a point mutation from adenine (A) to guanine (G) within a non-coding region of a gene. In some embodiments, the intended mutation is a point mutation from cytosine (C) to thymine (T) within a non-coding region of a gene. In some embodiments, the intended mutation is a mutation in a splice receptor in an intron of a gene associated with a disease or disorder. In some cases, the intended mutation is an indel mutation. In some embodiments, the intended mutation is a point mutation from adenine (A) to guanine (G) at a splice receptor site in an intron of a gene associated with the disease or disorder. The intended mutation may include introducing a stop codon into the polynucleotide sequence. In some embodiments, the intended mutation is a mutation that interferes with the normal splicing of the intact transcript of the gene, for example, a change from A to G at a splice receptor site within an intron of a gene causing or associated with the disease. In some embodiments, the intended mutation is a mutation at a splice receptor site that interferes with the splicing of the gene transcript and results in an alternative transcript that encodes a truncated or non-functional protein product.

[0179] In some embodiments, any base editor or nuclease provided herein may produce a ratio of intended mutations to unintended mutations greater than 1:1 (e.g., intended point mutation: unintended point mutation). In some embodiments, any base editor provided herein may generate a ratio of intended mutation to unintended mutation (e.g., intended point mutation:unintended point mutation) of at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 150:1, at least 200:1, at least 250:1, at least 500:1, or at least 1000:1 or greater. there is.

[0180] In some embodiments, editing multiple nucleobase pairs in one or more genes using the method provided herein forms at least one intended mutation. In some embodiments, the formation of at least one intended mutation is present at the splice receptor site and interferes with the splicing of the mRNA transcript of the disease-associated gene. In some embodiments, the formation of at least one intended mutation results in a reduction in the activity and / or expression of the disease-associated gene. It should be understood that multiple editing can be achieved using any method provided herein or a combination of methods.

[0181] The present disclosure provides a method for treating a subject diagnosed with dysregulation and / or overactivation of the complement system or any disease or disorder associated therewith. For example, in some embodiments, a method is provided comprising the step of altering the factor B polynucleotide sequence by administering an effective amount of a nucleobase editor (e.g., an adenosine deaminase base editor or a cytidine deaminase base editor) to a subject who suffers from dysregulation and / or overactivation of the complement system or who is prone to developing it.

[0182] Complement System and Factor B

[0183] The complement system is a system composed of numerous plasma and cell-binding proteins that play important roles in both innate and adaptive immunity. The proteins of the complement system act in a series of enzymatic chain reactions through various protein interactions and cleavage phenomena.

[0184] The complement system is a component of the innate immune system and is important for the elimination of pathogens and apoptotic or dying cells. Complement activation leads to the formation of membrane attack complexes and cytolysis, the targeting of opsonized exogenous substances for phagocytosis, and the activation of inflammation and various immune components. Many complement components are circulating factors primarily produced in the liver.

[0185] The complement system plays a crucial role in defending the body against infectious agents. The complement system comprises more than 30 serum and cellular proteins involved in three major pathways known as the classical pathway, the alternative pathway, and the lectin pathway. The classical pathway is generally triggered by the binding of a complex of an antigen and an IgM or IgG antibody to C1 (although other specific activators can also initiate this pathway). Activated C1 cleaves C4 and C2 to produce C4a and C4b, in addition to C2a and C2b. C4b and C2a combine to form C3-convertase, which cleaves C3 at defined cleavage sites to form C3a and C3b. When C3b binds to C3-convertase, C5-convertase is produced, which cleaves C5 into C5a and C5b. C3a, C4a, and C5a are anaphylotoxins and mediate numerous reactions in acute inflammatory responses. C3a and C5a are also chemotactic factors that attract immune system cells such as neutrophils. Further details regarding C3 are incorporated herein by reference in their entirety for all purposes [Ricklin, et al. "Complement component C3 - The 'Swiss Army Knife' of innate immunity and host defense." Immunol Rev. . [2016 Nov; 274(1):33-58] and provided in the literature [Janssen et al., "Structures of complement component C3 provide insights into the function and evolution of immunity." Nature. 2005 Sep 22;437(7058):505-11].

[0186] Alternative path (e.g., Fig. 1(Refer to [reference]) is generally initiated and amplified by microbial surfaces and various complex polysaccharides. The alternative pathway is triggered by C3b covalently binding to the surface of a pathogen or cell. Next, Factor B binds to the surface-bound C3b, making it sensitized to cleavage by plasma Factor D. This results in the production of Ba and the active protease Bb, which remain bound to C3b to generate C3bBb, the C3 convertase of the alternative complement pathway. This initiates an amplification loop as the C3 convertase generates more C3b on the cell surface, and this process is repeated. Ultimately, the cell surface becomes saturated with C3b, accompanied by the release of the anti-inflammatory mediator C3a. Consequently, some of the C3b binds to the existing C3 convertase, which generates C3b2Bb, the C5 convertase of the alternative pathway. This cleaves C5 into C5b to form the membrane attack complex (MAC) and produces C5a, a potent inflaming mediator. Complement-mediated endothelial cell damage induces thrombophilia. This exposes subendothelial collagen, releases vWF, and forms fibrinogen. Generally, the presence of complement regulatory proteins on the cell surface prevents significant activation of complement on it. For a more detailed description of alternative pathways, the full disclosure thereof is incorporated herein by reference for all purposes [Keir, L. and Coward, RJM, 2011. Pediatr. Nephrol]. . Provided in 26, 523-533.

[0187] Complement Factor B (CFB, alternatively 'Factor B') is a serine protease and a key component of the amplification loop of the alternative complement pathway (AP). Another serine protease, Complement Factor D (CFD), cleaves CFB to form Ba and Bb. Bb forms an essential part of the AP convertase complex, which plays a role in activating central complement proteins C3 and C5 through proteolytic cleavage.

[0188] C5-convertase generated in both pathways cleaves C5 to produce C5a and C5b. C5b then combines with C6, C7, and C8 to form C5b-8, which catalyzes the polymerization of C9 to form the C5b-9 membrane attack complex (MAC), also known as the terminal complement complex (TCC). The MAC self-inserts into the target cell membrane, inducing cell lysis. The presence of even small amounts of MAC on the cell membrane can lead to various consequences in addition to apoptosis. If the TCC is not inserted into the membrane, it can circulate in the blood as soluble sC5b-9. Blood levels of sC5b-9 can serve as an indicator of complement activation.

[0189] The lectin complement pathway can be initiated by mannose-binding lectin (MBL) and MBL-associated serine protease (MASP) binding to carbohydrates. The MB1-1 gene (known as LMAN-1 in humans) encodes a type I endogenous membrane protein located in the intermediate region between the endoplasmic reticulum and the Golgi apparatus. The MBL-2 gene encodes a soluble mannose-binding protein found in serum. In the human lectin pathway, MASP-1 and MASP-2 are involved in the proteolysis of C4 and C2, producing the aforementioned C3 convertase.

[0190] Accordingly, the present disclosure provides a method for interfering with complement activation by modifying a polynucleotide encoding factor B.

[0191] Diseases and / or disorders associated with an unwanted increase in the activity of the complement system

[0192] Inappropriate activation of the complement system can cause various diseases and / or disorders in subjects. For example, inappropriate activation of the complement system in subjects damages cells, resulting in increased inflammation, the presence of autoantibodies, neurodegeneration, and microthrombosis. Inappropriate activation of the complement system is associated with damage to the nervous system (e.g., central nervous system, CNS), circulatory system, kidneys, eyes, blood cells (e.g., red blood cells, white blood cells, and platelets), and transplanted organs, as well as damage to other organs or tissues, which may be associated with the presence of microembolisms. Therefore, effective treatment of these diseases and / or disorders may include a step of reducing the activation of the complement system in organs, cells, and / or tissues by modifying the Factor B nucleotide sequence to reduce and / or eliminate the expression and / or activity of the Factor B polypeptide in the subject. In an embodiment, the organ or tissue is the eye, kidney, nervous system component, heart, or thyroid. Without being bound by theory, the levels of complement protein in the eyes may vary depending on the circulating levels of complement protein produced in the liver.

[0193] Some important indications for patients requiring treatment for inappropriate activation of the complement system (e.g., overactivation or dysregulation) include paroxysmal nocturnal hemoglobinuria (PNH), atypical hemolytic uremic syndrome (aHUS), and IC-MPGN / C3 glomerulopathy. PNH is associated with red blood cell (RBC) hemolysis, leading to anemia and thrombosis. aHUS disorders are associated with thrombocytopenia and acute renal failure caused by RBC hemolysis as well as abnormal thrombus formation in renal small blood vessels. IC-MPGN and C3 glomerulopathy are associated with renal dysfunction and end-stage renal disease caused by renal glomerular damage.

[0194] Non-limiting examples of diseases associated with inappropriate activation of the complement system include blood disorders, transplant or graft rejection, inflammatory diseases or disorders, eye diseases or disorders, kidney diseases or disorders, heart disorders, respiratory / lung diseases or disorders, autoimmune disorders, inflammatory bowel diseases or disorders, arthritis, neurodegenerative diseases or disorders, musculoskeletal diseases or disorders associated with inflammation, disorders affecting the integumentary system, diseases or disorders affecting the central nervous system, diseases or disorders affecting the circulatory system, diseases or disorders affecting the gastrointestinal system, diseases or disorders affecting the thyroid, chronic pain, allergies, and lung diseases. Additional non-limiting examples of diseases associated with inappropriate activation of the complement system include acute antibody-mediated rejection, age-related macular degeneration (e.g., wet or dry age-related macular degeneration), allergic asthma, allergic bronchopulmonary aspergillosis, allergic neuritis, allergic rhinitis, Alzheimer's disease, amyotrophic lateral sclerosis, anaphylaxis, atopic dermatitis, atypical hemolytic syndrome [aHUS], autoimmune diseases, autoimmune hemolytic anemia, Behcet's disease, bronchiolitis, bronchiolitis obliterans, C3 glomerulopathy, cancer, central nervous system [CNS] inflammatory disorders, choroidal neovascularization [CNV], choroiditis, chronic allograft angiopathy, chronic hepatitis, chronic inflammation, chronic inflammatory diseases, chronic myoinflammatory disease, chronic pain, chronic pancreatitis, chronic rejection of grafts or transplants, chronic urticaria, Churg-Strauss syndrome, conjunctivitis, COVID-19, ciliary body inflammation, demyelinating disease, dermatitis, dermatomyositis, diabetic retinopathy, circulatory system disease, disorders associated with excessive or inappropriate activation of IgE-producing cells, disorders associated with high levels or inappropriate activation of Th17 subtype CD4+ helper T cells, encephalitis, eosinophilic pneumonia, eye disorders, geographic atrophy, giant cell arteritis, gingivitis, glaucoma, glomerulonephritis, glomerulonephritis (e.g., membranoproliferative glomerulonephritis or membranous glomerulonephritis),Graft rejection or failure, granulomatous polyangiitis / microscopic polyangiitis [granulomatosis with polyangiitis, microscopic, GPA / MPA], Hellep syndrome, Henoch-Schönlein purpura, hepatitis (e.g., Hepatitis C), Huntington's disease, hyperacute rejection, hypersensitivity pneumonitis, idiopathic pulmonary fibrosis [IPF], IgA nephropathy [IgAN], inflammatory bowel disease (e.g., Crohn's disease or ulcerative colitis), inflammatory joint conditions (e.g., arthritis such as rheumatoid arthritis or psoriatic arthritis, juvenile chronic arthritis, spondyloarthritis, Reiter's syndrome, or gout), inflammatory skin diseases, infusion rejection, interstitial pneumonia, iridocyclitis, iritis, ischemia / reperfusion injury, Kawasaki disease, keratitis, lupus nephritis, membranoproliferative Membranoproliferative glomerulonephritis (MPGN) (e.g., MPGN type I, II, or III), meningitis, microscopic polyangiitis, multiple sclerosis (MS), myasthenia gravis, myocarditis, nasal polyposis, neurodegenerative disease, neuromyelitis optica, neuromyelitis optica (NMO), neuropathic pain, ocular inflammation, osteoarthritis, pancreatitis, pancreatic inflammation, Parkinson's disease, paroxysmal nocturnal hemoglobinuria (PNH), planoplasmitis, pathological immune response to tissue / organ transplantation, pemphigus pseudopemphigus, pemphigus, periodontitis, persistent asthma, polyarteritis nodularis, polymyositis, primary membranous nephropathy, proliferative vitreoretinopathy, proteinuria, psoriasis, pulmonary fibrosis (e.g., idiopathic pulmonary fibrosis), radiation-induced lung injury, kidney disease, respiratory disease or disorder (e.g. For example, asthma or chronic obstructive pulmonary disease (COPD) or idiopathic pulmonary fibrosis or asthma), dyspnea syndrome, retinal neovascularization (RNV),Retinopathy of prematurity, rheumatoid arthritis [RA], sinusitis, sarcoidosis, sarcoidosis, scleritis, scleroderma, sclerodermatomyositis, sclerosis, sepsis, Sjögren's syndrome, Sjören syndrome, stroke, systemic lupus erythematosus, systemic scleroderma, Takayasu's arteritis, Th2-related disorders (e.g., disorders associated with high levels or high activation of CD4+ helper T cells of Th2 subtypes), thyroiditis (e.g., Hashimoto's thyroiditis, Graves' disease, or postpartum thyroiditis), thyroidosis, implant injury, implant rejection, ulcerative colitis, uveitis, vasculitis, and Wegener's granulomatosis. In an embodiment, the method of the present invention comprises the step of reducing complement-mediated hemolysis in a subject. Additional non-limiting examples of diseases include Creutzfeldt-Jakob disease, Pick's disease, mild cognitive impairment, fibromyalgia, frontotemporal dementia, Lewy body dementia, multiple system atrophy, chronic inflammation, demyelinating polyneuropathy, Guillain-Barré syndrome, multifocal motor neuropathy, non-alcoholic fatty liver disease (NAFLD), e.g., non-alcoholic steatohepatitis (NASH), and Stargardt macular degeneration. Paroxysmal nocturnal hemoglobinuria is associated with mutations in phosphatidyl inositol glycan anchor biosynthesis class a (PigA), and these mutations prevent GPI anchor production and the attachment of erythrocytes (RBCs) to CD59 and CD55, leading to erythrocyte lysis.

[0195] In some embodiments, inappropriate activation of the complement system is involved in the progression and pathogenesis of selected neurological diseases such as glaucoma, diabetic retinopathy, age-related macular degeneration and amyotrophic lateral sclerosis [ALS], multiple sclerosis [MS], Alzheimer's disease and various tauopathy.

[0196] Existing treatments for diseases associated with the inappropriate activation of the complement system often require regular, and in some cases invasive, administration regimens. Therefore, there is a need for improved treatments for diseases associated with the inappropriate activation of the complement system.

[0197] The methods and compositions of the present disclosure are suitable for use in embodiments for treating any of the diseases or disorders listed above associated with the inappropriate activation of the complement system. In various cases, the method comprises the step of introducing a modification to a factor B polynucleotide that results in a decrease in the expression and / or activity of the factor B polypeptide in cells.

[0198] Editing of target genes

[0199] Exemplary spacer sequences and guide polynucleotide sequences suitable for use in guide RNA that can be used to generate polynucleotide edits (e.g., introduction of stop codons, splice site disruption mutations, TATA box changes, start codon changes, etc.) described herein are as follows: Tables 1A to 2F It is listed in. To generate polynucleotide edits, cells (e.g., target-intravenous or target-derived cells) are below Tables 2A to 2FOne or more guide RNAs or fragments thereof containing one or more of the spacer sequences listed in [the example] are contacted with a nucleus-base editor polypeptide or complex containing a nucleic acid programmable DNA binding protein [napDNAbp] and one or more deaminases having cytidine deaminase and / or adenosine deaminase activity (e.g., a 'dual deaminase' having cytidine and adenosine deaminase activity). In the embodiments, the base editor and / or nucleus-intradissolvase are introduced into the cell using a polynucleotide sequence (e.g., mRNA) encoding the base editor and / or nucleus-intradissolvase. of the following Tables 1A to 1F This lists representative guide polynucleotide sequences suitable for use in the method of the present disclosure for modifying CFB polynucleotides. The following Tables 2A to 2F Lists representative guide RNA spacer sequences that can be used in various embodiments in combination with the indicated base editor. In the embodiments Tables 2A to 2F Using guide RNA containing the spacer sequences listed in Tables 2A to 2F The target sequences listed in can be targeted, and Tables 2A to 2F Any edits listed in (e.g., amino acid or nucleotide changes) may be selectively performed. In some cases, the gRNA is added directly to the cell. In some embodiments, the gRNA includes nucleotide analogs. These nucleotide analogs can inhibit the degradation of the gRNA during cellular processes. Tables 2A to 2F ... provides a target sequence to be used in gRNA. An additional suitable exemplary spacer sequence for use in the gRNA sequence to be used in the method provided herein is Tables 2A to 2F A fragment of any spacer provided in, as well as one modified to include an extension or truncation at the 3' and / or 5' ends, Tables 2A to 2FIncludes any spacer provided in. In the embodiment Tables 2A to 2F The spacer sequence of can be modified to include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotide extensions or truncations at the 3' and / or 5' end(s).

[0200] Variants of the spacer sequence provided herein are considered to have one, two, three, four, or five nucleotide changes. For example, variations in the target polynucleotide sequence within the population (e.g., single nucleotide polymorphisms) may require such changes in the spacer sequence so that the spacer in the target can bind better to the variant of the target sequence.

[0201] In various cases, it is advantageous for the spacer sequence to include 5' and / or 3''G' nucleotides. In some cases, for example, any spacer sequence or guide polynucleotide provided herein includes or further includes 5''G', in some embodiments, the 5''G' is complementary or not complementary to the target sequence. In some embodiments, the 5''G' is added to a spacer sequence that does not already contain 5''G'. For example, when a guide RNA is expressed under the control of a U6 promoter or others, it may be advantageous for the guide RNA to contain a 5' terminal 'G', because the U6 promoter prefers 'G' at the transcription initiation site (see Cong, L. et al. "Multiplex genome engineering using CRISPR / Cas systems. Science 339:819-823 (2013) doi: 10.1126 / science.1231143]). In some cases, the 5' terminal 'G' is added to the guide polynucleotide to be expressed under the control of the promoter, but is selectively not added to the guide polynucleotide when the guide polynucleotide is not expressed under the control of the promoter.

[0202] In some embodiments, the guide polynucleotide of the present disclosure comprises a spacer and a scaffold containing one of the following nucleotide modification schemes ('mod schemes'), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of nucleotide 'N', and 'Ns' represents nucleotide 'N' Indicates that it is linked to the next nucleotide by phosphorothioate (PS):

[0203] Terminal modification SpCas9 guide polynucleotide

[0204] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ ID NO: 440)

[0205] Terminal modification SaCas9 guide polynucleotide

[0206] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUsmUsmUsmU (SEQ ID NO: 441)

[0207] HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440)

[0208] HM07: mNsmNsmNsmNmNmNmNmNmNmNNNNNNNNNNNmGUUUUAGmAmGmCmUmAmGmAmAmAmUmAmGmCmAmAGUUmAAmAAmUAmAmGmGm CmUmAGUmCmCGUUAmUmCAAmCmUmUmGmAmAmAmAmAmGmUmGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440)

[0209] NLS(bpsv40): mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV(Sequence Nos 440 and 446)

[0210] LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGUGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmUsmU(Sequence No. 444)

[0211] NLS+LONGEST : mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmGmAmGmCmAmGmGmGmGmUmGmGmUsmUsmUsmU-NHC5-CrossL- CKRTADGSEFESPKKKRKV(SEQ Nos 445 and 446)

[0212] LONGEST+GOLD:

[0213] mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCU AGUCCGUUAmUmCAAmCmUmUGGACUUCGGUCCmAmAmGUGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 447)

[0214] In any of the embodiments, the number of N nucleotides (i.e., spacer sequence) of the aforementioned guide sequence may vary between 15 and 25. In some cases, the number of N nucleotides is 18, 19, 20, 21, 22, or 23.

[0215] The exemplary guide RNA sequence is as follows: Tables 1A–1F and 2A–2F Provided in. Throughout the table, ranges within guide polynucleotide names (e.g., 3–9) represent base editing ranges of exemplary base editors suitable for use with guide polynucleotides (e.g., nucleotides 3–9, where position 1 is the first nucleobase complementary to the spacer and adjacent to the proto-spacer adjacent motif).

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242]

[0243]

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258]

[0259]

[0260]

[0261]

[0262]

[0263]

[0264]

[0265]

[0266]

[0267]

[0268]

[0269]

[0270]

[0271]

[0272]

[0273]

[0274]

[0275]

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289]

[0290]

[0291]

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299]

[0300]

[0301]

[0302]

[0303]

[0304]

[0305]

[0306]

[0307]

[0308]

[0309]

[0310]

[0311]

[0312]

[0313]

[0314]

[0315]

[0316]

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328]

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335]

[0336]

[0337]

[0338]

[0339]

[0340]

[0341]

[0342]

[0343]

[0344]

[0345]

[0346]

[0347]

[0348]

[0349]

[0350]

[0351]

[0352]

[0353]

[0354]

[0355]

[0356]

[0357]

[0358]

[0359]

[0360]

[0361]

[0362]

[0363] sg23 / sgRNA_088 spacer sequence: CAGGAUCCGCACAGACUCCA (Sequence No. 3794) (Target gene: ALAS1). The target sequence corresponding to sg23 is CAGGATCCGCACAGACTCCA GGG (Sequence No. 3795), where the PAM sequence is thick It is displayed in text.

[0364] crRNA2 / gRNA2002 spacer sequence: UCCCCGUUCUCGAAGUCGUG (Sequence No. 3796). The target site corresponding to crRNA2 (TSBTx4946) is TCCCCGTTCTCGAAGTCGTG TGG (Sequence number 3797), where the PAM sequence is shown in bold.

[0365] Nuclear base editor

[0366] A nucleobase editor for editing, modifying, or altering a target nucleotide sequence of a polynucleotide is useful for the methods and compositions described herein. The nucleobase editors described herein generally comprise a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., adenosine deaminase, cytidine deaminase, or double deaminase). The polynucleotide programmable nucleotide binding domain specifically binds to the target polynucleotide sequence when combined with a bound guide polynucleotide (e.g., gRNA), thereby allowing the base editor to be localized to the target nucleic acid sequence desired to be edited.

[0367] Polynucleotide programmable nucleotide binding domain

[0368] The polynucleotide programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). The polynucleotide programmable nucleotide binding domain of the base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain comprises an endoneurase or a nucleotide degradase.

[0369] A base editor comprising a polynucleotide programmable nucleotide binding domain comprising all or part (e.g., a functional portion) of a CRISPR protein [i.e., a base editor comprising as a domain all or part (e.g., a functional portion) of a CRISPR protein (e.g., a Cas protein), also referred to as the 'CRISPR protein-derived domain' of the base editor] is disclosed herein. The CRISPR protein-derived domain incorporated into the base editor may be modified compared to the wild-type or natural version of the CRISPR protein. The CRISPR protein-derived domain may include one or more mutations, insertions, deletions, rearrangements, and / or recombinations compared to the wild-type or natural version of the CRISPR protein.

[0370] Cas proteins that can be used in the present invention include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, Includes CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 (e.g., SEQ ID NO. 232), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i and Cas12j / CasΦ, CARF, DinG, Turbo Cas9 (i.e., SpCas9 having amino acid changes of Q844R, V842L, F846Y, L847M and I852F), homologs thereof or modified versions thereof. CRISPR enzymes can direct the cleavage of one or both strands in a target sequence, such as within the target sequence and / or within the complement of the target sequence. For example, CRISPR enzymes can direct the cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.

[0371] A vector encoding a CRISPR enzyme mutated for a corresponding wild-type enzyme may be used such that the mutated CRISPR enzyme lacks the ability to cleave one or two strands of a target polynucleotide containing a target sequence. A Cas protein (e.g., Cas9, Cas12) or a Cas domain (e.g., Cas9, Cas12) may refer to a polypeptide or domain having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with respect to an exemplary wild-type Cas polypeptide or Cas domain. Cas (e.g., Cas9, Cas12) may refer to wild-type or modified forms of Cas proteins that may include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0372] In some embodiments, the CRISPR protein-derived domain of the base editor is Corynebacterium ullcerans ( Corynebacterium ulcerans )(NCBI Refs: NC_015683.1, NC_017317.1), Corynebacterium diphtheria( Corynebacterium diphtheria )(NCBI Refs: NC_016782.1, NC_016786.1), Spiroplasma cirpidicola( Spiroplasma syrphidicola )(NCBI Ref: NC_021284.1), Prevotella Intermedia( Prevotella intermedia )(NCBI Ref: NC_017861.1), Spiroplasma taiwanense( Spiroplasma taiwanense )(NCBI Ref: NC_021846.1), Streptococcus inie( Streptococcus iniae )(NCBI Ref: NC_021314.1), Beliela Baltica( Belliella baltica )(NCBI Ref: NC_018010.1), Cycloplexus torquis( Psychroflexus torquis)(NCBI Ref: NC_018721.1), Streptococcus thermophilus( Streptococcus thermophilus )(NCBI Ref: YP_820832.1), Listeria inocula( Listeria innocua )(NCBI Ref: NP_472073.1), Campylobacter jejuni( Campylobacter jejuni )(NCBI Ref: YP_002344900.1), Neisseria meningitidis( Neisseria meningitidis )(NCBI Ref: YP_002342100.1), Streptococcus paeogenes( Streptococcus pyogenes ), or Staphilococcus aureus ( Staphylococcus aureus It may include all or part of Cas9 (e.g., a functional part) in ).

[0373] Some aspects of the present disclosure provide high-fidelity Cas9 domains. High-fidelity Cas9 domains are known in the art, for example, in the literature incorporated herein by reference [Kleinstiver, BP, et.al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016)] and [Slaymaker, IM, et.al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015)]. An exemplary high-fidelity Cas9 domain is provided in the sequence list as SEQ ID NO. 233.

[0374] In some embodiments, any Cas9 fusion protein or complex provided herein comprises one or more of the D10A, N497X, R661X, Q695X and / or Q926X mutations, or a corresponding mutation in any amino acid sequence provided herein, wherein X is any amino acid.

[0375] Generally, S. Paeogenes ( S. pyogenes Cas9 proteins, such as Cas9 derived from (spCas9), require a 'protospacer adjacent motif (PAM)' or PAM-like motif, which is a DNA sequence of 2 to 6 base pairs immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of an NGG PAM sequence is required to bind to a specific nucleic acid region, where 'N' in 'NGG' is adenosine (A), thymidine (T), or cytosine (C), and 'G' is guanosine. In some embodiments, any fusion protein or complex provided herein may contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a standard [canonical] (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences are described in the art and will be obvious to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are referenced in the literature whose full contents are incorporated herein by reference [Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015)], and the literature [Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus It is described in "CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015).

[0376] In some embodiments, napDNAbp is a permutant (e.g., SEQ ID NO. 238).

[0377] In some embodiments, the polynucleotide programmable nucleotide binding domain comprises a cleavease domain. As used herein, the term "cleavease" refers to a polynucleotide programmable nucleotide binding domain comprising a nuclease domain capable of cleaving only one of the two strands in a dimerized nucleic acid molecule (e.g., DNA). For example, if the polynucleotide programmable nucleotide binding domain comprises a cleavease domain derived from Cas9, the Cas9-derived cleavease domain may comprise a D10A mutation and histidine at position 840. As another example, the Cas9-derived cleavease domain comprises an H840A mutation, while the amino acid residue at position 10 is retained as D.

[0378] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleaving domain, i.e., Cas9 is a cleaving enzyme referred to as the 'nCas9' protein (in the case of 'cleaving enzyme' Cas9, SEQ ID NO. 201). The Cas9 cleaving enzyme may be a Cas9 protein capable of cleaving only one strand of a dimerized nucleic acid molecule (e.g., a dimerized DNA molecule). In some embodiments, the Cas9 cleaving enzyme comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 cleaving enzymes provided herein. Additional suitable Cas9 cleaving enzymes will be apparent to a person skilled in the art based on the present disclosure and knowledge of the art, and are within the scope of the present disclosure.

[0379] The present invention also provides a base editor comprising a polynucleotide programmable nucleotide binding domain that is catalytically killed (i.e., incapable of cleaving the target polynucleotide sequence). For example, in the case of a base editor comprising a Cas9 domain, Cas9 may include both D10A mutations and H840A mutations. In further embodiments, the catalytically killed polynucleotide programmable nucleotide binding domain comprises point mutations (e.g., D10A or H840A) as well as deletions of all or part (e.g., functional parts) of the nuclease domain. The dCas9 domain is known in the art and is described, for example, in the literature [Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013; 152(5):1173-83], the entire contents of which are incorporated herein by reference.

[0380] The term 'protospacer adjacent motif (PAM)' or PAM-like motif refers to a DNA sequence of 2 to 6 base pairs immediately following the DNA sequence targeted by a nucleic acid programmable DNA-binding protein. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). The PAM sequence may be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGTT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAAC. Y is pyrimidine, N is any nucleotide base, and W is A or T.

[0381] The base editor provided herein may include a CRISPR protein-derived domain capable of binding a nucleotide sequence containing a standard or non-standard protospacer adjacent motif (PAM) sequence.

[0382] In some embodiments, PAM is 'NRN' PAM, where 'N' in 'NRN' is adenine (A), thymine (T), guanine (G) or cytosine (C) and R is adenine (A) or guanine (G), or PAM is 'NYN' PAM, where 'N' in NYN is adenine (A), thymine (T), guanine (G) or cytosine (C) and Y is cytidine (C) or thymine (T), for example, the entire contents of which are incorporated herein by reference [RT Walton et al., 2020, Science, 10.1126 / science.aba8853 (2020)].

[0383] Several PAM variants are as follows Table 3It is listed in.

[0384]

[0385] In some embodiments, PAM is NGC. In some embodiments, NGC PAM is recognized by a Cas9 variant. In some embodiments, the NGC PAM Cas9 variant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as 'MQKFRAER') of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135V, G1218R, R1335Q, and T1337R (collectively referred to as VRQR) of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from D1135V, G1218R, R1335E, and T1337R (collectively referred to as VRER) of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from E782K, N968K, and R1015H (collectively referred to as KHH) of saCas9 (SEQ No. 218). In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219S, R1335E, and T1337R (collectively referred to as 'MQKSER') of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9.

[0386] In some cases, the Cas9 variant has specificity for PAM 5'-NGC-3'. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135M, S1136Y, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135L, S1136Y, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135M, S1136Y, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135L, S1136Y, G1218K, E1219F, A1283D, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9.In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from A61R, L1111R, D1135L, S1136W, G1218K, E1219Q, N1317R, A1322R, R1333P, R1335Q, and T1337R of spCas9 (SEVENTEEN NO. 197) (SpRY), or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135L, S1136Q, G1218K, E1219F, E1250K, A1283D, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEVENTEEN NO. 197), or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from D1135M, S1136Y, G1218K, E1219F, E1250K, A1283D, A1322R, D1332A, R1335E, and T1337R of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9. In some embodiments, the Cas9 variant comprises one or more amino acid substitutions selected from R765A, Q768A, D1135L, S1136Y, G1218K, A1283D, E1219F, A1322R, D1332A, R1335E, and T1337K of spCas9 (SEQ No. 197) or a corresponding mutation of another Cas9.Any of the Cas9 proteins provided herein comprising SpCas9 in some embodiments are R765A, Q768A, W1126R, R1359W, E1250K, A1239T, A1239V, A1283D, R1335D, D1135L, D1135M, D1135R, D1135W, S1136H, S1136Q, S1136Y, G1218D, G1218K, G1218R, G1218E, G1218L, E1219F, E1219K, E1219N, A1322A, A1322R, A1322K, D1332A, R1335V, T1337K, T1337T, It includes any 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions of D1332A, D1135V, and T1337R.

[0387] In some embodiments, the CRISPR protein-derived domain of the base editor comprises all or part (e.g., a functional portion) of a Cas9 protein having a standard PAM sequence (NGG). In other embodiments, the Cas9-derived domain of the base editor may utilize a non-standard PAM sequence. Such sequences are described in the art and will be apparent to a person skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in the literature [Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015)], and the literature [Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015)], and the literature [RT Walton et al. “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants” Science 10.1126 / science.aba8853 (2020)], Hu et al. “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature, 2018 Apr. 5, 556(7699), 57-63], Miller et al., “Continuous evolution of SpCas9 variants compatible with non-G PAMs” Nat. Biotechnol., 2020 Apr;38(4):471-481].

[0388] A fusion protein or complex comprising NapDNAbp and cytidine deaminase and / or adenosine deaminase

[0389] Some aspects of the present disclosure provide a fusion protein or complex comprising a Cas9 domain or another nucleic acid programmable DNA binding protein (e.g., Cas12) and one or more cytidine deaminase, adenosine deaminase, or cytidine adenosine deaminase domains. It should be understood that the Cas9 domain may be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) may be fused with any cytidine deaminase and / or adenosine deaminase provided herein. The domains of the base editor disclosed herein may be arranged in any order.

[0390] In some embodiments, the fusion protein or complex comprising cytidine deaminase or adenosine deaminase and napDNAbp (e.g., Cas9 or Cas12 domain) does not include a linker sequence. In some embodiments, the linker is present between the cytidine or adenosine deaminase and the napDNAbp. In some embodiments, the cytidine or adenosine deaminase and the napDNAbp are fused through any linker provided herein. For example, in some embodiments, the cytidine or adenosine deaminase and the napDNAbp are fused through any linker provided herein.

[0391] It should be understood that the fusion protein or complex of the present disclosure may include one or more additional features. For example, in some embodiments, the fusion protein or complex may include an inhibitor, an extransport sequence such as a cytoplasmic localization sequence, a nuclear extransport sequence, or other localization sequence, as well as a sequence tag useful for solubilization, purification, or detection of the fusion protein or complex. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags also referred to as histidine tags or His tags, maltose binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), strep tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to a person skilled in the art. In some embodiments, the fusion protein or complex comprises one or more His tags.

[0392] Exemplary but non-limiting fusion proteins are described in international PCT applications PCT / US2017 / 045381, PCT / US2019 / 044935, and PCT / US2020 / 016288, each of which is incorporated herein by reference in its entirety.

[0393] Fusion protein or complex having an internal insertion

[0394] A fusion protein or complex comprising a nucleic acid programmable nucleic acid binding protein, for example, a heterologous polypeptide fused to a napDNAbp, is provided herein. The heterologous polypeptide may be fused to the napDNAbp at the C-terminus or N-terminus of the napDNAbp, or inserted into an internal location of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., cytidine or adenosine deaminase) or a functional fragment thereof. For example, the fusion protein may comprise Cas9 or Cas12 (e.g., Cas12b / C2c1), a deaminase located on the side of the N-terminal fragment and the C-terminal fragment of the polypeptide.

[0395] The deaminase may be a circular permutation deaminase. In some embodiments, the deaminase is TadA, which is a circular permutation at amino acid residues 116, 136, or 65 as numbered in the TadA reference sequence.

[0396] The fusion protein or complex may include more than one deaminase. The fusion protein or complex may include, for example, one, two, three, four, five, or more deaminases. In the fusion protein or complex, the deaminase may be an adenosine deaminase, a cytidine deaminase, or a combination thereof.

[0397] In some embodiments, the napDNAbp in the fusion protein or complex contains a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide may be a variant Cas9 polypeptide. The Cas9 polypeptide may be a Cas9 protein permuted into a circular form.

[0398] A heterogeneous polypeptide (e.g., a deaminase) can be inserted into a napDNA bp [e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)] at a suitable site, thereby allowing the napDNA bp to maintain, for example, the ability to bind to a target polynucleotide and a guide nucleic acid. The deaminase [e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase (dual deaminase)] can be inserted into the napDNA bp without impairing the function of the deaminase (e.g., base editing activity) or the napDNA bp (e.g., the ability to bind to a target nucleic acid and a guide nucleic acid).

[0399] In some embodiments, a deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of a Cas9 polypeptide containing a higher-than-average factor B (e.g., a higher-than-average factor B compared to a protein domain containing a total protein or a disordered region). The Cas9 polypeptide location containing a higher-than-average factor B may include residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, for example, as numbered in SEQ ID NO. 197. Cas9 polypeptide regions containing a B factor higher than average may include residues 792–872, 792–906, and 2–791, as numbered in SEQ ID NO. 197.

[0400] In some embodiments, a heterogeneous polypeptide (e.g., a deaminase) is inserted into a flexible loop of a Cas9 polypeptide. The flexible loop portion may be selected from the group consisting of corresponding amino acid residues from SEQ ID NO. 197, numbered 530–537, 569–570, 686–691, 943–947, 1002–1025, 1052–1077, 1232–1247, or 1298–1300 or another Cas9 polypeptide. The flexible loop portion may be selected from the group consisting of corresponding amino acid residues from 1 to 529, 538 to 568, 580 to 685, 692 to 942, 948 to 1001, 1026 to 1051, 1078 to 1231, or 1248 to 1297 or another Cas9 polypeptide, as numbered in SEQ ID NO. 197.

[0401] A heterogeneous polypeptide (e.g., adenine deaminase) may be inserted into a Cas9 polypeptide region corresponding to the corresponding amino acid residues in another Cas9 polypeptide, such as amino acid residues 1017–1069, 1242–1247, 1052–1056, 1060–1077, 1002–1003, 943–947, 530–537, 568–579, 686–691, 1242–1247, 1298–1300, 1066–1077, 1052–1056, or 1060–1077, as numbered in SEQ ID NO. 197.

[0402] Heterogeneous polypeptides (e.g., adenine deaminase) can be inserted in place of a deleted region of the Cas9 polypeptide. The deleted region may correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. An exemplary internal fusion base editor is shown below. Table 4A It is provided in.

[0403]

[0404] A heterogeneous polypeptide (e.g., a deaminase) may be inserted into a structural or functional domain of a Cas9 polypeptide. A heterogeneous polypeptide (e.g., a deaminase) may be inserted between two structural or functional domains of a Cas9 polypeptide. A heterogeneous polypeptide (e.g., a deaminase) may be inserted in place of a structural or functional domain of a Cas9 polypeptide, for example, after deleting a domain from the Cas9 polypeptide. The structural or functional domain of a Cas9 polypeptide may include, for example, RuvC I, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH.

[0405] The fusion protein may include a linker between the deaminase and the napDNAbp polypeptide. The linker may be a peptide or non-peptide linker. For example, the linker is XTEN, (GGGS) n (Sequence No. 246), SGGSSGGS(Sequence No. 330), (GGGGS) n (Sequence No. 247), (G) n , (EAAAK)n(sequence number 248), (GGS) n, may be SGSETPGTSESATPES (SEQ No. 249). In some embodiments, the fusion protein comprises a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N-terminal and C-terminal fragments of the napDNAbp are connected to the deaminase using a linker. In some embodiments, the N-terminal and C-terminal fragments are connected to the deaminase domain without a linker. In some embodiments, the fusion protein comprises a linker between the N-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the N-terminal Cas9 fragment and the deaminase.

[0406] In some embodiments, the napDNAbp within the fusion protein or complex is a Cas12 polypeptide, e.g., Cas12b / C2c1, or a functional fragment thereof capable of associating with a nucleic acid (e.g., gRNA) that guides Cas12 to a specific nucleic acid sequence. The Cas12 polypeptide may be a variant Cas12 polypeptide. In other embodiments, the N or C-terminal fragment of the Cas12 polypeptide comprises a nucleic acid programmable DNA binding domain or a RuvC domain. In other embodiments, the fusion protein contains a linker between the Cas12 polypeptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ No. 250) or GSSGSETPGTSESATPESSG (SEQ No. 251). In other embodiments, the linker is a rigid linker. In another embodiment of the above aspect, the linker is encoded by GGAGGCTCTGGAGGAAGC (SEQ No. 252) or GGCTCTTCTGGATCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ No. 253).

[0407] In another embodiment, the fusion protein or complex contains a nuclear localization signal (e.g., a bipartite nuclear localization signal). In another embodiment, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID No. 261). In another embodiment of the above aspect, the nuclear localization signal is the sequence

[0408] It is encoded by ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ No. 262). In another embodiment, the Cas12b polypeptide contains a mutation that inactivates the catalytic activity of the RuvC domain. In another embodiment, the Cas12b polypeptide contains D574A, D829A, and / or D952A mutations.

[0409] In some embodiments, the fusion protein or complex comprises a napDNAbp domain (e.g., a Cas12-derived domain) having an internally fused nucleobase editing domain [e.g., a deaminase domain, e.g., all or part of an adenosine deaminase domain (functional part)]. In some embodiments, the napDNAbp is Cas12b. In some embodiments, the base editor is as follows Table 4B It includes a BhCas12b domain having an internally fused TadA*8 domain inserted into the locus provided.

[0410]

[0411] In some embodiments, the base editing system described herein is an ABE having TadA inserted into Cas9. Polypeptide sequences of the associated ABE having TadA inserted into Cas9 are provided in the attached sequence list as SEQ ID NOs 263 to 308.

[0412] Exemplary but non-limiting fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Applications No. 62 / 852,228 and No. 62 / 852,224, the entire contents of which are incorporated herein by reference.

[0413] Editing from A to G

[0414] In some embodiments, the base editor described herein comprises an adenosine deaminase domain. This adenosine deaminase domain of the base editor can facilitate editing an A nucleus to a guanine (G) nucleus by deaminating adenine (A) to form inosine (I), which exhibits the base pairing properties of G. In some embodiments, the A-to-G base editor further comprises an inosine base excision repair inhibitor, e.g., a uracil glycosylation enzyme (UGI) domain, or a catalytically inactive inosine-specific nuclease. Without being bound by any particular theory, the UGI domain or the catalytically inactive inosine-specific nuclease can inhibit or prevent the base excision repair of a deaminated adenosine residue (e.g., inosine), which can improve the activity or efficiency of the base editor.

[0415] A base editor comprising an adenosine deaminase may act on any polynucleotide comprising DNA, RNA, and DNA-RNA hybrids. In one embodiment, the adenosine deaminase domain of the base editor comprises all or part (e.g., a functional portion) of ADAT containing one or more mutations that allow ADAT to deaminate target A in DNA. For example, the base editor may comprise all or part (e.g., a functional portion) of ADAT (EcTadA) from Escherichia coli containing one or more of the mutations D108N, A106V, D147Y, E155V, L84F, H123Y, I156F or a corresponding mutation in another adenosine deaminase. Exemplary ADAT homolog polypeptide sequences are provided in the sequence list as SEQ ID NOs 1 and 309–315.

[0416] Adenosine deaminase is used in any suitable organism (e.g., E. coli [ E.ColiIt can be derived from ]). In some embodiments, adenosine deaminase is E. coli [ Escherichia coli ], Staphilococcus aureus( Staphylococcus aureus ), Salmonella typhi( Salmonella typhi ), Shewanella putrepathiens( Shewanella putrefaciens ), Haemophilus influenzae ( Haemophilus influenzae ), Caullobacter crescentus( Caulobacter crescentus ) or Bacillus subtilis( Bacillus subtilis It originates from ). In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase comprising one or more mutations corresponding to any mutation provided herein (e.g., a mutation in ecTadA). Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Thus, a mutation may be generated in any naturally occurring adenosine deaminase (e.g., having homology to ecTadA) corresponding to any mutation described herein (e.g., any mutation identified in ecTadA).

[0417] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences presented in any adenosine deaminase provided herein. It should be understood that the adenosine deaminase provided herein may comprise one or more mutations (e.g., any mutation provided herein). The present disclosure provides any deaminase domain having a specific percentage of identity and any mutation or combination thereof described herein. In some embodiments, the adenosine deaminase is compared with the reference sequence or any adenosine deaminase provided herein, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, It includes an amino acid sequence having 46, 47, 48, 49, 50 or more mutations.

[0418] Any mutation provided herein [based on a TadA reference sequence, such as TadA*7.10 (Sequence No. 1)] is E. coli [ E. coli ] TadA[ecTadA], Es. Aureus TadA( S. aureusIt should be understood that it may be introduced into other adenosine deaminases, such as TadA (saTadA) or other adenosine deaminases (e.g., bacterial adenosine deaminase). In some embodiments, the TadA reference sequence is TadA*7.10 (Sequence No. 1). It will be apparent to those skilled in the art that additional deaminases may be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein. Thus, any mutation identified in the TadA reference sequence may be made in other adenosine deaminases (e.g., ecTada) having homologous amino acid residues. It should also be understood that any mutation provided herein may be made individually or in any combination with the TadA reference sequence or another adenosine deaminase.

[0419] In some embodiments, adenosine deaminase is as follows Table 5A~5G Includes a change or set of changes selected from those listed in.

[0420]

[0421]

[0422]

[0423]

[0424]

[0425] In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant further comprising an amino acid modification of F149Y. In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant further comprising amino acid modifications of R147D, F149Y, T166I, and D167N (TadA*8.10+). In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant further comprising amino acid modifications of S82T and F149Y (TadA*9v1). In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant further comprising amino acid modifications of Y147D, F149Y, T166I, D167N, and S82T (TadA*9v2).

[0426] In some embodiments, the adenosine deaminase is M1I, M1S, S2A, S2E, S2H, S2R, S2L, E3L, V4D, V4E, V4M, V4K, V4S, V4T, V4A, E5K, F6S, F6G, F6H, F6Y, F6I, F6E, S7K, H8E, H8Y, H8H, H8Q, H8E, H8G, H8S, E9Y, E9K, E9V, E9E, Y10F, Y10W, Y10Y, M12S, M12L, M12R, M12W, R13H, R13I, R13Y, R13R, R13G, R13S, in the TadA reference sequence (e.g., TadA*7.10,ecTadA or TadA8e). H14N, A15D, A15V, A15L, A15H, T17T, T17A, T17W, T17L, T17F, T17R, T17S, L18A, L18E, L18N, L18L, L18S, A19N, A19H, A19K, A19A, A19D, A19G, A19M, R21N, K20K, K20A, K20R, K20E, K20G, K20C, K20Q R21A, R21R, R21N, R21Y, R21C G22P, A22W, A22R, W23D, R23H, W23G, W23Q, W23L, W23R, W23H W23D W23M, W23W, W23I, D24E, D24G, D24W, D24D, D24R, E25F, E25M, E25D, E25A, E25G, E25R, E25E, E25H E25V, E25S, E25Y, R26D, R26E, R26G, R26N, R26Q, R26C, R26L, R26K, R26W, R26C, R26P, R26R, R26A, R26H, E27E, E27Q, E27H, E27C, E27G, E27K, E27S, E27P, E27R, E27L, E27V, E27D, V28V, V28A, V28C, V28G, V28P, V28S, V28T, P29V, P29P, P29A, P29G, P29K, P29L, V30V, V30I, V30L, V30F, V30G, V30A, V30M, L34S,L34V, L34L, L34M, L34W, L34G, H36E, H36V, L36H, H36L, H36N, N37N, N37H, N37R, N37T, N37S, N38G, N38R, N38N, N38E, V40I, W45A, W45W, W45R, W45L, W45N, N46N, N46M, N46P, N46G, N46L, N46R, N46V, R46W, R46F, R46Q, R46M, R47A, R47Q, R47F, R47K, R47P, R47W, R47M, R47R, R47G, R47S, R47V, R47H, P48T, P48L, P48A, P48I, P48S, P48R, P48K, P48D, P48E, P48H, P48G, P48P, P48N, I49G, I49H, I49V, I49F, I49H, I49I, I49M, I49N, I49K, I49Q, I49T, G50L, G50S, G50R, G50G, R51H, R51L, R51N, L51W, R51Y, R51G, R51V, R51R, H52D, H52Y, H52I, H52H, D53D, D53E, D53G, D53P, P54C, P54T, P54P, P54E, A55H, T55A, T55I, T55V, T55G, T55T, A56A, A56H, A56W, A56E, A56S, H57P, H57A, H57H, H57N, A58G, A58E, A58A, A58R, E59A, E59G, E59I, E59Q, E59W, E59E, E59T, E59H, E59P, M61A, M61I, M61L, M61V, M61P, M61G, M61I, L63S, L63V, L63T, L63R, L63H, L63A, R64A, R64Q, R64R, R64D, Q65V, Q65H, Q65G, Q65P, Q65F, Q65Q, Q65R, G66V, G66E, G66T, G66G, G66C, G67G, G67W, G67I, G67A, G67D, G67L, G67V, L68Q, L68M, L68V, L68H, L68L, L68G,V69A,V69M, V69V, M70V, M70L, E70A, M70A, M70M, M70E, M70T, M70v, Q71M, Q71N, Q71L, Q71R, Q71Q, Q71I, N72A, N72K, N72S, N72D, N72Y, N72N, N72H, N72G, N72M, Y73G, Y73I, Y73K, Y73R, Y73S, Y73Y, Y73H, Y73A, R74A, R74Q, R74G, R74K, R74L, R74N, R74G, R74K, R74R, I76H, I76R, I76W, I76Y, I76V, I76Q, I76L I76D, I76F, I76I, I76N, I76T, I76Y, D77G, D77D, D77A, D77Q, A78Y, A78T, A78G, A78A, A78I, T79M, T79R, T79L, T79T, L80M, L80Y, L80I, L80V, L80L, Y81D, Y81V, Y81Y, Y81M, V82A, V82S, V82G, V82T, V82V, V82Q, V82Y, T83L, T83F, T83T, T83N, L84E, L84F, L84Y, L84I, L84L, L84M, L84A, L84T, L84S, E85K E85G, E85P, E85S, E85E, E85F, E85V, E85R, P86T, P86C, P86P, P86L, P86N, P86K, P86H, C87M, C87I, C87S, C87N, C87P, S87C, S87L, S87V, V88A, V88M, V88V, V88T, V88E, V88D, V88S, C90S, C90P, C90A, C90T, C90M, A91A, A91G, A91S, A91V, A91T, A91C, A91L, G92T, G92M, G92A, G92Y, G92G, A93I, A93C, A93M, A93V, A93A, M94M, M94T, M94A, M94V, M94L, M94I, M94H, I95S, I95G, I95L, I95H, I95V, H96A, H96L, H96R, H96S, H96HH96N, H96E, S97C, S97G, S97I, S97M, S97R, S97S, S97P, R98K, R98I, R98N, R98Q, R98G, R98H, R98C, R98L, R98R, G100R, G100V, G100K, G100A, G100S, G100M, G100I, R101V, R101R, R101S, R101C, V102A, V102F, V102I, V102V, D103A, V103A, V103G, V103F, V103V, F104G, D104N, F104V, F104I, F104L, F104A, F104F, F104R, G105V, G105W, G105G, G105M, G105A, A106T, V106Q, V106F, V106W, V106M, A106A, A106Q, A106F, A106G, A106W, A106M, A106V, A106R, A106L, A106S, A106B, A106I, R107C, R107G, R107P, R107K, R107A, R107N, R107W, R107H, R107S, R107R, R107F, D108N, D108F, D108G, D108V, D108A, D108Y, D108H, D108I, D108K, D108L, D108M, D108Q, N108Q, N108F, N108W, N108M, N108K, D108K, D108F, D108M, D108Q, D108R, D108W, D108S, D108E, D108T, D108R, D108D, A109H, A109K, A109R, A109S, A109T, A109V, A109A, A109D, K110G, K110H, K110I, K110R, K110T, K110K, K110A, K110l, T111A, T111G, T111H, T111R, T111T, T111K, G112A, G112G, G112H, G112T, G112R, A113N, A114G, A114H, A114V, A114C, A114S, A114A, G115S, G115G, G115M, G115L,G115A, G115F, L117M, L117L, L117W, L117A, L117S, L117N, L117V, M118D, M118G, M118K, M118N, M118V, M118M, M118L, M118R, D119L, D119N, D119S, D119V, D119D, V120H, V120L, V120V, V120T, V120A, V120E, V120G, V120D, L121D, L121M, L121N, L121K, L121L, H122H, H122N, H122P, H122R, H122S, H122Y, H122G, H122T, H122L, H123C, H123G, H123P, H123V, H123Y, Y123H, H123Y, H123H, P124P, P124H, P124A, P124Y, P124D, P124G, P124I, P124L, P124W, G125H, G125I, G125A, G125M, G125K, G125G, G125P, M126D, M126H, M126K, M126I, M126N, M126O, M126S, M126Y, M126M, M126G, N127H, N127S, N127D, N127K, N127R, N127N, N127I, N127P, N127M, H128R, H128N, H128L, H128H, R129H, R129Q, R129V, R129I, R129E, R129V, R129R, R129M, R129P, V130R, V130V, V130E, V130D, E131E, E131I, E131V, E131K, I132I, I132F, I132T, I132L, I132V, I132E, T133V, T133E, T133G, T133K, T133T, T133A, T133H, T133F, T133I, E134A, E134E, E134G, E134I, E134H, E134K, E134T, G135G, G135V, G135I, G135P, G135E, I136G, I136L, I136T, I136I , l137A, l137D, l137E,L137M, l137S, L137L, L137I , A138D, A138E, A138G, S138A, A138N, A138S, A138T, A138V, A138Y, A138A, A138M, A138L, D139E, D139I, D139C, D139L, D139M, D139D, D139G, D139H, D139A, E140A, E140C, E140L, E140R, E140K, E140E, E140D, C141S, C141A, C141C, C141V, C141E, A142N, A142D, A142G, A142A, A142L, A142S, A142T, A142N, A142S, A142V, A142E, A142C, A143D, A143E, A143G, , A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, A143R, A143A, A143I, L144S, L144L, L144T, L144A, L145A, L145F, L145G, L145D, L145L, L145C, L145E, L145s, C146R, S146A, S146C, S146D, S146F, S146R, S146T, S146D, S146G, S146S, S146L, D147D, D147L, D147F, D147G, D147Y, Y147T, Y147R, Y147D, D147R, D147Y, D147A, D147T, D147H, D147F, D147U, D147V, D147I, D147C, F148L, F148F, F148R, F148Y, F148A, F148T, F149C, F149M, F149R, F149Y, F149N, F149F, F149A, F149T, F149V R150R, R150M, R150D, R150F, M151F, M151P, M151R, M151V, M151M, M151E, R152C, R152F, R152H, R152P, R152R, R152P, R152Q, R152M, R152O, R153C, R153Q, R153R, R153V,R153E, R153A, R153P, Q154E, Q154H, Q154M, Q154R, Q154L, Q154S, Q154V, Q154Q, Q154F, Q154I, Q154A, Q154K, E155F, E155G, E155I, E155K, E155P, E155V, E155D, E155E, E155L, E155Q, I156V, I156A, I156I, I156L, I156F, I156D, I156K, I156N, I156R, I156Y, E157A, E157F, E157I, E157P, E157T, E157V, N157K, K157N, K157V, K157P, K157I, K157F, K157F, K157T, K157A, K157S, K157R, A158Q, A158K, A158V, A158A, A158D, A158S, A158T, A158N, Q159S, Q159Q, Q159A, Q159F, Q159K, Q159L, Q159N, K160A, K160S, K160E, K160K, K160N, K160F, K160Q, K161T, K161K, K161R, K161I, K161A, K161N, K161Q, K161S, K161T, A162D, A162Q, R162H, R162P, A162S, A162A, A162N, A162M, A162K, Q163G, Q163S, Q163Q, Q163A, Q163H, Q163N, Q163R, S164F, S164S, S164Q, S164I, S164R, S164Y, S165S, S165P, S165Q, S165A, S165D, S165I, S165T, S165Y, T166T, T166Q, T166E, T166S, T166D, T166K, T166I, T166N, T166P, T166R, D167S D167D, D167I, D167G, D167T,It includes one or more of the D167A and / or D167N mutations and any alternative mutation at the corresponding position or one or more corresponding mutations in another adenosine deaminase. Additional mutations are described in these disclosures, U.S. Patent Application No. 2022 / 0307003 A1, U.S. Patent No. 11,155,803, and International Patent Application No. WO 2023 / 288304 A2, PCT / CN2022 / 143408, WO 2018 / 027078 A1, WO 2021 / 158921 A1, and WO 2023 / 034959 A2, the entirety of which is incorporated herein by reference for all purposes.

[0427] In various embodiments, the adenosine deaminase of the present disclosure lacks an N-terminal methionine.

[0428] In some embodiments, the present disclosure provides a TadA variant comprising a change in one or more amino acids selected from L36, I76, V82, Y147, Q154, and N157 compared to TadA*7.10. In some embodiments, the present disclosure provides a TadA variant comprising one or more of L36H, I76Y, V82T, Y147T, Q154S, and N157K changes to TadA*7.10. In some embodiments, the present disclosure provides a TadA variant comprising L36H, I76Y, V82T, Y147T, Q154S, and N157K changes to TadA*7.10. In some embodiments, the present disclosure provides TadA variants including F84Y, A109L, A109V, A109I, A109F, A109S, A109T, A109N, V155S, V155T, V155N, F156Y, F156W, F156R, F156N, and F156Q modifications to TadA*7.10. In some embodiments, the present disclosure relates to TadA*7.10 with respect to E3N, E3K, E3G, F6A, H14D, L18A, W23I, W23R, P29T, P29Y, P29Q, V35Q, L36S, N38D, G42M, N46Y, P48A, G50A, H52L, A62V, L63R, L63F, Q65R, G67N, L68V, M70I, N72Y, T79H, Y81V, V82S, M94R, G100V, V102E, V102S, R107A, A114C, G115E, M118L, D119L, H122T, P124H, P124K, P124Q, H128R, TadA variants including V130F, I132K, I132T, E140L, A142N, A142S, L144Q, L145R, L145N, Y147A, F149A, R152P, F156N, and K160E modifications are provided.

[0429] In some embodiments, the present disclosure provides a TadA variant comprising V82T, Y147T, and / or Q154S mutations. In some embodiments, the present disclosure provides a TadA variant comprising V82T, Y147T, and / or Q154S mutations. In some embodiments, the present disclosure provides a TadA*8.8 further comprising a V82T mutation. In some embodiments, the present disclosure provides a TadA*8.8 further comprising V82T, Y147T, and Q154S mutations. In some embodiments, the present disclosure provides a TadA*8.17 further comprising a V82T mutation. In some embodiments, the present disclosure provides a TadA*8.17 further comprising V82T, Y147T, and Q154S mutations. In some embodiments, the present disclosure provides a TadA*8.20 further comprising a V82T mutation. In some embodiments, the present disclosure provides TadA*8.20 further comprising V82T, Y147T, and Q154S mutations.

[0430] In the embodiments, variants of TadA*7.10 include one or more modifications selected from any of these modifications provided herein.

[0431] In certain embodiments, the adenosine deaminase heterodimer is Staphylococcus aureus ( Staphylococcus aureus , S. aureus ) TadA, Bacillus subtilis( Bacillus subtilis , B. subtilis ) TadA, Salmonella typhimurium( Salmonella typhimurium , S. typhimurium ) TadA, Shewanella putrepathiens( Shewanella putrefaciens , S. putrefaciens ) TadA, Haemophilus influenzae F3031( H. influenzae ) TadA, Caulobacter crescentus, C. crescentus ) TadA, Geobacter sulfur reducens( Geobacter sulfurreducens , G. sulfurreducensIt includes a TadA*8 domain selected from TadA or TadA*7.10 and an adenosine deaminase domain.

[0432] In some embodiments, TadA*8 is Table 5D It is a variant as shown in [figure]. Table 5D represents specific amino acid position numbers within the TadA amino acid sequence and the amino acids present at these positions within TadA-7.10 adenosine deaminase. Table 5D It also shows amino acid changes in TadA variants relative to TadA-7.10 after phage-assisted non-continuous evolution (PANCE) and phage-assisted continuous evolution (PACE), as described in the literature [M. Richter et al., 2020, Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z], the entire contents of which are incorporated herein by reference. In some embodiments, TadA*8 is TadA*8a, TadA*8b, TadA*8c, TadA*8d, or TadA*8e. In some embodiments, TadA*8 is TadA*8e. In one embodiment, the adenosine deaminase is TadA*8 comprising or essentially consisting of SEQ ID NO. 316 or a fragment thereof having adenosine deaminase activity.

[0433]

[0434] In some embodiments, the TadA variant Table 5E It is a variant as shown in [figure]. Table 5Erepresents a specific amino acid position number within the TadA amino acid sequence and the amino acids present at these positions within the TadA*7.10 adenosine deaminase. In some embodiments, the TadA variant is MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829. In some embodiments, the TadA variant is MSP828. In some embodiments, the TadA variant is MSP829.

[0435]

[0436]

[0437]

[0438]

[0439] In certain embodiments, the fusion protein or complex comprises a single (e.g., provided as a monomer) TadA* (e.g., TadA*8 or TadA*9). Throughout the present disclosure, adenosine deaminase base editors comprising a single TadA* domain are denoted by the terms ABEm or ABE#m, where '#' is an identification number (e.g., ABE8.20m) and 'm' indicates 'monomer'. In some embodiments, TadA* is linked to a Cas9 cleftase. In some embodiments, the fusion protein or complex of the present disclosure comprises a heteromer of wild-type TadA[TadA(wt)] linked to TadA*. Throughout this disclosure, an adenosine deaminase base editor and TadA(wt) comprising a single TadA* domain are denoted by the terms ABEd or ABE#d, where '#' is an identification number (e.g., ABE8.20d) and 'd' indicates 'dimer'. In other embodiments, the fusion protein or complex of this disclosure comprises a heteromer of TadA*7.10 linked to TadA*. In some embodiments, the base editor is ABE8 comprising a TadA* variant monomer. In some embodiments, the base editor is ABE comprising a heteromer of TadA* and TadA(wt). In some embodiments, the base editor is ABE comprising a heteromer of TadA* and TadA*7.10. In some embodiments, the base editor is ABE comprising a heteromer of TadA*. In some embodiments, TadA* is Tables 5A~5E It is selected from.

[0440] In some embodiments, the adenosine deaminase is expressed as a monomer. In other embodiments, the adenosine deaminase is expressed as a heteromer. In some embodiments, the deaminase or other polypeptide sequence lacks methionine, for example, when included as a component of a fusion protein. This may change the position numbering. However, a person skilled in the art will understand that these corresponding mutations refer to the same mutation.

[0441] Any mutation provided herein and any additional mutation (e.g., based on the ecTadA amino acid sequence) may be introduced into any other adenosine deaminase. Any mutation provided herein may be made individually or in any combination with the TadA reference sequence or another adenosine deaminase (e.g., ecTadA).

[0442] Detailed information regarding A to G nucleobase editing proteins is provided in International PCT Application No. PCT / US2017 / 045381 (No. WO2018 / 027078), the full contents of which are incorporated herein by reference, and the literature [Gaudelli, NM, et al., "Programmable base editing of A T to G It is described in [Nature, 551, 464-471 (2017)].

[0443] Editing from C to T

[0444] In some embodiments, the base editor disclosed herein comprises a fusion protein or complex comprising a cytidine deaminase capable of deaminating a target cytidine (C) base of a polynucleotide to produce uridine (U) having the base pair binding properties of thymine. In some embodiments, for example, if the polynucleotide is double-stranded (e.g., DNA), the uridine base may then be substituted with a thymidine base (e.g., by a cell repair mechanism) to cause a transition from C:G to T:A. In other embodiments, deamination from C to U in nucleic acid by the base editor cannot be accompanied by a substitution from U to T.

[0445] Deamination of target C in a polynucleotide to produce U is a non-limiting example of the type of base editing that can be performed by the base editor described herein. In another example, a base editor comprising a cytidine deaminase domain can mediate the conversion of the cytosine (C) base to the guanine (G) base. For example, U in the polynucleotide produced by the deamination of cytidine by the cytidine deaminase domain of the base editor can be excised from the polynucleotide by a base excision repair mechanism [e.g., a uracil DNA glycosylase (UDG) domain] to produce a base-free site. Then, the nucleus base opposite the base-free site can be substituted with another base, such as C, by, for example, subsequently by a translesion polymerase (e.g., by a base pairing mechanism). While it is common for the nucleus opposite the base-free site to be replaced by C, other substitutions (e.g., A, G, or T) can also occur.

[0446] Accordingly, in some embodiments, the base editor described herein comprises a deamination domain (e.g., a cytidine deamination domain) capable of deaminating a target C to U in a polynucleotide. Additionally, as described below, in some embodiments, the base editor may comprise an additional domain that facilitates the conversion of the U generated from deamination to T or G. For example, a base editor comprising a cytidine deaminase domain may further comprise a uracil glycosylation enzyme inhibitor (UGI) domain that mediates the substitution of U by T, thereby completing the base editing work from C to T. In another example, the base editor may comprise a uracil stabilizing protein as described herein. In another example, a base editor can improve the efficiency of base editing from C to G by incorporating damage-passing polymerase, because damage-passing polymerase can facilitate the incorporation of C on the opposite side of the base-free site (i.e., causing the incorporation of G at the base-free site to complete the base editing work from C to G).

[0447] A base editor containing cytidine deaminase as a domain can deaminate target C in any polynucleotide including DNA, RNA, and DNA-RNA hybrids.

[0448] In some embodiments, the cytidine deaminase of the base editor comprises all or part (e.g., a functional part) of the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of the APOBEC-like protein is a catalytic domain, while the C-terminal domain is a pseudo-catalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain and is important for cytidine deamination. Members of the APOBEC family include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D ('APOBEC3E' is now referred to as such), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminases.

[0449] Other exemplary deaminases that can be fused to Cas9 according to aspects of the present disclosure are provided below. In the embodiments, the deaminase is an activation-induced deaminase (AID). It should be understood that in some embodiments, the active domain of each sequence, for example, a domain without a localization signal (nuclear localization sequence, nuclear localization sequence without a nuclear release signal, cytoplasmic localization signal), may be used.

[0450] Some aspects of the present disclosure are based on the recognition that controlling the catalytic activity of the deaminase domain of any fusion protein or complex described herein, for example by generating a point mutation in the deaminase domain, affects the progression of the fusion protein (e.g., base editor) or complex. For example, a mutation that reduces but does not eliminate the catalytic activity of the deaminase domain within a base-editing fusion protein or complex will reduce the likelihood that the deaminase domain will catalyze the deaminization of residues adjacent to a target residue, thereby narrowing the deaminization window. The ability to narrow the deaminization window can prevent unwanted deaminization of residues adjacent to specific target residues, which can reduce or prevent off-target effects.

[0451] In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise one or more mutations selected from the group consisting of R33A, K34A, E63A, H102P, D104N, H121R, H122R, H122L, D124N, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1, and D316R, D317R, R320A, R320E, R313A, W285A, W285Y, and R326E of hAPOBEC3G, and any alternative mutation at a corresponding position in another APOBEC deaminase or one or more corresponding mutations. In some embodiments, the APOBEC deaminase incorporated into the base editor is K34A, H122L, and D124N (AALN); It may include one or more combinations of mutations selected from H102P and D104N (evoFERNY derived from FERNY); W90Y and R126E (YE1); W90Y and R132E (YE2); R126E and R132E (EE); W90Y, R126E and R132E (YEE), or rAPOBEC1, and any alternative mutation at a corresponding position in another APOBEC deaminase or one or more corresponding mutations.

[0452] A number of modified cytidine deaminases, including but not limited to SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, are commercially available and are available at Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177). In some embodiments, the deaminase incorporated into the base editor comprises all or part (e.g., a functional portion) of the APOBEC1 deaminase.

[0453] In some embodiments, the fusion protein or complex of the present disclosure comprises one or more cytidine deaminase domains. In some embodiments, the cytidine deaminase provided herein may deaminate cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytidine deaminase provided herein may deaminate cytosine from DNA. The cytidine deaminase may be derived from any suitable organism. In some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase comprising one or more mutations corresponding to any mutation provided herein. A person skilled in the art will be able to identify corresponding residues in any homologous protein, for example, by sequence alignment and determination of homologous residues. Thus, a person skilled in the art will be able to generate mutations in any naturally occurring cytidine deaminase corresponding to any mutation described herein. In some embodiments, the cytidine deaminase is derived from a prokaryote. In some embodiments, the cytidine deaminase is derived from bacteria. In some embodiments, cytidine deaminase is derived from mammals (e.g., humans).

[0454] In some embodiments, the cytidine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the cytidine deaminase amino acid sequences presented herein. It should be understood that the cytidine deaminase provided herein may comprise one or more mutations (e.g., any mutation provided herein). Some embodiments provide a polynucleotide molecule encoding a cytidine deaminase nucleobase editor polypeptide of any prior embodiment or a polypeptide as described herein. In some embodiments, the polynucleotide is codon optimized.

[0455] In an embodiment, the fusion protein of the present disclosure comprises two or more nucleic acid editing domains.

[0456] Detailed information on C-to-T nucleobase editing proteins is provided in International PCT Application No. PCT / US2016 / 058344 (No. WO2017 / 070632), the full contents of which are incorporated herein by reference, and in the literature [Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016)]. Additional non-limiting examples of C-to-T nucleobase editing proteins are provided in PCT Applications No. PCT / US2020 / 062428 and No. PCT / US2019 / 033848, the full contents of which are incorporated herein by reference.

[0457] Cytidine Adenosine Base Editor (CABE)

[0458] In some embodiments, the base editors described herein include adenosine deaminase variants with increased cytidine deaminase activity. Such base editors may be referred to as 'cytidine adenosine base editors (CABE)' or 'cytosine base editors derived from TadA* (CBE-T),' and the corresponding deaminase domain is 'TadA* (T ADIt may be referred to as a C)' domain or TadA-derived cytidine deaminases (TadA-CD). A base editor containing an adenosine deaminase variant having both cytidine deaminase and adenosine deaminase activities (i.e., TadA dual deaminase) may be referred to as a TadA-based dual editor [TadDE]. In some cases, the adenosine deaminase variant has both adenine and cytosine deaminase activities (i.e., it is a dual deaminase). In some embodiments, the adenosine deaminase variant deaminates adenine and cytosine of DNA. In some embodiments, the adenosine deaminase variant deaminates adenine and cytosine of single-stranded DNA. In some embodiments, the adenosine deaminase variant deaminates adenine and cytosine in RNA. In some embodiments, the adenosine deaminase variant primarily deaminates cytosine in DNA and / or RNA (e.g., more than 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of all deaminations catalyzed by the adenosine deaminase variant, or more than about or at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 75, 100, 500, or 1,000 times the number of adenine deaminations catalyzed by the variant). In some embodiments, the adenosine deaminase variant has approximately the same cytosine and adenosine deaminase activities (e.g., the two activities are within about 10% or 20% of each other). In some embodiments, the adenosine deaminase variant has mainly cytosine deaminase activity and little to no adenosine deaminase activity.In some embodiments, the adenosine deaminase variant has cytosine deaminase activity and no significant or detectable adenosine deaminase activity. In some embodiments, the target polynucleotide is present in vitro or in vivo in cells. In some embodiments, the cells are bacterial, yeast, fungal, insect, plant, or mammalian cells. Examples of adenosine deaminase variants having increased cytosine deaminase activity include those described in International Patent Applications No. WO 2024 / 040083 and WO 2022 / 204574, the entire disclosures of which are incorporated herein by reference for all purposes.

[0459] In some embodiments, CABE comprises a bacterial TadA deaminase variant (e.g., ecTadA). In some embodiments, CABE comprises a truncated TadA deaminase variant. In some embodiments, CABE comprises a fragment of a TadA deaminase variant. In some embodiments, CABE comprises a TadA*8.20 variant.

[0460] In some embodiments, the adenosine deaminase variant of the present disclosure is a TadA adenosine deaminase that includes one or more modifications that increase cytosine deaminase activity (e.g., at least about 10, 20, 30, 40, 50, 60, 70 or more), while maintaining adenosine deaminase activity [e.g., at least about 30%, 40%, 50% or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)]. In some cases, the adenosine deaminase variant comprises one or more modifications that increase cytosine deaminase activity relative to reference adenosine deaminase activity (e.g., at least about 10, 20, 30, 40, 50, 60, 70, or more), and comprises undetectable adenosine deaminase activity or adenosine deaminase activity of less than 30%, 20%, 10%, or 5% of the reference adenosine deaminase activity. In some embodiments, the reference adenosine deaminase is TadA*8.20 or TadA*8.19.

[0461] In some embodiments, the adenosine deaminase variant has two or more changes at an amino acid position selected from the group consisting of 2, 4, 6, 8, 13, 17, 23, 27, 29, 30, 47, 48, 49, 67, 76, 77, 82, 84, 96, 100, 107, 112, 114, 115, 118, 119, 122, 127, 142, 143, 147, 149, 158, 159, 162, 165, 166, and 167 of an amino acid sequence having at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or more identity with SEQ ID NO. 1. Or it is an adenosine deaminase containing a corresponding modification from another deaminase.

[0462] In some embodiments, the adenosine deaminase variant is an amino acid sequence S2H, V4K, V4S, V4T, V4Y, F6G, F6H, F6Y, H8Q, R13G, T17A, T17W, R23Q, E27C, E27G, E27H, E27K, E27Q, E27S, E27G, P29A, P29G, P29K, V30F, V30I, R47G, R47S, A48G, I49K, I49M, I49N, I49Q, I49T, G67W, I76H, I76R, I76W, having at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or more identity with SEQ ID NO. 1. It is an adenosine deaminase comprising one or more modifications selected from the group consisting of Y76H, Y76R, Y76W, F84A, F84M, H96N, G100A, G100K, T111H, G112H, A114C, G115M, M118L, H122G, H122R, H122T, N127I, N127K, N127P, A142E, R147H, A158V, Q159S, A162C, A162N, A162Q, and S165P, or a corresponding modification in another deaminase.

[0463] In some embodiments, the adenosine deaminase variant is any Tables 6A–6F It is an adenosine deaminase comprising an amino acid modification or a combination of amino acid modifications selected from those listed.

[0464] The residue identity of an exemplary adenosine deaminase variant capable of deaminating adenine and / or cytidine in target polynucleotides (e.g., DNA) is as follows: Tables 6A–6F Provided in. Additional examples of adenosine deaminase variants include variants of 1.17+E27H, 1.17+E27K, 1.17+E27S, 1.17+E27S+I49K, 1.17+E27G, 1.17+I49N, 1.17+E27G+I49N and 1.17+E27Q ( Table 6A(See [reference]). In some embodiments, any amino acid modification provided herein is substituted with a conservative amino acid. Additional mutations known in the art may be added to any adenosine deaminase variant provided herein.

[0465] In some embodiments, a base editor system comprising the CABE provided herein has at least about 30%, 40%, 50%, 60%, 70% or more of C-to-T editing activity in a target polynucleotide (e.g., DNA). In some embodiments, a base editor system comprising the CABE provided herein has increased C-to-T base editing activity compared to a reference base editor system comprising a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19) (e.g., at least about 30-fold, 40-fold, 50-fold, 60-fold, 70-fold or more of increase).

[0466]

[0467]

[0468]

[0469]

[0470]

[0471]

[0472]

[0473]

[0474]

[0475]

[0476]

[0477]

[0478]

[0479] A TadA-derived cytidine deaminase according to a specific embodiment (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO. 3594, wherein residue 27 of SEQ ID NO. 3594 is any amino acid excluding E (glutamic acid). TadA-CD having other sequence homology is also possible. For example, in a specific embodiment, a TadA-derived cytidine deaminase (e.g., TadA-CD) comprises an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO. 3594, wherein residue 28 of SEQ ID NO. 3594 is any amino acid excluding V (valine). In another exemplary embodiment, a TadA-derived cytidine deaminase is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical to the amino acid sequence of SEQ ID NO. 3594, wherein residue 96 of SEQ ID NO. 3594 is any amino acid excluding H (histidine). In another exemplary embodiment, the cytidine deaminase derived from TadA is at least 80% identical to the amino acid sequence of SEQ ID NO. 3594, at least 85% identical, at least 90% identical, at least 95% identical, at least 98% identical, at least 99% identical, and at least 99.5% identical, wherein residue 26 of SEQ ID NO. 3594 is any amino acid excluding R (arginine).In various embodiments, the TadA-derived cytidine deaminase includes a modification at one or more of positions 26, 27, 28, 48, 73, or 96 compared to SEQ ID NO. 3594.

[0480] As understood by a person skilled in the art, TadA-derived cytidine deaminases (e.g., TadA-CD) may contain a number of mutations compared to parent adenosine deaminases (e.g., TadA-8e). In some embodiments, the deaminase of the present application (e.g., TadA-CD) contains mutations at residues E27, V28, and H96. In some embodiments, the disclosed deaminase further comprises at least one mutation at residues selected from R26, M61, Y73, I76, M151, Q154, and A158 in the amino acid sequence of SEQ ID NO. 3594, or a corresponding mutation in a homologous adenosine deaminase.

[0481] In some embodiments, the deaminase comprises at least one mutation selected from E27A, E27K, V28G, V28A, and H96N in the amino acid sequence of SEQ ID NO. 3594 and additionally at least one mutation in residues selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S or a corresponding mutation in a homologous adenosine deaminase. Other mutations are also possible. For example, in a specific embodiment, the TadA-CD enzyme comprises a mutation selected from E27A, V28G, and H96N in the amino acid sequence of SEQ ID NO. 3594 and additionally at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S or a corresponding mutation in a homologous adenosine deaminase.

[0482] Another exemplary embodiment is a deaminase comprising (1) E27K, V28G, and H96N mutations in the amino acid sequence of SEQ ID NO. 3594 and additionally at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S or a corresponding mutation in a homologous adenosine deaminase; (2) a deaminase comprising E27A, V28A, and H96N mutations in the amino acid sequence of SEQ ID NO. 3594 and additionally at least one mutation selected from R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S or a corresponding mutation in a homologous adenosine deaminase; (3) It may include a deaminase comprising at least one mutation selected from the E27K, V28A, and H96N mutations in the amino acid sequence of SEQ ID NO. 3594 and additionally R26G, M61I, Y73H, Y73S, Y73C, I76F, M151I, Q154R, Q154H, and A158S or a corresponding mutation in a homologous adenosine deaminase.

[0483] In some embodiments, TadA-derived cytidine deaminase (TadA-CD) comprises at least two mutations in residues selected from R26, M61, Y73, I76, M151, Q154, and A158 (compared to reference adenosine deaminase). In other embodiments, TadA-CD comprises at least two mutations in residues selected from R26G, M61I, Y73H, I76F, M151I, Q154H, Q154R, and A158S.

[0484] In some embodiments, the addition of the V106W mutant improves selectivity by inhibiting A deamination to a greater extent than C deamination.

[0485] In some embodiments, the TadA-based dual editor comprises an adenosine deaminase variant comprising one, two, three, four, or five mutations selected from R26G, V28A, A48R, Y73S, and H96N (e.g., SEQ ID NO. 3600).

[0486] As such, in some embodiments, a deaminase is provided herein comprising mutations at residues R26, V28, A48, and Y73 in the amino acid sequence of SEQ ID NO. 3594 or corresponding mutations in a homologous adenosine deaminase. A deaminase is further provided herein comprising mutations at residues R26, E27, V28, A48, and Y73 in the amino acid sequence of SEQ ID NO. 3594 (e.g., further comprising a mutation at E27). In certain embodiments, these deaminases comprise R26G, V28A, A48R, Y73S, and H96N mutations. In some embodiments, these deaminases comprise R26G, V28G, A48R, and Y73C mutations.

[0487] TadA-CD variants are R26G, E27A, V28G, I76F, H96N, and M151I (e.g., TadA-CDa, SEQ ID NO. 3595); R26G, E27A, V28G, I76F, H96N, and A158S (e.g., TadA-CDb, SEQ ID NO. 3596); R26G, E27A, V28G, I76F, H96N, Q154R, and A158S (e.g., TadA-CDc, SEQ ID NO. 3597); E27A, V28G, Y73H, H96N, Q154H, and A158S (e.g., TadA-CDd, SEQ ID NO. 3598); It may include at least one mutation selected from R26G, V28A, A48R, Y73S, and H96N (e.g., TadA-CDe, SEQ ID NO. 3599); V28A, A48R, and Y73S (e.g., TadA-CDf, SEQ ID NO. 3600) and R26G, V28G, A48R, and Y73C (TadA-CDg, SEQ ID NO. 3601).

[0488] In some preferred embodiments, the deaminase is R26G, E27A, V28G, I76F, H96N and A158S (e.g., TadA-CDa, SEQ NO. 3595), R26G, E27A, V28G, I76F, H96N, Q154R and A158S (e.g., TadA-CDb, SEQ NO. 3596), R26G, E27A, V28G, I76F, H96N and M151I (e.g., TadA-CDc, SEQ NO. 3597), E27K, V28A, M61I and H96N (e.g., TadA-CDd, SEQ NO. 3598), E27A, V28G, Y73H, H96N, Q154H and A158S (e.g., TadA-CDe, SEQ NO. Includes mutations of 3599), R26G, V28A, A48R, Y73S, and H96N (e.g., TadA-CDf, SEQ ID NO. 3600) and R26G, V28G, A48R, and Y73C (e.g., TadA-CDg, SEQ ID NO. 3601).

[0489] In some embodiments, the TadA-CD variants described above and herein may also include the V106W mutation.

[0490] In some embodiments, the TadA-CD variant comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99.5% of any amino acid sequence of SEQ ID NOs 3594 to 3601.

[0491] In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46I, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-1, SEQ ID NO. 3602). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46T, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-2, SEQ ID NO. 3603). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46T, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-3, SEQ ID NO. 3604). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-4, SEQ ID NO. 3605). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-5, SEQ ID NO. 3606). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-6, SEQ ID NO. 3607). In some embodiments, the evolved TadA double deaminase comprises mutations of V28A, N46L, A48P, and Y73P (TadA-CD-7, SEQ ID NO. 3608) compared to the amino acid sequence of SEQ ID NO. 3594. In some embodiments, the evolved TadA double deaminase comprises mutations of V28A, N46C, A48P, and Y73P (TadA-CD-8, SEQ ID NO. 3609) compared to the amino acid sequence of SEQ ID NO. 3594.In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-9, SEQ ID NO. 3610). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-10, SEQ ID NO. 3611). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-11, SEQ ID NO. 3612). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-12, SEQ ID NO. 3613). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, Y73P, H96N, and A162V compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-13, SEQ ID NO. 3614).

[0492] In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46I, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-14, SEQ ID NO. 3615). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, A48R, Q71S, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-15, SEQ ID NO. 3616). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, and Y73P compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-16, SEQ ID NO. 3617). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-17, SEQ ID NO. 3618). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-18, SEQ ID NO. 3619). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-19, SEQ ID NO. 3620). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-20, SEQ ID NO. 3621). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G and N46L compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-21, SEQ ID NO. 3622).In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46I, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-22, SEQ ID NO. 3623). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-23, SEQ ID NO. 3624). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, A48P, Y73H, T79P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-24, SEQ ID NO. 3625). In some embodiments, the evolved TadA double deaminase includes mutations of R26G, N46I, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-25, SEQ ID NO. 3626).

[0493] In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-26, SEQ ID NO. 3627). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-27, SEQ ID NO. 3628). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, H96N, and A162V compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-28, SEQ ID NO. 3629). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Q71H, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-29, SEQ ID NO. 3630). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-30, SEQ ID NO. 3631). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, Y73P, H96N, and A162V compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-31, SEQ ID NO. 3632).

[0494] In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-32, SEQ ID NO. 3633). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48R, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-33, SEQ ID NO. 3634). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46V, A48P, Y73S, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-34, SEQ ID NO. 3635). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46C, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-35, SEQ ID NO. 3636). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, L34M, N46L, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-36, SEQ ID NO. 3637). In some embodiments, the evolved TadA double deaminase comprises mutations of R26G, V28A, N46L, A48R, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-37, SEQ ID NO. 3638). In some embodiments, the evolved TadA double deaminase includes mutations of R26G, V28A, N46L, A48P, R64K, Y73P, and H96N compared to the amino acid sequence of SEQ ID NO. 3594 (TadA-CD-38, SEQ ID NO. 3639).

[0495] In some embodiments, the evolved TadA double deaminase comprises mutations of N46I, S73P, and H154Q (TadA-CD-1, SEQ-NO. 3602) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises a mutation of N46T (TadA-CD-2, SEQ-NO. 3603) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46T and H154Q (TadA-CD-3, SEQ-NO. 3604) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V and H154Q (TadA-CD-4, SEQ-NO. 3605) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, S73P, G105S, and H154Q (TadA-CD-5, SEQ-NO. 3606) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L, S73P, and H154Q (TadA-CD-6, SEQ-NO. 3607) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of G26R N46L, R48P, S73P, N96H, and H154Q (TadA-CD-7, SEQ-NO. 3608) compared to the amino acid sequence of SEQ-NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46C, N96H, and H154Q (TadA-CD-8, SEQ No. 3609) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, S73P, and H154Q (TadA-CD-9, SEQ No. 3610) compared to the amino acid sequence of SEQ No. 3600.In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, Q71H, S73P, and H154Q (TadA-CD-10, SEQ No. 3611) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L and H154Q (TadA-CD-11, SEQ No. 3612) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46C, S73P, and H154Q (TadA-CD-12, SEQ No. 3613) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase includes mutations of N46C, S73P, H154Q, and A162V compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-13, SEQ ID NO. 3614).

[0496] In some embodiments, the evolved TadA double deaminase comprises mutations of N46I and H154Q (TadA-CD-14, SEQ No. 3615) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of Q71S and H154Q (TadA-CD-15, SEQ No. 3616) compared to the amino acid sequence of SEQ No. 3594. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L, S73P, N79T, and N96H (TadA-CD-16, SEQ No. 3617) compared to the amino acid sequence of SEQ No. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L, S73P, and N79T (TadA-CD-17, SEQ ID NO. 3618) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of R48A, S73P, and N79T (TadA-CD-18, SEQ ID NO. 3619) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V and N79T (TadA-CD-19, SEQ ID NO. 3620) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, S73P, and N79T compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-20, SEQ ID NO. 3621). In some embodiments, the evolved TadA double deaminase comprises mutations of A28V, N46L, R48A, S73Y, N79T, and N96H compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-21, SEQ ID NO. 3622). In some embodiments, the evolved TadA double deaminase comprises mutations of N46I, S73P, and N79T compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-22, SEQ ID NO. 3623).In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, S73P, N79T, and G106S compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-23, SEQ ID NO. 3624). In some embodiments, the evolved TadA double deaminase comprises mutations of R48P, S73H, and N79P compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-24, SEQ ID NO. 3625). In some embodiments, the evolved TadA double deaminase comprises mutations of A28V, N46I, R48A, S73Y, and N79T compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-25, SEQ ID NO. 3626).

[0497] In some embodiments, the evolved TadA double deaminase comprises mutations of N46V and S73P (TadA-CD-26, SEQ ID NO. 3627) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L (TadA-CD-27, SEQ ID NO. 3628) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46C, S73Y, and A162V (TadA-CD-28, SEQ ID NO. 3629) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V, Q71H, and S73P (TadA-CD-29, SEQ ID NO. 3630) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46C and S73P (TadA-CD-30, SEQ ID NO. 3631) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46C, S73P, and A162V (TadA-CD-31, SEQ ID NO. 3632) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46V and S73P (TadA-CD-32, SEQ ID NO. 3633) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises a mutation of N46V (TadA-CD-33, SEQ ID NO. 3634) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase includes mutations of N46V and R48P (TadA-CD-34, SEQ ID NO. 3635) compared to the amino acid sequence of SEQ ID NO. 3600.In some embodiments, the evolved TadA double deaminase comprises mutations of N46CV and S73P (TadA-CD-35, SEQ ID NO. 3636) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of L34M, N46L, and S73P (TadA-CD-36, SEQ ID NO. 3637) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase comprises mutations of N46L and S73P (TadA-CD-37, SEQ ID NO. 3638) compared to the amino acid sequence of SEQ ID NO. 3600. In some embodiments, the evolved TadA double deaminase includes mutations of N46L, r48P, R64K, and S73P compared to the amino acid sequence of SEQ ID NO. 3600 (TadA-CD-38, SEQ ID NO. 3639).

[0498] In some embodiments, TadA-CD evolved from TadA double is at least 80%, 85%, 90%, 95%, 98%, 99%, or 99.5% identical to any of the amino acid sequences of SEQ ID NOs 39, 41–54, and 359–383.

[0499] Exemplary amino acid sequences of TadA-derived cytosine base editors are TadA-CDa base editor (SpCas9n napDNAbp domain) (TadCBEa) (Sequence No. 3640), TadA-CDb base editor (SpCas9n napDNAbp domain) (TadCBEb) (Sequence No. 3641), TadA-CDc base editor (SpCas9n napDNAbp domain) (TadCBEc) (Sequence No. 3642), TadA-CDd base editor (SpCas9n napDNAbp domain) (TadCBEd) (Sequence No. 3643), TadA-CDe base editor (SpCas9n napDNAbp domain) (TadCBEe) (Sequence No. 3644), and TadA-CDa(V106W) base editor (SpCas9n napDNAbp domain)[TadCBEa(V106W)](Sequence No. Includes 3645), TadA-CDd(V106W) base editor (SpCas9n napDNAbp domain)[TadCBEd(V106W)](Sequence No. 3646), TadA-CDf base editor (SpCas9n napDNAbp domain)(TadCBEf)(Sequence No. 3647), TadA-CDg base editor (SpCas9n napDNAbp domain)(TadCBEg)(Sequence No. 3648), TadA-CDa:eNme2Cas9 base editor (Sequence No. 3649), TadA-CDa:SaCas9 base editor (Sequence No. 3650), TadA-CDa:SpCas9-NG base editor (Sequence No. 3651), and TadA-CDa:enCjCas9 base editor (Sequence No. 3652).

[0500] Exemplary polynucleotides encoding the TadA-derived cytosine base editor of the present disclosure include the TadCBEa-eNme2-C-BE4max vector (SEQ No. 3653), TadCBEa-enCjCas9-BE4max vector (SEQ No. 3654), TadCBEa-SpCas9-BE4max vector (SEQ No. 3655), TadCBEa-SaCas9-BE4max vector (SEQ No. 3656), and TadCBEa-SpCas9-NG-BE4max vector (SEQ No. 3657).

[0501] Guide Polynucleotide

[0502] The polynucleotide programmable nucleotide binding domain can specifically bind to a target polynucleotide sequence when bound to a bound guide polynucleotide (e.g., gRNA) (i.e., through the formation of complementary base pairs between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence), thereby allowing a base editor to be localized to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0503] In one embodiment, the guide polynucleotide described herein may be RNA or DNA. In one embodiment, the guide polynucleotide is gRNA.

[0504] In some embodiments, the guide polynucleotide is at least one single guide RNA ('sgRNA' or 'gRNA'). In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides, which may interact with each other, for example, through complementary base pairing (e.g., double guide polynucleotide, double gRNA). For example, the guide polynucleotide may comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA), or one or more trans-activating CRISPR RNAs (tracrRNA).

[0505] The guide polynucleotide may include natural or non-natural (or unnatural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some cases, the targeting region (e.g., spacer) of the guide nucleic acid sequence may be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotide lengths.

[0506] In some embodiments, the method described herein may utilize a genetically modified Cas protein. The guide RNA [gRNA] is a short synthetic RNA consisting of a user-defined approximately 20 nucleotide spacer that defines a scaffold sequence required for Cas binding and a genomic target to be modified. Exemplary gRNA scaffold sequences are provided in the Sequence List at SEQ ID NOs 317–327 and 425. Thus, a person skilled in the art can modify the genomic target of the Cas protein, the specificity of which is partly determined by how specific the gRNA targeting sequence is to the genomic target relative to the rest of the genome. In the embodiments, the spacer is approximately 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides in length. The spacers of gRNA can be 19, 20, or 21 nucleotides long, or approximately that length.

[0507] The gRNA or guide polynucleotide may target any exon or intron of a gene target. In some embodiments, the composition comprises multiple gRNAs targeting the same exon or multiple gRNAs targeting different exons. Exons and / or introns of a gene may be targeted. gRNA or guide polynucleotide can target a nucleic acid sequence at any point between about 20 nucleotides or fewer than about 20 nucleotides (e.g., at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 nucleotides) or about 1 to 100 nucleotides (e.g., 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 60, 70, 80, 90, 100). The target nucleic acid sequence may be 20 bases right next to the 5' of the first nucleotide of PAM or about that number of bases. gRNA may target the nucleic acid sequence. The target nucleic acid may be at least or at least about 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 70, 1 to 80, 1 to 90, or 1 to 100 nucleotides.

[0508] Guide polynucleotides may include standard ribonucleotides, modified ribonucleotides (e.g., pseudouridines), ribonucleotide isomers and / or ribonucleotide analogs.

[0509] In some embodiments, the base editor system may include multiple guide polynucleotides, e.g., gRNAs. For example, the gRNAs may target one or more target loci included in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences may be arranged side by side and separated by direct repeats.

[0510] Modified polynucleotide

[0511] To improve expression, stability, and / or genome / base editing efficiency and / or reduce possible toxicity, base editor coding sequences (e.g., mRNA) and / or guide polynucleotides (e.g., gRNA) may comprise one or more modified nucleotides and / or chemical modifications, e.g., pseudo-uridine, 5-methylcytosine, 2'- O -methyl-3'-phosphonoacetate, 2'- O -methylthiophase (MSP), 2'- O -methyl-PACE(MP), 2'-fluoroRNA(2'-F-RNA), =restricted ethyl[S-cEt], 2'-O-methyl('M'), 2'- O -methyl-3'-phosphorothioate('MS'), 2'- OIt can be modified using methyl-3'-thiophosphonoacetate ('MSP'), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine. Chemically protected gRNA can improve stability and editing efficiency in vivo and in vitro. Methods using chemically modified mRNA and guide RNA are known in the art, for example, the literature [Jiang et al., Chemical modifications of adenine base editor mRNA and guide RNA expand its application scope. Nat Commun 11, 1979 (2020). doi.org / 10.1038 / s41467-020-15892-8], each incorporated herein by reference in its entirety, and the literature [Callum et al., N It is described in the literature [Andries et al., Journal of Controlled Release, Volume 217, 10 November 2015, pp. 337-344] and [1-Methylpseudouridine substitution enhances the performance of synthetic mRNA switches in cells, Nucleic Acids Research, Volume 48, Issue 6, 06 April 2020, p. 35].

[0512] In some embodiments, the guide polynucleotide comprises one or more modified nucleotides at the 5' end and / or 3' end of the guide. In some embodiments, the guide polynucleotide comprises two, three, four, or more modified nucleosides at the 5' end and / or 3' end of the guide. In some embodiments, the guide polynucleotide comprises two, three, four, or more modified nucleosides at the 5' end and / or 3' end of the guide.

[0513] In some embodiments, the guide comprises at least about 50% to 75% of modified nucleotides. In some embodiments, the guide comprises at least about 85% or more of modified nucleotides. In some embodiments, at least about 1 to 5 nucleotides are modified at the 5' end of the gRNA and at least about 1 to 5 nucleotides are modified at the 3' end of the gRNA. In some embodiments, at least about 3 to 5 consecutive nucleotides are modified at each of the 5' and 3' ends of the gRNA. In some embodiments, at least about 20% of the nucleotides present in the direct repeat or anti-direct repeat are modified. In some embodiments, at least about 50% of the nucleotides present in the direct repeat or anti-direct repeat are modified. In some embodiments, at least about 50% to 75% of the nucleotides present in the direct repeat or anti-direct repeat are modified. In some embodiments, at least about 100% of the nucleotides present in the direct repeat or semi-direct repeat are modified. In some embodiments, at least about 20% or more of the nucleotides present in the hairpins present in the gRNA scaffold are modified. In some embodiments, at least about 50% or more of the nucleotides present in the hairpins present in the gRNA scaffold are modified. In some embodiments, the guide comprises a variable-length spacer. In some embodiments, the guide comprises a spacer of 20 to 40 nucleotides. In some embodiments, the guide comprises a spacer containing at least about 20 to 25 nucleotides or at least about 30 to 35 nucleotides. In some embodiments, the spacer comprises the modified nucleotides. In some embodiments, the guide comprises two or more of the following.

[0514] At least about 1 to 5 nucleotides are modified at the 5' end of the gRNA and at least about 1 to 5 nucleotides are modified at the 3' end of the gRNA;

[0515] At least about 20% of the nucleotides present in direct or semi-direct repeats are modified;

[0516] At least about 50–75% of the nucleotides present in direct or semi-direct repeats are modified;

[0517] At least about 20% or more of the nucleotides present in the hairpins present in the gRNA scaffold are modified;

[0518] Variable length spacer; and

[0519] A spacer containing a modified nucleotide.

[0520] In the embodiments, the gRNA contains a number of modified nucleotides and / or chemical modifications. Such modifications can increase base editing by about twofold in vivo or in vitro. In the embodiments, the gRNA is 2'- O - It includes methyl or phosphorothioate modifications. In one embodiment, the gRNA is 2'- O - Includes methyl and phosphorothioate modifications. In one embodiment, the modification increases base editing by at least about 2 times.

[0521] The guide polynucleotide may include one or more modifications to provide a nucleic acid having new or enhanced features. The guide polynucleotide may include a nucleic acid affinity tag. The guide polynucleotide may include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.

[0522] gRNA or guide polynucleotides are 5' adenylate, 5' guanosine-triphosphate cap, 5' N7-methylguanosine-triphosphate cap, 5' triphosphate cap, 3' phosphate, 3' thiophosphate, 5' phosphate, 5' thiophosphate, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d-spacer, PC spacer, r-spacer, spacer 18, spacer 9, 3'-3' modification, 2'- O -methylthiophase (MSP), 2'- O -Methyl-PACE(MP) and restricted ethyl[S-cEt], 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dibiotin, PC biotin, psoralen C2, psoralen C6, TINA, 3' DABCYL, Black Hole Quencher 1, Black Hole Quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'- O -Methyl ribonucleoside analogs, modified sugar analogs, wobble / universal bases, fluorescent dye labels, 2'-fluoroRNA, 2'- O - It may also be modified by methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.

[0523] In some cases, phosphorothioate-enhanced RNA gRNA can inhibit RNAase A, RNAase T1, calf serum nuclease, or any combination thereof. These properties may allow the use of PS-RNA gRNA in applications where exposure to nucleases occurs with high probability in vivo or in vitro. For example, a phosphorothioate (PS) conjugation can be introduced between the last 3 to 5 nucleotides at the 5' or 3' end of the gRNA to inhibit nuclease degradation. In some cases, a phosphorothioate conjugation can be added throughout the entire gRNA to reduce attack by nucleases.

[0524] Fusion protein or complex containing a nuclear localization sequence [NLS]

[0525] In some embodiments, the fusion protein or complex provided herein further comprises one or more (e.g., two, three, four, five) nuclear targeting sequences, e.g., nuclear localization sequences [NLS]. In one embodiment, a bipartite NLS is used. In some embodiments, the NLS comprises an amino acid sequence that facilitates the influx of the protein containing the NLS into the cell nucleus (e.g., by a nuclear transporter). In some embodiments, the NLS is fused to the N-terminus or C-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus or N-terminus of an nCas9 domain or a dCas9 domain. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas12 domain. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a cytidine or adenosine deaminase. In some embodiments, the NLS is fused to the fusion protein through one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises an amino acid sequence of any one of the NLS sequences provided or referenced herein. Additional nuclear localization sequences are known in the art and will be apparent to a person skilled in the art. For example, NLS sequences are described in Planck et al.’s PCT / EP2000 / 011690, the contents of which are incorporated herein by reference to the disclosure of exemplary nuclear localization sequences.

[0526] In some embodiments, the NLS is present in the linker, or the NLS is flanked by, for example, the linker described herein. The bipartite NLS comprises two clusters of basic amino acids separated by a relatively short spacer sequence (thus bipartite - 2 parts, whereas not monopartite NLS). The nucleoplasmin NLS KR[PAATKKAGQA]KKKK (SEQ No. 191) is a cluster of two basic amino acids separated by a spacer of about 10 amino acids, which is the prototype of a universal bipartite signal. An exemplary bipartite NLS sequence is as follows: PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ No. 328).

[0527] In some embodiments, any fusion protein or complex provided herein comprises an NLS comprising the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8 through 29 of SEQ ID NO. 328). In some embodiments, any adenosine base editor provided herein comprises an NLS comprising the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8 through 29 of SEQ ID NO. 328). In some embodiments, the NLS is located at the C-terminal portion of the adenosine base editor. In some embodiments, the NLS is located at the C-terminal of the adenosine base editor.

[0528] Additional domains

[0529] The base editor described herein may include any domain that helps facilitate the editing, modification, or alteration of the nucleobases of a polynucleotide. In some embodiments, the base editor includes a polynucleotide programmable nucleotide binding domain (e.g., Cas9), a nucleobase editing domain (e.g., a deaminase domain), and one or more additional domains. In some embodiments, the additional domain may be an inhibitor of a cellular mechanism (e.g., an enzyme) that facilitates the enzymatic or catalytic function of the base editor, the binding function of the base editor, or interferes with the desired base editing result. In some embodiments, the base editor may include a nuclease, a cleaver, a recombinase, a deaminase, a methyltransferase, a methylase, an acetyltransferase, an acetyltransferase, a transcription activator, or a transcription repressor domain.

[0530] In some embodiments, the base editor comprises a uracil glycosylation enzyme inhibitor (UGI) domain. In some cases, the base editor is expressed in cells as a UGI polypeptide and trans. In some embodiments, the cellular DNA repair response to the presence of U:G heterodimer DNA may be responsible for the reduction in nuclear base editing efficiency in the cell. In these embodiments, uracil DNA glycosylation enzyme (UDG) may catalyze the removal of U from the cell's DNA, which may initiate base excision repair (BER), resulting in the return from the U:G pair to the C:G pair, mostly. In these embodiments, BER may be inhibited in the base editor comprising one or more domains that bind to the single strand, block the edited base, inhibit UGI, inhibit BER, protect the edited base, and / or promote the repair of the unedited strand. Accordingly, the present disclosure considers a base editor fusion protein or complex comprising a UGI domain and / or a uracil stabilizing protein (USP) domain.

[0531] Base editor system

[0532] Systems, compositions, and methods for editing nucleosides using a base editor system are provided herein. In some embodiments, the base editor system comprises (1) a base editor [BE] comprising a polynucleotide programmable nucleotide binding domain and a nucleoside editing domain (e.g., a deaminase domain) for editing nucleosides, and (2) a guide polynucleotide (e.g., a guide RNA) together with the polynucleotide programmable nucleotide binding domain. In some embodiments, the base editor system is a cytidine base editor (CBE) or an adenosine base editor (ABE). In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA or RNA binding domain. In some embodiments, the nucleoside editing domain is a deaminase domain. In some embodiments, the deaminase domain may be a cytidine deaminase or a cytosine deaminase. In some embodiments, the deaminase domain may be an adenine deaminase or an adenosine deaminase. In some embodiments, the adenosine base editor can deaminate adenine in DNA. In some embodiments, the base editor can deaminate cytidine in DNA.

[0533] The use of the base editor system provided herein comprises the following steps: (a) bringing a target nucleotide sequence of a polynucleotide of a target (e.g., double or single-stranded DNA or RNA) into contact with a base editor system comprising a nucleobase editor (e.g., an adenosine base editor or a cytidine base editor) and a guide polynucleotide (e.g., gRNA), wherein the target nucleotide sequence comprises a targeted nucleobase pair; (b) inducing strand separation of the target region; (c) converting a first nucleobase of the target nucleobase pair to a second nucleobase on a single strand of the target region; and (d) cutting one or fewer strands of the target region, wherein a third nucleobase complementary to the first nucleobase is replaced with a fourth nucleobase complementary to the second nucleobase. It should be understood that in some embodiments, step (b) is omitted. In some embodiments, the targeted nucleobase pair is a plurality of nucleobase pairs in one or more genes. In some embodiments, the base editor system provided herein can perform multiple editing of multiple nucleobase pairs in one or more genes. In some embodiments, the multiple nucleobase pairs are located in the same gene. In some embodiments, the multiple nucleobase pairs are located in one or more genes, but at least one gene is located at a different locus.

[0534] Components of a base editor system (e.g., a deaminase domain, guide RNA, and / or a polynucleotide programmable nucleotide binding domain) may be associated with each other covalently or non-covalently. For example, in some embodiments, the deaminase domain may be targeted to a target nucleotide sequence by the polynucleotide programmable nucleotide binding domain, and optionally, the polynucleotide programmable nucleotide binding domain forms a complex with a polynucleotide (e.g., guide RNA). In some embodiments, the polynucleotide programmable nucleotide binding domain may be fused to or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide binding domain may target the deaminase domain to a target nucleotide sequence by interacting or associating with the deaminase domain non-covalently. For example, in some embodiments, a nucleobase editing component (e.g., a deaminase component) comprises a corresponding heterogeneous portion that is part of a polynucleotide programmable nucleotide binding domain and / or a guide polynucleotide (e.g., guide RNA) complexed therewith, and an additional heterogeneous portion or domain that can interact with, associate with, or form a complex with an antigen or domain. In some embodiments, a polynucleotide programmable nucleobase binding domain and / or a guide polynucleotide (e.g., guide RNA) complexed therewith comprises a corresponding heterogeneous portion that is part of a nucleobase editing domain (e.g., a deaminase component), and an additional heterogeneous portion or domain that can interact with, associate with, or form a complex with an antigen or domain. In some embodiments, the additional heterogeneous portion binds to, interacts with, or associates with the polypeptide,It may form a complex with it. In some embodiments, the additional heterogeneous portion may bind to, interact with, associate with, or form a complex with the polynucleotide. In some embodiments, the additional heterogeneous portion may bind to the guide polynucleotide. In some embodiments, the additional heterogeneous portion may bind to the polypeptide linker. In some embodiments, the additional heterogeneous portion may bind to the polynucleotide linker. The additional heterogeneous portion may be a protein domain. In some embodiments, additional heterogeneous portions include the 22-amino acid RNA binding domain of lambda bacteriophage antiterminator protein N [N22p], the 2G12 IgG homodimer domain, ABI, an antibody (e.g., an antibody binding to a component of a base editor system or a heterogeneous portion thereof) or a fragment thereof [e.g., heavy chain domain 2 [CH2] of IgM (MHD2) or IgE (EHD2), immunoglobulin Fc region, heavy chain domain 3 [CH3] of IgG or IgA, heavy chain domain 4 [CH4] of IgM or IgE, Fab, Fab2, miniantibodies and / or ZIP antibodies], a barnase-barsta dimer domain, a Bcl-xL domain, a calcineurin A [CAN] domain, cardiac phosphoramban Transmembrane pentameric domain, collagen domain, Com RNA binding domain (e.g., SfMu Com coated protein domain and SfMu Com binding protein domain), cyclophilin-Fas fusion protein [CyP-Fas] domain, Fab domain, Fe domain, fibrin fold-on domain, FK506 binding protein [FKBP] domain,mTOR FKBP binding domain (FRB), fold-on domain, fragment X domain, GAI domain, GID1 domain, glycoporin A transmembrane domain, GyrB domain, Halo tag, HIV Gp41 trimerization domain, HPV45 oncoprotein E7 C-terminal dimer domain, hydrophobic polypeptide, K Homology (KH) domain, Ku protein domain (e.g., Ku heteromer), leucine zipper, LOV domain, mitochondrial antiviral signaling protein CARD filament domain, MS2 coat protein domain (MCP), non-natural RNA aptamer ligand binding to the corresponding RNA motif / aptamer, parathyroid hormone dimerization domain, PP7 coat protein (PCP) domain, PSD95-Dlgl-zo-1 (PDZ) domain, PYL domain, SNAP tag, SpyCatcher It comprises polypeptides and / or fragments thereof, such as a moiety, SpyTag moiety, streptavidin domain, streptavidin binding protein domain, streptavidin binding protein [SBP] domain, telomerase Sm7 protein domain (e.g., Sm7 heterosarcoma or monomeric Sm-like protein). In the embodiments, additional heterogeneous portions comprise polynucleotides (e.g., RNA motifs), such as an MS2 phage operator stem loop (e.g., MS2, MS2 C-5 mutant or MS2 F-5 mutant), non-natural RNA motif, PP7 operator stem loop, SfMu phage Com stem loop, steryl alpha motif, telomerase Ku binding motif, telomerase Sm7 binding motif and / or fragments thereof. Non-limiting examples of additional heterogeneous parts are sequence numbers 380, 382, ​​384,It comprises a polypeptide or a fragment thereof having at least about 85% sequence identity with any one or more of 386–388. Non-limiting examples of additional heterogeneous parts include a polynucleotide or a fragment thereof having at least about 85% sequence identity with any one or more of SEQ ID NOs 379, 381, 383, and 385.

[0535] In some cases, the components of the base editor system are associated with each other through the interaction of leucine zipper domains (e.g., SEQ ID NOs 387 and 388). In some cases, the components of the base editor system are associated with each other through polypeptide domains (e.g., FokI domains) that are associated to form a protein complex comprising about, at least about, or about 1, 2 (i.e., dimerized), 3, 4, 5, 6, 7, 8, 9, 10 or fewer polypeptide domain units, and optionally the polypeptide domains may include modifications that reduce or eliminate their activity.

[0536] In some cases, components of a base editor system are associated with each other through the interaction of a polymeric antibody or a fragment thereof [e.g., heavy chain domain 2 [CH2] of IgG, IgD, IgA, IgM, IgE, IgM(MHD2) or IgE(EHD2), immunoglobulin Fc domain, heavy chain domain 3 [CH3] of IgG or IgA, heavy chain domain 4 [CH4] of IgM or IgE, Fab, and Fab2]. In some cases, the antibody is dimeric, trimeric, or tetrameric. In an embodiment, the dimeric antibody binds to a polypeptide or polynucleotide component of the base editor system.

[0537] In some cases, the components of a base editor system are associated with each other through interactions between polynucleotide binding protein domain(s) and polynucleotides(s). In some cases, the components of a base editor system are associated with each other through interactions between one or more polynucleotide binding protein domains and polynucleotides that are self-complementary and / or complementary to each other, such that when polynucleotides bind complementarily to each other, their respective bound polynucleotide binding protein domain(s) are associated.

[0538] In some cases, the components of a base editor system are associated with each other through the interaction of polypeptide domain(s) and small molecule(s) (e.g., chemical inducers of dimerization [CIDs], also known as 'dimerizing agents'). Non-limiting examples of CIDs include those disclosed in the literature [Amara, et al., "A versatile synthetic dimerizer for the regulation of protein-protein interactions," PNAS, 94:10618-10623 (1997)] and the literature [Voß, et al., "Chemically induced dimerization: reversible and spatiotemporal control of protein function in cells," Current Opinion in Chemical Biology, 28:194-201 (2015)], the whole of which is incorporated herein by reference for all purposes. In some embodiments, the base editor inhibits base excision repair [BER] of the edited strand. In some embodiments, the base editor protects or binds the unedited strand. In some embodiments, the base editor includes UGI activity or USP activity. In some embodiments, the base editor includes a catalytically inactive inosine-specific nuclease.

[0539] The base editor of the present disclosure may include any domain, feature, or amino acid sequence that facilitates editing of a target polynucleotide sequence. For example, in some embodiments, the base editor includes a nuclear localization sequence [NLS]. In some embodiments, the NLS of the base editor is localized between a deaminase domain and a polynucleotide programmable nucleotide binding domain. In some embodiments, the NLS of the base editor is localized to the C-terminus for the polynucleotide programmable nucleotide binding domain.

[0540] The protein domains included in the fusion protein may be heterogeneous functional domains. Non-limiting examples of protein domains that may be included in the fusion protein include deaminase domains (e.g., cytidine deaminase and / or adenosine deaminase), uracil glycosylation enzyme inhibitor (UGI) domains, epitope tags, and reporter gene sequences.

[0541] In some embodiments, the adenosine base editor (ABE) can deaminate adenine in DNA. In some embodiments, the ABE uses the APOBEC1 component of BE3 in natural or genetically modified E. coli [ E. coli ] is generated by replacing TadA with human ADAR2, mouse ADA, or human ADAT2. In some embodiments, ABE comprises an evolved TadA variant. In some embodiments, the base editor is ABE8.1 which comprises or is essentially composed of the sequence of SEQ ID NO. 331 or a fragment thereof having adenosine deaminase activity. Other ABE8 sequences are provided in the attached sequence list (SEQ ID NOs 332–354).

[0542] In some embodiments, the base editor includes an adenosine deaminase variant containing an amino acid sequence, which includes a modification to the ABE 7*10 reference sequence as described herein. Table 7The term 'monomer' used in [this document] refers to the monomeric form of TadA*7.10 containing the modifications described. Table 7 The term 'heteromer' used in [ refers to a specific wild-type E. coli fused to TadA*7.10 containing modifications as described] E. coli ] TadA refers to adenosine deaminase.

[0543]

[0544] In some embodiments, the base editor includes a domain comprising all or part (e.g., a functional part) of a uracil glycosylation enzyme inhibitor (UGI) or uracil stabilizing protein (USP) domain.

[0545] Linker

[0546] In certain embodiments, a linker may be used to link any peptide or peptide domain of the present disclosure. The linker may be as simple as a covalent bond or may be a polymeric linker of many atomic lengths. In certain embodiments, the linker is a polypeptide or is based on amino acids. In other embodiments, the linker is not similar to a peptide. In certain embodiments, the linker is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.).

[0547] In some embodiments, any fusion protein provided herein comprises a cytidine or adenosine deaminase and a Cas9 domain fused together via a linker. To achieve an optimal length for the activity of the cytidine or adenosine deaminase nucleobase editor, various linker lengths and flexibilities between the cytidine or adenosine deaminase and the Cas9 domain (e.g., highly flexible linker forms (GGGS)) n (Sequence No. 246), (GGGGS) n (Sequence No. 247) and (G)n, a form of linker that is more rigid (EAAAK) n (Sequence No. 248), (SGGS)n (Sequence No. 355), SGSETPGTSESATPES (Sequence No. 249) (e.g., see the document whose entire contents are incorporated herein by reference [Guilinger JP, et al. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82]) and (XP)n) may be used. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises a (GGS)n motif, wherein n is 1, 3, or 7. In some embodiments, the cytidine deaminase or adenosine deaminase and Cas9 domain of any fusion protein provided herein are fused through a linker comprising the amino acid sequence SGSETPGTSESATPES (SEQ No. 249), which may also be referred to as an XTEN linker.

[0548] In some embodiments, the domain of the base editor is an amino acid sequence

[0549] It is fused through a linker containing SGGSSGSETPGTSESATPESSGGS (SEQ No. 356), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ No. 357), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ No. 358), EGGSEEEEESGS (SEQ No. 3573), or KGPKPKKEESEK (SEQ No. 3574).

[0550] In some embodiments, the domain of the base editor is fused through a linker containing the amino acid sequence SGSETPGTSESATPES (SEQ No. 249), which may also be referred to as the XTEN linker. In some embodiments, the linker contains the amino acid sequence SGGS (SEQ No. 355). In some embodiments, the linker is 24 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES (SEQ No. 359). In some embodiments, the linker is 40 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGGS (SEQ No. 360). In some embodiments, the linker is 64 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ No. 361). In some embodiments, the linker is 92 amino acids long. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS (SEQ No. 362).

[0551] In some embodiments, the linker comprises a plurality of proline residues and has amino acid lengths of 5 to 21, 5 to 14, 5 to 9, or 5 to 7, e.g., PAPAP (SEQ No. 363), PAPAPA (SEQ No. 364), PAPAPAP (SEQ No. 365), PAPAPAPA (SEQ No. 366), P(AP)4 (SEQ No. 367), P(AP)7 (SEQ No. 368), P(AP)10 (SEQ No. 369) (e.g., see reference to the document whose entire contents are incorporated herein by reference [Tan J, Zhang F, Karcher D, Bock R. Engineering of high-precision base editors for site-specific single nucleotide replacement. Nat Commun. 2019 Jan 25;10(1):439]). Such proline-rich linkers are also referred to as 'rigid' linkers.

[0552] In various embodiments, the linker of the present disclosure comprises 8, 9, 10, 11, 12, 13, 14, or 15 amino acids.In some cases, the linkers are EGSSSKEEEEPG (SEQ No. 647), NSISSSNGQK (SEQ No. 648), GEEGEGSGGGEK (SEQ No. 649), EGEGGKESGSSE (SEQ No. 650), GGGGSSKSPGSE (SEQ No. 651), PIGSDQDD (SEQ No. 652), TEKGQVPHGS (SEQ No. 653), ESGEGGGGSEKK (SEQ No. 654), EEGKPKEGEGSG (SEQ No. 655), ASREPKDSS (SEQ No. 656), KQGSEHDE (SEQ No. 657), SESKSEKGSSEK (SEQ No. 658), QYDSGERSDQ (SEQ No. 659), PGANEEIPGQ (SEQ No. 660), NSPTDEK (SEQ No. 661), EGANEEIPGQ (Sequence No. 662), EGEKEKKKSGES (Sequence No. 663), PGRHEEVPGQ ​​(Sequence No. 664), SKHQTEQDDS (Sequence No. 665), ESEDDSSGRK (Sequence No. 666), KESEKKESESKS (Sequence No. 667), KGEGKSSIKD (Sequence No. 668), DRSQKQDQQD (Sequence No. 669), GPSSTSSS (Sequence No. 670), GSSGEKEEGEPS (Sequence No. 671), GEPKSKKSGSGS (Sequence No. 672), SSGEGGKSESGP (Sequence No. 673), SPQPTSSD (Sequence No. 674), EGGSEEEEESGS (Sequence No. 675), KGPKPKKEESEK (Sequence No. 676), It includes a sequence selected from one or more of SKSQQFVTYE (Sequence No. 677), TGNSKYQTGK (Sequence No. 678), PQPIPHTNPT (Sequence No. 679), ANAHSDISTG (Sequence No. 680), KSQQTEDQSK (Sequence No. 681), QSQDQKQKEH (Sequence No. 682), NQQRPSSD (Sequence No. 683), TTKDTSPKPQ (Sequence No. 684), EGKDNQQTGE (Sequence No. 685), or EPQPDSSE (Sequence No. 686).

[0553] Nucleic acid programmable DNA-binding protein with guide RNA

[0554] Compositions and methods for base editing in cells are provided herein. Compositions are further provided herein comprising a guide polynucleotide sequence, e.g., a guide RNA sequence, or combinations of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more guide RNAs as provided herein. In some embodiments, compositions for base editing as provided herein further comprise a polynucleotide encoding a base editor, e.g., a C base editor or an A base editor. For example, compositions for base editing may comprise an mRNA sequence encoding BE, BE4, ABE, and a combination of one or more guide RNAs as provided herein. Compositions for base editing may comprise a combination of a base editor polypeptide and one or more of any guide RNAs provided herein. These compositions may be used to perform base editing in cells via different delivery approaches, e.g., electroporation, nucleofection, viral transduction, or transfection. In some embodiments, the composition for base editing comprises a combination of an mRNA sequence encoding a base editor and one or more guide RNA sequences provided herein for electroporation.

[0555] Some aspects of the present disclosure provide a system comprising any fusion protein or complex provided herein and a guide RNA bound to the nucleic acid programmable DNA binding protein [napDNAbp] domain of the fusion protein or complex [e.g., Cas9 (e.g., dCas9, nuclease-active Cas9 or Cas9 cleftase) or Cas12]. These complexes are also referred to as ribonucleoproteins [RNP]. In some embodiments, the guide nucleic acid (e.g., guide RNA) is 15 to 100 nucleotides long and comprises a sequence of at least 10 adjacent nucleotides complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is an RNA sequence. In some embodiments, the target sequence is a sequence within the genome of bacteria, yeast, fungi, insects, plants, or animals. In some embodiments, the target sequence is a sequence within the human genome. In some embodiments, the 3' end of the target sequence is immediately adjacent to the standard PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is a non-standard PAM sequence (e.g., Table 3 or immediately adjacent to the sequence listed in 5'-NAA-3'). In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to the sequence within the target gene (e.g., a gene associated with a disease or disorder).

[0556] Some aspects of the present disclosure provide a method of using the fusion protein or complex provided herein. For example, some aspects of the present disclosure provide a method comprising the step of contacting a DNA molecule with any fusion protein or complex provided herein and at least one guide RNA, wherein the guide RNA is about 15 to 100 nucleotide long and comprises a sequence of at least 10 adjacent nucleotides complementary to the target sequence.

[0557] The domains of the base editor disclosed herein may be arranged in any order.

[0558] The defined target region may be a deamination range. The deamination range may be a defined region where the base editor acts on the target nucleotide and deaminates. In some embodiments, the deamination range is within a region of 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases. In some embodiments, the deamination range is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bases upstream of the PAM.

[0559] The base editor of the present disclosure may include any domain, feature, or amino acid sequence that facilitates the editing of a target polynucleotide sequence.

[0560] Method of using a fusion protein or complex containing cytidine or adenosine deaminase and a Cas9 domain

[0561] Some aspects of the present disclosure provide a method of using a fusion protein or complex provided herein. For example, some aspects of the present disclosure provide a method comprising the step of contacting a DNA molecule with any fusion protein or complex provided herein and at least one guide RNA described herein.

[0562] In some embodiments, the fusion protein or complex of the present disclosure is used to edit a target gene. In particular, the cytidine deaminase or adenosine deaminase nucleobase editors described herein can induce multiple mutations within a target sequence. These mutations may affect the function of the target. For example, when a cytidine deaminase or adenosine deaminase nucleobase editor is used to target a regulatory region, the function of the regulatory region is altered, and the expression of downstream proteins is reduced or eliminated.

[0563] Base editor efficiency

[0564] In some embodiments, the purpose of the method provided herein is to alter a gene and / or gene product through gene editing. The nucleobase editing proteins provided herein may be used in gene editing-based human therapeutics in vitro or in vivo. It will be understood by those skilled in the art that a fusion protein or complex comprising the nucleobase editing proteins provided herein, for example, a polynucleotide programmable nucleotide binding domain (e.g., Cas9) and a nucleobase editing domain (e.g., an adenosine deaminase domain or a cytidine deaminase domain), may be used to edit nucleotides from A to G or from C to T.

[0565] Advantageously, as provided herein, a base editor system provides genome editing without generating double-stranded DNA breaks as CRISPR can, without requiring a donor DNA template, and without inducing excessive stochastic insertions and deletions. In some embodiments, the present disclosure provides a base editor that efficiently generates intended mutations, such as stop codons in nucleic acids (e.g., nucleic acids within the target genome), without generating a significant number of unintended mutations, such as unintended point mutations.

[0566] The base editor of the present disclosure advantageously modifies specific nucleotide bases encoding proteins without generating a significant proportion of indels (i.e., insertions or deletions). These indels can lead to frameshift mutations within the coding region of a gene.

[0567] In some embodiments, the base editor provided herein can generate a ratio of intended mutation to indel greater than 1:1 (i.e., intended point mutation: unintended point mutation). In some embodiments, the base editor provided herein is at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 500:1, at least 600:1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1 or more. The ratio of intended mutations to indels can be generated. The number of intended mutations and indels can be verified using any suitable method.

[0568] In some embodiments, the base editor provided herein may limit the formation of indels in nucleic acid regions. In some embodiments, the region is located within a nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the base editor. In some embodiments, any base editor provided herein may limit the formation of indels in nucleic acid regions to less than 1%, less than 1.5%, less than 2%, less than 2.5%, less than 3%, less than 3.5%, less than 4%, less than 4.5%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, less than 10%, less than 12%, less than 15%, or less than 20%.

[0569] Base editing is often referred to as 'modification,' such as genetic modification, gene modification, and modification of nucleic acid sequence, and can be clearly understood based on the context that modification is a base editing modification. Accordingly, a base editing modification is a modification in nucleotide base counts as a result of deaminase activity, for example, discussed throughout this disclosure, which subsequently leads to a change in gene sequence and can affect the gene product.

[0570] In some embodiments, modification, e.g., single base editing, results in a reduction of about or at least about 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% of the gene target expression, or a reduction to an undetectable level.

[0571] The present disclosure provides adenosine deaminase variants (e.g., ABE8 variants) with increased efficiency and specificity. In particular, the adenosine deaminase variants described herein are more likely to edit a desired base within a polynucleotide and less likely to edit a base not intended to be changed (e.g., 'bystander').

[0572] In some embodiments, any base editor system comprising one of the ABE8 base editor variants described herein has bystander editing or mutation reduced by at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% compared to a base editor system comprising an ABE7 base editor, e.g., ABE7.10.

[0573] In some embodiments, any ABE8 base editor variant described herein has higher base editing efficiency compared to the ABE7 base editor. In some embodiments, any ABE8 base editor variant described herein has a base editing efficiency of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145%, 150%, 155%, 160%, 165%, 170%, 175% compared to an ABE7 base editor, e.g., ABE7.10. %, 180%, 185%, 190%, 195%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 450%, or 500% higher.

[0574] The ABE8 base editor variants described herein may be delivered to host cells via plasmids, vectors, LNP complexes, or mRNA. In some embodiments, any ABE8 base editor variant described herein is delivered to host cells as mRNA.

[0575] In some embodiments, the method described herein, for example, the base editing method, has minimal or no off-target effects. In some embodiments, the method described herein, for example, the base editing method, has minimal or no chromosomal translocation effects.

[0576] In some embodiments, the base editing method described herein successfully edits about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of a cell population.

[0577] In some embodiments, the percentage of viable cells in the cell population after base editing intervention exceeds at least 60%, 70%, 80%, or 90% of the starting cell population at the time of base editing. In some embodiments, the percentage of viable cells in the cell population after editing is about 70%. In some embodiments, the percentage of viable cells in the cell population after editing is about 75%. In some embodiments, the percentage of viable cells in the cell population after editing is about 80%. In some embodiments, the percentage of viable cells in the aforementioned cell population is about 85%. In some embodiments, the percentage of viable cells in the aforementioned cell population is about 90% or about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cells in the population at the time of base editing.

[0578] In an embodiment, the cell population is a cell population in contact with the base editor, complex, or base editor system of the present disclosure.

[0579] The number of intended mutations and indels is, for example, International PCT applications PCT / US2017 / 045381 (No. WO2018 / 027078) and PCT / US2016 / 058344 (No. WO2017 / 070632), the full contents of which are incorporated herein by reference, literature [Komor, AC, et.al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM, et.al., "Programmable base editing of A T to G It can be confirmed using any suitable method as described in [G in genomic DNA without DNA cleavage] Nature 551, 464-471 (2017)] and literature [Komor, AC, et.al., "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity" Science Advances 3:eaao4774 (2017)].

[0580] In some embodiments, to calculate indel frequency, a sequence reading is scanned for exact matches to two 10-bp sequences located on both sides of the range where an indel may occur. If no exact match is located, the reading is excluded from analysis. If the length of this indel range exactly matches the reference sequence, the reading is classified as not containing an indel. If the indel range consists of two or more bases that are longer or shorter than the reference sequence, the sequence reading is classified as an insertion or a deletion, respectively. In some embodiments, the base editor provided herein may restrict the formation of indels in a nucleic acid region. In some embodiments, the region is located within a nucleotide targeted by the base editor or within a region of 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides targeted by the base editor.

[0581] Multiple editing

[0582] In some embodiments, the base editor system provided herein may perform multiple editing of multiple nucleobase pairs in one or more genes or polynucleotide sequences. In some embodiments, the multiple nucleobase pairs are located in the same gene or one or more genes, but at least one gene is located at a different locus. In some embodiments, multiple editing comprises one or more guide polynucleotides. In some embodiments, multiple editing comprises one or more base editor systems. In some embodiments, multiple editing comprises one or more base editor systems having a single guide polynucleotide or multiple guide polynucleotides. In some embodiments, multiple editing comprises one or more guide polynucleotides having a single base editor system. It should be understood that the characteristics of multiple editing using any base editor as described herein may apply to any combination of methods using any base editor provided herein. It should also be understood that multiple editing using any base editor as described herein may include sequential editing of multiple nucleobase pairs.

[0583] In some embodiments, a base editor system capable of performing multiple editing of multiple nucleobase pairs in one or more genes includes one of ABE7, ABE8 and / or ABE9 base editors.

[0584] Expression of fusion proteins or complexes in host cells

[0585] The fusion protein or complex of the present disclosure comprising a deaminase may be expressed in substantially any target host cell, including but not limited to bacterial, yeast, fungal, insect, plant, and animal cells, using routine methods known to a person skilled in the art. For example, DNA encoding the adenosine deaminase of the present disclosure may be cloned by designing primers suitable for the upstream and downstream of the CDS based on the cDNA sequence. The cloned DNA may be linked to DNA encoding one or more additional components of a base editing system, either directly, or after cleavage with a restriction enzyme if desired, or after the addition of a suitable linker and / or nuclear localization signal. The base editing system is translated in the host cell to form a complex.

[0586] The polynucleotide encoding the polypeptide described herein can be obtained by chemically synthesizing the polynucleotide or by constructing a polynucleotide (e.g., DNA) encoding its full length by linking partially overlapping oligo short chains synthesized using PCR and Gibson Assembly methods. An advantage of constructing a full-length polynucleotide by chemical synthesis or a combination of PCR or Gibson Assembly methods is that the codons to be used can be selected depending on the host to which the polynucleotide is to be introduced. In the expression of heterologous DNA molecules, the protein expression level is expected to increase by converting its DNA sequence to codons that are very frequently used in the host organism. Codon usage data for host cells (e.g., codon usage data available at kazusa.or.jp / codon / index.html) can be used as a guideline for codon optimization for the polynucleotide sequence encoding the polypeptide. Codons with low usage frequency in the host can be converted into codons with high usage frequency while encoding the same amino acid.

[0587] An expression vector containing a polynucleotide and / or a nucleic acid base conversion enzyme encoding a nucleic acid sequence-recognition module can be generated, for example, by linking DNA downstream of a promoter in a suitable expression vector.

[0588] As an expression vector, Escherichia coli[ Escherichia coli ] derived plasmids (e.g., pBR322, pBR325, pUC12, pUC13); Bacillus subtilis ( Bacillus subtilis Plasmids derived from ) (e.g., pUB110, pTP5, pC194); yeast-derived plasmids (e.g., pSH19, pSH15); insect cell expression plasmids (e.g., pFast-Bac); animal cell expression plasmids (e.g., pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo); bacteriophages such as lambda phage, etc.; insect viral vectors such as baculovirus, etc. (e.g., BmNPV, AcNPV); animal viral vectors such as retrovirus, vaccinia virus, adenovirus, etc. are used.

[0589] Regarding the promoter to be used, any promoter suitable for the host to be used for gene expression may be used. Since the viability of host cells in methods using double-strand breaks is sometimes significantly reduced due to toxicity, it is desirable to increase the number of cells by initiating induction using an inducible promoter. However, since sufficient cell proliferation can be provided by expressing the nucleic acid modification enzyme complex of the present disclosure, constitutive promoters may be used without limitation.

[0590] For example, when the host is an animal cell, the SR.alpha. promoter, SV40 promoter, LTR promoter, cytomegalovirus (CMV) promoter, Rous sarcoma virus (RSV) promoter, Moloney mouse leukemia virus (MoMuLV), LTR, herpes simplex virus thymidine kinase (HSV-TK) promoter, etc., may be used. Among these, the CMV promoter, SR.alpha. promoter, etc., may be used.

[0591] In addition to those mentioned above, the expression vectors used in the present disclosure may include selection markers such as enhancers, splicing signals, termination factors, poly-A addition signals, drug resistance genes, nutritional requirement complementary genes, and replication origins, and such may be used.

[0592] RNA encoding the protein domain described herein may be prepared, for example, by in vitro transcription of a nucleic acid sequence encoding any fusion protein or complex disclosed herein.

[0593] The fusion protein or complex of the present disclosure can be expressed in a cell by introducing an expression vector comprising a nucleic acid sequence encoding the fusion protein or complex into the cell.

[0594] Expression ve...

Claims

Claim 1 A method for modifying a nucleobase of a complement factor B [CFB] polynucleotide, comprising the step of contacting a base editor system comprising one or more guide polynucleotides, or one or more polynucleotides encoding said one or more guide polynucleotides, and a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or one or more polynucleotides encoding said base editor, with said CFB polynucleotide, wherein (a) said one or more guide polynucleotides target said base editor, i. destroy a splice site of said CFB polynucleotide and / or, ii. modify a start codon of said CFB polynucleotide and / or, iii. modify a TATA box of said CFB polynucleotide and / or, iv. (b) introducing a new stop codon into the above CFB polynucleotide, and the deaminase domain comprises a TadA variant (TadA*) having an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1 (SEQ ID NO. 1), or a fragment thereof lacking only N-terminal methionine, wherein the TadA* comprises, compared with the TadA*7.10 amino acid sequence, i. I76Y, V82T, Y123H, Y147T, and Q154S, ii. iii. Y123H, Y147R and Q154R, iv. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v.(c) further comprising a combination of amino acid modifications selected from the group consisting of I76Y, V82S, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147R and Q154R, and / or, (c) one or more guide polynucleotides comprising a nucleic acid sequence selected from CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443) and UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467), and / or. Tables 2A to 2F A method for modifying the nucleobase of the CFB polynucleotide by including at least 10 to 23 consecutive nucleotides of the spacer nucleic acid sequence listed in any one of the methods. Claim 2 A method according to claim 1, wherein the splice site is located near the 5' end of exon 10, exon 1, exon 11, exon 12, exon 14, exon 15, or exon 16 of the CFB polynucleotide. Claim 3 A method according to claim 1, wherein the splice site is located near the 3' end of exon 5, exon 8, exon 9, exon 10, exon 11, exon 14, or exon 18 of the CFB polynucleotide. Claim 4 A method according to claim 1, wherein the one or more guide polynucleotides comprise a spacer complementary to both the human CFB polynucleotide and the non-human primate CFB polynucleotide. Claim 5 A method according to claim 1, wherein one or more guide polynucleotides comprise a spacer complementary to the human CFB polynucleotide. Claim 6 A method according to claim 1, wherein one or more guide polynucleotides comprise a spacer that is complementary to a human CFB polynucleotide but not complementary to a non-human primate CFB polynucleotide. Claim 7 The method of claim 1, wherein the one or more guide polynucleotides comprise a spacer composed of 20 or 21 nucleotides. Claim 8 The method of claim 1, wherein the one or more guide polynucleotides comprise a spacer comprising a nucleotide sequence selected from the group consisting of CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835). Claim 9 A method for modifying the nucleobase of a complement factor B [CFB] polynucleotide, comprising the step of contacting the CFB polynucleotide with one or more guide polynucleotides or one or more polynucleotides encoding said one or more guide polynucleotides, and a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or one or more polynucleotides encoding said base editor, wherein (a) the deaminase domain comprises an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ NO. 1). TadA variants (TadA*) or fragments thereof lacking only N-terminal methionine, wherein the TadA*, when compared to the TadA*7.10 amino acid sequence, comprises i. I76Y, V82T, Y123H, Y147T, and Q154S, ii. Y123H, Y147R, and Q154R, iii. I76Y, Y133H, Y147R, and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R, and Q154R, and vi.(b) further comprising a combination of amino acid modifications selected from the group consisting of I76Y, V82T, Y123H, Y147R, and Q154R, and (b) one or more guide polynucleotides are CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837) and UCUCACCUCUGCAAGUAUUG (SEQ No. A method for modifying the nucleobase of the CFB polynucleotide by including a spacer comprising a nucleotide sequence selected from the group consisting of 1529; TSBTx3835). Claim 10 A method according to claim 1 or 9, wherein the adenosine deaminase comprising the TadA*7.10 amino acid sequence, the deaminase domain further comprises a combination of amino acid changes selected from the group consisting of i. I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147R and Q154R. Claim 11 Method according to claim 1 or 9, wherein the napDNAbp is a splitting enzyme. Claim 12 In claim 11, the napDNAbp binds to a protospacer adjacent motif (PAM) selected from the group consisting of NGA, NGC, NGG, and NNNRRT, wherein 'N' is any nucleotide and 'R' is A or G. Claim 13 In paragraph 12, the above napDNAbp is a Cas9 polypeptide, method. Claim 14 A method according to claim 1 or 9, wherein one or more guide polynucleotides comprise a modified nucleotide. Claim 15 In claim 14, the above one or more guide polynucleotides are: terminal modified SpCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440), terminal modified SaCas9 guide polynucleotide mNsmNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUsmUsmUsmU (SEQ No. 3128);HM01:mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGC UAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),HM07: mNsmNsmNsmNmNmNmNmNmNmNmNNNNNNNNNNNNmGUUUUAGmAmGmCmUmAmGmAmAmAmUmAmGmCmAmAGUUmAAmAAmUAmAmGmGm CmUmAGUmCmCGUUAmUmCAAmCmUmUmGmAmAmAmAmAmGmUmGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),NLS(bpsv40): mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV(sequence numbers 440 and 446),LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGG CUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGUGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 445),NLS+LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmGmAmGmGmAmGmGmGmGmUmGmCmUsmUsmUsmU-NHC5-CrossL- CKRTADGSEFESPKKKRKV(SEQ Nos. 445 and 446), and LONGEST+GOLD: mNs mNs mNs NNNNNNNNNNNNNNNNN GUUUUAGA mGmCmCmGmGmCmGmGmAmAmAmCmGmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAmUmCAAmCmUmUGGACUUCGGUCC mAmAmGUGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmUsmUsmU, comprising a sequence selected from the group consisting of (SEQ No. 447), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' indicates that said nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), where N A method in which the number of nucleotides is 15 to 25. Claim 16 In claim 1 or 9, the method wherein the CFB polynucleotide is present in the cell. Claim 17 In paragraph 16, the above cell is a mammalian cell, method. Claim 18 In paragraph 17, the method wherein the cell is a retinal cell or another cell of the eye, a nerve cell or a liver cell. Claim 19 A cell produced by the method of any one of paragraphs 16 to 18. Claim 20 A base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or one or more polynucleotides encoding said base editor, and CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443) and UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467) and / or Tables 2A to 2F A base editor system comprising one or more guide polynucleotides comprising at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides of a spacer comprising a nucleotide sequence selected from, or one or more polynucleotides encoding said one or more guide polynucleotides. Claim 21 In claim 20, the deaminase domain comprises a TadA variant (TadA*) comprising an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1 (SEQ ID NO. 1), wherein, compared with the TadA*7.10 amino acid sequence, the TadA* comprises i. Y123H, Y147R, and Q154R, ii. I76Y, Y133H, Y147R, and Q154R, iii. V82S and Q164R, iv. A base editor system comprising a combination of amino acid changes selected from the group consisting of I76Y, V82S, Y123H, Y147R and Q154R, v. I76Y, V82T, Y123H, Y147R and Q154R, and vi. I76Y, V82T, Y123H, Y147T and Q154S, wherein the deaminase domain comprises cytidine deaminase. Claim 22 A base editor system according to claim 20, wherein one or more guide polynucleotides comprise a spacer comprising a nucleotide sequence selected from the group consisting of CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837), and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835). Claim 23 In claim 20, the above napDNAbp is a Cas9 clefting enzyme polypeptide that binds to a protospacer adjacent motif (PAM) selected from the group consisting of NGA, NGC, NGG, and NNNRRT, wherein 'N' is any nucleotide and 'R' is A or G, a base editor system. Claim 24 In claim 20, the above one or more guide polynucleotides are: terminal modified SpCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440), terminal modified SaCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUsmUsmUsmU (SEQ No. 3128); HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),HM07: mNsmNsmNsmNmNmNmNmNmNmNmNNNNNNNNNNNNmGUUUUAGmAmGmCmUmAmGmAmAmAmUmAmGmCmAmAGUUmAAmAAmUAmAmGmGm CmUmAGUmCmCGUUAmUmCAAmCmUmUmGmAmAmAmAmAmGmUmGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 440),NLS(bpsv40): mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV(sequence numbers 440 and 446),LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGG CUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGUGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmU (SEQ ID NO: 445),NLS+LONGEST: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmGmAmGmGmAmGmGmGmGmUmGmCmUsmUsmUsmU-NHC5-CrossL- CKRTADGSEFESPKKKRKV(SEQ Nos. 445 and 446), and LONGEST+GOLD: mNs mNs mNs NNNNNNNNNNNNNNNNNGUUUUAGA mGmCmCmGmGmCmGmGmAmAmAmCmGmCmCmGmGmCAAGUUAAAAUAAGGCUAGUCCGUUAmUmCAAmCmUmUGGACUUCGGUCC mAmAmGUGGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmUsmUsmUsmUsmUsmU, comprising a sequence selected from the group consisting of (SEQ No. 447), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' indicates that said nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), where the number of N nucleotides 15 to 25 base editor systems. Claim 25 A polynucleotide encoding a base editor system or a component thereof according to any one of paragraphs 20 through 24. Claim 26 A base editor system of any one of claims 20 to 24 or a vector comprising the polynucleotide of claim 25. Claim 27 A base editor system of any one of claims 20 to 24 or a lipid nanoparticle comprising the polynucleotide of claim 25. Claim 28 A pharmaceutical composition comprising an effective amount of a base editor system of any one of claims 20 to 24, a vector of claim 26, or lipid nanoparticles of claim 27, and a pharmaceutically acceptable excipient. Claim 29 A kit comprising a container containing a base editor system of any one of claims 20 to 24, a vector of claim 26, lipid nanoparticles of claim 27, or a pharmaceutical composition of claim 28. Claim 30 graph 1A to 2F A guide polynucleotide containing a sequence listed in any one of the following. Claim 31 A base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, or a base editor system comprising one or more polynucleotides encoding said base editor, and one or more guide polynucleotides, or one or more polynucleotides encoding said guide polynucleotides, wherein (a) the deaminase domain is a cytidine deaminase, or comprises a TadA variant (TadA*) comprising an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1 (SEQ ID NO. 1), wherein said TadA* Compared with the TadA*7.10 amino acid sequence, i. I76Y, V82T, Y123H, Y147T, and Q154S, ii. Y123H, Y147R, and Q154R, iii. I76Y, Y133H, Y147R, and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R, and Q154R, and vi. (b) further comprising a combination of amino acid modifications selected from the group consisting of I76Y, V82T, Y123H, Y147R and Q154R, and (b) one or more guide polynucleotides are CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534;(c) comprising a spacer having a nucleotide sequence selected from the group consisting of TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837) and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; TSBTx3835), and (c) one or more guide polynucleotides are end-modified SpCas9 guide polynucleotides mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440),HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmAmAmAmGmGmAmGmGmUmGmGmUmGmCmUsmUsmUsmUsmU(Sequence No. 440), and NLS(bpsv40): mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV(Sequence No. A base editor system comprising a sequence selected from the group consisting of 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' represents that said nucleotide 'N' is connected to the next nucleotide by phosphothioate (PS), wherein the number of N nucleotides is 15 to 25, and (d) said napDNAbp is an spCas9 clefting enzyme polypeptide that binds to a protospacer adjacent motif (PAM) selected from the group consisting of NGA, NGC, and NGG, wherein 'N' is any nucleotide.; Claim 32 Lipid nanoparticles comprising the base editor system of claim 31. Claim 33 As a composition, a) a polynucleotide encoding a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, wherein the deaminase domain is cytidine deaminase or comprises a TadA variant (TadA*) comprising an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ No. 1), wherein the TadA*, when compared with the TadA*7.10 amino acid sequence, i. ii. I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. a) further comprising combinations of amino acid modifications selected from the group consisting of I76Y, V82T, Y123H, Y147R, and Q154R, and b) CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837) and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; A nucleotide sequence selected from the group consisting of TSBTx3835) and / or Tables 2A to 2F A composition comprising a guide RNA comprising a spacer comprising at least 10 to 23 consecutive nucleotides of a spacer nucleic acid sequence listed in any one of the above. Claim 34 In claim 33, the guide RNA is a terminally modified SpCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440), HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmAmGmUmGmGmAmAmGmGmUmGmGmUmGmGmUsmUsmUsmU (SEQ No. 440), A composition comprising a sequence selected from the group consisting of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV (SEQ Nos. 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' represents that said nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), wherein the number of N nucleotides is 15 to 25. Claim 35 In paragraph 33, the composition wherein the polynucleotide is mRNA. Claim 36 A composition formulated with lipid nanoparticles (LNP) in any one of claims 33 to 35. Claim 37 As a lipid nanoparticle (LNP) composition, a) mRNA encoding a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, wherein the deaminase domain is cytidine deaminase or comprises a TadA variant (TadA*) comprising an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1 (SEQ ID NO. 1), wherein the TadA*, when compared with the TadA*7.10 amino acid sequence, i. ii. I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. a) further comprising combinations of amino acid modifications selected from the group consisting of I76Y, V82T, Y123H, Y147R, and Q154R, and b) CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837) and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; A nucleotide sequence selected from the group consisting of TSBTx3835) and / or Tables 2A to 2F A composition comprising a guide RNA comprising a spacer comprising at least 10 to 23 consecutive nucleotides of a spacer nucleic acid sequence listed in any one of the above. Claim 38 In claim 37, the guide RNA is a terminally modified SpCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440), HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmAmGmUmGmGmAmGmGmUmGmGmUmGmCmUsmUsmUsmU (SEQ No. 440), and NLS(bpsv40): A sequence selected from the group consisting of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac-CKRTADGSEFESPKKKRKV (SEQ Nos. 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' represents that said nucleotide 'N' is connected to the next nucleotide by phosphorothioate (PS), wherein the number of N nucleotides is 15 to 25, an LNP composition. Claim 39 As a therapeutic method, a) an mRNA encoding a base editor comprising a nucleic acid programmable DNA binding protein [napDNAbp] domain and a deaminase domain, wherein the deaminase domain is cytidine deaminase or comprises a TadA variant (TadA*) comprising an amino acid sequence having at least 90% sequence identity with respect to the TadA*7.10 amino acid sequence of SEQ ID NO. 1 (SEQ ID NO. 1), wherein the TadA* is, compared with the TadA*7.10 amino acid sequence, i. ii. I76Y, V82T, Y123H, Y147T and Q154S, ii. Y123H, Y147R and Q154R, iii. I76Y, Y133H, Y147R and Q154R, iv. V82S and Q164R, v. I76Y, V82S, Y123H, Y147R and Q154R, and vi. a) further comprising combinations of amino acid modifications selected from the group consisting of I76Y, V82T, Y123H, Y147R, and Q154R, and b) CCUCAGAUGUCUAUGUGUUU (SEQ No. 1524; TSBTx3826), UGCUCCCCAUGGCGUUGGAA (SEQ No. 3467; gRNA3657), UUGCUCCCCAUGGCGUUGGA (SEQ No. 3476; gRNA3658), CCCCAUGGCGUUGGAAGGCA (SEQ No. 3443; gRNA3660), GCUUACAAUGACUGAGAUCU (SEQ No. 1534; TSBTx3837), UGCUUACAAUGACUGAGAUCU (SEQ No. 1535; TSBTx3837) and UCUCACCUCUGCAAGUAUUG (SEQ No. 1529; A nucleotide sequence selected from the group consisting of TSBTx3835) and / or Tables 2A to 2F A method comprising the step of administering to a subject a lipid nanoparticle (LNP) comprising a guide RNA comprising a spacer having at least 10 to 23 consecutive nucleotides of a spacer nucleic acid sequence listed in any one of the methods. Claim 40 In claim 39, the guide RNA is a terminally modified SpCas9 guide polynucleotide mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUsmUsmUsmU (SEQ No. 440), HM01: mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmAmGmUmGmGmAmAmGmGmUmGmGmUmGmCmUsmUsmUsmU (SEQ No. 440) and NLS(bpsv40): a sequence selected from the group consisting of mNsmNsmNsNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmUsmUsmUsmU-NHC6-CrossL-ac- CKRTADGSEFESPKKKRKV (SEQ Nos. 440 and 446), wherein 'N' represents any nucleotide, 'mN' represents the 2'-OMe modification of said nucleotide 'N', and 'Ns' represents said nucleotide 'N' being connected to the next nucleotide by phosphorothioate (PS), wherein the number of N nucleotides is 15 to 25.