Compositions and methods for treating alpha-1 antitrypsin

JP2025037975A5Active Publication Date: 2025-10-14BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024208066
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2024-11-29
Publication Date
2025-10-14
Estimated Expiration
2040-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to solve the lung pathology and hepatotoxicity problems in patients with alpha-1 anti-test protein deficiency (A1AD). Traditional treatments such as protein replacement therapy and gene therapy have side effects.

Method used

By modifying the single nucleotide polymorphism (SNP) associated with A1AD, using a modified adenine deaminase called "ABE8", and binding to single guide RNA (guide RNA) and a programmed DNA binding domain, A·T to G·C gene editing is performed precisely in hepatocytes.

Benefits of technology

Implementing improved lung function and reduced hepatotoxicity in patients with A1AD provides a potential treatment option that can address both lung and liver pathological problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025037975000001
    Figure 2025037975000001
  • Figure 2025037975000002
    Figure 2025037975000002
  • Figure 2025037975000003
    Figure 2025037975000003
Patent Text Reader

Abstract

To provide a method of treating patients with alpha-1 anti-trypsin deficiency that addresses both lung pathology and liver toxicity.SOLUTION: The present invention features compositions and methods for editing deleterious mutations associated with alpha-1 anti-trypsin (A1AT) deficiency. In particular embodiments, the invention provides methods for correcting mutations in an A1AT polynucleotide using an adenosine deaminase base editor, ABE8, having unprecedented levels of efficiency.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Patent Application No. 62 / 805,238, filed February 13, 2019; No. 62 / 805,271 filed on May 23, 2019; No. 62 / 852,224 filed on May 23, 2019 No. 62 / 852,228 filed on November 6, 2019; No. 62 / 931,722 filed on November 6, 2019; No. 62 / 941,569 filed on January 27, 2019; and No. 62 / 966,526 filed on January 27, 2020. and international PCT applications claiming the benefit of the following documents, the entire contents of which are incorporated herein by reference: To be incorporated.

[0002] Incorporation by Reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. Unless specifically and individually indicated to be incorporated by reference herein, any patents or patent applications are expressly incorporated by reference. Unless otherwise indicated, this All publications, patents, and patent applications referred to in the specification are hereby incorporated by reference in their entirety. Be absorbed. [Background technology]

[0003] In healthy individuals, alpha-1 antitrypsin (A1AT) is secreted by hepatocytes in the liver. It is produced by the ATPase inhibitor A1AT and secreted into the systemic circulation where it functions as a protease inhibitor. is a particularly good inhibitor of neutrophil elastase and therefore protects tissues and organs such as the lungs Protects against elastin degradation. In patients with alpha-1 antitrypsin deficiency (A1AD), A1 A mutation in the gene encoding AT, which reduces protein production .

[0004] As a result, elastin in the lung is more easily degraded by neutrophil elastase and Over time, lung elasticity is impaired, leading to chronic obstructive pulmonary disease (COPD).

[0005] The most common pathogenic A1AT variant is a glutamic acid to lysine substitution at amino acid 342. This substitution is a guanine to adenine mutation that results in the conversion of the protein to adenine in liver cells. Proteins misfold and polymerize, eventually forming toxic aggregates that cause liver damage and Hepatotoxicity can be prevented by gene knockout (CRISPR / ZFN / TALEN) or genetic These can be addressed by gene knockdown (siRNA), but neither approach has any impact on lung pathology. Lung pathology can be addressed with protein replacement therapy, but this therapy also addresses hepatotoxicity. Gene therapy would also be inappropriate for addressing A1AT gene defects. The livers of A1AD patients are already under a severe disease burden caused by endogenous A1AT. Therefore, gene therapy to increase A1AT in the liver would be counterproductive. Summary of the Invention [Problem to be solved by the invention]

[0006] Therefore, there is a need for treatments for A1AD patients that address both lung pathology and liver toxicity. do. [Means for solving the problem]

[0007] As described below, the present invention provides a method for treating alpha-1 antitrypsin deficiency (A1AD)-associated The present invention features compositions and methods for editing deleterious mutations in a gene. In certain embodiments, The present invention provides a method for correcting mutations associated with A1AD, thereby providing exceptional levels of cytotoxicity (e.g., >60-70%). The modified adenosine deaminase, designated "ABE8," has the same efficiency and specificity as the We provide a treatment for A1AD.

[0008] In one aspect, the present invention provides a method for detecting a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency ( A method for editing an alpha-1 antitrypsin polynucleotide containing a SNP. , this polynucleotide, one or more guide RNAs, and a polynucleotide program possible DNA binding domains and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD adenosine deaminase variants containing an alteration at amino acid position 82 or 166 of contacting a base editor containing at least one base editor domain wherein the guide RNA targets the base editor to generate an alpha The method results in alteration of a SNP associated with α-1 antitrypsin deficiency.

[0009] In another aspect, the present invention provides a method for identifying a single nucleotide polymorphism (SN) associated with alpha-1 antitrypsin deficiency. 1. A method for editing an alpha-1 antitrypsin polynucleotide comprising: one or more guide RNAs, as well as the following sequence: JPEG2025037975000002.jpg171162 (The bold sequence indicates the Cas9-derived sequence, the italicized sequence indicates the linker sequence, and the underlined sequence indicates the bipartite sequence. Polynucleotide-programmable DNA binding containing a (bipartite) nuclear localization sequence Domain and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD adenosine deaminase variants containing an alteration at amino acid position 82 or 166 of contacting a fusion protein containing at least one base editor domain with The method further comprises the steps of:

[0010] In another aspect, the invention relates to a fusion protein of any of the previous aspects, and 5'-ACCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCA AC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAA C UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-UCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC U UGAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UU GAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' The present invention provides a base editing system comprising a guide RNA containing a nucleic acid sequence selected from the following:

[0011] In another embodiment, the cell contains: a base editor or a polynucleoside encoding the base editor, a polynucleotide-programmable DNA-binding domain and a base editor; and the base editor comprising an adenosine deaminase of any of the preceding aspects. or polynucleotides; and targeting base editors to produce alpha-1 anti- One or more guide polynucleotides that result in the A·T to G·C alteration of the SNP associated with trypsin deficiency cleotide In one embodiment, the present invention provides a cell or a progenitor cell thereof produced by introducing In another embodiment, the cells produced are hepatocytes or their progenitor cells. In another embodiment, the cells are derived from a subject with flu-1 antitrypsin deficiency. The cells may be animal or human cells.

[0012] In various embodiments of the above aspects, the gRNA comprises the nucleic acid sequence 5'- GUUUUAGAGC UAGAAAUAGC AAGUU It further contains AAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3.

[0013] In yet another aspect, the present invention provides a method for treating alpha-1 antitrypsin deficiency in a subject. a method of treating a subject, the method comprising administering to the subject cells of any of the preceding aspects. In one embodiment, the cells are autologous or allogeneic to said subject.

[0014] In yet another aspect, the present invention provides a method for producing a cell line comprising: The present invention provides isolated cells or cell populations that have been isolated or expanded.

[0015] In yet another aspect, the present invention provides a method for producing hepatocytes, comprising: (a) alpha-1 alpha a base editor or the like in liver cells containing a SNP associated with antitrypsin deficiency; A polynucleotide encoding a base editor, wherein the base editor is a polynucleotide Nucleotide-programmable nucleotide-binding domains and the above aspects and embodiments The base comprising the adenosine deaminase variant domain described in any one of an editor or polynucleotide; and one or more guide polynucleotides wherein the one or more guide polynucleotides target the base editor. Targeting the A·T to G·C modification of the SNP associated with alpha-1 antitrypsin deficiency The method for producing the change is provided.

[0016] In various embodiments, the hepatocytes are mammalian or human cells.

[0017] In other embodiments of the above aspects, the adenosine deaminase variant is and modifications at positions 166. In another embodiment of the above aspect, the adenosine deaminase barrier In another embodiment of the above aspect, the adenosine deaminase barrier In another embodiment of the above aspect, the adenosine deaminase variant comprises a T166R modification. The ant comprises a V82S and a T166R modification. The amine variants contain one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In other embodiments of the above aspects, the adenosine deaminase variant is The following modifications: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y1 47R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147 R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R. The adenosine deaminase variant includes Y147R + Q154R + Y123H. In one embodiment, the adenosine deaminase variant comprises Y147R + Q154R + I76Y. In other such embodiments, the adenosine deaminase variant is Y147R + Q154R + T166R In other embodiments of the above aspects, the adenosine deaminase variant is Y147T+ In other embodiments of the above aspects, the adenosine deaminase variant includes Y14 7T + Q154S. In other embodiments of the above aspects, the adenosine deaminase variant comprises , Y147R + Q154S. In another embodiment of the above aspect, the adenosine deaminase barrier In another embodiment of the above aspect, the adenosine deaminase inhibitor comprises V82S + Q154S. The variant contains V82S + Y147R. The variant comprises V82S + Q154R. In another embodiment of the above aspect, the enzyme variant comprises V82S + Y123H. The aminase variant comprises I76Y + V82S. The deaminase variant comprises V82S + Y123H + Y147T. In another embodiment, the adenosine deaminase variant comprises V82S + Y123H + Y147R. In an embodiment, the adenosine deaminase variant comprises V82S + Y123H + Q154R. In other embodiments of the above aspects, the adenosine deaminase variant is Y123H + Y147R + In another embodiment of the above aspect, the adenosine deaminase variant In another embodiment of the above aspect, the adenosine Aminase variants include I76Y + V82S + Y123H + Y147R + Q154R. In embodiments, the adenosine deaminase variant is 149, 150, 151, 152, 153, 154 , 155, 156, and 157. In other embodiments of the above aspects, the base editor domain comprises a single base containing V82S and T166R. In another embodiment of the above aspect, the base edited adenosine deaminase variant is The target domain is a wild-type adenosine deaminase domain and adenosine deaminase In other embodiments of the above aspects, the adenosine deaminase variant is , Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In other embodiments of the above aspects, the base editor domain is a TadA7.10 domain. and adenosine deaminase variants. Syndeaminase variants include Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In another embodiment of the above aspect, the base editor further comprises a modification selected from the group consisting of: TadA7.10 domain, as well as Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y 147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R In another embodiment of the above aspect, the adenosine deaminase variant contains a modification wherein the base editor is a sequence of the following sequence or a fragment thereof that has adenosine deaminase activity: ABE8, comprising or consisting essentially of a sequence or fragment thereof: MSEVEFSHEYWM RHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAG AMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD. above In another embodiment of the above aspect, the adenosine deaminase variant has a , 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N and truncated ABE8, in which the terminal amino acid residues are missing. In embodiments, the adenosine deaminase variant has 1, 2, 3, 4 or more amino acids relative to full-length ABE8. , 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acids Contains a truncated ABE8 in which residues are missing.

[0018] In other embodiments of the above aspects, the A· at a SNP associated with alpha-1 antitrypsin deficiency The T to G·C modification results in a reduced glutamate in the alpha-1 antitrypsin polypeptide. In another embodiment of the above aspect, the alpha-1 antitrypsin deficiency-associated The associated SNP is an alpha-1 antitrypsin polypeptide with a lysine at amino acid position 342. In another embodiment of the above aspect, the expression of The SNP results in a substitution of glutamic acid with lysine. Select cells for the A·T to G·C alteration of a SNP associated with 1-antitrypsin deficiency In other embodiments of the above aspects, the polynucleotide programmable DNA binding domain comprises: Modified Staphylococcus aureus Cas9(SaCas9), Streptococcus thermophilus 1 Cas9(St1 Cas9), modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In other embodiments of the above aspects, the polynucleotide-programmable DNA binding domain is having specificity for a variant protospacer adjacent motif (PAM) or for non-G PAMs; In other embodiments of the above aspects, the modified PAM comprises the nucleic acid sequence 5'-NGC In other embodiments of the above aspects, the modified SpCas9 has specificity for the -3' amino acid residue. Conversion, D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or In another embodiment of the above aspect, the polynucleotide protease comprises the corresponding amino acid substitution. Gram-capable DNA-binding domains are nuclease-inactive or nickase variants In other embodiments of the above aspects, the nickase variant has the amino acid substitution D10A or In other embodiments of the above aspects, the base editor comprises a zinc In another embodiment of the above aspect, the adenosine deaminase The ze domain can deaminate adenine in deoxyribonucleic acid (DNA). In other embodiments of the above aspects, the one or more guide RNAs are CRISPR RNA (crRNA) and It contains a trans-encoded small RNA (tracrRNA), which encodes alpha-1 antitrinucleotides. Nucleic acid sequences complementary to alpha-1 antitrypsin nucleic acid sequences containing SNPs associated with alpha-1 antitrypsin deficiency In other embodiments of the above aspects, the base editor and one or more guideposts In another embodiment of the above aspect, the base edta is a nucleotide that forms a complex in the cell. The filter is an alpha-1 antitrypsin gene containing a SNP associated with alpha-1 antitrypsin deficiency. Complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the trypsin nucleic acid sequence Form.

[0019] In another embodiment, for treating alpha-1 antitrypsin deficiency (A1AD) in a subject. The method of claim 1, further comprising administering to a subject an adenosine deoxyribonucleotide inserted into a Cas9 or Cas12 polypeptide. A fusion protein containing a serotype variant, or a polypeptide encoding the fusion protein. and a single nucleotide that targets the fusion protein and associates with A1AD. One or more guide polynucleotides that result in a single nucleotide polymorphism (SNP) A·T to G·C alteration are administered. administering to the subject a therapeutically effective amount of a compound selected from the group consisting of benzodiazepines, ... and benzodiazepines, thereby treating A1AD in the subject.

[0020] In another aspect, a method of treating alpha-1 antitrypsin deficiency (A1AD) in a subject an adenosine base editor, ABE8, or a gene encoding said base editor; ABE8 is an adenosine trinucleotide inserted into a Cas9 or Cas12 polypeptide. the base editor or polynucleotide, comprising a deaminase variant; and One or more of the following targeting ABE8 results in an A·T to G·C alteration of the SNP associated with A1AD: administering the above guide polynucleotide, thereby treating A1AD in the subject. The method includes:

[0021] In the embodiment of the above method, ABE8 is ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8 .5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, A BE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8 .20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, ABE8.24-m, ABE8.1-d, ABE8.2-d, ABE8.3-d , ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11 -d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d , ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d In an embodiment of the above method, the adenosine deaminase variant is: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD wherein the amino acid sequence comprises at least one modification. adenosine deaminase variants containing alterations at amino acid positions 82 and / or 166 In one embodiment, the at least one modification is: V82S, T166R, Y147T, Y147R, Q154S, Y12 Contains 3H, and / or Q154R.

[0022] In one embodiment of the above method, the adenosine deaminase variant has the following combination of modifications: Matching: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V 82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y14 7R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q1 54R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I7 6Y + V82S + Y123H + Y147R + Q154R. The deaminase variants are TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, and TadA*8.6. dA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13 , TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, T TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. The deaminase variants are 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the C-terminal deletion begins at a residue selected from the group consisting of adenosine The deaminase variant is an adenosine deaminase variant containing the TadA*8 adenosine deaminase variant domain. In one embodiment, adenosine deaminase is a monomer of adenosine deaminase. The adenosine deaminase variant contains the wild-type adenosine deaminase domain and the TadA*8 adenosine deaminase domain. Heterodimers of adenosine deaminase containing deaminase variant domains In one embodiment, the adenosine deaminase variant is a TadA domain. adenosine deaminase containing TadA and TadA*8 adenosine deaminase variant domains In one embodiment of the method, the A·T to G heterodimer at the SNP associated with A1AD is The modification to C changes the glutamic acid at amino acid position 342 to a lysine. In an embodiment, the SNP associated with A1AD is alpha-1 antitrinucleotide with lysine at amino acid position 342. In one embodiment of the method, the alpha-1 antithrombin gene is expressed as a psin polypeptide. The SNP associated with pusin deficiency substitutes glutamic acid for lysine.

[0023] In one embodiment of the above method, the adenosine deaminase variant is Cas9 or Cas12. flexible loops, alpha helical regions, unstructured parts, or solvent access of polypeptides In one embodiment of the method, the adenosine deaminase variant is inserted into the target region. The fragments are flanked by N- and C-terminal fragments of the Cas9 or Cas12 polypeptide.

[0024] In one embodiment of the above method, the fusion protein or ABE8 has the structure NH2-[Cas9 or Cas12 N-terminal fragment of polypeptide]-[adenosine deaminase variant]-[Cas9 or Cas12 polypeptide The C-terminal fragment of the peptide comprises ]-COOH, where each instance of "]-[" is an optional linker. In embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment is a Cas9 or Cas12 polypeptide. In one embodiment, the flexible loop comprises a portion of an adenosine deaminase (ADD) amino acid sequence. When the enzyme variant deaminates the target nucleobase, it deaminates the amino acids adjacent to the target nucleobase. Includes.

[0025] In one embodiment of the above method, the method comprises detecting deamination of a SNP target nucleobase associated with A1AD. The method further comprises administering to the subject a guide nucleic acid sequence to achieve this. In this state, deamination of the SNP target nucleobase is achieved by replacing the target nucleobase with a wild-type nucleobase or with a non- Deamination of the targeted nucleobase by substitution with the wild-type nucleobase alleviates the symptoms of A1AD. In one embodiment of the method, the deamination of the SNP associated with A1AD involves replacing glutamic acid with lysine. do.

[0026] In one embodiment of the above method, the target nucleobase is PA in the target polynucleotide sequence. In one embodiment, the target nucleobase is 2 to 12 nucleobases from the PAM sequence. In one embodiment of the method, the nucleobase is upstream of the N-terminus of the Cas9 or Cas12 polypeptide. The fragment or C-terminal fragment binds to the target polynucleotide sequence. The N- or C-terminal fragment contains the RuvC domain; the N- or C-terminal fragment contains the HNH Neither the N-terminal nor the C-terminal fragment contains the HNH domain; Alternatively, neither the N-terminal fragment nor the C-terminal fragment contains a RuvC domain. The Cas9 or Cas12 polypeptide may have a partial or complete deletion in one or more structural domains. wherein the deaminase is located at the position of the partial or complete deletion of the Cas9 or Cas12 polypeptide. In certain embodiments, the deletion is within the RuvC domain; or the deletion bridges the RuvC domain and the C-terminal domain.

[0027] In one embodiment of the above method, the fusion protein or ABE8 comprises a Cas9 polypeptide. In embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus occus aureus Cas9(SaCas9), Streptococcus thermophilus 1 Cas9(St1Cas9), and and variants thereof. In one embodiment, the Cas9 polypeptide has the following amino acid sequence: Cas9 reference sequence): JPEG2025037975000003.jpg177162 (single underline: HNH domain; double underline: RuvC domain; (Cas9 reference sequence), or its counterpart In certain embodiments, the Cas9 polypeptide is a region corresponding to a Cas9 polypeptide reference sequence. Cas9, containing a deletion of amino acids 1017 to 1069, or the corresponding amino acids; The polypeptide is selected from amino acids 792 to 872 of the Cas9 polypeptide reference sequence, as numbered; or the Cas9 polypeptide comprises a deletion of the corresponding amino acid; Deletion of amino acids 792 to 906, or the corresponding amino acids, in the reference sequence. In one embodiment of the above method, the adenosine deaminase variant is In one embodiment, the flexible loop is inserted within a flexible loop of the Cas9 reference sequence. Numbered as 530-537, 569-579, 686-691, 768-793, 943-947, 1002-10 40, 1052-1077, 1232-1248, and 1298-1300, or the corresponding amino acid positions The amino acid residues are selected from the group consisting of:

[0028] In one embodiment of the above method, the deaminase variant is a variant of the Cas9 deaminase that is numbered in the Cas9 reference sequence. The amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, and 1026 ~1027, 1029~1030, 1040~1041, 1052~1053, 1054~1055, 1067~1068, 1068~1069, It is inserted between 1247 and 1248, or 1248 and 1249, or the corresponding amino acid positions. In one embodiment of the method, the deaminase variants are numbered in the Cas9 reference sequence. Such as amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1040-1041, 1068- The amino acid sequence is inserted between amino acid positions 1069, 1247 and 1248, or the corresponding amino acid sequences. In one embodiment, the deaminase variant is a Cas9 variant as numbered in the Cas9 reference sequence. , amino acid positions 1016-1017, 1023-1024, 1029-1030, 1040-1041, 1069-1070, or 12 In one embodiment of the method, the amino acid sequence is inserted between amino acid positions 47 and 1248, or the corresponding amino acid sequence. The adenosine deaminase variants are expressed by the Cas9 polypeptide at the loci identified in Table 13A. In one embodiment, the N-terminal fragment is inserted within the amino acid residues of the Cas9 reference sequence. Units 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, and and / or residues 1248-1297, or the corresponding residues thereof. , amino acid residues 1301 to 1368, 1248 to 1297, 1078 to 1231, 1026 to 1051, and 948 of the Cas9 reference sequence containing residues 1001, 692-942, 580-685, and / or 538-568, or their corresponding residues nothing.

[0029] In one embodiment of the above method, the Cas9 polypeptide is a modified Cas9 and a modified PAM or non-GP In one embodiment of the above method, the Cas9 polypeptide has specificity for AM. or the Cas9 polypeptide is nuclease inactive. In one embodiment, the Cas9 polypeptide is a modified SpCas9 polypeptide. , amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T133 7R (SpCas9-MQKFRAER) and has specificity for the modified PAM 5'-NGC-3'.

[0030] In one embodiment of the above method, the fusion protein or ABE8 comprises a Cas12 polypeptide. In one embodiment, the adenosine deaminase variant is introduced into a Cas12 polypeptide. In one embodiment, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12 In one embodiment, the adenosine deaminase antibody is Cas12e, Cas12g, Cas12h, or Cas12i. The variants were at amino acid positions: a) 153–154, 255–256, 306–307, 980–981, and 101 in BhCas12b; 9-1020, 534-535, 604-605, or 344-345, or Cas12a, Cas12c, Cas12d, Cas b) the corresponding amino acid residues of BvCas12b; , 248–249, 299–300, 991–992, or 1031–1032, or Cas12a, Cas12c, Cas12d , the corresponding amino acid residues of Cas12e, Cas12g, Cas12h, or Cas12i; or c) AaCas 157–158, 258–259, 310–311, 1008–1009, or 1044–1045 of 12b, or Cas12a , Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i corresponding amino acid residues In one embodiment, the adenosine deaminase variant is one of the variants identified in Table 13B. In one embodiment, the Cas12 polypeptide is inserted at a locus corresponding to the target gene. In one embodiment, the Cas12 polypeptide is a BhCas12b domain, a BvCas12b domain, or a BhCas12b domain. b domain, or AACas12b domain.

[0031] In one embodiment of the above method, the guide RNA is a CRISPR RNA (crRNA) and a transactivating crRNA. In one embodiment of the above method, the subject is a mammal or a human. .

[0032] In another aspect, the base editing system of any one of the above methods, aspects and embodiments, and a pharmaceutically acceptable carrier, vehicle, or excipient. Provide.

[0033] In one aspect, the cells of the above aspects and embodiments and a pharmaceutically acceptable carrier. The present invention provides a pharmaceutical composition comprising a compound, a vehicle, or an excipient.

[0034] In another aspect, a method for the preparation of a nucleic acid comprising the base editing system of any one of the above methods, aspects and embodiments. We provide a kit that includes:

[0035] In another aspect, a kit is provided comprising a cell of any one of the above aspects and embodiments. In one embodiment of the kit, the kit further comprises a package containing instructions for use. Includes insert.

[0036] In one aspect, provided herein are polynucleotide-programmable DNA-binding domains. Inn, and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD adenosine deaminase variants comprising a modification at amino acid position 82 or 166 of a base editor comprising at least one base editor domain.

[0037] In one embodiment, a base editor system comprises the base editor described above and a guide RNA; The guide RNA targets the base editor to encode alpha-1 antitrypsin. In some embodiments, the adenosine deaminase gene is a nucleotide sequence that results in an alteration of a SNP associated with steroid deficiency. In some embodiments, the enzyme variant comprises a V82S modification and / or a T166R modification. The adenosine deaminase variants contain the following modifications: Y147T, Y147R, Q154S, Y123H, and and Q154R. In some embodiments, the base editor domain further comprises one or more of: Contains wild-type adenosine deaminase domain and adenosine deaminase variants In some embodiments, the adenosine deaminase heterodimer comprises an adenosine deaminase heterodimer. The aminase variants are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, Truncated Ta lacking 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues In some embodiments, the adenosine deaminase variant is full-length TadA8. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or a truncated TadA8 lacking the 20 C-terminal amino acid residues. , a polynucleotide-programmable DNA-binding domain is engineered Staphylococcus aureus Cas 9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), modified Streptococcus pyo In some embodiments, the polynucleotide is a nucleotide sequence encoding the nucleotide sequence of interest. The oxid programmable DNA binding domain contains engineered protospacer adjacent motif (PAM) features SpCas9 variants with specificity for heteromeric or non-G PAMs. In this form, the polynucleotide-programmable DNA-binding domain binds to a nuclease-inactive Cas 9. In some embodiments, the polynucleotide programmable DNA binding domain is , Cas9 nickase.

[0038] In one embodiment, one or more guide RNAs as well as the following sequence: JPEG2025037975000004.jpg195164 (The bold sequence indicates the Cas9-derived sequence, the italicized sequence indicates the linker sequence, and the underlined sequence indicates the bipartite sequence. a polynucleotide-programmable DNA binding domain comprising a nucleic acid sequence (which represents a nuclear localization sequence); and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD adenosine deaminase variants containing modifications at amino acid positions 82 and / or 166 of a fusion protein comprising at least one base editor domain, -Provide a system.

[0039] In one aspect, a cell comprising any one of the above base editor systems is provided. In some embodiments, the cell is a human cell or a mammalian cell. , the cells may be ex vivo, in vivo, or in vitro.

[0040] The present invention provides a method for editing mutations associated with alpha-1 antitrypsin deficiency (A1AD). The compositions and articles defined by the present invention are: isolated or otherwise prepared in connection with the examples provided herein. The features and advantages of the present invention will be apparent from the detailed description and claims.

[0041] definition The following definitions supplement those in the art and are intended for the present application and are not intended to be limiting. Related or unrelated matters of interest, such as commonly owned patents or applications, Any methods and materials similar or equivalent to those described herein are not intended to be limiting. Although preferred materials and methods may be used in carrying out the tests of the present disclosure, Accordingly, the terminology used herein is intended to be illustrative and not restrictive. are for illustrative purposes only and are not intended to be limiting.

[0042] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the meaning commonly understood by one of ordinary skill in the art to which the invention pertains. , provides those skilled in the art with general definitions of many of the terms used in this invention: n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The C ambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary o f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0043] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, the singular forms "a," "an," and "the" are used unless the context clearly indicates otherwise. It should be noted that unless specifically indicated, plural referents are included. In this regard, the use of "or" means "and / or" and is inclusive unless otherwise stated. Furthermore, the terms "including" and "include" are understood to be generic. The use of other types such as "," "includes," and "included" is non-limiting. It is target.

[0044] As used in this specification and claims, the term "comprising" (and "comprise" and "comprises" and all its forms), "hav "ing" (any of its forms, such as "have" and "has"), "including" "including" ("include" and "includes") or "Containing" (such as "contains" and "contain") Any form thereof) is inclusive or open-ended and may include additional, unlisted elements or does not exclude any method steps. Any embodiment discussed herein may be combined with any method of the present disclosure. It is contemplated that the same may be performed on the composition, or vice versa. The compositions of the present disclosure can be used to achieve the methods of the present disclosure.

[0045] The terms "about" or "approximately" refer to a range of values ​​as determined by one of ordinary skill in the art. This means that the value is within an acceptable margin of error for a particular value, which indicates how How it is measured or determined depends in part on the limitations of the measurement system. For example, "about" means within 1 or more than 1 standard deviation, according to the practice in the art. Alternatively, "about" can mean up to 20%, up to 10%, up to 5%, or can mean a range of up to 1%. Alternatively, particularly with respect to biological systems or processes , the term can mean values ​​within the same order of magnitude, for example, within 5-fold or within 2-fold. Where a particular value is recited in the application and claims, that value is used unless otherwise stated. The term "about" means that a particular value is estimated to be within an acceptable margin of error. It should be.

[0046] The ranges provided herein relate to all values ​​within the range, inclusive of the first and last values. For example, the range 1 to 50 is understood to be 1, 2, 3, 4, 5, 6, 7, 8, 9. , 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 , 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50. It is understood.

[0047] In the specification, "some embodiments," "an embodiment," "one embodiment," or Reference to "another embodiment" may include any particular features, structures, or features are included in at least some embodiments of the present disclosure, but not necessarily all. This means that the embodiments are not necessarily included.

[0048] "Adenosine deaminase" refers to the enzyme that hydrolyzes the deamination of adenine or adenosine. In some embodiments, the term "polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing The deaminase or deaminase domain converts adenosine to inosine or deo Adenosine dehydrogenase catalyzing the hydrolytic deamination of hydroxyadenosine to deoxyinosine In some embodiments, adenosine deaminase is a deoxyribonuclease. Catalyzes the hydrolytic deamination of adenine or adenosine in nucleic acids (DNA). Adenosine deaminase (e.g., genetically engineered adenosine deaminase) provided in the literature The adenosine deaminase (evolved adenosine deaminase) can be from any organism, such as a bacterium.

[0049] In some embodiments, the adenosine deaminase comprises a modification in the following sequence: : MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNY RLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLC YFFR MPRQVFNAQK KAQSSTD (also known as TadA*7.10).

[0050] In some embodiments, TadA*7.10 comprises at least one modification. In certain embodiments, TadA*7.10 contains modifications at amino acid positions 82 and / or 166. In some embodiments, variants of the above reference sequences include one or more of the following modifications: Y147T, Y1 47R, Q154S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H is also referred to herein. In this study, it is also called H123H (a modification in which the modification H123Y in TadA*7.10 is reverted to Y123H (wt)). In other embodiments, the variants of the TadA*7.10 sequence are Y147T + Q154R; Y147T + Q154 S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154 R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154 R

[0051] In other embodiments, the present invention provides adenosine deaminase variants comprising deletions, e.g. , containing C-terminal deletions beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157 In another embodiment, the adenosine deaminase variant is one or more T containing the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R In another embodiment, the adenosine deaminase Ant is: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y14 7R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R It is a monomer.

[0052] In still other embodiments, the adenosine deaminase variants each comprise one or more of the following: Two with the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R It is a homodimer containing the adenosine deaminase domain of TadA*8 (e.g., TadA*8). In embodiments, the adenosine deaminase variants are: Y147T + Q154R; Y147T + Q15 4S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q15 4R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q15 Two adenosine deaminase domains having a combination of modifications selected from the group of 4R It is a homodimer containing a nucleotide sequence (e.g., TadA*8).

[0053] In other embodiments, the adenosine deaminase variant is a wild-type TadA adenosine deaminase. The amino acid sequence of the ... and / or adenosine deaminase variant domains (e.g., T In another embodiment, the adenosine deaminase gene is a heterodimer comprising adenosine deaminase (adA*8). The variants contain the wild-type TadA adenosine deaminase domain, as well as: Y147T + Q154R; 147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y1 23H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q15 4R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y 147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y1 Adenosine deaminase barrier comprising a combination of alterations selected from the group of 47R + Q154R It is a heterodimer containing a target domain (e.g., TadA*8).

[0054] In another embodiment, the adenosine deaminase variant is a TadA*7.10 domain, and one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. Heterodimers containing adenosine deaminase variant domains (e.g., TadA*8) containing the above In another embodiment, the adenosine deaminase variant is a TadA*7.10 domain. Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154 S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147 R + Q154R; or adenosine containing the combination I76Y + V82S + Y123H + Y147R + Q154R It is a heterodimer containing a deaminase variant domain (e.g., TadA*8).

[0055] In one embodiment, the adenosine deaminase has the following sequence or adenosine deaminase and TadA*8, comprising or consisting essentially of a fragment thereof having enzyme activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0056] In some embodiments, TadA*8 is truncated. The truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and 13 amino acids compared to the full-length TadA*8. 3, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues are missing. In embodiments, the truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues are missing In some embodiments, the adenosine deaminase variant is full-length TadA*8. .

[0057] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and and an adenosine deaminase domain selected from one of the following:

[0058] Staphylococcus aureus (S. aureus)TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTL YVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0059] Bacillus subtilis (B. subtilis)TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTL EPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLS E

[0060] Salmonella typhimurium(S. typhimurium)TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLV LQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFF I DRAW RMRRQEICALKCADE

[0061] Shewanella putrefaciens(S. putrefaciens)TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEP CAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFKRRRDEKKALQRAQ QGIE

[0062] Haemophilus influenzae F3031(H. influenzae)TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYR LLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRRE EKKIEKALLKSLSDK

[0063] Caulobacter crescentus(C. crescentus)TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLT DLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAK I

[0064] Geobacter sulfurreducens(G. sulfrreducens)TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLT GATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRK KAKATPALFIDERKVPPEP

[0065] TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0066] An "adenosine deaminase base editor (ABE8) polypeptide" is defined as any polypeptide having the following reference sequence: : MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD adenosine deaminase variants containing modifications at amino acid positions 82 and / or 166 of "BE" means a base editor (BE) as defined and / or described herein, including a base editor (BE).

[0067] In some embodiments, ABE8 contains additional modifications compared to the reference sequence.

[0068] "Adenosine deaminase base editor 8 (ABE8) polynucleotide" means an ABE8 polynucleotide. It refers to a polynucleotide (polynucleotide sequence) that encodes a peptide.

[0069] "Administering" as used herein refers to administering to a patient or subject one or more of the compositions described herein. By way of example and without limitation, administration of a composition, e.g., injection, refers to intravenous administration of a composition. intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, or It may be by intramuscular (im) injection. One or more of these routes may be used. Parenteral administration may be, for example, by bolus injection or by gradual perfusion over time. Alternatively, or concurrently, administration may be by the oral route.

[0070] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or Rui means a fragment of it.

[0071] "Alpha-1 antitrypsin (A1AT) protein" is a small molecule identified in UniProt accession number P01009. It refers to a polypeptide or fragment thereof having at least about 95% amino acid sequence identity. In certain embodiments, the A1AT protein has one or more modifications compared to the following reference sequences: In one particular embodiment, the A1AT protein associated with A1AD comprises an E342K mutation. An exemplary A1AT amino acid sequence is provided below. >sp|P01009|A1AT_HUMAN Alpha-1-Antitrypsin OS=Homo sapiens OX=9606 GN=SERP INA1 PE=1 SV=3 MPSSVSWGILLLAGLCCLVPVSLAEDPQGDAAQKTDTSHHDQDHPTFNKITPNLAEFAFFSLYRQLAHQSNSTNIFFSPVS IATAFAMLSLGTKADTHDEILEGLNFNLTEIPEAQIHEGFQELLRTLNQPDSQLQLTTGNGLFLSEGLKLVDKFLEDVKK LYHSEAFTVNFGDTEEAKKQINDYVEKGTQGKIVDLVKELDRDTVFALVNYIFFKGKWERPFEVKDTEEEDFHVDQVTTV KVPMMKRLGMFNIQHCKKLSSWVLLMKYLGNATAIFFLPDEGKLQHLENELTHDIITKFLENEDRRSASLHLPKLSITGT YDLKSVLGQLGITKVFSNGADLSGVTEEAPLKLSKAVHKAVLTIDEKGTEAAGAMFLEAIPMSIPPEVKFNKPFVFLMIE QNTKSPLFMGKVVNPTQK In the above A1AT protein sequence, the first 24 amino acids constitute a signal peptide ( (Underlined). Position 342 of the sequence mutated in A1AD (i.e., E342K) is the The amino acid sequence is determined based on the amino acid residue "E" which is set as amino acid "1" in the

[0072] "Alteration" refers to any alteration that can be detected by standard art known methods such as those described herein. Such a change in the structure, expression level or activity of a gene or polypeptide (e.g., an increase or As used herein, modification means a change or reduction in a polynucleotide or polypeptide. Changes in polypeptide sequence or expression level, e.g., 25%, 40%, 50% This includes changes in the expression levels of the above.

[0073] "Ameliorate" means to reduce, inhibit, or attenuate the occurrence or progression of a disease. "To cause, reduce, stop, or stabilize" means to cause, reduce, stop, or stabilize.

[0074] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polynucleotide or polypeptide analog may be a derivative of a corresponding naturally occurring polypeptide. a naturally occurring polynucleotide while retaining the biological activity of the oligonucleotide or polypeptide A specific biochemical compound that enhances the function of the analog compared to a peptide or polypeptide. Such modifications can be used to modify the structure of the analog without altering, for example, ligand binding. Affinity for DNA, potency, specificity, protease or nuclease resistance, membrane The analogs can increase permeability and / or half-life. may include:

[0075] A "base editor (BE)" or "nucleobase editor (NBE)" is a In various embodiments, the base modifying agent is an agent that binds to a base and has nucleobase modifying activity. Determinants include nucleic acid base-modifying polypeptides (e.g., deaminases) and nucleic acid programs. The nucleotide-binding domain is then coupled to a guide polynucleotide (e.g., a guide RNA). In various embodiments, the agent comprises a protein domain with base editing activity, i.e., That is, bases (e.g., A, T, C, G, U) within a nucleic acid molecule (e.g., DNA) can be modified. In some embodiments, the polynucleotide protease is a biomolecular complex comprising a domain. The gram-capable DNA binding domain is fused or linked to the deaminase domain. In one embodiment, the agent is a fusion protein comprising a domain with base editing activity. In embodiments, the protein domain having base editing activity is linked to a guide RNA (e.g., For example, an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase In some embodiments, the domain having base editing activity is located within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating bases. One or more bases in the molecule can be deaminated. In some embodiments, the bases The editor can deaminate adenosines (A) in DNA. In some embodiments, the base editor is an adenosine base editor (ABE).

[0076] In some embodiments, the base editor is a circular permutant Cas9 (e.g., spCas9 or saCas9) and adenosine deoxyribonucleic acid (AD) within a scaffold containing a bipartite nuclear localization sequence. It is generated by cloning a cytosine aminotransferase variant (e.g., TadA*8) (e.g., Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176 , 254-267, 2019. Exemplary circular permutations are as follows, with the bolded sequences: Columns indicate Cas9-derived sequences, italicized sequences indicate linker sequences, and underlined sequences indicate bipartite nuclear localization sequences. The sequence is shown.

[0077] CP5 (MSP; NGC = Pam variant containing the mutation; normal Cas9 prefers NGG; PID = protein (including the protein-interacting domain and "D10A" nickase): JPEG2025037975000005.jpg169162

[0078] In some embodiments, ABE8 is a base editor selected from the group consisting of a base editor from Tables 6-9, 13, or 14 below. In some embodiments, ABE8 is an adenosine deoxyribonucleotide that has evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is The variant is a TadA*8 variant as described in Tables 7, 9, 13, or 14 below. In some embodiments, the adenosine deaminase variant is Y147T, Y147R, Q154 Ta containing one or more modifications selected from the group of S, Y123H, V82S, T166R, and / or Q154R. dA*7.10 variant (e.g., TadA*8). In various embodiments, ABE8 is Y147T + Q15 4R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123 H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123 TadA*7.10 variant ( In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. is an array: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD Includes.

[0079] In some embodiments, the polynucleotide programmable DNA binding domain is CRISP In some embodiments, the base editor is an R-associated (e.g., Cas or Cpf1) enzyme. A catalytically inactive (dead) Cas9 (dCas9) fused to a deaminase domain In some embodiments, the base editor comprises a Ca 2+ -binding domain fused to a deaminase domain. The base editor is s9 nickase (nCas9). For details of the base editor, see International PCT Application No. PCT / 2017 / 0453 81 (WO 2018 / 027078) and PCT / US 2016 / 058344 (WO 2017 / 070632), Komor, AC, et al., "Programma ble editing of a target base in genomic DNA without double-stranded DNA cleavage ”Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editin g of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (20 17); Komor, AC, et al., “Improved base excision repair inhibition and bacteri ophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living ce lls.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 also See, the entire contents of which are incorporated herein by reference.

[0080] By way of example, base editing compositions, systems, and methods described herein include: Such an adenine base editor (ABE) encodes the nucleic acid sequence (8877 base pairs) provided below. , (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681) :464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct ;36(9):843-846. doi: 10.1038 / nbt.4172.) ABE nucleic acid sequence with at least 95% or more Also included are polynucleotide sequences having identity to the sequence: ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGG GACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGG GCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAA ATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGAGGTC TATATAAGCAGAGCTGGTTTAGTGAACCGTCAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAG CCGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAAGTCTCTGAAGTCGAGTT AGCCACGAGTATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCACACGCAGAGA TCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCCTGTATGTGACACTGGAGCCA TGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGTGTTCGGAGCACGGGACGCCAAGACCGGCGC AGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATGAACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACG AGTGCGCCGCCCTGCTGAGCGATTTCTTTAGAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACC GACTCTGGAGGATCTAGCGGAGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGG CGGCTCCTCCGGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAGCC ATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACT GATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCCACTCTAGGATCGGCCGCG TGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGGCTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCAC CGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGT GTTCAATGCTCAGAAGAAGGCCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTG GCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTGGGCAA CACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACAGCCGAGGCCACCCGGC TGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATG GCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCC CATCTTCGGCAACATCGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGG ACAGCACCGACAAGGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATC GAGGGCGACCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGACGGCTGGAAA ATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCCTGAGCCTGGGCCTGACCCCC AACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACGACCTGGACAA CCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCG ACATCCTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAG GACCTGACCCTGCTGAAAGCTCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAA CGGCTACGCCGGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATC CCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAGGACAACCG GGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCCAGGGGAAACAGCAGATTCGCCT GGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGC TTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTA CTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGC AGAAAAAGGCCATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATACCACGATCT GCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGAAGATATCGTGCTGACCCTGA CACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCCCACCTGTTCGACGACAAAGTGATGAAGCAG CTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAA GACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCT TTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGC CCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGAATGAAGCGGA TCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAACACCAGCTGCAGAACGAAG CTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGAACTGGACATCAACCGGCTGTCCGACTACGA TGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACC GGGGCAAGAGCGACAACGTGCCCTCCGAAGAGGTCGTGAAGAAGATGAAGAACTACTGCGGCAGCTGCTGAACGCCAAG CTGATTACCCAGAGAAAGTTCGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCAT CAAGAGACAGCTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCGGAAGGATTTC CAGTTTTACAAAGTGCGCGAGATCCAACAACTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACGCCCT GATCAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCA AGAGCGGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATT ACCCTGGCCAACGGCGAGATCCGGAAGCGGCCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGG CCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCG GCTTCAGCAAAGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGTCCAAGAAACT GAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAAGAGCAGCTTCGAGAAGAATCCCATCGACTTTCTGGAAG CCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCGG AAGAGAATGCTGGCCTTGCCGGCGAACTGCAGAAGGGAACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTA CCTGGCCAGCACTATGAGAAGCTGAAGGGCTCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGC ACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTG CTGTCCGCCTACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGG ACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAGGTGACTCTGGC GGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAACCGGTCATCATCACCATCAC CATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCC TTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGT GTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGAT GCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGAAGCATAAAGT GTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAAC CTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCTCTTCCGCTTCCTC GCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGAGCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCA CAGAATCAGGGGATAACGCAGGAAAGAACATGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTG CTGGCGTTTTTCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGAC AGGACTATAAAGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTC GTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTATCCGGTAACTATCGTCTTGA GTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTAGGCG GTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAG CCAGTTACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTG CAAGCAGCAGATTACGCGCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGA ACGAAAACTCACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGGCACCTATCTC AGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTACGATACGGGAGGGCTTACCAT CTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCAGATTTATCAGCAATAAACCAGCCAGCCGGA AGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCTCCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAG TAGTTCGCCAGTTAATAGTTTGCGCAACGTTGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGG CTTCATTCAGCTCCGGTTCCCAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTC GGTCCTCCGATCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGTATGCGGCGAC CGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAAAAGTGCTCATCATTGGAAAA CGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAGTTCGATGTAACCCACTCGTGCACCCAACTG ATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGAGCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAA GGGCGACACGGAAATGTTGAATACTCATACTCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATG AGCGGATACATATTTGAATGTATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGA CGTCGACGGATCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAACAAGGCAAGGC TTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGATGTACGGGCCAGATATACGCG TTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCG TTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTT CCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACA TCAAGTGTATC

[0081] "Base editing activity" refers to the ability to chemically modify bases within a polynucleotide. In one embodiment, the first base is converted to the second base. In DNA editing, adenosine or adenine deaminase activity, e.g., converting A·T to G·C, Base editing activity is also referred to as adenosine or adenine deaminase activity. activity, e.g., A·T to G·C conversion activity, and cytidine deaminase activity, e.g., targeting C In some embodiments, the base editing activity may involve the activity of converting a .GAMMA. to a T.A. The base editing efficiency can be assessed by any suitable means, for example, by measuring the base editing efficiency of the target gene. In some embodiments, the nucleotide sequence can be determined by next generation sequencing or next generation sequencing. Base editing efficiency is the total sequencing reads including nucleobase conversions achieved by the base editor. by the proportion of total sequencing reads containing the target AT base pair converted to a GC base pair, e.g. In some embodiments, the base editing efficiency is measured by determining the number of base editing sites in the cell population. The percentage of total cells containing nucleobase conversions achieved by base editors was It is measured by

[0082] The term "base editor system" refers to a system that uses nucleic acid bases to edit a target nucleotide sequence. In various embodiments, the base editor system comprises: (1) a polynucleotide sequence; (1) a nucleotide-programmable nucleotide-binding domain (e.g., Cas9); and (2) the nucleobase a deaminase domain (e.g., adenosine deaminase) for deaminating 3) one or more guide polynucleotides (e.g., guide RNAs). In this embodiment, the polynucleotide-programmable nucleotide binding domain is In some embodiments, the base editor is a programmable DNA binding domain. , adenine or adenosine base editors (ABEs). The base editor system is ABE8.

[0083] In some embodiments, the base editor system comprises more than one base editing component. For example, the base editor system may include more than one deaminase. In some embodiments, the base editor system may comprise one or more adenosine deoxyribonucleic acid (ADC) bases. In some embodiments, the single guide polynucleotide may comprise an aminase. may be utilized to target different deaminases to a target nucleic acid sequence. In some embodiments, a single pair of guide polynucleotides is utilized to target nucleic acid sequences with different The deaminase may be targeted.

[0084] Deaminase domains of base editor systems and polynucleotide programmability The functional nucleotide binding entities may be covalently or non-covalently associated with each other. or may be associated by any combination of such associations and interactions. For example, in some embodiments, the deaminase domain is a polynucleotide protease. The target nucleotide sequence is targeted by a programmable nucleotide binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain The domain may be fused or linked to the deaminase domain. In one embodiment, the polynucleotide programmable nucleotide binding domain is a deaminase The target nucleotide sequence is targeted by non-covalently interacting with or associating with the For example, in some embodiments, the deaminase domain may be targeted to the The amino acid sequence of the amino acid sequence is a part of the polynucleotide-programmable nucleotide-binding domain. interacting with, associating with, or complexing with an additional heterologous moiety or domain, In some embodiments, the polypeptide may include additional heterologous moieties or domains capable of forming In the present invention, the additional heterologous moiety binds to, interacts with, associates with, or forms a complex with the polypeptide. In some embodiments, the additional heterologous moiety may be a polynucleotide. It may be capable of binding to, interacting with, associating with, or forming a complex with a protease. In this embodiment, the additional heterologous moiety may be capable of binding to the guide polynucleotide. In some embodiments, additional heterologous moieties may be attached to the polypeptide linker. In some embodiments, the additional heterologous moiety can be attached to the polynucleotide linker. The additional heterologous moiety may be a protein domain. In this study, additional heterologous moieties were identified, including the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein domain. coat protein domain, SfMu Com coat protein domain, sterile alpha motif fu, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and and Sm7 protein, or an RNA recognition motif.

[0085] The base editor system can further include a guide polynucleotide component. The components of the base editor system may be linked by covalent bonds, non-covalent interactions, or both. It is understood that the molecules may be associated with one another through any combination of associations and interactions. In some embodiments, the deaminase domain should The target nucleotide sequence can be targeted by a sequence. In embodiments, the deaminase domain is a portion or segment of a guide polynucleotide. (e.g., polynucleotide motifs) that can interact with, associate with, or form complexes with an additional heterologous moiety or domain (e.g., a polynucleotide such as an RNA or DNA binding protein) In some embodiments, additional heterologous moieties or A domain (e.g., a polynucleotide-binding domain such as an RNA or DNA-binding protein) In some embodiments, the additional heterologous deaminase domain may be fused or linked to the deaminase domain. The moiety binds to, interacts with, associates with, or forms a complex with the polypeptide. In some embodiments, the additional heterologous moiety may be a polynucleotide. capable of binding to, interacting with, associating with, or forming a complex with a phosphodiesterase In some embodiments, the additional heterologous moiety is attached to the guide polynucleotide. In some embodiments, the additional heterologous moiety is attached to the polypeptide linker. In some embodiments, the additional heterologous moiety may be a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety is a K homology (KH) domain, an MS2 coat protein Protein domain, PP7 coat protein domain, SfMu Com coat protein domain, Ste Lyle alpha motif, telomerase Ku binding motif and Ku protein, telomerase The motif may be an Sm7 binding motif and an Sm7 protein, or an RNA recognition motif.

[0086] In some embodiments, the base editor system comprises a base excision repair (BER) component. The components of the base editor system may further comprise a covalently bound, non- Related to each other through shared interactions or any combination of their associations and interactions It should be understood that the BER component inhibitor may be a BER inhibitor. In some embodiments, the inhibitor of BER may comprise a uracil DNA glycosylase. In some embodiments, the inhibitor of BER may be a boar enzyme inhibitor (UGI). In some embodiments, the inhibitor of BER may be a polynucleotide BER inhibitor. The nucleotide-programmable nucleotide-binding domain targets the target nucleotide sequence. In some embodiments, the polynucleotide programmable The functional nucleotide binding domain may be fused or linked to the BER inhibitor. In some embodiments, the polynucleotide-programmable nucleotide binding domain is It may be fused or linked to an aminase domain and an inhibitor of BER. In this embodiment, the polynucleotide-programmable nucleotide binding domain is Nucleotide sequences that target inhibitors of BER by non-covalent interaction or association with the inhibitor. For example, in some embodiments, the inhibitors of BER components may be targeted to The molecule contains additional heterologous sequences that are part of the polynucleotide-programmable nucleotide binding domain. A molecule capable of interacting with, associating with, or forming a complex with a species moiety or domain. Possibly additional heterologous moieties or domains may be included.

[0087] In some embodiments, the inhibitor of BER is targeted by a guide polynucleotide. For example, in some embodiments, the nucleotide sequence of the BER gene may be targeted to a nucleotide sequence. The inhibitor may be a portion or segment of the guide polynucleotide (e.g., a polynucleotide capable of interacting with, associating with, or forming a complex with (a) a nucleotide sequence (e.g., a nucleotide sequence motif) Additional heterologous moieties or domains (e.g., polynucleotides such as RNA or DNA binding proteins) In some embodiments, the guide polynucleotide may comprise a nucleotide-binding domain. Additional heterologous moieties or domains (e.g., polynucleotides such as RNA or DNA binding proteins) The ATP-binding domain may be fused or linked to an inhibitor of BER. The additional heterologous moiety may bind to, interact with, associate with, or or may be capable of forming a complex. In some embodiments, the additional heterologous moiety The molecule may be capable of binding to the guide polynucleotide. In some embodiments, an additional The heterologous moiety may be capable of being attached to a polypeptide linker. Additional heterologous moieties may be capable of being attached to the polynucleotide linker. In some embodiments, the additional heterologous moiety may be a K-phase protein domain. (KH) domain, MS2 coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif - and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or RN A recognition motif.

[0088] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Cas9 Active, inactive, or partially active DNA cleavage domains of Cas9 and / or gRNA binding of Cas9 Cas9 nuclease refers to an RNA-guided nuclease containing a nucleotide sequence (a protein containing a nucleotide sequence). , Casnl nuclease or CRISPR (clustered regularly interspaced short palindromic CRISPR is a mobile genetic element (U)-associated nuclease. the adaptive immune system, which provides protection against viruses, transposable elements, and conjugative plasmids A CRISPR cluster consists of a spacer, a sequence complementary to the preceding mobile element, and a target The CRISPR cluster contains the target invasive nucleic acid. It is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA is essential for transfection. Coding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein tracrRNA guides the RNase 3-assisted processing of pre-crRNA. Then, Cas9 / crRNA / tracrRNA assembles a linear or circular dsD that is complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved with an endonuclease. It is cleaved nucleolytically and then exonucleolytically trimmed to 3'-5'. In the mitochondrial world, both proteins and RNAs are required for DNA binding and cleavage. However, crRNA and tRNA are Single guide RNAs ("sgRNAs") are used to incorporate both aspects of racrRNA into a single RNA species. or simply "gRNA") can be engineered. See, e.g., Jinek M., Chylinski K., Fonfa See ra I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012) Cas9 is a CRISPR repeat sequence. It recognizes a short motif (PAM or protospacer adjacent motif) in the The sequence and structure of Cas9 nuclease are well known to those skilled in the art. known (e.g., "Complete genome sequence of an M1 strain of Streptococcus p yogenes.”Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qia n Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001 ); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. ” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, E ckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and “A progr ammable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jine k M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 33 7:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 orthologs include, but are not limited to, S. pyogenes and S. thermophilus It has been described in a variety of species. Additional suitable Cas9 nucleases and sequences are described in the present disclosure. Such Cas9 nucleases and sequences will be apparent to those skilled in the art based on the disclosures of Chylinsk i, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas Organisms and genes disclosed in “immunity systems” (2013) RNA Biology 10:5, 726-737 The Cas9 sequences from the locus are included, the entire contents of which are incorporated herein by reference.

[0089] An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), whose amino acid sequence is is provided below. JPEG2025037975000006.jpg168162 (single underline: HNH domain; double underline: RuvC domain)

[0090] Nuclease-inactivated Cas9 proteins are interchangeably referred to as "dCas9" proteins (nuclease- It may also be referred to as "dead" Cas9 or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) having the same structure are known (e.g., Jinek et al., J. Med. Soc., 1999). et al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Gui ded Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference. For example, the DNA cleavage domain of Cas9 consists of the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain is complementary to the gRNA. The RuvC1 subdomain cleaves the non-complementary strand. Mutations can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A inhibit the nuclease activity of S. completely inactivates the nuclease activity of C. pyogenes Cas9 (Jinek et al., Science. 337: 816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, In the present study, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., That is, Cas9 is a nickase called the "nCas9" protein (for "nickase" Cas9). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, In some embodiments, the protein comprises two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) one of the DNA cleavage domains of Cas9. In some embodiments, Cas9 or Proteins containing the fragments are referred to as "Cas9 variants." or a fragment thereof. For example, a Cas9 variant may share homology with a wild-type Cas9, At least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical 1. At least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 9 9% identical, at least about 99.5% identical, or at least about 99.9% identical. In terms of morphology, Cas9 variants have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 , 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30 , 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 Or it may have more amino acid changes. In some embodiments, the Cas9 variant is a fragment of Cas9 (e.g., a gRNA binding domain or is a DNA cleavage domain), and the fragment is at least about 70% identical to the corresponding fragment of wild-type Cas9. 1. At least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 9 6% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least In some embodiments, the fragment is at least about 99.5% identical, or at least about 99.9% identical. , at least about 30%, at least about 35%, or at least about 30% of the amino acid length of the corresponding wild-type Cas9. at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85% %, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least At least about 98%, at least about 99%, or at least about 99.5%.

[0091] In some embodiments, the fragment is at least 100 amino acids in length. In embodiments, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 5 50, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids.

[0092] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAGAAGTTTTAGATGCCACTCTTATCCATCCATCCATGGTCTTTGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025037975000007.jpg166162(Picture:HNHドメイン;Picture:RuvCドメイン)

[0093] In some embodiments, the wild-type Cas9 comprises the following nucleotides and / or amino acids: Corresponding to or containing sequences: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAACCTGATCGCACAATTACCCGGAGAGAAGAAAAAATGGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGGTTCGCCAGCCATCAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025037975000008.jpg167164 (single underline: HNH domain; double underline: RuvC domain)

[0094] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_002737.2 (nucleotide sequence is as follows) and Uniprot reference sequence: Sequence: Q99ZW2 (amino acid sequence is as follows): ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025037975000009.jpg166162 (SEQ ID NO: 1. Single underline: HNH domain; double underline: RuvC domain)

[0095] In some embodiments, Cas9 is derived from Corynebacterium ulcerans (NCBI reference: NC_015683. 1, NC_017317.1); Corynebacterium diphtheria (NCBI reference: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI reference: NC_021284.1); Prevotella intermedia (NCBI Reference: NC_017861.1); Spiroplasma taiwanense (NCBI reference: NC_021846.1); Streptococc us iniae (NCBI reference: NC_021314.1); Belliella baltica (NCBI reference: NC_018010.1); Psy chroflexus torquisI (NCBI reference: NC_018721.1); Streptococcus thermophilus (NCBI reference: NC_018721.1); Reference: YP_820832.1), Listeria innocua (NCBI reference: NP_472073.1), Campylobacter jejuni (NCBI Reference: YP_002344900.1) or Neisseria meningitidis (NCBI Reference: YP_002342 100.1), or any other organism-derived Cas9.

[0096] In some embodiments, the Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or In some embodiments, NmeCas9 has specificity for the NNNNGAYW PAM. wherein Y is C or T and W is A or T. In some embodiments, NmeCas9 has the formula: In some embodiments, NmeC has specificity for the NNNGYTT PAM, where Y is C or T. As9 has specificity for the NNNNGTCT PAM. In some embodiments, NmeCas9 is In some embodiments, NmeCas9 has the amino acid sequence NNNNGATT PAM, NNNNCCTA PAM, NNNNCC TC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNN NCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PAM, or NNNGATT In some embodiments, Nme1Cas9 has specificity for the PAM, NNNNGATT PAM, NN Specific for NNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, or NNNNCCTG PAM In some embodiments, NmeCas9 has specificity for CAA PAM, CAAA PAM, or CCA PAM. In some embodiments, the NmeCas9 is Nme2 Cas9. In this study, NmeCas9 has specificity for the NNNNCC (N4CC) PAM, where N is A, G, C, or T. In some embodiments, NmeCas9 is one of NNNNCCGT PAM, NNNNCCGGPAM, N NNNCCCA PAM, NNNCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PA In some embodiments, NmeCas9 has specificity for NmeM, NNNGATT, or NNNGATT PAM. In some embodiments, the NmeCas9 is a NNNNCAAA PAM, a NNNNCC PAM, or NNNNCNNN PAM. In some embodiments, Nme1, Nme2, or Nme3 The PAM-interacting domains of NmeC are N4GAT, N4CC, and N4CAAA, respectively. The characteristics and PAM sequence of as9 are described in Edraki et al., A Compact, High-Accuracy Cas9 with a Di nucleotide PAM for In Vivo Genome Editing, Mol. Cell. (2019) 73(4): 714-726. No. 6,239,999, which is incorporated herein by reference in its entirety.

[0097] An exemplary Neisseria meningitidis Cas9 protein, Nme1Cas9, (NCBI reference: WP_002235 162.1; Type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: : 1 maafkpnpin yilgldigia svgwamveid edenpiclid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlspelqd eigtafslfk tdeditgrlk driqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlgrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy ​​eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkernlndt ryvnrflcqf vadrmrltgk gkkrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgevlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 qghmetvksa krldegvsvl rvpltqlklk dlekmvnrer epklyealka rleahkddpa 901 kafaepfyky dkagnrtqqv kavrveqvqk tgvwvrnhng iadnatmvrv dvfekgdkyy 961 lvpiyswqva kgilpdravv qgkdeedwql iddsfnfkfs lhpndlvevi tkkarmfgyf 1021 aschrgtgni nirihdldhk igkngilegi gvktalsfqk yqidelgkei rpcrlkkrpp 1081vr

[0098] Another exemplary Neisseria meningitidis Cas9 protein, Nme2Cas9, (NCBI reference: WP_00 2230835; Type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: する: 1 maafkpnpin yilgldigia svgwamveid eeenpirlid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvannahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlsselqd eigtafslfk tdeditgrlk drvqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlvrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkecnlndt ryvnrflcqf vadhilltgk gkrrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgkvlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 ahkdtlrsak rfvkhnekis vkrvwlteik ladlenmvny kngreielye alkarleayg 901 gnakqafdpk dnpfykkggq lvkavrvekt qesgvllnkk naytiadngd mvrvdvfckv 961 dkkgknqyfi vpiyawqvae nilpdidckg yriddsytfc fslhkydlia fqkdekskve 1021 fayyincdss ngrfylawhd kgskeqqfri stqnlvliqk yqvnelgkei rpcrlkkrpp 1081vr

[0099] In some embodiments, dCas9 is provided with one or more abrupt changes that inactivate Cas9 nuclease activity. The Cas9 amino acid sequence may correspond in part or in whole to a mutated Cas9 amino acid sequence, or For example, in some embodiments, the dCas9 domain comprises, in part or in whole, a sequence The clone contains the D10A and H840A mutations, or the corresponding mutations in another Cas9. In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG2025037975000010.jpg169163 (single underline: HNH domain; double underline: RuvC domain)

[0100] In some embodiments, the Cas9 domain comprises a D10A mutation, and is located at position 840, or The residues at the corresponding positions in any of the amino acid sequences provided herein are The histidine remains.

[0101] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided. This results in, for example, nuclease-inactivated Cas9 (dCas9). For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., the HNH nuclease subdomain and / or the RuvC1 subdomain) In some embodiments, at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical, dC In some embodiments, variants or homologs of as9 are provided. 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids , about 50 amino acids, about 75 amino acids, about 100 amino acids or more, short or long amino acids Variants of dCas9 having the amino acid sequence are provided.

[0102] In some embodiments, the Cas9 fusion proteins as provided herein comprise a Cas9 protein. The full-length amino acid sequence of a protein, such as one of the Cas9 sequences provided herein, includes, but is not limited to, the full-length amino acid sequence of a protein, such as one of the Cas9 sequences provided herein. In other embodiments, the fusion proteins as provided herein do not include the full-length Cas9 sequence. , including only one or more fragments thereof. Exemplary amino acids of suitable Cas9 domains and Cas9 fragments are: The sequence of the Cas9 domains and fragments is provided herein, and additional suitable sequences for the Cas9 domains and fragments are within the skill of the art. It should be obvious.

[0103] Additional Cas9 proteins (e.g., nuclease-inactive Cas9 (dCas9), Cas9 nickase (n Cas9), or nuclease-active Cas9), including its variants and homologs, is used herein. It should be appreciated that within the scope of the disclosure, exemplary Cas9 proteins include, but are not limited to, In some embodiments, the Cas9 protein is In some embodiments, the Cas9 protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is a nuclease. It is enzyme-active Cas9.

[0104] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0105] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0106] An exemplary catalytically active Cas9 amino acid sequence is as follows: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0107] In some embodiments, Cas9 is used in the archaic microorganisms that comprise the domain and kingdom of unicellular prokaryotic microorganisms. It refers to Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, Cas9 is derived from, e.g., , Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res . 2017 Feb 21. doi: 10.1038 / cr.2017.21 refers to CasX or CasY described in the document. The entire contents of this publication are incorporated herein by reference. Many CRISPR-Cas systems, including Cas9, which was first reported in the archaeal domain of life, have been This diverse Cas9 protein is a poorly studied nanoarchitecture. It was discovered as part of an active CRISPR-Cas system in chia. Two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, which Among the most compact systems ever discovered. In some embodiments In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. Nucleic acid programmable DNA binding protein (napDNAbp) Other RNA-guided DNA binding proteins may also be used and are within the scope of the present disclosure. It should be understood.

[0108] In some embodiments, the Cas9 is a Cas9 variant with specificity for the engineered PAM sequence. In some embodiments, the additional Cas9 variant and PAM sequence are described in Miller et al. al., Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat Bi otechnol (2020). doi.org / 10.1038 / s41587-020-0412-8, the full contents of which are available by reference. In some embodiments, the Cas9 variants are designed to bind specific PAM elements. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, It has specificity for NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant has the PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, or In some embodiments, the SpCas9 variant has specificity for CAC. 1114, 1134, 1135, 1137, 1139, 1151, 1180, as numbered relative to the reference sequence 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 13 Amino acid substitutions at positions 23, 1332, 1333, 1335, 1337, or 1339, or the corresponding positions include. JPEG2025037975000011.jpg166162 (single underline: HNH domain; double underline: RuvC domain)

[0109] In some embodiments, the SpCas9 variants are numbered relative to the reference sequence. Like, 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or an amino acid substitution at position 1337, or the corresponding position. , SpCas9 variants are 1114, 1134, 1135 as numbered relative to the above reference sequence. , 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, Contains an amino acid substitution at positions 1320, 1323, 1333, or the corresponding positions. In some embodiments, the SpCas9 variants are numbered as follows: 1114, 1131, 1132, 1133, 1134, 1135, 1136, 1137, 1138, 1139, 1140, 1141, 1142, 1143, 1144, 1145, 1146, 1147, 1148, 1149, 1150, 1151, 1152, 1153, 1154, 1155, 1156, 1157, 1158, 1159, 1160, 1161, 1162, 1163, 1164, 1165, 1166, 1167, 1168, 116 , 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, Contains amino acid substitutions at positions 1320, 1321, 1332, 1335, 1339, or the corresponding positions. In some embodiments, the SpCas9 variants are numbered relative to the reference sequence. , 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, Contains an amino acid substitution at position 1349. Exemplary Amino Acid Substitutions and PAM Specificities of SpCas9 Variants The properties are shown in Tables A to D and FIG.

[0110] Table A JPEG2025037975000012.jpg140162

[0111] Table B JPEG2025037975000013.jpg124162

[0112] Table C JPEG2025037975000014.jpg97166

[0113] Table D JPEG2025037975000015.jpg83162

[0114] In certain embodiments, napDNAbps useful in the methods of the present invention include those known in the art. , for example, the circular permutation described by Oakes et al., Cell 176, 254-267, 2019. Exemplary circular permutations are as follows, with bold indicating Cas9-derived sequences and italicized sequences: indicates the linker sequence, and the underlined sequence indicates the bipartite nuclear localization sequence. CP5 (MSP; NGC = Pam variant containing the mutation; normal Cas9 prefers NGG; PID = protein (including the protein-interacting domain and "D10A" nickase): JPEG2025037975000016.jpg169162

[0115] Polynucleotide programmable nucleotide binding domains that can be incorporated into base editors Main, non-limiting examples include domains from CRISPR proteins, restriction nucleases, megakaryons, and Nucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs) ) is included.

[0116] In some embodiments, the nucleic acid sequences of any of the fusion proteins provided herein The tunable DNA binding protein (napDNAbp) may be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. The napDNAbp is a CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX protein. or CasY protein, at least 85%, at least 90%, at least 91%, at least at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least or at least 97%, at least 98%, at least 99%, or at least 99.5% identical In some embodiments, the napDNAbp comprises a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a CasX or CasY protein described herein. At least 85%, at least 90%, at least 91%, at least 92%, or at least at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, contain an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical Cas12b / C2c1, CasX, and CasY from other bacterial species may also be used in accordance with the present disclosure. We should recognize this.

[0117] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS= Alicyclobacillus acido - terrestris (ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B stock) GN=c2c1 PE =1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI

[0118] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR associated Casx protein OS = Sulfolobus islandicus (HVE1 0 / 4 strain) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG

[0119] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandicus (REY15A strain) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0120] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAK PLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPK KPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMD EKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTD GTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIG RDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQA AKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKL AYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELS AELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYK SGKQPFVGAWQAFYKRRLKEVWKPNA

[0121] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR conjugation protein CasY [uncultured Parcubacteria bacterium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0122] The term "conservative amino acid substitution" or "conservative mutation" refers to a mutation in which an amino acid is shared by two or more amino acids. It refers to the substitution of an amino acid with another amino acid that has the same properties as the amino acid. A functional method for defining amino acid changes between corresponding proteins of homologous organisms is The analysis of normalized frequencies (Schulz, GE and Schirmer, RH, Principles Es of Protein Structure, Springer-Verlag, New York (1979). For example, amino acids within a group are preferentially exchanged with each other, thus affecting the overall protein structure. Define groups of amino acids that are most similar to each other in their effect on structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include: Examples include, for example, the amino acids lysine and arginine, which are capable of maintaining a positive charge. The opposite is true: glutamic acid and aspartic acid can maintain a negative charge. Conversely, serine for threonine, which can maintain a free OH; and serine for threonine, which can maintain a free NH Examples of amino acid substitutions that can be made include glutamine for asparagine.

[0123] The terms "coding sequence" or "protein-coding sequence" are used interchangeably herein. " refers to a segment of a polynucleotide that encodes a protein. The sequence is bounded by a start codon near the 5' end and a stop codon near the 3' end. A coding sequence may also be referred to as an open reading frame.

[0124] As used herein, the terms "deaminase" or "deaminase domain" and "deaminase domain" refer to refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, Deaminase is an adenine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine or Adenosine deamination catalyzes the hydrolytic deamination of adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is adenosine deaminase, which converts adenosine to inosine or deoxyinosine, respectively In some embodiments, the ATP catalyzes the hydrolytic deamination of adenosine or deoxyadenosine. Adenosine deaminase hydrolyzes adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase enzymes provided herein (e.g., those derived from the gene Engineered adenosine deaminase (evolved adenosine deaminase) is a type of enzyme that is used in bacteria, In some embodiments, the adenosine deaminase can be derived from any organism. cherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens , Haemophilus influenzae, or Caulobacter crescentus.

[0125] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. The amino acid sequence of the ... or variants of naturally occurring deaminases from organisms such as mice. In embodiments, the deaminase or deaminase domain is not naturally occurring. For example, In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase. At least 50%, at least 55%, at least 60%, at least 65%, at least At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 1%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, At least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, At least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99 0.7%, at least 99.8%, or at least 99.9% identical. For example, the deaminase domain The invention is disclosed in International PCT Application No. PCT / US2016 / 0583 (WO 2018 / 027078) and PCT / US2016 / 0583. 44 (WO 2017 / 070632), the entire contents of each of which are incorporated herein by reference. Komor, AC, et al., “Programmable editing of a target base n genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “I improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Adva nces 3:eaao4774 (2017) ), and Rees, HA, et al., “Base editing: precision ch emistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. See also doi: 10.1038 / s41576-018-0059-1. The contents of which are incorporated herein by reference.

[0126] "Detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In this embodiment, sequence variations in a polynucleotide or polypeptide are detected. In another embodiment, the presence of indels is detected.

[0127] A "detectable label" is a label that, when attached to a molecule of interest, is capable of detecting a given molecule spectroscopically, photochemically, means a composition that renders the latter detectable through biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, and colloidal particles. , fluorescent dyes, electron-dense reagents, enzymes (e.g., those commonly used in ELISA), bio These include thiazol-1, ...

[0128] "Disease" means any condition that damages or interferes with the normal function of a cell, tissue, or organ. or disorder. In one embodiment, the disease is A1AD.

[0129] The term "effective amount" as used herein means an amount sufficient to elicit a desired biological response. In certain embodiments, an effective amount refers to the amount of biologically active agent present in a patient to achieve a therapeutic effect. and a base editor system (e.g., a nucleotide sequence encoding ... Fusion proteins containing programmable DNA-binding proteins, nucleobase editors and gRNs A) amount of A. Such therapeutic effects are achieved by modifying A1AD in all cells of the tissue or organ. It need not be sufficient to transform, for example, about 1%, 5%, or even more of the cells present in a subject, tissue, or organ. , 10%, 25%, 50%, 75% or more may be modified. An effective amount is sufficient to alleviate one or more symptoms of A1AD. The effective amount of active agent(s) used for this purpose will vary depending on the mode of administration, the age, weight, and general condition of the subject. Ultimately, your doctor or veterinarian will determine the appropriate dosage and administration. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is sufficient to introduce a modification into a gene of interest in a cell (e.g., in vitro or in vivo). The base editors of the invention (e.g., fusion proteins comprising programmable DNA binding proteins) In one embodiment, the effective amount is the amount of a therapeutic agent (protein, nucleobase editor, and gRNA). achieve an effect (e.g., reduce or control a disease or its symptoms or conditions) ) is the amount of base editor required.

[0130] "Fragment" means a portion of a polypeptide or nucleic acid molecule, which portion is identical to a reference nucleic acid At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the entire length of the molecule or polypeptide , or 90%. The fragments may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 20 Contains 0, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids It is possible.

[0131] A "guide RNA" or "gRNA" is a polynucleotide program specific for a target sequence. Forms a complex with a gram-capable nucleotide-binding domain protein (e.g., Cas9 or Cpf1) In one embodiment, a guide polynucleotide refers to a polynucleotide that can The first gene is a guide RNA (gRNA). gRNAs may exist as a complex of two or more RNAs. For example, gRNAs may exist as a single RNA molecule. Although it is sometimes called a single guide RNA (sgRNA), as a single molecule, "gRNA" or is used interchangeably to refer to guide RNAs that exist as a complex of two or more molecules Typically, gRNAs exist as a single RNA species that (1) share homology with the target nucleic acid; (2) a domain (e.g., that directs binding of the Cas9 complex to a target); and In some embodiments, the domain (2 ) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. In some embodiments, domain (2) is selected from the group consisting of the amino acid sequence of Jinek et al., Science 337:816-821 (2012) (and The entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) include the "Switchable Cas9 Nucleotide" and "Switchable Cas9 Nucleotide" gRNAs. U.S. Provisional Patent Application USSN 61 / 01 / 2013, filed September 6, 2013, entitled "Cleases and Uses Thereof" 874,682, and the September 2013 issue of "Delivery System For Functional Nucleases" and U.S. Provisional Patent Application No. 61 / 874,746, filed on the same day, the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises domains (1) and (2). and (2) may be referred to as an "extended gRNA." An extended gRNA is a As described herein, the target nucleic acid may be bound to two or more Cas9 proteins and targeted at two or more distinct regions. The gRNA contains a nucleotide sequence that is complementary to the target site, which binds to the target nucleic acid. It mediates the binding of the nuclease / RNA complex to the target site and determines the sequence specificity of the nuclease:RNA complex. As will be appreciated by those skilled in the art, RNA polynucleotide sequences, such as gRNA, The A sequence contains more pyrimidines than thymine (T), a nucleic acid base contained in a DNA polynucleotide sequence. The nucleic acid base uracil (U) is a derivative of adenine. It base pairs with thymine and replaces thymine during DNA transcription.

[0132] "Hybridization" means hydrogen bonding between complementary nucleobases, as defined by Watson- Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds. Denine and thymine are complementary nucleobases that form hydrogen bonds to form pairs.

[0133] The term "inhibitor of base repair," or "IBR," refers to an inhibitor of a nucleic acid repair enzyme, such as base excision repair. (BER) refers to a protein that can inhibit the activity of an enzyme. , IBR is an inhibitor of inosine base excision repair. Examples of inhibitors of base repair include A PE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG In some embodiments, the IBR is an inhibitor of En. In some embodiments, the IBR is an inhibitor of catalytically inactive E In some embodiments, the base repair inhibitor is ndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. , catalytically inactive EndoV or catalytically inactive hAAG.

[0134] In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UG I) UGI is a DNA repair enzyme that inhibits uracil-DNA glycosylase base excision repair. In some embodiments, the UGI domain refers to a protein that can encode wild-type UGI or a wild-type UGI. In some embodiments, the UGI proteins provided herein include fragments of native UGI. , fragments of UGI, and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In its form, the base repair inhibitor acts as a "catalytically inactive inosine-specific nuclease." or "inactive inosine-specific nuclease." Without being bound by any particular theory, However, catalytically inactive inosine glycosylases (e.g., alkylated inosine glycosylases) are not desirable. Denning glycosylase (AAG) can bind to inosine but does not create an abasic site. Neither the inosine can be removed, thereby preventing the newly formed inosine moiety. In some embodiments, the catalytically inactive molecule sterically blocks the molecule from DNA damage / repair machinery. An active inosine-specific nuclease can bind to inosine in nucleic acids, but not to the nucleus. Non-limiting exemplary catalytically inactive inosine-specific nucleases do not cleave the acid. For example, catalytically inactive alkyl adenosine glycosylase (AAG nuclease) from human and catalytically inactive endonuclease V (EndoV nuclease) from, for example, E. coli. In some embodiments, the catalytically inactive AAG nuclease is an E1 25Q mutation or a corresponding mutation in another AAG nuclease.

[0135] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0136] The intein excises itself and the remaining fragment (extein) tein) are linked by peptide bonds in a process known as protein splicing. Inteins are fragments of proteins that can bind to other proteins. The intein excises itself and joins the rest of the protein. The process is referred to herein as "protein splicing" or "intein-mediated transcription." In some embodiments, the precursor protein (intron) is spliced. Inteins (intein-containing proteins before intein-mediated protein splicing) Such an intein is referred to herein as a split intein. These are called split inteins (e.g., split intein-N and split intein-C). For example, cyano In bacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is expressed in two separate genes. The dnaE-n gene encodes the dnaE-c gene. The intein may be referred to herein as "intein N." The intein delivered may be referred to herein as "intein C."

[0137] Other intein systems can also be used, for example, the dnaE intein, i.e., Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) Synthetic inteins based on the nucleotide sequence have been described (see, e.g., ref. 1, which is incorporated herein by reference). Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Used in accordance with this disclosure. Non-limiting examples of intein pairs that can be used include: Cfa DnaE intein, Ssp GyrB Intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma Dna B intein, and Cne Prp8 intein (see, e.g., the references herein incorporated by reference). such as those described in U.S. Pat. No. 8,394,604.

[0138] Exemplary nucleotide and amino acid sequences of inteins are provided.

[0139] DnaE Intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGA ATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGG AAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAG ATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT

[0140] DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDG QMLPIDEIFERELDLMRVDNLPN

[0141] DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTG CTCTGAAGAACGGATTCATAGCTTCTAAT

[0142] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0143] Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGA ATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAG AAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAG ATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA

[0144] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP

[0145] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCT TGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC

[0146] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0147] Intein N and intein B were used to link the N-terminal part of split Cas9 with the C-terminal part of split Cas9. The tein C can be fused to the N-terminal end of the split-Cas9 and the C-terminal end of the split-Cas9, respectively. For example, In some embodiments, an intein-N is fused to the C-terminus of the N-terminal portion of the split-Cas9; That is, the structure N--[N-terminal portion of split Cas9]-[intein-N]--C is formed. In embodiments, intein-C is fused to the N-terminus of the C-terminal portion of the split-Cas9, i.e., N- The structure is formed by [intein-C]-[C-terminal part of split Cas9]-C. Intein-mediated protein splicing for joining proteins (e.g., split-Cas9) The mechanism of this binding is described, for example, in Shah et al., Chem Sci. 2002, incorporated herein by reference. 014; 5(1):446-461. Methods for designing and using such antibodies are known in the art, see, for example, WO2014004336, WO201713 2580, US20150344549 and US20180127780, each of which is The entirety of which is incorporated herein by reference.

[0148] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its native state. from which, to varying degrees, components normally associated with the product when found in its original form have been removed. "Isolated" refers to the degree of separation from the original source or surrounding environment. "Purified" refers to "Purified" or "biologically pure" proteins impurities do not materially affect the biological properties of the protein or cause other adverse consequences Other substances have been sufficiently removed so as not to cause any damage to the nucleic acid of the present invention. or peptides, if produced by recombinant DNA technology, may be derived from cellular material, viral material, or or substantially free of culture medium, or, if chemically synthesized, free of chemical precursors or Purity and homogeneity are typically determined by the Typically, analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography, are used. The term "purified" means that a nucleic acid or protein is purified to a high degree. This can mean that the protein produces essentially one band in a gel. For proteins that can undergo phosphorylation or glycosylation, the different modifications are This can result in different isolated proteins that can be purified separately.

[0149] An "isolated polynucleotide" is a polynucleotide that is not present in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. In this context, it refers to nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene. The term refers to, for example, incorporated into a vector; incorporated into an autonomously replicating plasmid or virus; embedded in the genomic DNA of prokaryotes or eukaryotes; independent of other sequences Separate molecules (e.g., cDNA generated by PCR or restriction endonuclease digestion) or genomic or cDNA fragments). RNA molecules transcribed from DNA molecules, as well as hybrids encoding additional polypeptide sequences. It contains recombinant DNA that is part of the hybrid gene.

[0150] An "isolated polypeptide" is a polypeptide of the present invention that has been separated from components that naturally accompany it. Typically, a polypeptide refers to a polypeptide that is present in a natural state. at least 60% by weight free from associated proteins and naturally occurring organic molecules; Preferably, the preparation is at least 75%, more preferably 90%, and most preferably 10% by weight of the The isolated polypeptide of the present invention is preferably 99% or more of the polypeptide of the present invention. For example, extraction from natural sources, expression of recombinant nucleic acids encoding such polypeptides, or by chemically synthesizing the protein. Purity can be achieved by any suitable method. Suitable methods, such as column chromatography, polyacrylamide gel electrophoresis, or can be measured by HPLC analysis.

[0151] As used herein, the term "linker" refers to a molecule or moiety that connects two molecules or moieties, e.g., a Two components of a protein or ribonucleocomplex, or two domains of a fusion protein A domain, e.g., a polynucleotide programmable DNA binding domain (e.g., dCas9) and a determinant adenosine deaminase domain (e.g., as described in PCT / US19 / 44935) , or a covalent linker ( It can refer to a linker, chemical group, or molecule. Carriers connect different components or parts of components of a base editor system. For example, in some embodiments, the linker may be a polynucleotide protease. Guide polynucleotide binding domain and deactivation of grammable nucleotide binding domain In some embodiments, the linker can link the catalytic domains of the amines. In some embodiments, the CRISPR polypeptide and the deaminase can be linked. In some embodiments, the linker can link the Cas9 and the deaminase. In some embodiments, the linker can link the dCas9 and the deaminase. In embodiments, the linker can link the nCas9 and the deaminase. In this embodiment, the linker connects the guide polynucleotide and the deaminase. In some embodiments, the linker can be a deaminase inhibitor of the base editor system. ligating the polynucleotide-programmable nucleotide-binding component and the polynucleotide-programmable nucleotide-binding component In some embodiments, the linker can be used to deamidate the base editor system. RNA-binding moieties of polynucleotide components and polynucleotide-programmable nucleotide-binding structures In some embodiments, the linker can link the components. The RNA-binding portion of the deamination component of the system and the polynucleotide-programmable nucleic acid The RNA-binding moieties of the nucleotide-binding components can be linked together. The linker can be composed of two groups: molecules or other moieties, located between or adjacent to them, and either covalently or non-covalently bonded They are connected to each other through bonded interactions, and can therefore connect the two. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In embodiments, the linker can be a DNA linker. In some embodiments, the linker is an R In some embodiments, the linker can be a NA linker. In some embodiments, the ligand can include an aptamer that can decompose. In some embodiments, the linker may be a molecule, a peptide, a protein, or a nucleic acid. The aptamer can include an aptamer that may be derived from a riboswitch. The riboswitches that are used are theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, Chi, adenosine cobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch switch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahybrid Folate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch The riboswitch may be selected from a GlmS riboswitch, a GlmS riboswitch, or a prequeosin 1 (PreQ1) riboswitch. In some embodiments, the linker is a polypeptide or polypeptide ligand, etc. In some embodiments, the polynucleotide may comprise an aptamer bound to a protein domain of the polypeptide. The peptide ligands are the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein domain. Protein domain, SfMu Com coat protein domain, sterile alpha motif , telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and and Sm7 protein, or an RNA recognition motif. The peptide ligand can be part of a base editor system component. For example, a nucleic acid salt The base editing component may include a deaminase domain and an RNA recognition motif.

[0152] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide). In some embodiments, the linker may be about 5 to 100 nm in length. amino acids, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 , 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90 or 90-100 amino In some embodiments, the linker may be about 100-150, 150-200, 200 Can be ~250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids Longer or shorter linkers are also contemplated.

[0153] In some embodiments, the linker comprises an RNA program containing a Cas9 nuclease domain. The gRNA-binding domain of a transducible nuclease and a nucleic acid-editing protein (e.g., adenosine In some embodiments, the linker links the catalytic domains of dCas9 and dCas9 deaminase. and the nucleic acid editing protein. For example, the linker may be a linker that connects two groups, molecules, or other are located between or adjacent to the moieties and are connected to each other via covalent bonds. In some embodiments, the linker is an amino acid or multiple amino acids. In some embodiments, the amino acid is a phospholipid. The car is an organic molecule, group, polymer, or chemical moiety. The linker may be 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 69, 68, 69, 7 5, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 9 0, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190 , or 200 amino acids. Longer or shorter linkers are also contemplated.

[0154] In some embodiments, the nucleobase editor domain is SGGSSGSETPGTSESATPESSG GS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTE Amino acid PSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS In some embodiments, the nucleobase editor is fused via a linker comprising the sequence The domains are linked via a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n、 (EAAAK) n , (GGS)n , SGSETPGTSESATPES, or (XP) n motif, or any combination of these n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. do.

[0155] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the phosphoryl group comprises the amino acid sequence SGGSSGGSSGSETPGTESATPESSGGSSGGSSGSSGGS. In some embodiments, the linker is 64 amino acids in length. Contains several In embodiments, the linker is 92 amino acids in length. The amino acid sequence is PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GT Includes STEPSEGSAPGTSESATPESGPGSEPATS.

[0156] A "marker" is any molecule that has an altered expression level or activity that is associated with a disease or disorder. The term "protein" refers to any protein or polynucleotide.

[0157] As used herein, the term "mutation" refers to a change in a sequence, e.g., a nucleic acid or amino acid sequence. The substitution of a residue in an amino acid sequence by another residue, or the substitution of one or more residues in the sequence Mutations, as used herein, typically refer to deletions or insertions of the original residues. Then, the position of the residue in the sequence is identified, and the identity of the newly substituted residue is determined. Thus, the methods for making the amino acid substitutions (mutations) provided herein are described. A variety of methods are well known in the art, see, for example, Green and Sambrook, Mo. lecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Pre ss, Cold Spring Harbor, NY (2012). Therefore, the base editors of the present disclosure can eliminate a significant number of unintended mutations, e.g., unintended point mutations. "intended" in a nucleic acid (e.g., a nucleic acid in a subject's genome) without generating mutations In some embodiments, mutations, such as point mutations, can be efficiently generated. , the intended mutation is specifically designed to produce the intended mutation, A specific base editor (e.g., adenosine triphosphate) bound to a guide polynucleotide (e.g., gRNA) These are mutations caused by DNA base editors.

[0158] Generally, a sequence (e.g., an amino acid sequence described herein) is made or identified. The mutation to be detected is compared to the reference (or wild-type) sequence, i.e., the sequence not containing the mutation. Those skilled in the art will recognize the differences in amino acid and nucleic acid sequences relative to a reference sequence. It will be readily apparent to those skilled in the art how to determine the location of the mutation.

[0159] The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g. For example, lysine for tryptophan or phenylalanine for serine. In this case, the non-conservative amino acid substitutions do not disrupt or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions are preferred because they do not impair the biological activity of the functional variant. The biological activity of the functional variant is enhanced so that it is increased compared to the wild-type protein. It can be done.

[0160] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to the localization of a protein. The term "nuclear localization sequence" refers to an amino acid sequence that promotes import into the cell nucleus. It is known in the art, for example, as WO / 2001 / 038547 filed on November 23, 2000 and published on May 31, 2001. Plank et al. in published international PCT application PCT / EP 2000 / 011690, the contents of which are incorporated herein by reference. is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In embodiments, the NLS may be any of the NLSs described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.417 In some embodiments, the NLS is an optimized NLS as described by Acid sequence: KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENG Contains RKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0161] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a nucleic acid molecule that is a nucleic acid having a nucleic acid base and an acid. Compounds containing a cyclic moiety, such as a nucleoside, nucleotide, or polymer of nucleotides Typically, polymeric nucleic acids, e.g., nucleic acid molecules containing three or more nucleotides, are , a linear molecule in which adjacent nucleotides are linked to each other via phosphodiester linkages In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and In some embodiments, a "nucleic acid" refers to a group of three or more amino acids. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing individual nucleotide residues. The terms "oligonucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., can be used interchangeably to refer to a sequence of at least three nucleotides. In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids include, for example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, Naturally occurring nucleic acid in the context of a mid, chromosome, chromatid, or other naturally occurring nucleic acid molecule Alternatively, the nucleic acid molecule may be present in any suitable form, e.g., recombinant DNA or RNA, artificial chromosome, engineered the genome, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or non-natural It may be a non-naturally occurring molecule that contains a naturally occurring nucleotide or nucleoside. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms refer to nucleic acid analysis. Nucleic acids include analogs, e.g., analogs having other than a phosphodiester backbone. purified from recombinant expression systems; produced using recombinant expression systems and optionally purified; chemically synthesized In the case of chemically synthesized molecules, nucleic acids may, where appropriate, be Nucleotides, such as analogs with chemically modified bases or sugars, and backbone modifications, are also used. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. In some embodiments, the nucleic acid is composed of natural nucleosides (e.g., adenosine, thymidine, Guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine adenosine, and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine , 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-amino Adenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxo Guanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically Modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose , ribose, 2'-deoxyribose, arabinose, and hexose); and / or Modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) , or including them.

[0162] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to and can be used interchangeably with "programmable nucleotide binding domain" a guide nucleic acid or guide polynucleotide (e.g., It refers to a protein that associates with a nucleic acid (e.g., DNA or RNA) such as a gRNA. In this embodiment, the polynucleotide programmable nucleotide binding domain is a polynucleotide In some embodiments, the polynucleotide is a nucleotide-programmable DNA binding domain. The polynucleotide-programmable nucleotide binding domain is In some embodiments, the polynucleotide programmable RNA binding domain The nucleotide-binding domain is the Cas9 protein. The Cas9 protein binds to the guide RNA. It can assemble with guide RNA, which guides the Cas9 protein to specific complementary DNA sequences. In some embodiments, the napDNAbp comprises a Cas9 domain, e.g., a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas 12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, C as9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl , Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy 3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2 , Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17 , Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Cs d1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effect CAR protein, Cas V effector protein, Cas VI effector protein, CAR F, DinG, their homologs, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins may also be used, even if not specifically listed in this disclosure. It is possible that this is not possible, but is within the scope of this disclosure. nd Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:3 25-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science. See aav7271, the entire contents of each of which are incorporated herein by reference.

[0163] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. refers to nitrogen-containing biological compounds that form nucleosides, and nucleosides are The ability of nucleobases to base pair and stack with each other is directly related to the It gives rise to long helical structures such as nucleic acids (RNA) and deoxyribonucleic acid (DNA). Adenine The five main nucleic acid bases are amino acid (A), cytosine (C), guanine (G), thymine (T), and uracil (U). These are called primary bases or canonical nucleobases. Guanine is derived from purine, while cytosine, uracil, and thymine are derived from pyrimidine. The RNA may also contain other (minor) bases that are modified. Non-limiting exemplary modified nucleobases These include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, and 5-methyl These include hypoxanthine and xanthine. Both amines can be produced by deamination (calcification of amine groups) in the presence of mutagens. Hypoxanthine can be produced by modification of adenine. Xanthine can be modified from guanine. Uracil is formed by deamination of cytosine. A "nucleoside" is a nucleic acid base and a pentose sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-amino-2-methyl-2-propanol ... -Methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxy Nucleosides with modified nucleobases include cytosine, cytosine, and deoxycytidine. Examples include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine Nucleotides include 5-methylcytidine (m5C), 5-methylcytidine (m5C), and pseudouridine (Ψ). " is a nucleic acid consisting of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least It also consists of one phosphate group.

[0164] The terms "nucleobase editing domain" or "nucleobase editing protein" are used herein. When used in deamination to ribonucleotides (ribonucleotides), and non-templated nucleotide addition and insertion. Proteins or enzymes capable of catalyzing nucleic acid base modifications in RNA or DNA In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., For example, adenine deaminase or adenosine deaminase. In some embodiments, the nucleobase editing domain may be linked to more than one deaminase domain (e.g., adenine Deaminase, or adenosine deaminase and cytidine or cytosine deaminase , e.g., as described in PCT / US19 / 44935). In some embodiments, the nucleobase The editing domain can be a naturally occurring nucleobase-editing domain. In this study, nucleobase-editing domains were engineered or developed from naturally occurring nucleobase-editing domains. The nucleobase-editing domain may be a modified nucleobase-editing domain. The nucleobase-editing domain may be a modified nucleobase-editing domain derived from bacteria, humans, It can be derived from any organism, such as a pansy, gorilla, monkey, cow, dog, rat, or mouse. It is possible.

[0165] As used herein, "obtaining," as in "obtaining an agent," refers to obtaining the agent. Synthesize, purchase, produce, prepare, or otherwise obtain This includes:

[0166] As used herein, a "patient" or "subject" refers to a person who has been diagnosed with a disease or disorder. have, are at risk of having or developing, or may have or develop refers to a mammalian subject or individual suspected of developing "Patient" refers to a mammalian subject who has a higher than average likelihood of developing a disease or disorder. Possible patients include humans, non-human primates, cats, dogs, pigs, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs) and and other mammals that can benefit from the therapies disclosed herein. The patient may be male and / or female.

[0167] A "patient in need thereof" or a "subject in need thereof" is used herein to refer to a person with a disease or diagnosed with, having, being at risk of having, or being susceptible to a disability or disorder , patients who have been determined to have it or are suspected of having it, or or individuals.

[0168] "Pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant" The terms "mutant," "deleterious mutation," or "predisposing mutation" refer to a specific disease or refers to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a disorder In some embodiments, the pathogenic mutation is a mutation in a protein encoded by the gene. at least one wild-type amino acid in the This includes those that have been

[0169] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and are linked together by a peptide (amide) bond The term refers to a polymer of amino acid residues bound together in a single chain. It typically refers to a protein, peptide, or polypeptide. A protein, peptide, or polypeptide is at least three amino acids in length. Polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may be conjugated, functionalized, or For functionalization or other modifications, e.g., carbohydrate groups, hydroxyl groups, phosphate groups, pharmacophores, etc. By adding chemical entities such as nesyl groups, isofarnesyl groups, fatty acid groups, linkers, etc. The protein, peptide, or polypeptide may also be a single molecule. or multi-molecular complexes. Proteins, peptides, or polypeptides The fragment may simply be a fragment of a naturally occurring protein or peptide. The peptide, or polypeptide, may be naturally occurring, recombinant, or synthetic. or any combination thereof. As used herein, the term "fusion tag" A "protein" is a hybrid that contains protein domains from at least two different proteins. A hybrid polypeptide is a polypeptide that is a fusion protein. One protein is a fusion protein that is a fusion protein. The other protein ... It can be located in the (terminal) or carboxy-terminal (C-terminal) portion of a protein, and therefore These form amino-terminal or carboxy-terminal fusion proteins, respectively. Proteins contain distinct domains, e.g., nucleic acid binding domains (e.g., The Cas9 gRNA binding domain (which guides the binding of the target gene) and the nucleic acid cleavage domain, or nucleic acid editing domain, In some embodiments, the protein may comprise a catalytic domain of a protein. The polymeric portion, e.g., the amino acid sequence constituting the nucleic acid binding domain, and the organic compound, e.g., For example, compounds that can act as nucleic acid cleaving agents. In some embodiments, the protein , complexed with or associated with a nucleic acid, such as RNA or DNA. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and The protein can be produced through purification and synthesis by fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and are particularly suitable. Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) and the entire contents of which are incorporated herein by reference.

[0170] The polypeptides and proteins disclosed herein (including functional portions thereof and their functional variants) A synthetic amino acid may be substituted for one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and can be used, for example, aminocyclo Hexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylacetone aminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenyl phenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine phenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-Naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, Aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6 -Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclo cyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornene) α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine Polypeptides and proteins include alpha-tert-butylglycine, alpha-tert-butylglycine, and alpha-tert-butylglycine. , can be associated with post-translational modification of one or more amino acids of a polypeptide construct. Non-limiting examples of post-modifications include acylation, including phosphorylation, acetylation, and formylation. , glycosylation (including N-linked and O-linked), amidation, hydroxylation, methylation and alkylation, including ethylation, ubiquitination, addition of pyrrolidone carboxylic acid, disulfide Formation of amide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation These include ionization, geranylation, glypiation, lipoylation, and iodination. can be done.

[0171] The term "recombinant" as used herein with respect to a protein or nucleic acid refers to a protein or nucleic acid that is naturally occurring. refers to proteins or nucleic acids that do not exist in the genome and are the product of human manipulation. For example, some In embodiments, the recombinant protein or nucleic acid molecule is a recombinant protein or nucleic acid molecule that is amplified compared to any naturally occurring sequence. At least one, at least two, at least three, at least four, at least five, an amino acid or nucleotide sequence containing at least six, or at least seven mutations Includes.

[0172] By "reduced" is meant a negative alteration of at least 10%, 25%, 50%, 75%, or 100%. .

[0173] "Reference" refers to a standard or control condition. In one embodiment, a reference are wild-type or healthy cells. In other embodiments, without limitation, the reference is Unexposed or placebo or saline, medium, buffer, and / or , untreated cells exposed to a control vector that does not carry the polynucleotide of interest. .

[0174] A "reference sequence" is a defined sequence used as a basis for sequence comparison. It may be a subset or the entirety of a particular sequence; for example, a full-length cDNA or gene sequence. A segment of a sequence, or the complete cDNA or gene sequence. For polypeptides, the reference polypeptide The length of the polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, It may be at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, Each of the nucleotides is about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides. Nucleotide or about these values ​​or any integer therebetween. In embodiments, the reference sequence is the wild-type sequence of the protein of interest. In the above example, the reference sequence is a polynucleotide sequence encoding a wild-type protein.

[0175] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to Used in conjunction with (e.g., binds to or associates with) one or more non-target RNAs. In some embodiments, the RNA-programmable nuclease, when complexed with RNA, Typically, the bound RNA(s) is a guide RNA ( gRNAs can exist as a complex of two or more RNAs, or as a single RNA. gRNAs that exist as a single RNA molecule can be called single-guide Although sometimes referred to as sgRNA, "gRNA" can be used as a single molecule or as two or more are used interchangeably to refer to guide RNAs that exist as either a molecular complex or Typically, gRNAs exist as a single RNA species, consisting of (1) a domain that shares homology with the target nucleic acid; (2) a target domain (e.g., directing binding of the Cas9 complex to the target); and (3) a target domain (e.g., directing binding of the Cas9 complex to the target); and In some embodiments, the domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. In some embodiments, domain (2) is a polypeptide of the invention as described in Jinek et al., Science 337:816-821 (2012) (the The entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are "Switchable Cas9 Nucleotide" and "Switchable Cas9 Nucleotide". U.S. Provisional Patent Application USSN 61, filed September 6, 2013, entitled "Eases and Uses Thereof" / 874,682 and "Delivery System For Functional Nucleases" on September 6, 2013 and US Provisional Patent Application No. USSN 61 / 874,746 filed on 2004 / 010364, each of which is incorporated herein by reference. The entire contents of each are incorporated herein by reference. In some embodiments, the gRNA is It may contain two or more of domains (1) and (2) and be referred to as an "extended gRNA." For example, The extended gRNA can then be coupled to, for example, two or more Cas9 proteins, as described herein. The gRNA binds to the target nucleic acid at two or more distinct regions. a nucleic acid sequence that mediates binding of the nuclease / RNA complex to the target site; Provides sequence specificity for the nuclease:RNA complex.

[0176] In some embodiments, the RNA programmable nuclease (CRISPR-associated system) Cas9 endonuclease, e.g., Cas9 from Streptococcus pyogenes (Casnl) (For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes.") Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Na jar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, Mc Laughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA mat uration by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chy linski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J. , Charpentier E., Nature 471:602-607(2011)).

[0177] RNA-programmable nucleases (e.g., Cas9) target DNA cleavage sites with RNA: Because DNA hybridization is used, these proteins are, in principle, Any sequence specified by the RNA can be targeted for site-specific cleavage. Methods that use RNA-programmable nucleases such as Cas9 to modify genomes (for example, Cong, L. et al., Multiplex genome sequencing) are known in the art (see, e.g., Cong, L. et al., Multiplex genome sequencing). e engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas sy stem. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Gen ome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic ac ids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes u See Nature biotechnology 31, 233-239 (2013). the entire contents of each of which are incorporated herein by reference).

[0178] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific position in the genome. Each mutation is present to a discernible extent (e.g., >1%) in the population. At certain base positions in the genome, C nucleotides can occur in most individuals, but in a few individuals In the human genome, this position is occupied by A. This means that there is a SNP at this particular position, and it can be either C or A. This means that there are two possible nucleotide variations at this position, which are alleles. SNPs are the cause of disease. Genetic variation also underlies differences in susceptibility to certain diseases. The severity of the disease and the body's response to treatment also depend on genetic variation. SNPs can occur in the coding region of a gene, the non-coding region of a gene, or in intergenic regions. In some embodiments, the region may be located within a coding sequence. SNPs do not necessarily change the amino acid sequence of the protein produced due to the degeneracy of the genetic code. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Nonsynonymous SNPs do not affect the protein sequence, but nonsynonymous SNPs change the amino acid sequence of the protein. There are two types of NPs: missense and nonsense. SNPs that are not in protein-coding regions are , gene splicing, transcription factor binding, messenger RNA degradation, or non-coding R The gene expression affected by this type of SNP can be determined by the sequence of the NA. P (expressed SNP) and can be upstream or downstream of a gene. Single nucleotide variants (SNVs) A somatic single-nucleotide mutation is a single-nucleotide mutation that occurs at any frequency and can occur in somatic cells. It may also be called modification.

[0179] "Specifically binds" means that the polypeptide and / or nucleic acid molecule of the present invention is recognized and binds to it but does not substantially recognize other molecules in the sample, e.g., biological sample, Unbound nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programs) "DNA-binding domain and guide nucleic acid"), compound, or molecule.

[0180] Nucleic acid molecules useful in the methods of the present invention may encode a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. Typically, but not necessarily, substantial identity is shown. A polynucleotide having a nucleotide sequence typically hybridizes with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention can be modified to The present invention also includes any nucleic acid molecule encoding an endogenous nucleotide sequence or a fragment thereof. It need not be 100% identical to the nucleic acid sequence, but typically will show substantial identity. Polynucleotides having "substantial identity" to a double-stranded nucleic acid molecule typically have a small number of double-stranded nucleic acid molecules. "Hybridize" means to hybridize with at least one strand. A complementary polynucleotide sequence (e.g., a polynucleotide sequence as described herein) is synthesized under various stringency conditions. This means that a double-stranded molecule is formed between the two genes (or a part of them). For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (1987) Methods Enzymol. 152:507).

[0181] For example, a stringent salt concentration is typically less than about 750 mM NaCl and 75 mM citrate. trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, More preferably, it is less than about 250 mM NaCl and 25 mM trisodium citrate. Sequential hybridization is achieved in the absence of organic solvents, such as formamide. whereas high stringency hybridization is at least about 35%. It can be obtained in the presence of formamide, more preferably at least about 50% formamide. Stringent temperature conditions are generally at least about 30°C, more preferably at least Preferably, the hybridization temperature will be at least about 37°C, and most preferably at least about 42°C. The reaction time, concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and carrier A variety of additional parameters, such as the inclusion or exclusion of DNA, are well known to those of skill in the art. By combining these various conditions as needed, various levels of stringency can be achieved. In one embodiment, hybridization is performed at 30° C. in 750 mM N In another embodiment, the hybridization occurs in 75 mM aCl, 75 mM trisodium citrate, and 1% SDS. Dimerization was performed at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% HCl. In another embodiment, the amplification is carried out in fluoramide and 100 μg / ml denatured salmon sperm DNA (ssDNA). Hybridization was performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, and 1% SDS. The reaction occurs in 50% formamide and 200 μg / ml ssDNA. Useful variations of these conditions are The various options will be readily apparent to one skilled in the art.

[0182] For most applications, the washing steps that follow hybridization are also stringent. Wash stringency conditions are defined by salt concentration and temperature. As mentioned above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. This can be increased by increasing the thickness of the strip for the cleaning process. The appropriate salt concentration is preferably less than about 30 mM NaCl and 3 mM trisodium citrate. and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash step are usually at least about 25°C, more preferably Preferably, the temperature is at least about 42°C, and even more preferably at least about 68°C. In an embodiment, the wash steps are performed at 25° C. using 30 mM NaCl, 3 mM trisodium citrate, and 0.1% In a more preferred embodiment, the wash steps are performed in 15 mM NaCl, 1.5 mM SDS at 42°C. In a more preferred embodiment, washing is carried out in 0.5 mM trisodium citrate and 0.1% SDS. The cleaning step was carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Dav is (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 7 2:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Int. erscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Technique) ques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: Described in A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York There are.

[0183] "Split" means divided into two or more pieces.

[0184] A "split Cas9 protein" or "split-Cas9" is a protein that is split into two separate nucleotides. Cas9 protein provided as N-terminal and C-terminal fragments encoded by the sequences The polypeptides corresponding to the N-terminal and C-terminal parts of the Cas9 protein are spliced ​​together. In certain embodiments, the Cas9 protein can be reconstituted to form a "reconstituted" Cas9 protein. The protein may be prepared by the method described in, for example, Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014 or as described in Jiang et al. (2016) Science 351: 867-871. PDB fil e. Proteins as described in 5F9R (each incorporated herein by reference). In some embodiments, the protein is split into two fragments within a disordered region of the protein. within the region of SpCas9 between approximately amino acids A292 to G364, F445 to K483, or E565 to T637. or any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other nap DNA fragments at the corresponding position in the DNA fragment. In some embodiments, the protein is SpCas9 T310, T313, A456, S469, or C57. 4 into two fragments. In some embodiments, the protein is divided into two fragments. The process is called "splitting" the protein.

[0185] In other embodiments, the N-terminal portion of the Cas9 protein is S. pyogenes Cas9 wild type (SpCas9 ) (NCBI Reference Sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2) The C-terminal part of the Cas9 protein, containing ~637 or the corresponding position / mutation, is located in the SpCas9 domain. It contains the portion of amino acids 574 to 1368 or 638 to 1368 of the wild type.

[0186] The C-terminal part of the split Cas9 is ligated to the N-terminal part of the split Cas9 to form the complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein can form a C It begins where the N-terminal part of the as9 protein ends. In embodiments, the C-terminal portion of the split-Cas9 comprises amino acids (551-651) to 1368 of spCas9. "(551-651)-1368" refers to the amino acids between amino acids 551 and 651 (inclusive). For example, the C-terminal part of the split Cas9 is spCas 9 amino acids 551–1368, 552–1368, 553–1368, 554–1368, 555–1368, 556–1368, 557 ~1368, 558~1368, 559~1368, 560~1368, 561~1368, 562~1368, 563~1368, 564~1 368, 565~1368, 566~1368, 567~1368, 568~1368, 569~1368, 570~1368, 571~1368 , 572~1368, 573~1368, 574~1368, 575~1368, 576~1368, 577~1368, 578~1368, 5 79~1368, 580~1368, 581~1368, 582~1368, 583~1368, 584~1368, 585~1368, 586 ~1368, 587~1368, 588~1368, 589~1368, 590~1368, 591~1368, 592~1368, 593~1 368, 594~1368, 595~1368, 596~1368, 597~1368, 598~1368, 599~1368, 600~1368 , 601~1368, 602~1368, 603~1368, 604~1368, 605~1368, 606~1368, 607~1368, 6 08~1368, 609~1368, 610~1368, 611~1368, 612~1368, 613~1368, 614~1368, 615 ~1368, 616~1368, 617~1368, 618~1368, 619~1368, 620~1368, 621~1368, 622~1 368, 623~1368, 624~1368, 625~1368, 626~1368, 627~1368, 628~1368, 629~1368 , 630~1368, 631~1368, 632~1368, 633~1368, 634~1368, 635~1368, 636~1368, 6 37~1368, 638~1368, 639~1368, 640~1368, 641~1368, 642~1368, 643~1368, 644 ~1368, 645~1368, 646~1368, 647~1368, 648~1368, 649~1368, 650~1368, or In some embodiments, the split Cas9 protein may comprise any one of portions 651 to 1368. The C-terminal portion of the protein includes amino acids 574 to 1368 or 638 to 1368 of SpCas9.

[0187] "Serpin1A polynucleotide" refers to a nucleic acid molecule that encodes the A1AT protein or a fragment thereof. An exemplary Serpin1A polynucleotide is available under NCBI accession number NM_000295. The sequence of the nucleotides is provided below: 1 acaatgactc ctttcggtaa gtgcagtgga agctgtacac tgcccaggca aagcgtccgg 61 gcagcgtagg cgggcgactc agatcccagc cagtggactt agcccctgtt tgctcctccg 121 ataactgggg tgaccttggt taatattcac cagcagcctc ccccgttgcc cctctggatc 181 cactgcttaa atacggacga ggacagggcc ctgtctcctc agcttcaggc accaccactg 241 acctgggaca gtgaatcgac aatgccgtct tctgtctcgt ggggcatcct cctgctggca 301 ggcctgtgct gcctggtccc tgtctccctg gctgaggatc cccagggaga tgctgcccag 361 aagacagata catcccacca tgatcaggat cacccaacct tcaacaagat cacccccaac 421 ctggctgagt tcgccttcag cctataccgc cagctggcac accagtccaa cagcaccaat 481 atcttcttct ccccagtgag catcgctaca gcctttgcaa tgctctccct ggggaccaag 541 gctgacactc acgatgaaat cctggagggc ctgaatttca acctcacgga gattccggag 601 gctcagatcc atgaaggctt ccaggaactc ctccgtaccc tcaaccagcc agacagccag 661 ctccagctga ccaccggcaa tggcctgttc ctcagcgagg gcctgaagct agtggataag 721 tttttggagg atgttaaaaa gttgtaccac tcagaagcct tcactgtcaa cttcggggac 781 accgaagagg ccaagaaaca gatcaacgat tacgtggaga agggtactca agggaaaatt 841 gtggatttgg tcaaggagct tgacagagac acagtttttg ctctggtgaa ttacatcttc 901 tttaaaggca aatgggagag accctttga gtcaaggaca ccgaggaga ggacttccac 961 gtggaccagg tgaccaccgt gaaggtgcct atgatgaagc gtttaggcat gtttaacatc 1021 cagcactgta agaagctgtc cagctgggtg ctgctgatga aatacctggg aatgccacc 1081 gccatcttct tcctgcctga tgaggggaa ctacagcacc tggaaaatga ctcacccac 1141 gatatcatca ccaagttcct ggaaatga gacagaaggt ctgccagctt catttaccc 1201 aaactgtcca ttactggaac ctatgatctg aagagcgtcc tgggtcaact ggcatcact 1261 aaggtcttca gcaatggggc tgacctctcc ggggtcacag aggaggcacc ctgaagctc 1321 tccaaggccg tgcataaggc tgtgctgacc atcgac g aga aagggactga gc tgctggg 1381 gccatgtttt tagaggccat acccatgtct atccccccg aggtcaagtt aacaaaccc 1441 tttgtcttct tatgattga acaaaatacc aagtctcccc tcttcatggg aaagtggtg 1501 aatcccaccc aaaaataact gcctctcgct cctcaacccc tcccctccat cctggcccc 1561 ctccctggat gacattaaag aagggttgag ctggtccctg cctgcatgtg ctgtaaatc 1621 cctcccatgt tttctctgag tctccctttg cctgctgagg ctgtatgtgg ctccaggta 1681 acagtgctgt cttcgggccc cctgaactgt gttcatggag catctggctg gtaggcaca 1741 tgctgggctt gaatccaggg gggactgaat cctcagctta cggacctggg ccatctgtt 1801 tctggagggc tccagtcttc cttgtcctgt cttggagtcc ccaagaagga tcacagggg 1861 aggaaccaga taccagccat gaccccaggc tccaccaagc atcttcatgt cccctgctc 1921 atcccccact cccccccacc cagagttgct catcctgcca gggctggctg gcccacccc 1981 aaggctgccc tcctgggggc cccagaactg cctgatcgtg ccgtggccca ttttgtggc 2041 atctgcagca acacaagaga gaggacaatg tcctcctctt gacccgctgt acctaacca 2101 gactcgggcc ctgcacctct caggcacttc tggaaaatga ctgaggcaga tcttcctga 2161 agcccattct ccatggggca acaaggacac ctattctgtc cttgtccttc atcgctgcc 2221 ccagaaagcc tcacatatct ccgtttagaa tcaggtccct tctccccaga gaagaggag 2281 ggtctctgct ttgttttctc tatctcctcc tcagacttga ccaggcccag aggccccag 2341 aagaccatta ccctatatcc cttctcctcc ctagtcacat ggccataggc tgctgatgg 2401 ctcaggaagg ccattgcaag gactcctcag ctatgggaga ggaagcacat acccattga 2461 cccccgcaac ccctcccttt cctcctctga gtcccgactg gggccacatg agcctgact 2521 tctttgtgcc tgttgctgtc cctgcagtct tcagagggcc accgcagctc agtgccacg 2581 gcaggaggct gttcctgaat agcccctgtg gtaagggcca ggagagtcct ccatcctcc 2641 aaggccctgc taaaggacac agcagccagg aagtcccctg ggcccctagc gaaggacag 2701 cctgctccct ccgtctctac caggaatggc cttgtcctat ggaaggcact ccccatccc 2761 aaactaatct aggaatcact gtctaaccac tcactgtcat gaatgtgtac taaaggatg 2821 aggttgagtc ataccaaata gtgatttcga tagttcaaaa tggtgaaatt gcaattcta 2881 catgattcag tctaatcaat ggataccgac tgtttcccac acaagtctcc gttctctta 2941 agcttactca ctgacagcct ttcactctcc acaaatacat taaagatatg ccatcacca 3001 agcccccctag gatgacacca gacctgagag tctgaagacc tggatccaag tctgacttt 3061 tccccctgac agctgtgtga ccttcgtgaa gtcgccaaac ctctctgagc ccagtcatt 3121 gctagtaaga cctgcctttg agttggtatg atgttcaagt tagataacaa atgtttata 3181 cccattagaa cagagaataa atagaactac atttcttgca The PAM sequence is highlighted to show the correct sequence after adenine base editing.

[0188] "Subject" means a mammal, whether human or non-human, e.g., a bovine, equine, This includes, but is not limited to, the following families: Canidae, Ovine, or Feline. livestock, livestock raised to produce labor and provide goods, such as food Domesticated animals, including but not limited to cows, goats, and chickens , horses, pigs, rabbits, and sheep.

[0189] "Substantially identical" means that the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein) is substantially identical to the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein). any one of the sequences) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) It means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to In this form, such sequences are identified at the amino acid or nucleic acid level as sequences used for comparison. having at least 60%, 80%, 85%, 90%, 95% or even 99% identity at the level. Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE UP / PRETTYBOX program). Such software is available in various permutations. By assigning degrees of homology to the sequences, deletions, and / or other modifications, it is possible to determine whether the sequences are identical or Similar sequences are matched. Conservative substitutions typically include substitutions within the following groups: leucine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, Paragine, glutamine; serine, threonine; lysine, arginine; phenylalanine In an exemplary approach for determining the degree of identity, closely related Showing the sequence -3 and e -100 You can use the BLAST program, which includes a probability score between Cut.

[0190] COBALT can be used, for example, with the following parameters: a) Alignment parameters: Gap penalty -11, -1 and end gap penalty Ruti-5, -1 b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find conserved columns Continue recalculating c) Query Clustering Parameters: Use query clustering on; Code size 4; max cluster distance 0.8; regular alphabet. The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) Gap open: 10; c) Gap extension: 0.5; d) Output format: vs; e) End Gap Penalty: False; f) Terminal gap open: 10; and g) Terminal gap extension: 0.5.

[0191] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase or deaminase (e.g., It is deaminated by a fusion protein containing an enzyme such as adenine deaminase.

[0192] As used herein, the terms "treat," "treating," "Treatment" and the like means the alleviation or amelioration of a disorder and / or its associated symptoms. refers to the process of treating or obtaining a desired pharmacological and / or physiological effect. Treating a condition or disorder requires the complete elimination of the associated disorder, condition or symptom. It will be understood that the effect is not necessarily therapeutic. In some embodiments, the effect is therapeutic, That is, without limitation, the effect may be to alleviate the disease and / or adverse effects resulting from the disease. Partially or completely reduce, diminish, eliminate, alleviate, relieve, lessen the intensity of, or cure symptoms In some embodiments, the effect is preventative, i.e., the effect prevents the onset of the disease or condition. To this end, the disclosed methods include those described herein. In one embodiment, the method comprises administering a therapeutically effective amount of a composition as described above to a patient suffering from a disease. is alpha-1 antitrypsin deficiency (A1AD).

[0193] "Uracil glycosylase inhibitor," or "UGI," refers to a compound that inhibits the uracil excision repair system. In one embodiment, the agent inhibits a host uracil-DNA glycosylase. It is a protein or fragment thereof that binds to and prevents the removal of uracil residues from DNA. In embodiments, the UGI may inhibit uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain is a protein, fragment, or domain thereof that can The domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain The domain comprises a fragment of the exemplary amino acid sequence provided below. In some embodiments, UGI fragments may be at least 60%, at least 65%, or at least 69% of the exemplary UGI sequences provided below. At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 Contains 5%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% In some embodiments, the UGI comprises an exemplary amino acid sequence, as described below. The amino acid sequence of the target UGI polypeptide includes an amino acid sequence homologous to the target UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof may be a wild-type UGI or UGI sequence or a portion thereof, as described below. or part thereof at least 70%, at least 75%, at least 80%, at least 85%, At least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least Exemplary U have at least 99%, at least 99.5%, at least 99.9% or 100% identity. GI comprises the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGEN KIKML.

[0194] The term "vector" refers to a means for introducing a nucleic acid sequence into a cell, resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, An "expression vector" includes a vector that is expressed in a recipient cell. An expression vector is a nucleic acid sequence that contains a nucleotide sequence to be expressed. Additional nucleic acid sequences, such as start, stop, enhancer sequences, that promote and / or facilitate expression The gene may include a promoter and a secretory sequence.

[0195] Any composition or method provided herein may be used in combination with any other composition or method provided herein. More than one composition and method may be combined.

[0196] As used herein, the terms "some embodiments," "an embodiment," "one embodiment," and "an A reference to "another embodiment" may refer to the specific features, structures, or aspects described in connection with that embodiment. or features are included in at least some embodiments of the present disclosure, but not necessarily all. This means that the embodiments are not necessarily included.

[0197] The recitation of a list of chemical groups in any definition of a variable herein means that the listed groups The definitions of the variables as any single group or combination of: The description of an embodiment with respect to a variable or aspect does not necessarily mean that it is a single embodiment or that it is a part of any single embodiment. This includes any embodiment in combination with any other embodiment or portion thereof.

[0198] DNA editing alters disease states by correcting pathogenic mutations at the genetic level Until recently, all DNA editing platforms induces DNA double-strand breaks (DSBs) at specific genomic sites and relies on endogenous DNA repair pathways to repair the DSBs. They function by determining product outcomes in a probabilistic manner, resulting in a complex population of gene products. Accurate, user-defined repair results are achieved through the homology-directed repair (HDR) pathway. Although high-efficiency repair using HDR is achievable in therapeutically relevant cell types, However, this path has been hindered by several difficulties. In fact, this path is dominated by competing, Furthermore, HDR is less efficient than the non-homologous end joining pathway, which is more likely to occur in the G1 and G2 stages of the cell cycle. This restriction of DSBs to the S phase prevents accurate repair of DSBs in post-mitotic cells. , in these populations, we can efficiently and programmatically analyze genomes. Altering the sequence has proven difficult or impossible. [Brief explanation of the drawings]

[0199] [Figure 1]Figures 1A-1C show plasmids. Figure 1A is an expression vector encoding the TadA7.10-dCas9 base editor. Figure 1B is a plasmid containing a nucleic acid molecule encoding a protein that confers chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene that has been disabled by two point mutations. Figure 1C is a plasmid containing a nucleic acid molecule encoding a protein that confers chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene that has been disabled by three point mutations. [Figure 2] Figure 2 shows images of bacterial colonies transduced with the expression vector shown in Figure 1A-C, which contains a deletion of the kanamycin resistance gene. The vector contained an ABE7.10 variant generated using error-prone PCR. Bacterial cells expressing these "evolved" ABE7.10 variants were selected for kanamycin resistance using increasing concentrations of kanamycin. Bacteria expressing ABE7.10 variants with adenosine deaminase activity were able to correct the mutation introduced into the kanamycin resistance gene, thereby restoring kanamycin resistance. Kanamycin-resistant cells were selected for further analysis. [Figure 3] Figure 3 is a graph quantifying the potency and specificity of selected ABE8s listed in Table 6. Editing was assayed at the alpha-1 antitrypsin locus in HEK293T cells. [Figure 4] Figures 4A and 4B illustrate the editing efficiency and specificity of ABE8. Figures 4A and 4B are graphs quantifying the base editing and specificity of selected ABE8s listed in Table 6. Single variant TadA deaminase domains or wild-type TadA deaminase were evaluated. [Figure 5]Figure 5 presents a graph showing the efficacy of ABE8 in editing the on-target adenine (A) base pair bystander A. Notably, ABE8 produces a five-fold increase in editing (i.e., A·T to G·C conversion) at the A1AD site compared to the effective TadA deaminase, ABE7.10. [Figure 6] Figures 6A-6D show nucleic acid sequences, tables, and bar graphs for the generation of improved rates of nucleobase correction in primary PiZ fibroblasts through base editor engineering. Figure 6A shows the target site DNA sequence encoding the PiZ mutation associated with A1AD. This sequence includes a 20-nucleotide protospacer and the non-canonical spCas9 NGC PAM. Figure 6B presents a table listing both the TadA deaminase and Cas9 PAM variant components of various editors used to correct the PiZ mutation. Figures 6C and 6D present bar graphs showing the editing rates observed in patient-derived PiZZ fibroblasts (GM11423 Corriel Biorepository) transfected with base editing reagents using the Neon electroporation system. Each treatment consisted of 70,000 fibroblasts, 100 ng mRNA encoding a base editor, and 50 ng alpha-1 corrected gRNA in 10 μl electroporation buffer. After 48 hours of recovery, cells were lysed and the locus of interest was examined by sequencing the targeted amplicon. Data were obtained from two independent experiments. These data and results demonstrate improved editing efficiency with both NGC PAM recognition optimization (Variant 1-3, Figures 6B and 6C) and TadA deaminase optimization through the incorporation of ABE8 / 9 mutations (Variant 4-9, Figures 6B-6D). [Figure 7]Figures 7A-7D present nucleic acid sequences, tables, and graphs related to the increase in serum A1AT produced by lipid nanoparticle (LNP)-mediated delivery and base editing in NSG-PiZ transgenic mice. Figure 7A shows the target site DNA sequence, including the 20-nucleotide protospacer and noncanonical spCas9 NGC PAM. Figure 7B presents a table listing both the TadA deaminase and Cas9 PAM variant components of the various editors used to correct PiZ mutations. Figure 7C presents a graph showing the editing rate observed in whole liver gDNA from an NSG-PiZ transgenic mouse model 7 days after treatment with 1.5 mg / kg LNP containing a 1:1 weight ratio of gRNA and mRNA encoding the base editor. Commercially available NSG-PiZ mice express mutant human SERPINA1 (Glu342Lys mutation) on an immunodeficient NOD-SCID gamma (NSG) background, which provides a stable background for human hepatocytes after partial hepatectomy (The Jackson Laboratory, Mount Desert Island, ME). Results demonstrated that ngcABEvar9 produces higher editing rates than the earlier version, variant 8. Figure 7D presents a graph showing that editing rates correlate with increases in serum alpha-1 antitrypsin compared to pretreatment samples, as measured by MSD sandwich immunoassay. Based on these results, base editing using the ABE8 reagent can address alpha-1 antitrypsin deficiency and its potential pulmonary sequelae. [Figure 8]Figure 8 is a table showing Cas9 variants that can access all possible PAMs within the NRNN PAM space. Only Cas9 variants that require recognition of three or fewer defined nucleotides in the PAM are listed. Non-G PAM variants include SpCas9-NRRH, SpCas9-NRTH, and SpCas9-NRCH (Miller, SM, et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), ( / / doi.org / 10.1038 / s41587-020-0412-8), the contents of which are incorporated herein by reference in their entirety. DETAILED DESCRIPTION OF THE INVENTION

[0200] As described below, the present invention provides a method for treating alpha-1 antitrypsin deficiency (A1AD)-associated In some embodiments, the present invention features compositions and methods for modifying mutations in a gene comprising: The editing corrects the deleterious mutation and ensures that the edited polynucleotide is identical to the wild-type reference polynucleotide. In another embodiment, the editing amends the deleterious mutation. The modified polynucleotide is then altered so that it contains a benign mutation.

[0201] The present invention relates, at least in part, to a base editor comprising an adenosine deaminase variant. Based on the discovery that a marker can effectively and precisely edit deleterious mutations associated with A1AD, Made.

[0202] [Alpha-1 Antitrypsin Deficiency (A1AD)] Alpha-1 antitrypsin (A1A) is encoded by the SERPINA1 gene on chromosome 14. This glycoprotein is synthesized primarily in the liver and distributed into the bloodstream. The serum concentration in healthy adults is 1.5-3.0 g / L (20-52 μmol / L). Proteins diffuse into the lung interstitium and alveolar lining fluid, where they inactivate neutrophil elastase. thereby protecting lung tissue from protease-mediated damage. Trypsin deficiency (A1AD) is inherited in an autosomal codominant manner. Over 100 mutations in the SERPINA1 gene Genetic variants that can cause HIV infection have been described, but not all are disease-associated. The alphabetical designation of these variants is based on their migration speed on gel electrophoresis. The common variant is the M (intermediate mobility) allele (PiM), which is the two most frequent defective alleles. The alleles are PiS and PiZ (the latter has the slowest migration velocity). Several mutations have been described that do not produce a gene; these are called "null" alleles. The most common genotype is MM, which has normal serum levels of alpha-1 antitrypsin. Most people with severe deficiency are homozygous for the Z allele (ZZ). More than 60,000 A1AD patients in the United States have the severe ZZ phenotype. The Z protein is expressed in hepatocytes. During production in the endoplasmic reticulum, it misfolds and polymerizes; these abnormal polymers are secreted into the liver. This leads to a severe decrease in serum levels of alpha-1 antitrypsin. Unstable A1AT production leads to liver and / or lung lesions in patients with A1AD. The liver disease seen in patients with alpha-1 antitrypsin deficiency is caused by the accumulation of abnormal enzymes in liver cells. The accumulation of trypsin-1 protein results in autophagy and endoplasmic reticulum storage. This is caused by cellular reactions, including the cell death response and apoptosis. Decreased circulating levels of α-1 antitrypsin inhibit neutrophil elastase activity in the lungs This imbalance of proteases and antiproteases results in Lung disease occurs.

[0203] Alpha-1 antitrypsin deficiency ("A1AD") is most common in Caucasians and It most frequently affects the lungs and liver. In the lungs, the most common signs are ulcers most evident at the lung bases. It is a markedly early-onset (patients in their 30s and 40s) panacinolar emphysema. Lobar emphysema may occur, as may bronchiectasis. Symptoms include dyspnea, wheezing, and coughing. Pulmonary function tests in affected individuals show known symptoms consistent with COPD. however, a bronchodilator response may be observed and a diagnosis of asthma may be made. The liver disease caused by the ZZ genotype manifests in a variety of ways. Symptoms may appear during the neonatal period, including cholestatic jaundice and sometimes acholestasis (pale or Clay-colored liver and hepatomegaly. Conjugated bilirubin, transaminases and cancer in the blood Maglutamyltransferase levels are elevated in older children and adults. Liver disease in this setting may be discovered incidentally as elevated transaminases or as a result of variceal bleeding. Symptoms may include established signs of cirrhosis, including ascites. -1 antitrypsin deficiency also predisposes patients to hepatocellular carcinoma. The heterozygous Z mutation is necessary for liver disease to develop, but the heterozygous Z mutation is necessary for hepatitis C infection and and cystic fibrosis liver disease, placing patients at greater risk of more severe liver disease. and may act as genetic modifiers of other diseases.

[0204] The two most common clinical variants in A1AD are the E264V (PiS) and E342K (PiZ) alleles. The clinical single nucleotide variant E342K (PiZ) results in an unstable and / or inactive A1AT protein. Proteins are produced, resulting in liver and lung toxicity. Inheritance is autosomal codominant. More than half of A1AD patients harbor at least one copy of the E342K mutation.

[0205] JPEG2025037975000017.jpg66163

[0206] In some embodiments, the disease or disorder is alpha-1 antitrypsin deficiency (A1AD). In some embodiments, the pathogenic mutation is in the gene SERPINA1. In some embodiments, the SERPINA1 mutation is E342K (PiZ allele). In , the A at position 7 is edited to G, restoring the PiZ allele to the wild-type allele.

[0207] [Nucleobase Editor] A base editor for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Disclosed herein are nucleic acid base editors or nucleobase editors. Polynucleotide Programmable Nucleotide Binding Domains and Nucleobase Editing Domains a nucleobase editor or base editor comprising an enzyme such as adenosine deaminase The polynucleotide-programmable nucleotide binding domain binds to the bound guide polynucleotide. When combined with a nucleotide (e.g., gRNA), the target polynucleotide sequence (i.e., Complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence. and capable of specifically binding (through synthesis) to the target nucleic acid sequence for which editing is desired. In some embodiments, the base editor can be localized to the target polynucleotide. The nucleic acid sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. - RNA hybrids.

[0208] [Polynucleotide-programmable nucleotide-binding domain] Polynucleotide programmable nucleotide binding domains also bind to nuclear RNA. It is understood that the present invention may include an acid programmable protein. The polynucleotide-programmable nucleotide binding domain is The nucleotide-binding domain can be associated with a nucleic acid that guides the RNA. Other nucleic acid programs are possible. Functional DNA binding proteins are also within the scope of this disclosure, although they are not specifically listed in this disclosure. It has not been done.

[0209] The polynucleotide-programmable nucleotide-binding domain of the base editor is The polynucleotide programmable vector itself can contain one or more domains. The active nucleotide binding domain can include one or more nuclease domains. In some embodiments, the polynucleotide-programmable nucleotide binding domain The nuclease domain can include an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to an enzyme that can degrade nucleic acids (e.g., RNA or DNA). The term "endo" refers to a protein or polypeptide that can be digested from its free end. A "nuclease" is an enzyme that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease refers to a protein or polypeptide that can The endonuclease is capable of cleaving a single strand of a double-stranded nucleic acid. Nucleases are capable of cleaving both strands of a double-stranded nucleic acid molecule. In this embodiment, the polynucleotide-programmable nucleotide binding domain binds to a deoxyribonucleic acid. In some embodiments, the polynucleotide programmable nucleic acid may be a nucleic acid fragment. The protease binding domain can be a ribonuclease.

[0210] In some embodiments, the polynucleotide-programmable nucleotide binding domain The nuclease domain of cleaves zero, one, or two strands of the target polynucleotide In some embodiments, the polynucleotide programmable nucleotide The binding domain may comprise a nickase domain. A "deoxyglucose" can cleave only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). Polynucleotide-programmable nucleotide-binding domain containing a nuclease domain capable of In some embodiments, the nickase is an active polynucleotide program. Polynucleotides can be synthesized by introducing one or more mutations into the target nucleotide-binding domain. Fully catalytically active (e.g., native) nucleotide-programmable nucleotide-binding domains For example, polynucleotide-programmable nucleotide binding domains can be obtained. If the gene contains a nickase domain derived from Cas9, the nickase domain derived from Cas9 can include a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 retains catalytic activity, thereby cleaving one strand of a nucleic acid duplex. In another example, the nickase domain from Cas9 can include an H840A mutation. while the amino acid residue at position 10 remains D. The enzyme removes all or part of the nuclease domain that is not required for nickase activity. By doing so, the polynucleotide-programmable nucleotide binding domain can be fully It can be obtained from a catalytically active (e.g., naturally occurring) form. For example, a polynucleotide programmable When the functional nucleotide-binding domain contains a nickase domain derived from Cas9, The resulting nickase domain is a deletion of all or part of the RuvC or HNH domain. This may include losses.

[0211] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0212] Thus, polynucleotides containing a nickase domain can programmably bind nucleotides. The base editor containing the domain can be used to target a specific polynucleotide target sequence (e.g., a bound generating a single-stranded DNA break (nick) at a site (determined by the complementary sequence of the target nucleic acid) In some embodiments, a nickase domain (e.g., a nickase derived from Cas9) can be used. Nucleic acid double-stranded target polynucleotide sequences cleaved by base editors containing The strand in the column is the strand that is not edited by the base editor (i.e., Thus, the strand that is cleaved is the opposite strand from the strand containing the base to be edited. base editors containing a nickase domain (e.g., the nickase domain from Cas9) , can cleave the strand of the DNA molecule targeted for editing. In this case, the non-target strand is not cleaved.

[0213] catalytically inactive (i.e., incapable of cleaving a target polynucleotide sequence) Base editors containing polynucleotide programmable nucleotide binding domains are also disclosed. As used herein, the terms "catalytically inactive" and "nucleic acid" are used interchangeably. "Serine inactive" refers to one or more mutations and / or mutations that result in the inability to cleave a strand of nucleic acid. Polynucleotide-programmable nucleotide binding domains having deletions and / or deletions In some embodiments, catalytically inactive poly(methylenediamine)s are used interchangeably to refer to Nucleotide-programmable nucleotide-binding domain base editors are designed to lack of nuclease activity as a result of specific point mutations in the nuclease domain For example, in the case of a base editor containing a Cas9 domain, Cas9 can Such mutations can include both the H840A and H840A mutations. In another embodiment, the catalytically active nuclease domain is inactivated, resulting in the loss of nuclease activity. The inactive polynucleotide-programmable nucleotide binding domain is These may include one or more deletions of all or part of the RuvC1 and / or HNH domains. In a further embodiment, a catalytically inactive polynucleotide program can be used. Potential nucleotide binding domains may contain point mutations (e.g., D10A or H840A) as well as nucleotide It includes deletion of all or part of the ase domain.

[0214] Also provided herein are the following polynucleotide-programmable nucleotide binding domains: Catalytically inactive polynucleotides can be programmed from previously functional versions. Mutations that can generate functional nucleotide binding domains are also contemplated. In the case of catalytically inactive Cas9 ("dCas9"), mutations other than D10A and H840A were present. Variants are provided that result in nuclease-inactivated Cas9. Such mutations include: For example, other amino acid substitutions in D10 and H840, or the nuclease domain of Cas9 Other substitutions within the HNH nuclease subdomain and / or RuvC1 subdomain (e.g., Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and the art. It will be apparent to those skilled in the art based on knowledge in the technical field and will be within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to: Although not intended, D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A Mutant domains are included (e.g., Prashant et al., CAS9 transcriptional activity ators for target specificity screening and paired nickases for cooperative genomes e engineering. Nature Biotechnology. 2013; 31(9): 833-838 (in its entirety) The contents of which are incorporated herein by reference).

[0215] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains from CRISPR proteins, restriction nucleases, Meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (Z In some embodiments, the base editor includes a CRISPR (i.e. Clustered Regularly Interspaced Short Palindromic Repeats (CIR)-mediated modification Natural or modified proteins that can bind to nucleic acid sequences via guide nucleic acids. A polynucleotide comprising a programmable nucleotide binding domain, Such proteins are referred to herein as "CRISPR proteins." Disclosed herein are polynucleotides comprising all or a portion of a CRISPR protein. Base editors containing nucleotide-programmable nucleotide-binding domains (i.e., A base editor containing all or part of the CRISPR protein as a domain, The CRISPR protein-derived domain of the target gene is also called the CRISPR protein-derived domain. The inserted CRISPR protein-derived domains were compared to wild-type or native CRISPR proteins. For example, as described below, a domain from a CRISPR protein can be , one or more mutations, insertions, or deletions compared to the wild-type or native CRISPR protein , rearrangement and / or recombination.

[0216] CRISPR provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids) CRISPR clusters are composed of spacers, preceding mobile elements, and The CRISPR cluster contains complementary sequences and the target invader nucleic acid. The CRISPR cluster is transcribed and produces CRISPR RNA (cr In the type II CRISPR system, the correct processing of pre-crRNA is required. The markers are transcoding small RNAs (tracrRNAs), endogenous ribonuclease 3 (rnc), and Ca tracrRNA requires ribonuclease 3-assisted processing of pre-crRNA. Cas9 / crRNA / tracrRNA then synthesizes a linear or nucleotide sequence complementary to the spacer. The target strand that is not complementary to the crRNA is cleaved with an endonuclease. First endonucleolytic cleavage, then exonucleolytic trimming to 3'-5' In nature, DNA binding and cleavage requires both a protein and RNA. Single-guide RNA incorporating both crRNA and tracrRNA aspects into a single RNA species ("sgRNA", or simply "gRNA") can be produced. See, e.g., Jinek M., Chylins ki K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(20 12), the entire contents of which are incorporated herein by reference. It recognizes a short motif (PAM or protospacer adjacent motif) in the SPR repeat sequence and Helps distinguish between "self" and "non-self."

[0217] In some embodiments, the methods described herein involve using engineered Cas proteins. Guide RNA (gRNA) contains a scaffold sequence required for Cas binding and a modified It is a short synthetic RNA consisting of a user-defined approximately 20-base spacer that defines the genomic target. Therefore, one skilled in the art can alter the genomic target of Cas protein specificity, which , how specific the gRNA targeting sequence is to its genomic target compared to the rest of the genome The amount of data that is used is determined in part by whether the data is

[0218] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.

[0219] In some embodiments, a domain from a CRISPR protein incorporated within a base editor The nucleic acid binds to a target polynucleotide when combined with a binding guide nucleic acid. Endonucleases (e.g., deoxyribonucleases or ribonucleases) that can In some embodiments, the base editor is derived from a CRISPR protein incorporated within the base editor. The binding domain binds to the target polynucleotide when combined with the bound guide nucleic acid. In some embodiments, a nickase is a nickase that can be incorporated into a base editor. The embedded CRISPR protein-derived domain, when combined with the bound guide nucleic acid, It is a catalytically inactive domain that can bind to a target polynucleotide when In some embodiments, the target that binds to the CRISPR protein-derived domain of the base editor The polynucleotide is DNA. In some embodiments, the base editor CRISPR The target polynucleotide that binds to the protein-derived domain is RNA.

[0220] Cas proteins that can be used herein include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, C as5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10 , Csy1 , Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1 , Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO , Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and and Cas12i, CARF, DinG, their homologs, or modified versions thereof. CRISPR enzymes, like Cas9, contain two functional endonuclease domains, RuvC and HNH. CRISPR enzymes can have DNA cleavage activity within and / or across target sequences. A target sequence, such as within the complementary strand of a sequence, can induce cleavage of one or both strands. For example, the CRISPR enzyme may target sequences approximately 1, 2, 3, 4, 5, 6, or 10 nucleotides from the first or last nucleotide of the target sequence. , 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs or more apart, can induce cleavage of both strands.

[0221] Lacking the ability to cleave one or both strands of a target polynucleotide containing the target sequence To achieve this, vectors encoding CRISPR enzymes mutated relative to the corresponding wild-type enzymes were used. Cas9 can be used in conjunction with a wild-type exemplary Cas9 polypeptide (e.g., derived from S. pyogenes). (previously known as Cas9) and at least approximately 50%, 60%, 70%, 80%, 90%, 91%, 92% , 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology Cas9 can refer to a polypeptide having the properties of a wild-type exemplary Cas9 polypeptide. For example, those derived from S. pyogenes, the maximum or maximum is approximately 50%, 60%, 70%, %, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or polypeptides having sequence homology. or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any of these It can refer to modified forms of the Cas9 protein that may contain amino acid changes, such as combinations.

[0222] In some embodiments, the CRISPR protein-derived domain of the base editor includes Cory nebacterium ulcerans (NCBI reference: NC_015683.1, NC_017317.1); Corynebacterium dipht heria (NCBI references: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI references: NC_021284.1); Prevotella intermedia (NCBI reference: NC_017861.1); Spiroplasma taiwane nse (NCBI reference: NC_021846.1); Streptococcus iniae (NCBI reference: NC_021314.1); Bellie lla baltica (NCBI reference: NC_018010.1); Psychroflexus torquis (NCBI reference: NC_018721. 1); Streptococcus thermophilus (NCBI reference: YP_820832.1); Listeria innocua (NCBI reference: YP_820832.1); Reference: NP_472073.1); Campylobacter jejuni (NCBI reference: YP_002344900.1); Neisseria men ingitidis (NCBI reference: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus The Cas9 vector may include all or part of Cas9 from Cucumis aureus.

[0223] [Cas9 domain, a nucleobase editor] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J. J., McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S ., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, R en Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin R. E., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K. , Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpen tier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA e ndonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I ., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012). (The entire contents of each are incorporated herein by reference.) Cas9 orthologs are Although not widely known, it has been described in a variety of species, including S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpenti. er, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2 013) Cas9 sequences from the organisms and loci disclosed in RNA Biology 10:5, 726-737 No. 6,229,693, the entire contents of which are incorporated herein by reference.

[0224] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas 9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domains are available as nuclease-active Cas9 domains, nuclease-inactive Cas9 domains (dCas9), or or Cas9 nickase (nCas9). In some embodiments, the Cas9 domain is Nuclease activity domains. For example, the Cas9 domain binds both strands of a double-stranded nucleic acid (e.g., For example, it may be a Cas9 domain that cleaves both strands of a double-stranded DNA molecule. In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain is directed to any one of the amino acid sequences described herein. At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, At least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least The amino acid sequence may be at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of the target gene. In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 , 41 , 42 , 43 , 44 , 45 , 46 , 47 , 48 , 49 , 50 or more mutations In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. at least 10, at least 15, at least 20, or fewer amino acid sequences compared to any one of the amino acid sequences At least 30, at least 40, at least 50, at least 60, at least 70, at least 80, At least 90, at least 100, at least 150, at least 200, at least 250, less At least 300, at least 350, at least 400, at least 500, at least 600, at least 7 00, at least 800, at least 900, at least 1000, at least 1100, or at least also contain amino acid sequences with 1200 identical consecutive amino acid residues.

[0225] In some embodiments, proteins comprising fragments of Cas9 are provided. For example, In some embodiments, the protein comprises one of the following two Cas9 domains: (1) Ca (2) the gRNA binding domain of Cas9; and (3) the DNA cleavage domain of Cas9. In some embodiments, Cas9 or Proteins containing fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with wild-type Cas9 or a fragment thereof. For example, Cas9 variants may share at least one of the following homology with wild-type Cas9: At least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, At least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical , at least about 99.5% identical, or at least about 99.9% identical. Cas9 variants have the following advantages over wild-type Cas9: 2, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 3 2, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, and In some embodiments, the Cas9 variant may have more amino acid changes. , a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), At least about 70% identical, at least about 80% identical, at least about 90% identical to the corresponding fragment of live Cas9 % identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical In some embodiments, the fragment is at least the amino acid length of the corresponding wild-type Cas9. At least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 5 5%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, At least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, At least 98%, at least 99%, or at least 99.5%. In some embodiments, the fragment is at least 100 amino acids in length. At least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 mesh The length of the amino acid.

[0226] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The Cas9 sequence may comprise a full-length amino acid sequence of interest, such as one of the Cas9 sequences provided herein, but In other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments provided herein, additional suitable sequences for Cas9 domains and fragments will be apparent to those of skill in the art. Perhaps.

[0227] The Cas9 protein guides the protein to a specific DNA sequence complementary to its guide RNA. In some embodiments, the polynucleotide The programmable nucleotide binding domain can be a Cas9 domain, e.g., a nuclease-active Ca s9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, and CasY. , Cpf1, Cas12b / C2Cl, and Cas12c / C2C3.

[0228] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025037975000018.jpg169163 (single underline: HNH domain; double underline: RuvC domain)

[0229] In some embodiments, the wild-type Cas9 comprises the following nucleotides and / or amino acids: Corresponding to or containing sequences: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGTCCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAAACATGAACGGCACCCCATCTTTGGAAACATAGGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAGTTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGA AAATATTATCCATTTGTTTACTCTTCCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025037975000019.jpg166162 (single underline: HNH domain; double underline: RuvC domain)

[0230] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NCBI Reference Reference sequence: NC_002737.2 (nucleotide sequence is as follows) and Uniprot reference sequence: Q99ZW2 (The amino acid sequence corresponds to: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAGGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025037975000020.jpg167161 (single underline: HNH domain; double underline: RuvC domain)

[0231] In some embodiments, Cas9 is derived from Corynebacterium ulcerans (NCBI reference: NC_015683. 1, NC_017317.1); Corynebacterium diphtheria (NCBI reference: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI reference: NC_021284.1); Prevotella intermedia (NCBI Reference: NC_017861.1); Spiroplasma taiwanense (NCBI reference: NC_021846.1); Streptococc us iniae (NCBI reference: NC_021314.1); Belliella baltica (NCBI reference: NC_018010.1); Psy chroflexus torquisI (NCBI reference: NC_018721.1); Streptococcus thermophilus (NCBI reference: NC_018721.1); Reference: YP_820832.1), Listeria innocua (NCBI reference: NP_472073.1), Campylobacter jejuni (NCBI Reference: YP_002344900.1) or Neisseria meningitidis (NCBI Reference: YP_002342 100.1), or any other organism-derived Cas9.

[0232] Additional Cas9 proteins (e.g., nuclease-inactive (dead) Cas9 (dCas9), Cas9 Nickase (nCas9), or nuclease-active Cas9, its variants and homologs It is understood that within the scope of this disclosure are any and all Cas9 proteins, including but not limited to: In some embodiments, Cas9 transcription factors are used in the transcription of HIV-1-associated ... The protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 The protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein The quality is nuclease-active Cas9.

[0233] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can cleave either strand of a double-stranded nucleic acid molecule without cleaving either strand. It can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule). The nuclease-inactive dCas9 domain is a D10X mutation of the amino acid sequence described herein and and H840X mutations, or the corresponding mutations in any of the amino acid sequences provided herein. and X is any amino acid change. The enzyme-inactive dCas9 domain comprises the D10A mutation and H of the amino acid sequence provided herein. 840A mutation, or the corresponding mutation in any of the amino acid sequences described herein. As an example, a nuclease-inactive Cas9 domain is incorporated into the cloning vector pPla tTET-gRNA 2 (accession number BAV54124) contains the amino acid sequence shown in:

[0234] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-s specific control of gene expression.” Cell. 2013; 152(5):1173-83 (see the entire contents of which are incorporated herein by reference).

[0235] Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. These additional examples will be apparent to those skilled in the art based on their knowledge and are within the scope of this disclosure. Exemplary suitable nuclease-inactive Cas9 domains include, without limitation, D10A / H840A, D10A / D83 9A / H840A, and D10A / D839A / H840A / N863A mutant domains (e.g., Pra shant et al., CAS9 transcriptional activators for target specificity screening a nd paired nickases for cooperative genome engineering. Nature Biotechnology. 201 3; 31(9): 833-838, the entire contents of which are incorporated herein by reference. ).

[0236] In some embodiments, the Cas9 nuclease is an inactive (e.g., inactivated) DNA cleavage site. Cas9 has a cleavage domain, i.e., Cas9 is called the "nCas9" protein (for "nickase" Cas9). The nuclease-inactivated Cas9 protein is interchangeably referred to as "dCas" It is also called the "nuclease-"dead" Cas9 protein or catalytically inactive Cas9. A method for generating a Cas9 protein (or a fragment thereof) with an inactive DNA cleavage domain The method is known (e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing osing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Exp See “Ression” (2013) Cell. 28; 152(5): 1173-83 (the contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is a HNH nuclease subdomain. It is known to contain two subdomains: the HNH subdomain and the RuvC1 subdomain. The RuvC1 subdomain cleaves the complementary strand of the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within the domain can suppress the nuclease activity of Cas9. For example, mutations D10A and and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al. l, Science. 337:816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, the dCas9 domain is any of the dCas9 domains provided herein. At least 60%, at least 65%, at least 70%, at least 75%, or at least at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least and amines having 97%, at least 98%, at least 99%, or at least 99.5% identity with each other. In some embodiments, the Cas9 domain comprises an amino acid sequence as set forth herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, compared to any one of the sequences 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 domain comprises an amino acid sequence having a mutation. At least 10, at least 1, 5, at least 20, at least 30, at least 40, at least 50, at least 60, less At least 70, at least 80, at least 90, at least 100, at least 150, at least 200 , at least 250, at least 300, at least 350, at least 400, at least 500, At least 600, at least 700, at least 800, at least 900, at least 1000, at least and amino acid sequences having at least 1100 or at least 1200 identical consecutive amino acid residues. include.

[0237] In some embodiments, dCas9 is provided with one or more abrupt changes that inactivate Cas9 nuclease activity. It corresponds to, or contains part or all of, a mutated Cas9 amino acid sequence. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or another Ca mutation. Contains the corresponding mutation in s9.

[0238] In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG2025037975000021.jpg166161 (single underline: HNH domain; double underline: RuvC domain)

[0239] In some embodiments, the Cas9 domain comprises a D10A mutation, while the Residue 840 in the amino acid sequence, or any of the amino acid sequences provided herein The residue at the corresponding position in either remains a histidine.

[0240] In other embodiments, D10A and H84, e.g., yield a nuclease-inactivated Cas9 (dCas9), dCas9 variants with mutations other than 0A are provided. Such mutations can be used in, for example, For example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9. Substitutions (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9, at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, those having about 5 amino acids, about 10 amino acids, or about 99.9% identity are provided. amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about Amino acid sequences shorter or longer than 50 amino acids, about 75 amino acids, about 100 amino acids or more Variants of dCas9 are provided, having the following structure: In some embodiments, the Cas9 domain is a Cas9 nickase. Cas9 proteins are capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase can be a target protein of a double-stranded nucleic acid molecule. The Cas9 nickase cleaves the target strand, which allows the Cas9 to separate the gRNA (e.g., sgRNA) bound to the Cas9. This means cutting the base-paired (complementary) strand. In this embodiment, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule. This is because the Cas9 nickase base pairs with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase cleaves the untranslated strand. containing the H840A mutation and having an aspartic acid residue at position 10, or the corresponding mutation In some embodiments, the Cas9 nickase has the structure described herein. At least 60%, at least 65%, at least 70%, at least 75% of any one of the cases , at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical Additional suitable Cas9 nickases may be identified based on this disclosure and knowledge in the art. Such modifications will be apparent to those skilled in the art and are within the scope of this disclosure.

[0241] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD

[0242] In some embodiments, Cas9 is used in the archaic microorganisms that comprise the domain and kingdom of unicellular prokaryotic microorganisms. Refers to Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, programmable Nucleotide-binding proteins are described, for example, in Burstein et al., "New CRISPR-Cas systems for from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 Alternatively, the CasX or CasY protein may be a CasX or CasY protein, as described herein, the entire contents of which are incorporated by reference. Incorporated into the technical document. Using genome-resolved metagenomics to characterize the archaeal domain of life. Many CRISPR-Cas systems have been identified, including Cas9, which was first reported in this field. The as9 protein is a key regulator of active CRISPR-Cas systems in the little-studied nanoarchaea. Two previously unknown systems were discovered in bacteria. CRISPR-CasX and CRISPR-CasY were discovered, and they are the most complex proteins discovered to date. In some embodiments, the base editors described herein In the TA system, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. Other RNA-induced DNA-binding proteins may also be used as the protein (napDNAbp). It should be understood that is within the scope of this disclosure.

[0243] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The gram-capable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. The pDNAbp is a CasY protein. In some embodiments, the napDNAbp is a naturally occurring At least 85%, at least 90%, at least 91%, or at least at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least amino acid sequences that are 97%, at least 98%, at least 99%, or at least 99.5% identical to each other In some embodiments, the programmable nucleotide binding protein comprises a sequence of In some embodiments, the programmable CasX or CasY protein is a naturally occurring CasX or CasY protein. The functional nucleotide binding protein can be any of the CasX or CasY proteins described herein. At least 85%, at least 90%, at least 91%, at least 92%, or less At least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 9 8%, at least 99%, or at least 99.5% identical amino acid sequence. It should be understood that CasX and CasY derived from can also be used in accordance with the present disclosure. .

[0244] Exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN 87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) G The amino acid sequence of N = SiH_0402 PE=4 SV=1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG.

[0245] Exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS = Sulfolob The amino acid sequence of Streptococcus islandicus (REY15A strain) GN=SiRe_0771 PE=4 SV=1) is as follows: R: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.

[0246] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGEVHAAEQAALNIARSWLFLNSNSTEFKSY KSGKQPFVGAWQAFYKRRLKEVWKPNA

[0247] An exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein The amino acid sequence of the protein CasY (an uncultured Parcubacteria group bacterium) is: MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI.

[0248] Cas9 nuclease contains two functional endonuclease domains, RuvC and HNH Upon binding to its target, Cas9 undergoes a conformational change that transforms the nuclease domain The end result of Cas9-mediated DNA cleavage is that the target The resulting double-strand break (DSB) occurs in the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). DSBs are repaired by one of two general repair pathways: (1) an efficient but error-prone repair pathway; (2) the less efficient but more fidelity homology-directed repair pathway; Return (HDR) path.

[0249] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) is a function of any It can be calculated by a convenient method. For example, in some embodiments, the efficiency is: It can be expressed as a percentage of successful HDR. For example, Cleavage products were generated using enzyme assays, and the ratio of product to substrate was used to calculate percentages. For example, the newly incorporated constraints that result from the success of HDR can be calculated. A test nuclease enzyme can be used that directly cleaves the DNA containing the sequence. The more substrates used, the higher the rate of HDR (the more efficient HDR). As a general example, the HDR percentage can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (e.g., (b+c) / (a+b+c) where "a" is the band intensity, and "b" and "c" are cleavage products).

[0250] In some embodiments, efficiency can be expressed as the success rate of NHEJ. For example, T7 endotoxin Nuclease I assay was used to generate cleavage products, and the ratio of product to substrate was used to characterize NHEJ. The percentage of T7 endonuclease I that is expressed in wild-type and exonuclease-specific fragments can be calculated. NHEJ involves the insertion of small random insertions or deletions at the initial break site. mismatches resulting from hybridization of nucleotides (leading to indels) The more breaks there are, the higher the rate of NHEJ (high efficiency of NHEJ). As an illustrative example, the rate (percentage) of NHEJ can be calculated using the formula (1-(1-(b+c) / (a +b+c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate. and "b" and "c" are cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6): 1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308).

[0251] The NHEJ repair pathway is the most active repair mechanism, and small nucleotide insertion or The random nature of NHEJ-mediated DSB repair is due to the frequent occurrence of deletions (indels). Cell populations expressing NA or guide polynucleotides yield a diverse array of mutations In most embodiments, NHEJ involves small fragments of the target DNA, which has important practical implications. This results in a significant indel within the open reading frame (ORF) of the target gene. Amino acid deletions, insertions, or frameshift mutations resulting in premature stop codons occur in The ideal end result is a loss-of-function mutation in the target gene.

[0252] NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, but Homologous directed repair (HDR) is the process of repairing a single nucleotide change with a gene, such as the addition of a fluorophore or tag. HDR can be used to generate specific nucleotide changes ranging from large insertions to To utilize gRNA for gene editing, a DNA repair template containing the desired sequence is introduced into the gRNA ( and Cas9 or Cas9 nickase into a cell type of interest. The repair template is designed to identify the desired edit and the target immediately upstream and downstream (left). The length of each homologous arm is The length may depend on the magnitude of the change introduced, with larger insertions requiring longer homologous sequences. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a The efficiency of HDR depends on the amount of Cas9, gRNA, and exogenous gene. Even in cells expressing the repair template, the percentage of corrected alleles is generally low (less than 10% corrected alleles). Because DR occurs between the S and G2 phases of the cell cycle, synchronizing cells can improve the efficiency of HDR. Chemical or genetic inhibition of genes involved in NHEJ can also enhance HDR. The frequency may be increased.

[0253] In some embodiments, the Cas9 is a modified Cas9. There may be additional sites of partial homology throughout the genome. These targets must be considered when designing gRNAs. Optimizing gRNA design In addition, the specificity of CRISPR can be increased by modifying Cas9. RuvC amplifies double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. The Cas9 nickase, a D10A mutant of SpCas9, generates one nuclease domain. This preserves the DNA fragment and generates DNA nicks rather than DSBs. Gene editing can also be combined with the nickase system.

[0254] In some embodiments, the Cas9 is a variant Cas9 protein. The peptides differ by a single amino acid compared to the amino acid sequence of the wild-type Cas9 protein. In some examples, the amino acid sequence may be a sequence having a sequence similar to that of the amino acid sequence of the target gene (e.g., a sequence having deletions, insertions, substitutions, or fusions). In this case, the variant Cas9 polypeptide reduces the nuclease activity of the Cas9 polypeptide. have amino acid changes (e.g., deletions, insertions, or substitutions) that result in Thus, the variant Cas9 polypeptides do not possess the nuclease activity of the corresponding wild-type Cas9 protein. have less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the In some embodiments, the variant Cas9 protein has no substantial nuclease activity. The target Cas9 protein may be a variant Cas9 protein that does not have substantial nuclease activity. If it is a protein, it may be referred to as "dCas9."

[0255] In some embodiments, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein may be a wild-type Cas9 protein, e.g., a wild-type Cas9 less than about 20%, less than about 15%, less than about 10%, less than about 5% of the endonuclease activity of the protein; It shows less than about 1%, or less than about 0.1%.

[0256] In some embodiments, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. capable of cleaving the non-complementary strand of the double-stranded guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence For example, variant Cas9 proteins contain mutations that reduce the function of the RuvC domain (A). As a non-limiting example, in some embodiments, The modified Cas9 protein is D10A (aspartic acid to alanine at amino acid position 10). and thus can cleave the complementary strand of the double-stranded guide target sequence, but The variant Cas9 protein has a reduced ability to cleave the non-complementary strand of the target sequence (hence, When a protein cleaves a double-stranded target nucleic acid, it creates a single-strand break (SSB) instead of a double-strand break (DSB). (See, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) ).

[0257] In some embodiments, the variant Cas9 protein is a non-correlated double-stranded guide target sequence. Able to cleave the complementary strand, but has a reduced ability to cleave the complementary strand of the guide target sequence For example, variant Cas9 proteins contain mutations that reduce the function of the HNH domain (amino acid residues). Non-limiting examples include: In some embodiments, the variant Cas9 protein is H840A (an amino acid residue at position 840). The guide has a histidine to alanine mutation, which prevents cleavage of the non-complementary strand of the guide target sequence. can cleave the complementary strand of the guide target sequence, but with a reduced ability to cleave the complementary strand of the guide target sequence (and thus When this variant Cas9 protein cleaves the double-stranded guide target sequence, it generates an SSB instead of a DSB. Such a Cas9 protein can be used to target a guide target sequence (e.g., a single-stranded guide target sequence). ) but binds to a guide target sequence (e.g., a single-stranded guide target sequence). It retains the ability to merge.

[0258] In some embodiments, the variant Cas9 protein amplifies the complementary strand of the double-stranded target DNA and As a non-limiting example, some implementations have a reduced ability to cleave both the complementary and non-complementary strands. In this form, the variant Cas9 protein harbors both the D10A and H840A mutations. As a result, the polypeptide has the ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). have reduced ability to bind to target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA do.

[0259] As another non-limiting example, in some embodiments, the variant Cas9 protein is 76A and W1126A mutations, resulting in the polypeptide having reduced ability to cleave target DNA. These Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). Reduced capacity, but retains the ability to bind to target DNA (e.g., single-stranded target DNA) .

[0260] As another non-limiting example, in some embodiments, the variant Cas9 protein is P4 75A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, resulting in Such Cas9 proteins have a reduced ability to cleave target DNA. a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but It retains the ability to bind to DNA.

[0261] As another non-limiting example, in some embodiments, the variant Cas9 protein is H8 40A, W476A, and W1126A mutations, so that the polypeptide cleaves the target DNA. These Cas9 proteins have a reduced ability to target DNA (e.g., single-stranded target DNA). have a reduced ability to cleave target DNA but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein The protein harbors H840A, D10A, W476A, and W1126A mutations, resulting in a polypeptide Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., The target DNA (e.g., single-stranded target DNA) has a reduced ability to cleave, but In some embodiments, the variant Cas9 retains the ability to bind to the Cas9 HNH domain. The catalytic His residue at main position 840 is restored (A840H).

[0262] As another non-limiting example, in some embodiments, the variant Cas9 protein is H8 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, resulting in The polypeptide has a reduced ability to cleave target DNA. Reduced ability to cleave target DNA (e.g., single-stranded target DNA), but As another non-limiting example, in some embodiments, The variant Cas9 proteins are D10A, H840A, P475A, W476A, N477A, ​​D1125A, and W1126A. and D1127A mutation, resulting in the polypeptide having reduced ability to cleave target DNA. Such Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). have a reduced ability to bind to target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA In some embodiments, the variant Cas9 protein contains the W476A and W1126A mutations. or variant Cas9 proteins containing P475A, W476A, N477A, ​​D1125A, W1126 A, and variant Cas9 proteins harboring D1127A mutations efficiently target PAM sequences. Therefore, in such cases, the variant Cas9 protein does not bind. When used in this method, the method does not require a PAM sequence. When such variant Cas9 proteins are used in conjugation methods, the method involves the use of guide RNA. However, this method can be performed in the absence of a PAM sequence (thus limiting the specificity of binding). (The targeting segment of the guide RNA is responsible for this). residues can be mutated (i.e., inactivating one or the other nuclease moiety) Non-limiting examples include residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A98 4, D986, and / or A987 can be modified (i.e., substituted). Mutations other than recombination are also suitable.

[0263] In some embodiments, variant Cas9 proteins with reduced catalytic activity (e.g., If the Cas9 protein is D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986 , and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A , H982A, H983A, A984A, and / or D986A) interacts with the guide RNA As long as it retains its ability to act, it can still bind to the target DNA in a site-specific manner ( (Because it is still guided to the target DNA sequence by the guide RNA).

[0264] In some embodiments, the variant Cas protein is spCas9, spCas9-VRQR, spCas9-V RER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9- It can be LRVSQL.

[0265] In some embodiments, the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1 A modified SpC containing 332A, R1335E, and T1337R and with specificity for the modified PAM 5'-NGC-3' as9 (SpCas9-MQKFRAER) was used.

[0266] Alternatives to S. pyogenes Cas9 include Cpf1 family-derived cleavage enzymes that exhibit cleavage activity in mammalian cells. RNA-guided endonucleases derived from Prevotella and Francisella may also be included. The current CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. 1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This mechanism is found in the bacteria Prevotella and Francisella. The Cpf1 gene is associated with the CRISPR locus. It contains an endonuclease that uses a guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and encodes the CRISPR / Cas9 Overcoming some of the limitations of the system: Unlike Cas9 nuclease, Cpf1-mediated DNA fragmentation The result of cleavage is a double-strand break with a short 3' overhang. The cleavage pattern of the (red) fragments allows for directional genetic modification, similar to traditional restriction enzyme cloning. This opens up the possibility of introducing new genes, which may increase the efficiency of gene editing. Like its variants and orthologues, Cpf1 is a CRISPR-targetable gene. Expanding the number of SpCas9-preferred NGG PAM sites to AT-rich regions or AT-rich genomes lacking these sites The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I, and the subsequent helix. The Cpf1 protein contains the RuvC-II domain, RuvC-II, and zinc finger-like domain of Cas9. Cpf1 has a RuvC-like endonuclease domain similar to the vC domain. It lacks an endonuclease domain, and the N-terminus of Cpf1 is the alpha helix recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain organization makes Cpf1 functionally unique and class 2. We showed that the Cpf1 locus is classified as a type V CRISPR system. It encodes Cas1, Cas2, and Cas4 proteins that are more similar to types I and III than to stem Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA) and therefore does not function as a CRISPR promoter. Cpf1 is not only smaller than Cas9, but also requires a smaller sgRNA molecule (C This is beneficial for genome editing because Cas9 has about half the number of nucleotides. In contrast to the G-rich PAMs it targets, the Cpf1-crRNA complex binds to the protospacer-adjacent motif Cpf1 cleaves target DNA or RNA by identifying the 5'-YTN-3' PAM. introduces sticky end-like DNA double-strand breaks with five-nucleotide overhangs do.

[0267] [Cas12 domain, a nucleobase editor] Typically, microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multisubunit effector complexes, while class 2 systems have The system has a single protein effector. For example, Cas9 and Cpf1 are different. Although they are of different types (types II and V, respectively), they are class 2 effectors. In addition to Cpf1, class 2, type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, and Ca Also included are s12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. For example, Shmakov et al., “Discovery and Functional Characterization of Diverse Cla ss 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 60(3): 385-397; Makarova et a l., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1(5): 325-336; and Yan et al., “Functionally Diverse Typ e V CRISPR-Cas Systems,” Science, 2019 Jan. 4; 363: 88-91 (see these (The references are incorporated herein by reference in their entirety.) V-type Cas proteins include RuvC ( (or RuvC-like) endonuclease domain. Production of mature CRISPR RNA (crRNA). Generally, Cas12b / C2c1 is tracrRNA-independent, but for example, Cas12b / C2c1 requires tracrRNA for crRNA production. A. Cas12b / C2c1 depends on both crRNA and tracrRNA for DNA cleavage .

[0268] Nucleic acid programmable DNA binding proteins contemplated by the present invention include those classified as Class 2, Type V. Cas proteins (Cas12 proteins) are included. Cas class 2 and V proteins are limited. Examples of genes that do not function include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, and Cas12e / CasX. , Cas12g, Cas12h, and Cas12i, their homologs or modified forms. In this case, the Cas12 protein may also be a Cas12 nuclease, a Cas12 domain, or a Cas12 tag. In some embodiments, the Cas12 proteins of the present invention may be referred to as a protein domain. is interrupted by an internally fused protein domain, e.g., a deaminase domain. It contains an amino acid sequence.

[0269] In some embodiments, the Cas12 domain is a nuclease-inactive Cas12 domain or In some embodiments, the Cas12 domain exhibits nuclease activity. For example, the Cas12 domain is a domain that binds to one side of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). In some embodiments, the Cas12 domain may be a Cas12 domain that nicks the target strand. In some embodiments, the C The as12 domain may comprise at least one amino acid sequence selected from the group consisting of: 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, At least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the amino acid sequence is at least 99%, or at least 99.5%, identical to the amino acid sequence of the target gene. In the present specification, the Cas12 domain is 1 to 100% identical to any one of the amino acid sequences described herein. , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21 , 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43 44, 45, 46, 47, 48, 49, 50, or more amino acid sequences In some embodiments, the Cas12 domain comprises the amino acid sequence described herein. At least 10, at least 15, at least 20, at least 30 compared to any one At least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, At least 350, at least 400, at least 500, at least 600, at least 700, at least at least 800, at least 900, at least 1000, at least 1100, or at least 1200 It includes amino acid sequences that have identical consecutive amino acid residues.

[0270] In some embodiments, proteins comprising fragments of Cas12 are provided. For example, some In some embodiments, the protein comprises one of the following two Cas12 domains: (1) (2) the gRNA binding domain of Cas12; and (3) the DNA cleavage domain of Cas12. In some embodiments, Proteins containing Cas12 or fragments thereof are referred to as "Cas12 variants." The Cas12 variant shares homology with Cas12 or a fragment thereof. For example, the Cas12 variant may be a variant of wild-type Cas12. At least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least are also about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, the Cas12 variant has 1, 2, 3, 4, 5, 6, 7, 8 or more amino acids compared to wild-type Cas12. , 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28 , 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 In some embodiments, the Cas1 may have 49, 50, or more amino acid changes. Two variants include a fragment of Cas12 (e.g., the gRNA binding domain or the DNA cleavage domain), The fragment is at least about 70% identical, at least about 80% identical to the corresponding fragment of wild-type Cas12. , at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, the fragment is at least about 99.9% identical to the corresponding wild-type Cas12. At least 30%, at least 35%, at least 40%, at least 45%, at least At least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% , at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96% , at least 97%, at least 98%, at least 99%, or at least 99.5%. In some embodiments, the fragment is at least 100 amino acids in length. In this embodiment, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600 , 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids in length.

[0271] In some embodiments, Cas12 is provided with one or more amplified polypeptides that alter Cas12 nuclease activity. The amino acid sequence of Cas12 may correspond in part or in whole to or contain a mutation. Such mutations include, for example, the Ruv of Cas12. In some embodiments, the wild-type C At least about 70% identical to as12, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or provides variants or homologs of Cas12 that are at least about 99.9% identical. In some embodiments, the amino acid sequence may be about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, or the like. amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more , short, or long variants of Cas12 are provided.

[0272] In some embodiments, the Cas12 fusion protein as provided herein comprises a Cas12 The full-length amino acid sequence of the protein, for example, one of the Cas12 sequences provided herein. However, in other embodiments, the fusion proteins as provided herein contain the full-length Cas12 sequence. Exemplary amino acid sequences of suitable Cas12 domains are: Additional suitable sequences for Cas12 domains and fragments are provided herein and will be apparent to those skilled in the art. It would be.

[0273] Generally, class 2, V-type Cas proteins contain a single functional RuvC endonuclease domain. (See, e.g., Chen et al., "CRISPR-Cas12a target binding unleashes See “Scriminate single-stranded DNase activity,” Science 360:436-439 (2018) In some cases, the Cas12 protein is a variant Cas12b protein (Strec ker et al., Nature Communications, 2019, 10(1): Art. No.: 212). In embodiments, the variant Cas12 polypeptide has the amino acid sequence of a wild-type Cas12 protein: differs in 1, 2, 3, 4, 5 or more amino acids (e.g., deletions, insertions, substitutions, fusions) when compared to In some instances, the variant Cas12 polypeptide has an amino acid sequence , amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the activity of the Cas12 polypeptide. For example, in some instances, the variant Cas12 has a corresponding wild-type Cas12b tag. Less than 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the protein It is a Cas12b polypeptide with cytoplasmic enzyme activity.

[0274] In some cases, variant Cas12b proteins have reduced nickase activity. For example, the variant Cas12b protein may have a structural identity that is less than about 20%, less than about 15%, or both, of the wild-type Cas12b protein. , less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% nickase activity.

[0275] In some embodiments, the Cas12 protein includes a Cas12 protein that is active in mammalian cells. RNA-guided endonucleases from the 12a / Cpf1 family are included. Prevotella and F CRISPR derived from Rancisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease in the class II CRISPR / Cas system. This adaptive immune mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is involved in CRI Endoplasmic reticulum (ER) associates with the SPR locus and uses guide RNA to find and cleave viral DNA. Cpf1 encodes a nuclease that is smaller and simpler than Cas9. Unlike Cas9 nuclease, it overcomes some of the limitations of the CRISPR / Cas9 system. The result of Cpf1-mediated DNA cleavage is a double-strand break with a short 3' overhang. The staggered cleavage pattern allows for directional gene transfer similar to traditional restriction enzyme cloning. This could open up new possibilities for gene editing, which could increase the efficiency of gene editing. Similar to Cas9 variants and orthologues, Cpf1 also contains sites that can be targeted by CRISPR. The number of nucleotides was increased to AT-rich regions or AT-rich genomes lacking the NGG PAM site preferred by SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I, followed by a helix. The Cpf1 protein contains the RuvC-II domain and a zinc finger-like domain. Furthermore, Cpf1 possesses a RuvC-like endonuclease domain similar to the RuvC domain of Unlike Cas9, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 is the alpha domain of Cas9. The Cpf1 CRISPR-Cas domain construction is a key step in Cpf1's functional unity. The Cpf1 locus is a class II, type V CRISPR system. It encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III systems than to type III systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA) and therefore does not function as a CRISPR Only PR (crRNA) is required. Cpf1 is not only smaller than Cas9, but also smaller than sgR. It also contains an NA molecule (roughly half the number of nucleotides of Cas9), which is beneficial for genome editing. The Cpf1-crRNA complex targets protozoa, in contrast to the G-rich PAM targeted by Cas9. Identification of the pacer adjacent motif 5'-YTN-3' or 5'-TTTN-3' allows target DNA or RNA to be identified. After PAM identification, Cpf1 cleaves the attached fragment with a 4- or 5-nucleotide overhang. Introduces end-like DNA double-strand breaks.

[0276] In some embodiments of the invention, the vectors are mutated with respect to the corresponding wild-type enzyme. The mutated CRISPR enzyme encodes a target gene containing the target sequence. Cas12 lacks the ability to cleave one or both strands of a target polynucleotide. At least about 50%, 60% of the target Cas12 polypeptide (e.g., Cas12 from Bacillus hisashii) , 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity The term "Cas12" may refer to a polypeptide having sequence identity and / or sequence homology. target Cas12 polypeptides (e.g., derived from Bacillus hisashii (BhCas12b), derived from Bacillus species V3-13) (BvCas12b), and Alicyclobacillus acidiphilus (AaCas12b). Approximately 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and may refer to a polypeptide with 100% sequence identity and / or sequence homology. may be a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof. It may also refer to wild-type or modified forms of the Cas12 protein, which may contain amino acid changes such as cleavage. good.

[0277] [Nucleic acid programmable DNA binding protein] Some embodiments of the present disclosure provide nucleic acid-programmable DNA binding proteins. The present invention provides a fusion protein comprising a domain, which specifically binds a protein such as a base editor. The present invention can be used to guide a specific nucleic acid (e.g., DNA or RNA) sequence. In this study, the fusion protein was constructed by combining a nucleic acid-programmable DNA-binding protein domain and a deaminase Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Non-limiting examples of Cas enzymes include Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas 8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas1 2a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas 12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Cs n1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1 , Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa 5. Type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins tar proteins, CARF, DinG, their homologs, or modified or engineered versions thereof Other nucleic acid programmable DNA binding proteins are also included, as specifically listed in this disclosure. Although not all of the above may be applicable, it is within the scope of this disclosure. cation and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally divers e type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / See science.aav7271 (the entire contents of each of which are incorporated herein by reference). ).

[0278] An example of a nucleic acid programmable DNA binding protein with a PAM specificity different from Cas9 is Pr Clustered Regularly Interspaced Short Palindromes from evotella and Francisella1 Similar to Cas9, Cpf1 is a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with characteristics distinct from those of Cas9. , a single RNA-guided endonuclease lacking tracrRNA and adjacent to a T-rich protospacer Furthermore, Cpf1 utilizes a junction motif (TTN, TTTN, or YTN) to separate DNA into a staggered DNA duplex. Of the 16 Cpf1 family proteins, Acidaminococcus and Lachnospira Two enzymes from ceae have been shown to have efficient genome editing activity in human cells. The Cpf1 protein is known in the art and has been previously described, for example, by Yamano et al., "C rystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, pp. 949-962, the entire contents of which are incorporated herein by reference. do.

[0279] The guide nucleotide sequence can be used as a programmable DNA binding protein domain. Nuclease-inactive Cpf1 (dCpf1) variants that are useful in the compositions and methods of the invention. The Cpf1 protein is a RuvC-like endonuclease that is similar to the RuvC domain of Cas9. The N-terminus of Cpf1 is the same as that of Cas9. It does not have an alpha helix recognition lobe. Zetsche et al., Cell, 163, 759-771, 2015 (see (incorporated herein by reference) demonstrated that the RuvC-like domain of Cpf1 is responsible for cleavage of both DNA strands. showed that inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. Mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 In some embodiments, the dCpf1 of the present disclosure is D91 7A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006 Any mutation that inactivates the RuvC domain of Cpf1, including the mutation corresponding to A / D1255A. It will be understood that, for example, substitution mutations, deletions, or insertions may be used in accordance with the present disclosure. I want to be.

[0280] In some embodiments, the nucleic acid promoter of any of the fusion proteins provided herein The gram-competent DNA-binding protein (napDNAbp) may be the Cpf1 protein. In some embodiments, the Cpf1 protein is Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is nuclease-inactive Cpf1 (dCpf1). In some embodiments, Cpf1, nCpf1, or dCpf1 is at least as similar to the Cpf1 sequences disclosed herein. At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or or an amino acid sequence at least 99.5% identical to dCpf1. In some embodiments, dCpf1 is , at least 85%, at least 90%, at least 91% relative to the Cpf1 sequences disclosed herein %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical The sequence includes D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or containing the mutations corresponding to D917A / E1006A / D1255A. Cpf1 from other bacterial species also contains the mutations It should be understood that the same may be used in accordance with the disclosure. 【02...

Claims

1. 1. A pharmaceutical composition for use in a method for treating alpha-1 antitrypsin deficiency in a subject comprising a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency, the pharmaceutical composition comprising lipid nanoparticles comprising a guide RNA and an mRNA encoding a base editor; the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and an adenosine deaminase domain, and the SNP associated with alpha-1 antitrypsin deficiency results in expression of an alpha-1 antitrypsin polypeptide having a lysine at amino acid position 342 with reference to SEQ ID NO: 21, where the first E in SEQ ID NO: 21 is position 1; and the guide RNA targets the base editor to result in modification of the SNP associated with alpha-1 antitrypsin deficiency and comprises the nucleotide sequence AUCGACAAGAAAGGGACUGA. Pharmaceutical compositions.

2. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD 2. The pharmaceutical composition of claim 1, comprising an amino acid sequence having at least 85% identity to the TadA*7.10 amino acid sequence and further containing an arginine or threonine at amino acid position 147 compared to the TadA*7.10 amino acid sequence.

3. The pharmaceutical composition of claim 2 , wherein the adenosine deaminase domain further comprises an I76Y amino acid modification and / or a Q154S amino acid modification.

4. 3. The pharmaceutical composition of claim 2, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to the TadA*7.10 amino acid sequence.

5. 3. The pharmaceutical composition of claim 2, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 95% identity to the TadA*7.10 amino acid sequence.

6. The guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 2. The pharmaceutical composition of claim 1, comprising:

7. a) the napDNAbp is SpCas9 with specificity for a protospacer adjacent motif comprising the nucleotide sequence 5'-NGC-3'; and / or b) said napDNAbp comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 13 and further comprises the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to SEQ ID NO: 13; The pharmaceutical composition of claim 1.

8. 1. A lipid nanoparticle comprising a guide RNA and mRNA encoding a base editor, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and an adenosine deaminase domain; wherein the guide RNA targets the base editor to result in modification of a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency, the SNP comprising the nucleotide sequence AUCGACAAGAAAGGGACUGA, and wherein the SNP associated with alpha-1 antitrypsin deficiency results in expression of an alpha-1 antitrypsin polypeptide having a lysine at amino acid position 342 with reference to SEQ ID NO: 21, wherein the first E in SEQ ID NO: 21 is position 1.

9. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD The lipid nanoparticle of claim 8, comprising an amino acid sequence having at least 85% identity to the TadA*7.10 amino acid sequence and further containing arginine or threonine at amino acid position 147 compared to the TadA*7.10 amino acid sequence.

10. The guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' The lipid nanoparticle of claim 8, comprising:

11. a) the napDNAbp is SpCas9 with specificity for a protospacer adjacent motif comprising the nucleotide sequence 5'-NGC-3'; and / or b) said napDNAbp comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 13 and further comprises the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to SEQ ID NO: 13; The lipid nanoparticle of claim 8.

12. A pharmaceutical composition comprising the lipid nanoparticles of any one of claims 8 to 11 and a pharmaceutically acceptable carrier, vehicle, or excipient.

13. A guide RNA comprising the nucleic acid sequence AUCGACAAGAAAGGGACUGA.

14. The guide RNA has the following nucleic acid sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' The guide RNA of claim 13, comprising:

15. 15. The guide RNA of claim 13 or 14, wherein the guide RNA comprises 2'-O-methyl RNA and phosphorothioate (PS) internucleotide linkages.

16. 1. A base editor system comprising a guide RNA and an mRNA encoding a base editor, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and an adenosine deaminase domain; wherein the guide RNA targets the base editor to modify a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency, the SNP associated with alpha-1 antitrypsin deficiency being in an alpha-1 antitrypsin polynucleotide and resulting in expression of an alpha-1 antitrypsin polypeptide having a lysine at amino acid position 342 with reference to SEQ ID NO:21, wherein the first E in SEQ ID NO:21 is position 1, and wherein the guide RNA comprises the nucleotide sequence AUCGACAAGAAAGGGACUGA.

17. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD 17. The base editor system of claim 16, comprising an amino acid sequence having at least 85% identity to the TadA*7.10 amino acid sequence and further containing an arginine or threonine at amino acid position 147 compared to the TadA*7.10 amino acid sequence.

18. The guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 17. The base editor system of claim 16, comprising:

19. a) the napDNAbp is SpCas9 with specificity for a protospacer adjacent motif comprising the nucleotide sequence 5'-NGC-3'; and / or b) said napDNAbp comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 13 and further comprises the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to SEQ ID NO: 13; 17. The base editor system of claim 16.