Compositions and methods for treating alpha-1 antitrypsin deficiency

The base editing system corrects A1AD mutations by converting A·T to G·C in alpha-1 antitrypsin polynucleotides, addressing both lung and liver issues in A1AD patients, enhancing alpha-1 antitrypsin production and improving lung and liver health.

JP7861090B2Active Publication Date: 2026-05-18BEAM THERAPEUTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024208066
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2024-11-29
Publication Date
2026-05-18
Estimated Expiration
2040-02-13

AI Technical Summary

Technical Problem

Current treatments for alpha-1 antitrypsin deficiency (A1AD) fail to address both lung pathology and hepatotoxicity effectively, as gene therapy increases liver burden and protein replacement therapy does not correct underlying mutations.

Method used

A base editing system using modified adenosine deaminase (ABE8) and guide RNAs targets specific SNPs in alpha-1 antitrypsin polynucleotides to convert A·T to G·C, correcting harmful mutations and producing functional alpha-1 antitrypsin in hepatocytes.

Benefits of technology

The method effectively corrects A1AD mutations, reducing lung pathology and hepatotoxicity by enhancing alpha-1 antitrypsin production, thereby improving lung elasticity and liver health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861090000062
    Figure 0007861090000062
  • Figure 0007861090000063
    Figure 0007861090000063
  • Figure 0007861090000064
    Figure 0007861090000064
Patent Text Reader

Abstract

To provide a method of treating patients with alpha-1 anti-trypsin deficiency that addresses both lung pathology and liver toxicity.SOLUTION: The present invention features compositions and methods for editing deleterious mutations associated with alpha-1 anti-trypsin (A1AT) deficiency. In particular embodiments, the invention provides methods for correcting mutations in an A1AT polynucleotide using an adenosine deaminase base editor, ABE8, having unprecedented levels of efficiency.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application is U.S. Provisional Patent Application No. 62 / 805,238, filed on February 13, 2019; February 13, 2019 Filing No. 62 / 805,271 on May 23, 2019; Filing No. 62 / 852,224 on May 23, 2019 Filing No. 62 / 852,228 on [date]; Filing No. 62 / 931,722 on November 6, 2019; November 2019 Filing No. 62 / 941,569 on the 27th of the month; and Filing No. 62 / 966,526 on the 27th of January 2020 This is an international PCT application claiming the interests of, and the entire contents of these documents are referred herein by reference. It will be incorporated.

[0002] Embedding by reference All publications, patents, and patent applications referred to herein are the same as, in each individual publication. A patent or patent application is specifically and individually incorporated herein by reference. It is incorporated herein by reference to the same extent as it is incorporated herein by reference. Unless otherwise indicated, this Publications, patents, and patent applications referenced in this specification are incorporated herein by reference in their entirety. To be included. [Background technology]

[0003] In healthy individuals, alpha-1 antitrypsin (A1AT) is produced by hepatocytes in the liver. It is produced and secreted into the systemic circulation, where it functions as a protease inhibitor. A1AT It is a particularly excellent inhibitor of neutrophil elastase, and therefore, tissues and organs such as the lungs It protects against elastin degradation. In patients with alpha-1 antitrypsin deficiency (A1AD), A1 There is a mutation in the gene encoding AT, resulting in a decrease in protein production. .

[0004] As a result, elastin in the lungs is more easily degraded by neutrophil elastase, and over time, the elasticity of the lungs is impaired and it develops into chronic obstructive pulmonary disease (COPD).

[0005] The most common pathogenic A1AT variant is a mutation from guanine to adenine that results in a substitution of glutamic acid with lysine at amino acid 342. This substitution misfolds and polymerizes the protein in hepatocytes, and ultimately toxic aggregates can cause liver injury and cirrhosis. The hepatotoxicity can be addressed by gene knockout (CRISPR / ZFN / TALEN) or gene knockdown (siRNA), but neither approach addresses the lung pathology. The lung pathology can be addressed by protein replacement therapy, but this therapy also cannot address the hepatotoxicity. Gene therapy would also be inappropriate for addressing A1AT gene deficiency. The livers of A1AD patients are already under the burden of severe diseases caused by endogenous A1AT, so gene therapy to increase A1AT in the liver would have a reverse effect.

Summary of the Invention

Problems to be Solved by the Invention

[0006] Therefore, there is a need for a treatment method for A1AD patients that addresses both lung pathology and hepatotoxicity.

Means for Solving the Problems

[0007] As described below, the present invention relates to alpha-1 antitrypsin deficiency (A1AD). The present invention features compositions and methods for editing harmful mutations. In certain embodiments, The present invention aims to correct mutations associated with A1AD at exceptional levels (e.g., >60-70%). Using a modified adenosine deaminase called "ABE8" which has the efficiency and specificity of ) They offer A1AD treatment.

[0008] In one embodiment, the present invention relates to a single nucleotide polymorphism associated with alpha-1 antitrypsin deficiency ( A method for editing alpha-1 antitrypsin polynucleotides containing SNPs, This polynucleotide, along with one or more guide RNAs, and a polynucleotide program Possible DNA-binding domain and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD It is an adenosine deaminase variant that contains a modification at amino acid position 82 or 166. The step of contacting a base editor containing at least one base editor domain The guide RNA includes the base editor and targets (targets) the ALF The present invention provides a method that results in modification of an SNP associated with a-1 antitrypsin deficiency.

[0009] In another aspect, the present invention relates to a single nucleotide polymorphism (SN) associated with alpha-1 antitrypsin deficiency. A method for editing an alpha-1 antitrypsin polynucleotide containing P), 1 One or more guide RNAs, and the following sequences: JPEG0007861090000001.jpg171162 (Bold sequences indicate sequences derived from Cas9, italic sequences indicate linker sequences, and underlined sequences indicate two parts) Polynucleotide programmable DNA binding containing (bipartite) nuclear localization sequence. Domain and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD Contains an adenosine deaminase variant with a modification at amino acid position 82 or 166. The process involves contacting a fusion protein containing at least one base editor domain. The present invention provides the method including the steps described above.

[0010] In another embodiment, the present invention relates to a fusion protein and any of the previously described embodiments. 5'-ACCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCA AC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAA C UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-UCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC U UGAAAAGU GGCACCGAGU CGGUGCUUUU-3' 5'-CGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UU GAAAAAAGU GGCACCGAGU CGGUGCUUUU-3' We provide a base editing system containing a guide RNA containing a more selectable nucleic acid sequence.

[0011] In another embodiment, the following is found within the cell: A base editor or a polynucleo encoding the base editor for the aforementioned cells. It is a DNA-binding domain with a base editor and a polynucleotide programmable DNA-binding domain. The base editor containing the adenosine deaminase described in any of the preceding embodiments, or polynucleotides; and targeting alpha-1 antinucleotides with a base editor. One or more guidepolynes that result in A·T to G·C modification of SNPs associated with trypsin deficiency Cleotide The present invention provides cells or their progenitor cells produced by introducing [a specific substance]. In one embodiment, The cells produced are hepatocytes or their progenitor cells. In another embodiment, the cells are A The cells are derived from subjects with rufa-1 antitrypsin deficiency. In another embodiment, the cells are derived from mammals. These are animal cells or human cells.

[0012] In various embodiments of the above-described model, the gRNA has the nucleic acid sequence 5'- GUUUUAGAGC UAGAAAUAGC AAGUU It further contains AAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3.

[0013] In yet another aspect, the present invention treats alpha-1 antitrypsin deficiency in a subject. A method of treatment, comprising the step of administering cells of any of the above embodiments to the subject, A method is provided. In one embodiment, the cells are self or homogeneous with respect to the subject.

[0014] In yet another aspect, the present invention relates to cells grown from the embodiments and models listed above. or provides isolated cells or cell populations that have been expanded.

[0015] In yet another aspect, the present invention relates to a method for producing hepatocytes, wherein (a) alpha-1 In hepatocytes containing SNPs associated with thitrypsin deficiency, a base editor or the aforementioned A polynucleotide that encodes a base editor, wherein the base editor is a polynucleotide Cleotide programmable nucleotide-binding domain, as well as the above aspects and embodiments. The base comprising an adenosine deaminase variant domain as described in any one of the following Editor or polynucleotide; and introduction of one or more guide polynucleotides. The process includes the step of using one or more guide polynucleotides to select the base editor. Targeting and modifying SNPs associated with alpha-1 antitrypsin deficiency from A·T to G·C The present invention provides the method for bringing about change.

[0016] In various embodiments, hepatocytes are mammalian cells or human cells.

[0017] In other embodiments of the above-described model, the adenosine deaminase variant is at amino acid position 82 This includes modifications to position 166. In other embodiments of the above-described model, an adenosine deaminase barrier The nt includes the V82S modification. In other embodiments of the above embodiment, an adenosine deaminase barrier The formula includes the T166R modification. In other embodiments of the above embodiment, adenosine deaminase variant Ant includes V82S and T166R modifications. In other embodiments of the above aspects, adenosine dea Minase variants include one of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. The above further includes. In other embodiments of the above embodiments, the adenosine deaminase variant is The following modifications: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y1 47R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147 R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and includes I76Y + V82S + Y123H + Y147R + Q154R. In other embodiments of the above configuration, The denosine deaminase variant includes Y147R + Q154R + Y123H. Other embodiments of the above embodiment Morphologically, the adenosine deaminase variant contains Y147R + Q154R + I76Y. In other embodiments of the same, the adenosine deaminase variant is Y147R + Q154R + T166R Includes. In other embodiments of the above embodiments, the adenosine deaminase variant is Y147T + Includes Q154R. In other embodiments of the above embodiments, the adenosine deaminase variant is Y14 Contains 7T + Q154S. In other embodiments of the above embodiment, the adenosine deaminase variant is , containing Y147R + Q154S. In other embodiments of the above embodiment, an adenosine deaminase barrier The nucleotide includes V82S + Q154S. In other embodiments of the above configuration, adenosine deaminase nucleotides are used. The riant includes V82S + Y147R. In other embodiments of the above configuration, adenosine deaminator The zevariant includes V82S + Q154R. In other embodiments of the above embodiment, adenosine deami The enzyme variant includes V82S + Y123H. In other embodiments of the above embodiment, adenosine The aminase variant includes I76Y + V82S. In other embodiments of the above aspects, adenosy The endoaminase variant includes V82S + Y123H + Y147T. In other embodiments of the above-described model The adenosine deaminase variant contains V82S + Y123H + Y147R. In this embodiment, the adenosine deaminase variant comprises V82S + Y123H + Q154R. In other embodiments of the above-described model, the adenosine deaminase variant is Y123H + Y147R + Includes Q154R + I76Y. In other embodiments of the above embodiment, an adenosine deaminase variant This includes V82S + Y123H + Y147R + Q154R. In other embodiments of the above configuration, adenosine The aminase variant includes I76Y + V82S + Y123H + Y147R + Q154R. In this embodiment, the adenosine deaminase variants are 149, 150, 151, 152, 153, 154 Includes a C-terminal deletion beginning with a residue selected from the group consisting of 155, 156, and 157. In other embodiments of the described aspect, the base editor domain contains single V82S and T166R. Contains adenosine deaminase variant. In other embodiments of the above embodiment, base eddy The tar domain is the wild-type adenosine deaminase domain and adenosine deaminase Includes a variant. In other embodiments of the above embodiment, the adenosine deaminase variant is Modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. It also includes. In other embodiments of the above-described set, the base editor domain is the TadA7.10 domain. and an adenosine deaminase variant. In other embodiments of the above embodiment, adenosine Syntheaminase variants include Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. Further includes modifications selected from the group. In other embodiments of the above embodiment, a base editor This includes the TadA7.10 domain, as well as Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y 147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H Selected from the group consisting of + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R The present invention contains an adenosine deaminase variant that includes the modified form. The base editor is a fragment of which has the following sequence or adenosine deaminase activity. ABE8 containing, or essentially composed of, such sequences or fragments:MSEVEFSHEYWM RHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAG AMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD. above In other embodiments of the described aspect, the adenosine deaminase variant is compared to full-length ABE8. , N of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 Includes truncated ABE8, in which terminal amino acid residues are lost. Other embodiments of the above. In this embodiment, the adenosine deaminase variant is 1, 2, 3, 4 compared to full-length ABE8. , 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acids Includes partially excised ABE8, where residues are missing.

[0018] In another embodiment of the above-described aspect, A at SNPs associated with alpha-1 antitrypsin deficiency. The modification from T to G·C involves the addition of glutamic acid in the alpha-1 antitrypsin polypeptide. Convert to din. In other embodiments of the above configuration, alpha-1 antitrypsin deficiency is associated with The associated SNP is an alpha-1 antitrypsin polypeptide containing lysine at amino acid position 342. This results in the expression of alpha-1 antitrypsin deficiency. In other embodiments of the above model, it is associated with alpha-1 antitrypsin deficiency. In the SNP, glutamic acid is replaced with lysine. In other embodiments of the above configuration, alpha- Regarding the A·T to G·C modification of SNPs associated with antitrypsin deficiency, cell selection In other embodiments of the above-described model, the polynucleotide programmable DNA-binding domain is Modified Staphylococcus aureus Cas9(SaCas9), Streptococcus thermophilus 1 Cas9(St1 Cas9, modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In other embodiments of the above-described model, the polynucleotide programmable DNA-binding domain is modified Having specificity for variant protospacer adjacent motifs (PAMs) or non-G PAMs, It includes a variant of SpCas9. In other embodiments of the above aspects, the modified PAM includes the nucleic acid sequence 5'-NGC It has specificity for -3'. In other embodiments of the above embodiment, the modified SpCas9 has amino acid position Exchange, D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or This includes the corresponding amino acid substitution. In other embodiments of the above embodiments, polynucleotide pro The Gram-capable DNA-binding domain is either nuclease-inactive or a nickasase variant. In other embodiments of the above-described model, the nickerse variant is amino acid substitution D10A or Includes corresponding amino acid substitutions. In other embodiments of the above aspects, the base editor is zinc Further comprising a finger domain. In other embodiments of the above embodiment, adenosine deaminer Zedmain is capable of deaminating adenine in deoxyribonucleic acid (DNA). In other embodiments of the above configuration, one or more guide RNAs are CRISPR RNA (crRNA) and It contains transencoded small RNA molecules (tracrRNA), and the crRNA is alpha-1 antitri Nuclei complementary to alpha-1 antitrypsin nucleic acid sequences containing SNPs associated with psin deficiency Includes an acid sequence. In other embodiments of the above embodiment, a base editor and one or more guide points The nucleotides form a complex in the cell. In other embodiments of the above configuration, the base ed The IV contains SNPs associated with alpha-1 antitrypsin deficiency. A complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the trypsin nucleic acid sequence. It forms.

[0019] In another embodiment, to treat alpha-1 antitrypsin deficiency (A1AD) in a subject A method in which the target is an adenosine dea inserted into a Cas9 or Cas12 polypeptide. A fusion protein containing a minase variant, or a port encoding such a fusion protein. renucleotides; and single nucleotides associated with A1AD that target this fusion protein. One or more guide polynucleotides are introduced that result in the modification of a polymorphism (SNP) from A·T to G·C. The present invention provides a method comprising the step of administering a drug and thereby treating A1AD in a subject.

[0020] In another aspect, a method for treating alpha-1 antitrypsin deficiency (A1AD) in a subject. and adenosine base editor, ABE8, or a port encoding the base editor A dinucleotide in which ABE8 is inserted into a Cas9 or Cas12 polypeptide. The base editor or polynucleotide comprising a ndeaminase variant; and Targeting ABE8 to bring about a change from A·T to G·C in SNPs associated with A1AD, one or more The process involves administering the above guide polynucleotide to treat A1AD in the subject. The present invention provides the method including the above.

[0021] In the embodiment of the method described above, ABE8 is ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8 .5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, A BE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8 .20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, ABE8.24-m, ABE8.1-d, ABE8.2-d, ABE8.3-d , ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11 -d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d Select from ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d. In the embodiment of the above method, the adenosine deaminase variant is: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD The amino acid sequence comprises; this amino acid sequence comprises at least one modification. In one embodiment, The adenosine deaminase variant contains modifications at amino acid positions 82 and / or 166. In one embodiment, at least one modification is: V82S, T166R, Y147T, Y147R, Q154S, Y12 Includes 3H and / or Q154R.

[0022] In one embodiment of the above method, the adenosine deaminase variant is modified in the following combinations: Waseda: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V 82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y14 7R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q1 54R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I7 Includes one of 6Y + V82S + Y123H + Y147R + Q154R. In one embodiment of the above method, adenos The endoaminase variants are TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, Ta dA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13 , TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, T adA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, adenos The endoaminase variants are 149, 150, 151, 152, 153, 154, 155, 156, and 157. This includes a C-terminal deletion beginning with a residue selected from the following group. In one embodiment, adenosine The deaminase variant contains the TadA*8 adenosine deaminase variant domain. It is a monomer of denosine deaminase. In one embodiment, adenosine deaminase The minase variant contains the wild-type adenosine deaminase domain and TadA*8 adenosine. Heterodimer of adenosine deaminase containing the deaminase variant domain (rodimer). In one embodiment, the adenosine deaminase variant is TadA domain Adenosine deaminase containing the TadA*8 adenosine deaminase variant domain It is a zeheterodimer. In one embodiment of the above method, A·T to G at SNPs associated with A1AD The modification to C changes the glutamic acid at amino acid position 342 to lysine. One example of the above method. In the application morphology, the SNP associated with A1AD is alpha-1 antitriglyceride, which has lysine at amino acid position 342. This results in the expression of a psin polypeptide. In one embodiment of the above method, alpha-1 antigen SNPs associated with psin deficiency substitute glutamate with lysine.

[0023] In one embodiment of the above method, the adenosine deaminase variant is Cas9 or Cas12 Flexible loops, alpha-helical regions, unstructured parts, or solvent access in polypeptides. It is inserted into the possible portion. In one embodiment of the above method, adenosine deaminase varian The 't' is adjacent to the N-terminal and C-terminal fragments of the Cas9 or Cas12 polypeptide.

[0024] In one embodiment of the above method, the fusion protein or ABE8 has the structure NH2-[Cas9 or Cas12 [N-terminal fragment of polypeptide]-[adenosine deaminase variant]-[Cas9 or Cas12 poly] The peptide contains the C-terminal fragment []-COOH, where each example of "]-[" is an arbitrary linker. In the application form, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment is Cas9 or Cas12 polypeptide. It includes a portion of the flexible loop. In one embodiment, the flexible loop is adenosine deamin. - When a ze variant deaminates a target nucleic acid base, the amino acids in the vicinity of the target nucleic acid base... Includes.

[0025] In one embodiment of the above method, the method involves deamination of an A1AD-related SNP target nucleic acid base. The process further includes administering a guide nucleic acid sequence to achieve the goal. One embodiment of the above method. In this state, deamination of SNP target nucleic acid bases is performed by replacing the target nucleic acid base with a wild-type nucleic acid base or a non- Substitution with wild-type nucleic acid bases and deamination of target nucleic acid bases alleviates the symptoms of A1AD. In one embodiment of the method, deamination of an A1AD-associated SNP is performed by substituting glutamic acid with lysine. do.

[0026] In one embodiment of the above method, the target nucleic acid base is PA in the target polynucleotide sequence. It is located 1 to 20 nucleic acid bases away from the M sequence. In one embodiment, the target nucleic acid bases are 2 to 12 of the PAM sequence. It is upstream of the nucleic acid base. In one embodiment of the above method, the N-terminus of the Cas9 or Cas12 polypeptide The fragment or C-terminal fragment binds to the target polynucleotide sequence. In certain embodiments, N The terminal or C-terminal fragment contains a RuvC domain; the N-terminal or C-terminal fragment contains an HNH domain. Does it contain a domain?; Does neither the N-terminal nor the C-terminal fragment contain an HNH domain? Alternatively, neither the N-terminal nor the C-terminal fragment contains the RuvC domain. In one embodiment... Cas9 or Cas12 polypeptides have partial or complete deletions in one or more structural domains. It contains a deaminase at the site of a partial or complete deletion of a Cas9 or Cas12 polypeptide. It is inserted. In certain embodiments, the deletion is within the RuvC domain; the deletion is within the HNH domain. If present or deleted, it bridges the RuvC domain and the C-terminal domain.

[0027] In one embodiment of the above method, the fusion protein or ABE8 contains a Cas9 polypeptide. In this embodiment, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphyloc occus aureus Cas9(SaCas9), Streptococcus thermophilus 1 Cas9(St1Cas9), and This is a variant. In one embodiment, the Cas9 polypeptide has the following amino acid sequence ( Includes: Cas9 reference array: JPEG0007861090000002.jpg177162 (Single underline: HNH domain; Double underline: RuvC domain; (Cas9 reference sequence), or its pair) (corresponding region). In certain embodiments, the Cas9 polypeptide is located in the Cas9 polypeptide reference sequence. Numbered as such, it includes amino acids 1017-1069, or the deletion of the corresponding amino acid; Cas9 The polypeptide is numbered from amino acids 792 to 872 in the Cas9 polypeptide reference sequence, This includes the deletion of the corresponding amino acid; or the Cas9 polypeptide is Cas9 polypeptide In the numbering of the reference sequence, the deletion of amino acids 792-906, or the corresponding amino acids, Includes. In one embodiment of the above method, the adenosine deaminase variant is Cas9 polypept It is inserted into the flexible loop of the cyd. In one embodiment, the flexible loop is in the Cas9 reference sequence. Numbers such as 530-537, 569-579, 686-691, 768-793, 943-947, 1002-10 At positions 40, 1052-1077, 1232-1248, and 1298-1300, or the corresponding amino acid positions It includes a region selected from the group consisting of amino acid residues.

[0028] In one embodiment of the above method, the deaminase variant is numbered in the Cas9 reference sequence For example, amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026 ~1027, 1029~1030, 1040~1041, 1052~1053, 1054~1055, 1067~1068, 1068~1069, It is inserted between amino acid positions 1247-1248, or 1248-1249, or their corresponding amino acid positions. In one embodiment of the notation method, the deaminase variant is numbered in the Cas9 reference sequence. For example, amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1040-1041, 1068- It is inserted between amino acid positions 1069, 1247-1248, or their corresponding positions. (Method described above) In one embodiment, the deaminase variant is numbered in the Cas9 reference sequence as , amino acid positions 1016-1017, 1023-1024, 1029-1030, 1040-1041, 1069-1070, or 12 It is inserted between amino acid positions 47-1248, or the corresponding amino acid positions. In one embodiment of the above method The adenosine deaminase variant is located at the locus identified in Table 13A, and the Cas9 polypeptide It is inserted (ins) into the cydode. In one embodiment, the N-terminal fragment is the amino acid residue of the Cas9 reference sequence. Base 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, approximately It includes 1248-1297, or the corresponding residue. In one embodiment, the C-terminal fragment is , amino acid residues 1301-1368, 1248-1297, 1078-1231, 1026-1051, 948 of the Cas9 reference sequence Includes residues ~1001, 692~942, 580~685, and / or 538~568, or their corresponding residues. nothing.

[0029] In one embodiment of the above method, the Cas9 polypeptide is modified Cas9, and is modified PAM or non-GP. It has specificity for AM. In one embodiment of the above method, the Cas9 polypeptide is nicasse. Either the Cas9 polypeptide is nuclease-inactive. In one embodiment of the above method In this embodiment, the Cas9 polypeptide is a modified SpCas9 polypeptide. , amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T133 It contains 7R (SpCas9-MQKFRAER) and has specificity for modified PAM 5'-NGC-3'.

[0030] In one embodiment of the above method, the fusion protein or ABE8 comprises a Cas12 polypeptide. In one embodiment, an adenosine deaminase variant is introduced into the Cas12 polypeptide. In one embodiment, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12 e, Cas12g, Cas12h, or Cas12i. In one embodiment, adenosine deaminase b The amino acid positions of rianto are: a) BhCas12b 153-154, 255-256, 306-307, 980-981, 101 9-1020, 534-535, 604-605, or 344-345, or Cas12a, Cas12c, Cas12d, Cas The corresponding amino acid residues of 12e, Cas12g, Cas12h, or Cas12i; b) 147-148 of BvCas12b , 248-249, 299-300, 991-992, or 1031-1032, or Cas12a, Cas12c, Cas12d , the corresponding amino acid residue of Cas12e, Cas12g, Cas12h, or Cas12i; or c) AaCas 12b 157-158, 258-259, 310-311, 1008-1009, or 1044-1045, or Cas12a , the corresponding amino acid residues of Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i It is inserted in between. In one embodiment, the adenosine deaminase variant is identified in Table 13B. It is inserted into the Cas12 polypeptide at the gene locus. In one embodiment, the Cas12 polypeptide The domain is Cas12b. In one embodiment, the Cas12 polypeptide has a BhCas12b domain and a BvCas12 Includes the b domain or the AACas12b domain.

[0031] In one embodiment of the above method, the guide RNA is CRISPR RNA (crRNA) and transactivated cr It includes RNA (tracrRNA). In one embodiment of the above method, the subject is a mammal or a human. .

[0032] In another embodiment, any one of the above methods, embodiments and embodiments of a base editing system, Furthermore, a pharmaceutical composition comprising a pharmaceutically acceptable carrier, vehicle, or excipient is proposed. To provide.

[0033] In one embodiment, the cells of the above embodiment and embodiment, as well as a pharmaceutically acceptable carrier, are used. - Provides a pharmaceutical composition comprising a vehicle or excipient.

[0034] In another embodiment, the base editing system includes any one of the above methods, embodiments, and embodiments. We will provide a kit.

[0035] In another embodiment, a kit is provided that includes cells from any one of the embodiments described above. In one embodiment of the kit, the kit further includes a package containing instructions for use. Includes inserts.

[0036] In one embodiment, the herein provides a polynucleotide programmable DNA-binding domain In, and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD Contains adenosine deaminase variants with modifications at amino acid positions 82 or 166, It is a base editor that contains at least one base editor domain.

[0037] In one embodiment, the base editor system includes the base editor and guide RNA. The guide RNA targets the base editor and performs alpha-1 antitripsy This results in modifications of SNPs associated with adenosine deaminase. In some embodiments, adenosine deaminase is used. - The variants include the V82S modification and / or the T166R modification. In some embodiments, Adenosine deaminase variants include the following modifications: Y147T, Y147R, Q154S, Y123H, and It further comprises one or more of Q154R. In some embodiments, the base editor domain is Contains wild-type adenosine deaminase domain and adenosine deaminase variant Contains adenosine deaminase heterodimer. In some embodiments, adenosine The aminase variants, compared to full-length TadA8, are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, Partially excised Ta is missing N-terminal amino acid residues 12, 13, 14, 15, 6, 17, 18, 19, or 20. It is dA8. In some embodiments, the adenosine deaminase variant is full-length TadA8. In comparison, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, Alternatively, it is a partially excised TadA8 in which 20 C-terminal amino acid residues are missing. In some embodiments, The polynucleotide programmable DNA-binding domain is modified Staphylococcus aureus Cas 9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), modified Streptococcus pyo The genes (SpCas9) or its variants. In some embodiments, polynucleotides The Otido programmable DNA-binding domain is a modified protospacer adjacent motif (PAM). It is a variant of SpCas9 that has specificity for heterosexual or non-G PAM. Several implementations Morphologically, the polynucleotide programmable DNA-binding domain is nuclease-inactive Cas 9. In some embodiments, the polynucleotide programmable DNA-binding domain is , Cas9 nickas.

[0038] In one embodiment, one or more guide RNAs, and the following sequences: JPEG0007861090000003.jpg195164 (Bold sequences indicate sequences derived from Cas9, italic sequences indicate linker sequences, and underlined sequences indicate two parts) A polynucleotide programmable DNA-binding domain containing a nuclear localization sequence, and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD Contains an adenosine deaminase variant with modifications at amino acid positions 82 and / or 166. A fusion protein comprising at least one base editor domain, - Provide the system.

[0039] In one embodiment, a cell is provided that contains any one of the above-mentioned base editor systems. In some embodiments, the cells are human cells or mammalian cells. In some embodiments, The cells are ex vivo, in vivo, or in vitro.

[0040] This invention edits mutations associated with alpha-1 antitrypsin deficiency (A1AD). The present invention provides compositions and methods for the following compositions and articles as defined by the present invention. Isolated or otherwise manufactured in connection with the examples provided. The features and advantages will be evident from the detailed description and claims.

[0041] definition The following definitions supplement the definitions in the technical field and apply to this application. Cases related or unrelated to the intent, for example, patents or applications relating to common ownership It is not the same as any other method and material described herein. While the materials and methods described herein may be used in conducting the tests, preferred materials and methods may be used in this specification. This will be explained in writing. Therefore, the terms used herein are intended to describe specific embodiments. This is for general use only and is not intended to be limiting.

[0042] Unless otherwise defined, all technical and scientific terms used herein are defined in accordance with the Code of Civil Code. The meaning is as generally understood by those skilled in the art to which the invention pertains. The following references are To those skilled in the art, we provide general definitions of many terms used in this invention: Singleto n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The C ambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary o f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0043] In this application, the use of the singular form includes the plural form unless otherwise specified. As used, the singular forms "a," "an," and "the" are used when the context clearly indicates otherwise. It should be noted that unless otherwise specified, it will include multiple reference subjects. The use of "or" means "and / or" unless otherwise specified, and includes It is understood to be comprehensive. Furthermore, the terms "including" and "include" The use of other types such as "includes" and "included" is not limited. It is.

[0044] As used in this specification and in the claims, the term “comprising” (and (all forms of "comprise" and "comprises"), "have" (ing) (all forms of "have" and "has"), "include (incl) (including) (all forms thereof such as "include" and "includes") or "Containment" (including "contains" and "contain") Its various forms) are comprehensive or open, and include additional, unlisted elements or This does not exclude any method steps. Any embodiment discussed herein is any method of this disclosure. Alternatively, this can be done with respect to the composition, and vice versa. Furthermore, The methods disclosed can be achieved using the compositions disclosed herein.

[0045] The terms "about" or "approximately" are determined by those skilled in the art. This means that the value is within an acceptable margin of error, and this is because the value is within an acceptable margin of error. It is measured or determined in a manner that partially depends on the limitations of the measurement system. In this context, "approximately" means, according to practice in this technical field, within a standard deviation of 1 or greater than 1. It could mean that... Alternatively, "about" could mean up to 20%, up to 10%, up to 5% of a given value, or... This can mean a range of up to 1%. Or, especially in relation to biological systems or processes. This term can mean values ​​within the same number of digits, for example, within 5 times or 2 times. If a specific value is stated in the application and claims, unless otherwise stated, The term "approximately" means that a particular value is within an acceptable margin of error. It should be done.

[0046] The range provided herein includes all values ​​within the range, including the first and last values. It is understood to be an abbreviated expression. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9 , 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 ,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49 or any number, combination of numbers, or sub-range from a group of 50. It is understood.

[0047] In the specification, “several embodiments,” “a certain embodiment,” “one embodiment,” or The phrase "other embodiments" refers to specific features, structures, etc., described in relation to those embodiments. Or, the characteristics are included in at least some embodiments of this disclosure, but not necessarily all. This means that it is not necessarily included in the embodiment.

[0048] "Adenosine deaminase" refers to the hydrolytic deamination of adenine or adenosine. This means a polypeptide or fragment thereof that can catalyze a substance. In some embodiments, Deaminase or the deaminase domain converts adenosine to inosine, or deo Adenosine deoxyinosine catalyzes the hydrolytic deamination of xyadenosine to deoxyinosine. It is an aminase. In some embodiments, adenosine deaminase is deoxyribo This invention catalyzes the hydrolytic deamination of adenine or adenosine in nucleic acids (DNA). Adenosine deaminase provided in the book (for example, genetically modified adenosine deaminase) Adenosine deaminase (the evolved form of adenosine deaminase) can originate from any organism, such as bacteria.

[0049] In some embodiments, adenosine deaminase includes the following sequence modifications. : MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNY RLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLC YFFR MPRQVFNAQK KAQSSTD (Also known as TadA*7.10).

[0050] In some embodiments, TadA*7.10 includes at least one modification. In this state, TadA*7.10 includes modifications at amino acid positions 82 and / or 166. Specific implementation Morphologically, variants of the above reference sequence include one or more of the following modifications: Y147T, Y1 47R, Q154S, Y123H, V82S, T166R, and / or Q154R. Modified Y123H is also specified herein. In this context, it is also referred to as H123H (a modification in TadA*7.10 where the modified H123Y was reverted to Y123H (wt)). In other embodiments, variants of the TadA*7.10 sequence are Y147T + Q154R; Y147T + Q154 S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154 R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154 This includes combinations of modifications selected from the group consisting of R.

[0051] In other embodiments, the present invention relates to an adenosine deaminase variant containing a deletion, for example , including C-terminal deletions beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157. TadA*8 is provided. In other embodiments, there is one or more adenosine deaminase variants. The following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or T including Q154R It is an adA monomer (e.g., TadA*8). In other embodiments, adenosine deaminase variant Ants are: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y14 7R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; o This includes combinations of modifications selected from the group I76Y + V82S + Y123H + Y147R + Q154R. It is a monomer.

[0052] In yet another embodiment, the adenosine deaminase variant is one or more Modifications below: Two having Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R It is a homodimer containing the adenosine deaminase domain (e.g., TadA*8). In the application form, the adenosine deaminase variants are: Y147T + Q154R; Y147T + Q15 4S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q15 4R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q15 Two adenosine deaminase drugs having a combination of modifications selected from the 4R group It is a homodimer containing n (e.g., TadA*8).

[0053] In other embodiments, the adenosine deaminase variant is wild-type TadA adenosine deaminase Minase domain, and the following modified Y147T, Y147R, Q154S, Y123H, V82S, T166R, adenosine deaminase variant domains containing one or more Q154R (e.g., T It is a heterodimer containing adA*8). In other embodiments, adenosine deaminase b The rianto contains the wild-type TadA adenosine deaminase domain, as well as:Y147T + Q154R;Y 147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y1 23H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q15 4R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y 147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y1 Adenosine deaminase barriers containing modified combinations selected from the 47R + Q154R group It is a heterodimer containing a nucleotide domain (e.g., TadA*8).

[0054] In other embodiments, the adenosine deaminase variant is the TadA*7.10 domain, In addition, any of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R A heterogeneous adenosine deaminase variant domain (e.g., TadA*8) containing the above In other embodiments, the adenosine deaminase variant is TadA*7.10. In, and the following modifications: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154 S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147 Adenosine containing the combinations R + Q154R; or I76Y + V82S + Y123H + Y147R + Q154R. It is a heterodimer containing a deaminase variant domain (e.g., TadA*8).

[0055] In one embodiment, adenosine deaminase is the following sequence or adenosine deaminase TadA*8 is a fragment containing or essentially derived from the following fragments that have xenergic activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0056] In some embodiments, TadA*8 is partially removed. In some embodiments, In partial resection TadA*8, compared to full-length TadA*8, the values ​​are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 1 The N-terminal amino acid residues 3, 14, 15, 6, 17, 18, 19, or 20 are missing. In the embodiment, in the partially excised TadA*8, compared to the full-length TadA*8, 1, 2, 3, 4, 5, 6, 7, 8, If C-terminal amino acid residues 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 are lost In some embodiments, the adenosine deaminase variant is full-length TadA*8. .

[0057] In certain embodiments, the adenosine deaminase heterodimer has a TadA*8 domain, and containing an adenosine deaminase domain selected from one of the following:

[0058] Staphylococcus aureus (S. aureus)TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTL YVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0059] Bacillus subtilis (B. subtilis)TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTL EPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLS E

[0060] Salmonella typhimurium(S. typhimurium)TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLV LQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFF I DRAW RMRRQEICALKCADE

[0061] Shewanella putrefaciens(S. putrefaciens)TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEP CAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFKRRRDEKKALQRAQ QGIE

[0062] Haemophilus influenzae F3031(H. influenzae)TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYR LLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRRE EKKIEKALLKSLSDK

[0063] Caulobacter crescentus(C. crescentus)TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLT DLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAK I

[0064] Geobacter sulfurreducens(G. sulfrreducens)TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLT GATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRK KAKATPALFIDERKVPPEP

[0065] TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0066] The "adenosine deaminase base editor (ABE8) polypeptide" is based on the following reference sequence : MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD Contains an adenosine deaminase variant with modifications at amino acid positions 82 and / or 166. This means a base editor (BE) as defined and / or described herein.

[0067] In some embodiments, ABE8 includes further modifications compared to the reference sequence.

[0068] "Adenosine deaminase base editor 8 (ABE8) polynucleotide" is ABE8 poly This refers to a polynucleotide (polynucleotide sequence) that codes for a peptide.

[0069] "Administer" in this specification means administering to a patient or subject one or more of the combinations described herein. This refers to providing a product. For example, and without limitation, the administration of a composition, such as by injection, is static. intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, or It may be administered by intramuscular (im) injection. One or more such routes may be used. Parenteral administration may be, for example, by bolus injection or by prolonged gradual perfusion. Alternatively, administration may also be carried out via the oral route.

[0070] "The agent" can be any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide. Rui means fragment.

[0071] "Alpha-1 antitrypsin (A1AT) protein" is found in UniProt deposit number P01009. This refers to polypeptides or fragments that have at least approximately 95% amino acid sequence identity. In certain embodiments, the A1AT protein undergoes one or more modifications compared to the following reference sequence. Includes. In one particular embodiment, the A1AT protein associated with A1AD includes the E342K mutation. An example A1AT amino acid sequence is provided below. >sp|P01009|A1AT_HUMAN Alpha-1-antitrypsin OS=Homo sapiens OX=9606 GN=SERP INA1 PE=1 SV=3 MPSSVSWGILLLAGLCCLVPVSLAEDPQGDAAQKTDTSHHDQDHPTFNKITPNLAEFAFFSLYRQLAHQSNSTNIFFSPVS IATAFAMLSLGTKADTHDEILEGLNFNLTEIPEAQIHEGFQELLRTLNQPDSQLQLTTGNGLFLSEGLKLVDKFLEDVKK LYHSEAFTVNFGDTEEAKKQINDYVEKGTQGKIVDLVKELDRDTVFALVNYIFFKGKWERPFEVKDTEEEDFHVDQVTTV KVPMMKRLGMFNIQHCKKLSSWVLLMKYLGNATAIFFLPDEGKLQHLENELTHDIITKFLENEDRRSASLHLPKLSITGT YDLKSVLGQLGITKVFSNGADLSGVTEEAPLKLSKAVHKAVLTIDEKGTEAAGAMFLEAIPMSIPPEVKFNKPFVFLMIE QNTKSPLFMGKVVNPTQK In the above A1AT protein sequence, the first 24 amino acids constitute a signal peptide. (Underlined). Position 342 of the mutated sequence in A1AD (i.e., E342K) follows the signal sequence. It is determined based on amino acid residue "E" which is set as amino acid "1".

[0072] "Modification" means detection by known methods of standard art as described herein. Changes in the structure, expression level, or activity of a gene or polypeptide (e.g., increased activity) (or reduction) means. When used herein, modification means polynucleotide or poly Changes in the lipeptide sequence, or changes in expression levels, e.g., a 25% change, a 40% change, a 50% change. This includes changes in the expression levels mentioned above.

[0073] "To ameliorate" means to reduce, suppress, or attenuate the onset or progression of a disease. It means to cause, reduce, stop, or stabilize.

[0074] "Analog" refers to a molecule that is not identical but possesses similar functional or structural characteristics. It tastes like... For example, polynucleotide or polypeptide analogs are the same as the corresponding naturally occurring polynucleotides. While retaining the biological activity of the dinucleotide or polypeptide, naturally occurring polynucleotides Certain biochemical properties enhance the function of its analog compared to ocides or polypeptides. It has modifications. Such modifications, for example, do not alter ligand binding, and the analog... DNA affinity, efficacy, specificity, protease or nuclease resistance, membrane It can increase permeability and / or half-life. The analog is a non-natural amino acid. It may include.

[0075] A "base editor (BE)" or "nucleic acid base editor (NBE)" is a polynucleotide editor. This refers to an agent that binds to ocides and has nucleic acid base modification activity. In various embodiments, the base e Diter is a nucleic acid base modified polypeptide (e.g., deaminase) and nucleic acid program The possible nucleotide-binding domain, along with a guide polynucleotide (e.g., guide RNA), It contains. In various embodiments, the agent is a protein domain having base editing activity, that is, In other words, it is possible to modify the bases (e.g., A, T, C, G, U) within nucleic acid molecules (e.g., DNA). It is a biomolecular complex containing a domain. In some embodiments, polynucleotide pro The Gram-capable DNA-binding domain is fused to or ligated to the deaminase domain. Morphologically, the agent is a fusion protein containing a domain with base-editing activity. In the application form, the protein domain with base editing activity is linked to the guide RNA (e.g.) For example, RNA-binding motifs on guide RNA and RNA-binding domains fused to deaminase. (via). In some embodiments, the domain having base editing activity is located within the nucleic acid molecule. Bases can be deaminated. In some embodiments, the base editor can deaminate DNA. One or more bases within the molecule can be deaminated. In some embodiments, the base The editor can deaminate adenosine (A) in DNA. Several implementations Morphologically, the base editor is an adenosine base editor (ABE).

[0076] In some embodiments, the base editor is a circular permutant Cas9 (e.g., spCas9 or saCas9) and adenosine dea within a scaffold containing a bipartite nuclear localization sequence It is produced by cloning a minase variant (e.g., TadA*8) (for example) (ABE8). The circulating replacement Cas9 is known in the art, for example, Oakes et al., Cell 176 This is described in 254-267, 2019. Exemplary cyclic substitutions are as follows, indicated in bold. The columns show sequences derived from Cas9, the italicized sequences show linker sequences, and the underlined sequences show binuclear localization. Show the array.

[0077] CP5(MSP "NGC = Pam variant containing mutations, normal Cas9 prefers NGG", PID = protein (Contains quality interaction domain and "D10A" nickase): JPEG0007861090000004.jpg169162

[0078] In some embodiments, ABE8 is a base editor derived from Tables 6-9, 13, or 14 below. More selected. In some embodiments, ABE8 is an adenosine dea evolved from TadA. Contains a minase variant. In some embodiments, ABE8 adenosine deaminator The Zevariant is a TadA*8 variant as described in Tables 7, 9, 13, or 14 below. In some embodiments, the adenosine deaminase variants are Y147T, Y147R, Q154 Ta includes one or more modifications selected from the group S, Y123H, V82S, T166R, and / or Q154R. This is a dA*7.10 variant (e.g., TadA*8). In various embodiments, ABE8 is Y147T + Q15 4R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123 H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123 TadA*7.10 variants include combinations of modifications selected from the group H + Y147R + Q154R. For example, it includes TadA*8). In some embodiments, ABE8 is a monomer construct. In one embodiment, ABE8 is a heterodimer construct. In some embodiments, ABE8 is an array: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD Includes.

[0079] In some embodiments, the polynucleotide programmable DNA-binding domain is CRISP It is an R-associating enzyme (e.g., Cas or Cpf1). In some embodiments, the base editor is A catalytically inactive (dead)Cas9 (dCas9) fused to the deaminase domain. In some embodiments, the base editor is Ca fused to the deaminase domain. This is s9 nickase (nCas9). For details on the base editor, see International PCT application number PCT / 2017 / 0453. It is described in 81 (WO 2018 / 027078) and PCT / US 2016 / 058344 (WO 2017 / 070632), The entirety of each is incorporated herein by reference. Komor, AC, et al., “Programma ble editing of a target base in genomic DNA without double-stranded DNA cleavage ”Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editin g of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (20 17); Komor, AC, et al., “Improved base excision repair inhibition and bacteri ophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and "Product purity," Science Advances 3:eaao4774 (2017), and Rees, HA, et al. “Base editing: precision chemistry on the genome and transcriptome of living ce lls.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 also See reference (the entire content of which is incorporated into this specification by reference).

[0080] For example, the base editing compositions, systems, and methods described herein use The Una Adenine Base Editor (ABE) provides nucleic acid sequences (8877 base pairs) such as the following: , (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681) :464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct It has at least 95% of the ABE nucleic acid sequence. (36(9):843-846. doi: 10.1038 / nbt.4172.) This also includes polynucleotide sequences that share the same identity. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGG GACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGG GCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAA ATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGAGGTC TATATAAGCAGAGCTGGTTTAGTGAACCGTCAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAG CCGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAAGTCTCTGAAGTCGAGTT AGCCACGAGTATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCACACGCAGAGA TCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCCTGTATGTGACACTGGAGCCA TGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGTGTTCGGAGCACGGGACGCCAAGACCGGCGC AGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATGAACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACG AGTGCGCCGCCCTGCTGAGCGATTTCTTTAGAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACC GACTCTGGAGGATCTAGCGGAGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGG CGGCTCCTCCGGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTGGAACAGAGCC ATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTGGTCATGCAGAACTACAGACT GATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGCGCCATGATCCACTCTAGGATCGGCCGCG TGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGGCTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCAC CGCGTCGAAATTACCGAGGGAATCCTGGCAGATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGT GTTCAATGCTCAGAAGAAGGCCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTG GCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTGGGCAA CACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACAGCCGAGGCCACCCGGC TGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATG GCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCC CATCTTCGGCAACATCGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGG ACAGCACCGACAAGGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATC GAGGGCGACCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGACGGCTGGAAA ATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCCTGAGCCTGGGCCTGACCCCC AACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACGACCTGGACAA CCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCG ACATCCTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAG GACCTGACCCTGCTGAAAGCTCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAA CGGCTACGCCGGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATC CCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAGGACAACCG GGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCCAGGGGAAACAGCAGATTCGCCT GGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGC TTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTA CTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGC AGAAAAAGGCCATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATACCACGATCT GCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGAAGATATCGTGCTGACCCTGA CACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCCCACCTGTTCGACGACAAAGTGATGAAGCAG CTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAA GACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCT TTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGC CCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGAATGAAGCGGA TCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAACACCAGCTGCAGAACGAAG CTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGAACTGGACATCAACCGGCTGTCCGACTACGA TGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACC GGGGCAAGAGCGACAACGTGCCCTCCGAAGAGGTCGTGAAGAAGATGAAGAACTACTGCGGCAGCTGCTGAACGCCAAG CTGATTACCCAGAGAAAGTTCGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCAT CAAGAGACAGCTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCGGAAGGATTTC CAGTTTTACAAAGTGCGCGAGATCCAACAACTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACGCCCT GATCAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCA AGAGCGGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATT ACCCTGGCCAACGGCGAGATCCGGAAGCGGCCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGG CCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCG GCTTCAGCAAAGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGTCCAAGAAACT GAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAAGAGCAGCTTCGAGAAGAATCCCATCGACTTTCTGGAAG CCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCGG AAGAGAATGCTGGCCTTGCCGGCGAACTGCAGAAGGGAACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTA CCTGGCCAGCACTATGAGAAGCTGAAGGGCTCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGC ACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTG CTGTCCGCCTACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGG ACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAGGTGACTCTGGC GGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAACCGGTCATCATCACCATCAC CATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCC TTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGT GTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGAT GCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGAAGCATAAAGT GTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAAC CTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCTCTTCCGCTTCCTC GCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGAGCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCA CAGAATCAGGGGATAACGCAGGAAAGAACATGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTG CTGGCGTTTTTCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGAC AGGACTATAAAGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTC GTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTATCCGGTAACTATCGTCTTGA GTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTAGGCG GTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAG CCAGTTACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTG CAAGCAGCAGATTACGCGCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGA ACGAAAACTCACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGGCACCTATCTC AGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTACGATACGGGAGGGCTTACCAT CTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCAGATTTATCAGCAATAAACCAGCCAGCCGGA AGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCTCCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAG TAGTTCGCCAGTTAATAGTTTGCGCAACGTTGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGG CTTCATTCAGCTCCGGTTCCCAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTC GGTCCTCCGATCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGTATGCGGCGAC CGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAAAAGTGCTCATCATTGGAAAA CGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAGTTCGATGTAACCCACTCGTGCACCCAACTG ATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGAGCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAA GGGCGACACGGAAATGTTGAATACTCATACTCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATG AGCGGATACATATTTGAATGTATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGA CGTCGACGGATCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAACAAGGCAAGGC TTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGATGTACGGGCCAGATATACGCG TTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCG TTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTT CCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACA TCAAGTGTATC

[0081] "Base editing activity" refers to the ability to chemically modify the bases within a polynucleotide. This means that. In one embodiment, the first base is converted to the second base. So, base editing activity is adenosine or adenine deaminase activity, for example, A·T to G·C This is the activity that converts to adenosine or adenine deaminase activity. Sexual activity, for example, the activity to convert A·T to G·C, and cytidine deaminase activity, for example, target C The activity may also involve converting G to T and A. In some embodiments, the base editing activity is It is evaluated by editing efficiency. Base editing efficiency is evaluated by any appropriate means, for example, by Sun It can be measured by Gar sequencing or next-generation sequencing. In some embodiments, salt Base editing efficiency is the total sequencing read, including nucleic acid base conversion, achieved by the base editor. The proportion, for example, the proportion of total sequencing reads that include target AT base pairs converted to GC base pairs. It is measured as follows. In some embodiments, base editing efficiency is measured in a cell population. When this is done, the percentage of total cells that include nucleic acid base conversion achieved by the base editor is... It is measured.

[0082] The term "base editor system" refers to a system that edits the nucleic acid bases of a target nucleotide sequence. This refers to a system. In various embodiments, the base editor system is (1) polynucleotide (2) the nucleic acid base (2) a programmable nucleotide-binding domain (e.g., Cas9) A deaminase domain (e.g., adenosine deaminase) for deamination of; ( 3) Containing one or more guide polynucleotides (e.g., guide RNA). Several implementations In this state, the polynucleotide programmable nucleotide-binding domain is polynucleotide It is a programmable DNA-binding domain. In some embodiments, the base editor is , adenine or adenosine base editor (ABE). In some embodiments, The base editor system is ABE8.

[0083] In some embodiments, the base editor system has more than one base editing component. It may include more than one deaminase. For example, a base editor system may contain more than one deaminase. However, in some embodiments, the base editor system may use one or more adenosine de It may also contain aminase. This can be used to target different deaminases to the target nucleic acid sequence. In that embodiment, a single pair of guide polynucleotides is used to differentiate the target nucleic acid sequence. You may also target the deaminase.

[0084] The base editor system allows programming of the deaminase domain and polynucleotides. The nucleotide-binding components may be associated with each other covalently or non-covalently. or they may be associated by any combination of such associations and interactions. For example, in some embodiments, the deaminase domain is a polynucleotide prog The nucleotide-binding domain allows for targeting of the target nucleotide sequence. In some embodiments, polynucleotide programmable nucleotide bonded The main component may be fused to or ligated to the deaminase domain. Several implementations Morphologically, the polynucleotide programmable nucleotide-binding domain is a deaminase By interacting with or associating with the domain non-covalently, target nucleotides are distributed The deaminase domain can be targeted to the column. For example, in some embodiments, The minase domain is part of the polynucleotide programmable nucleotide-binding domain. This involves interacting with, associating with, or forming complexes with additional heterogeneous parts or domains. It may include additional heterogeneous parts or domains that can form a certain embodiment. So, the additional heterogeneous part binds to, interacts with, associates with, or forms a complex with the polypeptide. It may be possible to form additional heterogeneous parts. In some embodiments, additional heterogeneous parts are polynucleotides. It may be possible to bind, interact with, associate with, or form complexes with rheotides. In this embodiment, additional heterogeneous portions may be able to bind to the guide polynucleotide. In some embodiments, additional heterogeneous parts may be bondable to the polypeptide linker. In some embodiments, additional heterologous parts are capable of binding to a polynucleotide linker. Possible. The additional heterogeneous portion may be a protein domain. Several embodiments So, the additional heterogeneous parts are the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat. Protein domain, SfMu Com coated protein domain, sterile alpha-motif F, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif It could be the Sm7 protein or an RNA recognition motif.

[0085] The base editor system may further include guide polynucleotide components. The components of the base editor system are covalent bonds, non-covalent interactions, or their It is understood that they can be related to each other through any combination of association and interaction. It should be done. In some embodiments, the deaminase domain is a guide polynucleotide. Tide can be used to target specific nucleotide sequences. For example, several In the application form, the deaminase domain is part or segment of the guide polynucleotide. (e.g., polynucleotide motifs) capable of interacting, associating, or forming complexes A certain additional heterogeneous part or domain (e.g., a polynucleus such as an RNA or DNA-binding protein) It may include a cleotide-binding domain. In some embodiments, additional heterogeneous parts or Domains (for example, polynucleotide-binding domains such as RNA or DNA-binding proteins) It can be fused to or linked to the deaminase domain in some embodiments. The portion binds to, interacts with, associates with, or forms a complex with the polypeptide. It may be possible to do so. In some embodiments, additional heterogeneous parts are polynucleotides. It is possible to bind to, interact with, associate with, or form complexes with tides. It is possible. In some embodiments, additional heterogeneous parts are bound to the guide polynucleotide. It may be possible. In some embodiments, additional heterogeneous parts are added to the polypeptide linker. It may be possible to bind to it. In some embodiments, additional heterologous parts are polynucleotides. It may be possible to bind to a linker. The additional heterogeneous portion may be a protein domain. In some embodiments, additional heterogeneous parts include K homology (KH) domains, MS2 coat proteins. Quality domain, PP7 coated protein domain, SfMu Com coated protein domain, ste Lyle-alpha motif, telomerase Ku-binding motif and Ku protein, telomer This could be a Sm7 binding motif and / or an RNA recognition motif, such as the Sm7 protein or RNA recognition motif.

[0086] In some embodiments, the base editor system is a base excision repair (BER) component. It may further contain inhibitors of the base editor system. The components of the base editor system are covalent, non Related to one another through shared interactions, or any combination thereof, of association and interaction. It should be understood that it may be added. BER component inhibitors are BER inhibitors It may also include. In some embodiments, the BER inhibitor is uracil DNA glycosylation. It may also be a bisphosphonate inhibitor (UGI). In some embodiments, the BER inhibitor is a wild boar It may also be a BER inhibitor. In some embodiments, the BER inhibitor is a polynucleotide. The Otid programmable nucleotide-binding domain allows the target nucleotide sequence to be targeted. It may be programmed. In some embodiments, polynucleotide programmed The nucleotide-binding domain may be fused to or linked to the BER inhibitor. In several embodiments, the polynucleotide programmable nucleotide-binding domain is The aminase domain and BER inhibitors may be fused to or linked to several. In this embodiment, the polynucleotide programmable nucleotide-binding domain inhibits BER Non-covalent interactions or associations with harmful factors target BER inhibitors to the target nucleotide sequence. It may also be targeted to the inhibitors of BER components. The child is an additional heterogene that is part of the polynucleotide programmable nucleotide-binding domain. The species part or domain may interact, associate, or form a complex with it. It may include any additional heterogeneous parts or domains that are possible.

[0087] In some embodiments, BER inhibitors are targeted by guide polynucleotides. It can be targeted to a nucleotide sequence. For example, in some embodiments, the BER The inhibitor is a part or segment of the guide polynucleotide (e.g., polynucleotide) (Motif) is either interactable, associateable, or capable of forming a complex with it. Additional heterogeneous parts or domains (e.g., polynucleotides such as RNA or DNA-binding proteins) It may include an ocidal binding domain. In some embodiments, the guide polynucleotide Additional heterogeneous parts or domains (e.g., polynucleos such as RNA or DNA-binding proteins) The tide-binding domain can be fused to or linked to a BER inhibitor. In some embodiments... The additional heterogeneous portion either binds to, interacts with, or associates with the polynucleotide. Alternatively, it may be possible to form a composite. In some embodiments, additional heterogeneous parts The fraction may be able to bind to the guide polynucleotide. In some embodiments, additional The heterogeneous portion may be able to bind to the polypeptide linker. In some embodiments, The additional heterogeneous portion may be able to bind to a polynucleotide linker. It may be a protein domain. In some embodiments, the additional heterogeneous portion is the K phase. Same (KH) domain, MS2 coat protein domain, PP7 coat protein domain, SfMu Com coated protein domain, steril alpha motif, telomerase Ku-binding motif ---- and Ku proteins, telomerase Sm7 binding motif and Sm7 protein, or RN A. This could be a recognizable motif.

[0088] The terms "Cas9" or "Cas9 domain" refer to the Cas9 protein or a fragment of it (e.g., Cas9 The active, inactive, or partially active DNA cleavage domain, and / or gRNA binding of Cas9. This refers to RNA-induced nucleases containing proteins (including the domain). Cas9 nuclease is Casnl nuclease or CRISPR (clustered regularly interspaced short palindrom) It is sometimes called an associative nuclease (CRISPR). Adaptive immune system provides protection against viruses, translocation factors, and zygosity plasmids. The CRISPR cluster consists of a spacer, a complementary arrangement to the preceding movable element, and a target. It contains target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA is performed in the transendence. coding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein TracrRNA is required. TracrRNA is a guide in the processing of pre-crRNA assisted by ribonuclease 3. Next, Cas9 / crRNA / tracrRNA produces a linear or circular dsD complementary to the spacer. The NA target is cleaved with an endonuclease. Target strands that are not complementary to the crRNA are first endonucleated. It is cleaved nuclease-wise and then trimmed to 3'-5' exonuclease-wise. In the world, both proteins and RNA are required for DNA binding and cleavage. However, crRNA and t To incorporate both aspects of racrRNA into a single RNA species, a single guide RNA ("sgRNA") is used. Alternatively, it is possible to manipulate "gRNA" (or simply "gRNA"). For example, Jinek M., Chylinski K., Fonfa See ra I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012). (The entire content of which is incorporated herein by reference). Cas9 is a CRISPR repeat sequence Recognize the short motifs inside (PAM or protospacer adjacent motifs) and the "self" and It helps to distinguish "non-self". The sequence and structure of Cas9 nuclease are well known to those skilled in the art. Known (for example, “Complete genome sequence of an Ml strain of Streptococcus p yogenes.”Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qia n Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001 ); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. ” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, E ckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and “A progr ammable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jine k M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 33 See 7:816–821 (2012), the entire contents of which are incorporated herein by reference. Cas9 orthologs include, but are not limited to, S. pyogenes and S. thermophilus. It has been described in a variety of species. Additional suitable Cas9 nucleases and sequences are disclosed herein. As will become apparent to those skilled in the art, such Cas9 nucleases and sequences are Chylinsk i, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas The organisms and genes disclosed in “immune systems” (2013) RNA Biology 10:5, 726-737 It includes the Cas9 sequence derived from the constellation Densigma. Its entire contents are incorporated herein by reference.

[0089] An example of Cas9 is Streptococcus pyogenes Cas9 (spCas9), and its amino acid sequence It is provided below. JPEG0007861090000005.jpg168162 (Single underline: HNH domain; Double underline: RuvC domain)

[0090] Nuclease-inactivating Cas9 protein can be replaced with the "dCas9" protein (nuclease- It can also be called “dead” Cas9 or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) containing the 'in' gene are known (e.g., Jinek et al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Gui ded Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; See 152(5): 1173–83 (the full contents of each are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 consists of the HNH nuclease subdomain and the RuvC1 subdomain. It is known to contain two subdomains. The HNH subdomain is complementary to the gRNA. By cleaving the chain, the RuvC1 subdomain cleaves non-complementary chains within these subdomains. Mutations can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A suppress Cas9 activity. Complete inactivation of the pyogenes Cas9 nuclease activity (Jinek et al, Science. 337: 816-821 (2012); Qi et al, Cell. 28;152(5): 1173-83 (2013). In some embodiments Cas9 nucleases have an inactive (e.g., inactivated) DNA cleavage domain, that is, In fact, Cas9 is a nickase called the "nCas9" protein (meaning "nickase" Cas9). In some embodiments, a protein containing a Cas9 fragment is provided. For example, several In that embodiment, the protein has two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) comprising one of the DNA cleavage domains of Cas9. In some embodiments, Cas9 or Proteins containing this fragment are called "Cas9 variants." or they share homology to that fragment. For example, a Cas9 variant shares homology to wild-type Cas9, At least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical 1. At least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 9 They are 9% identical, at least about 99.5% identical, or at least about 99.9% identical. Morphologically, Cas9 variants differ from wild-type Cas9 by 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. , 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30 ,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50 Alternatively, it may have more amino acid changes. In some embodiments, the Cas9 variant is a fragment of Cas9 (e.g., the gRNA-binding domain or The fragment contains a DNA cleavage domain, and the fragment is at least approximately 70% identical to the corresponding fragment of wild-type Cas9. 1. At least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 9 6% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least They are also approximately 99.5% identical, or at least approximately 99.9% identical. In some embodiments, the fragments are , at least about 30%, at least about 35%, and at least the amino acid length of the corresponding wild-type Cas9. Also about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, little At least 65%, at least 70%, at least 75%, at least 80%, at least 85% %, at least about 90%, at least about 95%, at least about 96%, at least about 97%, and at least It is approximately 98%, at least approximately 99%, or at least approximately 99.5%.

[0091] In some embodiments, the fragment is at least 100 amino acids in length. In the application form, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 5 50, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, Or at least 1300 amino acids.

[0092] In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes. (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows). ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCCATCCATGGTCTTTGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG0007861090000006.jpg166162(Picture:HNHドメイン;Picture:RuvCドメイン)

[0093] In some embodiments, wild-type Cas9 comprises the following nucleotides and / or amino acids Corresponding to or containing arrays: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAACCTGATCGCACAATTACCCGGAGAGAAGAAAAAATGGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGGTTCGCCAGCCATCAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG0007861090000007.jpg167164 (Single underline: HNH domain; Double underline: RuvC domain)

[0094] In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes. (NCBI reference sequence: NC_002737.2 (nucleotide sequence is as follows) and Uniprot reference sequence) Column: Q99ZW2 (amino acid sequence is as follows): ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCAGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG0007861090000008.jpg166162 (Sequence ID 1. Single underline: HNH domain; Double underline: RuvC domain)

[0095] In some embodiments, Cas9 is Corynebacterium ulcerans (see NCBI: NC_015683). 1, NC_017317.1); Corynebacterium diphtheria (NCBI reference: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI reference: NC_021284.1); Prevotella intermedia (NCBI Reference: NC_017861.1); Spiroplasma taiwanense (NCBI reference: NC_021846.1); Streptococc us iniae (NCBI reference: NC_021314.1); Belliella baltica (NCBI reference: NC_018010.1); Psy chroflexus torquisI (NCBI reference: NC_018721.1); Streptococcus thermophilus (NCBI reference: NC_018721.1); Reference: YP_820832.1), Listeria innocua (NCBI reference: NP_472073.1), Campylobacter jejuni (See NCBI: YP_002344900.1) or Neisseria meningitidis (See NCBI: YP_002342) 100.1) Represents Cas9 of origin, or Cas9 of any other biological origin.

[0096] In some embodiments, Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or It is a variant. In some embodiments, NmeCas9 has specificity for NNNNGAYW PAM. It has such that Y is C or T and W is A or T. In some embodiments, NmeCas9 is N It has specificity for NNNGYTT PAM, where Y is C or T. In some embodiments, NmeC as9 has specificity for NNNNGTCT PAM. In some embodiments, NmeCas9 is Nme1 C as9 is in some embodiments. TC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNN NCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PAM, or NNNGATT It has specificity for PAM. In some embodiments, Nme1Cas9 is NNNNGATT PAM, NN It has specificity for NNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, or NNNNCCTG PAM. In some embodiments, NmeCas9 is specific to CAA PAM, CAAA PAM, or CCA PAM. It has properties. In some embodiments, NmeCas9 is Nme2 Cas9. Therefore, NmeCas9 has specificity for NNNNCC (N4CC) PAM, where N is A, G, C, or T. There is one difference. In some embodiments, NmeCas9 is NNNNCCGT PAM, NNNNCCGGPAM, N NNNCCCA PAM, NNNCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PA It has specificity for M or NNNGATT PAM. In some embodiments, NmeCas9 is Nme It is 3Cas9. In some embodiments, NmeCas9 is NNNNCAAA PAM, NNNNCC PAM, or It has specificity for NNNNCNNN PAM. In some embodiments, Nme1, Nme2, or Nme3 The PAM interaction domains are N4GAT, N4CC, and N4CAAA, respectively. Further NmeC The characteristics of as9 and the PAM sequence are described in Edraki et al., A Compact, High-Accuracy Cas9 with a Di nucleotide PAM for In Vivo Genome Editing, Mol. Cell. (2019) 73(4): 714-726. The document is included, and the entire document is incorporated herein by reference.

[0097] Exemplary Neisseria meningitidis Cas9 protein, Nme1Cas9, (NCBI reference: WP_002235) 162.1; Type II CRISPR RNA-induced endonuclease Cas9) has the following amino acid sequence : 1 maafkpnpin yilgldigia svgwamveid edenpiclid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlspelqd eigtafslfk tdeditgrlk driqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlgrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy ​​eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkernlndt ryvnrflcqf vadrmrltgk gkkrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgevlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 qghmetvksa krldegvsvl rvpltqlklk dlekmvnrer epklyealka rleahkddpa 901 kafaepfyky dkagnrtqqv kavrveqvqk tgvwvrnhng iadnatmvrv dvfekgdkyy 961 lvpiyswqva kgilpdravv qgkdeedwql iddsfnfkfs lhpndlvevi tkkarmfgyf 1021 aschrgtgni nirihdldhk igkngilegi gvktalsfqk yqidelgkei rpcrlkkrpp 1081 vr

[0098] Another exemplary Neisseria meningitidis Cas9 protein, Nme2Cas9, (see NCBI: WP_00) 2230835; Type II CRISPR RNA-induced endonuclease Cas9) has the following amino acid sequence する: 1 maafkpnpin yilgldigia svgwamveid eeenpirlid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvannahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlsselqd eigtafslfk tdeditgrlk drvqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlvrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkecnlndt ryvnrflcqf vadhilltgk gkrrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgkvlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 ahkdtlrsak rfvkhnekis vkrvwlteik ladlenmvny kngreielye alkarleayg 901 gnakqafdpk dnpfykkggq lvkavrvekt qesgvllnkk naytiadngd mvrvdvfckv 961 dkkgknqyfi vpiyawqvae nilpdidckg yriddsytfc fslhkydlia fqkdekskve 1021 fayyincdss ngrfylawhd kgskeqqfri stqnlvliqk yqvnelgkei rpcrlkkrpp 1081 vr

[0099] In some embodiments, dCas9 inactivates one or more nucleotides that inactivate Cas9 nuclease activity. It partially or entirely corresponds to the mutated Cas9 amino acid sequence, or so It includes the sequence partially or entirely. For example, in some embodiments, dCas9 dome The genes include D10A and H840A mutations, or corresponding mutations in other Cas9 sequences. In one embodiment, dCas9 includes the amino acid sequence of dCas9 (D10A and H840A): JPEG0007861090000009.jpg169163 (Single underline: HNH domain; Double underline: RuvC domain)

[0100] In some embodiments, the Cas9 domain contains a D10A mutation at position 840, or in this specification The amino acid sequence provided in this document contains the corresponding residue in the amino acid sequence provided above. It remains histidine.

[0101] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided. This, for example, produces nuclease-inactivated Cas9 (dCas9). For example, other amino acid substitutions at D10 and H840, or the Cas9 nuclease dome Other substitutions within the nuclease (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) This includes substitutions within the code. In some embodiments, at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, less They are at least 99% identical, at least 99.5% identical, or at least 99.9% identical, dC Provides a variant or homolog of as9. In some embodiments, about 5 amino acids, about 10 amino acids, approximately 15 amino acids, approximately 20 amino acids, approximately 25 amino acids, approximately 30 amino acids, approximately 40 amino acids Approximately 50 amino acids, approximately 75 amino acids, approximately 100 amino acids or more, short or long amino acids This provides a variant of dCas9 having a no-acid sequence.

[0102] In some embodiments, the Cas9 fusion protein provided herein is a Cas9 fusion protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, In other embodiments, the fusion protein provided herein does not contain a full-length Cas9 sequence. , including only one or more of the fragments. Exemplary amino acids of appropriate Cas9 domains and Cas9 fragments. Acid sequences are provided herein, and additional suitable sequences of the Cas9 domain and fragments are available to those skilled in the art. It should be obvious.

[0103] Additional Cas9 proteins (e.g., nuclease-inactive Cas9 (dCas9), Cas9 nickase (n Cas9), or nuclease-active Cas9, including its variants and homologs, It must be recognized that this is within the scope of disclosure. The example Cas9 protein is limited Without any, the following is provided. In some embodiments, the Cas9 protein is This is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is C It is as9 nickase (nCas9). In some embodiments, the Cas9 protein is a nucleate It is a ze-active Cas9.

[0104] The amino acid sequence of an example catalytically inactive Cas9 (dCas9) is as follows: DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0105] The amino acid sequence of an example catalytic Cas9 nikkase (nCas9) is as follows: DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0106] An example of a catalytically active Cas9 amino acid sequence is as follows: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0107] In some embodiments, Cas9 is an ancient element that constitutes a domain and kingdom of a single-celled prokaryotic microorganism. This refers to Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, Cas9 is, for example, , Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res . February 21, 2017. Refers to CasX or CasY described in doi: 10.1038 / cr.2017.21, and the relevant text The entire contents of the reference are incorporated herein by reference. Using genome degradation metagenomics And many CRISPR-Cas cells, including Cas9, which was first reported in the archaeal domain of life. The stem was identified. This diverse Cas9 protein is a little-studied nano-protein. It was discovered in KIA as part of the active CRISPR-Cas system. In bacteria, it had previously been discovered. Two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, and they are It is among the most compact systems discovered to date. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY Or it refers to a variant of CasY. Nucleic acid programmable DNA-binding protein (napDNAbp) and Other RNA-inducible DNA-binding proteins may also be used, provided that they fall within the scope of this disclosure. It should be understood.

[0108] In some embodiments, Cas9 is a Cas9 varian having specificity for the modified PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are used in Miller et al. al., Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat Bi This is described in otechnol (2020).doi.org / 10.1038 / s41587-020-0412-8, and the full content can be found at [reference]. More incorporated herein. In some embodiments, the Cas9 variant is a specific PAM requirement. It does not have any. In some embodiments, Cas9 variants, such as SpCas9 variants, It has specificity for NRNH PAM, where R is A or G, and H is A, C, or T. How many? In that embodiment, the SpCas9 variant is a PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, and It has specificity for CAC. In some embodiments, the SpCas9 variant is as follows Numbers such as 1114, 1134, 1135, 1137, 1139, 1151, 1180, are assigned to the reference array. 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 13 Amino acid substitution at positions 23, 1332, 1333, 1335, 1337, or 1339, or their corresponding positions. include. JPEG0007861090000010.jpg166162 (Single underline: HNH domain; Double underline: RuvC domain)

[0109] In some embodiments, the SpCas9 variant is numbered relative to the above reference sequence. Like 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or include amino acid substitutions at position 1337 or the corresponding position. In some embodiments, The SpCas9 variants are numbered 1114, 1134, 1135, etc., relative to the above reference sequence. ,1137,1139,1151,1180,1188,1211,1219,1221,1256,1264,1290,1318,1317, Includes amino acid substitutions at positions 1320, 1323, 1333, or their corresponding positions. Several implementations In this state, the SpCas9 variants are numbered 1114, 1131, etc., relative to the above reference sequence. ,1135,1150,1156,1180,1191,1218,1219,1221,1227,1249,1253,1286,1293, Includes amino acid substitutions at positions 1320, 1321, 1332, 1335, 1339, or their corresponding positions. In one embodiment, the SpCas9 variant is numbered with respect to the above reference sequence. ,1114,1127,1135,1180,1207,1219,1234,1286,1301,1332,1335,1337,1338, Includes amino acid substitution at position 1349. Exemplary amino acid substitutions and PAM-specific features of the SpCas9 variant. The sexes are shown in Tables A-D and Figure 8.

[0110] Table A JPEG0007861090000011.jpg140162

[0111] Table B JPEG0007861090000012.jpg124162

[0112] Table C JPEG0007861090000013.jpg97166

[0113] Table D JPEG0007861090000014.jpg83162

[0114] In certain embodiments, napDNAbps useful in the method of the present invention are known in the art and This includes, for example, the cyclic substitution described by Oakes et al., Cell 176, 254-267, 2019. The following are exemplary cyclic substitutions, where bold indicates sequences derived from Cas9, and italicized sequences. The arrows indicate linker sequences, and the underlined sequences indicate binuclear localization sequences. CP5(MSP "NGC = Pam variant containing mutations, normal Cas9 prefers NGG", PID = protein (Contains quality interaction domain and "D10A" nickase): JPEG0007861090000015.jpg169162

[0115] Polynucleotide programmable nucleotide bonds that can be incorporated into a base editor Main, non-exclusive examples include CRISPR protein-derived domains, restriction nucleases, and mega Nucleases, TAL nucleases (TALEN), and zinc finger nucleases (ZFN) ) is included.

[0116] In some embodiments, one of the nucleic acid programs of the fusion protein provided herein The napDNAbp (nap-binding DNA-binding protein) may be either a CasX or CasY protein. In some embodiments, napDNAbp is the CasX protein. In some embodiments, napDNAbp is a CasY protein. In some embodiments, napDNAbp is a naturally occurring CasX protein. Or the CasY protein contains at least 85%, at least 90%, at least 91%, and 92%, at least 93%, at least 94%, at least 95%, at least 96%, and at least Amino acids that are 97%, at least 98%, at least 99%, or at least 99.5% identical Includes a sequence. In some embodiments, napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, napDNAbp is the CasX or CasY protein described herein. At least 85%, at least 90%, at least 91%, at least 92%, and a small percentage of any of the following: At least 93%, at least 94%, at least 95%, at least 96%, at least 97%, small Contains amino acid sequences that are at least 98%, at least 99%, or at least 99.5% identical. Cas12b / C2c1, CasX, and CasY derived from other bacterial species may also be used in accordance with this disclosure. We should recognize this.

[0117] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS= Alicyclobacillus acido - terrestris (ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B stock) GN=c2c1 PE =1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI

[0118] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR associated Casx protein OS = Sulfolobus islandicus (HVE1 0 / 4 strain) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG

[0119] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandicus (REY15A strain) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0120] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAK PLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPK KPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMD EKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTD GTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIG RDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQA AKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKL AYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELS AELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYK SGKQPFVGAWQAFYKRRLKEVWKPNA

[0121] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR associated protein CasY [uncultured Parcubacteria bacterium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0122] The terms "conservative amino acid substitution" or "conservative mutation" refer to the phenomenon where a certain amino acid is common to both amino acids. This refers to replacing an amino acid with another amino acid that possesses the same characteristics. A functional method for defining this is the amino acid changes between corresponding proteins in homologous organisms. This involves analyzing normalized frequencies (Schulz, GE and Schirmer, RH, Principles). (Es of Protein Structure, Springer-Verlag, New York (1979)). According to such analysis... If amino acids within a group are preferentially exchanged with each other, the overall protein structure is thus determined. Define the group of amino acids that are most similar to each other in their effects on structure. This is possible (Schulz, GE and Schirmer, RH above). Non-limited conservative mutations Examples include, for example, lysine and the amino acid arginine, which can maintain a positive charge. Conversely, glutamic acid and aspartic acid can maintain a negative charge. Conversely, serine can maintain free OH for threonine, and also maintain free NH2. Possible amino acid substitutions include glutamine for asparagine.

[0123] The terms "coding sequence" or "protein coding sequence" are interchangeable within this specification. " refers to a segment of polynucleotides that codes for a protein. This region or The sequence is bounded by the start codon near the 5' end and by the stop codon near the 3' end. It is bounded. A code array can also be called an open reading frame.

[0124] As used herein, the terms “deaminase” or “deaminase domain” are used interchangeably. This refers to a protein or enzyme that catalyzes the deamination reaction. In some embodiments, Deaminase is an enzyme that catalyzes the hydrolytic deamination of adenine to hypoxanthine. It is a deaminase. In some embodiments, the deaminase is adenosine or Adenosine deamination catalyzes the hydrolytic deamination of adenine (A) to inosine (I). It is an enzyme. In some embodiments, the deaminase or deaminase domain is These are adenosine deaminases, which convert inosine or deoxyinosine to adenosine. It catalyzes the hydrolytic deamination of ¹⁶ or deoxyadenosine. In some embodiments, Adenosine deaminase is the enzyme that hydrolyzes adenosine in deoxyribonucleic acid (DNA). It catalyzes deamination. Adenosine deaminases provided herein (e.g., gene Manipulated adenosine deaminase (evolved adenosine deaminase) is used in bacteria, etc. It may be of any biological origin. In some embodiments, adenosine deaminase is Es cherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens It originates from bacteria such as Haemophilus influenzae or Caulobacter crescentus.

[0125] In some embodiments, adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. Minase domains are found in, for example, humans, chimpanzees, gorillas, monkeys, cattle, dogs, rats, Alternatively, it is a variant of a naturally occurring deaminase derived from organisms such as mice. In applied forms, deaminase or deaminase domains do not exist naturally. For example, In some embodiments, the deaminase or deaminase domain is a naturally occurring deaminase. At least 50%, at least 55%, at least 60%, at least 65%, less than 70% each, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 1%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, At least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, At least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99 They are identical by 0.7%, at least 99.8%, or at least 99.9%. For example, deaminase inhibitors The patent is in International PCT Application No. PCT / / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 0583 This is described in document No. 44 (WO 2017 / 070632), and the full contents of each of these documents are referenced in this specification. It is incorporated into the book. Komor, AC, et al., “Programmable editing of a target base i n genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “I improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Adva nces 3:eaao4774 (2017)), and Rees, HA, et al., “Base editing: precision ch emistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 See also Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1, all of these The contents are incorporated herein by reference.

[0126] "Detection" refers to identifying the presence, absence, or quantity of an analyte to be detected. In this embodiment, sequence modifications in polynucleotides or polypeptides are detected. In another embodiment, the presence of an indel is detected.

[0127] A "detectable label" is one that, when attached to a molecule of interest, exhibits spectroscopic, photochemical, This refers to a composition that makes the latter detectable by biochemical, immunochemical, or chemical means. For example, useful labels include radioactive isotopes, magnetic beads, metal beads, and colloidal particles. , fluorescent dyes, high electron density reagents, enzymes (e.g., those commonly used in ELISA), bio It contains tin, digoxigenin, or hapten.

[0128] "Disease" is any condition that damages or interferes with the normal functioning of cells, tissues, or organs. Or it means a disorder. In one embodiment, the disease is A1AD.

[0129] The term "effective dose" in this specification means a quantity sufficient to induce a desired biological response. This refers to the amount of biological activator used. In a particular embodiment, the effective amount is used to achieve the therapeutic effect. To modify the A1AT mutation in cells, a base editor system (e.g., Fusion proteins including programmable DNA-binding proteins, nucleic acid base editors, and gRNs. This is the amount of A). These therapeutic effects modify A1AD in all cells of the tissue or organ. It does not need to be enough to cause change, but about 1% to 5% of the cells present in the subject, tissue, or organ. Only 10%, 25%, 50%, 75%, or more may be modified. In one embodiment, The effective dose is sufficient to alleviate one or more symptoms of A1AD. The effective dose of the activator(s) used depends on the method of administration, the age and weight of the target, and overall It varies depending on the individual's health condition. Ultimately, the attending physician or veterinarian will determine the appropriate dosage and medication regimen. Determine the amount. Such an amount is called the "effective" amount. In one embodiment, the effective amount is the cell To introduce a modification to a gene of interest within (for example, in vitro or in vivo cells) A base editor of the present invention (e.g., a fusion protein comprising a programmable DNA-binding protein) This refers to the amount of protein, nucleic acid base editor, and gRNA. In one embodiment, the effective amount is the therapeutic amount. Achieve an effect (for example, reduce or control a disease or its symptoms or condition) This is the amount of base editor needed for ).

[0130] A "fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion is a reference nucleic acid. At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the total length of the molecule or polypeptide. It contains 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 20. Contains 0, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. It is possible to have.

[0131] "Guide RNA" or "gRNA" is specific to the target sequence and is a polynucleotide progesterone. It forms a complex with lammable nucleotide-binding domain proteins (e.g., Cas9 or Cpf1). This means a polynucleotide that can be used as a guide. In one embodiment, guide polynucleotide The first RNA is a guide RNA (gRNA). A gRNA can also exist as a complex of two or more RNAs. In some cases, it may exist as a single RNA molecule. gRNAs that exist as single RNA molecules are It is sometimes called single guide RNA (sgRNA), but "gRNA" refers to a single molecule. Alternatively, it is used interchangeably to refer to guide RNA that exists as a complex of two or more molecules. Typically, gRNAs that exist as a single RNA species (1) share homology with the target nucleic acid. (For example, a domain that directs the binding of the Cas9 complex to a target); and (2) the Cas9 protein It includes two domains, which are domains that bind to each other. In some embodiments, the domain (2 ) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, In one embodiment, domain (2) is Jinek et al., Science 337:816-821(2012)( The entire contents of are incorporated herein by reference to the same tracrRNA as provided in These are homologous. Another example of a gRNA (e.g., one containing domain 2) is "Switchable Cas9 Nu U.S. Provisional Patent Application USSN 61 / , filed September 6, 2013, titled "Clases and Uses Thereof" 874,682, and a document titled "Delivery System For Functional Nucleases" dated September 6, 2013. This can be found in the US provisional patent application USSN 61 / 874,746 filed in Japan, and the full contents of each of them refer to This is incorporated herein by means of. In some embodiments, the gRNA has domain (1) It may be called an "extended gRNA" if it contains two or more of the above (2). The extended gRNA is, As described in the document, it binds to two or more Cas9 proteins and targets two or more different regions. It binds to the target nucleic acid. gRNA contains a nucleotide sequence that complements the target site, and this is the target site. It mediates the binding of the nuclease / RNA complex to the site, and enhances the sequence specificity of the nuclease:RNA complex. Provides, as will be recognized by those skilled in the art, RNA polynucleotide sequences, e.g., gRN The A sequence contains more pyrimyl nucleotides than thymine (T), a nucleic acid base found in DNA polynucleotide sequences. It contains the nucleic acid base uracil (U), which is a din derivative. In RNA, uracil is adenine It forms base pairs with thymine and substitutes thymine during DNA transcription.

[0132] "Hybridization" refers to hydrogen bonding between complementary nucleic acid bases, as defined by Watson- It can be a click, Hougsteen, or reverse Hougsteen hydrogen bond. For example, Denine and thymine are complementary nucleic acid bases that form a pair by forming a hydrogen bond.

[0133] The term "base repair inhibitor," or "IBR," refers to nucleic acid repair enzymes, such as base excision repair enzymes. (BER) refers to a protein that can inhibit the activity of an enzyme. In some embodiments, IBR is an inhibitor of inosine base excision repair. An example of a base repair inhibitor is A PE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG Examples include inhibitors of UDG, hSMUGl, and hAAG. In some embodiments, IBR is En It is an inhibitor of doV or hAAG. In some embodiments, IBR is catalytically inactive E ndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is , Endo V or hAAG inhibitors. In some embodiments, base repair inhibitors are This is either catalytically inactive EndoV or catalytically inactive hAAG.

[0134] In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UG I) UGI is the inhibition of the uracil-DNA glycosylase base excision repair enzyme. It refers to the protein that can be produced. In some embodiments, the UGI domain is wild-type UGI or wild It contains fragments of live UGI. In some embodiments, the UGI protein provided herein is , containing UGI fragments and UGI or proteins homologous to UGI fragments. Several implementations In this context, base repair inhibitors are inhibitors of inosine base excision repair. Several implementations Morphologically, base repair inhibitors are "catalystically inactive inosine-specific nucleases." Alternatively, it is an "inactive inosine-specific nuclease." It is not bound by any particular theory. Although we do not want this to happen, catalytically inactive inosing glycosylases (e.g., alkyla Denning glycosylase (AAG) can bind to inosine, but it creates an abasic site. It is not possible to remove or eliminate inosine, and as a result, the newly formed inosine part It sterically blocks the DNA damage / repair mechanism. In some embodiments, it catalytically blocks Active inosine-specific nucleases can bind to inosine in nucleic acids, but the nucleus It does not cleave the acid. Non-restrictive and illustrative catalytically inactive inosine-specific nucleases. For example, a catalytically inactive alkyladenosing glycosylase (AAG nuclea) of human origin. EndoV nucleases), and catalytically inactive endonucleases V (EndoV nucleases) derived from, for example, E. coli. It contains (AAG nuclease). In some embodiments, catalytically inactive AAG nuclease is E1 This includes the 25Q mutation or the corresponding mutation in another AAG nuclease.

[0135] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0136] "Intein" is a term used to describe cutting out oneself and leaving behind fragments (extein). In a process known as protein splicing, tein is linked by peptide bonds. It is a fragment of protein that can bind together. Inteins are called "protein introns" It is also called a protein. The protein cuts itself out and connects to the rest of the protein. In this specification, Seth refers to "protein splicing" or "intrinsing-mediated protein splicing." This is called "protein splicing." In some embodiments, the precursor protein (in The intein of the intein-containing protein before intein-mediated protein splicing is , derived from two genes. Such an intein is referred to herein as a split intein. These are called inteins (for example, split inteins-N and split inteins-C). For example, cyano In bacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is derived from two separate genes. Encoded by the offspring dnaE-n and dnaE-c. Encoded by the dnaE-n gene In this specification, intein may be referred to as "inteiin N". The dnaE-c gene The Intein used may be referred to as "Intein C" in this specification.

[0137] Other intein systems may also be used. For example, the dnaE intein, namely Cfa-N For example, in the intein pairs of (split intein-N) and Cfa-C (split intein-C) The synthetic inteins based on this are described (for example, incorporated herein by reference). Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Use in accordance with this disclosure. Non-restrictive examples of possible intein pairs include: Cfa DnaE intein, Ssp GyrB Intein, Ssp DnaX Intein, Ter DnaE3 Intein, Ter ThyX Intein, Rma Dna B-intene and Cne Prp8-intene (for example, incorporated herein by reference) Examples include those described in U.S. Patent No. 8,394,604.

[0138] This provides exemplary nucleotide and amino acid sequences of the intein.

[0139] DnaE Intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGA ATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGG AAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAG ATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT

[0140] DNAE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDG QMLPIDEIFERELDLMRVDNLPN

[0141] DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTG CTCTGAAGAACGGATTCATAGCTTCTAAT

[0142] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0143] Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGA ATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAG AAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAG ATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA

[0144] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP

[0145] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCT TGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC

[0146] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0147] To connect the N-terminus of the divided Cas9 and the C-terminus of the divided Cas9, intein N and intein The tein C can be fused to the N-terminal and C-terminal portions of the segmented Cas9, respectively. In some embodiments, intein-N is fused to the C-terminus of the N-terminus portion of the divided Cas9, That is, it forms the structure N--[N-terminal portion of partitioned Cas9]-[intei-N]--C. In this embodiment, intein-C is fused to the N-terminus of the C-terminus portion of the divided Cas9, i.e., N- The structure [intene-C]--[C-terminal portion of split Cas9]-C is formed. Intein-mediated protein splicing for linking proteins (e.g., split Cas9) The mechanism of the process is, for example, incorporated herein by reference by Shah et al., Chem Sci. 2 As described in 014; 5(1):446-461, it is known in the art. The methods for designing and using the are publicly known in the art, for example, WO2014004336, WO201713 It is described by 2580, US20150344549 and US20180127780, each of which is The entire text is incorporated herein by reference.

[0148] The terms “isolated,” “purified,” or “biologically pure” refer to their natural state. When found in a certain state, the components that usually accompany it have been removed to varying degrees. It refers to quality. "Isolation" indicates the degree of separation from the original source or surrounding environment. "Purification" refers to quality. It shows a higher degree of separation than isolation. "Purified" or "biologically pure" protein This means that impurities can substantially affect the biological properties of proteins or cause other harmful consequences. Other substances are sufficiently removed so as not to cause any problems. That is, the nucleic acid of the present invention Alternatively, peptides, when produced by recombinant DNA technology, can be cellular or viral substances. Alternatively, if the culture medium is substantially absent, or if it is chemically synthesized, the chemical precursor may be present. If it is substantially free of other chemicals, it is considered purified. Purity and uniformity are standard. In terms of type, analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography, are used. Determined using matrixing. The term "purified" means that nucleic acids or proteins have been electrolyzed. This could essentially mean the formation of a single band in a pneumatophoresis gel. Modification, for example, For proteins that can undergo oxidation or glycosylation, different modifications are, This can result in different isolated proteins that can be purified separately.

[0149] "Isolated polynucleotides" are derived from the naturally occurring genomes of organisms from which the nucleic acid molecules of the present invention originate. In this context, it refers to nucleic acids (e.g., DNA) that do not contain genes adjacent to the gene in question. Therefore, this The term can refer to, for example, a vector incorporated into a plasmid or virus that replicates autonomously. Incorporated; integrated into the genomic DNA of prokaryotes or eukaryotes; or independently of other sequences. A separate molecule (e.g., cDNA produced by PCR or restriction endonuclease digestion) This includes recombinant DNA that exists as a genome or cDNA fragment. Furthermore, this term means High-frequency transcription of RNA molecules from DNA molecules, as well as encoding further polypeptide sequences. It contains recombinant DNA, which is part of the hybrid gene.

[0150] "Isolated polypeptide" refers to the polypeptide that is separated from its naturally occurring components in this invention. This refers to polypeptides. Typically, a polypeptide is a polypeptide that, in its natural state, If the combined protein and naturally occurring organic molecules do not make up at least 60% by weight, It is isolated. Preferably, the preparation contains at least 75%, more preferably 90%, by weight. Preferably, 99% is the polypeptide of the present invention. The isolated polypeptide of the present invention is For example, extraction from natural sources, expression of recombinant nucleic acids encoding such polypeptides. Alternatively, it can be obtained by chemically synthesizing proteins. Purity is determined by any appropriate method. Essential methods, such as column chromatography, polyacrylamide gel electrophoresis, and This can be measured by HPLC analysis.

[0151] As used herein, the term "linker" refers to two molecules or parts, for example, a Two components of a protein complex or ribonuclear complex, or two components of a fusion protein Main, for example, a polynucleotide programmable DNA binding domain (e.g., dCas9) and Dear Minase domain (for example, adenosine deaminase as described in PCT / US19 / 44935) A covalent linker that links adenosine deaminase and cytidine deaminase. It can refer to a linker (for example, a covalent bond), a non-covalent linker, a chemical group, or a molecule. Kerr connects different components or different parts of components of a base editor system. This is possible. For example, in some embodiments, the linker is a polynucleotide pro Guide polynucleotide binding domain and der The catalytic domain of the minase can be linked. In some embodiments, the linker CRISPR polypeptides and deaminases can be linked. Several embodiments The linker can then link Cas9 and deaminase. Several implementations In this state, the linker can link dCas9 and deaminase. In the application form, the linker can link nCas9 and deaminase. In this embodiment, the linker links the guide polynucleotide and the deaminase. This is possible. In some embodiments, the linker deaminates the base editor system. Linking the chemical components and the polynucleotide programmable nucleotide binding components. It is possible. In some embodiments, the linker deaminates the base editor system. RNA-binding portion of the nucleotide component and polynucleotide programmable nucleotide-binding portion Components can be linked together. In some embodiments, the linker is a base editor. RNA binding portion and polynucleotide programmable portion of the deamination component of the system The RNA-binding portion of the cleotide-binding component can be linked. The linker has two groups. A molecule, or a molecule, located between or adjacent to other parts, and bonded covalently or noncovalently. They are linked to each other through bonded interactions, and therefore can be linked together. In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. Obtain. In some embodiments, the linker may be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker is R It may be an NA linker. In some embodiments, the linker binds to the ligand. It may include an aptamer that can perform carbonation. In some embodiments, the ligand is a carbon atomizer. It may be a substance, peptide, protein, or nucleic acid. In some embodiments, a linker It may include aptamers that originate from the riboswitch. The riboswitch is the theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch. Adenosine cobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch Switches, SAH riboswitches, flavin mononucleotide (FMN) riboswitches, tetrahedronyl riboswitches Drofolate Riboswitch, Lysine Riboswitch, Glycine Riboswitch, Purine Riboswitch The GlmS riboswitch or the Prequeosin-1 (PreQ1) riboswitch can be selected. In some embodiments, the linker is a polypeptide or polypeptide ligand, etc. It may include an aptamer bound to the protein domain of the protein. In some embodiments, poly The peptide ligand consists of a K homology (KH) domain, an MS2 coat protein domain, and a PP7 coat. Protein domain, SfMu Com coated protein domain, sterile alpha motif , telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and It may be the Sm7 protein or RNA recognition motif. In some embodiments, the polyp Plitter ligands can be part of a base editor system. For example, nucleic acid salts The basic editing components may include a deaminase domain and an RNA recognition motif.

[0152] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., peptides) It may be a dendritic or protein. In some embodiments, the linker is about 5 to 100 cm in length. Amino acids, for example, with lengths of approximately 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90 or 90-100 amino acids It may be acidic. In some embodiments, the linker is about 100-150, 150-200, 200 in length. Possible ranges: ~250, 250~300, 300~350, 350~400, 400~450, or 450~500 amino acids. Longer or shorter linkers are also possible.

[0153] In some embodiments, the linker is an RNA programmer containing a Cas9 nuclease domain. gRNA-binding domains of nucleases and nucleic acid editing proteins (e.g., adenosine) The catalytic domain of deaminase is linked. In some embodiments, the linker is dCas9 and link nucleic acid editing proteins. For example, a linker can link two groups, molecules, or other It is positioned between or adjacent to the parts and connected to each other via covalent bonds. Therefore, the two are linked. In some embodiments, the linker is an amino acid or multiple It is an amino acid (e.g., a peptide or protein). In some embodiments, phosphorus A KAR is an organic molecule, group, polymer, or chemical part. In some embodiments, Linkers are 5 to 200 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 1 5, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 9 0, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190 , or 200 amino acids. Longer or shorter linkers are also intended.

[0154] In some embodiments, the domain of the nucleic acid base editor is SGGSSGSETPGTSESATPESSG GS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTE Amino acids PSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS They are fused via a linker containing the sequence. In some embodiments, a nucleic acid base editor The domain contains the amino acid sequence SGSETPGTSESATPES, which may also be called the XTEN linker. The fusion occurs via a linker. In some embodiments, the linker includes the amino acid sequence SGGS. In some embodiments, the linker is (SGGS) n (GGGS) n (GGGGS) n , (G) n、 (EAAAK) n (GGS)n , SGSETPGTSESATPES, or (XP) n motifs, or any combination thereof. The set includes n, where n is an integer between 1 and 30, and X is any amino acid. How many? In that embodiment, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. ru.

[0155] In some embodiments, the linker is 24 amino acids long. The linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, The linker is 40 amino acids long. In some embodiments, the linker is amino The acid sequence SGGSSGGSSGSETPGTESATPESSGGSSGGSSGSSGGS is included. In some embodiments, phosphorus The linker has a length of 64 amino acids. In some embodiments, the linker is the amino acid sequence SG Includes GSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids long. The amino acid sequence is PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GT Includes STEPSEGSAPGTSESATPESGPGSEPATS

[0156] A "marker" is a substance that has altered expression levels or activity associated with a disease or disorder. It means a protein or polynucleotide.

[0157] As used herein, the term "mutation" means a mutation within a sequence, for example, in a nucleic acid or A residue within a amino acid sequence is substituted by another residue, or one or more residues within a sequence are substituted. This refers to the deletion or insertion of a residue. Mutations, in this specification, typically involve identifying the original residue. Next, the position of the residues within the sequence is identified, and the identity of the newly substituted residues is determined. Therefore, it is described. For the production of amino acid substitutions (mutations) provided herein. Various methods are well known in this field, for example, Green and Sambrook, Mo lecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Pre Provided by ss, Cold Spring Harbor, NY (2012). Several embodiments Therefore, the base editor of this disclosure can handle a significant number of unintentional mutations, for example, unintentional point Without generating mutations, the "intended" nucleic acids (e.g., nucleic acids within the target genome) Mutations, such as point mutations, can be efficiently generated. In some embodiments, The intended mutation is specifically designed to produce that intended mutation. A specific base editor (e.g., adenosyl) is bound to a guide polynucleotide (e.g., gRNA). This is a mutation caused by a base editor.

[0158] Generally, this is done or identified in a sequence (for example, the amino acid sequence described herein). The mutation is applied to the reference (or wild-type) sequence, i.e., the sequence that does not contain the mutation. Numbered in relation to the reference sequence. Those skilled in the art will know that in the amino acid and nucleic acid sequences relative to the reference sequence You will easily understand how to determine the location of the mutation.

[0159] The term "non-conservative mutation" refers to amino acid substitutions between different groups, for example. For example, lysine for tryptophan, or phenylalanine for serine. In this case, non-conservative amino acid substitutions interfere with or inhibit the biological activity of the functional variant. It is preferable that it does not cause harm. Non-conservative amino acid substitutions affect the biological activity of the functional variant. The functional variant enhances the biological activity of the protein, increasing it compared to the wild-type protein. It can be done.

[0160] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to the structure of a protein. This refers to an amino acid sequence that promotes translocation into the cell nucleus. Nuclear localization sequences are used in this technology. It is publicly known, for example, filed on November 23, 2000, and registered as WO / 2001 / 038547 on May 31, 2001. It is described in Plank et al.'s published international PCT application PCT / EP 2000 / 011690, and its contents The disclosure of exemplary nuclear localization sequences is incorporated herein by reference. In terms of application methods, NLS is, for example, described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.417 This is an optimized NLS as described in 2. In some embodiments, the NLS is an optimized NLS. No acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENG Includes RKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0161] As used herein, the terms “nucleic acid” and “nucleic acid molecule” refer to nucleic acid bases and acids. Compounds containing a sexual moiety, such as nucleosides, nucleotides, or polymers of nucleotides - refers to polymeric nucleic acids, for example, nucleic acid molecules containing three or more nucleotides. A linear molecule in which adjacent nucleotides are linked to each other via phosphodiester linkages. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides). This refers to a nucleic acid (or nucleoside). In some embodiments, "nucleic acid" refers to three or more individuals. This refers to an oligonucleotide chain containing several nucleotide residues. When used herein, The terms "oligonucleotide" and "polynucleotide" refer to polymers of nucleotides (e.g.) For example, it can be used interchangeably to refer to a sequence of at least three nucleotides. In that embodiment, “nucleic acid” includes RNA and single-stranded and / or double-stranded DNA. Nucleic acids include, for example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, etc. In relation to smids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules, It can exist in [location]. On the other hand, nucleic acid molecules include, for example, recombinant DNA or RNA, artificial chromosomes, and manipulated [unclear]. A genome, or a fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or non-genome. It may be a naturally occurring molecule containing a nucleotide or nucleoside. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms are used to refer to nucleic acid ana. This includes analogs with logs, for example, a phosphodiester backbone. Nucleic acids are from natural sources or It is purified using a recombinant expression system and purified as needed, or chemically synthesized. It is possible to do, etc. In the case of chemically synthesized molecules, nucleic acids, in appropriate cases, for example For example, nucleos such as chemically modified bases or sugars, and analogs having main chain modifications. May contain side analogs. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise specified. In some embodiments, nucleic acids are natural nucleosides (e.g., adenosine, thymidine, Guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyg Anosine and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine) 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcythididine C5-bromouridine, C5-fluorouridine, C5-iodouridine C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-amino Adenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxo Guanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically Modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose) , ribose, 2'-deoxyribose, arabinose, and hexose); and / or Is it a modified phosphate group (e.g., phosphorothioate and 5'-N-phosphoamidite linkage)? , or including them.

[0162] The term "nucleic acid programmable DNA-binding protein" or "napDNAbp" is a term used by "po It is interchangeable with the "renucleotide programmable nucleotide-binding domain". , a guide nucleic acid or guide polynucleotide that guides napDNAbp to a specific nucleic acid sequence (e.g.) It refers to proteins that associate with nucleic acids (e.g., DNA or RNA), such as gRNA. In this embodiment, the polynucleotide programmable nucleotide binding domain is polynucleotide It is a cleotide programmable DNA-binding domain. In some embodiments, polynucleotides The rheotide programmable nucleotide-binding domain is polynucleotide programmable. It is an RNA-binding domain. In some embodiments, it is polynucleotide programmable. The nucleotide-binding domain is the Cas9 protein. The Cas9 protein binds to the guide RNA. It can associate with a guide RNA that guides the Cas9 protein to a specific complementary DNA sequence. In some embodiments, napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9. Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of lambable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas 12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Examples include Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, C as9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl , Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy 3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2 , Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17 , Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Cs d1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effect Cas effector protein, V-type Cas effector protein, VI-type Cas effector protein, CAR This includes F, DinG, their homologs, or their modified or manipulated versions. Other nucleic acid programmable DNA-binding proteins are also included, but are not specifically listed in this disclosure. There is a possibility of this, but it is within the scope of this disclosure. For example, Makarova et al. “Classification a nd Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:3 25-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science. See aav7271 (all of its contents are incorporated herein by reference).

[0163] The terms “nucleic acid base,” “nitrogen base,” or “base” are used interchangeably in this specification. This refers to nitrogen-containing biological compounds that form nucleosides, and nucleosides are nucleosides. It is a component of ribonucleotides. The ability of nucleic acid bases to form base pairs and stack on top of each other is directly related to ribonucleotides. It gives rise to long helical structures such as nucleic acids (RNA) and deoxyribonucleic acid (DNA). Adenine The five nucleic acid bases (A), cytosine (C), guanine (G), thymine (T), and uracil (U) are the main ones. These are called primary bases or canonical nucleic acid bases. Adenine and Guanine is derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. RNA may also contain other modified (non-major) bases. Non-restrictive exemplary modified nucleic acid bases It contains hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, and 5-methyl It contains lucitosine (m5C) and 5-hydromethylcytosine. Hypoxanthine and xan Chin can be produced in the presence of a mutagen, and both involve deamination (removal of the amine group). It is produced by substitution of a bonyl group. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil is produced by the deamination of cytosine. To obtain. A "nucleoside" is a nucleic acid base and a pentose sugar (either ribose or deoxyribose). It consists of (k). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5 -Methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxy It contains siuridine and deoxycytidine. Nucleosides with modified nucleic acid bases Examples include inosine (I), xanthosine (X), 7-methylguanosine (m7G), and dihydrouridine. It contains nucleotides (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). " is a nucleic acid base, a pentose (either ribose or deoxyribose), and at least It also consists of one phosphate group.

[0164] The terms "nucleic acid base editing domain" or "nucleic acid base editing protein" are defined herein. When used in this context, adenine (or adenosine) is converted to hypoxanthine (or i Deamination of (nosine), as well as addition and insertion of non-templated nucleotides. Furthermore, a protein or enzyme that can catalyze nucleic acid base modification in RNA or DNA. It refers to. In some embodiments, the nucleic acid base editing domain is a deaminase domain (for example) (For example, adenine deaminase or adenosine deaminase). Several implementation forms In this state, the nucleic acid base editing domain has more than one deaminase domain (for example, adenine Deaminase, or adenosine deaminase and cytidine or cytosine deaminase (For example, as described in PCT / US19 / 44935). In some embodiments, nucleic acid bases The editing domain may be a naturally occurring nucleic acid base editing domain. Several embodiments So, are nucleic acid editing domains modified or advanced from naturally occurring nucleic acid editing domains? It may be a modified nucleic acid base editing domain. Nucleic acid base editing domains are found in bacteria, humans, and Chinese. It is derived from any organism such as pansies, gorillas, monkeys, cows, dogs, rats, or mice. It is possible.

[0165] When used in this specification, "to obtain" as in "to obtain the agent" means to obtain the agent. To synthesize, purchase, generate, prepare, or otherwise obtain This includes the following.

[0166] As used herein, “patient” or “subject” means a person diagnosed with a disease or disorder. to have or have the risk of developing such things, or may have them This refers to a mammalian subject or individual suspected of being a potential candidate for development. In some embodiments, the term "Patient" refers to a mammalian subject that is more likely than average to develop a disease or disorder. (Exemplary) Patients include humans, non-human primates, cats, dogs, pigs, cats, horses, camels, llamas, Goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs) and native animals Other mammals that may benefit from the therapies disclosed in the document may include exemplary individuals. The patient may be male and / or female.

[0167] In this specification, "patients who need it" or "subjects who need it" refers to those with a disease. or diagnosed with a disorder, have it, are at risk of having it, are susceptible to it , patients who have been predetermined to have it, or who are suspected to have it. They are referred to as individuals.

[0168] "Pathogenic mutation," "Pathogenic variant," "Disease-causing mutation," "Disease-causing variant" The terms "ant," "harmful mutation," or "predisposing mutation" refer to certain diseases. This refers to genetic modification or mutation that increases an individual's susceptibility or predisposition to a disorder. In some embodiments, the pathogenic mutation is a protein encoded by a gene. At least one wild-type amino acid in it is replaced by at least one pathogenic amino acid. This includes things that have been done.

[0169] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents. These are interchangeable as used herein and are linked together by peptide (amide) bonds. This term refers to a polymer of amino acid residues of any size, structure, or function. Refers to proteins, peptides, or polypeptides. Typically, proteins, peptides, and polypeptides. A polypeptide is at least 3 amino acids in length. Proteins, peptides, or A polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide are conjugated, ductal For activating or other modifications, for example, carbohydrate groups, hydroxyl groups, phosphate groups, and phosphate groups. By adding chemical entities such as necil groups, isofarnesyl groups, fatty acid groups, and linkers, They can be modified. Proteins, peptides, or polypeptides can also be modified, even as single molecules. Often, or sometimes as multimolecular complexes, these may be proteins, peptides, or polypeptides. These may be mere fragments of naturally occurring proteins or peptides. Butides, or polypeptides, can be naturally occurring, recombinant, or synthetic. Or any combination thereof. When used herein, the term "fusion ta "protein" refers to a protein containing protein domains derived from at least two different proteins. This refers to a hybrid polypeptide. One protein is the amino terminus (N terminus) of the fusion protein. It can be located at the end (part) or carboxy-terminal (C-terminal) of a protein, and therefore it They form amino-terminal fusion proteins or carboxy-terminal fusion proteins. Proteins have different domains, for example, nucleic acid binding domains (for example, protein binding to target sites) The Cas9 gRNA-binding domain that guides the binding of substances, and the nucleic acid cleavage domain, or nucleic acid editing domain. It may contain a catalytic domain of protein. In some embodiments, the protein is a protein The qualitative parts, for example, the amino acid sequence that constitutes the nucleic acid binding domain, and the organic compound, for example The compound may also be a nucleic acid cleavage agent. In some embodiments, the protein is It is complexed with nucleic acids, such as RNA or DNA, or associated with nucleic acids. Any protein provided in this document may be produced by any method known in the art. This is possible. For example, the proteins provided herein are recombinant protein expression and It can be produced through purification, which results in a fusion protein containing a peptide linker. It is particularly suitable. Methods for the expression and purification of recombinant proteins are well known. Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring As described by Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). This includes, and its entire contents are incorporated herein by reference.

[0170] Polypeptides and proteins disclosed herein (functional portions and their functional parts) (Containing rianto) contains synthetic amino acids instead of one or more naturally occurring amino acids. This can be done. Such synthetic amino acids are known in the art, for example, aminocyclo Hexanecarboxylic acid, norleucine, α-aminon-decanoic acid, homoserine, S-acetylated Minomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophen Lalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine Nylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-Naphthylalanine, Cyclohexylalanine, Cyclohexylglycine, Indoline 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, A Minomaronate monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6 -Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocycline Roxenic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbolol) Nan-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylamine Examples include ranine and α-tert-butylglycine. Polypeptides and proteins are This can be associated with post-translational modifications of one or more amino acids in the polypeptide construct. Non-limiting examples of post-modification include acylation, which includes phosphorylation, acetylation, and formylation. , glycosylation (including N-linking and O-linking), amidation, hydroxylation, methylation and Alkylation including biethylation, ubiquitination, addition of pyrrolidone carboxylic acid, disulfide Cross-linking, sulfation, myristoylation, palmitoylation, isoprenylation, farnesia Examples include geranylation, glypiation, lipoylation, and iodation. It is possible.

[0171] As used herein in relation to proteins or nucleic acids, the term “recombinant” means a naturally occurring organism. It refers to proteins or nucleic acids that do not exist naturally and are products of artificial manipulation. For example, several In the embodiment, the recombinant protein or nucleic acid molecule is compared to any naturally occurring sequence. And at least one, at least two, at least three, at least four, at least five, an amino acid or nucleotide sequence containing at least six or at least seven mutations Includes.

[0172] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%. .

[0173] "Reference" means a standard or contrasting condition. In one embodiment, the reference These are wild-type or healthy cells. In other embodiments, without limitation, reference is to the test conditions. Unexposed, or placebo or saline, culture medium, buffer, and / or Untreated cells are exposed to a control vector that does not contain the polynucleotide of interest. .

[0174] A "reference sequence" is a predefined sequence used as the basis for sequence comparison. A reference sequence is: This can be a subset or the whole of a specific sequence; for example, a full-length cDNA or gene sequence A segment of a column, or a complete cDNA or gene sequence. For polypeptides, see Reference The length of a lipeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids. It is at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. Regarding nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, and less Each has approximately 60 nucleotides, at least approximately 75 nucleotides, approximately 100 nucleotides, or approximately 300 nucleotides. Cleotide, or approximately those values, or any integer between them. In this embodiment, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments... The reference sequence is a polynucleotide sequence that encodes the wild-type protein.

[0175] The terms "RNA programmable nuclease" and "RNA-inducible nuclease" refer to cleavage It is used in conjunction with one or more non-target RNAs (e.g., by binding to or associating with them). In some embodiments, the RNA programmable nuclease is a complex with RNA, Clease: Can be called an RNA complex. Typically, the bound RNA(s) are guide RNAs ( gRNA is called a gRNA. gRNA can exist as a complex of two or more RNAs, or as a single RN. It can also exist as an A molecule. gRNAs that exist as a single RNA molecule are single guide It is sometimes called RNA (sgRNA), but "gRNA" can refer to a single molecule or two or more molecules. It is interchangeable to refer to the guide RNA, which exists as one of the molecular complexes. Typically, gRNAs that exist as a single RNA species are (1) drugs that share homology with the target nucleic acid. (1) Main (for example, directing the binding of the Cas9 complex to the target); and (2) Cas9 protein It includes two domains, namely the domains to be joined. In some embodiments, domain (2) This corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, several In that embodiment, domain (2) is Jinek et al., Science 337:816-821(2012)(that The tracrRNA is identical to the one provided in (the entire content of which is incorporated herein by reference) They are homologous. Another example of a gRNA (e.g., one containing domain 2) is "Switchable Cas9 Nucle U.S. Provisional Patent Application USSN61, filed on September 6, 2013, is titled "eases and Uses Thereof". / 874,682 and titled "Delivery System For Functional Nucleases" were published on September 6, 2013. This can be found in the filed U.S. provisional patent application USSN61 / 874,746, and each of them The entire contents of the above are incorporated herein by reference. In some embodiments, gRNA is It may be said to be an "extended gRNA" that contains two or more domains (1) and (2). The elongated gRNA, as described herein, for example, with two or more Cas9 proteins It binds and binds to the target nucleic acid in two or more different regions. gRNA is a nucleic acid that complements the target site. It contains a rheotide sequence, which mediates the binding of the nuclease / RNA complex to the target site, Nucleases provide sequence specificity for RNA complexes.

[0176] In some embodiments, RNA programmable nucleases are used in (CRISPR-related systems) Cas9 endonucleases, such as Cas9 (Casnl) derived from Streptococcus pyogenes. (For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes.") Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Na jar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, Mc Laughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA mat uration by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chy linski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J. (See Charpentier E., Nature 471:602-607 (2011)).

[0177] RNA programmable nucleases (e.g., Cas9) use RNA to target DNA cleavage sites: Because they use DNA hybridization, these proteins are, in principle, It is possible to target any sequence identified by the doRNA. Due to site-specific cleavage Methods using RNA programmable nucleases like Cas9 (for example, to modify genomes) (for this purpose) is known in the art (e.g., Cong, L. et al., Multiplex genom e engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas sy stem. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Gen ome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic ac ids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes u See CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013). (The full contents of each of these are incorporated herein by reference.)

[0178] The term "single nucleotide polymorphism (SNP)" refers to a mutation in a single nucleotide that occurs at a specific position in the genome. Yes, each mutation exists to a recognizable extent within the population (e.g., >1%). For example, human hair At certain base positions in the nucleotide, C nucleotides can appear in most individuals, but only a small number of individuals... In the body, that position is occupied by A. This means that there is an SNP at this particular position, and it is either C or A. This means that the allele is of this magnitude, with two possible nucleotide mutations. SNPs are in disease. This underlies differences in susceptibility. The severity of the disease and the body's response to treatment are also influenced by genetic variations. This is a manifestation of the following: SNPs are the coding regions, non-coding regions, or inter-gene regions of a gene. It can exist in a region (the region between genes). In some embodiments, it is present in the coding sequence. SNPs, due to the degeneracy of the genetic code, do not necessarily alter the amino acid sequence of the resulting protein. Don't let that happen. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs are tampers Non-synonymous SNPs do not affect the protein sequence, but they alter the amino acid sequence of proteins. There are two types of NPs: missense and nonsense. SNPs that are not located in protein-coding regions are... gene splicing, transcription factor binding, messenger RNA degradation, or non-coding R It can affect the sequence of NAs. Gene expression affected by this type of SNP is eSN P (expressed SNP) is called a single nucleotide variant (SNV) and can be located upstream or downstream of a gene. Somatic single nucleotide mutations are single nucleotide mutations with no limit on frequency and can occur in somatic cells. It can also be called a modification.

[0179] "Specifically binding" means recognizing polypeptides and / or nucleic acid molecules of the present invention. It binds to this, but does not substantially recognize other molecules in the sample, such as a biological sample, and Nucleic acid molecules, polypeptides, or complexes thereof that do not bind (e.g., nucleic acid programmes) This refers to a potential DNA-binding domain and guide nucleic acid, compound, or molecule.

[0180] Nucleic acid molecules useful in the method of the present invention are polypeptides or fragments thereof. It contains any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. It is not necessary, but it typically shows substantial identity. "Substantial identity" for endogenous sequences. Polynucleotides typically consist of at least one strand of a double-stranded nucleic acid molecule and a hive. It can be reduced. Nucleic acid molecules useful in the method of the present invention are the polypeptide of the present invention. It includes any nucleic acid molecule encoding a cytoplasm or a fragment thereof. Such nucleic acid molecules are endogenous It does not need to be 100% identical to the nucleic acid sequence, but typically shows substantial identity. In contrast, polynucleotides that possess "substantial identity" are typically composed of a small number of double-stranded nucleic acid molecules. Even without one, it can hybridize with one chain. "Hybridizing" means multiple Complementary polynucleotide sequences under various stringency conditions (e.g., as described herein) This means forming a pair that creates a double-stranded molecule between (the genes) or between (parts of) them. For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (See Methods Enzymol. 152:507, 1987).

[0181] For example, stringent salt concentrations are typically less than 750 mM NaCl and 75 mM citrate. Trisodium citrate, preferably less than 500 mM NaCl and 50 mM trisodium citrate, Preferably, the NaCl is less than about 250 mM and the trisodium citrate is 25 mM. Gency hybridization is obtained in the absence of organic solvents, such as formamide. This is possible, while high stringency hybridization is at least about 35% It can be obtained in the presence of formamide, more preferably at least about 50% formamide. The stringent temperature conditions are typically at least about 30°C, more preferably less than The temperature will also include approximately 37°C, most preferably at least approximately 42°C. The control time, the concentration of the surfactant, such as sodium dodecyl sulfate (SDS), and the carrier. Various additional parameters, such as the inclusion or exclusion of DNA, are well known to those skilled in the art. By combining these various conditions as needed, various levels of stringing can be achieved. A phenotype is achieved. In one embodiment, hybridization is performed at 30°C with 750 mM N This occurs in aCl, 75 mM trisodium citrate, and 1% SDS. In another embodiment, hybrid Disization was performed at 37°C with 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, and 35% chlorine. This occurs in lumamide and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another embodiment, Hybridization was performed at 42°C using 250 mM NaCl, 25 mM trisodium citrate, and 1% SDS. This occurs in 50% formamide and 200 μg / ml ssDNA. Useful variations of these conditions The reason will be readily apparent to those skilled in the art.

[0182] In most applications, the washing process following hybridization is also done by stringen The conditions differ. Washing stringency conditions are defined by salt concentration and temperature. This can be done. As mentioned above, washing stringency reduces the salt concentration or warm It can be increased by raising the degree. For example, strips for the washing process The ideal salt concentration is preferably less than 30 mM NaCl and 3 mM trisodium citrate. The most preferred components are less than approximately 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions for the washing process are usually at least about 25°C, more preferably. The temperature includes at least about 42°C, and more preferably at least about 68°C. In the application method, the washing step is performed at 25°C with 30 mM NaCl, 3 mM trisodium citrate, and 0.1% This is carried out in SDS. In a more preferred embodiment, the washing step is performed at 42°C with 15 mM NaCl, 1.5 This is carried out in mM trisodium citrate and 0.1% SDS. In a more preferred embodiment, washing The purification process was carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art, for example, Benton and Dav is (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 7 2:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Int. erscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Technique) Ques, 1987, Academic Press, New York; and Sambrook et al., Molecular Cloning: As described in A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York. Yes, they are.

[0183] "Split" means to be divided into two or more pieces.

[0184] "Split Cas9 protein" or "split Cas9" refers to two separate nucleotides. The Cas9 protein is provided as an N-terminal and C-terminal fragment encoded by the sequence. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein are splices. It can be remodeled to form a "reconstituted" Cas9 protein. In certain embodiments, Cas9 For example, see Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949. As described in 2014, or Jiang et al. (2016) Science 351: 867-871. PDB fil e: As described in 5F9R (each incorporated herein by reference), protein It is divided into two fragments within a disordered region of quality. In some embodiments, the protein This is within the region of SpCas9 between approximately amino acids A292-G364, F445-K483, or E565-T637. In any C, T, A, or S of, or in any other Cas9, Cas9 variant (for example) It splits into two fragments at the corresponding position in nCas9, dCas9, or other napDNAbp. In some embodiments, the protein is SpCas9 T310, T313, A456, S469, or C57 In step 4, it is divided into two fragments. In some embodiments, the protein is divided into two fragments. The process is referred to as "splitting" the protein.

[0185] In other embodiments, the N-terminal portion of the Cas9 protein is S. pyogenes Cas9 wild-type (SpCas9). )(NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2) Amino acids 1-573 or 1 ~637, or including the corresponding position / mutation, the C-terminal portion of the Cas9 protein is in the SpCas9 field It contains the portion of amino acids 574-1368 or 638-1368 in its raw form.

[0186] The C-terminal portion of the fragmented Cas9 protein is joined to the N-terminal portion of the fragmented Cas9 protein to form a complete Cas9 protein. It can form quality. In some embodiments, the C-terminal portion of the Cas9 protein is C It starts from where the N-terminus of the as9 protein ends. In this embodiment, the C-terminal portion of split Cas9 is the portion of spCas9 from amino acids (551-651) to 1368. Includes. "(551~651)~1368" refers to amino acids between amino acids 551 and 651 (including both ends). This means it starts and ends at amino acid 1368. For example, the C-terminal portion of split Cas9 is spCas 9 amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557 ~1368, 558~1368, 559~1368, 560~1368, 561~1368, 562~1368, 563~1368, 564~1 368, 565~1368, 566~1368, 567~1368, 568~1368, 569~1368, 570~1368, 571~1368 , 572~1368, 573~1368, 574~1368, 575~1368, 576~1368, 577~1368, 578~1368, 5 79~1368, 580~1368, 581~1368, 582~1368, 583~1368, 584~1368, 585~1368, 586 ~1368, 587~1368, 588~1368, 589~1368, 590~1368, 591~1368, 592~1368, 593~1 368, 594~1368, 595~1368, 596~1368, 597~1368, 598~1368, 599~1368, 600~1368 , 601~1368, 602~1368, 603~1368, 604~1368, 605~1368, 606~1368, 607~1368, 6 08~1368, 609~1368, 610~1368, 611~1368, 612~1368, 613~1368, 614~1368, 615 ~1368, 616~1368, 617~1368, 618~1368, 619~1368, 620~1368, 621~1368, 622~1 368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368 , 630~1368, 631~1368, 632~1368, 633~1368, 634~1368, 635~1368, 636~1368, 6 37-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644 ~1368, 645~1368, 646~1368, 647~1368, 648~1368, 649~1368, 650~1368, or It may include any one of the parts from 651 to 1368. In some embodiments, the split Cas9 tamper The C-terminal portion of the protein contains the amino acid segments 574-1368 or 638-1368 of SpCas9.

[0187] "Serpin1A polynucleotide" is a nucleic acid that encodes the A1AT protein or a fragment thereof. Meaning "child." An exemplary Serpin1A polynucleo is available under NCBI deposit number NM_000295. The sequence of CHIDS is provided below: 1 acaatgactc ctttcggtaa gtgcagtgga agctgtacac tgcccaggca aagcgtccgg 61 gcagcgtagg cgggcgactc agatcccagc cagtggactt agcccctgtt tgctcctccg 121 ataactgggg tgaccttggt taatattcac cagcagcctc ccccgttgcc cctctggatc 181 cactgcttaa atacggacga ggacagggcc ctgtctcctc agcttcaggc accaccactg 241 acctgggaca gtgaatcgac aatgccgtct tctgtctcgt ggggcatcct cctgctggca 301 ggcctgtgct gcctggtccc tgtctccctg gctgaggatc cccagggaga tgctgcccag 361 aagacagata catcccacca tgatcaggat cacccaacct tcaacaagat cacccccaac 421 ctggctgagt tcgccttcag cctataccgc cagctggcac accagtccaa cagcaccaat 481 atcttcttct ccccagtgag catcgctaca gcctttgcaa tgctctccct ggggaccaag 541 gctgacactc acgatgaaat cctggagggc ctgaatttca acctcacgga gattccggag 601 gctcagatcc atgaaggctt ccaggaactc ctccgtaccc tcaaccagcc agacagccag 661 ctccagctga ccaccggcaa tggcctgttc ctcagcgagg gcctgaagct agtggataag 721 tttttggagg atgttaaaaa gttgtaccac tcagaagcct tcactgtcaa cttcggggac 781 accgaagagg ccaagaaaca gatcaacgat tacgtggaga agggtactca agggaaaatt 841 gtggatttgg tcaaggagct tgacagagac acagtttttg ctctggtgaa ttacatcttc 901 tttaaaggca aatgggagag accctttga gtcaaggaca ccgaggaga ggacttccac 961 gtggaccagg tgaccaccgt gaaggtgcct atgatgaagc gtttaggcat gtttaacatc 1021 cagcactgta agaagctgtc cagctgggtg ctgctgatga aatacctggg aatgccacc 1081 gccatcttct tcctgcctga tgaggggaa ctacagcacc tggaaaatga ctcacccac 1141 gatatcatca ccaagttcct ggaaatga gacagaaggt ctgccagctt catttaccc 1201 aaactgtcca ttactggaac ctatgatctg aagagcgtcc tgggtcaact ggcatcact 1261 aaggtcttca gcaatggggc tgacctctcc ggggtcacag aggaggcacc ctgaagctc 1321 tccaaggccg tgcataaggc tgtgctgacc atcgac g aga aagggactga gc tgctggg 1381 gccatgtttt tagaggccat acccatgtct atccccccg aggtcaagtt aacaaaccc 1441 tttgtcttct taatgattga acaaaatacc aagtctcccc tcttcatggg aaagtggtg 1501 aatcccaccc aaaaataact gcctctcgct cctcaacccc tcccctccat cctggcccc 1561 ctccctggat gacattaaag aagggttgag ctggtccctg cctgcatgtg ctgtaaatc 1621 cctcccatgt tttctctgag tctccctttg cctgctgagg ctgtatgtgg ctccaggta 1681 acagtgctgt cttcgggccc cctgaactgt gttcatggag catctggctg gtaggcaca 1741 tgctgggctt gaatccaggg gggactgaat cctcagctta cggacctggg ccatctgtt 1801 tctggagggc tccagtcttc cttgtcctgt cttggagtcc ccaagaagga tcacagggg 1861 aggaaccaga taccagccat gaccccaggc tccaccaagc atcttcatgt cccctgctc 1921 atcccccact cccccccacc cagagttgct catcctgcca gggctggctg gcccacccc 1981 aaggctgccc tcctgggggc cccagaactg cctgatcgtg ccgtggccca ttttgtggc 2041 atctgcagca acacaagaga gaggacaatg tcctcctctt gacccgctgt acctaacca 2101 gactcgggcc ctgcacctct caggcacttc tggaaaatga ctgaggcaga tcttcctga 2161 agcccattct ccatggggca acaaggacac ctattctgtc cttgtccttc atcgctgcc 2221 ccagaaagcc tcacatatct ccgtttagaa tcaggtccct tctccccaga gaagaggag 2281 ggtctctgct ttgttttctc tatctcctcc tcagacttga ccaggcccag aggccccag 2341 aagaccatta ccctatatcc cttctcctcc ctagtcacat ggccataggc tgctgatgg 2401 ctcaggaagg ccattgcaag gactcctcag ctatgggaga ggaagcacat acccattga 2461 cccccgcaac ccctcccttt cctcctctga gtcccgactg gggccacatg agcctgact 2521 tctttgtgcc tgttgctgtc cctgcagtct tcagagggcc accgcagctc agtgccacg 2581 gcaggaggct gttcctgaat agcccctgtg gtaagggcca ggagagtcct ccatcctcc 2641 aaggccctgc taaaggacac agcagccagg aagtcccctg ggcccctagc gaaggacag 2701 cctgctccct ccgtctctac caggaatggc cttgtcctat ggaaggcact ccccatccc 2761 aaactaatct aggaatcact gtctaaccac tcactgtcat gaatgtgtac taaaggatg 2821 aggttgagtc ataccaaata gtgatttcga tagttcaaaa tggtgaaatt gcaattcta 2881 catgattcag tctaatcaat ggataccgac tgtttcccac acaagtctcc gttctctta 2941 agcttactca ctgacagcct ttcactctcc acaaatacat taaagatatg ccatcacca 3001 agcccccctag gatgacacca gacctgagag tctgaagacc tggatccaag tctgacttt 3061 tccccctgac agctgtgtga ccttcgtgaa gtcgccaaac ctctctgagc ccagtcatt 3121 gctagtaaga cctgcctttg agttggtatg atgttcaagt tagataacaa atgtttata 3181 cccattagaa cagagaataa atagaactac atttcttgca The PAM sequence is highlighted, showing the correct sequence after adenine base editing.

[0188] "Target" means mammals, including humans and non-human mammals, such as the Bovidae subfamily and horses. This includes, but is not limited to, the families Canidae, Sheepidae, or Felidae. Livestock, animals raised to produce labor and provide goods, such as food. This includes, but is not limited to, domesticated animals such as cows, goats, and chickens. This includes, but is not limited to, horses, pigs, rabbits, and sheep.

[0189] "Substantially identical" means that the reference amino acid sequence (for example, the amino acids listed herein) is substantially identical. Any of the sequences) or nucleic acid sequences (for example, any of the nucleic acid sequences described herein) This means a polypeptide or nucleic acid molecule that exhibits at least 50% identity with respect to [the specified substance]. In terms of morphology, such sequences are used for comparison with sequences at the amino acid level or nucleic acid level. Have at least 60%, 80%, 85%, 90%, 95% or even 99% identity at the level. Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705's Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE / PRETTYBOX program). Such software assigns a degree of homology by accounting for various substitutions , deletions, and / or other modifications to match identical or similar sequences. Conservative substitutions typically include substitutions within the following groups: Gly cine, Alanine; Valine, Isoleucine, Leucine; Aspartic Acid, Glutamic Acid, As paragine, Glutamine; Serine, Threonine; Lysine, Arginine; Phenylalanine , Tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program, including the probability score between closely related sequences represented by e -3 and e -100 , can be used.

[0190] COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalty -11, -1 and end gap penalty -5, -1 b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Keep recalculating to find conserved columns c) Query clustering parameters: Use query clustering on; W Code size 4; Maximum cluster distance 0.8; Alphabet is normal. EMBOSS Needle is used with the following parameters, for example: a) Matrix: BLOSUM62; b) Gap open: 10; c) Gap extension: 0.5; d) Output format: Pair; e) Terminal gap penalty: False; f) Terminal gap open: 10; and g) Terminal gap elongation: 0.5.

[0191] The term "target site" refers to a sequence within a nucleic acid molecule that has been modified by a nucleic acid base editor. This refers to a sequence that is subjected to a deaminase. In one embodiment, the target site is a deaminase or a deaminase (e.g. It is deaminated by a fusion protein containing (for example, adenine deaminase).

[0192] As used herein, the terms “treat” and “treating” "Treatment" refers to reducing or improving a disorder and / or related symptoms. This refers to doing or obtaining the desired pharmacological and / or physiological effect. Treating a condition requires the complete elimination of the associated disorder, condition, or symptoms. It will be understood that this is not the case. In some embodiments, the effect is therapeutic, In other words, the effects are not limited to, but include, the disease and / or the harmful effects caused by the disease. To partially or completely reduce, decrease, eliminate, alleviate, soothe, lessen, or cure symptoms. In some embodiments, the effect is preventive, that is, the effect is the onset of disease or condition. To protect or prevent the disease or recurrence. For this purpose, the methods disclosed herein are described herein. This includes administering a therapeutically effective amount of the composition as described. In one embodiment, the disease This is alpha-1 antitrypsin deficiency (A1AD).

[0193] "Uracil glycosylase inhibitor," or "UGI," refers to the uracil removal and repair system. This refers to an agent that inhibits host uracil-DNA glycosylase. In one embodiment, the agent is a host uracil-DNA glycosylase. It is a protein or fragment that binds to DNA and prevents the removal of uracil residues from DNA. In this embodiment, UGI inhibits uracil-DNA glycosylase base excision repair enzyme. It is a protein, a fragment thereof, or a domain. In some embodiments, UGI The main component includes wild-type UGI or modified versions thereof. In some embodiments, the UGI is The main component includes the exemplary amino acid sequence fragments presented below. In some embodiments, The UGI fragment is at least 60%, at least 65%, and less than the exemplary UGI sequence provided below. 70% each, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 Including 5%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% It includes an amino acid sequence. In some embodiments, the UGI is as illustrated below. The UGI amino acid sequence contains an amino acid sequence homologous to the amino acid sequence or a fragment thereof. Several implementations Morphologically, UGIs or parts thereof may be wild-type UGIs or UGI sequences as described below. or at least 70%, at least 75%, at least 80%, at least 85% of a portion thereof, At least 90%, at least 95%, at least 96%, at least 97%, at least 98%, less It has at least 99%, at least 99.5%, at least 99.9%, or 100% identity. Exemplary U GI contains the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGEN KIKML.

[0194] The term "vector" refers to a means of introducing nucleic acid sequences into cells to produce transformed cells. It refers to. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes are included. The "expression vector" is expressed in recipient cells. It is a nucleic acid sequence containing the nucleotide sequence to be expressed. The expression vector contains the introduced sequence Additional nucleic acid sequences, e.g., start, stop, enhancer, to promote and / or facilitate the process. —, promoter, and secretory sequence may be included.

[0195] Any composition or method provided herein may be used with any other composition or method provided herein. It may be combined with one or more compositions and methods.

[0196] In this specification, "several embodiments," "a certain embodiment," "one embodiment," and The reference to "other embodiments" refers to specific features, structures, etc., described in relation to those embodiments. , or characteristics, are included in at least some embodiments of this disclosure, but not necessarily all. This means that it is not necessarily included in the embodiments.

[0197] In any definition of a variable substance in this specification, the listing of chemical groups is based on the enumerated groups. This includes the definition of its variable as any single base or combination thereof. The description of the embodiments relating to the variable object or aspect may be presented as any single embodiment, or as an arbitrary single embodiment. This includes embodiments of the same, combined with other embodiments or parts thereof.

[0198] DNA editing modifies disease conditions by correcting pathogenic mutations at the genetic level. It emerged as a viable means to do so. Until recently, all DNA editing platforms It induces DNA double-strand breaks (DSBs) at specific genomic sites, relying on the endogenous DNA repair pathway, and semi- The probabilistic method of determining the product result results in a complex population of gene products, thereby contributing to their function. Through the Homology Directional Restoration (HDR) pathway, accurate, user-defined restoration results are achieved. While the result is achievable, highly efficient repair using HDR is not possible in therapeutically appropriate cell types. However, this has been hindered by several difficulties. In fact, this path is conflicting and error-prone. It is less efficient compared to the more common non-homologous end linkage pathway. Furthermore, HDR is less efficient than the G1 of the cell cycle. This is strictly limited to the S phase and hinders the precise repair of DSBs in postmittal cells. In these populations, the genome is efficiently programmed in a user-defined manner. It has been shown that modifying sequences is difficult or impossible. [Brief explanation of the drawing]

[0199] [Figure 1]Figures 1A–1C show plasmids. Figure 1A is an expression vector encoding the TadA7.10-dCas9 base editor. Figure 1B is a plasmid containing nucleic acid molecules encoding proteins that confer chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene rendered inoperable by two point mutations. Figure 1C is a plasmid containing nucleic acid molecules encoding proteins that confer chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene rendered inoperable by three point mutations. [Figure 2] Figure 2 shows images of bacterial colonies transduced with the expression vectors shown in Figures 1A-C, including those lacking the kanamycin resistance gene. The vectors contained ABE7.10 variants generated using error-prone PCR. Using increasing concentrations of kanamycin, bacterial cells expressing these "evolved" ABE7.10 variants were selected for kanamycin resistance. Bacteria expressing ABE7.10 variants with adenosine deaminase activity were able to correct the mutation introduced into the kanamycin resistance gene, thereby restoring kanamycin resistance. Kanamycin-resistant cells were selected for further analysis. [Figure 3] Figure 3 is a graph quantifying the efficacy and specificity of the selected ABE8s listed in Table 6. Editing was assayed at the alpha-1 antitrypsin locus in HEK293T cells. [Figure 4] Figures 4A and 4B illustrate the editing efficiency and specificity of ABE8. Figures 4A and 4B are graphs quantifying the base editing and specificity of selected ABE8s, as listed in Table 6. Single variant TadA deaminase domains or wild-type TadA deaminases were evaluated. [Figure 5]Figure 5 presents a graph demonstrating the effectiveness of ABE8 in editing on-target adenine (A) base pairs bystander A. In particular, ABE8 results in a 5-fold increase in editing at the A1AD site (i.e., A·T to G·C conversion) compared to the effective TadA deaminase, ABE7.10. [Figure 6] Figures 6A–6D show nucleic acid sequences, tables, and bar graphs relating to the generation of improved rates of nucleic acid base modification in primary PiZ fibroblasts through base editor operations. Figure 6A shows the target site DNA sequence encoding the A1AD-associated PiZ mutation. This sequence includes a 20-nucleotide protospacer and a non-canonical spCas9 NGC PAM. Figure 6B presents a table listing both the TadA deaminase and Cas9 PAM variant components of various editors used to modify the PiZ mutation. Figures 6C and 6D present bar graphs showing the editing rates observed in patient-derived PiZZ fibroblasts (GM11423 Corriel Biorepository) transfected with base editing reagents using a Neon electroporation system. Each treatment consisted of 10 μl of electroporation buffer containing 70,000 fibroblasts, 100 ng mRNA encoding the base editor, and 50 ng alpha-1 modification gRNA. After 48 hours of recovery, cells were lysed, and the locus of interest was identified by sequencing of the target amplicon. Data were obtained from two independent experiments. These data and results demonstrate improvements in editing efficiency for both NGC PAM recognition optimization (Variant 1-3, Figures 6B and 6C) and TadA deaminase optimization through the incorporation of ABE8 / 9 mutations (Variant 4-9, Figures 6B-6D). [Figure 7]Figures 7A–7D present nucleic acid sequences, tables, and graphs related to the increase in serum A1AT produced by lipid nanoparticle (LNP) mediated delivery and base editing in NSG-PiZ transgenic mice. Figure 7A shows the target site DNA sequence, including a 20-nucleotide protospacer and a non-canonical spCas9 NGC PAM. Figure 7B presents a table listing both the TadA deaminase and Cas9 PAM variant components of the various editors used to correct PiZ mutations. Figure 7C presents a graph showing the editing rates observed in whole liver gDNA from NSG-PiZ transgenic mouse models 7 days after treatment with 1.5 mg / kg of LNP containing a 1:1 weight ratio of gRNA and mRNA encoding the base editor. Commercially available NSG-PiZ mice express mutant human SERPINA1 (Glu342Lys mutation) on an immunodeficient NOD-SCID gamma (NSG) background, which provides a stable background for human hepatocytes after partial hepatectomy (The Jackson Laboratory, Mount Desert Island, ME). The results demonstrated that ngcABEvar9 produced a higher editing rate than the earlier variant 8. Figure 7D presents a graph showing that the editing rate correlated with an increase in serum alpha-1 antitrypsin compared to the untreated sample, as measured by the MSD sandwich immunoassay. Based on these results, base editing using the ABE8 reagent can address alpha-1 antitrypsin deficiency and its potential pulmonary complications. [Figure 8]Figure 8 is a table showing Cas9 variants for accessing all possible PAMs within the NRNN PAM space. Only Cas9 variants requiring recognition of three or fewer defined nucleotides in the PAM are listed. Non-G PAM variants include SpCas9-NRRH, SpCas9-NRTH, and SpCas9-NRCH (Miller, SM, et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), ( / / doi.org / 10.1038 / s41587-020-0412-8), the contents of which are incorporated herein by reference in their entirety. [Modes for carrying out the invention]

[0200] As described below, the present invention relates to alpha-1 antitrypsin deficiency (A1AD). The present invention features compositions and methods for modifying mutations. In some embodiments, Editing corrects harmful mutations, and the edited polynucleotides are compared to the wild-type reference polynucleotides. The editing is made indistinguishable from the rheotide sequence. In another embodiment, the editing modifies harmful mutations. The goal is to ensure that the modified and edited polynucleotides contain benign mutations.

[0201] The present invention relates at least in part to a base eddy containing an adenosine deaminase variant. Based on the discovery that the ter can effectively and accurately edit harmful mutations associated with A1AD Duku.

[0202] [Alpha-1 antitrypsin deficiency (A1AD)] Alpha-1 antitrypsin (A1A) is encoded by the SERPINA1 gene on chromosome 14. It is a protease inhibitor. This glycoprotein is mainly synthesized in the liver and distributed into the bloodstream. It is secreted, and the serum concentration in healthy adults is 1.5-3.0 g / L (20-52 μmol / L). The protein diffuses into the pulmonary interstitium and alveolar lining fluid, where it inactivates neutrophil elastase. This protects lung tissue from protease-mediated damage. Alpha-1 Anti Trypsin deficiency (A1AD) is inherited in an autosomal codominant manner. The SERPINA1 gene has over 100 genes. While various gene variants have been described, not all of them are associated with disease. The alphabetical naming of these variants is based on their electrophoretic velocity in gel electrophoresis. The most common variant is the M (moderate mobility) allele (PiM), and the two most frequent failures The alleles are PiS and PiZ (the latter having the slowest electrophoretic velocity). Measurable serum proteins Several mutations that do not produce this compound have been described; these are called "null" alleles. The most common genotype is MM, and normal serum levels of alpha-1 antitrypsin are present. This results in a bell syndrome. Most humans with severe deficiency are homozygous for the Z allele (ZZ). More than 60,000 A1AD patients in the United States have a severe ZZ phenotype. The Z protein is found in hepatocytes. During production in the endoplasmic reticulum, misfolding occurs and polymerization takes place; these abnormal polymers are found in the liver. It is captured, and serum levels of alpha-1 antitrypsin decrease significantly. Insufficient or Due to unstable A1AT production, patients with A1AD develop liver and / or lung lesions. It is said that liver disease seen in patients with alpha-1 antitrypsin deficiency is caused by abnormalities in hepatocytes. The alpha-1 antitrypsin protein accumulates, resulting in autophagy and endoplasmic reticulum suppression. It is caused by cellular reactions, including the resuscitation reaction and apoptosis. A decrease in the circulating level of pha-1 antitrypsin leads to a decrease in neutrophil elastase activity in the lungs. This increases; as a result of this imbalance between proteases and antiproteases, this condition is associated with This can lead to lung disease.

[0203] Alpha-1 antitrypsin deficiency ("A1AD") is most common in Caucasians. It frequently affects the lungs and liver. In the lungs, the most common signs are most pronounced at the base of the lungs. It is a markedly early-onset (patients in their 30s and 40s) panlobar emphysema. However, it is diffuse or upper Lobar emphysema may occur, as may bronchiectasis. These are the most frequently reported symptoms. These include shortness of breath, wheezing, and coughing. Lung function tests in affected individuals are consistent with those for COPD. It shows signs of a bronchodilator reaction, and there is a possibility of being diagnosed with asthma. Liver diseases caused by the ZZ genotype manifest in various forms. In affected infants, new Symptoms may appear in the newborn period, including cholestatic jaundice, and sometimes acholestial feces (pale color or (Clay-colored) and accompanied by hepatomegaly. Conjugated bilirubin, transaminase, and cancer in the blood. Maglutamyltransferase levels are elevated in older children and adults. In liver disease, elevated transaminase levels are discovered incidentally, or varicose bleeding may occur. The symptoms may be discovered through established signs of cirrhosis, including ascites. -1 antitrypsin deficiency also increases the likelihood of developing hepatocellular carcinoma in patients. Homozygous ZZ inheritance While the subtype is necessary for the development of liver disease, heterozygous Z mutations are associated with hepatitis C infection. This increases the risk of more severe liver disease, as seen in cystic fibrosis liver disease. Therefore, it may act as a gene modifying factor in other diseases.

[0204] The two most common clinical variants of A1AD are the E264V (PiS) and E342K (PiZ) alleles. The clinical single nucleotide variant E342K(PiZ) is unstable and / or inactive A1AT tan. It produces proteins, which in turn cause liver and lung toxicity. The inheritance is autosomal codominant. More than half of A1AD patients harbor at least one copy of the E342K mutation.

[0205] JPEG0007861090000016.jpg66163

[0206] In some embodiments, the disease or disorder is alpha-1 antitrypsin deficiency (A1AD). In some embodiments, the pathogenic mutation is in the gene SERPINA1. In some embodiments, the SERPINA1 mutation is E342K (PiZ allele). So, A, which is in 7th place, is edited to G, restoring the PiZ allele to the wild-type allele.

[0207] [Nucleic Acid Editor] Base editing for editing, modifying, or altering the target nucleotide sequence of polynucleotides. A nucleotide editor or nucleic acid base editor is disclosed herein. What is described herein is: Polynucleotide programmable nucleotide-binding domain and nucleic acid base editing domain Nucleic acid base editors or base editors that include (for example, adenosine deaminase) The polynucleotide programmable nucleotide-binding domain binds to the guide polynucleotide. When together with nucleotides (e.g., gRNA), the target polynucleotide sequence (that is, Complementary base pair between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence It can specifically bind (via a mechanism), thereby allowing editing of the desired target nucleic acid sequence. The base editor can be localized to the target polynucleus. In some embodiments, the target polynucleus The rheotide sequence includes single-stranded or double-stranded DNA. In some embodiments, the target polygon The creotide sequence contains RNA. In some embodiments, the target polynucleotide sequence is DNA -Includes RNA hybrids.

[0208] [Polynucleotide programmable nucleotide-binding domain] The polynucleotide programmable nucleotide-binding domain also binds to the nucleus of RNA. It should be understood that it can contain acid-programmable proteins. For example, polynuclear The rheotide programmable nucleotide-binding domain is polynucleotide programmable. The nucleotide-binding domain can associate with nucleic acids that guide RNA. Other nucleic acid programs are also possible. DNA-binding proteins are also within the scope of this disclosure, but they are not specifically listed in this disclosure. It hasn't been done.

[0209] The polynucleotide programmable nucleotide-binding domain of the base editor is It can contain one or more domains. For example, polynucleotide programmed The nucleotide-binding domain may contain one or more nuclease domains. In some embodiments, the polynucleotide programmable nucleotide-binding domain Nuclease domains may contain endonucleases or exonucleases. In this specification, the term "exonuclease" means to process nucleic acids (e.g., RNA or DNA). The term "end" refers to a protein or polypeptide that can be digested from its free ends. Nucleases are enzymes that catalyze (e.g., cleave) the internal regions of nucleic acids (e.g., DNA or RNA). This refers to a protein or polypeptide that can be produced. In some embodiments, it refers to an endonucleus. -ase can cleave a single strand of double-stranded nucleic acid. In some embodiments, end Nucleases can cleave both strands of a double-stranded nucleic acid molecule. Several implementations In this state, the polynucleotide programmable nucleotide-binding domain is deoxyribonu It may be a crease. In some embodiments, a polynucleotide programmable nucleus The rheotide-binding domain may be a ribonuclease.

[0210] In some embodiments, polynucleotide programmable nucleotide binding domain The nuclease domain cleaves zero, one, or two strands of the target polynucleotide. This is possible. In some embodiments, polynucleotide programmable nucleotides The binding domain may include a niccasse domain. In this specification, the term "niccasse" is used. "Cackase" is a enzyme that can cleave only one strand of a double-stranded nucleic acid molecule (such as DNA). Polynucleotide programmable nucleotide bond containing a nuclease domain This refers to the main component. In some embodiments, nickase is used to activate polynucleotide programming. By introducing one or more mutations into the nucleotide-binding domain, polynuclear Fully catalytically active (e.g., naturally occurring) creotide programmable nucleotide binding domain It can be obtained from the morphology. For example, polynucleotide programmable nucleotide bond domain If the n contains a nickarse domain derived from Cas9, the nickarse domain derived from Cas9 This can include the D10A mutation and histidine at position 840. In such embodiments, Therefore, residue H840 retains catalytic activity, thereby enabling the cleavage of a single strand of a double-stranded nucleic acid. Yes, it is possible. In another example, the Cas9-derived nickase domain may contain the H840A mutation. In some embodiments, Nikka -ase removes all or part of the nuclease domain that is not necessary for nickas activity. By doing so, the polynucleotide programmable nucleotide binding domain can be fully It can be obtained from catalytically active (e.g., natural) forms. For example, polynucleotide programmeable If the nucleotide-binding domain contains a nickase domain derived from Cas9, The resulting nickase domain is a deficiency of all or part of the RuvC domain or HNH domain. It is possible to lose.

[0211] The amino acid sequence of an example catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0212] Therefore, programmable nucleotide bonds containing a nickase domain in polynucleotides A base editor containing a domain targets a specific polynucleotide target sequence (e.g., a bound nucleotide). The generation of single-strand DNA breaks (nicks) (determined by the complementary sequence of the id nucleic acid). This can be achieved. In some embodiments, the nickase domain (e.g., a nickase derived from Cas9) is used. Nucleic acid double-stranded target polynucleotides cleaved by a base editor containing the domain The chain of columns is the chain that is not edited by the base editor (i.e., the base editor Therefore, the strand that is cut is the strand opposite to the strand containing the base being edited. A base editor containing a nickase domain (e.g., a nickase domain derived from Cas9) is This allows for the cleavage of the DNA molecule strand targeted for editing. Therefore, the non-target strand is not cleaved.

[0213] Catalytically inactive (i.e., unable to cleave the target polynucleotide sequence) A base editor containing a polynucleotide programmable nucleotide-binding domain is also available. Provided in the details. In this specification, the terms “catalytically inert” and “nucleate” are used. "Ase inactivity" is a condition in which one or more mutations result in the inability to cleave nucleic acid chains. Polynucleotide programmable nucleotide-binding domains having and / or deletions Used interchangeably for pointing. In some embodiments, catalytically inert poly The nucleotide programmable nucleotide-binding domain base editor allows you to edit one or more nucleotides. Lack of nuclease activity as a result of specific point mutations in the clease domain. This is possible. For example, in the case of a base editor that includes the Cas9 domain, Cas9 can handle the D10A mutation. This can include both the H840A mutation and the H840A mutation. Such mutations can include both nuclei. The zed domain is inactivated, resulting in a loss of nuclease activity. In other embodiments, catalytically The inactive polynucleotide programmable nucleotide-binding domain is a catalytic domain. This includes one or more deletions of all or part of the nucleotide (e.g., RuvC1 and / or HNH domain). This can be done. In a further embodiment, a catalytically inert polynucleotide program Possible nucleotide-binding domains include point mutations (e.g., D10A or H840A) and nucleonucleotides. Includes deletion of all or part of the ase domain.

[0214] Furthermore, in this specification, the polynucleotide programmable nucleotide-binding domain From the previously working version, the catalytically inactive polynucleotide programmable Mutations that can generate nucleotide-binding domains are also being considered. For example, In the case of a mediatedly inactive Cas9 ("dCas9"), mutations other than D10A and H840A are present. A variant is provided that results in nuclease inactivation of Cas9. Such mutations For example, other amino acid substitutions in D10 and H840, or the nuclease domain of Cas9. Other substitutions within (e.g., in the HNH nuclease subdomain and / or RuvC1 subdomain) This includes substitutions in the above. Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and the above. This can be made clear to those skilled in the art based on knowledge in the technical field, and the scope of this disclosure It is inside. Such additional exemplary suitable nuclease-inactive Cas9 domains are limited Although not intended, D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A It contains a mutant domain (e.g., Prashant et al., CAS9 transcriptional activ ators for target specificity screening and paired nickases for cooperative genomes See e engineering. Nature Biotechnology. 2013; 31(9): 833-838 (in its entirety The contents are incorporated herein by reference.

[0215] Polynucleotide programmable nucleotides that can be incorporated into a base editor Non-restrictive examples of binding domains include CRISPR protein-derived domains, restriction nucleases, Meganuclease, TAL nuclease (TALEN), and zinc finger nuclease (Z FN) is included. In some embodiments, the base editor is CRISPR (i.e.) of nucleic acids. During the intermediation and modification of Clustered Regularly Interspaced Short Palindromic Repeats, the combination Natural or modified proteins that can bind to nucleic acid sequences via guide nucleic acids Contains a polynucleotide programmable nucleotide-binding domain, which includes the quality or a part thereof. Yes. Such proteins are referred to as "CRISPR proteins" in this specification. Disclosed herein are polynuclear molecules containing all or part of the CRISPR protein. A base editor containing a creotide programmable nucleotide binding domain (i.e., A base editor containing all or part of a CRISPR protein as a domain. It is also called the "CRISPR protein-derived domain" of the protein. The incorporated CRISPR protein-derived domain is compared to wild-type or native CRISPR protein. It can be compared and modified. For example, as described below, the CRISPR protein-derived domain is Compared to the wild-type or natural-type CRISPR protein, one or more mutations, insertions, or deletions This may include rearrangement and / or reconfiguration.

[0216] CRISPR provides protection against mobile genetic elements (viruses, transposable elements, and zygosity plasmids). It is an adaptive immune system. CRISPR clusters interact with spacers, preceding mobile elements. It contains a complementary sequence and a target entry nucleic acid. The CRISPR cluster is transcribed and CRISPR RNA (cr It is processed into RNA. In the type II CRISPR system, the correct processing of pre-crRNA The proteins involved are transcoding small RNA molecules (tracrRNA), endogenous ribonuclease 3 (rnc), and Ca It requires the s9 protein. TracrRNA is a pre-crRNA process assisted by ribonuclease 3. It acts as a guide for the thi. Subsequently, Cas9 / crRNA / tracrRNA connects to the spacer in a linear fashion. Alternatively, a circular dsDNA target is cleaved with an endonuclease. Target strands that are not complementary to crRNA are... First, it is endonuclease-likely cleaved, and then exonuclease-likely trimmed to 3'-5'. In nature, both proteins and RNA are required for DNA binding and cleavage. However, To incorporate aspects of both crRNA and tracrRNA into a single RNA species, a single guide RNA It is possible to produce (sgRNA, or simply gRNA). For example, Jinek M., Chylins ki K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(20 See 12) (the entire contents of which are incorporated herein by reference). Cas9 is CRI Recognizing short motifs (PAM or protospacer adjacency motifs) within SPR repeat sequences, It helps to distinguish between "self" and "non-self."

[0217] In some embodiments, the methods described herein involve manipulating the Cas protein It can be used. Guide RNA (gRNA) is a scaffold sequence necessary for Cas binding and is modified It is a short synthetic RNA consisting of a user-defined spacer of approximately 20 base pairs that defines the genome target. Therefore, those skilled in the art can alter the genomic target of Cas protein specificity, which is How specific gRNA target-directed sequences are to genomic targets compared to the rest of the genome. It is partially determined by whether or not.

[0218] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.

[0219] In some embodiments, CRISPR protein-derived base editors are incorporated into the base editor. The 'in' binds to the target polynucleotide when combined with a binding guide nucleic acid. Endonucleases that can produce this (e.g., deoxyribonuclease or ribonuclease) In some embodiments, the CRISPR protein incorporated into the base editor is used. The target domain, when combined with the bound guide nucleic acid, binds to the target polynucleotide. It is a nickas that can be combined. In some embodiments, it is combined within a base editor. The embedded CRISPR protein-derived domain, when combined with the bound guide nucleic acid, It is a catalytically inactive domain that can bind to the target polynucleotide. In some embodiments, the target binds to the CRISPR protein-derived domain of the base editor. Polynucleotides are DNA. In some embodiments, a base editor CRISPR is used. The target polynucleotide that binds to the protein-derived domain is RNA.

[0220] The Cas proteins that can be used herein include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, and C. as5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10 , Csy1 , Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1 , Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO , Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and This includes Cas12i, CARF, DinG, their homologs, or their modifications. CRISPR enzymes, like Cas9, have two functional endonuclease regions, RuvC and HNH. It can have DNA cleavage activity. CRISPR enzymes cleave DNA within the target sequence and / or within the target sequence. It is possible to induce cleavage of one or both strands at a target sequence, such as within the complementary strand of a column. For example, the CRISPR enzyme extracts approximately 1, 2, 3, 4, 5, 6 nucleotides from the first or last nucleotide of the target sequence. , 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs or more apart, on the other hand This can induce the break of both chains.

[0221] Lacking the ability to cleave one or both strands of a target polynucleotide containing the target sequence. For example, a vector encoding a CRISPR enzyme that has been mutated relative to the corresponding wild-type enzyme. It can be used. Cas9 is derived from the wild-type exemplary Cas9 polypeptide (e.g., S. pyogenes). (Cas9) and at least, or at least approximately, 50%, 60%, 70%, 80%, 90%, 91%, 92% 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology It can refer to a polypeptide that has sex. Cas9 is an example of the wild-type Cas9 polypeptide. For example, those derived from S. pyogenes, up to or approximately up to 50%, 60%, 70% Sequence identity of %, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% It can refer to polypeptides having sequence homology with and / or Cas9, wild type, or deletion, insertion, substitution, variant, mutation, fusion, chimera, or any of these This can refer to modified versions of the Cas9 protein that may include amino acid changes such as combinations of amino acids.

[0222] In some embodiments, the CRISPR protein-derived domain of the base editor contains Cory nebacterium ulcerans (NCBI reference: NC_015683.1, NC_017317.1); Corynebacterium dipht heria (NCBI reference: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI reference: NC_021284.1); Prevotella intermedia (NCBI reference: NC_017861.1); Spiroplasma taiwane nse (NCBI reference: NC_021846.1); Streptococcus iniae (NCBI reference: NC_021314.1); Bellie lla baltica (NCBI reference: NC_018010.1); Psychroflexus torquis (NCBI reference: NC_018721. 1); Streptococcus thermophilus (NCBI reference: YP_820832.1); Listeria innocua (NCBI reference: YP_820832.1); Reference: NP_472073.1); Campylobacter jejuni (NCBI reference: YP_002344900.1); Neisseria men ingitidis (NCBI reference: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus It may contain all or part of Cas9 derived from Cus aureus.

[0223] [Cas9 domain of nucleic acid base editor] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (e.g., "Complete"). genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J. J., McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S ., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, R en Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin R. E., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K. , Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpen tier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA e ndonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I See Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012). (Each of these is incorporated herein by reference in its entirety.) Cas9 orthologs are limited Although not described, it is described in a variety of species, including S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences are evident to those skilled in the art based on this disclosure. Therefore, such Cas9 nucleases and sequences were identified by Chylinski, Rhun, and Charpenti. er, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2 013) Cas9 sequences from organisms and loci disclosed in RNA Biology 10:5, 726-737 Included. Its entire contents are incorporated herein by reference.

[0224] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is Cas There are 9 domains. Non-exclusive, exemplary Cas9 domains are provided herein. The nuclease consists of a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain (dCas9), and It may be Cas9 nickase (nCas9). In some embodiments, the Cas9 domain is This is the nuclease activity domain. For example, the Cas9 domain is involved in both strands of a double-stranded nucleic acid (e.g. For example, it could be a Cas9 domain that cleaves both strands of a double-stranded DNA molecule. Several implementations In this state, the Cas9 domain contains one of the amino acid sequences described herein. In one embodiment, the Cas9 domain corresponds to any one of the amino acid sequences described herein. and at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, At least 85%, at least 90%, at least 95%, at least 96%, at least 97%, and less It contains amino acid sequences that are identical by at least 98%, at least 99%, or at least 99.5%. In some embodiments, the Cas9 domain is one of the amino acid sequences described herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 , 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations It contains an amino acid sequence. In some embodiments, the Cas9 domain is the amino acid sequence described herein. Compared to any one of the no-acid sequences, at least 10, at least 15, at least 20, less At least 30, at least 40, at least 50, at least 60, at least 70, at least 80, At least 90, at least 100, at least 150, at least 200, at least 250, less 300 each, at least 350, at least 400, at least 500, at least 600, at least 7 00, at least 800, at least 900, at least 1000, at least 1100, or at least It also contains an amino acid sequence having 1200 identical consecutive amino acid residues.

[0225] In some embodiments, a protein containing a fragment of Cas9 is provided. For example, several In that embodiment, the protein includes one of the following two Cas9 domains: (1) Ca (2) gRNA binding domain of Cas9; (3) DNA cleavage domain of Cas9. In some embodiments, Cas9 also Proteins containing this fragment are called "Cas9 variants." Cas9 variants are Ca It shares homology with s9 or a fragment thereof. For example, Cas9 variants share homology with wild-type Cas9. Both are approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, and less At least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical They are at least approximately 99.5% identical, or at least approximately 99.9% identical. In some embodiments, The Cas9 variants are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 1 compared to wild-type Cas9. 2, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 3 2, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, also It may have more amino acid changes. In some embodiments, the Cas9 variant is , containing a Cas9 fragment (e.g., a gRNA-binding domain or a DNA-cleaving domain), and that fragment is in the field It is at least about 70% identical, at least about 80% identical, and at least about 90% identical to the corresponding fragment of bio-Cas9. % identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at Approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or at least approximately 99.9% identical One. In some embodiments, the fragment has a shorter amino acid length than the corresponding wild-type Cas9. 30% each, at least 35%, at least 40%, at least 45%, at least 50%, at least 5 5%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, At least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, It is at least 98%, at least 99%, or at least 99.5%. In some embodiments The fragment has a length of at least 100 amino acids. In some embodiments, the fragment is At least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 ami This is the length of the acid.

[0226] In some embodiments, the Cas9 fusion protein provided herein is a Cas9 protein This includes high-quality full-length amino acid sequences, such as one of the Cas9 sequences provided herein. However, In other embodiments, the fusion protein provided herein does not contain a full-length Cas9 sequence, Includes only one or more fragments. Exemplary amino acid sequences of appropriate Cas9 domains and Cas9 fragments. The following are provided herein, and additional suitable sequences of Cas9 domains and fragments are obvious to those skilled in the art. It is likely.

[0227] The Cas9 protein guides itself to a specific DNA sequence complementary to its guide RNA. It can then associate with guide RNA. In some embodiments, polynucleotides The programmable nucleotide-binding domain is a Cas9 domain, for example, a nuclease-active Ca Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, and CasY. This includes, but is not limited to, Cpf1, Cas12b / C2Cl, and Cas12c / C2C3.

[0228] In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes. (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows). ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG0007861090000017.jpg169163 (Single underline: HNH domain; Double underline: RuvC domain)

[0229] In some embodiments, wild-type Cas9 comprises the following nucleotides and / or amino acids Corresponding to or containing arrays: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGTCCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAAACATGAACGGCACCCCATCTTTGGAAACATAGGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAGTTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGA AAATATTATCCATTTGTTTACTCTTCCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG0007861090000018.jpg166162 (Single underline: HNH domain; Double underline: RuvC domain)

[0230] In some embodiments, wild-type Cas9 is derived from Cas9 (NCBI) derived from Streptococcus pyogenes. Reference sequence: NC_002737.2 (nucleotide sequence is as follows) and Uniprot reference sequence: Q99ZW2 (The amino acid sequence corresponds to the following): ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAGGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG0007861090000019.jpg167161 (Single underline: HNH domain; Double underline: RuvC domain)

[0231] In some embodiments, Cas9 is Corynebacterium ulcerans (see NCBI: NC_015683). 1, NC_017317.1); Corynebacterium diphtheria (NCBI reference: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI reference: NC_021284.1); Prevotella intermedia (NCBI Reference: NC_017861.1); Spiroplasma taiwanense (NCBI reference: NC_021846.1); Streptococc us iniae (NCBI reference: NC_021314.1); Belliella baltica (NCBI reference: NC_018010.1); Psy chroflexus torquisI (NCBI reference: NC_018721.1); Streptococcus thermophilus (NCBI reference: NC_018721.1); Reference: YP_820832.1), Listeria innocua (NCBI reference: NP_472073.1), Campylobacter jejuni (See NCBI: YP_002344900.1) or Neisseria meningitidis (See NCBI: YP_002342) 100.1) Represents Cas9 of origin, or Cas9 of any other biological origin.

[0232] Additional Cas9 proteins (e.g., nuclease-inactive (dead) Cas9 (dCas9), Cas9 Nicakase (nCas9, or nuclease-active Cas9) is a variant and homolog of nCas9. Please understand that this is included within the scope of this disclosure. The example Cas9 protein is limited This does not include the following, but is provided below. In some embodiments, Cas9 The protein is nuclease-inactive Cas9 (dCas9). In some embodiments, Cas9 The protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein The substance is Cas9 with nuclease activity.

[0233] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain does not cleave either strand of a double-stranded nucleic acid molecule. It can bind to double-stranded nucleic acid molecules (e.g., via gRNA molecules). In some embodiments, The crease-inactive dCas9 domain is a D10X mutation in the amino acid sequence described herein and The H840X mutation, or the corresponding amino acid sequence in any of the amino acid sequences provided herein This includes mutations where X is any amino acid change. In some embodiments, the nuclea - The enzyme-inactive dCas9 domain is a D10A mutation and H of the amino acid sequence provided herein. 840A mutation, or the corresponding mutation in any of the amino acid sequences described herein. It contains differences. For example, the nuclease-inactive Cas9 domain is used in the cloning vector pPla Includes the amino acid sequence shown in tTET-gRNA 2 (deposit number BAV54124):

[0234] An example of a catalytically inactive Cas9 (dCas9) amino acid sequence is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-s See "Specific control of gene expression." Cell. 2013; 152(5):1173-83. The entire content of is incorporated herein by reference.

[0235] Additional suitable nuclease-inactive dCas9 domains are relevant to this disclosure and the art. Such additional examples would be obvious to those skilled in the art and are within the scope of this disclosure. Exemplary and appropriate nuclease-inactive Cas9 domains include, without limitation, D10A / H840A, D10A / D83 This includes the 9A / H840A and D10A / D839A / H840A / N863A mutant domains (e.g., Pra shant et al., CAS9 transcriptional activators for target specificity screening a nd paired nickases for cooperative genome engineering. Nature Biotechnology. 201 See 3; 31(9): 833–838 (the entire contents of that document are incorporated herein by reference). ).

[0236] In some embodiments, Cas9 nuclease is used to cleave inactive (e.g., inactivated) DNA. It has a cleavage domain, that is, Cas9 is the "nCas9" protein (meaning "nickase" Cas9) This is called nickase. The nuclease-inactivating Cas9 protein is replaceable with "dCas Cas9 is also called the nuclease (meaning "dead" Cas9) or catalytically inactive Cas9. Obtain. A method that produces a Cas9 protein (or a fragment thereof) with an inactive DNA cleavage domain. The law is publicly known (e.g., Jinek et al, Science. 337:816-821(2012); Qi et al, “Repurp osing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Exp See “ression” (2013) Cell. 28; 152(5): 1173-83 (For details, please refer to this specification). It is incorporated into the document). For example, the DNA cleavage domain of Cas9 is the HNH nuclease subdomain. It is known to contain two subdomains: the HNH subdomain and the RuvC1 subdomain. The 'in' subdomain cleaves the complementary strand of the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. These subdomains... Mutations within the domain can suppress the nuclease activity of Cas9. For example, the D10A mutation... H840A completely inactivates the nuclease activity of S. pyogenes Cas9 (Jinek et al.). l, Science. 337:816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, the dCas9 domain is one of the dCas9 domains provided herein. For each deviation, at least 60%, at least 65%, at least 70%, at least 75%, At least 80%, at least 85%, at least 90%, at least 95%, at least 96%, and at least Ami who have 97%, at least 98%, at least 99%, or at least 99.5% identity It contains an amino acid sequence. In some embodiments, the Cas9 domain is an amino acid sequence as shown herein. Compared to any one of the acid sequences, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more It contains an amino acid sequence having several mutations. In some embodiments, the Cas9 domain is Compared to any one of the amino acid sequences shown herein, at least 10, at least 1 5, at least 20, at least 30, at least 40, at least 50, at least 60, less 70, at least 80, at least 90, at least 100, at least 150, at least 200 , at least 250, at least 300, at least 350, at least 400, at least 500, small At least 600, at least 700, at least 800, at least 900, at least 1000, less An amino acid sequence having 1100 or at least 1200 identical consecutive amino acid residues include.

[0237] In some embodiments, dCas9 inactivates one or more nucleotides that inactivate Cas9 nuclease activity. This corresponds to, or includes, a part or all of, a mutated Cas9 amino acid sequence. However, in some embodiments, the dCas9 domain is a D10A and H840A mutation or another Ca Includes the corresponding mutation in s9.

[0238] In some embodiments, dCas9 includes the amino acid sequences of dCas9 (D10A and H840A): JPEG0007861090000020.jpg166161 (Single underline: HNH domain; Double underline: RuvC domain)

[0239] In some embodiments, the Cas9 domain contains the D10A mutation, while the one provided above... The residue at position 840 in the amino acid sequence, or any of the amino acid sequences provided herein. The residue at the corresponding position in the compound remains histidine.

[0240] In other embodiments, for example, D10A and H84 produce nuclease-inactivated Cas9 (dCas9). A dCas9 variant with mutations other than 0A is provided. Such mutations are, for example For example, other amino acid substitutions at D10 and H840, or other amino acid substitutions within the Cas9 nuclease domain. Substitution (for example, in the HNH nuclease subdomain and / or RuvC1 subdomain) Includes (exchange). In some embodiments, a variant or homolog of dCas9, with at least At least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, At least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or at least The provided product is approximately 99.9% identical. In some embodiments, approximately 5 amino acids, approximately 10 amino acids Mino acids, approximately 15 amino acids, approximately 20 amino acids, approximately 25 amino acids, approximately 30 amino acids, approximately 40 amino acids, approximately 50 amino acids, approximately 75 amino acids, approximately 100 amino acids or more: Short or long amino acid sequences A variant of dCas9 having the following characteristics is provided. In some embodiments, the Cas9 domain is Cas9 nickase. Cas9 nickase is Cas9 can cleave only one strand of a double-stranded nucleic acid molecule (for example, a double-stranded DNA molecule). It can be an protein. In some embodiments, Cas9 nickase is used as a label for double-stranded nucleic acid molecules. The target strand is cleaved, and this is because Cas9 nickase cleaves the gRNA (e.g., sgRNA) that is bound to Cas9. This means cutting the strands that form base pairs (they are complementary). Several implementation forms In this state, Cas9 nickase contains the D10A mutation and has histidine at position 840. In that embodiment, Cas9 nickase cleaves the non-target, non-base-edited strand of a double-stranded nucleic acid molecule. This is because Cas9 nickase forms a base pair with the gRNA (e.g., sgRNA) bound to Cas9. This means cutting the chain that has not been cut. In some embodiments, Cas9 nickase, Contains the H840A mutation and has an aspartic acid residue at position 10, or the corresponding mutation. It has. In some embodiments, Cas9 nickase is Cas9 nickase provided herein. One of the gauzes and at least 60%, at least 65%, at least 70%, at least 75% , at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, small At least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acids Contains an acid sequence. Additional suitable Cas9 nickases are available based on this disclosure and art knowledge. This is obvious to those skilled in the art and falls within the scope of this disclosure.

[0241] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD

[0242] In some embodiments, Cas9 is an ancient element that constitutes a domain and kingdom of a single-celled prokaryotic microorganism. This refers to Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, it is programmable. Nucleotide-binding proteins are, for example, Burstein et al., "New CRISPR-Cas systems f "rom uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 It may also be the CasX or CasY protein, and its entirety is clearly explained in the references. It will be incorporated into the detailed report. Using genome degradation metagenomics, the archaeal domain of life Many CRISPR-Cas systems have been identified, including Cas9, which was the first to be reported. The as9 protein is found in the largely unstudied nanoarchaea, where it activates CRISPR-Cas. It was discovered as part of the stem. In bacteria, two previously unknown systems Then, CRISPR-CasX and CRISPR-CasY were discovered, and they are the most powerful of all the chemotherapeutic cells discovered to date. Enter a compact system. In some embodiments, the base eddy described herein In the tar system, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, the base editor system described herein uses Cas9 This is replaced by CasY or a variant of CasY. Nucleic acid programmable DNA binding Other RNA-inducible DNA-binding proteins may be used as the protein (napDNAbp), and these It should be understood that this falls within the scope of this disclosure.

[0243] In some embodiments, any nucleic acid protozoon of the fusion proteins provided herein Gram-capable DNA-binding proteins (napDNAbp) can be CasX or CasY proteins. In some embodiments, napDNAbp is a CasX protein. pDNAbp is a CasY protein. In some embodiments, napDNAbp is naturally occurring At least 85%, at least 90%, at least 91%, and less than the CasX or CasY protein. At least 92%, at least 93%, at least 94%, at least 95%, at least 96%, and at least The amino acid composition is 97%, at least 98%, at least 99%, or at least 99.5% identical. Includes a column. In some embodiments, the programmable nucleotide-binding protein is heaven These are naturally occurring CasX or CasY proteins. In some embodiments, they are programmable. The nucleotide-binding protein is one of the CasX or CasY nucleotides described herein. At least 85%, at least 90%, at least 91%, at least 92%, and less than 85% of the protein. 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 9 Contains amino acid sequences that are 8%, at least 99%, or at least 99.5% identical. Other bacterial species It should be understood that the derived CasX and CasY may also be used in accordance with this disclosure. .

[0244] Exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN 87|F0NN87_SULIHCRISPR-associated Casx protein OS = Sulfolobus islandicus (HVE10 / 4 strain) G The amino acid sequence of N = SiH_0402 PE=4 SV=1 is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG.

[0245] Exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS = Sulfolob The amino acid sequence of *Cyperus usiformis* (REY15A strain) GN=SiRe_0771 PE=4 SV=1 is as follows: ru: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.

[0246] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSY KSGKQPFVGAWQAFYKRRLKEVWKPNA

[0247] Example CasY ((ncbi.nlm.nih.gov / protein / APG80656.1) > APG80656.1 CRISPR-associated protein The amino acid sequence of the protein CasY (uncultured Parcubacteria bacteria) is as follows: MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI.

[0248] Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. When Cas9 binds to its target, it undergoes a conformational change, which is the nuclease domain. It positions itself and cuts the opposite strand of the target DNA. The final result of Cas9-mediated DNA cleavage is the target This is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). Subsequently, DSB is repaired by one of the following two common repair pathways: (1) Efficient but error-free (2) Non-homologous end joining (NHEJ) pathway; or (2) Homologous directional repair, which is less efficient but has higher fidelity. HDR (recovery) route.

[0249] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) is any It can be calculated by a simple method. For example, in some embodiments, the efficiency is It can be expressed as a percentage of successful HDR. For example, a surveyor nuclea. -A cleavage assay was used to generate the cleavage product, and the percentage was calculated using the ratio of the product to the substrate. It is possible to calculate the new limitations that are incorporated as a result of successful HDR implementation. A nuclease enzyme for testing that directly cuts DNA containing the sequence can be used. This indicates that a higher amount of substrate is used results in a higher percentage of HDR (higher HDR efficiency). As a typical example, the percentage of HDR can be calculated using the following formula. [(Cleavage product) / (Substrate + Cleavage product)] (For example, (b+c) / (a+b+c) where "a" is the DNA substrate) This represents band strength, where "b" and "c" are cleavage products.

[0250] In some embodiments, efficiency can be expressed as the success rate of the NHEJ. For example, T7 end A nuclease I assay was used to generate cleavage products, and the ratio of the product to the substrate was used to determine the NHEJ (Nuclease Emissions). The percentage can be calculated. T7 endonuclease I is wild type and phagocytic. Natural mutant DNA strands (NHEJ) have small random insertions or deletions at the initial break site. Mismatch resulting from hybridization of (indels) It cleaves heterodouble-stranded DNA. The more cleavage there is, the higher the percentage of NHEJ (higher efficiency of NHEJ). This shows that... For example, the percentage of NHEJ is given by the formula (1-(1-(b+c) / (a (+b+c)) 1 / 2 This can be calculated using ) × 100, where "a" is the band intensity of the DNA substrate. Yes, "b" and "c" are cleavage products (Ran et. al., Cell. 2013 Sep. 12; 154(6): 1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308).

[0251] The NHEJ repair pathway is the most active repair mechanism, and involves small molecule nucleotide insertion at the DSB site. This frequently causes deletions (indels). The randomness of NHEJ-mediated DSB repair is due to Cas9 and gR Cell populations expressing NA or guide polynucleotides produce a diverse array of mutations. Therefore, it has important practical implications. In most embodiments, the NHEJ is small in the target DNA. This generates indels, and as a result, within the open reading frame (ORF) of the target gene. This can result in an amino acid deletion, insertion, or frameshift mutation that leads to an immature stop codon. The ideal final outcome is a loss-of-function mutation within the target gene.

[0252] NHEJ-mediated DSB repair often disrupts the open reading frame of genes, but Same-Sex Directed Repair (HDR) involves the addition of fluorophores or tags from single nucleotide changes. It can be used to generate specific nucleotide changes ranging from large insertions to large ones. HDR To use it for gene editing, a DNA repair template containing the desired sequence is created using gRNA( (Multiple doses are possible) and delivery to the cell type of interest together with Cas9 or Cas9 nickase. This is possible. The repair template allows for the desired edit as well as the immediate upstream and downstream (left) of its target. The right homologous arm (called the right homologous arm) may contain additional homologous arrangements. The length of each homologous arm The size may depend on the magnitude of the change being introduced, and a larger insertion will result in a longer homologous insertion. A repair template is required. The repair template consists of single-stranded oligonucleotides and double-stranded oligonucleotides. It can be a rheotide or a double-stranded DNA plasmid. The efficiency of HDR depends on Cas9, gRNA, and exogenous factors. Even in cells expressing sex repair templates, the percentage is generally low (less than 10% modified alleles). Since DR occurs between the S phase and G2 phase of the cell cycle, synchronizing cells improves the efficiency of HDR. It can be enhanced. Chemically or genetically inhibiting the genes involved in NHEJ can also enhance HDR. The frequency can be increased.

[0253] In some embodiments, Cas9 is a modified Cas9. A predetermined gRNA targeting sequence is used. The entire mass may have additional parts that exhibit partial homology. These parts are off This is called a target, and it needs to be considered when designing gRNA. Optimizing gRNA design In addition, the specificity of CRISPR can be enhanced by modifying Cas9. It performs double-strand cleavage (DSB) via the combined activity of two nuclease domains, RuvC and HNH. Cas9 nicasse, a D10A mutant of SpCas9, produces one nuclease dominance. It retains the DNA and generates DNA nicks instead of DSBs. HDR-mediated genes for specific gene editing. The Nickase system can also be combined with the data editing process.

[0254] In some embodiments, Cas9 is a variant Cas9 protein. The lipeptide differs from the amino acid sequence of the wild-type Cas9 protein at the single amino acid level. It has an amino acid sequence that includes (for example, deletions, insertions, substitutions, and fusions). Furthermore, the variant Cas9 polypeptide reduces the nuclease activity of the Cas9 polypeptide. It has amino acid changes (e.g., deletion, insertion, or substitution). For example, in some cases The variant Cas9 polypeptide exhibits the nuclease activity of the corresponding wild-type Cas9 protein. It has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. In some embodiments, the variant Cas9 protein has virtually no nuclease activity. The target Cas9 protein is a variant Cas9 protein that does not have substantial nuclease activity. If it is of the chromosome type, it may be referred to as "dCas9".

[0255] In some embodiments, the variant Cas9 protein has reduced nuclease activity. For example, a variant Cas9 protein is a wild-type Cas9 protein, for example, wild-type Cas9 Less than approximately 20% of the protein's endonuclease activity, less than approximately 15%, less than approximately 10%, less than approximately 5%, This indicates less than approximately 1%, or less than approximately 0.1%.

[0256] In some embodiments, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. While cleavage is possible, the ability to cleave the non-complementary strand of the double-stranded guide target sequence is reduced. For example, variant Cas9 protein is a mutation that reduces the function of the RuvC domain (A It may have (amino acid substitution). As a non-limiting example, in some embodiments, The rianto-Cas9 protein is D10A (from aspartic acid to alanine at amino acid position 10). It has such characteristics that it can cleave the complementary strand of the double-stranded guide target sequence, but the double-stranded guide The ability to cleave the non-complementary strand of the target sequence is reduced (therefore, this variant Cas9 chain When a protein cleaves a double-stranded target nucleic acid, it produces single-strand breaks (SSBs) instead of double-strand breaks (DSBs). (Occurs) (See, for example, Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) ).

[0257] In some embodiments, the variant Cas9 protein is non-phase of the double-stranded guide target sequence. It can cleave complementary strands, but its ability to cleave complementary strands of guide target sequences is reduced. For example, variant Cas9 protein is a mutation that reduces the function of the HNH domain (Ammonia). It can have (no acid substitution) (RuvC / HNH / RuvC domain motif). Non-limiting examples include In some embodiments, the variant Cas9 protein has H840A (amino acid at position 840) It has a histidine-to-alanine mutation, and therefore cleaves the non-complementary strand of the guide target sequence. It is possible, but the ability to cleave the complementary strand of the guide target sequence is reduced (therefore, When this variant Cas9 protein cleaves the double-stranded guide target sequence, it produces single-stranded bones (SSBs) instead of double-stranded bones (DSBs). (This occurs). Such Cas9 proteins have a guide target sequence (for example, a single-stranded guide target sequence). Although the ability to cleave ) is reduced, it is still able to bind to guide target sequences (e.g., single-stranded guide target sequences). It possesses the ability to combine.

[0258] In some embodiments, the variant Cas9 protein is the complementary strand of the double-stranded target DNA and The ability to cleave both complementary and non-complementary chains is reduced. As a non-limiting example, several implementations Morphologically, the variant Cas9 protein harbors both D10A and H840A mutations, and As a result, polypeptides have the ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. It is decreasing. Such Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). Although its ability to do so is reduced, it retains its ability to bind to target DNA (e.g., single-stranded target DNA). ru.

[0259] As another non-limiting example, in some embodiments, the variant Cas9 protein is W4 The polypeptide harbors the 76A and W1126A mutations, and as a result, its ability to cleave target DNA is impaired. It is decreasing. These Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). Although its ability to bind to target DNA (e.g., single-stranded target DNA) is reduced, it still retains the ability to bind to target DNA. .

[0260] As another non-limiting example, in some embodiments, the variant Cas9 protein is P4 It harbors the 75A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, and as a result, polyp The plutidops has reduced ability to cleave target DNA. Such Cas9 proteins are unable to cleave target DNA. The ability to cleave NA (e.g., single-stranded target DNA) is reduced, but the ability to cleave target DNA (e.g., single-stranded target DNA) is reduced. It retains the ability to bind to DNA.

[0261] As another non-limiting example, in some embodiments, the variant Cas9 protein is H8 The polypeptide harbors the 40A, W476A, and W1126A mutations, and as a result, the polypeptide cleaves the target DNA. The ability to do this is reduced. These Cas9 proteins target DNA (e.g., single-stranded target DNA) Although its ability to cleave is reduced, it retains its ability to bind to target DNA (e.g., single-stranded target DNA). It possesses. As another non-limiting example, in some embodiments, the variant Cas9 tamper The polypeptide harbors H840A, D10A, W476A, and W1126A mutations, and as a result, the polypeptide , the ability to cleave target DNA is reduced. Such Cas9 proteins cleave target DNA (for example Although the ability to cleave single-stranded target DNA is reduced, the ability to cleave target DNA (e.g., single-stranded target DNA) It retains the ability to bond. In some embodiments, the variant Cas9 is Cas9 HNH The catalytic His residue is restored at position 840 of the main molecule (A840H).

[0262] As another non-limiting example, in some embodiments, the variant Cas9 protein is H8 It harbors the 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, and as a result, Polypeptides have a reduced ability to cleave target DNA. These Cas9 proteins, The ability to cleave target DNA (e.g., single-stranded target DNA) is reduced, but the ability to cleave target DNA (e.g., single-stranded target DNA) is reduced. It retains the ability to bind to the target DNA strand. As another non-limiting example, in some embodiments... The variant Cas9 proteins are D10A, H840A, P475A, W476A, N477A, ​​D1125A, and W1126A. , and harboring the D1127A mutation, as a result the polypeptide has the ability to cleave target DNA It is decreasing. Such Cas9 proteins cleave target DNA (e.g., single-stranded target DNA). Although its ability to do so is reduced, it retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 protein has W476A and W1126A mutations. If the variant Cas9 protein is P475A, W476A, N477A, ​​D1125A, W1126 When harboring A and D1127A mutations, the variant Cas9 protein efficiently transforms into the PAM sequence. It does not bind. Therefore, in such cases, this variant Cas9 protein does not bind. When used in law, this method does not require a PAM sequence. In other words, in some embodiments When using such variant Cas9 proteins in the binding method, this method uses guide RNA. This may include, but this method can be performed in the absence of the PAM sequence (hence the specificity of binding). (This is brought about by the target segment of the guide RNA). In order to achieve the above effect, It is possible to mutate the residues (i.e., inactivate one or the other nuclease moiety). As a non-limiting example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A98 4, D986, and / or A987 may be modified (i.e., replaced). Also, the alanine position Mutations other than substitutions are also acceptable.

[0263] In some embodiments, variant Cas9 proteins having reduced catalytic activity (for example) Cas9 protein D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986 , and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A (If it has H982A, H983A, A984A, and / or D986A) it interacts with the guide RNA. As long as it retains its ability to act, it can still bind to target DNA in a site-specific manner. (Because the target DNA sequence is still induced by the guide RNA.)

[0264] In some embodiments, the variant Cas protein is spCas9, spCas9-VRQR, spCas9-V RER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9- It could be LRVSQL.

[0265] In some embodiments, amino acid substitutions include D1135M, S1136Q, G1218K, E1219F, A1322R, and D1 Modified SpC containing 332A, R1335E, and T1337R, having specificity for modified PAM 5'-NGC-3'. I used as9 (SpCas9-MQKFRAER).

[0266] As an alternative to Cas9 in S. pyogenes, the Cpf1 family, which exhibits cleavage activity in mammalian cells, is used. It may also contain RNA-induced endonucleases from Prevotella and Francisella 1. The next generation of CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. 1 is an RNA-induced endonuclease of the class II CRISPR / Cas system. This adaptive immunity The mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus. It uses an endonuclease that employs guide RNA to locate and cleave viral DNA. It codes for CRISPR / Cas9. Cpf1 is a smaller and simpler endonuclease than Cas9. Overcoming some of the system's limitations. Unlike Cas9 nucleases, DNA mediated by Cpf1. The result of the cleavage is a double-strand break with a short 3' overhang. (Cpf1 stagge) The cleavage pattern of (red) is directional, similar to traditional restriction enzyme cloning. This opens up the possibility of introducing offspring, which could increase the efficiency of gene editing. Similar to variants and orthologues, Cpf1 is a site that CRISPR can target. The number of AT-rich regions or AT-rich genomes that lack the NGG PAM site preferred by SpCas9 are expanded. It can also be said that the Cpf1 locus is an alpha / beta mixed domain, RuvC-I and the helix that follows it. It contains the region, RuvC-II and zinc finger-like domain. The Cpf1 protein contains the Ru of Cas9 It possesses a RuvC-like endonuclease domain similar to the vC domain. Furthermore, Cpf1 is HNH It lacks an endonuclease domain, and the N-terminus of Cpf1 is the alpha helix recognition lobe of Cas9. It does not have the Cpf1 CRISPR-Cas domain configuration, which means that Cpf1 is functionally unique and class 2. It was shown to be classified as a type V CRISPR system. The Cpf1 locus is a type II system locus. It encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III than the stem. Functional Cpf1 does not require trans-activated CRISPR RNA (tracrRNA), and therefore, CR It requires only ISPR (crRNA). Cpf1 is not only smaller than Cas9, but also smaller than sgRNA molecules (C Because it has about half the number of nucleotides as Cas9, this is beneficial for genome editing. In contrast to the targeted G-rich PAM, the Cpf1-crRNA complex has a protospacer adjacent motif. The target DNA or RNA is cleaved by the identification of f5'-YTN-3'. After PAM identification, Cpf1 also This introduces sticky end-like DNA double-strand breaks with a 5-nucleotide overhang. do.

[0267] [Cas12 domain of nucleic acid base editor] Typically, microbial CRISPR-Cas systems are divided into Class 1 and Class 2 systems. Class 1 systems have a multi-subunit effects complex, while Class 2 The system has a single protein effector. For example, Cas9 and Cpf1 are different Despite being of different types (Type II and Type V respectively), they are Class 2 effects pedals. In addition to Cpf1, Class 2, Type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, and Ca This also includes s12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. For example, Shmakov et al., “Discovery and Functional Characterization of Diverse Cla ss 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 60(3): 385-397; Makarova et a l., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1(5): 325-336; and Yan et al., “Functionally Diverse Typ See e V CRISPR-Cas Systems, Science, 2019 Jan. 4; 363: 88-91 (these The literature is incorporated herein by reference in its entirety. The V-type Cas protein is RuvC ( It contains an endonuclease domain (or RuvC-like). It is involved in the production of mature CRISPR RNA (crRNA). Generally, tracrRNA is independent, but for example, Cas12b / C2c1 is independent of tracrRNA in the production of crRNA. A is required. Cas12b / C2c1 relies on both crRNA and tracrRNA for DNA cleavage. .

[0268] The nucleic acid programmable DNA-binding proteins intended in this invention are classified as Class 2, Type V. It contains Cas proteins (Cas12 proteins). Limited Cas class 2, type V proteins. Examples of those that are not affected include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, and Cas12e / CasX. This includes Cas12g, Cas12h, and Cas12i, their homologs or modifications. When present, the Cas12 protein also contains Cas12 nuclease, Cas12 domain, or Cas12 nuclease. It may also be called a protein domain. In some embodiments, the Cas12 protein of the present invention It is interrupted by a protein domain fused inside, such as a deaminase domain. It contains the amino acid sequence.

[0269] In some embodiments, the Cas12 domain is a nuclease-inactive Cas12 domain or This is a Cas12 nickase. In some embodiments, the Cas12 domain is nuclease activity It is a domain. For example, the Cas12 domain is one of the two strands of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). It may be a Cas12 domain that inserts a nick into the chain. In some embodiments, the Cas12 domain n comprises any one of the amino acid sequences described herein. In some embodiments, C The as12 domain has at least one amino acid sequence described herein. 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, At least 90%, at least 95%, at least 96%, at least 97%, at least 98%, less Containing amino acid sequences that are identical by at least 99%, or at least 99.5%. Several embodiments Therefore, the Cas12 domain is compared to any one of the amino acid sequences described herein. , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21 , 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43 , containing amino acid sequences with 44, 45, 46, 47, 48, 49, 50 or more mutations In some embodiments, the Cas12 domain is the amino acid sequence described herein. Compared to any one of them, at least 10, at least 15, at least 20, at least 30, At least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, At least 350, at least 400, at least 500, at least 600, at least 700, fewer at least 800, at least 900, at least 1000, at least 1100, or at least 1200 It contains an amino acid sequence having identical consecutive amino acid residues.

[0270] In some embodiments, a protein containing a fragment of Cas12 is provided. For example, In some embodiments, the protein contains one of the following two Cas12 domains: (1) (2) gRNA binding domain of Cas12; (3) DNA cleavage domain of Cas12. In some embodiments, Cas Proteins containing Cas12 or fragments of Cas12 are called "Cas12 variants." It shares homology with Cas12 or a fragment thereof. For example, Cas12 variants share homology with wild-type Ca s12 is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least They are approximately 99% identical, at least approximately 99.5% identical, or at least approximately 99.9% identical. In this embodiment, the Cas12 variant is 1, 2, 3, 4, 5, 6, 7, 8 compared to wild-type Cas12. , 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28 , 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 It may have 49, 50 or more amino acid changes. In some embodiments, Cas1 The two variants contain a fragment of Cas12 (e.g., a gRNA-binding domain or a DNA-cleaving domain), The fragment is at least approximately 70% identical to the corresponding fragment of wild-type Cas12, and at least approximately 80% identical. , at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical Identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or slightly identical. At least they are approximately 99.9% identical. In some embodiments, the fragment is the same as the corresponding wild-type Cas12. At least 30%, at least 35%, at least 40%, at least 45%, of the amino acid length 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96% , at least 97%, at least 98%, at least 99%, or at least 99.5%. In some embodiments, the fragment has a length of at least 100 amino acids. In this state, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600 , 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids long.

[0271] In some embodiments, Cas12 is modified by one or more mutations that alter the Cas12 nuclease activity. This corresponds partially or entirely to the mutated Cas12 amino acid sequence, or These mutations include, for example, the Ruv of Cas12. This includes amino acid substitutions within the C nuclease domain. In some embodiments, wild-type C as12 is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at Approximately 95% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, and This provides a variant or homolog of Cas12 that is at least approximately 99.9% identical. In this embodiment, approximately 5 amino acids, approximately 10 amino acids, approximately 15 amino acids, approximately 20 amino acids, approximately 25 amino acids 0 amino acids, approximately 30 amino acids, approximately 40 amino acids, approximately 50 amino acids, approximately 75 amino acids, approximately 100 amino acids or more It provides short or long Cas12 variants.

[0272] In some embodiments, a Cas12 fusion protein such as the one provided herein is Cas12 This includes the full-length amino acid sequence of a protein, for example, one of the Cas12 sequences provided herein. In other embodiments, the fusion protein provided herein contains a full-length Cas12 sequence. It does not contain, but contains only one or more fragments thereof. The appropriate exemplary amino acid sequence of the Cas12 domain is provided. Additional suitable sequences of the Cas12 domain and fragments provided herein will be obvious to those skilled in the art. It is likely.

[0273] Generally, class 2, type V Cas proteins are single functional RuvC endonucleases. Having (for example, Chen et al., “CRISPR-Cas12a target binding unleashes indi See "scriminate single-stranded DNase activity," Science 360:436-439 (2018). (This is the case). In some cases, the Cas12 protein is the variant Cas12b protein (Strec See Ker et al., Nature Communications, 2019, 10(1): Art. No.: 212. In this embodiment, the variant Cas12 polypeptide has the amino acid sequence of the wild-type Cas12 protein. When compared, there are differences in 1, 2, 3, 4, and 5 or more amino acids (e.g., deletion, insertion, substitution, fusion). It has an amino acid sequence that has ( ). In some cases, the variant Cas12 polypeptide is Amino acid changes (e.g., deletion, insertion, or substitution) that reduce the activity of the Cas12 polypeptide. ) has. For example, in some cases, the variant Cas12 has the corresponding wild-type Cas12b. Less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less...

Claims

1. A pharmaceutical composition for use in a method for treating alpha-1 antitrypsin deficiency in subjects containing a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency, wherein the pharmaceutical composition comprises lipid nanoparticles containing a guide RNA and mRNA encoding a base editor, The base editor comprises a nickase and a nucleic acid programmable DNA-binding protein (napDNAbp) domain containing an amino acid sequence having at least 90% identity with SEQ ID NO: 13, and an adenosine deaminase domain, wherein the SNP associated with alpha-1 antitrypsin deficiency refers to SEQ ID NO: 21 and results in the expression of an alpha-1 antitrypsin polypeptide having lysine at amino acid position 342, where the first E of SEQ ID NO: 21 is taken as position 1, and the guide RNA targets the base editor to result in the modification of the SNP associated with alpha-1 antitrypsin deficiency and comprises the nucleotide sequence AUCGACAAGAAAGGGACUGA. Pharmaceutical composition.

2. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD The pharmaceutical composition according to claim 1, comprising an amino acid sequence that has at least 90% identity with the above TadA*7.10 amino acid sequence and further contains arginine or threonine at amino acid position 147.

3. The pharmaceutical composition according to claim 2, wherein the adenosine deaminase domain further comprises an I76Y amino acid modification and / or a Q154S amino acid modification.

4. The pharmaceutical composition according to claim 2, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 95% identity with the TadA*7.10 amino acid sequence.

5. The aforementioned guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' A pharmaceutical composition according to claim 1, comprising:

6. The pharmaceutical composition according to claim 1, wherein the napDNAbp domain has specificity for a protospacer adjacent motif containing the nucleotide sequence 5'-NGC-3'.

7. The pharmaceutical composition according to claim 1, wherein the napDNAbp domain further comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to Sequence ID No.

13.

8. Lipid nanoparticles comprising a guide RNA and mRNA encoding a base editor, wherein the base editor comprises a nickasase and a nucleic acid programmable DNA-binding protein (napDNAbp) domain containing an amino acid sequence having at least 90% identity to SEQ ID NO: 13, and an adenosine deaminase domain, the guide RNA targets the base editor to result in a modification of a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency and comprises the nucleotide sequence AUCGACAAGAAAGGGACUGA, the SNP associated with alpha-1 antitrypsin deficiency referring to SEQ ID NO: 21 and resulting in the expression of an alpha-1 antitrypsin polypeptide having lysine at amino acid position 342, where the first E of SEQ ID NO: 21 is at position 1, the lipid nanoparticles.

9. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD Lipid nanoparticles according to claim 8, which have at least 90% identity with respect to and further include an amino acid sequence that contains arginine or threonine at amino acid position 147 compared to the above TadA*7.10 amino acid sequence.

10. The aforementioned guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' Lipid nanoparticles according to claim 8, comprising:

11. The lipid nanoparticle according to claim 8, wherein the napDNAbp domain has specificity for a protospacer adjacent motif containing the nucleotide sequence 5'-NGC-3'.

12. The lipid nanoparticle according to claim 8, wherein the napDNAbp domain further comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to SEQ ID NO:

13.

13. A pharmaceutical composition comprising lipid nanoparticles according to any one of claims 8 to 12 and a pharmaceutically acceptable carrier, vehicle, or excipient.

14. A base editor system comprising a guide RNA and an mRNA encoding a base editor, wherein the base editor comprises a nickase and a nucleic acid programmable DNA-binding protein (napDNAbp) domain containing an amino acid sequence having at least 90% identity to SEQ ID NO: 13, and an adenosine deaminase domain, the guide RNA targets the base editor to produce a modification of a single nucleotide polymorphism (SNP) associated with alpha-1 antitrypsin deficiency, the SNP associated with alpha-1 antitrypsin deficiency being in an alpha-1 antitrypsin polynucleotide, and resulting in the expression of an alpha-1 antitrypsin polypeptide having lysine at amino acid position 342 with reference to SEQ ID NO: 21, where the first E in SEQ ID NO: 21 is taken as position 1, and the guide RNA comprises the nucleotide sequence AUCGACAAGAAAGGGACUGA.

15. The adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD The base editor system according to claim 14, comprising an amino acid sequence that has at least 90% identity with the above TadA*7.10 amino acid sequence and further contains arginine or threonine at amino acid position 147.

16. The aforementioned guide RNA has the following nucleotide sequence: 5'-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3' A base editor system according to claim 14, including the above.

17. The base editor system according to claim 14, wherein the napDNAbp domain has specificity for a protospacer adjacent motif containing the nucleotide sequence 5'-NGC-3'.

18. The base editor system according to claim 14, wherein the napDNAbp domain further comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R with reference to Sequence ID No. 13.