A deaminase, base editor comprising the same, and application thereof

CN122161928APending Publication Date: 2026-06-05YOLTECH THERAPEUTICS CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YOLTECH THERAPEUTICS CO LTD
Filing Date
2024-10-25
Publication Date
2026-06-05

Smart Images

  • Figure 00000063_0000
    Figure 00000063_0000
  • Figure 00000063_0001
    Figure 00000063_0001
  • Figure 00000064_0000
    Figure 00000064_0000
Patent Text Reader

Abstract

Disclosed are a deaminase, a base editor comprising the same, and applications thereof. The deaminase comprises the following sequence: (i) an amino acid sequence as shown in SEQ ID NO: 10; or (ii) an amino acid sequence having at least 80% sequence identity with the amino acid sequence as shown in SEQ ID NO: 10. The provided deaminase can improve editing efficiency when used in a base editor and a base editing system, and has a prospect of clinical application.
Need to check novelty before this filing date? Find Prior Art

Description

A deaminase, a base editor containing the same, and applications thereof

[0001] This application claims the benefit of Chinese Patent Application No. 2023114029610, filed on October 25, 2023. This application incorporates the entirety of the aforementioned Chinese Patent Application. Technical Field

[0002] The present invention belongs to the field of gene editing, and specifically relates to a deaminase, a base editor containing the same, and applications thereof. Background Art

[0003] How to accurately and efficiently modify the genome is a key research goal in the life sciences. Traditional CRISPR / Cas9 technology creates double-strand breaks (DSBs) at target sites, thereby inducing homologous recombination (HR) and non-homologous end joining (NHEJ) repair pathways within cells, thereby achieving site-specific modifications such as knockout, replacement, and insertion of genomic DNA. However, DSB-induced DNA repair struggles to achieve efficient and stable single-base mutations.

[0004] Currently available base editors include cytidine base editors that convert the target C·G base pair to T·A (e.g., BE4) and adenine base editors that convert A·T to G·C (e.g., ABE8e). For the use of higher editing efficiency, the use of the editing efficiency of existing base editors may be limited. This field requires base editors with higher specificity and editing efficiency, and it is necessary to improve base editors.

[0005] Summary of the Invention

[0006] The technical problem to be solved by the present invention is the defect that the existing technology lacks base editors with higher specificity and editing efficiency. A deaminase, a base editor containing the same, and its application are provided. The deaminase provided by the present invention can improve the editing efficiency when constituting a base editor and being used in a base editing system, and has the prospect of clinical application.

[0007] The present invention solves the above technical problems through the following technical solutions.

[0008] The first aspect of the present invention provides a deaminase comprising the following sequence:

[0009] (i) the amino acid sequence shown in SEQ ID NO: 10; or

[0010] (ii) an amino acid sequence that has at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% sequence identity to the amino acid sequence of SEQ ID NO: 10, and which retains the deamination activity of the deaminase of the amino acid sequence of SEQ ID NO: 10; and is not SEQ ID NO: 1.

[0011] In some embodiments, the amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 10 is an amino acid sequence obtained by adding, substituting, deleting or inserting one or more amino acid residues into the amino acid sequence shown in SEQ ID NO: 10.

[0012] In some embodiments, the substitution occurs at one or more of the following positions of the amino acid sequence as shown in SEQ ID NO: 10: C46, ​​Y47, G48, H49, C144, Q145, F146, Y147, Q148, Q149, P150, R151, E152, V153, F154, N155, A156, E157, R158, E159, A160, R161, R162, L163, N164, Q165, P166, D167, R168, A169, and D170.

[0013] In some preferred embodiments, the substitution is a combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10:

[0014] (1) 7 or 10 of C46, ​​Y47, G48, H49, Q148, P150, E152, V153, F154, and N155;

[0015] (2) 6, 7, or 8 of Q148, Q149, P150, R151, E152, V153, F154, and N155;

[0016] (3) 7 or 8 of A156, E157, R158, E159, A160, R161, R162, and L163;

[0017] (4) 6 or 7 of N164, Q165, P166, D167, R168, A169, and D170;

[0018] (5) 6 or 7 of C144, Q145, F146, Y147, Q148, Q149, and P150; or

[0019] (6) 2 or 9 of C144, Q145, Q148, Q149, P150, E152, V153, F154, and N155.

[0020] In some preferred embodiments, the substitution is a substitution at any combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10:

[0021] (1) C46, ​​Y47, G48, H49, Q148, P150, E152, V153, F154, and N155;

[0022] (2) G48, Q148, P150, E152, V153, F154, and N155;

[0023] (3) Q148, Q149, P150, E152, V153, F154, and N155;

[0024] (4) Q148, P150, E152, V153, F154, and N155;

[0025] (5) Q149, P150, R151, E152, V153, F154, and N155;

[0026] (6) Q148, Q149, P150, R151, E152, V153, F154, and N155;

[0027] (7) Q148, Q149, P150, E152, F154, and N155;

[0028] (8) A156, E157, R158, E159, A160, R161, R162, and L163;

[0029] (9) A156, E157, R158, E159, A160, R162, and L163;

[0030] (10) N164, Q165, P166, D167, R168, A169, and D170;

[0031] (11) N164, Q165, D167, R168, A169, and D170;

[0032] (12) C144, Q145, F146, Y147, Q148, Q149, and P150;

[0033] (13) C144, Q145, F146, Y147, Q148, and P150;

[0034] (14) C144, Q145, Q148, Q149, P150, E152, V153, F154, and N155;

[0035] (15)C144 and Q145.

[0036] In some embodiments, the substitutions at the positions are selected from the group consisting of: C46P, Y47I, G48A / T, H49R, C144T / L / W, Q145L / K, F146A / R, Y147S / F, Q148R / G / T / S / C, Q149N / P / R / V / G / F / C / K, P150A / L / R / I / G / S / T, R151K / P, E152P / L / Q / H / S, V153T / A / F / Y / K / P, F154S / P / V / N / L / D / H, N155P / G / T / S / Y / R / A, A156L / T, E157F, R158N / L, E159L / H, A1 60K / T, R161K, R162L / K, L163D / I, N164G / R, Q165T / L, P166Q, D167L, R168L / P, A169N / T and D170R / H.

[0037] In the present invention, " / " indicates that the solutions before and after the symbol are in an "or" relationship. For example, C144T / L / W means that the C at position 144 can be substituted with T, L, or W.

[0038] In some specific embodiments, the substitutions occurring at the positions are selected from the group consisting of: C46P, Y47I, G48A / T, H49R, C144T / L / W, Q145L / K, F146A / R, Y147S / F, Q148R / T / C, Q149N / P / R / G / C / K, P150A / L / R / G / S / T, R151P, E152P / L / Q / H, V153T / A / F / K, F154S / P / V / L / D / H, N155P / G / T / S / A, A156L / T, E157F, R158N / L, E159L / H, A160K / T, R161K, R162L / K, L163D / I, N164G / R, Q165T / L, P166Q, D167L, R168L / P, A169N / T and D170R / H.

[0039] In some specific embodiments, the substitution is a combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10:

[0040] (1)Q148R、Q149N、P150A、E152P、V153T、F154S、N155P;

[0041] (2)Q148R、P150L、E152L、V153A、F154P、N155G;

[0042] (3)Q148R、Q149R、P150R、E152P、V153F、F154V、N155T;

[0043] (4)Q148G、Q149G、P150I、E152L、V153Y、F154N、N155S;

[0044] (5)Q149P、P150G、R151K、E152Q、V153K、F154L、N155P;

[0045] (6)Q148R、Q149V、P150S、R151P、E152L、V153F、F154P、N155Y;

[0046] (7)Q148T、Q149G、P150R、E152H、V153A、F154D、N155S;

[0047] (8)Q148S、Q149F、P150L、E152S、V153P、F154L、N155R;

[0048] (9)Q148G、Q149C、P150S、E152P、F154H、N155A;

[0049] (10)A156L、E157F、R158N、E159L、A160K、R161K、R162L、L163D;

[0050] (11)A156T、E157F、R158L、E159H、A160T、R162K、L163I;

[0051] (12)N164G、Q165T、P166Q、D167L、R168L、A169N、D170R;

[0052] (13)N164R、Q165L、D167L、R168P、A169T、D170H;

[0053] (14)C144T、Q145L、F146A、Y147S、Q148R、Q149K、P150S;

[0054] (15)C144L, Q145L, F146R, Y147F, Q148C, P150T;

[0055] (16)C144W, Q145K, Q148R, Q149R, P150R, E152P, V153F, F154V, N155T;

[0056] (17) C144W, Q145K;

[0057] (18)G48A, Q148R, P150L, ​​E152L, V153A, F154P, N155G;

[0058] (19)C46P, Y47I, G48T, H49R, Q148R, P150L, ​​E152L, V153A, F154P, N155G.

[0059] The second aspect of the present invention provides a base editor fusion protein, which comprises the deaminase as described in the first aspect, and a nucleic acid programmable nucleotide binding domain.

[0060] In some embodiments, the nucleic acid programmable nucleotide binding domain is a Cas protein or an AGO protein.

[0061] In some embodiments, the Cas protein is selected from Cas9, CasX, CasY, Cpf1, C2c1, C2c2, and C2c3.

[0062] In some embodiments, the AGO protein is selected from the group consisting of pAgo, eAgo, Ago1, Ago2, Ago3, and Ago4.

[0063] In some embodiments, the deaminase is linked to one end of the nucleic acid programmable nucleotide binding domain or is embedded in the nucleic acid programmable nucleotide binding domain.

[0064] In some preferred embodiments, the connection is direct connection or connection via a linker. The linker preferably comprises an amino acid sequence as shown in one or more of SEQ ID NOs: 32-41.

[0065] In some preferred embodiments, the chimeric site is located in the carboxyl terminal domain of the nucleic acid programmable nucleotide binding domain.

[0066] In some preferred embodiments, the nucleic acid programmable nucleotide binding domain retains partial or no nucleotide chain cleavage activity.

[0067] In some embodiments, the base editor fusion protein further comprises a nuclear localization signal sequence; the nuclear localization signal sequence is connected to the N-terminus and / or C-terminus of the base editor fusion protein, and / or, to the N-terminus and / or C-terminus of the deaminase.

[0068] In some preferred embodiments, the nuclear localization signal sequence is connected to the N-terminus and C-terminus of the base editor fusion protein.

[0069] In some specific embodiments, the structure of the base editor fusion protein from N-terminus to C-terminus is: nuclear localization signal sequence-deaminase-nucleic acid programmable nucleotide binding domain-nuclear localization signal sequence.

[0070] In other specific embodiments, the structure of the base editor fusion protein from N-terminus to C-terminus is: nuclear localization signal sequence-deaminase-nucleic acid programmable nucleotide binding domain-nuclear localization signal sequence.

[0071] In some embodiments, when the nucleic acid programmable nucleotide binding domain is a Cas protein, such as a Cas9 protein, the chimeric site is located between positions 1249-1250 of Cas9.

[0072] In some embodiments, the base editor fusion protein comprises an amino acid sequence as shown in any one of SEQ ID NO:10, 18, 20 and 22.

[0073] A third aspect of the present invention provides a base editing system comprising:

[0074] (i) a deaminase and a nucleic acid programmable nucleotide binding domain as described in the first aspect;

[0075] or (ii) the base editor fusion protein as described in the second aspect,

[0076] and a guide polynucleotide;

[0077] The nucleic acid programmable nucleotide binding domain or the base editor fusion protein forms a ribonucleoprotein complex with the guide polynucleotide and binds to the target nucleic acid under the guidance of the guide polynucleotide.

[0078] The fourth aspect of the present invention provides a polynucleotide encoding the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, or the base editing system as described in the third aspect.

[0079] In some specific embodiments, the polynucleotide encoding the base editor fusion protein comprises a nucleotide sequence as shown in any one of SEQ ID NO:11, 19, 21 and 23.

[0080] The fifth aspect of the present invention provides a vector comprising the polynucleotide according to the fourth aspect.

[0081] In some embodiments of the present invention, the polynucleotide is located on one or more vectors.

[0082] In some embodiments, the polynucleotide is operably linked to a promoter.

[0083] In some embodiments, the promoter is selected from one or more of a constitutive promoter, an inducible promoter, a ubiquitin promoter, a cell type-specific promoter, and a tissue-specific promoter.

[0084] The sixth aspect of the present invention provides an isolated cell comprising the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the polynucleotide as described in the fourth aspect and / or the vector as described in the fifth aspect.

[0085] In some embodiments, the cell is a prokaryotic cell or a eukaryotic cell; for example, selected from an animal cell, a plant cell, and a fungal cell.

[0086] In some preferred embodiments, the cell is a vertebrate cell or an invertebrate cell; the vertebrate cell is preferably a mammalian cell.

[0087] In some preferred embodiments, the mammalian cell is selected from a rodent cell, a primate cell, and a non-primate cell; the primate cell is, for example, a human cell.

[0088] The seventh aspect of the present invention provides a pharmaceutical composition, which comprises the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect and / or the cell as described in the sixth aspect, and optionally a pharmaceutically acceptable carrier and / or excipient.

[0089] The eighth aspect of the present invention provides a kit comprising the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect, the cell as described in the sixth aspect and / or the pharmaceutical composition as described in the seventh aspect.

[0090] The ninth aspect of the present invention provides a delivery system comprising the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect, the cell as described in the sixth aspect, the pharmaceutical composition as described in the seventh aspect and / or the kit as described in the eighth aspect.

[0091] In some embodiments, the delivery vehicle is selected from the group consisting of liposomes, nanoparticles, viral vectors, exosomes, microvesicles, and cell-penetrating peptides.

[0092] The tenth aspect of the present invention provides a base editing method, which comprises the steps of contacting the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, or the base editing system as described in the third aspect with the target nucleic acid and causing a deamination reaction.

[0093] In some embodiments, the base editing method is in vivo or in vitro.

[0094] In some embodiments, the base editing method is for non-diagnostic or non-therapeutic purposes.

[0095] The eleventh aspect of the present invention provides the use of the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect, the cell as described in the sixth aspect, the pharmaceutical composition as described in the seventh aspect, the kit as described in the eighth aspect, or the delivery system as described in the ninth aspect in the preparation of a medicament for treating diseases associated with or caused by point mutations.

[0096] In some embodiments, the disease is selected from one or more of hypercholesterolemia, transthyretin amyloidosis, alpha 1 -antitrypsin deficiency, and beta-hemoglobinopathies.

[0097] The twelfth aspect of the present invention provides a method for treating a condition or disease, comprising administering to a subject in need thereof an effective amount of the deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect, the cell as described in the sixth aspect, the pharmaceutical composition as described in the seventh aspect, the kit as described in the eighth aspect and / or the delivery system as described in the ninth aspect.

[0098] The thirteenth aspect of the present invention provides a deaminase as described in the first aspect, the base editor fusion protein as described in the second aspect, the base editing system as described in the third aspect, the polynucleotide as described in the fourth aspect, the vector as described in the fifth aspect, the cell as described in the sixth aspect, the pharmaceutical composition as described in the seventh aspect, the kit as described in the eighth aspect, or the delivery system as described in the ninth aspect as a drug.

[0099] The fourteenth aspect of the present invention provides a deaminase as described in the first aspect, a base editor fusion protein as described in the second aspect, a base editing system as described in the third aspect, a polynucleotide as described in the fourth aspect, a vector as described in the fifth aspect, a cell as described in the sixth aspect, a pharmaceutical composition as described in the seventh aspect, a kit as described in the eighth aspect, or a delivery system as described in the ninth aspect for the treatment of a condition or disease. In some embodiments, the condition or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the condition or disease includes the diseases shown in the following table:

[0100] In some embodiments, the disease or condition comprises one or more of hypercholesterolemia, transthyretin amyloidosis, alpha 1 -antitrypsin deficiency, and beta-hemoglobinopathies.

[0101] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.

[0102] The reagents and raw materials used in the present invention are commercially available.

[0103] The positive progress effect of the present invention is:

[0104] The deaminase provided by the present invention greatly improves the editing efficiency when constituting a base editor and being used in a base editing system, and can be used to modify pathogenic DNA targets. For example, a base editor can be used to direct mutation of adenine (A) in nucleic acids (such as DNA) to guanine (G). By changing the amino acid sequence of the protein through such changes, the start codon is destroyed or a new one is created, or a stop codon is created, to destroy the splice donor, to destroy the splice acceptor or to edit the regulatory sequence, thereby correcting the pathogenic gene to achieve the purpose of treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 shows the editing efficiency of 005V1-nCas9 and 5V3354-nCas9 base editors for PCSK9 targets.

[0106] Figure 2 shows the editing efficiency of various mutant base editors on the PCSK9 target site.

[0107] Figure 3 shows the base editing efficiency of 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, and 5V22.2-1249-nCas9 at the PCSK9 gene targeting site (A6 site). DETAILED DESCRIPTION

[0108] the term

[0109] The term "mutant" refers to a protein produced by mutation or recombinant DNA procedures.

[0110] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. Deaminases herein are nucleobase deaminases, and the terms "deaminase" and "nucleobase deaminase" are used interchangeably herein. A deaminase can be a naturally occurring deaminase or an active fragment or variant thereof. A deaminase can be active on single-stranded nucleic acids, such as ssDNA or ssRNA, or on double-stranded nucleic acids, such as dsDNA or dsRNA. In some embodiments, a deaminase can only deaminate ssDNA and has no effect on dsDNA. In some embodiments, the deaminase is an adenosine deaminase or a cytidine deaminase.

[0111] The term "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, a polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze the hydrolytic deamination reaction of converting adenine (or the adenine portion of a molecule) into hypoxanthine (or the hypoxanthine portion of a molecule). In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). Adenosine deaminases include, but are not limited to, members of the enzyme family known as adenosine deaminases acting on RNA (ADARs), members of the enzyme family known as adenosine deaminases acting on tRNA (ADATs), and other family members containing adenosine deaminase domains (ADADs). According to the present disclosure, adenosine deaminases can target adenine in RNA / DNA and RNA duplexes. In specific embodiments, adenosine deaminases have been modified to increase their ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.

[0112] The term "base editor" is a fusion protein including a nucleic acid programmable nucleotide binding protein (napDNAbp) (such as a nuclease) and a deaminase. "Base Editor (BE)" or "Core Base Editor" refers to a reagent that binds to a polynucleotide and has a core base modification activity. In various embodiments, the base editor comprises a core base modification polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide binding domain (e.g., a nucleic acid programmable DNA binding protein) bound to a guide polynucleotide (e.g., a guide RNA). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute protein (AGO). In various embodiments, the reagent is a biomolecular complex comprising a protein domain with base editing activity, i.e., capable of modifying bases (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA, RNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or connected to a deaminase domain. In one embodiment, the reagent is a fusion protein comprising a domain having base editing activity. In some embodiments, the domain having base editing activity is capable of deaminating bases within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is an adenosine base editor (ABE).

[0113] The term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. In some embodiments, the DNA binding polypeptide is an endonuclease that is capable of cleaving phosphodiester bonds between nucleotides in a nucleic acid molecule. In certain embodiments, the DNA binding polypeptide is an exonuclease that is capable of cleaving nucleotides at either end (5' or 3') of a nucleic acid molecule. In some embodiments, the nuclease is selected from the group consisting of a homing endonuclease (Meganuclease), a zinc finger nuclease (ZFN), a TAL effector DNA nuclease fusion protein (TALEN), and an RNA-guided nuclease or homologs or variants thereof, wherein the nuclease activity is reduced or inhibited.

[0114] The term "homing endonuclease" or "meganuclease" refers to an endonuclease that binds to a recognition site within dsDNA of 12 to 40 bp in length. Exemplary, non-limiting examples of homing endonucleases include the LAGLIDADG series. "Homing endonuclease" can refer to either a dimeric or single-chain meganuclease.

[0115] The term "zinc finger nuclease" or "Zinc Finger Nuclease (ZFN)" refers to a chimeric protein comprising a zinc finger DNA binding domain and a nuclease domain.

[0116] The term "TAL effector DNA nuclease fusion protein" or "TALEN" refers to a chimeric protein comprising a TAL effector DNA binding domain and a nuclease domain.

[0117] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide programmable nucleotide binding domain" and "nucleic acid programmable nucleotide binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or a guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the nucleic acid programmable nucleotide binding protein is an RNA-guided nucleic acid programmable nucleotide binding protein. In some embodiments, the RNA-guided nucleic acid programmable nucleotide binding protein is an RNA-guided nuclease.

[0118] In some embodiments, the RNA-guided nuclease is selected from type II CRISPR-Cas polypeptides, type I CRISPR-Cas polypeptides, type III CRISPR-Cas polypeptides, type IV CRISPR-Cas polypeptides, type V CRISPR-Cas polypeptides, type VI CRISPR-Cas polypeptides, type VII CRISPR-Cas polypeptides, IscB polypeptides, TnpB polypeptides, IsrB polypeptides. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can be associated with a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactivated Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cas13a (C2c2), Cas13b, Cas13c, Cas13d.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure. See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” (CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033); Yan et al., “Functionally diverse type V CRISPR-Cas systems” (Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271), the entire contents of which are incorporated herein by reference.

[0119] As used herein, "base editing activity" refers to the activity of chemically altering a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is adenosine or adenine deaminase activity, for example, converting a target A·T to C·G.

[0120] In some embodiments, base editing activity is assessed by editing efficiency. Base editing efficiency can be measured by any suitable means, for example, by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured by the percentage of total sequencing reads with nuclear base conversions affected by the base editor, for example, the percentage of total sequencing reads with target C·G base pairs converted to A·T base pairs. In some embodiments, when base editing is performed in a cell population, base editing efficiency is measured by the percentage of total cells with nuclear base conversions affected by the base editor.

[0121] "Guide polynucleotide," "guide RNA," or "gRNA" refers to a polynucleotide that can specifically target a target sequence and can form a complex with a nucleic acid programmable nucleotide binding domain protein (e.g., Cas9). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species includes two domains: (1) a domain that has homology to the target nucleic acid (e.g., guides the binding of the Cas9 complex to the target nucleic acid); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical to or homologous to the tracrRNA provided in Jinek et al, Science 337:816-821 (2012). In some embodiments, the gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA". The extended gRNA will bind to two or more Cas9 proteins and bind to the target nucleic acid at two or more different regions. The gRNA includes a nucleotide sequence complementary to the target site that mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity to the nuclease:RNA complex.

[0122] The term "identity" is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. "Identity" represents the percentage of the number of identical residues between the polypeptide or nucleic acid sequences to the total number of residues, and the calculation of the total number of residues is determined based on the mutation type. Mutation types include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence. Taking a polypeptide sequence as an example, if the mutation type is one or more of the following: substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence, the total number of residues is calculated based on the larger of the molecules being compared. If the mutation type also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., the number of insertions or deletions at both ends is less than 20) is not included in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps (if any) in the alignment are resolved by a specific algorithm. The same principle applies to the calculation of nucleotide identity.

[0123] The terms "sequence identity" and "sequence homology" are used interchangeably herein and, as used in conjunction with polynucleotides or polypeptides, refer to the percentage of bases or amino acids that are identical and in the same relative position when two sequences of polypeptides or polynucleotides are compared or aligned. Sequence identity can be determined in a variety of different ways. For example, sequences can be aligned using various methods and computer programs (e.g., BLAST, T-COFFEE, MUSCLE, MAFFT, etc.).

[0124] The term "DNA sequence or DNA polynucleotide sequence encoding a specific RNA is a sequence of DNA that can be transcribed into RNA. A DNA polynucleotide can encode an RNA (mRNA) that is translated into a protein, or a DNA polynucleotide can encode an RNA that is not translated into a protein (e.g., tRNA, rRNA, or guide RNA; also referred to as "non-coding" RNA or "ncRNA"). A DNA sequence or DNA polynucleotide sequence can also "encode" a specific polypeptide or protein sequence, wherein, for example, DNA directly encodes an mRNA that can be translated into a polypeptide or protein sequence. A "protein coding sequence" or a sequence encoding a specific protein or polypeptide is a nucleic acid sequence that can be transcribed into mRNA (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vitro or in vivo when placed under the control of an appropriate regulatory sequence. The boundaries of the coding sequence can be determined by a translation termination nonsense codon at the 5' end (N-terminus) and a 3' end (C-terminus). The coding sequence can include, but is not limited to, cDNA from prokaryotic or eukaryotic mRNA, genomic DNA sequences from prokaryotic or eukaryotic DNA, and synthetic nucleic acids. A transcription termination sequence will usually be located 3' to the coding sequence.

[0125] The term "promoter" or "promoter sequence" is a DNA regulatory sequence that is capable of promoting transcription of an operably linked coding or non-coding sequence (e.g., downstream (3' direction) coding or non-coding sequence) (e.g., capable of causing detectable levels of transcription and / or increasing detectable levels of transcription (relative to the level provided in the absence of the promoter)), such as by binding RNA polymerase. Various promoters, including inducible promoters and constitutive promoters, can be used to drive the vectors disclosed herein. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in the cell under most or all physiological conditions of the cell. Examples of promoters known in the art that can be used in certain embodiments (e.g., in the viral vectors disclosed herein) include CMV promoters, CBA promoters, smCBA promoters, and promoters derived from immunoglobulin genes, SV40, or other tissue-specific genes (e.g., RLBP1, RPE, VMD2). In addition, standard techniques for generating functional promoters by mixing and matching known regulatory elements are known in the art. Fragments of a promoter can also be used, such as those that retain at least a minimum number of bases or elements to initiate transcription at levels detectable above background.

[0126] "Operably linked" is the connection between genetic elements, meaning that the target nucleotide sequence is connected to the regulatory element in a manner that allows the expression of the nucleotide sequence (for example, in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). For example, a promoter needs to be located before the coding sequence of the gene it controls in order to initiate transcription of the gene. The promoter must be correctly placed in front of the coding sequence, and there needs to be an appropriate connection between them so that the promoter can effectively initiate transcription of the coding sequence. Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can also be selected to target specific types of cells.

[0127] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to thereby produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. The vector may also contain a replication initiation site. Vectors include plasmids and viral vectors. The plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Viral vector, wherein the DNA or RNA sequence derived from the virus is present in the vector for packaging the virus, including for example retrovirus, replication defective retrovirus, adenovirus, replication defective adenovirus and adeno-associated virus. Viral vector also comprises the polynucleotide carried by the virus for transfection into a host cell. Some vectors (for example, bacterial vectors and additional mammalian vectors with bacterial replication origin) can replicate autonomously in the host cell into which they are introduced. Other vectors (for example, non-additional mammalian vectors) are integrated into the genome of the host cell after introducing the host cell, and thus replicate together with the host genome. Moreover, some vectors can instruct the expression of the gene that they are operably connected. Such vectors are referred to as "expression vectors".

[0128] The term "wild type" has the meaning generally understood by those skilled in the art, which refers to the typical form of an organism, strain, gene, protein or the characteristics that distinguish it from mutant or variant forms when it exists in nature, which can be isolated from a source in nature and has not been intentionally modified by man.

[0129] The terms "variant," "derivative," and "analog" refer to polypeptides that substantially retain the function or activity of a protein. Generally, derivatization of a protein does not adversely affect the desired activity of the protein, i.e., the derivative of the protein has the same activity as the protein. Modified forms of "derivatives" include those in which one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted.

[0130] The terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort.

[0131] As used herein, a "functional fragment" or "active fragment" of a polynucleotide or polypeptide may refer to any subset of consecutive nucleotides or consecutive amino acids, respectively, that retains the original (e.g., wild-type) activity (or substantially similar activity) of the polynucleotide or polypeptide. In some embodiments, the "functional fragment" or "active fragment" comprises any portion or subsequence of the original (e.g., wild-type) or mutant polynucleotide or polypeptide. In some embodiments, the activity of the "functional fragment" or "active fragment" of the polynucleotide or polypeptide, for example, relative to the activity of the original (e.g., wild type), the "functional fragment" or "active fragment" can be about 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or less than 10% of the activity.

[0132] The nucleic acid cleavage of the present invention includes: DNA or RNA breakage in the target nucleic acid generated by the Cas protein (Cis cleavage), and DNA or RNA breakage in a side branch nucleic acid substrate (single-stranded nucleic acid substrate) caused by the Cas protein's side cutting activity (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0133] The terms “clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” or “CRISPR-Cas system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of the Cas genes.

[0134] The terms "target nucleic acid" and "target sequence" are used interchangeably and refer to a specific nucleic acid that comprises a nucleic acid sequence that is fully or partially complementary to the guide sequence in the gRNA. A "target sequence" refers to a polynucleotide targeted by the guide sequence in the gRNA, such as a sequence complementary to the guide sequence, wherein the hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including Cas protein and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter or terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts) of the cell. The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be related to a protospacer adjacent motif (PAM).

[0135] The detection method of the present invention can be used for quantitative detection of target nucleic acids. The quantitative detection index can be quantified based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group, or the width of the color band.

[0136] The term "regulatory element" includes promoters, enhancers, internal ribosome entry sites (IRES) and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences). In some cases, regulatory elements include those that direct the constitutive expression of a nucleotide sequence in many types of host cells and those that direct the expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a timing-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type-specific.

[0137] The term "host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism cultured as a unicellular entity (e.g., a cell line), which serves as a recipient of nucleic acid (e.g., an expression vector) and includes the descendants of the original cell that has been genetically modified by the nucleic acid.

[0138] It is understood that the progeny of a single cell may not necessarily have completely the same morphology, genome, etc. as the original parent cell due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also called a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.

[0139] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.

[0140] The term "NLS" refers to a "nuclear localization sequence" or "nuclear localization signal," which refers to an amino acid sequence that promotes protein entry into the cell nucleus. Nuclear localization sequences are known in the art (e.g., Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, and described in WO / 2001 / 038547, published on May 31, 2001), which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, such as described in Koblan et al., Nature Biotech. 2018 doi: 10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence: KRTADGSEFESPKKKRKV (SEQ ID NO:24), AVKRPAATKKAGQAKKKKLD (SEQ ID NO:25), KRPAATKKAGQAKKKK (SEQ ID NO:26), KKTELQTTNAENKTKKL (SEQ ID NO:27), KRGINDRNFWRGENGRKTR (SEQ ID NO:28), RKSGKIAAIVVKRPRK (SEQ ID NO:29), PKKKRKV (SEQ ID NO:30), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO:31).

[0141] The term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence using traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 out of 10 complementarity represents 50%, 60%, 70%, 80%, 90%, and 100% complementarity). "Perfect complementarity" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in the other nucleic acid sequence. "Substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.

[0142] Hybridization of the target sequence with the gRNA indicates that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complement each other and hybridize to form a complex.

[0143] The term "delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present invention can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein RNA. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.

[0144] The term "linker" refers to a linear polypeptide formed by connecting multiple amino acid residues through peptide bonds. The linker can be an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence.

[0145] The term "effective amount" or "therapeutically effective amount" refers to a dosage sufficient to achieve beneficial or desired results. A therapeutically effective amount may depend on the individual and disease condition being treated, the individual's weight and age, the severity of the disease condition, the mode of administration, etc., and can be readily determined by one skilled in the art.

[0146] The terms "treatment," "treating," and the like refer to obtaining a desired pharmacological and / or physiological effect, e.g., treating or curing a condition in a subject, delaying the onset of symptoms of a condition, and / or delaying the severity of a condition. The effect may be prophylactic, in terms of completely or partially preventing a disease or its symptoms, and / or therapeutic, in terms of partially or completely curing a disease and / or side effects attributable to the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal (e.g., a human), and includes: (a) preventing the occurrence of a disease in a subject who may be susceptible to the disease but has not yet been diagnosed with the disease; (b) inhibiting the disease, i.e., arresting its development; and (c) relieving the disease, i.e., causing the disease to regress.

[0147] The terms "individual," "subject," "host," and "patient" refer to individual organisms, including but not limited to various animals, plants, and microorganisms. Animals include mammals, including but not limited to bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), apes, non-human primates (e.g., macaques or cynomolgus monkeys), humans, mammalian farm animals, mammalian sports animals, and mammalian pets. In certain embodiments, the subject (e.g., a human) suffers from a disorder (e.g., a disorder caused by a disease-related gene defect). "Plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development.

[0148] For any embodiment of the present invention described herein, including an embodiment described only in the examples or claims or described only in one aspect / section below, it should be understood that unless expressly denied or the combination is inappropriate, the embodiment can be combined with any other one or more embodiments of the present invention.

[0149] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.

[0150] Where specific conditions are not specified in the examples, the experiments were carried out under conventional conditions or the conditions recommended by the manufacturer. Reagents or instruments used, for which the manufacturer is not specified, are commercially available conventional products. Those skilled in the art will appreciate that the examples describe the present invention by way of example and are not intended to limit the scope of the invention. All publications and other references mentioned herein are incorporated herein by reference in their entirety.

[0151] Example 1: Obtaining deaminase mutants

[0152] (1) Obtaining the deaminase mutant 5V3354

[0153] In order to construct an adenosine deaminase with higher editing efficiency and specificity, the amino acid site function prediction was performed on the amino acid sequence of the known adenosine deaminase 005V1 (deaminase 005V1 of CN114634923A, whose amino acid sequence is shown in SEQ ID NO: 1 and the nucleotide coding sequence is shown in SEQ ID NO: 2), and multiple positions that may improve editing efficiency and specificity were found. By site-directed PCR mutagenesis of the base editor 005V1-nCas9 expression vector containing the wild-type deaminase 005V1 (the amino acid sequence of the base editor 005V1-nCas9 is shown in SEQ ID NO: 3 and the nucleotide coding sequence is shown in SEQ ID NO: 4), an expression vector containing the adenosine deaminase 005V1 variant-nCas9 was constructed.

[0154] Base editors with different deaminase variants were generated by PCR-based site-directed mutagenesis. The specific method was to amplify the DNA sequence encoding the base editor 005V1-nCas9 (SEQ ID NO: 4) centered on multiple amino acids near the mutation site, and at the same time introduce the sequence to be mutated on the primer. Different mutant base editors were obtained by homologous recombination and connection of the amplified fragments (Table 1), and mutant 5V3354 was obtained.

[0155] Table 1 Mutation pattern of deaminase mutant 5V3354

[0156] The specific mutation methods are as follows:

[0157] The plasmid expressing the 005V1-nCas9 base editor was used as a template, and the plasmid of the 005V1-nCas9 base editor was amplified using amplification primers containing the mutant sequence using Vazyme's high-fidelity enzyme kit (Vazyme, P501-d2).

[0158] The amplification system is shown in Table 2:

[0159] Table 2 Plasmid mutation amplification system for 005V1-nCas9 base editor

[0160] The PCR amplification program is shown in the table below:

[0161] Table 3 Plasmid mutation amplification PCR program for 005V1-nCas9 base editor

[0162] The amplified PCR product was recovered and purified using a kit (Tiangen, Universal DNA Purification and Recovery Kit, DP214). The purified PCR product was transformed into Escherichia coli DH5a competent cells (Weidi Biotechnology, DL1001) and cultured. Single colonies were picked and, after sequencing confirmation, positive clones were shaken and plasmids were extracted using an endotoxin-free plasmid extraction kit (TIANGEN: DP120-01) and stored at -20°C until use.

[0163] (2) Verification of editing activity of deaminase mutant 5V3354

[0164] To test the editing activity of the deaminase mutant 5V3354, the PCSK9 target sequence PCSK9-sgRNA was designed for the PCSK9 gene: cccgcaccttggcgcagcgg (SEQ ID NO: 5).

[0165] The construction process of sgRNA expression vector (sgRNA plasmid) is as follows:

[0166] sgRNAs were designed and oligonucleotides (oligos) were synthesized based on the target sequence. The sgRNA targeting sequence used was shown in SEQ ID NO: 5. A CACC sequence was added to the 5' end of the upstream sequence of each sgRNA targeting sequence, and an AAAC sequence was added to the 5' end of the downstream sequence. Therefore, the upstream and downstream primer sequences used for synthesis were PCSK9-sgRNA-F (SEQ ID NO: 6) and PCSK9-sgRNA-R (SEQ ID NO: 7), respectively.

[0167] After synthesis, the upstream and downstream sequences were annealed using a preset PCR program (95°C, 5 min; 95°C–85°C at −2°C / s; 85°C–25°C at −0.1°C / s; maintained at 4°C), and the annealed products were ligated into the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) linearized with BbsI (NEB, R3539S).

[0168] Among them, the system used in the construction of sgRNA plasmid is as follows:

[0169] The linearization system of lenti U6-sgRNA / EF1a-mCherry vector is as follows: 3 μg of vector; 6 μL of buffer (NEB: R0539L); 2 μL of BbsI; ddH2O is added to 60 μL, and enzyme digestion is carried out at 37°C overnight.

[0170] The sgRNA annealing product was ligated with the linearized vector using the following system: 1 μL of T4 ligase buffer (NEB, M0202L), 20 ng of linearized vector, 5 μL of annealed oligo fragment (10 μM), 0.5 μL of T4 ligase (NEB: M0202L), and ddH2O was added to make up to 10 μL. The ligation was carried out at 16°C overnight.

[0171] The ligated vector was transformed into Escherichia coli DH5α competent cells (Weidi Biotech, DL1001). The specific process is as follows: Remove the DH5α competent cells from -80°C and quickly place them on ice. After 5 minutes, allow the bacterial mass to thaw. Add the ligated product and gently mix by flicking the bottom of the centrifuge tube. Let it rest on ice for 25 minutes. Heat shock the cells in a 42°C water bath for 45 seconds, quickly return them to ice, and let them rest for 2 minutes. Add 700 μL of sterile LB medium without antibiotics to the centrifuge tube, mix thoroughly, and then recover the cells at 37°C, 200 rpm, for 60 minutes. Centrifuge at 5000 rpm for one minute to harvest the cells. Collect approximately 100 μL of the supernatant, gently pipette to resuspend the bacterial mass, and spread the supernatant onto LB medium containing the Amp antibiotic. Incubate the plate upside down in a 37°C incubator overnight. Single colonies were picked, and after sequencing confirmation, the positive clones were shaken and the sgRNA plasmid was extracted using an endotoxin-free plasmid extraction kit (TIANGEN, DP120-01) and the concentration was measured. The plasmids were then stored in a -20°C refrigerator for later use.

[0172] HEK293T cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a 37°C cell culture incubator with 5% CO2. Cells for transfection were seeded in 24-well cell culture plates the day before and cultured. The cells were observed the next day and transfected when the cells grew to a cell density of approximately 80%. The amount of editor fusion protein plasmid transfected per well of the 24-well plate was 0.4 μg, and the amount of sgRNA plasmid transfected was 0.4 μg.

[0173] After mixing the plasmids, dilute them with 25 μL of reduced serum medium (Source Bio, L530KJ), add 2 μL of p3000 reagent, pipette and mix well as reagent A, and let it stand for 5 minutes. At the same time, dilute 2 μL of Lipofectamine 3000 transfection reagent (Thermo, 11668019) with 25 μL of reduced serum medium and mix well as reagent B, and let it stand for 5 minutes. Mix the above reagents A and B and pipette evenly, and let it stand for 20 minutes. After standing, add the mixed reagent dropwise to the 24-well plate cells to be transfected and return them to the 37°C incubator for culture. After 6 hours of transfection, replace the culture medium with DMEM medium containing 10% FBS. 48 hours after transfection, collect cells for editing efficiency detection.

[0174] The collected cells were subjected to genomic extraction (TIANGEN, DP304-03), and primers were designed according to experimental requirements. The identification primer sequences used were PCSK9-F (SEQ ID NO: 8) and PCSK9-R (SEQ ID NO: 9).

[0175] Using the genome as a template, PCR amplification was performed on the sequence near the target site. The system used for target site sequence amplification was as follows: 2×Taq Master Mix (Vazyme, P112-03) 25 μL; Primer-F (10 pmol / μL) 1 μL; Primer-R (10 pmol / μL) 1 μL; template 1 μL; ddH2O was added to make up to 50 μL.

[0176] The amplified PCR products were used for high-throughput deep sequencing (Genwizhi Biotechnology Co., Ltd.) or Sanger sequencing (Boshang Biotechnology (Shanghai) Co., Ltd.) to identify the editing efficiency.

[0177] Gene editing effect detection:

[0178] For the method of calculating gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun; 1(3): 239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262; PMCID: PMC6694769.

[0179] In the same way, the 005V1-nCas9 base editor plasmid and sgRNA plasmid were co-transfected into HEK293T cells, and their editing efficiency was calculated.

[0180] In this embodiment, the structure of the base editor comprising adenosine deaminase and nCas9 provided by the present invention is as follows:

[0181] NH2-[NLS]-[adenosine deaminase]-linker-[nCas9]-[NLS]-COOH, the example amino acid sequence of this structure can be found in SEQ ID NO: 3. However, it is only used for illustration and is not used to limit the structure of the base editor. This example statistics the editing efficiency of the 005V1-nCas9 and 5V3354-nCas9 base editors for the PCSK9 target (A6 site, the efficiency of mutation from adenine A to guanine G), among which 005V1-nCas9 has a base editing efficiency of 21% at this site, while the base editing efficiency of 5V3354-nCas9 at this site reaches 28% (Figure 1).

[0182] Example 2: Acquisition of other deaminase mutants

[0183] (1) To obtain a base editor with higher editing efficiency, further mutations were performed based on the deaminase 5V3354 (amino acid sequence shown in SEQ ID NO: 10, nucleotide coding sequence shown in SEQ ID NO: 11). In the same manner as in Example 1, PCR mutagenesis of the deaminase 5V3354-nCas9 expression vector was performed to obtain a variety of deaminase mutants (Table 4).

[0184] Table 4 Mutation patterns of various mutants of deaminase 5V3354

[0185] The following base editors were obtained: 5V17.1-nCas9, 5V17.2-nCas9, 5V17.3-nCas9, 5V17.4-nCas9, 5V17.5-nCas9, 5V17.6-nCas9, 5V17.7-nCas9, 5V17.8-nCas9, 5V17.9-nCas9, 5V18.1-nCas9, 5V18.2-nCas9, 5V19.1-nCas9, 5V19.2-nCas9, 5V20.1-nCas9, 5V20.2-nCas9, 5V21.1-nCas9, 5V21.2-nCas9, 5V22.1-nCas9, 5V22.2-nCas9.

[0186] (2) The PCSK9-sgRNA expression plasmid was constructed in the same manner as in Example 1. The expression plasmids of each mutant base editor obtained were co-transfected with the PCSK9-sgRNA expression plasmid into HEK293T cells, and the editing efficiency was detected ( Figure 2 ).

[0187] By comparison, it can be seen that compared with the base editor composed of wild-type 5V3354 (editing efficiency 28%), the base editors composed of mutants 5V17.1, 5V17.2, 5V17.5, 5V17.7, 5V18.1, 5V18.2, 5V19.1, 5V19.2, 5V20.2, 5V21.1, 5V21.2, 5V22.1, and 5V22.2 have significantly improved editing efficiency in the PCSK9 target sequence, which are 37%, 39%, 36%, 32%, 29%, 33%, 34%, 33%, 34%, 36%, 29%, 31% and 35%, respectively; especially 5V17.2-nCas9, the base editing efficiency is close to 40%.

[0188] In addition, the editing efficiency of the base editor composed of 5V17.3 reached 26%, and the editing efficiency of 5V17.9 reached 27%, both of which were significantly better than 005V1; the editing efficiency of 5V20.1 reached 21%, which is comparable to the editing efficiency of 005V1.

[0189] Example 3: Determination of editing efficiency of chimeric base editors for some mutants

[0190] To further investigate the activity of these mutants, chimeric recombinant base editors were designed for the deaminase mutants 5V17.2, 5V22.1, and 5V22.2, embedding the deaminase between sites 1249 and 1250 of nCas9. These base editors were named 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, and 5V22.2-1249-nCas9, respectively. The structures of these base editors are as follows: NH2-[NLS]-[nCas9 N-terminal fragment]-[adenosine deaminase]-[nCas9 C-terminal fragment]-[NLS]-COOH.

[0191] The experimental steps are as follows:

[0192] 1. Construction of nCas9 Plasmid

[0193] (1) Design nCas9 cloning primers (synthesized by Shanghai Boshang Biotechnology Co., Ltd.), the upstream and downstream primers are nCas9-F (SEQ ID NO: 12) and nCas9-R (SEQ ID NO: 13), respectively.

[0194] ABE8e (Addgene, #138489) was amplified by PCR using a high-fidelity enzyme kit (Vazyme, P501-d2) from Novagen. The amplification system is shown in Table 5:

[0195] Table 5 ABE8e (Addgene, #138489) PCR amplification system

[0196] The PCR amplification program is shown in the table below:

[0197] Table 6 ABE8e (Addgene, #138489) PCR amplification procedure

[0198] The amplified PCR product was recovered according to the kit instructions (Tiangen, Universal DNA Purification and Recovery Kit, DP214). The purified PCR product was transformed into Escherichia coli DH5a competent cells (Weidi Biotechnology, DL1001).

[0199] The specific process is as follows:

[0200] Remove the DH5α competent cells from -80°C and quickly place them on ice. After 5 minutes, allow the bacterial mass to thaw. Add the ligation product and gently mix by hand by tapping the bottom of the centrifuge tube. Let it rest on ice for 25 minutes. Heat shock the tube in a 42°C water bath for 45 seconds, quickly return it to ice, and let it rest for 2 minutes. Add 700 μL of sterile LB medium without antibiotics to the centrifuge tube, mix thoroughly, and then recover at 37°C, 200 rpm for 60 minutes. Harvest the cells by centrifugation at 5000 rpm for one minute. Retain approximately 100 μL of the supernatant, gently pipette to resuspend the bacterial mass, and spread it onto LB medium supplemented with Amp antibiotics. Incubate the plate upside down at 37°C in a humidified incubator overnight. Single colonies were selected. After sequencing confirmation, positive clones were shaken and the nCas9 plasmid was extracted using an endotoxin-free plasmid extraction kit (TIANGEN: DP120-01). The concentration was then determined and stored in a -20°C refrigerator until needed.

[0201] (2) Obtaining DNA sequences encoding deaminases 5V17.2, 5V22.1, and 5V22.2

[0202] PCR amplification was performed on the 5V17.2-nCas9, 5V22.1-nCas9, and 5V22.2-nCas9 plasmids using primers. The amplification system and PCR procedures were the same as in Tables 5 and 6. The amplified PCR products were recovered according to the kit instructions (Tian Gen, Universal DNA Purification and Recovery Kit, DP214). PCR products of adenosine deaminase 5V17.2, 5V22.1, and 5V22.2 were obtained. The upstream and downstream PCR primers used were ADA-F (SEQ ID NO: 14) and ADA-R (SEQ ID NO: 15).

[0203] (3) Design of different chimeric base editors and primer sequences corresponding to nCas9 plasmids

[0204] The primer sequences corresponding to nCas9 were designed according to the insertion position of the deaminase, as shown below:

[0205] nCas9-1249-F:ggggcagcagcggggggtcacccgaggataatgagcagaaacagctgt (SEQ ID NO: 16)

[0206] nCas9-1249-R:CCGCCGCTAGATCCTCCAGAggagcccttcagcttctcatagtggct (SEQ ID NO: 17)

[0207] The nCas9 plasmid obtained in step (1) was amplified using the above primer sequences. The amplification system and PCR program were the same as in Tables 5 and 6. The amplified PCR products were recovered according to the kit instructions (Tian Gen, Universal DNA Purification and Recovery Kit, DP214). The amplified nCas9 PCR products were homologously recombined with the PCR products of adenosine deaminase 5V17.2, 5V22.1, and 5V22.2 obtained in step (2) to construct chimeric base editors 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, and 5V22.2-1249-nCas9 (Table 7). The kit used for homologous recombination was the Gibson Assembly Master Mix Recombination Kit (NEB, E2611S).

[0208] Table 7 Chimeric base editors 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, 5V22.2-1249-nCas9

[0209] (4) Amplification and sequencing of chimeric base editors 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, and 5V22.2-1249-nCas9.

[0210] The homologous recombination product was transformed into Escherichia coli DH5α competent cells (Weidi Biotechnology, DL1001). The specific process is as follows: Remove the DH5α competent cells from -80°C and quickly place them on ice. After 5 minutes, allow the bacterial mass to thaw. Add the ligation product and gently mix by flicking the bottom of the centrifuge tube. Let it rest on ice for 25 minutes. Heat shock the cells in a 42°C water bath for 45 seconds, quickly return them to ice, and let them rest for 2 minutes. Add 700 μL of sterile LB medium without antibiotics to the centrifuge tube, mix thoroughly, and then recover at 37°C, 200 rpm, for 60 minutes. Centrifuge at 5000 rpm for one minute to harvest the cells. Collect approximately 100 μL of the supernatant, gently pipette to resuspend the bacterial mass, and spread the supernatant onto a plate of LB medium containing the Amp antibiotic. Incubate the plate upside down in a 37°C incubator overnight. Single colonies were picked, and after sequencing confirmation, the positive clones were shaken and the chimeric recombinant plasmid was extracted using an endotoxin-free plasmid extraction kit (TIANGEN, DP120-01) and the concentration was measured. The plasmid was then stored in a -20°C refrigerator for later use.

[0211] (5) Cell culture and transfection

[0212] HEK293T cells (purchased from ATCC) were inoculated in DMEM medium supplemented with 10% FBS (v / v) (Gibco, 11965092), containing 1% Penicillin Streptomycin (v / v) (Gibco, 15140122), and cultured in a 37°C cell culture incubator containing 5% CO2. The cells used for transfection were inoculated in a 24-well cell culture plate the day before and cultured. The cells were observed the next day and transfected when the cells grew to a cell density of about 80%. The amount of plasmid transfected in each well of the 24-well plate was 0.4 μg of the chimeric recombinant plasmid and 0.4 μg of the sgRNA plasmid. The same PCSK9-sgRNA as in Example 2 was used for activity testing.

[0213] After mixing the chimeric recombinant plasmid and sgRNA plasmid, dilute with 25 μL of reduced serum medium (Source Bio, L530KJ) and add 2 μL of p3000 reagent. Mix by pipetting as reagent A and let it stand for 5 minutes. At the same time, dilute 2 μL of Lipofectamine 3000 transfection reagent (Thermo, 11668019) with 25 μL of reduced serum medium and mix well as reagent B. Let it stand for 5 minutes. Mix the above reagents A and B and pipette evenly. Let it stand for 20 minutes. After the standing period, add the mixed reagent dropwise to the 24-well plate cells to be transfected and return to the 37°C incubator for culture. After 6 hours of transfection, the culture medium was replaced with DMEM medium containing 10% FBS. 48 hours after transfection, the cells were collected for testing of editing efficiency.

[0214] (6) Editing efficiency detection

[0215] HEK293T cells were subjected to genomic extraction using a genomic DNA extraction kit (TIANGEN, DP304-03). The primer sequence identification and editing efficiency determination methods were the same as in Example 2. The editing efficiency of the chimeric base editor was also compared with that of the base editor before modification. The base editing efficiencies of 5V17.2-1249-nCas9, 5V22.1-1249-nCas9, and 5V22.2-1249-nCas9 at the PCSK9 gene target site (A6 site) were 44%, 46%, and 56%, respectively. Compared with the base editing efficiency before modification (39%, 31%, and 35%, respectively), the base editing efficiency was greatly improved (Figure 3).

[0216] Example 4: Targeted Editing of α1-Antitrypsin Deficiency and Hemoglobinopathy Gene Targets Using Base Editors

[0217] HEK293T cells harboring the E342K mutation in the A1AT (alpha-1 antitrypsin) gene were used for testing. The 5V17.9 variant was selected to replace the deaminase region of the plasmid pCMV-SpRY-ABE8e (Addgene, 185671) for detection. The A1AT target sequence was designed based on the target: ATCGACAAGAAAGGGACTGA (SEQ ID NO: 42).

[0218] The A1AT-sgRNA plasmid was constructed using the method of Example 1 and co-transfected with the base editor 5V17.9-nCas9 (the experimental method was the same as in Example 3) into the aforementioned mutated HEK293T cells. PCR was performed using primers (A1AT-F: (SEQ ID NO: 43), A1AT-R: (SEQ ID NO: 44)) to detect the editing efficiency. After detection and analysis, it was found that the base editor 5V17.9-nCas9 had a significant editing efficiency at the target position, with an editing efficiency of 51% for this site (AG).

[0219] HEK293T cells harboring the E6V mutation in the β-globin gene were used for testing. The 5V17.9 variant was selected to replace the deaminase region of the plasmid pCMV-SpRY-ABE8e (Addgene, 185671). The β-globin target sequence was designed based on the target: ACTTCTCCACAGGAGTCAGA (SEQ ID NO: 45).

[0220] The β-globin-sgRNA plasmid was constructed using the method of Example 1 and co-transfected with the base editor 5V17.9-nCas9 (the same transfection method as in Example 3) into the aforementioned mutated HEK293T cells. The editing efficiency was detected after PCR using primers (β-globin-F: (SEQ ID NO: 46), β-globin-R: (SEQ ID NO: 47)). After detection and analysis, it was found that the base editor 5V17.9-nCas9 had a significant editing efficiency at the target position, and the editing efficiency of this site (AG) was 35%.

[0221] Therefore, the base editor of the present invention can be applied to the treatment of α-trypsin deficiency and hemoglobinopathy.

[0222] Example 5: Using base editors to treat base mutation diseases

[0223] The base editors provided in the present disclosure (e.g., 5V17.2-nCas9) can be used to modify pathogenic DNA targets. For example, base editors can be used to direct mutations from adenine (A) in nucleic acids (e.g., DNA) to guanine (G). Such changes can alter the amino acid sequence of a protein to destroy or create a new start codon, or create a stop codon, to disrupt splice donors, to disrupt splice acceptors, or to edit regulatory sequences, thereby correcting pathogenic genes and achieving therapeutic purposes.

[0224] The disease is obtained from the NCBI ClinVar database available on the NCBI ClinVar website, for example, can be selected from the base editing disease targets shown in Table 34 in WO2022056254A2 published on March 17, 2022.

[0225] The partial sequences used in the present invention are shown in Table 8 below.

[0226] Table 8 Partial sequences used in the present invention

[0227] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.

Claims

1. A deaminase comprising the following sequence: (i) the amino acid sequence shown in SEQ ID NO: 10; or (ii) an amino acid sequence that has at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% sequence identity to the amino acid sequence of SEQ ID NO: 10, and which retains the deamination activity of the deaminase of the amino acid sequence as shown in SEQ ID NO: 10; and is not SEQ ID NO:

1.

2. The deaminase according to claim 1, characterized in that The amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% sequence identity with the amino acid sequence shown in SEQ ID NO: 10 is an amino acid sequence obtained by adding, substituting, deleting or inserting one or more amino acid residues into the amino acid sequence shown in SEQ ID NO: 10; Preferably, the substitution is a substitution occurring at one or more of the following positions of the amino acid sequence as shown in SEQ ID NO: 10: C46, ​​Y47, G48, H49, C144, Q145, F146, Y147, Q148, Q149, P150, R151, E152, V153, F154, N155, A156, E157, R158, E159, A160, R161, R162, L163, N164, Q165, P166, D167, R168, A169 and D170; More preferably, the substitution is a combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10: (1) 7 or 10 of C46, ​​Y47, G48, H49, Q148, P150, E152, V153, F154, and N155; (2) 6, 7, or 8 of Q148, Q149, P150, R151, E152, V153, F154, and N155; (3) 7 or 8 of A156, E157, R158, E159, A160, R161, R162, and L163; (4) 6 or 7 of N164, Q165, P166, D167, R168, A169, and D170; (5) 6 or 7 of C144, Q145, F146, Y147, Q148, Q149, and P150; or (6) 2 or 9 of C144, Q145, Q148, Q149, P150, E152, V153, F154, and N155.

3. The deaminase according to claim 2, characterized in that The substitution is a substitution occurring at any combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10: (1) C46, ​​Y47, G48, H49, Q148, P150, E152, V153, F154, and N155; (2) G48, Q148, P150, E152, V153, F154, and N155; (3) Q148, Q149, P150, E152, V153, F154 and N155; (4) Q148, P150, E152, V153, F154 and N155; (5) Q149, P150, R151, E152, V153, F154 and N155; (6) Q148, Q149, P150, R151, E152, V153, F154 and N155; (7) Q148, Q149, P150, E152, F154 and N155; (8) A156, E157, R158, E159, A160, R161, R162 and L163; (9) A156, E157, R158, E159, A160, R162 and L163; (10) N164, Q165, P166, D167, R168, A169 and D170; (11) N164, Q165, D167, R168, A169 and D170; (12) C144, Q145, F146, Y147, Q148, Q149 and P150; (13) C144, Q145, F146, Y147, Q148 and P150; (14) C144, Q145, Q148, Q149, P150, E152, V153, F154 and N155; (15) C144 and Q145; Preferably, the substitution occurring at the site is selected from: C46P, Y47I, G48A / T, H49R, C144T / L / W, Q145L / K, F146A / R, Y147S / F, Q148R / G / T / S / C, Q149N / P / R / V / G / F / C / K, P150A / L / R / I / G / S / T, R151K / P, E152P / L / Q / H / S, V153T / A / F / Y / K / P, F154S / P / V / N / L / D / H, N155P / G / T / S / Y / R / A, A156L / T, E157F, R158N / L , E159L / H, A160K / T, R161K, R162L / K, L163D / I, N164G / R, Q165T / L, P166Q, D167L, R168L / P, A169N / T and D170R / H; preferably selected from C46P, Y47I, G48A / T, H49R, C144T / L / W, Q145L / K, F146A / R, Y147S / F, Q148R / T / C, Q149N / P / R / G / C / K, P150A / L / R / G / S / T, R151P, E152P / L / Q / H, V153T / A / F / K, F154S / P / V / L / D / H, N155P / G / T / S / A, A156L / T, E157F, R158N / L, E159L / H, A160K / T, R 161K, R162L / K, L163D / I, N164G / R, Q165T / L, P166Q, D167L, R168L / P, A169N / T and D170R / H; More preferably, the substitution is a combination of the following positions in the amino acid sequence as shown in SEQ ID NO: 10: (1)Q148R, Q149N, P150A, E152P, V153T, F154S, N155P; (2)Q148R, P150L, ​​E152L, V153A, F154P, N155G; (3)Q148R, Q149R, P150R, E152P, V153F, F154V, N155T; (4) Q148G, Q149G, P150I, E152L, V153Y, F154N, N155S; (5)Q149P, P150G, R151K, E152Q, V153K, F154L, N155P; (6)Q148R, Q149V, P150S, R151P, E152L, V153F, F154P, N155Y; (7) Q148T, Q149G, P150R, E152H, V153A, F154D, N155S; (8)Q148S, Q149F, P150L, ​​E152S, V153P, F154L, N155R; (9)Q148G, Q149C, P150S, E152P, F154H, N155A; (10)A156L, E157F, R158N, E159L, A160K, R161K, R162L, L163D; (11)A156T, E157F, R158L, E159H, A160T, R162K, L163I; (12)N164G, Q165T, P166Q, D167L, R168L, A169N, D170R; (13)N164R, Q165L, D167L, R168P, A169T, D170H; (14)C144T, Q145L, F146A, Y147S, Q148R, Q149K, P150S; (15)C144L, Q145L, F146R, Y147F, Q148C, P150T; (16)C144W, Q145K, Q148R, Q149R, P150R, E152P, V153F, F154V, N155T; (17) C144W, Q145K; (18)G48A, Q148R, P150L, ​​E152L, V153A, F154P, N155G; (19)C46P, Y47I, G48T, H49R, Q148R, P150L, ​​E152L, V153A, F154P, N155G.

4. A base editor fusion protein comprising the deaminase according to any one of claims 1 to 3, and a nucleic acid programmable nucleotide binding domain; Preferably, the nucleic acid programmable nucleotide binding domain is a Cas protein or an AGO protein; the Cas protein is, for example, selected from Cas9, CasX, CasY, Cpf1, C2c1, C2c2 and C2c3; and / or, the AGO protein is, for example, selected from pAgo, eAgo, Ago1, Ago2, Ago3 and Ago4; and / or, the deaminase is connected to one end of the nucleic acid programmable nucleotide binding domain or is embedded in the nucleic acid programmable nucleotide binding domain; More preferably, the connection is a direct connection or a connection through a linker; and / or the chimeric site is located in the carboxyl terminal domain of the nucleic acid programmable nucleotide binding domain; and / or the nucleic acid programmable nucleotide binding domain retains part or no nucleotide chain cleavage activity; the linker comprises an amino acid sequence preferably as shown in one or more of SEQ ID NOs:32-41.

5. The base editor fusion protein according to claim 4, characterized in that The base editor fusion protein further comprises a nuclear localization signal sequence; the nuclear localization signal sequence is connected to the N-terminus and / or C-terminus of the base editor fusion protein, and / or, to the N-terminus and / or C-terminus of the deaminase; Preferably, the nuclear localization signal sequence is connected to the N-terminus and C-terminus of the base editor fusion protein; for example, the structure of the base editor fusion protein from N-terminus to C-terminus is: nuclear localization signal sequence-deaminase-nucleic acid programmable nucleotide binding domain-nuclear localization signal sequence; or nuclear localization signal sequence-deaminase-nucleic acid programmable nucleotide binding domain-nuclear localization signal sequence.

6. The base editor fusion protein according to claim 4 or 5, characterized in that When the nucleic acid programmable nucleotide binding domain is a Cas protein, such as a Cas9 protein, the chimeric site is located between positions 1249-1250 of Cas9; Preferably, the base editor fusion protein comprises an amino acid sequence as shown in any one of SEQ ID NO: 10, 18, 20 and 22.

7. A base editing system, comprising: (i) a deaminase or a nucleic acid programmable nucleotide binding domain as claimed in any one of claims 1 to 3; or (ii) the base editor fusion protein according to any one of claims 4 to 6, and a guide polynucleotide; The nucleic acid programmable nucleotide binding domain or the base editor fusion protein forms a ribonucleoprotein complex with the guide polynucleotide and binds to the target nucleic acid under the guidance of the guide polynucleotide.

8. A polynucleotide encoding the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, or the base editing system according to claim 7; Preferably, the polynucleotide encoding the base editor fusion protein comprises a nucleotide sequence as shown in any one of SEQ ID NO:11, 19, 21 and 23.

9. A vector comprising the polynucleotide according to claim 8; Preferably, the polynucleotide is located on one or more vectors; and / or, the vector further comprises a promoter, and the polynucleotide is operably linked to the promoter; More preferably, the promoter is selected from one or more of a constitutive promoter, an inducible promoter, a ubiquitin promoter, a cell type-specific promoter and a tissue-specific promoter.

10. An isolated cell comprising the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, the polynucleotide according to claim 8 and / or the vector according to claim 9; Preferably, the cell is a prokaryotic cell or a eukaryotic cell; for example, selected from animal cells, plant cells and fungal cells; More preferably, the cell is a vertebrate cell or an invertebrate cell; the vertebrate cell is preferably a mammalian cell; Further preferably, the mammalian cell is selected from rodent cells, primate cells and non-primate cells; the primate cells are, for example, human cells.

11. A pharmaceutical composition comprising the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, the base editing system according to claim 7, the polynucleotide according to claim 8, the vector according to claim 9 and / or the cell according to claim 10, and optionally a pharmaceutically acceptable carrier and / or excipient.

12. A kit comprising the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, the base editing system according to claim 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10 and / or the pharmaceutical composition according to claim 11.

13. A delivery system comprising the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, the base editing system according to claim 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10, the pharmaceutical composition according to claim 11 and / or the kit according to claim 12, and a delivery medium; Preferably, the delivery medium is selected from liposomes, nanoparticles, viral vectors, exosomes, microvesicles and cell-penetrating peptides.

14. A base editing method, comprising the step of contacting the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, or the base editing system according to claim 7 with a target nucleic acid and causing a deamination reaction; Preferably, the base editing method is in vivo or in vitro; and / or, the base editing method is for non-diagnostic or non-therapeutic purposes.

15. Use of the deaminase according to any one of claims 1 to 3, the base editor fusion protein according to any one of claims 4 to 6, the base editing system according to claim 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10, the pharmaceutical composition according to claim 11, the kit according to claim 12 or the delivery system according to claim 13 in the preparation of a medicament for treating a disease associated with or caused by a point mutation; Preferably, the disease is selected from one or more of hypercholesterolemia, transthyretin amyloidosis, α1-antitrypsin deficiency and β-hemoglobinopathies.