Adenosine deaminase, base editors and applications

JP2025512996A5Pending Publication Date: 2026-03-02ヨルテック セラピューティクス カンパニー リミテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024559441
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-07
Filing Date
2023-02-24
Publication Date
2026-03-02

AI Technical Summary

Technical Problem

Current base editors have site-dependent differences in editing efficiency and wide editing windows that lead to non-essential editing, necessitating improvements for more accurate and efficient genome modification.

Method used

Site-directed mutations to the basal sequence of deaminase and optimization of the structure of base editor fusion proteins, particularly identifying the preferred chimeric position of adenosine deaminase in the nuclease domain, to enhance editing efficiency in eukaryotic cells.

Benefits of technology

The improved base editors exhibit increased editing efficiency in eukaryotic cells, with potential for better clinical applications and treatment of diseases associated with point mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of biotechnology, and provides a deaminase and an adenine base editor, and also provides a mutant deaminase and a corresponding adenine base editor, which has multiple amino acid mutations compared to the parent deaminase, and therefore has improved base editing efficiency and is expected to be applicable in the future.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of biotechnology. More specifically, the present invention relates to adenosine deaminase, base editor fusion proteins, base editor systems and uses. [Background technology]

[0002] How to modify genomes precisely and efficiently is an important goal in life science research, and gene editing technology mediated by CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas9 is the most powerful tool to achieve this goal. Conventional CRISPR / Cas9 technology generates DNA double strand breaks (DSBs) in the target, induces homologous recombination (HR) and non-homologous end joining (NHEJ) in cells to repair the pathway, and further realizes modifications such as site-specific knockout, substitution, and insertion of genomic DNA. However, DNA repair by DSBs is difficult to achieve efficient and stable single-base mutations. Single-base mutations cause the occurrence of approximately two-thirds of human genetic diseases and are also the genetic basis of important trait mutations in many animals and plants. For this reason, the development of technology that can achieve single-base substitutions precisely and efficiently is particularly important. The David R. Liu Laboratory's base editors have been developed for this purpose. The David R. Liu Laboratory has developed three different base editors, the Cytosine Base Editor (CBE), the Adenine Base Editor (ABE), and the Prime Editor. These base editors do not depend on the occurrence of DSBs when operating, and do not require the involvement of donor DNA.

[0003] Adenosine deaminase-based adenine base editing technology mainly utilizes the nicking enzyme Cas9n (D10A) or dCas9 combined with adenosine deaminase to form a fusion protein, which guides sgRNA to deaminate the target base adenine A located within the base editing active window to form hypoxanthine I, which is then gradually replaced with G through DNA repair and copying, ultimately forming a directional substitution of A to G (A → G).

[0004] Base editors can also be used to treat some diseases (e.g., hypercholesterolemia, transthyretin amyloidosis, β-hemoglobinopathy) by editing gene targets. Therefore, optimization of conventional base editors is expected to have a better future in clinical applications.

[0005] However, current base editors have site-dependent differences in editing efficiency and wide editing windows that lead to non-essential editing, making improvements to conventional base editors necessary. Summary of the Invention

[0006] The present invention discloses a deaminase, a base editor containing the deaminase, and its use. Through experimentation and exploration, site-specific mutations are performed on the basic sequence of the deaminase to obtain an adenine base editor with improved editing efficiency, and the editing efficiency in eukaryotic cells is improved. In addition, a number of experiments are performed on the structure of the base editor fusion protein to find a suitable chimeric position of adenosine deaminase in the nuclease domain, and a base editor fusion protein structure with improved editing efficiency is obtained, which is expected to have better prospects. [Definition] As is well known to those skilled in the art, a protein can undergo amino acid alterations (eg, substitutions, deletions, or additions) and the resulting protein can retain its function or activity.

[0007] The above-mentioned "substitution" refers to the replacement of an amino acid residue at a certain position in an amino acid sequence with another amino acid residue. Among these, the "substitution" may be a conservative amino acid substitution.

[0008] A "conservative modification," "conservative substitution," or "conservative substitution" is the replacement of amino acids in a protein with other amino acids having similar characteristics (e.g., charge, side chain size, hydrophobicity / hydrophilicity, main-chain conformation and rigidity, etc.), allowing frequent changes without altering the biological activity of the protein.

[0009] As those skilled in the art know, generally, single amino acid substitutions in non-essential regions of a polypeptide result in little or no change in biological activity (see Watson et al. (1987), Molecular Biology of the Gene, The Benjamin / Cummings Pub. Co., p. 224, (4th ed.)). Also, substitution of amino acids with similar structures and functions are less likely to destroy biological activity. Exemplary conservative substitutions are described in "Exemplary Amino Acid Conservative Substitutions." Exemplary Conservative Amino Acid Substitutions [Table 1] Furthermore, amino acids with similar characteristics are shown below. [Table 2] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art. As used herein, the following terms have the meanings ascribed to them below, unless otherwise specified.

[0010] "Polynucleotide" and "nucleic acid" are polymeric forms of nucleotides of any length (e.g., RNA or DNA). The terms include, but are not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers that contain purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.

[0011] A "DNA sequence encoding a particular RNA" is a DNA nucleotide sequence that is transcribed into RNA.

[0012] A "regulatory element," "regulatory element," or "DNA regulatory sequence," or "control element" refers to transcriptional and translational control sequences, such as promoters, enhancers, terminators, and the like.

[0013] A "promoter (also called promoter sequence)" is a DNA regulatory region capable of binding RNA polymerase and initiating transcription of a downstream (3' direction) coding or non-coding sequence. In general, various promoters, including inducible promoters, are used to drive expression of the various vectors of the invention.

[0014] "Codon optimization" refers to modifying a nucleic acid sequence so that said nucleic acid sequence is better expressed in a host cell. This is generally done by replacing at least one codon (e.g., it may be one, or multiple codons, e.g., 10, 20 or more) in the original nucleic acid sequence with a codon that is more frequently or most frequently used in the genes of said host cell, while maintaining the natural amino acid sequence that is expressed.

[0015] "Naturally-occurring (also called unmodified / unmodified / wild type (wt)")" is a nucleic acid, polypeptide, cell, or organism that occurs in nature. For example, a polypeptide or polynucleotide sequence that is present in an organism and that can be isolated from a natural source is naturally-occurring.

[0016] The terms "polypeptide", "peptide" and "protein" as used in this disclosure are used interchangeably herein and refer to amino acid polymers of any length. The polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. The terms also include amino acid polymers that have been modified (e.g., disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling moiety). A "fusion protein" is a heteropolypeptide that contains protein domains derived from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion or the carboxyl-terminal (C-terminal) portion of the fusion protein, forming an amino-terminal fusion protein or a carboxyl-terminal fusion protein, respectively.

[0017] A "CRISPR-Cas system" (also referred to as a "CRISPR system" or "CRISPR / Cas system") typically comprises transcription products or other elements associated with the expression of CRISPR-associated ("Cas") genes, or capable of directing the activity of said Cas genes. In some embodiments, the components of a CRISPR system may include nucleic acids (e.g., vectors) encoding one or more components of the system, the components in protein form, or a combination thereof.

[0018] "Cas protein" refers to a CRISPR-associated protein (Cas) (also referred to as a "CRISPR-associated protein", "CRISPR effector", "effector", "Cas protein", "Cas enzyme" or "CRISPR enzyme"), which is a protein that performs an enzymatic activity and / or binds to a target site on a nucleic acid specified by an RNA guide. In some embodiments, a Cas protein has endonuclease activity, nicking enzyme activity, exonuclease activity, transposase activity, and / or cleavage activity. In some other examples, the Cas protein may be nuclease inactive or partially inactive.

[0019] A "guide RNA (also referred to as leader RNA, gRNA or crRNA)" is any RNA molecule that is advantageous in targeting a Cas protein to a target nucleic acid (e.g., DNA and / or RNA). A "guide RNA" includes one or more guide RNAs and their equivalents known to those of skill in the art, including but not limited to RNA-based molecules that can form a complex with a Cas protein (e.g., tandem repeat (DR) sequences), and includes sequences (e.g., spacer sequences) that are sufficiently complementary to hybridize with a target nucleic acid sequence and guide the sequence-specific binding of the complex to the target nucleic acid sequence.

[0020] The term "crRNA" as used in this disclosure includes repeat and spacer sequences. CRISPR transcription forms a long pre-CRISPR RNA (pre-crRNA), and pre-crRNA processing yields a short crRNA that contains repeat and spacer sequences. In some CRISPR / Cas systems, the crRNA is derived by the action of Cas protein on the pre-crRNA. In other CRISPR / Cas systems, the crRNA is derived by the joint action of Cas protein and tracrRNA (trans-activating crRNA) on the pre-crRNA.

[0021] In different CRISPR / Cas systems, the crRNA can act alone as a guide RNA (gRNA) to guide the Cas protein to localize to a target sequence located near the PAM sequence, or can be combined with the tracrRNA to form a single guide RNA (sgRNA) to guide the Cas protein to localize to a target sequence located near the PAM sequence.

[0022] As used in the present disclosure, a "guide sequence of crRNA" is a sequence in the crRNA that hybridizes with a target sequence of a target nucleic acid, and is correspondingly formed by the spacer sequence of the crRNA.

[0023] The term "target sequence" as used in the present disclosure is a nucleotide sequence that is complementary or at least partially complementary to the crRNA in a target nucleic acid. After the Cas protein, the crRNA form a ternary complex with the target sequence, the Cas protein exerts specific cleavage activity on the target nucleic acid strand and / or non-nucleotide strand in the target nucleic acid. In the present disclosure, "target sequence" may be used interchangeably with "target nucleic acid", "target polynucleotide", "target sequence", and "target nucleic acid sequence".

[0024] As used in this disclosure, the term "target strand" refers to a nucleotide strand in a target nucleic acid that hybridizes with crRNA, and the term "non-target strand" refers to a nucleotide strand in a target nucleic acid that does not hybridize with crRNA.

[0025] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes a Cas9 protein or a fragment thereof (e.g., a protein that includes an active, inactive or partially active DNA cleavage domain of Cas9, and / or a binding domain of gRNA Cas9). Cas9 nuclease may also be referred to as CRISPR-associated nuclease 9. As described above, CRISPR is an adaptive immune system that can provide defense against mobile genetic elements (viruses, trans-elements and conjugative plasmids). CRISPR clusters contain multiple short, conserved repeat regions and spacer regions. CRISPR clusters are transcribed and processed into pre-crRNA. In type II CRISPR / cas9 systems, the correct processing of pre-crRNA requires a transcoding small RNA (tracrRNA), endogenous ribonuclease 3 (RNase III) and Cas9 protein. As ribonuclease 3, tracrRNA assists in processing the guide of pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the spacer sequence. Target strands that are not complementary to the crRNA are endonucleolytically cleaved. In nature, DNA binding and cleavage usually requires a protein and two types of RNA. However, single guide RNA (sgRNA) can be engineered to incorporate aspects of the crRNA and tracrRNA into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes one short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.The sequence and structure of Cas9 nuclease are well known to those of skill in the art (see, "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E. et al., Nature 471:602-607 (2011); and "Aprogrammable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M. et al., Science 337:816-821 (2012)). Orthologues of Cas9 have already been described in each species, including but not limited to S. pyogenes and S. thermophilus. Other suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0026] Nuclease-inactive Cas9 proteins may be interchangeably referred to as "dCas9" proteins (nuclease "dead" Cas9 or Cas9 without nuclease activity) or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28, 152(5):1173-83). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains, the HNH nuclease subdomain and the RuvC subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC subdomain cleaves the strand that is not complementary. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28:152(5):1173-83 (2013)). Based on the knowledge in the art, other suitable nuclease-free Cas9 domains will be obvious to the skilled artisan. Other exemplary suitable nuclease-free Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., Nature Biotechnology. 2013, 31(9):833-838).

[0027] Cas9 nicking enzyme can cleave one strand of double-stranded DNA. Cas9 nicking enzyme can be generated by introducing an inactive mutation into the HNH or RuvC subdomain. For example, an inactive mutation (D10A) can be introduced into the RuvC domain of S. pyogenes Cas9 while retaining the activity of the HNH domain, i.e., retaining the 840th residue as histidine. Such Cas9 variants can generate single-stranded DNA breaks (nicking) at specific positions based on the target sequence specified by the gRNA. One skilled in the art can identify catalytic residues in the RuvC and HNH domains of any known Cas9 protein and introduce inactive mutations to generate the corresponding dCas9 or nCas9.

[0028] Similarly, for other Cas proteins, those skilled in the art can use the same method to obtain corresponding Cas proteins without nuclease activity and nicking enzymes that cleave one strand of double-stranded DNA.

[0029] The term "deaminase" or "deaminase domain" as used in this disclosure is a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA).

[0030] The term "nucleic acid programmable nucleotide binding domain" and "nucleic acid programmable DNA-binding protein (napDNAbp)" as used in this disclosure is a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide polynucleotide (e.g., gRNA), which guides the napDNAbp to a specific nucleic acid sequence, e.g., by hybridizing to a target nucleic acid sequence. For example, a Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nicking enzyme (nCas9), or a Cas9 without nuclease activity (dCas9). Examples of nucleic acid programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute proteins (AGO). However, it should be understood that the nucleic acid programmable DNA binding protein also includes the nucleic acid programmable protein that binds to RNA. For example, napDNAbp can bind to the nucleic acid that guides napDNAbp to RNA. Although not specifically described in this disclosure, other nucleic acid programmable DNA binding proteins are also within the scope of this disclosure.

[0031] The nCas9 domain comprises nCas9 or a fragment thereof, wherein the nCas9 fragment has a certain identity to nCas9 (e.g., at least about 70% identity, or at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity) and retains its basic function.

[0032] A "base editor (BE)" or "nucleobase editor" as used in this disclosure is a reagent that binds to a polynucleotide and has nucleobase modifying activity. In embodiments, the base editor comprises a nucleobase modifying polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide binding domain (e.g., a nucleic acid programmable DNA binding protein) that binds to a guide polynucleotide (e.g., a guide RNA). In embodiments, the reagent is a biomolecular complex that comprises a protein domain with base editing activity, i.e., that can modify a base (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA, RNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the reagent is a fusion protein that comprises a domain with base editing activity. In some embodiments, the domain with base editing activity can deaminate a base in a nucleic acid molecule. In some embodiments, the base editor can deaminate one or more bases in a DNA molecule. In some embodiments, the base editor is an adenine base editor (ABE).

[0033] As used herein, a "base editing activity" is one for chemically modifying a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is an adenosine or adenine deaminase activity, e.g., to convert a target A·T to C·G.

[0034] In some embodiments, base editing activity is assessed by editing efficiency. Base editing efficiency can be measured by any suitable method, such as sanger sequencing or next generation sequencing. In some embodiments, base editing efficiency is measured by the percentage of total sequencing reads that have a nucleobase conversion affected by a base editor, e.g., the percentage of total sequencing reads that have a target C·G base pair that is converted to an A·T base pair. In some embodiments, when base editing is performed in a cell population, base editing efficiency is measured by the percentage of total cells that have a nucleobase conversion affected by a base editor.

[0035] The term "base editor system" as used in this disclosure refers to a system for editing a nucleobase of a target nucleotide sequence. In various embodiments, the base editor system includes: (1) a nucleic acid programmable nucleotide binding domain (e.g., Cas9); (2) a deaminase domain (e.g., adenosine deaminase) for deaminating the nucleobase; and (3) one or more guide polynucleotides (e.g., guide RNAs).

[0036] A "guide polynucleotide", "guide RNA" or "gRNA" is a polynucleotide that can specifically target a target sequence and can form a complex with a nucleic acid programmable nucleotide binding domain protein (e.g., Cas9). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), although "gRNA" can be used interchangeably to refer to a guide RNA that exists as a single molecule or a complex of two or more molecules.

[0037] Typically, gRNA, which exists as a single RNA species, contains two domains: (1) a domain that has identity to the target nucleic acid (e.g., a domain that guides binding of the Cas9 complex to the target nucleic acid) and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence referred to as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012). Other examples of gRNAs may be those disclosed in U.S. Provisional Patent Application USSN 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof," and U.S. Provisional Patent Application USSN 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases." In some embodiments, the gRNA comprises two or more of domains (1) and (2) and can be referred to as an "extended gRNA." The extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions. The gRNA comprises a nucleotide sequence complementary to the target site that mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease / RNA complex.

[0038] According to the present invention, the three-letter and one-letter codes of amino acids used are as described in J. Biol. Chem, 243, p. 3558 (1968).

[0039] A "vector" is a polynucleotide composition for transferring, delivering or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phage vectors, viral vectors (e.g., retroviral vectors, adeno-associated viral vectors, herpes simplex viral vectors, AAV vectors, lentiviral vectors, baculoviral vectors), and the like.

[0040] "Delivery system" includes a delivery vector including one or more of a liposome, a nanoparticle, an exosome, an exosome, a microvesicle, a viral vector, a gene gun or an electroporation device, and the like.

[0041] A "functional fragment" is a protein or polypeptide sequence that contains fewer amino acids compared to the original sequence of a protein or polypeptide, but the remaining amino acid sequence retains a certain percentage (e.g., 10%, 20%, 30%, 40%, 50% or 60-99%, 100%) of the functional activity of the original reference sequence (e.g., a protein or polypeptide can be modified by substitution, insertion, deletion and / or addition of one or more amino acids while retaining a certain percentage of enzymatic activity).

[0042] "Identity" refers to matching sequences between two polypeptides or two nucleic acids. If a position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., if a position in each of the two DNA molecules is occupied by adenine, or if a position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions common to the two sequences, divided by the number of positions being compared, multiplied by 100. For example, if 6 of 10 positions in two sequences are matched, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (3 of a total of 6 positions are matched). Generally, the comparison is performed by aligning the two sequences to produce the maximum identity.

[0043] "Host cell" includes in vitro, ex vivo or in vivo cells or cell lines, or their progeny, including but not limited to CHO, BHK, 293, 293T cell lines, etc. Said cells or cell lines, or their progeny, contain the Cas13 protein, fusion protein, CRISPR-Cas system, polynucleotide, vector or delivery system described in the present invention.

[0044] "Operably linked" means that the nucleotide sequence of interest is linked to a regulatory sequence in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a target cell when the vector is introduced into the target cell).

[0045] The terms "coding sequence" or "protein coding sequence," which may be used interchangeably herein, are a polynucleotide fragment that codes for a protein. The region or sequence has a start codon near the 5' end and a stop codon near the 3' end. A coding sequence may also be referred to as an open reading frame.

[0046] The term "nuclear localization sequence", "nuclear localization signal (NLS)" refers to an amino acid sequence that promotes the import of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International PCT application to Plank et al. (PCT / EP2000 / 011690, filed November 23, 2000, published May 31, 2001 as WO / 2001 / 038547), the contents of which are incorporated herein by reference for disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech.2018doi:10.1038 / nbt.4172.

[0047] The term "linker" as used herein refers to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that links two molecules or moieties (e.g., two components of a protein complex or a ribonucleic acid complex, e.g., two domains of a fusion protein, such as a polynucleotide programmable DNA binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase)). A linker can link different components of a base editor system or different parts of a component. For example, in some embodiments, a linker can link a guide polynucleotide binding domain of a polynucleotide programmable nucleotide binding domain and a catalytic domain of a deaminase. A linker can be located between or on either side of two groups, molecules, or other moieties and can link them together by linking them to each other through a covalent bond or a non-covalent interaction. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker.

[0048] In some embodiments, the linker may be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker may be about 5-100 amino acids long, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids long. In some embodiments, the linker may be about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids long. Longer or shorter linkers are also contemplated.

[0049] The term "cleavage" as used in this disclosure refers to the cleavage of a phosphodiester bond in a nucleotide chain. The type of cleavage may be a single-stranded break or a double-stranded break.

[0050] The terms "complementary" or "hybridizing" as used in this disclosure refer to "polynucleotide" and "oligonucleotide" in relation to base pairing rules (these are interchangeable terms referring to nucleotide sequences). For example, the sequence "CAGT" and the sequence "GTCA" are complementary. Complementary can be "partial" or "full". "Partial" complementary means that one or more nucleic acid bases are mismatched according to the base pairing rules. "Full" or "complete" complementary between nucleic acids means that the base pairing of each nucleic acid base with another base all matches the base pairing rules. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions and detection methods based on binding between nucleic acids.

[0051] As used herein, the term "hybridize" refers to the pairing of complementary nucleic acids by any process in which a strand of nucleic acid joins with a complementary strand through base pairing to form a hybridized complex.

[0052] The terms "nucleic acid sequence" and "nucleotide sequence" as used in this disclosure refer to oligonucleotides or polynucleotides and fragments or portions thereof, which may be DNA or RNA of single- or double-stranded, genomic or synthetic origin, and which refer to the sense or antisense strand.

[0053] The terms "sequence identity" and "percent identity" as used in this disclosure refer to the percentage of nucleotides or amino acids that are the same (i.e., identical) between two or more polynucleotides or polypeptides. Sequence identity between two or more polynucleotides or polypeptides can be measured by aligning the nucleotide or amino acid sequences of the polynucleotides or polypeptides, scoring the number of positions in the aligned polynucleotides or polypeptides that contain the same nucleotide or amino acid residue, and comparing it to the number of positions in the aligned polynucleotides or polypeptides that contain different nucleotides or amino acid residues. Polynucleotides can differ at a position, for example, by containing a different nucleotide (i.e., substitution or mutation) or by deleting a nucleotide (i.e., nucleotide insertion or deletion in one or two polynucleotides). Polypeptides can differ at a position, for example, by containing a different amino acid (i.e., substitution or mutation) or by deleting an amino acid (i.e., amino acid insertion or deletion in one or two polypeptides). Sequence identity can be calculated by dividing the number of positions that contain the same nucleotide or amino acid residue by the total number of amino acid residues in the polynucleotides or polypeptides. By way of example, the percent identity can be calculated by dividing the number of positions containing the same nucleotide or amino acid residue by the total number of nucleotides or amino acid residues in the polynucleotide or polypeptide and multiplying by 100.

[0054] Illustratively, when compared and aligned for maximum correspondence using a sequence alignment algorithm or by visual inspection, two or more sequences or subsequences have at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% nucleotide "sequence identity" or "percent identity." In some embodiments, the sequences are substantially identical over the entire length of any one or two of the compared biopolymers (e.g., polynucleotides).

[0055] The term "vector" refers to a means for introducing a nucleic acid sequence into a cell to produce a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence that is expressed in a recipient cell. An expression vector may contain additional nucleic acid sequences, such as, for example, initiators, terminators, enhancers, promoters and secretory sequences, to enhance and / or promote expression of the introduced sequence.

[0056] As used in this disclosure, the terms "individual" and "subject" may be used interchangeably and are mammals, including, but not limited to, domestic animals (e.g., cows, sheep, cats, dogs, and horses), primates (e.g., humans and non-human primates such as monkeys), rabbits, and rodents (e.g., mice and rats). In particular, an individual is a human.

[0057] The methods disclosed herein can be performed in vitro, ex vivo, or in vivo, or the products can be present in vitro, ex vivo, or in vivo. The term "in vitro" refers to experiments using materials, biological materials, cells, and / or tissues in laboratory conditions or culture media, and the term "in vivo" refers to experiments and processes using intact multicellular organisms. In some embodiments, methods performed in vivo can be performed in a non-human animal. "Ex vivo" refers to events that exist or occur outside an organism, e.g., outside the human or animal body, e.g., in tissues (e.g., whole organs) or cells removed from an organism.

[0058] The term "pharmaceutical acceptable carrier" as used in this disclosure refers to a pharmaceutically acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc powder, magnesium, calcium or zinc stearate or stearic acid) or solvent encapsulating material, involved in the transport or transportation of a compound from one site of the body (e.g., a delivery site) to another site (e.g., an organ, tissue or body part). A pharmaceutically acceptable carrier is "acceptable" in that it is compatible with the other ingredients of the formulation and is non-toxic to the subject's tissues (e.g., physiological compatibility, sterility, physiological pH, etc.). Examples of materials that can be used as pharma- ceutically acceptable carriers include: (1) sugars, such as lactose, glucose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium dodecyl sulfate, and talc powder; (8) excipients, such as cocoa butter and suppository wax; (9) peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and the like; (10) oils, such as oil and soybean oil; (11) polyhydric alcohols, such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffers, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethanol; (20) pH buffers; (21) polyesters, polycarbonates, and / or polyanhydrides; (22) bulking agents, such as polypeptides and amino acids; (23) serum components, such as serum albumin, high density lipoprotein (HDL), and low density lipoprotein (LDL); (22) C2-C12 alcohols, such as ethanol; and (23) other non-toxic compatible substances used in pharmaceutical formulations.Wetting agents, coloring agents, releasing agents, coating agents, sweeteners, flavorings, fragrances, preservatives and antioxidants may also be present in the formulation. The terms "excipient", "pharmaceutical acceptable carrier" and the like may be used interchangeably herein.

[0059] As used herein, the term "effective amount" refers to an amount of a bioactive agent sufficient to cause a desired biological response. For example, in some embodiments, an effective amount of a base editor is an amount of the base editor sufficient to induce mutations at a target site of specific binding by the base editor. As one of skill in the art can appreciate, the effective amount of a reagent, such as a base editor fusion protein, deaminase, polynucleotide, etc., can vary depending on a variety of factors, including the desired biological response, such as the particular allele, genome, or target site to be edited, the cell or tissue to be targeted, and the reagent used.

[0060] The term "treatment" refers to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. The term "treatment" refers to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have already formed and / or after a disease has already been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms, or to inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual before symptoms develop (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may continue to be administered after symptoms have subsided, for example, to prevent or delay recurrence.

[0061] Specifically, the present invention includes the following.

[0062] <Adenosine deaminase> In some embodiments according to the first aspect of the invention, there is provided an adenosine deaminase comprising one or more of the following sequences:

[0063] an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to the amino acid sequence shown in SEQ ID NO:1 or SEQ ID NO:170, wherein the amino acid sequence retains the deamination activity of the amino acid sequence shown in SEQ ID NO:1; An amino acid sequence having one or more amino acid residues added, substituted, deleted or inserted in the amino acid sequence shown in SEQ ID NO:1 or SEQ ID NO:170, which retains the deamination activity of the amino acid sequence shown in SEQ ID NO:1 or SEQ ID NO:170, or An amino acid sequence encoded by a nucleotide sequence that hybridizes under stringent conditions to a polynucleotide sequence encoding the amino acid sequence shown in SEQ ID NO:1 or SEQ ID NO:170, and which retains the deamination activity of the amino acid sequence shown in SEQ ID NO:1 or SEQ ID NO:170 (the stringent conditions are medium stringency conditions, medium to high stringency conditions, high stringency conditions, or very high stringency conditions).

[0064] The present invention also relates to adenosine deaminase variants having modifications at any two or more of the following amino acid positions relative to the amino acid sequence shown in SEQ ID NO:1: 33, 35, 36, 46, 47, 48, 49, 104, 105, 107, 148, 149, 150, 151, 152, 153, 154, and 155.

[0065] Further, the adenosine deaminase variant may have any of the following amino acid sequences with respect to the amino acid sequence shown in SEQ ID NO:1: V33I, D35G, D35R, D36N, D36G, A46C, I47Y, I47R, I47V, T48G, T48R, L49T, L49H, L49K, V104A, V104M, S105C, S105G, S107R, S107K, S107A, Q148P, Q148C, Q148A, Q148G, The amino acid sequence includes a modification at at least one amino acid position selected from Q149M, Q149G, Q149L, P150R, P150L, ​​P150C, R151K, E152R, E152T, E152G, E152W, V153P, V153F, V153I, V153T, F154H, F154K, F154L, N155T, N155R, and N155H. The modifications at specific positions may be one or a combination of multiple modifications.

[0066] Specifically, the adenosine deaminase variants have the following amino acid sequences with respect to the amino acid sequence shown in SEQ ID NO:1: (1) Q148G+Q149M+P150R, (2) S107A, (3) E152R+V153P+F154H+N155T, (4) S107K, (5) Q148P+Q149G+P150L, ​​(6) A46C+I47Y+T48G+L49H, (7) E152T+V153F+N155T, (8) S107R, (9) V104M, (10) E152G+V153I+F154K+N15 The amino acid sequence includes any one of the following alterations: 5R, (11) S105C, (12) Q148C+Q149L+P150R+R151K, (13) Q148A+Q149G+P150C, (14) E152W+V153T+F154L+N155H, (15) S105G, (16) I47R+T48G+L49T, (17) D35G+D36N, (18) V33I+D35R+D36G, (19) V104A, or (20) I47V+T48R+L49K.

[0067] The present invention provides adenosine deaminase variants having modifications at any of a number of amino acid positions relative to the amino acid sequence set forth in SEQ ID NO: 170. In some embodiments, the substitutions are at any of the following amino acid positions in the amino acid sequence set forth in SEQ ID NO: 170: S15, D16, H17, E18, F19, N20, D21, E22, Y23, W24, M25, R26, H27, A28, L29, T30, K33, R34, A35, R36, V41, V43, L47, L49, N51, N59, A61, I62, L64, A69, E72, G80, L81, V82, L83, Q84, N85, Y86, I89, D9 and substitutions occurring at one or more of the following positions: 0, A91, T92, V95, F97, A106, R111, I112, S113, R114, L115, F117, V119, R120, N121, S122, K123, R124, N132, V133, L134, N135, P137, G138, M139, N140, H141, R142, E144, D160, V168, F169, N170.

[0068] Further, in some embodiments, the substitutions are selected from the group consisting of S15T, S15G, D16E, D16N, H17C, H17K, E18D, E18K, F19C, F19Y, F19S, N20Q, N20L, N20S, D21Q, E22D, Y23F, W24F, M25L, M25V, R26K, R26T, H27R, A28C, L29I, T30E, K33R, K33S, R34K, A35S, R36Q, L49F, N51G, L64M, L64Q, L64P, G80A, L81N, V82A, V82T, L83I, L83Q in the amino acid sequence set forth in SEQ ID NO:170. , Q84N, N85S, Y86W, I89E, I89L, D90G, A91C, A91T, T92D, A106C, I112L, S113K, R114K, V119L, S122N, S122P, R124H, R124T, N132K, V133I, L134F, N135H, N135S, P137F, G138A, M139L, N140K, H141A, R142L, R142S, E144H, D160E, V168A, F169C, F169V, N170D.

[0069] In some specific embodiments, the substitution is in the amino acid sequence set forth in SEQ ID NO: 170: (1) S15T+D16E+H17K+F19Y+N20Q (adenosine deaminase 004V2), (2) E22D+Y23F+W24F+R26K+H27R+L29I (adenosine deaminase 004V3), (3) K33R+R34K+A35S (adenosine deaminase 004V4), (4) G80A+L81N+V82A+L83I+Q84N+N85S+Y86W (adenosine deaminase 004V7), (5) I89L+D90G+A91T+T92D (adenosine deaminase 004V8), (6) I112L+S113K+R114K (adenosine deaminase 004V10), (7) N132K+V133I+L134F+N135H (adenosine deaminase 004V12), (8) P137F+G138A+M139L (adenosine deaminase 004V13), (9) E18K+F19S+N20L (adenosine deaminase 004V14), (10) M25L (adenosine deaminase 004V15), (11) S15G+D16N+H17C (adenosine deaminase 004V16), (12) L64Q+S122P+V168A (adenosine deaminase 004V17), (13) R36Q+L64Q+S122P+F169V (adenosine deaminase 004V18), (14) R36Q+V119L+V168A (adenosine deaminase 004V19), (15) R36Q+V119L+N170D (adenosine deaminase 004V20), (16) L64Q+V119L+V168A (adenosine deaminase 004V21), (17) F169C (adenosine deaminase 004V22), (18) L64Q+V119L+V168A (adenosine deaminase 004V23), (19) L64P + R124T + F169V (adenosine deaminase 004V24), (20) I89E+A91C (adenosine deaminase 004V25), (21) L64Q+V119L+F169V (adenosine deaminase 004V26), (22) L64M (adenosine deaminase 004V27), (23) L64Q+S122P+N170D (adenosine deaminase 004V28), (24) E18D+F19C+N20S (adenosine deaminase 004V29), (25) L64P+V119L+N170D (adenosine deaminase 004V30), (26) S122N + R124H + N135S + D160E (adenosine deaminase 004V31), (27) R26T + H27R + A28C (adenosine deaminase 004V32), (28) V82T+L83Q (adenosine deaminase 004V33), (29) M25V (adenosine deaminase 004V34), (30) D21Q+T30E+K33S (adenosine deaminase 004V35), (31) N140K+H141A+R142S (adenosine deaminase 004V36), (32) L49F+N51G (adenosine deaminase 004V37), (33) A106C (adenosine deaminase 004V38), (34) R142L+E144H (adenosine deaminase 004V39), (35) E22R+Y23F+W24P+M25H (adenosine deaminase 004V40), and / or (36) L64Q+R124T+V168A (adenosine deaminase 004V41), This is a substitution that occurs in a combination of the following positions:

[0070] In the present invention, "retaining deamination activity", for example, "retaining the deamination activity of the amino acid sequence shown in SEQ ID NO:1" or "retaining the deamination activity of the amino acid sequence shown in SEQ ID NO:170" may mean completely retaining the deamination activity of adenosine deaminase of the original sequence, or may mean partially retaining the deamination activity of adenosine deaminase of the original sequence. For example, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the deamination activity may be retained. In some other embodiments, adenosine deaminase having a modified sequence, for example, adenosine deaminase having a sequence with amino acid substitution, may have a higher deamination activity than adenosine deaminase of the original sequence.

[0071] The adenosine deaminase provided by the present invention can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In some embodiments, the adenosine deaminase can deaminate the target adenine (A) of a polynucleotide that includes DNA. In some embodiments, the adenosine deaminase can deaminate the target adenine (A) of a polynucleotide that includes RNA.

[0072] In some embodiments, the substitution is a conservative substitution.

[0073] In the present invention, "moderate stringency conditions", "moderate to high stringency conditions", "high stringency conditions" or "very high stringency conditions" describe nucleic acid hybridization and washing conditions. For guidelines on performing hybridization, see Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1-6.3.6, the contents of which are incorporated herein by reference. This document describes both hydrated and non-hydrated methods, both of which can be used. For example, specific hybridization conditions are as follows: (1) Low stringency hybridization conditions are in 6x sodium chloride / sodium citrate (SSC) at about 45°C, followed by two washes at least 50°C with 0.2x SSC, 0.1% SDS (for low stringency conditions, the wash temperature may be increased to 55°C). (2) Moderate stringency hybridization conditions are 6×SSC at about 45° C., followed by one or more washes at 60° C. with 0.2×SSC, 0.1% SDS. (3) High stringency hybridization conditions are 6×SSC at about 45° C., followed by one or more washes at 65° C. with 0.2×SSC, 0.1% SDS, and are preferred. (4) Ultra-high stringency hybridization conditions are 0.5 M sodium phosphate, 7% SDS at 65° C., followed by one or more washes at 65° C. with 0.2×SSC, 1% SDS.

[0074] The amino acid sequence of adenosine deaminase 005V1 is as follows:

[0075] MSELNDAYWMKQALALAQKAREQGEVPVGAILVLDDEVIGQGWNRAITLHDPTAHAEIMALQQGGQIVQNYRLLNATLYVTFEPCVMCAGAMVHSRIKRLVYGVSNSKRGAAGSLLNVLNYPGMNHQIEITAGVMANECSEMLCQFYQQPREVFNAEREARRLNQPDRAD(SEQ ID NO.1) The amino acid sequence of adenosine deaminase 004V1 is as follows:

[0076] MYNAPRFFCRSAAVSDHEFNDEYWMRHALTLAKRAREEGEVPVGAVLVLNNQVIGEGWNRAIGLHDPTAHEIMALRQGGLVLQNYRLIDATLYVTFEPCVMCAGAMVHSRISRLVFGVRNSKRGAAGSLINVLNYPGMNHRVEITEGILAESCSAMLCDFYRWPREVFNALKKARQEEG(SEQ ID NO:170) The amino acid sequence of adenosine deaminase 005V1 is as follows:

[0077] ATGAGTGAGCTGAATGATGCTTACTGGATGAAACAGGCACTCGCTTTAGCTCAGAAGGCCCGGGAACAGGGAGAAGTTCCAGTGGGCGCTATTCTGGTGCTGGATGATGAAGTGATAGGACAGGGATG GAATAGAGCCATCACCCTGCACGACCCCCCGCCCACGCCGAGATCATGGCCCTGCAGCAGGGCGGCCAGATCGTGCAGAACTACCGGCTGCTGAACGCCACCCTGTACGTGACCTTCGAGCCCTGCGT GATGTGCGCCGGCGCCATGGTGCACAGCAGGATCAAGAGACTGGTCTACGGCGTGAGCAACTCAAAAAGAGGCGCCGCCGGCAGCCTGCTGAACGTGCTGAACTACCCCGGCATGAACCACCAGATCG AGATCACCGCCGGCGTGATGGCCAACGAGTGCAGCGAGATGCTGTGCCAGTTCTATCAGCAGCTAGGGAAGTGTTCAATGCTGAGCGTGAGGCTAGGCGGCTGAACCAACCTGATAGAGCTGAC(SEQ ID NO.2)

[0078] The amino acid sequence of adenosine deaminase 004V1 is as follows: (SEQ ID NO.156)

[0079] <Base editor fusion protein> In a second aspect, the present invention provides a method for producing a composition comprising the steps of: There is provided a base editor fusion protein comprising an adenosine deaminase according to the first aspect of the invention and a nucleic acid programmable nucleotide binding domain. In the present invention, the nucleic acid programmable nucleotide binding domain, when bound to a guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence (i.e., via complementary base pairing sequences between the bases of the bound guide nucleic acid and the target polynucleotide) and localize the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. It is understood that the nucleic acid programmable nucleotide binding domain can also comprise a nucleic acid programmable protein that binds RNA.

[0080] In some embodiments of the invention, the nucleic acid programmable nucleotide binding domain in the base editor is a Cas protein or an AGO protein. The Cas protein or AGO protein includes naturally occurring Cas proteins or AGO proteins, and their homologs or modified or engineered forms. For example, in some embodiments, the Cas protein or AGO protein may be a protein comprising an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the amino acid sequence of a naturally occurring Cas protein or AGO protein. In some other embodiments, the Cas protein also includes a form of the protein that lacks its nicking enzyme or nuclease activity.

[0081] In some embodiments, non-limiting examples of Cas proteins that can be used as a nucleic acid programmable nucleotide binding domain include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f (C2c10 / Cas14), Cas12g, Cas12h, Cas12i, Cas12j, Cas12k / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f (C2c10 / Cas14), Cas12g, Cas12h, Cas12i, Cas12j, Cas12k / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f (C2c10 / Cas14), Cas12k / C2c3, ... c5, Cas12l, Cas12m, Cas12n, Cas13a(C2c2), cas13b, Cas13c, Cas13d, Csy1, Csy2, Csy3, Csy4 , Css1, Css2, Cse5e, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr2, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Class II Cas effector, Type V Cas effector, Class VI Cas protein, CARF, DinG, homologs thereof, or modified or engineered forms thereof. Other nucleic acid programmable nucleotide binding domains, which may not be specifically mentioned in this disclosure, are also within the scope of this disclosure.

[0082] In some specific embodiments of the invention, the Cas protein is selected from the Cas9 family, the Cas12 family, and the Cas13 family, including, but not limited to, Cas9, Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f (C2c10 / Cas14), Cas12g, Cas12h, Cas12i, Cas12j, Cas12k (C2c5), Cas12l, Cas12m, Cas12n, Cas13a (C2c2), Cas13b, Cas13c, Cas13d, homologs thereof, or modified or engineered forms thereof. In some specific embodiments of the invention, the Cas proteins include nuclease-free forms of the above Cas proteins, such as dCas9, dCas12a, dCas12b, dCas12c, dCas12d, dCas12e, dCas12f, dCas12g, dCas12h, dCas12i, dCas12j, dCas12k, dCas12l, dCas12m, dCas12n, dCas13a, dcas13b, dCas13c, and dCas13d. For example, in some specific embodiments of the invention, the Cas proteins also include nicking enzyme forms of the above proteins, such as, but not limited to, nCas9.

[0083] In some preferred embodiments of the invention, the nucleic acid programmable nucleotide binding domain is Cas9. In some specific embodiments of the invention, the Cas9 is Cas9 from Streptococcus pyogenes (SpCas9), Cas9 from Staphylococcus aureus (SaCas9), or Cas9 from Streptococcus thermophilus 1 (St1Cas9). In some preferred embodiments of the invention, the Cas9 is Cas9 from Streptococcus pyogenes (SpCas9).

[0084] In some more preferred embodiments of the invention, the Cas9 may be a nuclease-active Cas9, a Cas9 nicking enzyme (nCas9) or a Cas9 without nuclease activity (dCas9).

[0085] In some further preferred embodiments of the invention, the nucleic acid programmable nucleotide binding domain is a Cas9 nicking enzyme (nCas9). In some further preferred embodiments of the invention, the nucleic acid programmable nucleotide binding domain comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the amino acid sequence of a Cas9 nicking enzyme (nCas9) provided herein.

[0086] In some embodiments of the present invention, adenosine deaminase is directly fused / linked to a nucleic acid programmable nucleotide binding domain to form a fusion protein, or fused / linked via a linker to form a fusion protein. The order of fusion / linkage of adenosine deaminase and nucleic acid programmable nucleotide is not particularly limited, for example, adenosine deaminase may be at the N-terminus of the base editor, or the nucleic acid programmable nucleotide binding domain may be at the N-terminus of the base editor.

[0087] In embodiments in which the adenosine deaminase and the nucleic acid programmable nucleotide binding domain are directly fused, an exemplary base editor fusion protein has the following structure:

[0088] NH2-[adenosine deaminase]-[napDNAbp]-COOH, NH2-[napDNAbp]-[adenosine deaminase]-COOH, or NH2-[N-terminal fragment of napDNAbp]-[adenosine deaminase]-[C-terminal fragment of napDNAbp]-COOH.

[0089] In embodiments in which the adenosine deaminase and the nucleic acid programmable nucleotide binding domain are fused via a linker, an exemplary base editor fusion protein has the following structure:

[0090] NH2-[adenosine deaminase]-[optional linker]-[napDNAbp]-COOH, NH2-[napDNAbp]-[optional linker]-[adenosine deaminase]-COOH, or NH2-[N-terminal fragment of napDNAbp]-[optional linker]-[adenosine deaminase]-[optional linker]-[C-terminal fragment of napDNAbp]-COOH. In some embodiments, the nucleobase editor domain is fused via a linker comprising the amino acid sequence: [Table 3]

[0091] In some embodiments of the invention, the base editor comprises at least one nuclear localization signal sequence (NLS sequence). The NLS sequence may be selected from the amino acid sequences in the table below. [Table 4]

[0092] In some embodiments of the invention, the nuclear localization signal sequence may be at the N-terminus, C-terminus, both termini of the base editor, or may be between the adenosine deaminase and the nucleic acid programmable nucleotide binding domain. In some embodiments of the invention, the nuclear localization signal sequence may be fused directly to the base editor or fused to the base editor via a linker.

[0093] Exemplary structures of base editors that contain a nuclear localization signal sequence are as follows: NH2-[NLS]-[adenosine deaminase]-[napDNAbp]-COOH, NH2-[adenosine deaminase]-[NLS]-[napDNAbp]-COOH, NH2-[adenosine deaminase]-[napDNAbp]-[NLS]-COOH, NH2-[NLS]-[napDNAbp]-[adenosine deaminase]-COOH, NH2-[napDNAbp]-[NLS]-[adenosine deaminase]-COOH, NH2-[napDNAbp]-[adenosine deaminase]-[NLS]-COOH, NH2-[NLS]-[adenosine deaminase]-[optional linker]-[napDNAbp]-COOH, NH2-[adenosine deaminase]-[optional linker]-[NLS]-[optional linker]-[napDNAbp]-COOH, NH2-[adenosine deaminase]-[optional linker]-[napDNAbp]-[NLS]-COOH, NH2-[NLS]-[napDNAbp]-[optional linker]-[adenosine deaminase]-COOH, NH2-[napDNAbp]-[any linker]-[NLS]-[any linker]-[adenosine deaminase]-COOH, NH2-[napDNAbp]-[optional linker]-[adenosine deaminase]-[NLS]-COOH, NH2-[NLS]-[adenosine deaminase]-[optional linker]-[napDNAbp]-[NLS]-COOH.

[0094] In some preferred embodiments of the invention, a base editor has the following structure:

[0095] NH2-[NLS]-[adenosine deaminase]-[optional linker]-[napDNAbp]-[NLS]-COOH, or NH2-[NLS]-[N-terminal fragment of napDNAbp]-[optional linker]-[adenosine deaminase]-[optional linker]-[C-terminal fragment of napDNAbp]-[NLS]-COOH.

[0096] In some further preferred embodiments of the invention, the base editor comprises one or more of the following sequences: (i) the amino acid sequence shown in SEQ ID NO:190 or SEQ ID NO:3; (ii) an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3, wherein the amino acid sequence retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; (iii) an amino acid sequence in which one or more amino acid residues are added, substituted, deleted, or inserted in the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3, and which retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; or (iv) An amino acid sequence encoded by a nucleotide sequence that hybridizes under stringent conditions to a polynucleotide sequence encoding the amino acid sequence shown in SEQ ID NO:190 or SEQ ID NO:3, and which retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO:190 or SEQ ID NO:3 (the stringent conditions are moderate stringency conditions, medium to high stringency conditions, high stringency conditions, or ultra-high stringency conditions).

[0097] The sequences of the base editors consisting of adenosine deaminases 004V1 and 005V1 and nCas9 are as follows:

[0098] Amino acid sequence of 004V1-nCas9: [ka] (Note that the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V1 sequence, and the asterisk at the C-terminus indicates the position of the stop codon.)

[0099] 005V1-nCas9 amino acid sequence: [ka] (Here, the bold sequence indicates a sequence derived from nCas9, the italicized sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates an adenosine deaminase 005V1 sequence, and the asterisk indicates the position of the stop codon.)

[0100] The nucleotide sequence of 005V1-nCas9 is as follows:

[0101] Furthermore, the present invention provides mutants obtained by performing amino acid substitution based on adenosine deaminase 004V1: adenosine deaminase 004V2, adenosine deaminase 004V3, adenosine deaminase 004V4, adenosine deaminase 004V7, adenosine deaminase 004V8, adenosine deaminase 004V10, adenosine deaminase 004V12, and adenosine deaminase 004V13 to 004V41. Specifically, they are as follows. [Table 5] Exemplary sequences of the base editors that each of the adenosine deaminases constitute are as follows:

[0102] The amino acid sequence of 004V2-nCas9 is as follows: [ka] (Here, the bold sequence indicates a sequence derived from nCas9, the italicized sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates an adenosine deaminase 004V2 sequence, and the asterisk indicates the position of the stop codon.)

[0103] The amino acid sequence of 004V3-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V3 sequence, and the asterisk indicates the position of the stop codon.)

[0104] The amino acid sequence of 004V4-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V4 sequence, and the asterisk indicates the position of the stop codon.)

[0105] The amino acid sequence of 004V7-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V7 sequence, and the asterisk indicates the position of the stop codon.)

[0106] The amino acid sequence of 004V8-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V8 sequence, and the asterisk indicates the position of the stop codon.)

[0107] The amino acid sequence of 004V10-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V10 sequence, and the asterisk indicates the position of the stop codon.)

[0108] The amino acid sequence of 004V12-nCas9 is as follows: [ka] (wherein the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates a deaminase 004V12 sequence, and the asterisk indicates the position of the stop codon.)

[0109] The amino acid sequence of 004V13-nCas9 is as follows: [ka] (Note that the bold sequence indicates a sequence derived from nCas9, the italic sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates an adenosine deaminase 004V13 sequence, and the asterisk at the C-terminus indicates the position of the stop codon.)

[0110] The present invention further relates to a mutant of adenosine deaminase 005V1 obtained by performing amino acid substitution at one or more sites based on the deaminase 005V1, the mutant being 005V1-10-3, 005V1-11-5, 005V1-10-1, 005V1-11-2, 005V1-15-1, 005V1-11-3, 005V1-10-4, 005V1-11-4, 005V1-15-2, 005V1-2-7, 005V1-1-1, 005V 1-1-2, 005V1-1-3, 005V1-1-4, 005V1-1-5, 005V1-1-6, 005V1-1-7, 005V1-1-8, 005V1-2-1, 005V1-2-2, 005V1-2-3 , 005V1-2-5, 005V1-2-6, 005V1-2-8, 005V1-3-1, 005V1-3-2, 005V1-3-3, 005V1-3-4, 005V1-3-5, 005V1-3-7, 005V1 -3-8, 005V1-4-1, 005V1-4-2, 005V1-4-3, 005V1-15-5, 005V1-5-4, 005V1-5-8, 005V1-10-5, 005V1-5-2, 005V1-3- 6, 005V1-4-5, 005V1-4-7, 005V1-4-8, 005V1-5-1, 005V1-5-3, 005V1-5-5, 005V1-5-6, 005V1-5-7, 005V1-6-1, 005V The present invention offers multiple variants including 005V1-1, 005V1-1-2, 005V1-1-3, 005V1-1-4, 005V1-1-5, 005V1-1-6, 005V1-6-5, 005V1-6-6, 005V1-6-8, 005V1-7-1, 005V1-7-2, 005V1-7-3, 005V1-8-1, 005V1-8-2, 005V1-8-3, 005V1-8-4, 005V1-8-5, 005V1-9-1, 005V1-9-2, 005V1-9-3, 005V1-9-4, 005V1-9-5, 005V1-10-2.

[0111] By searching for a suitable adenosine deaminase insertion site in the nCas9 protein, a number of chimeric base editor fusion proteins with improved editing efficiency can also be obtained, and the adenosine deaminases used include adenosine deaminases 005V1, 005V1-10-1, and 005V1-10-3. In the base editor fusion proteins comprising the above three adenosine deaminase sequences, the nucleic acid programmable nucleotide binding domain is an nCas9 domain, and the adenosine deaminase insertion position is between amino acid positions 583-584, 768-769, 770-771, 776-777, 793-794, 905-906, 919-920, 1048-1063, 1049-1062, 1249-1250, 1263-1264, or 1276-1277 of SEQ ID NO:61.

[0112] In the present invention, "retaining the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO:190" and "retaining the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO:3" may mean completely retaining the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO:190 or SEQ ID NO:3, or may only partially retain the activity. In some other embodiments, a base editor having a modified sequence may have higher polynucleotide binding and base editing activity than a base editor of the amino acid sequence shown in SEQ ID NO:190 or SEQ ID NO:3.

[0113] It is understood that the base editor fusion proteins of the present disclosure may include one or more additional features. For example, in some embodiments, the fusion protein may include an inhibitor, a cytoplasmic localization sequence, a trafficking sequence (e.g., a nuclear export sequence or other localization sequence), and a tag useful for lysis, purification, or fusion detection. Suitable tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tag, myc tag, calcitonin tag, FLAG tag, hemagglutinin (HA) tag, polyhistidine tag (also called histidine tag or His-tag), maltose binding protein (MBP) tag, nus tag, glutathione-S-transferase (GST) tag, green fluorescent protein (GFP) tag, thioredoxin tag, S tag, Softags (e.g., Softag 1, Softag 3), strand tag, biotin ligase tag, Flash tag, V5 tag, and SBP tag. Other suitable sequences will be apparent to those of skill in the art. In some embodiments, the fusion protein comprises one or more His tags.

[0114] <Polynucleotide> In a third aspect, the invention provides a polynucleotide encoding an adenosine deaminase according to the first aspect of the invention, or encoding a base editor fusion protein according to the second aspect of the invention.

[0115] <Expression vector> In a fourth aspect, the present invention provides a vector comprising a polynucleotide according to the third aspect of the invention. In some embodiments of the invention, the vector is a mammalian expression vector. In some embodiments, the expression vector is selected from one or more of an adeno-associated virus, a retroviral vector, an adenoviral vector, a lentiviral vector, a Sendai virus vector, and a herpes virus vector. In some embodiments, the vector comprises a promoter.

[0116] <cell> In a fifth aspect, the present invention provides a cell comprising one or more of an adenosine deaminase according to the first aspect of the invention, a base editor fusion protein according to the second aspect of the invention, a polynucleotide according to the third aspect of the invention, and a vector according to the fourth aspect of the invention. In some embodiments of the invention, the cell is a prokaryotic cell, a eukaryotic cell, and may further be a bacterial cell, a plant cell, an insect cell, a human cell or a mammalian cell.

[0117] <Nucleotide Editor System> In a sixth aspect, the present invention provides a base editor system. In some embodiments, the base editor system comprises an adenosine deaminase according to the first aspect of the invention, a nucleic acid programmable nucleotide binding domain, and a guide polynucleotide.

[0118] In some alternative embodiments, the base editor system comprises a base editor fusion protein according to the second aspect of the invention and a guide polynucleotide.

[0119] In some embodiments of the present invention, the guide polynucleotide is a guide RNA (gRNA), a short synthetic RNA consisting of a backbone sequence required for Cas protein binding and a user-defined spacer sequence of about 20 nucleotides. The spacer sequence defines the genome target to be modified. Thus, one skilled in the art can modify the genome target-specific portion of the Cas protein depending on the specificity of the gRNA target sequence for the genome target compared to other parts of the genome.

[0120] In some more specific embodiments of the invention, the guide polynucleotide is an sgRNA, which consists of a backbone sequence required for Cas protein binding and a user-defined spacer sequence of about 20 nucleotides.

[0121] Different backbone sequences can be selected for different origins or types of Cas proteins. In some more specific embodiments of the invention, the backbone sequence of the domain that binds to the Cas9 protein (SpCas9), i.e., the sgRNA, is: GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO:199).

[0122] Pharmaceutical compositions, kits, delivery systems, uses and methods In a seventh aspect, the present invention provides a pharmaceutical composition comprising one or more of an adenosine deaminase according to the first aspect of the invention, a base editor fusion protein according to the second aspect of the invention, a polynucleotide according to the third aspect of the invention, an expression vector according to the fourth aspect of the invention, a cell according to the fifth aspect of the invention and a base editor system according to the sixth aspect of the invention, and a pharma- ceutically acceptable carrier.

[0123] In some embodiments, the pharma- ceutically acceptable carrier may be a delivery vehicle, such as a lipid, a cationic lipid, or a polymer with other drug delivery functionality.

[0124] In an eighth aspect, the present invention provides a kit, particularly a disease treatment kit, comprising one or more of the adenosine deaminase according to the first aspect of the invention, the base editor fusion protein according to the second aspect of the invention, the polynucleotide according to the third aspect of the invention, the expression vector according to the fourth aspect of the invention, the cell according to the fifth aspect of the invention, the base editor system according to the sixth aspect of the invention and the pharmaceutical composition according to the seventh aspect of the invention.

[0125] In a ninth aspect, the present invention provides a delivery system comprising one or more of an adenosine deaminase according to the first aspect of the invention, a base editor fusion protein according to the second aspect of the invention, a polynucleotide according to the third aspect of the invention, a vector according to the fourth aspect of the invention, a cell according to the fifth aspect of the invention and a base editor system according to the sixth aspect of the invention, and a delivery vehicle.

[0126] In some embodiments, the delivery vehicle may be a nanoparticle, a liposome, an exosome, a microvesicle, a gene gun, or a cell membrane penetrating peptide.

[0127] In a tenth aspect, the present invention provides the use of an adenosine deaminase according to the first aspect of the invention in the manufacture of a base editor or a base editor system.

[0128] In an eleventh aspect, the present invention provides the use of an adenosine deaminase according to the first aspect of the invention, a base editor fusion protein according to the second aspect of the invention, a polynucleotide according to the third aspect of the invention, an expression vector according to the fourth aspect of the invention, a cell according to the fifth aspect of the invention, a base editor system according to the sixth aspect of the invention or a pharmaceutical composition according to the seventh aspect of the invention, or a delivery system according to the ninth aspect of the invention in the manufacture of a medicament for treating a disease associated with or caused by a point mutation.

[0129] In some embodiments, the drug is capable of correcting the point mutation, hi some embodiments, the point mutation is G→A and / or C→T.

[0130] In some embodiments, diseases associated with or caused by point mutations include hypercholesterolemia, transthyretin amyloidosis, and β-hemoglobinopathy. In some further embodiments, non-limiting examples of the disease include Meier-Gorlin syndrome, Seckel syndrome, Joubert syndrome, Leber congenital amaurosis, Charcot-Marie-Tooth disease type 2, Usher syndrome type 2C, Spinocerebellar Ataxias, long QT syndrome type 2, Sjogren-Larsson syndrome, genetic diabetes, neuroblastoma, Kallmann syndrome type 1, Kallmann syndrome, metachromatic leukodystrophy, Rett syndrome, amyotrophic lateral sclerosis, and glaucoma. Sclerosis type 10 and Li-Fraumeni syndrome.

[0131] In a twelfth aspect, the present invention provides a method for base editing of a nucleic acid, comprising the step of contacting the nucleic acid with a base editor system described in the sixth aspect of the present invention.

[0132] In some embodiments, the nucleic acid is DNA. Further, the nucleic acid is double-stranded DNA.

[0133] In some embodiments, the nucleic acid comprises a target sequence associated with a disease.

[0134] In some embodiments, the target sequence comprises a point mutation associated with a disease.

[0135] In some specific embodiments, the target sequence contains a disease or disorder-associated G→A or C→T point mutation, in which deamination of the mutated A base results in a sequence that is not associated with the disease or disorder.

[0136] In some embodiments, the target sequence encodes a protein and the point mutation is in a codon that results in an alteration in the amino acid encoded by the mutated codon compared to the wild-type codon.

[0137] In some embodiments, the target sequence is located at a splice site and the point mutation results in a splicing alteration in the mRNA transcript compared to a wild-type transcript.

[0138] In some embodiments, the target sequence is located in a promoter of a gene and the point mutation results in increased gene expression.

[0139] In some embodiments, the target sequence is located in a promoter of a gene and the point mutation results in reduced gene expression.

[0140] In some embodiments, the nucleic acid is located in the genome of the organism.

[0141] In some embodiments, the organism is a prokaryote, a eukaryote, a vertebrate, or a mammal.

[0142] In some embodiments, deamination of the mutant A base results in an alteration of the amino acid encoded by the mutant codon, an alteration of the codon encoding the wild-type amino acid, an alteration of the mRNA transcript, an alteration of the wild-type mRNA transcript, increased gene expression, or decreased gene expression.

[0143] In some embodiments, said contacting is performed in vitro.

[0144] In some embodiments, the contacting is performed within a subject's body (in vivo).

[0145] In some embodiments, the subject has already been diagnosed with the disease or disorder.

[0146] In some embodiments, the disease or disorder is associated with a point mutation in the proprotein convertase subtilisin / kexin type 9 (PCSK9) gene.

[0147] In some embodiments, the disease comprises hypercholesterolemia, transthyretin amyloidosis, β-hemoglobinopathy. In some further embodiments, non-limiting examples of the disease include Meier-Gorlin syndrome, Seckel syndrome, Joubert syndrome, Leber congenital amaurosis, Charcot-Marie-Tooth disease type 2, Usher syndrome type 2C, Spinocerebellar Ataxias, long QT syndrome type 2, Sjogren-Larsson syndrome, genetic diabetes, neuroblastoma, Kallmann syndrome, metachromatic leukodystrophy, Rett syndrome, amyotrophic lateral sclerosis, and glaucoma. Sclerosis type 10 and Li-Fraumeni syndrome.

[0148] The present invention provides in a thirteenth aspect a method for treating a disease associated with or caused by a point mutation. In some embodiments, the method provided comprises administering to a subject suffering from such a disease an effective amount of a base editor fusion protein according to the second aspect of the invention, a base editor system according to the sixth aspect of the invention, a pharmaceutical composition according to the seventh aspect of the invention, a kit according to the eighth aspect of the invention, or a delivery system according to the ninth aspect of the invention to correct the point mutation or introduce an inactivating mutation in the disease-associated gene.

[0149] <Control> In the following examples, a highly efficient base editor ABE8e evolved by David R. Liu team is used for comparison (Richter MF, Zhao KT, Eton E, Lapinaite A, Newby GA, Thuronyi BW, Wilson C, Koblan LW, Zeng J, Bauer DE, Doudna JA, Liu DR. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat Biotechnol. 2020 Jul;38(7):883-891. doi: 10.1038 / s41587-020-0453-z. Epub 2020 Mar 16. Erratum in: Nat Biotechnol. 2020 May 20;: PMID: 32433547; PMCID: PMC7357821.) are used and compared with adenosine deaminase-containing base editors provided by the present invention. ABE8e is an optimized deaminase component of ABE7.10 (Gaudelli NM, Komor AC, Rees HA, Packer MS, Badran AH, Bryson DI, Liu DR. Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage. Nature. 2017 Nov 23;551(7681):464-471. doi: 10.1038 / nature24644. Epub 2017 Oct 25. Erratum in: Nature. 2018 May 2;: PMID: 29160308; PMCID: PMC5726555.), and according to the experimental results obtained by the above team, its activity (first-order kinetics. deamination rate constants, kapp) is 590-fold higher than that of ABE7.10.

[0150] Specifically, in the Examples below, the amino acid sequence of ABE8e is as follows:

[0151] [ka] (Note that the bold sequence indicates a sequence derived from nCas9, the italicized sequence indicates a linker sequence, the double underlined sequence indicates a nuclear localization sequence, the single underlined sequence indicates an ecTadA* deaminase sequence, and the * at the C-terminus indicates the position of the stop codon.)

[0152] Correspondingly, the nucleotide sequence of ABE8e is as follows:

[0153] In the following examples, ABE8e and napDNAbp used in the adenosine deaminase-containing base editor provided by the present invention are both nCas9, and their amino acid sequences are as follows:

[0154]

[0155] The nucleotide sequence of nCas9 is as follows:

[0156] Effect of the Invention

[0157] The present invention improves the editing efficiency of the corresponding base editor by performing site-specific mutation on the parent adenosine deaminase, and is expected to have broad future potential and valuable use. The base editor of the present invention has a significantly higher editing efficiency for PCSK9 target and other sites than the wild type, and has great potential in the field of treating diseases related to or caused by point mutations (e.g., hypercholesterolemia, transthyretin amyloidosis, β-hemoglobinopathy). Through extensive exploratory research, we have found suitable chimeric positions within the nCas9 protein suitable for the adenosine deaminase of the present invention for base editors consisting of some adenosine deaminases and nucleases, and further improved the base editing efficiency.

[0158] For any embodiment of the invention described herein, including those described only in the examples or claims or those described in only one aspect / part below, it is understood that, unless expressly denied or in appropriate combination, said embodiment may be combined with any other one or more embodiments of the invention. [Brief description of the drawings]

[0159] [Figure 1] Figure 1 is the 005V1-nCas9 plasmid map.

[0160] [Diagram 2] 2 shows the editing efficiency of 005V1-nCas9 and each mutant base editor at the PCSK9 site. The editing target site is the PCSK9 gene, and the cells used are 293T cells.

[0161] [Diagram 3]Figure 3 shows a comparison of the editing efficiency at the PCSK9 site between 005V1-nCas9 and each mutant base editor. To more clearly show the comparison of the editing efficiency between the mutant base editor and 005V1-nCas9, the editing efficiency value of 005V1-nCas9 is set to 1, and the editing efficiencies of other mutant base editors are calculated proportionally.

[0162] [Figure 4] Figure 4 shows a comparison of the editing efficiency of mutant base editors with editing efficiency and 005V1-nCas9. 005V1-10-3-nCas9, 005V1-10-1-nCas9, 005V1-11-2-nCas9, 005V1-15-1-nCas9, 005V1-11-3-nCas9, 005V1-11-5-nCas9, 005V1-10-4-nCas9, 005V1-11-4-nCas9, 005V1-15-2-nCas9, 005V1 This includes a comparison of the editing efficiency of base editors such as 005V1-3-3-nCas9, 005V1-15-5-nCas9, 005V1-5-4-nCas9, 005V1-5-8-nCas9, 005V1-10-5-nCas9, 005V1-5-2-nCas9, 005V1-3-6-nCas9, 005V1-2-7-nCas9, and 005V1-nCas9. For convenience of illustration, only the adenosine deaminase name is given in the figure to indicate the corresponding base editor.

[0163] [Diagram 5] Figure 5 is a comparison of the editing efficiency at different sites between 005V1-nCas9 and some mutant base editors. Different depths of color in the figure represent different editing efficiencies. For convenience of illustration, only the adenosine deaminase name is given in the figure to indicate the corresponding base editor.

[0164] [Figure 6A] FIG. 6A is the PHK09 plasmid map. [Figure 6B] Figure 6B is a diagram of the 004V1-nCas9 structure (Figure 6B).

[0165] [Figure 7] Figure 7 shows the A·T→G·C editing efficiency at site 1 of 004V1-nCas9, including panels A-C. Figure 7A is a comparison of the editing efficiency at site 1 of 004V1-nCas9 and ABE8e, where the editing positions are +3, +5, +7, and +8 adenine deoxynucleotides from the 5' end of the sgRNA, and the error bars indicate the mean ± SEM of three biological replicates for each group of samples. Figure 7B is the sequencing result at site 1 after transfection of the base editor 004V1-nCas9. Figure 7C is the sequencing result at site 1 after transfection of the base editor ABE8e.

[0166] [Figure 8] Figure 8 shows the A·T→G·C editing efficiency at site 17 of 004V1-nCas9, including panels A-C. Figure 8A is a comparison of the editing efficiency at site 17 of 004V1-nCas9 and ABE8e, where the editing positions are +3, +4, +5, and +7 adenine deoxynucleotides from the 5' end of the sgRNA, and the error bars indicate the mean ± SEM of three biological replicates for each group of samples. Figure 8B is the sequencing result at site 17 after transfection of the base editor 004V1-nCas9. Figure 8C is the sequencing result at site 1 after transfection of the base editor ABE8e.

[0167] [Figure 9] FIG. 9 shows the A·T→G·C editing efficiency at site 18 of 004V1-nCas9, including panels A-C. FIG. 9A shows a comparison of the editing efficiency at site 18 of p004V1-nCas9 and ABE8e, where the editing positions are +3, +5, +9 adenine deoxynucleotides from the 5' end of the sgRNA, and the error bars show the mean ± SEM of three biological replicates for each group of samples. FIG. 9B shows the sequencing results at site 18 after transfection of the base editor 004V1-nCas9. FIG. 9C shows the sequencing results at site 1 after transfection of the base editor ABE8e.

[0168] [Figure 10] FIG. 10 shows the A·T→G·C editing efficiency at the PCSK9 site of 004V1-nCas9, including panels A-C. FIG. 10A shows a comparison of the editing efficiency at the PCSK9 site of 004V1-nCas9 and ABE8e, where the editing position is +6 adenine deoxynucleotides from the 5' end of the sgRNA, and the error bars show the mean ± SEM of three biological replicates for each group of samples. FIG. 10B shows the sequencing results at the PCSK9 site after transfection of the base editor 004V1-nCas9. FIG. 10C shows the sequencing results at the PCSK9 site after transfection of the base editor ABE8e.

[0169] [Figure 11] Figure 11 shows the editing efficiency of each adenosine deaminase 004V1 mutant-nCas9 and ABE8e at different sites.

[0170] [Figure 12] FIG. 12 shows the editing efficiencies at different sites of base editors consisting of several other mutants of adenosine deaminase 004V1.

[0171] [Figure 13] Figure 13 shows the base editing efficiency at site 1 of some chimeric base editors and base editors in which a deaminase is linked to the N-terminus or C-terminus of nCas9.

[0172] [Figure 14] FIG. 14 shows the base editing efficiency at PCSK9 sites of some chimeric base editors and base editors in which a deaminase is linked to the N-terminus or C-terminus of nCas9. [Figure 15] Figure 15 shows the base editing efficiency at FANCF sites of some chimeric base editors and base editors in which a deaminase is linked to the N-terminus or C-terminus of nCas9. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0173] The experiments and methods described in the examples are basically performed according to conventional methods well known in the art and described in various references, unless otherwise specified. For example, the conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, edited by FMAusubel et al. (1987); METHODS IN ENZYMOLOGY series (Academic Press): PCR 2: A PRACTICAL APPROACH, edited by MJMacPherson, BDHames and GRTaylor (1995); ANTIBODIES, A LABORATORY MANUAL, edited by Harlow and Lane (1988); and ANIMAL CELL CULTURE, edited by RIFreshney (1987). In the examples, the specific conditions are not specified, but are carried out according to the usual conditions or conditions recommended by the manufacturer. Any reagents or equipment used that are not specified by the manufacturer are commercially available and usual. It will be apparent to those skilled in the art that the examples are intended to illustrate the invention in an exemplary manner and are not intended to limit the scope of the claims of the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety. EXAMPLES

[0174] Obtaining base editors consisting of adenosine deaminase 005V1 and its mutants The applicant used bioinformatics to predict key amino acid sites that may affect the biological function, and mutated the amino acid sites to obtain multiple adenosine deaminase mutants whose editing activity is significantly improved compared to the base editor consisting of adenosine deaminase 005V1. The specific site amino acid mutation patterns of 005V1 deaminase are shown in Tables 1a and 1b, and the mutation patterns of each mutant are shown in Table 2. [Table 6] [Table 7] [Table 8]

[0175] Different deaminase variant base editors were generated by site-specific mutagenesis using PCR. Specifically, the DNA sequence encoding the 005V1-nCas9 base editor was amplified around 4 to 6 amino acids near the mutation site, and the sequence to be mutated was introduced into a primer. The amplified fragments were subjected to homologous recombination or enzymatic cleavage and ligation to obtain different mutant base editors. PCR primers for mutant base editors are shown in Tables 3-1, 3-2, and 3-3. [Table 9-1] [Table 9-2] [Table 10-1] [Table 10-2] [Table 11-1] [Table 11-2]

[0176] The applicant obtained the corresponding mutant base editor by site-directed mutagenesis of deaminase 005V1-nCas9 by PCR. The specific method is as follows: As mentioned above, the applicant designed two mutation primers near the mutation site and inserted the sequence to be mutated into the mutation primers. The mutation primers used are shown in Tables 2 and 3. The full-length DNA sequence of the plasmid was amplified (according to the instructions) using 2xPhanta Flash Master Mix enzyme (Vazyme, P520), and the remaining original template was digested after amplification using DpnI enzyme (NEB, R1076L), and then the fragments of interest were enzymatically cut and ligated at the ends using BsaI-HF® v2 (NEB, R3733L) and T4 DNA ligase (NEB, M0202L), and sequenced at ▲Hakushin Biotechnology Co., Ltd. after transformation. The sequencing results showed that different mutants of 005V1-nCas9 with correct sequences were obtained. EXAMPLES

[0177] Verification of the editing activity of base editors constructed with 005V1-nCas9 and each deaminase mutant (1) The procedure for constructing an sgRNA expression vector (sgRNA plasmid) is as follows. PCSK9-sgRNA was designed for PCSK9 target, and site1-sgRNA, site8-sgRNA, site16-sgRNA, and site18-sgRNA were designed based on the targets site1, site8, site16, and site18. (Gaudelli NM, Komor AC, Rees HA, Packer MS, Badran AH, Bryson DI, Liu DR. Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage. Nature. 2017 Nov 23;551(7681):464-471. doi: 10.1038 / nature24644. Epub 2017 Oct 25. Erratum in: Nature. 2018 May 2;: PMID: 29160308; PMCID: (See, e.g., PMC5726555.) Specific sgRNAs are shown in Table 4. [Table 12] Based on the target sequence, sgRNAs were designed and oligonucleotides (oligos) were synthesized. The sgRNA sequences used are shown in SEQ ID NO: 141-145. A CACC sequence was added to the 5' end of the upstream sequence of each sgRNA, and an AAAC sequence was added to the 5' end of the downstream sequence, so that the upstream sequence of each sgRNA used for synthesis was 5'-CACCXXXXXXXXXXXXXXXXXXXX (20 nt)-3', and the downstream sequence was 5'-AAACXXXXXXXXXXXXXXXXXXXX (20 nt)-3'. After synthesis, the upstream and downstream sequences were annealed by a preset program (95 °C, 5 min; -2 °C / s from 95 °C to 85 °C; -0.1 °C / s from 85 °C to 25 °C; hold at 4 °C), and the annealed product was ligated into the PHK09 vector linearized with BsmBI (NEB: R0739L) (the plasmid map is shown in Figure 6A, and it is a homemade product that already contains the backbone sequence of sgRNA). The sequence of the PHK09 vector is as follows: The system used to construct the sgRNA plasmid is as follows.

[0178] The linearization system for the PHK09 vector was as follows: 3 μg of PHK09 vector; 6 μL of buffer (NEB:R0539L); 2 μL of BsmBI; supplemented to 60 μL with ddH2O and digested with enzyme overnight at 37°C.

[0179] The ligation system for the sgRNA annealed product and the linearized vector was as follows: 1 μL of T4 ligase buffer (NEB:M0202L), 20 ng of linearized vector, 5 μL of annealed oligo fragment (10 μM), 0.5 μL of T4 ligase (NEB:M0202L), supplemented with ddH2O to 10 μL, and ligated overnight at 16 °C.

[0180] The ligated vector was transformed into E. coli DH5a competent cells (Weidi Biotechnology Co., Ltd., DL1001). The specific procedure was as follows: DH5α competent cells were removed from -80°C and quickly inserted into ice. After 5 minutes, the bacterial pellet melted, and the ligation product was added, the bottom of the centrifuge tube was tapped by hand to mix gently, and the tube was left to stand in ice for 25 minutes. The tube was heat-shocked in a 42°C water bath for 45 seconds, and then quickly returned to ice and left to stand for 2 minutes. 700μl of antibiotic-free sterile LB medium was added to the centrifuge tube, mixed uniformly, and then resuscitated at 37°C and 200 rpm for 60 minutes. The tube was centrifuged at 5000 rpm for 1 minute to recover the bacteria, leaving about 100μl of the supernatant, and the bacterial pellet was resuspended by gently pipetting, and then applied to LB medium containing Amp antibiotics. The plate was inverted and placed in a 37°C incubator to culture overnight. Single colonies were picked and confirmed by sequencing. The positive clones were then cultured with shaking to extract the plasmid (TIANGEN, DP120-01), after which the concentration was measured and the plasmid was stored in a -20°C refrigerator until use.

[0181] (2) Cell culture and transfection HEK293 T cells (purchased from ATCC) were seeded in DMEM medium containing 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) (Gibco, 11965092) and cultured in a cell incubator at 5% CO2 and 37°C. The cells for transfection were seeded and cultured on a 24-well cell culture plate the day before, and the cells were observed the next day. Transfection was performed when the cells became about 80% confluent. The amount of plasmid transfected per well of the 24-well plate was 0.4 μg of 005V1-nCas9 plasmid and each 005V1 mutant-nCas9, and 0.4 μg of sgRNA plasmid. After mixing the plasmids, they were diluted with 25 μl of serum-reduced medium (Genbai Seibu, L530 KJ), and 2 μl of p3000 reagent was added, mixed uniformly by pipetting, and left to stand for 5 minutes. At the same time, 2 μl of Lipofectamine3000 transfection reagent (Thermo, 11668019) was diluted with 25 μl of serum-reduced medium, mixed uniformly, and left to stand for 5 minutes. The above-mentioned reagents A and B were mixed, pipetted uniformly, and left to stand for 20 minutes. After the standing, the mixed reagent was dropped drop by drop onto the cells in the 24-well plate for transfection, and the cells were returned to the incubator at 37 ° C. and cultured. After 6 hours of transfection, the medium was replaced with DMEM medium containing 10% FBS. After 48 hours of transfection, the cells were harvested and the editing efficiency was detected. (3) The harvested cells were subjected to genome extraction (TIANGEN, DP304-03), and primers were designed according to experimental needs. The identification primer sequences used are shown in Table 5. [Table 13] The target adjacent sequence was amplified by PCR using the genome as a template, and the amplified PCR product was used to identify the editing efficiency by high-throughput sequencing (GENEWIZ, Inc.) or Sanger sequencing (Hakushin Biotechnology (Shanghai) Co., Ltd.). The target site sequence amplification system consisted of 25 μL of 2×Taq Master Mix (Vazyme, P112-03); 1 μL of Primer-F (10 pmol / μL); 1 μL of Primer-R (10 pmol / μL); 1 μL of template; supplemented with ddH2O to 50 μL.

[0182] Detecting gene editing effects: For methods to calculate gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun;1(3):239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262;PMCID:PMC6694769. In this example, the structure of a base editor comprising adenosine deaminase and nCas9 provided by the present invention is as follows: NH2-[NLS]-[adenosine deaminase]-linker-[nCas9]-[NLS]-COOH. However, this is merely an example and does not limit the structure of the base editor. The structural example can be seen in SEQ ID NO:3 and SEQ ID NO:190. In this example, the editing efficiency of each base editor for the PCSK9 target (efficiency of mutation from adenine A to guanine G) was statistically measured and compared. The editing efficiencies of the base editors composed of 005V1-nCas9 and each deaminase mutant are shown in Tables 6 to 9. [Table 14] [Table 15] [Table 16] [Table 17] The editing efficiencies and comparisons of editors consisting of 005V1-nCas9 and each mutant are shown in Figures 2 to 5.

[0183] Figure 2 shows the A·T→G·C editing efficiency at the PCSK9 site for 005V1-nCas9 and each mutant of 005V1-nCas9. The editing position is +6 adenine deoxynucleotides from the 5' end of the sgRNA, and error bars indicate the mean ± SEM of three biological replicates for each group of samples.

[0184] As shown in Figure 3, 005V1-10-3-nCas9, 005V1-11-2-nCas9, 005V1-10-1-nCas9, 005V1-10-4-nCas9, 005V1-11-3-nCas9, 005V1-15-1-nCas9, 005V1-11-5-nCas9, 005V1-11-4-nCas9, 005V1-15-2-nCas9, 005V1-3-3-nCas9, 005V1-5-4 The editing efficiency at the PCSK9 site of 005V1-nCas9, 005V1-15-5-nCas9, 005V1-5-8-nCas9, 005V1-10-5-nCas9, 005V1-5-2-nCas9, 005V1-3-6-nCas9, and 005V1-2-7-nCas9 was significantly improved, with the average editing efficiency being over 20%, which is 1.07 to 2.41 times higher than the editing efficiency of the original 005V1-nCas9.

[0185] As shown in Figure 4, among the mutants with significantly improved editing efficiency at the PCSK9 site, the editing efficiency of 005V1-10-3-nCas9 can reach 46.67%. The editing efficiencies of 005V1-11-2-nCas9 and 005V1-10-1-nCas9 also reach 35.67% and 35.33%. The editing efficiencies of most of the other mutant base editors in the figure are also higher than that of 005V1-nCas9.

[0186] As shown in Figure 5, 005V1-5-4-nCas9, 005V1-3-3-nCas9, 005V1-5-2-nCas9, 005V1-2-7-nCas9, 005V1-3-6-nCas9, 005V1-2-8-nCas9, 005V1-5-3-nCas9, and 005V1-3-4-nCas9 have clearly narrower editing windows at site 1 (+5 positions, +7 positions), site 8 (+2 positions, +3 positions, +4 positions, +6 positions), and site 16 (+3 positions, +4 positions, +6 positions). In contrast, the above-mentioned base editor has a narrower editing window at the site 1 site, but still maintains very high editing activity for the adenine deoxynucleotide at position +5 from the 5' end of the sgRNA within the window, enabling efficient and precise editing, and has great potential application value. EXAMPLES

[0187] In this example, adenosine deaminase 004V1 was used and expressed in fusion with nCas9 to construct a novel adenine base editor 004V1-nCas9 with a narrower editing window. In this example, the construction strategy for the adenine base editor 004V1-nCas9 is to obtain the novel adenine base editor 004V1-nCas9 by replacing the adenosine deaminase in ABE8e with 004V1. In the following examples, the structures of base editors comprising adenosine deaminase and nCas9 provided by the present invention are as follows: NH2-[NLS]-[adenosine deaminase]-linker-[napDNAbp]-[NLS]-COOH. However, this is merely an example and does not limit the structure of the base editor. In the following examples, the sgRNAs used refer to Table 10. The identifying primer sequences used are shown in Table 11. [Table 18] [Table 19]

[0188] The specific construction of the sgRNA and base editor is as follows. 1. The method for constructing the sgRNA expression vector (sgRNA plasmid) was the same as in Example 2. 2. Construction of adenine base editor 004V1-nCas9 expression vector (004V1-nCas9 plasmid) In this example, an adenine base editor expression vector 004V1-nCas9 was prepared. The nucleotide sequence of adenosine deaminase 004V1 is shown in SEQ ID NO:156.

[0189] Nucleotide sequence of 004V1: (SEQ ID NO:156) The nucleotide sequence of the deaminase 004V1 was codon-optimized according to the bias of human codon usage frequency, and the artificial synthesis of the 540bp deaminase 004V1 gene was entrusted to Sangon Biotech (Shanghai) Co., Ltd., which replaced the 63rd to 560th nucleotides of the ABE8e sequence with the synthesized gene. The 004V1-nCas9 plasmid map is shown in FIG. 6B. (The amino acid sequence of 004V1-nCas9 is correspondingly shown in SEQ ID NO:190.)

[0190] The nucleotide sequence of 004V1-nCas9 is as follows:

[0191] 3. Cell Culture and Transfection HEK293 T cells (purchased from ATCC) were seeded in DMEM medium containing 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) (Gibco, 11965092) and cultured in a cell incubator at 5% CO2 and 37°C. The cells for transfection were seeded and cultured on a 24-well cell culture plate the day before, and the cells were observed the next day. Transfection was performed when the cells became about 80% confluent. The amount of plasmid transfected per well of the 24-well plate was 0.4 μg of p004V1-nCas9 plasmid and 0.4 μg of sgRNA plasmid, respectively. After mixing the plasmids, they were diluted with 25 μl of serum-reduced medium (Genbai Seibu, L530 KJ), and 2 μl of p3000 reagent was added, mixed uniformly by pipetting, and left to stand for 5 minutes. At the same time, 2 μl of Lipofectamine3000 transfection reagent (Thermo, 11668019) was diluted with 25 μl of serum-reduced medium, mixed uniformly, and left to stand for 5 minutes. The above-mentioned reagents A and B were mixed, pipetted uniformly, and left to stand for 20 minutes. After the standing, the mixed reagent was dropped drop by drop onto the cells in the 24-well plate for transfection, and the cells were returned to the incubator at 37 ° C. and cultured. After 6 hours of transfection, the medium was replaced with DMEM medium containing 10% FBS. After 48 hours of transfection, the cells were harvested and the editing efficiency was detected.

[0192] 4. Detection of editing efficiency of the adenine base editor 004V1-nCas9 in an endogenous gene site according to this embodiment The cells described in "3. Cell culture and transfection" were subjected to genome extraction (TIANGEN, DP304-03). Primers were designed according to experimental needs, and the identified primer sequences used are shown in SEQ ID NO: 148-149, 166-167, 154-155 in Table 11. The genome was used as a template to PCR amplify the target adjacent sequence, and the amplified PCR product was used to identify the editing efficiency by high-throughput sequencing (GENEWIZ, Inc.) or Sanger sequencing (▲Hakushin Biotechnology (Shanghai) Co., Ltd.). The system for amplifying the target site sequence was 25 μL of 2×Taq Master Mix (Vazyme, P112-03); 1 μL of Primer-F (10 pmol / μL); 1 μL of Primer-R (10 pmol / μL); 1 μL of template; supplemented with ddH2O to 50 μL.

[0193] The procedure for testing the gene editing effect is as follows:

[0194] When the 004V1-nCas9 plasmid was co-transfected into HEK293T cells (purchased from ATCC) together with sgRNA plasmids at different sites, we found that the editing efficiency of 004V1-nCas9 was similar to that of ABE8e at site 1 and site 17 compared to the ABE8e plasmid (addgene, Plasmid #138489) (Figures 7 and 8), but the editing efficiency of 004V1-nCas9 was clearly superior to that of ABE8e at site 18 (Figure 9).

[0195] Regarding the editing window, the editing window of 004V1-nCas9 is smaller than that of ABE8e at sites 1, 17, and 18 (see Figures 7 to 9).

[0196] For the calculation method of gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun;1(3):239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262; PMCID: PMC6694769. The results of this example are shown in Figures 7 to 9. EXAMPLES

[0197] In this example, the adenine base editor 004V1-nCas9 obtained in Example 3 was applied to the treatment of a disease.

[0198] Proprotein convertase subtilisin / kexin type 9 (PCSK9) is the ninth member of the kexin-like proprotein convertase subtilisin family and consists of 692 amino acid residues. As a negative regulator of the low-density lipoprotein receptor (LDLR), excess PCSK9 can accelerate the degradation of LDLR on the surface of hepatocytes after binding to it, which reduces the uptake of low-density lipoprotein cholesterol (LDL-C) by hepatocytes, further increasing the LDL-C level in the peripheral circulation, and ultimately increasing the cholesterol level in the blood.

[0199] 1. Construction of sgRNA expression vector (sgRNA plasmid) In this example, the construction of the sgRNA plasmid targeting PCSK9 was as described in Example 1. The sgRNA sequence used is shown in SEQ ID NO:141.

[0200] PCSK9-sgRNA: Cccgcaccttggcgcagcgg (SEQ ID NO:141).

[0201] 2. Cell Culture and Transfection The culture and transfection methods of HEK293T cells in this example were the same as those in Example 3.

[0202] 3. Detection of editing efficiency at PCSK9 sites using optimized base editing tools In this example, the method for detecting editing efficiency is as described in Example 1. The identification primer sequences used are shown in SEQ ID NO:146-147.

[0203] PCSK9-forward primer Gctagccttgcgttccg (SEQ ID NO:146); PCSK9-reverse primer Gtccccaagatcgtgccaa (SEQ ID NO:147).

[0204] In this example, the adenine base editor expression vector p004V1-nCas9 was co-transfected into HEK293T cells together with a sgRNA plasmid targeting PCSK9. As shown in FIG. 10, compared with the ABE8e adenine base editor, the editing efficiency of 004V1-nCas9 is superior to ABE8e. It can be seen that the base editor 004V1-nCas9 can target and treat hypercholesterolemia caused by high expression of PCSK9. EXAMPLES

[0205] In this example, adenosine deaminase 004V2, adenosine deaminase 004V3, adenosine deaminase 004V4, adenosine deaminase 004V7, adenosine deaminase 004V8, adenosine deaminase 004V10, adenosine deaminase 004V12, and adenosine deaminase 004V13, which are mutants of adenosine deaminase 004V1, were used to construct an adenine base editor expression vector and an sgRNA expression vector in the same manner as in Example 3 or Example 4, and cell culture and transfection were performed to detect the editing efficiency.

[0206] The amino acid sequence information of the base editor consisting of the above mutant is as follows.

[0207] 004V2-nCas9(SEQ ID NO:191), 004V3-nCas9(SEQ ID NO:192), 004V4-nCas9(SEQ ID NO:193), 004V7-nCas9(SEQ ID NO:194), 004V8-nCas9(SEQ ID NO:195), 004V10-nCas9(SEQ ID NO: 196), 004V12-nCas9 (SEQ ID NO: 197), 004V13-nCas9 (SEQ ID NO: 198).

[0208] The results of base editing efficiency are shown in Figure 11. As shown in Figure 11, compared with ABE8e, 004V2-nCas9, 004V3-nCas9, 004V4-nCas9, 004V7-nCas9, 004V8-nCas9, 004V10-nCas9, 004V12-nCas9, and 004V13-nCas9 have obviously narrower editing windows at site 1 (+5, +7). Among them, the editing window of 004V12-nCas9 is the narrowest, with editing efficiency only at position +5, and a base editing efficiency of 46%, which can achieve efficient precise editing. At site 18, the editing efficiencies of different mutants of 004V1-nCas9 are all obviously better than ABE8e. For the PCSK9 site, the editing efficiency of 004V1-nCas9 and 004V3-nCas9 is superior to ABE8e.

[0209] The nucleotide sequences of each adenosine deaminase base editor are as follows:

[0210] Nucleotide sequence of 004V2-nCas9:

[0211] Nucleotide sequence of 004V3-nCas9:

[0212] Nucleotide sequence of 004V4-nCas9:

[0213] Nucleotide sequence of 004V7-nCas9:

[0214] Nucleotide sequence of 004V8-nCas9:

[0215] Nucleotide sequence of 004V10-nCas9:

[0216] Nucleotide sequence of 004V12-nCas9:

[0217] Nucleotide sequence of 004V13-nCas9: EXAMPLES

[0218] Several other amino acid sites of adenosine deaminase 004V1 were also mutated, resulting in adenosine deaminase 004V14, adenosine deaminase 004V15, adenosine deaminase 004V16, adenosine deaminase 004V17, adenosine deaminase 004V18, adenosine deaminase 004V19, adenosine deaminase 004V20, adenosine deaminase 004V21, adenosine deaminase 004V22, adenosine deaminase 004V23, adenosine deaminase 004V24, adenosine deaminase 004V25, adenosine deaminase 004V26, and adenosine deaminase 004V27. adenosine deaminase 004V27, adenosine deaminase 004V28, adenosine deaminase 004V29, adenosine deaminase 004V30, adenosine deaminase 004V31, adenosine deaminase 004V32, adenosine deaminase 004V33, adenosine deaminase 004V34, adenosine deaminase 004V35, adenosine deaminase 004V36, adenosine deaminase 004V37, adenosine deaminase 004V38, adenosine deaminase 004V39, adenosine deaminase 004V40 and adenosine deaminase 004V41 were obtained. For the amino acid mutation patterns, see Table 12, and for the amino acid mutation patterns of each deaminase mutant, see Table 13. [Table 20] [Table 21]

[0219] An adenine base editor expression vector consisting of adenosine deaminase 004V14-004V41 and a PCSK9-sgRNA expression vector were constructed in the same manner as in Examples 2 to 4, and cell culture and transfection were performed in the same manner as in the above examples. The test procedures for the gene editing effect are as follows. For methods to calculate gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun;1(3):239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262; PMCID: PMC6694769. The editing efficiency of each of the base editors consisting of adenosine deaminase 004V1 and nCas9 described above is compared with the editing efficiency of deaminase 004V1-nCas9 as follows. [Table 22] The results of this example are shown in FIG. 12. Compared to 004V1-Cas9, the base editors consisting of deaminases 004V14-004V32 and nCas9 all have higher editing efficiency at the PCSK9 site than 004V1-nCas9. The editing efficiency of the base editors consisting of adenosine deaminases 004V14-004V29 and nCas9 is significantly improved compared to 004V1-nCas9, and the editing efficiency of 004V17-nCas9 to 004V29-nCas9 is improved by more than 20% compared to 004V1-nCas9. In particular, the editing efficiency of 004V14-nCas9, 004V15-nCas9, and 004V16-nCas9 is more than twice that of 004V1-nCas9. EXAMPLES

[0220] In order to further improve the editing efficiency of the base editor fusion protein, the applicant explored the structure of the base editor fusion protein. Using adenosine deaminases including adenosine deaminases 005V1, 005V1-10-1, and 005V1-10-3, the applicant searched for a suitable adenosine deaminase insertion site in the nCas9 protein. For the sequences of the above three adenosine deaminases, see Example 1. The experimental procedure is as follows.

[0221] 1. Construction of nCas9 Plasmid (1) Primer design (primers were synthesized by Shanghai Bosun Biotechnology Co., Ltd.) [Table 23] ABE8e (Addgene, #138489) was PCR amplified using a high-fidelity enzymes kit (Vazyme, P501-d2) from Vazyme Biotech Co., Ltd. The amplification system is shown in Table 15. [Table 24] The PCR amplification process is shown in the table below. [Table 25] The amplified PCR products were recovered according to the kit instructions (Amane, General-Purpose DNA Purification and Recovery Kit, DP214). The purified PCR product was transformed into E. coli DH5a competent cells (Weidi Biotechnology Co., Ltd., DL1001). The specific procedure was as follows: DH5α competent cells were removed from -80°C and quickly inserted into ice. After 5 minutes, the bacterial pellet melted, and the ligation product was added, the bottom of the centrifuge tube was tapped by hand to mix gently, and the tube was left in ice for 25 minutes. The tube was heat-shocked in a 42°C water bath for 45 seconds, and then quickly returned to ice and left to stand for 2 minutes. 700μl of antibiotic-free sterile LB medium was added to the centrifuge tube, mixed uniformly, and then resuscitated at 37°C and 200 rpm for 60 minutes. The tube was centrifuged at 5000 rpm for 1 minute to recover the bacteria, leaving about 100μl of the supernatant, and the bacterial pellet was resuspended by gently pipetting, and then applied to LB medium containing Amp antibiotics. The plate was inverted and placed in a 37°C incubator to culture overnight. Single colonies were picked and confirmed by sequencing, and the positive clones were then cultured with shaking. The nCas9 plasmid was extracted using an endotoxin-free plasmid megaprep kit (TIANGEN: DP120-01), the concentration was measured, and the plasmid was stored in a -20°C refrigerator until use.

[0222] (2) Obtaining DNA sequences encoding deaminases 005V1, 005V1-10-1, and 005V1-10-3 Using primers, 005V1-nCas9, 005V1-10-1-nCas9, and 005V1-10-3-nCas9 plasmids were PCR amplified, and the amplification system and PCR process are shown in Tables 15 and 16. The amplified PCR products were recovered according to the kit instructions (Tengen, General-Purpose DNA Purification and Recovery Kit, DP214). The PCR primers used are shown in the table below. [Table 26]

[0223] (3) Design of different chimeric base editors and the primer sequences of the corresponding nCas9 plasmids The insertion positions of deaminases 005V1, 005V1-10-1, and 005V1-10-3 in nCas9 were examined, and multiple chimeric base editors were designed, and primer sequences corresponding to nCas9 were designed according to the different insertion positions of the above deaminases. Specifically, they are shown in Table 18. [Table 27] The primer sequence designs for the nCas9 plasmids corresponding to the different chimeric base editors listed in Table 18 are as follows: [Table 28-1] [Table 28-2] The nCas9 plasmids obtained in step (1) were each amplified using the primer sequences in Table 19. The amplification system and PCR process are shown in Tables 15 and 16. The amplified PCR products were collected according to the kit instructions (Tengen, Universal DNA Purification and Collection Kit, DP214). The primers used to amplify the nCas9 plasmid corresponding to different insertion positions are shown in Table 19. For example, the amino acid sequence of 005V1-1249-ABE is as follows: [ka] (Note that italics indicate NLS, bold indicates nCas9 fragment, underline indicates linker, double underline indicates deaminase, and the C-terminal asterisk indicates the position of the stop codon. The deaminase sequence is deaminase 005V1, and the position of incorporation of deaminase 005V1 in nCas9 is between 1249 and 1250.)

[0224] (4) The PCR products of adenosine deaminase 005V1, 10-1, and 10-3 obtained in step (2) and the linearized PCR products of different insertion positions obtained in step (3) were subjected to homologous recombination using the Gibson Assembly Master Mix recombination kit (NEB, E2611S) to obtain different chimeric recombinant plasmids. The reaction system is shown in the table below. [Table 29] The homologous recombination product was transformed into E. coli DH5a competent cells (Weidi Biotechnology Co., Ltd., DL1001). The specific procedure was as follows: DH5α competent cells were removed from -80°C and quickly inserted into ice. After 5 minutes, the bacterial pellet melted, and the ligation product was added, the bottom of the centrifuge tube was tapped by hand to mix gently, and the tube was left to stand in ice for 25 minutes. The tube was heat-shocked in a 42°C water bath for 45 seconds, and then quickly returned to ice and left to stand for 2 minutes. 700μl of antibiotic-free sterile LB medium was added to the centrifuge tube, mixed uniformly, and then resuscitated at 37°C and 200 rpm for 60 minutes. The tube was centrifuged at 5000 rpm for 1 minute to recover the bacteria, leaving about 100μl of the supernatant, and the bacterial pellet was resuspended by gently pipetting, and then applied to LB medium containing Amp antibiotics. The plate was turned upside down and placed in a 37°C incubator to culture overnight. Single colonies were picked and confirmed by sequencing, and the positive clones were then cultured with shaking to extract the chimeric recombinant plasmid using an endotoxin-free plasmid megaprep kit (TIANGEN: DP120-01), after which the concentration was measured and the plasmid was stored in a -20°C refrigerator until use.

[0225] (5) Cell culture and transfection HEK293 T cells (purchased from ATCC) were seeded in DMEM medium containing 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) (Gibco, 11965092) and cultured in a cell incubator at 5% CO2 and 37°C. The cells for transfection were seeded and cultured on a 24-well cell culture plate the day before, and the cells were observed the next day. Transfection was performed when the cells became about 80% confluent. The amount of plasmid transfected per well of the 24-well plate was 0.4 μg of chimeric recombinant plasmid and 0.4 μg of sgRNA plasmid, respectively. The sgRNA used is shown in the table below. [Table 30] The PCSK9-sgRNA and site1-sgRNA sequences used in this example are identical to the corresponding sequences in other examples, and FANCF-sgRNA is an sgRNA targeting the FANCF gene. After mixing the chimeric recombinant plasmid and the sgRNA plasmid, the mixture was diluted with 25 μl of serum-reduced medium (Genbai Seibutsu, L530 KJ), and 2 μl of p3000 reagent was added, mixed uniformly by pipetting, and left to stand for 5 minutes. At the same time, 2 μl of Lipofectamine3000 transfection reagent (Thermo, 11668019) was diluted with 25 μl of serum-reduced medium, mixed uniformly, and left to stand for 5 minutes. The above-mentioned reagents A and B were mixed, pipetted uniformly, and left to stand for 20 minutes. After the end of the stand, the mixed reagent was dropped drop by drop onto the cells in the 24-well plate for transfection, and the cells were returned to the incubator at 37 ° C. for culture. After 6 hours of transfection, the medium was replaced with DMEM medium containing 10% FBS. After 48 hours of transfection, the cells were collected and the editing efficiency was detected.

[0226] (6) Detection of editing efficiency Genomic DNA extraction kit (TIANGEN, DP304-03) was used to extract genome from HEK293T cells. Primers were designed according to experimental needs, and the identified primer sequences used are shown in Table 21. [Table 31]

[0227] Using the genome as a template, sequences adjacent to the sgRNA target were amplified by PCR using identified primers, and the amplified PCR products were used to identify the editing efficiency by Sanger sequencing (Hakushin Biotechnology (Shanghai) Co., Ltd.). The system for amplifying the target site sequence consisted of 25 μL of 2×Taq Master Mix (Vazyme, P112-03); 1 μL of Primer-F (10 pmol / μL); 1 μL of Primer-R (10 pmol / μL); 1 μL of template; and supplemented to 50 μL with ddH2O. For methods to calculate gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun;1(3):239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262; PMCID: PMC6694769. The results of this example are shown in Figures 13 to 15. At the site 1 site, the base editing efficiency of each chimeric base editor and the base editor in which deaminase is linked to the N-terminus or C-terminus of nCas9 is shown in Figure 13. At the site 1 site, it can be seen that the base editors 005V1-C-ABE, 005V1-10-1-C-ABE, and 005V1-10-3-C-ABE all have no editing activity. For the other base editors, 005V1-1047-1064-ABE, 005V1-1048-1063-ABE, and 005V1-1249-ABE have significantly higher editing efficiency at the A5 position at the site 1 site than 005V1-N-ABE, and in particular, the editing efficiency of 005V1-1249-ABE is nearly twice that of 005V1-N-ABE. 005V1-776-ABE, 005V1-793-ABE, 005V1-905-ABE, and 005V1-919-ABE have higher base editing efficiency at specific positions in site1 than 005V1-N-ABE, and can be used to edit specific gene positions to treat diseases. 005V1-10-1-1249-ABE has higher editing efficiency at all positions in site1 than 005V1-10-1-N-ABE. 005V1-10-3-1249-ABE has higher editing efficiency at all positions in site1 than 005V1-10-3-N-ABE. It can be seen that for 005V1 and its mutants 005V1-10-1 and 005V1-10-3, better editing efficiency can be obtained when the nCas9 insertion position is selected between positions 1249 and 1250. The base editing efficiency of each chimeric base editor and a base editor in which a deaminase is linked to the N-terminus or C-terminus of nCas9 at the PCSK9 site is shown in Figure 14. It can be seen that at the PCSK9 site, the editing efficiency of 005V1-1249-ABE is higher than 005V1-N-ABE, the editing efficiency of 005V1-10-1-1249-ABE at the site is higher than 005V1-10-1-N-ABE, and the editing efficiency of 005V1-10-3-1249-ABE at the site is higher than 005V1-10-3-N-ABE. The base editing efficiency of each chimeric base editor and the base editor in which deaminase is linked to the N-terminus or C-terminus of nCas9 at the FANCF site is shown in Figure 15. At the FANCF site, 005V1-1249-ABE shows higher editing efficiency than 005V1-N-ABE at certain positions, and the editing characteristics at some positions are different from 005V1-N-ABE. 005V1-10-1-1249-ABE shows editing efficiency at the A10 and A12 positions of the site, but 005V1-10-1-N-ABE does not show editing activity at these positions, and the editing activity at other positions is not significantly different from 005V1-10-1-N-ABE. 005V1-10-3-1249-ABE exhibits similar properties to 005V1-10-1-1249-ABE at the site, showing editing efficiency at positions A10 and A12, whereas 005V1-10-3-N-ABE does not show editing activity at these positions and editing activity at other positions is not significantly different from 005V1-10-3-N-ABE.

Claims

1. The following array: (i) the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170; (ii) an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170, wherein the amino acid sequence retains the deaminating activity of the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170; (iii) an amino acid sequence having one or more amino acid residues added, substituted, deleted, or inserted in the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170, which retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170; or (iv) an amino acid sequence encoded by a nucleotide sequence that hybridizes under stringent conditions with a polynucleotide sequence encoding the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170, wherein the stringent conditions are medium stringency conditions, medium to high stringency conditions, high stringency conditions, or very high stringency conditions, and the amino acid sequence retains the deamination activity of the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 170; Adenosine deaminase characterized by comprising one or more selected from the following.

2. 2. The adenosine deaminase of claim 1, wherein the adenosine deaminase has modifications at any two or more of the following amino acid positions relative to the amino acid sequence set forth in SEQ ID NO: 1: 33, 35, 36, 46, 47, 48, 49, 104, 105, 107, 148, 149, 150, 151, 152, 153, 154, and 155.

3. The substitutions are S15, D16, H17, E18, F19, N20, D21, E22, Y23, W24, M25, R26, H27, A28, L29, T30, K33, R34, A35, R36, V41, V43, L47, L49, N51, N59, A61, I62, L64, A69, E72, G80, L81, V82, L83, Q84, N85, Y86, I89, D90, A91, T92, 2. The adenosine deaminase of claim 1, wherein the substitution occurs at one or more of the following positions: V95, F97, A106, R111, I112, S113, R114, L115, F117, V119, R120, N121, S122, K123, R124, N132, V133, L134, N135, P137, G138, M139, N140, H141, R142, E144, D160, V168, F169, N170.

4. The adenosine deaminase may comprise any one of the following amino acids with respect to the amino acid sequence shown in SEQ ID NO: 1: V33I, D35G, D35R, D36G, D36N, A46C, I47Y, I47R, I47V, T48G, T48R, L49T, L49K, L49H, V104A, V104M, S105C, S105G, S107R, S107K, S107A, Q148P, Q148C, Q148A, Q148G, Q149M, 3. The adenosine deaminase of claim 2, comprising an alteration at an amino acid position selected from Q149G, Q149L, P150R, P150L, ​​P150C, R151K, E152R, E152T, E152G, E152W, V153P, V153F, V153I, V153T, F154H, F154K, F154L, N155T, N155R, and N155H.

5. The adenosine deaminase has the following amino acid sequences with respect to the amino acid sequence shown in SEQ ID NO: 1: (1) Q148G + Q149M + P150R, (2) S107A, (3) E152R + V153P + F154H + N155T, (4) S107K, (5) Q148P + Q149G + P150L, ​​(6) A46C + I47Y + T48G + L49H, (7) E152T + V153F + N155T, (8) S107R, (9) V104M, (10) E152G + V153I + F154K + N155R, (11) S105C, The adenosine deaminase of claim 4, comprising an alteration selected from any one of the following sets: (12) Q148C + Q149L + P150R + R151K, (13) Q148A + Q149G + P150C, (14) E152W + V153T + F154L + N155H, (15) S105G, (16) I47R + T48G + L49T, (17) D35G + D36N, (18) V33I + D35R + D36G, (19) V104A, and (20) I47V + T48R + L49K.

6. The substitutions are selected from the group consisting of S15T, S15G, D16E, D16N, H17C, H17K, E18D, E18K, F19C, F19Y, F19S, N20Q, N20L, N20S, D21Q, E22D, Y23F, W24F, M25L, M25V, R26K, R26T, H27R, A28C, L29I, T30E, K33R, K33S, R34K, A35S, R36Q, L49F, N51G, L64M, L64Q, L64P, G80A, L81N, V82A, V82T, L83I, and L83Q in the amino acid sequence shown in SEQ ID NO:

170. , Q84N, N85S, Y86W, I89E, I89L, D90G, A91C, A91T, T92D, A106C, I112L, S113K, R114K, V119L, S122N, S122P, R124H, R124T, N132K, V133I, L134F, N135H, N135S, P137F, G138A, M139L, N140K, H141A, R142L, R142S, E144H, D160E, V168A, F169C, F169V, N170D; Preferably, the substitutions are: (1) S15T + D16E + H17K + F19Y + N20Q, (2) E22D + Y23F + W24F + R26K + H27R + L29I, (3) K33R + R34K + A35S, (4) G80A + L81N + V82A + L83I + Q84N + N85S + Y86W, (5) I89L + D90G + A91T + T92D, (6) I112L + S113K + R114K, (7) I112L + S113K + R114K, (8) I112L + S113K + R114K, (9) I112L + S113K + R114K, (10) I112L + S113K + R114K, (11) I112L + S113K + R114K, (12) I112L + S113K + R114K, (13) I112L + S113K + R114K, (14) I112L + S113K + R114K, (15) I112L + S113K + R114K, (16) I112L + S113K + R114K, (17) I112L + S113K + R114K, (18) I112L + S113K + R114K, (19) I112L + S113K + R114K, (20) I112L + S113K + R114K, (21) I112L + S113K + R114K, (22) I112L + S113K + 7) N132K+V133I+L134F+N135H, (8) P137F+G138A+M139L, (9) E18K+F19S+N20L, (10) M25L, (11) S15G+D16N+H17C , (12) L64Q+S122P+V168A, (13) R36Q+L64Q+S122P+F169V, (14) R36Q+V119L+V168A, (15) R36Q+V119L+N170D, (1 6) L64Q+V119L+V168A, (17) F169C, (18) L64Q+V119L+V168A, (19) L64P+R124T+F169V, (20) I89E+A91C, (21) L64 Q+V119L+F169V, (22) L64M, (23) L64Q+S122P+N170D, (24) E18D+F19C+N20S, (25) L64P+V119L+N170D, (26) S122 The adenosine deaminase of claim 3, wherein the substitution occurs in a combination of the following positions: N + R124H + N135S + D160E, (27) R26T + H27R + A28C, (28) V82T + L83Q, (29) M25V, (30) D21Q + T30E + K33S, (31) N140K + H141A + R142S, (32) L49F + N51G, (33) A106C, (34) R142L + E144H.

7. 10. A base editor fusion protein comprising the adenosine deaminase of claim 1 and a nucleic acid-programmable nucleotide binding domain.

8. 8. The base editor fusion protein of claim 7, wherein the nucleic acid-programmable nucleotide binding domain is a Cas protein or an AGO protein.

9. further comprising at least one nuclear localization signal sequence, and optionally a linker; Optionally, the linker comprises one or more of the sequences set forth in SEQ ID NOs: 171-180; 8. The base editor fusion protein of claim 7, optionally wherein the nuclear localization signal sequence comprises one or more of the sequences set forth in SEQ ID NOs: 181-189.

10. 9. The base editor fusion protein of claim 8, wherein the Cas protein comprises a Cas9, Cas12, or Cas13 family.

11. 11. The base editor fusion protein of claim 10, wherein the Cas protein comprises SpCas9, SaCas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, or Cas13.

12. 11. The base editor fusion protein of claim 10, wherein the Cas protein is an inactive nuclease or a nicking nuclease.

13. The following array: (i) the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; (ii) an amino acid sequence having at least 80%, 82%, 85%, 87%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3, wherein the amino acid sequence retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; (iii) an amino acid sequence having one or more amino acid residues added, substituted, deleted, or inserted in the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3, which retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; or (iv) an amino acid sequence encoded by a nucleotide sequence that hybridizes under stringent conditions with a polynucleotide sequence encoding the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3, wherein the stringent conditions are medium stringency conditions, medium to high stringency conditions, high stringency conditions, or very high stringency conditions, and the amino acid sequence retains the polynucleotide binding and base editing activity of the amino acid sequence shown in SEQ ID NO: 190 or SEQ ID NO: 3; 8. The base editor fusion protein of claim 7, comprising one or more selected from:

14. 14. The base editor fusion protein of claim 13, comprising any of the sequences set forth in SEQ ID NOs: 191-198.

15. NH 2 8. The base editor fusion protein of claim 7, having the structure: -[N-terminal fragment of nucleic acid programmable nucleotide binding domain]-[deaminase]-[C-terminal fragment of nucleic acid programmable nucleotide binding domain]-COOH.

16. Selecting an nCas9 domain as the nucleic acid programmable nucleotide binding domain; 16. The base editor fusion protein of claim 15, wherein the adenosine deaminase is inserted at amino acid positions 583-584, 768-769, 770-771, 776-777, 793-794, 905-906, 919-920, 1048-1063, 1049-1062, 1249-1250, 1263-1264, or 1276-1277 of SEQ ID NO:

61.

17. The adenosine deaminase according to claim 1, or A base editor fusion protein comprising the adenosine deaminase of claim 1 and a nucleic acid-programmable nucleotide binding domain. A polynucleotide encoding the nucleotide sequence of the present invention, preferably comprising any of SEQ ID NOs. 2, 4, 156 to 165.

18. A vector comprising the polynucleotide of claim 17.

19. The adenosine deaminase according to claim 1 .

10. A base editor fusion protein comprising the adenosine deaminase of claim 1 and a nucleic acid-programmable nucleotide binding domain. a polynucleotide encoding the adenosine deaminase or the base editor fusion protein; or A vector containing the polynucleotide including, cells.

20. A base editor system comprising the adenosine deaminase of claim 1, a nucleic acid programmable nucleotide binding domain, and a guide polynucleotide.

21. An adenosine deaminase according to any one of claims 1 to 6, a base editor fusion protein according to any one of claims 7 to 16, a polynucleotide according to claim 17, a vector according to claim 18, a cell according to claim 19 or a base editor system according to claim 20, and A pharmaceutical composition comprising a pharmaceutically acceptable vector.

22. 21. A kit comprising the adenosine deaminase of any one of claims 1 to 6, the base editor fusion protein of any one of claims 7 to 16, the polynucleotide of claim 17, the vector of claim 18, the cell of claim 19, or the base editor system of claim 20.

23. An adenosine deaminase according to any one of claims 1 to 6, a base editor fusion protein according to any one of claims 7 to 16, a polynucleotide according to claim 17, a vector according to claim 18, a cell according to claim 19 or a base editor system according to claim 20, and A delivery system comprising a delivery vehicle.

24. 21. A method for base editing of a nucleic acid, comprising contacting a nucleic acid of which bases are to be edited with the base editor system of claim 20.

25. The pharmaceutical composition of claim 21, for treating a disease associated with or caused by a point mutation.

26. 26. The pharmaceutical composition of claim 25, which is primarily for the treatment of diseases related to PCSK9 targets, and / or for the treatment of hypercholesterolemia, transthyretin amyloidosis, or β-hemoglobinopathy.